Machine Unlearning
Lucas BourtouleVarun ChandrasekaranChristopher A. Choquette-ChooHengrui JiaAdelin TraversBaiwu ZhangDavid LieNicolas Papernot
Introduces the SISA framework to efficiently remove user data from trained machine learning models by strategically partitioning datasets and saving intermediate model states, drastically cutting the computational cost of complete retraining to support practical data deletion rights.
Modern privacy regulations, such as the European Union's General Data Protection Regulation and the California Consumer Privacy Act, legally mandate a "right to be forgotten," requiring organizations to completely erase personal data upon request. Standard machine learning models, especially deep neural networks, inherently memorize training data, leaving organizations vulnerable to privacy attacks if sensitive information remains embedded in model parameters. Provably removing a user's data requires "machine unlearning," ensuring the final model is statistically indistinguishable from a model never trained on that data. Traditionally, this demands retraining entire models from scratch, which creates prohibitive computational costs and operational delays for large-scale production pipelines.
The article introduces and evaluates "SISA training" (Sharded, Isolated, Sliced, and Aggregated), a practical framework designed to dramatically accelerate machine unlearning for stateful learning algorithms like deep neural networks without sacrificing strong, verifiable privacy guarantees.
The framework partitions a training dataset into multiple disjoint, isolated shards and trains separate constituent sub-models on each shard without exchanging parameter updates. Each shard is further divided into incremental slices, with model parameter checkpoints saved sequentially after training on each additional slice. When an unlearning request arrives, only the sub-model containing the target data point must be retrained, beginning directly from the last checkpoint saved before the target data slice was introduced. The outputs of all sub-models are combined during inference via aggregation strategies such as majority voting or prediction vector averaging. The authors evaluate this approach analytically and empirically across simple and complex datasets (including MNIST, Purchase, SVHN, CIFAR-100, and ImageNet) under both sequential and batched deletion scenarios, comparing it against naive retraining from scratch and fractional data baselines.
The evaluation reveals several core findings. First, SISA training provides substantial speed-ups over retraining from scratch: for simple tasks, it achieves a 4.63x retraining speed-up on the Purchase dataset and 2.45x on SVHN (configured with 20 shards and 50 slices) while incurring a negligible accuracy degradation of less than 2 percentage points. Second, the computational speed-up over the baseline is maintained as long as the number of deletion requests remains below three times the number of shards, meaning it operates effectively in realistic request regimes. Third, for complex tasks like ImageNet classification, pure sharding degrades top-5 accuracy by roughly 12 to 19.5 percentage points because sub-models lack sufficient training examples per class; however, applying transfer learning from pre-trained base models recovers this gap, reducing the accuracy loss to under 1 percentage point for top-5 accuracy while preserving retraining speed-ups. Finally, if an organization has prior knowledge of user deletion distributions across jurisdictions, distribution-aware sharding further reduces the expected volume of data requiring retraining with minimal impact on aggregate predictive performance.
These findings demonstrate that organizations can achieve strict, certifiable compliance with data erasure laws at a fraction of standard retraining costs. Rather than maintaining massive, monolithic models that require complete recomputation upon every deletion request, organizations can deploy isolated sub-models to trade modest additional storage for significant computational efficiency and operational scalability. While slicing requires additional disk storage to save intermediate parameter states, this trade-off is highly cost-effective given the comparatively low cost of storage relative to graphics processor compute time.
Organizations handling personal data in machine learning pipelines should adopt SISA training to satisfy data governance and legal erasure requirements. For complex models, engineering teams should combine SISA partitioning with transfer learning and prediction vector aggregation to mitigate potential accuracy loss. When operating across international regions with disparate privacy request rates, teams should implement distribution-aware sharding to group high-probability deletion requests into smaller, dedicated shards.
Confidence in these findings is high for standard iterative gradient descent models, though certain boundary conditions apply. Slicing benefits are restricted to stateful algorithms and do not apply to greedy non-stateful algorithms like decision trees. Furthermore, if the volume of simultaneous unlearning requests vastly exceeds the number of shards, the computational performance of SISA training naturally degrades back to that of full retraining from scratch.
- Paper: Membership Inference Attacks Against Machine Learning Models, Reza Shokri et al. (2016). This seminal paper introduces membership inference attacks, establishing the core privacy risk of model memorization that machine unlearning seeks to mitigate.
- Paper: A Closer Look at Memorization in Deep Networks, Devansh Arpit et al. (2017). This work analyzes how deep neural networks prioritize patterns versus memorizing individual training examples, providing foundational context on why data influence persists in trained weights.
- Paper: Deep Learning with Differential Privacy, Martín Abadi et al. (2016). This foundational paper develops differentially private stochastic gradient descent to bound the influence of individual training points, serving as a primary theoretical baseline for data-privacy guarantees in machine learning.
- Paper: Exploiting Unintended Feature Leakage in Collaborative Learning, Luca Melis et al. (2018). This study demonstrates how intermediate training updates leak specific training samples and unintended features, underscoring the necessity of strict data partitioning and unlearning protocols.
- Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). This paper demonstrates practical verbatim training data extraction from large language models, illustrating the high-stakes memorization vulnerabilities that downstream machine unlearning methods must confront.
