Machine Unlearning of Pre-trained Large Language Models
Jin YaoEli ChienMinxin DuXinyao NiuTianhao WangZezhou ChengXiang Yue
Presents a unified framework and benchmark evaluating seven machine unlearning methods on pre-trained large language models, demonstrating that targeted gradient-based updates can remove copyrighted or private training data over 100,000 times faster than full retraining while preserving overall model utility.
Large language models are trained on massive datasets that often include copyrighted, private, or sensitive material, exposing developers to growing legal liability and regulatory pressure regarding the "right to be forgotten." While retraining a model from scratch without the targeted data is the ideal remedy, it is computationally prohibitive for pre-trained foundation models. Consequently, previous research has largely focused on unlearning within small fine-tuned models, leaving the challenge of removing information directly from pre-trained base models largely unaddressed.
The article aims to evaluate the feasibility, effectiveness, and efficiency of machine unlearning directly on pre-trained language models. Specifically, it establishes a unified mathematical framework for unlearning and benchmarks seven distinct first-order unlearning techniques across diverse real-world domains.
The authors conducted empirical experiments using the open-source Yi-6B model, which was trained on three trillion tokens. They evaluated unlearning across thousands of long-context samples (4,096 tokens each) drawn from three domains: 500 academic papers from arXiv, 2,000 code files from GitHub, and 100 copyrighted books. To assess success without the exorbitant cost of full retraining, the researchers introduced an approximate retraining benchmark using unseen, in-distribution data to simulate the baseline performance of a model that had never encountered the target material. The evaluation examined target data forgetting, general retained data preservation, downstream task benchmarks (such as reasoning and coding capabilities), and resistance to privacy leakage measured via membership inference attacks.
The investigation produced several key findings. First, approximate unlearning methods are over 100,000 times more computationally efficient than retraining from scratch, requiring roughly 3 to 6 times 10 to the 17th floating-point operations compared to over 10 to the 23rd for retraining. Second, all evaluated methods successfully degraded the model's ability to recall targeted sequences while reducing membership inference vulnerability toward baseline levels. Third, combining gradient ascent on the forget dataset with gradient descent on in-distribution retained data demonstrated superior stability and robustness against hyperparameter variations, avoiding severe utility collapse. Finally, unlearning highly structured data poses noticeable domain-specific trade-offs; for instance, certain methods reduced coding proficiency on the HumanEval benchmark by several percentage points, with the worst-performing configurations dropping coding accuracy significantly.
These findings indicate that machine unlearning is a viable, highly cost-effective operational tool for addressing copyright disputes and data privacy mandates without sacrificing overall model utility. However, practitioners face critical trade-offs between forgetting efficacy and the unintended degradation of specialized downstream capabilities. Implementing unlearning requires careful hyperparameter control, as aggressive updates can destabilize foundational capabilities and induce catastrophic forgetting across retained tasks.
Organizations seeking to implement data deletion should avoid retraining from scratch and instead adopt first-order unlearning methods. The evidence specifically supports pairing gradient ascent on the targeted removal data with gradient descent on domain-matched retained data to maintain model stability. When tuning hyperparameters, practitioners should limit the process to approximately four optimization steps and employ a coarse-to-fine learning rate search starting within the range of 5 times 10 to the negative 6th to 5 times 10 to the negative 5th. Further development is recommended before deploying unlearning in mission-critical coding or reasoning systems to ensure domain-specific performance is safeguarded.
Confidence in these findings is moderate to high for models of similar scale, but several limitations warrant caution. The empirical evaluations were conducted on a single 6-billion parameter architecture and three specific text domains. Further research is necessary to confirm whether these methods scale consistently to larger models (such as 70-billion parameter systems or mixture-of-experts architectures), broader domains like news and general web text, and non-copyright applications such as removing harmful outputs or embedded social biases.
- Paper: Machine Unlearning, Lucas Bourtoule et al. (2019). This paper establishes the formal framework and practical motivations for machine unlearning and data deletion guarantees that the source directly adapts and benchmarks on pre-trained foundation models.
- Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). This study demonstrates how large language models memorize and leak training data, supplying the core privacy and membership-inference attack methodology used in the source to evaluate unlearning success.
- Paper: Membership Inference Attacks Against Machine Learning Models, Reza Shokri et al. (2016). It introduces the foundational framework of membership inference attacks against machine learning models, which serves as the primary metric for measuring privacy leakage in the source's unlearning benchmark.
- Paper: The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks, Nicholas Carlini et al. (2018). It formulates quantitative metrics and testing methods for unintended memorization in generative neural sequence models, providing key concepts for assessing data retention and removal.
- Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). It synthesizes the mechanics of catastrophic forgetting and stability-plasticity trade-offs in neural networks, laying essential theoretical foundations for balancing targeted forgetting with general model utility.
- Paper: Modular Pretraining Enables Access Control, Ethan Roland et al. (2026). This work explores architectural modular pretraining as an alternative mechanism for controlled capability removal and access control, directly comparing its efficacy and recovery resistance against post-hoc unlearning.
- Paper: Extracting alignment data in open models, Federico Barbero et al. (2025). It extends the study of data memorization and extraction in open models from pre-training data to post-training alignment data, presenting new extraction risks for fine-tuned LLMs.
- Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). It provides a mechanistic interpretability perspective on parameter-level modifications during fine-tuning, explaining why shallow updates often suppress rather than fundamentally delete underlying capabilities.
