Machine Unlearning of Pre-trained Large Language Models

Jin YaoEli ChienMinxin DuXinyao NiuTianhao WangZezhou ChengXiang Yue

article2024ACL106 citations

Presents a unified framework and benchmark evaluating seven machine unlearning methods on pre-trained large language models, demonstrating that targeted gradient-based updates can remove copyrighted or private training data over 100,000 times faster than full retraining while preserving overall model utility.

Listen

Large language models are trained on massive datasets that often include copyrighted, private, or sensitive material, exposing developers to growing legal liability and regulatory pressure regarding the "right to be forgotten." While retraining a model from scratch without the targeted data is the ideal remedy, it is computationally prohibitive for pre-trained foundation models. Consequently, previous research has largely focused on unlearning within small fine-tuned models, leaving the challenge of removing information directly from pre-trained base models largely unaddressed.

The article aims to evaluate the feasibility, effectiveness, and efficiency of machine unlearning directly on pre-trained language models. Specifically, it establishes a unified mathematical framework for unlearning and benchmarks seven distinct first-order unlearning techniques across diverse real-world domains.

The authors conducted empirical experiments using the open-source Yi-6B model, which was trained on three trillion tokens. They evaluated unlearning across thousands of long-context samples (4,096 tokens each) drawn from three domains: 500 academic papers from arXiv, 2,000 code files from GitHub, and 100 copyrighted books. To assess success without the exorbitant cost of full retraining, the researchers introduced an approximate retraining benchmark using unseen, in-distribution data to simulate the baseline performance of a model that had never encountered the target material. The evaluation examined target data forgetting, general retained data preservation, downstream task benchmarks (such as reasoning and coding capabilities), and resistance to privacy leakage measured via membership inference attacks.

The investigation produced several key findings. First, approximate unlearning methods are over 100,000 times more computationally efficient than retraining from scratch, requiring roughly 3 to 6 times 10 to the 17th floating-point operations compared to over 10 to the 23rd for retraining. Second, all evaluated methods successfully degraded the model's ability to recall targeted sequences while reducing membership inference vulnerability toward baseline levels. Third, combining gradient ascent on the forget dataset with gradient descent on in-distribution retained data demonstrated superior stability and robustness against hyperparameter variations, avoiding severe utility collapse. Finally, unlearning highly structured data poses noticeable domain-specific trade-offs; for instance, certain methods reduced coding proficiency on the HumanEval benchmark by several percentage points, with the worst-performing configurations dropping coding accuracy significantly.

These findings indicate that machine unlearning is a viable, highly cost-effective operational tool for addressing copyright disputes and data privacy mandates without sacrificing overall model utility. However, practitioners face critical trade-offs between forgetting efficacy and the unintended degradation of specialized downstream capabilities. Implementing unlearning requires careful hyperparameter control, as aggressive updates can destabilize foundational capabilities and induce catastrophic forgetting across retained tasks.

Organizations seeking to implement data deletion should avoid retraining from scratch and instead adopt first-order unlearning methods. The evidence specifically supports pairing gradient ascent on the targeted removal data with gradient descent on domain-matched retained data to maintain model stability. When tuning hyperparameters, practitioners should limit the process to approximately four optimization steps and employ a coarse-to-fine learning rate search starting within the range of 5 times 10 to the negative 6th to 5 times 10 to the negative 5th. Further development is recommended before deploying unlearning in mission-critical coding or reasoning systems to ensure domain-specific performance is safeguarded.

Confidence in these findings is moderate to high for models of similar scale, but several limitations warrant caution. The empirical evaluations were conducted on a single 6-billion parameter architecture and three specific text domains. Further research is necessary to confirm whether these methods scale consistently to larger models (such as 70-billion parameter systems or mixture-of-experts architectures), broader domains like news and general web text, and non-copyright applications such as removing harmful outputs or embedded social biases.

  • Paper: Machine Unlearning, Lucas Bourtoule et al. (2019). This paper establishes the formal framework and practical motivations for machine unlearning and data deletion guarantees that the source directly adapts and benchmarks on pre-trained foundation models.
  • Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). This study demonstrates how large language models memorize and leak training data, supplying the core privacy and membership-inference attack methodology used in the source to evaluate unlearning success.
  • Paper: Membership Inference Attacks Against Machine Learning Models, Reza Shokri et al. (2016). It introduces the foundational framework of membership inference attacks against machine learning models, which serves as the primary metric for measuring privacy leakage in the source's unlearning benchmark.
  • Paper: The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks, Nicholas Carlini et al. (2018). It formulates quantitative metrics and testing methods for unintended memorization in generative neural sequence models, providing key concepts for assessing data retention and removal.
  • Paper: A Comprehensive Survey of Continual Learning: Theory, Method and Application, Liyuan Wang et al. (2023). It synthesizes the mechanics of catastrophic forgetting and stability-plasticity trade-offs in neural networks, laying essential theoretical foundations for balancing targeted forgetting with general model utility.
  • Paper: Modular Pretraining Enables Access Control, Ethan Roland et al. (2026). This work explores architectural modular pretraining as an alternative mechanism for controlled capability removal and access control, directly comparing its efficacy and recovery resistance against post-hoc unlearning.
  • Paper: Extracting alignment data in open models, Federico Barbero et al. (2025). It extends the study of data memorization and extraction in open models from pre-training data to post-training alignment data, presenting new extraction risks for fine-tuned LLMs.
  • Paper: Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, Samyak Jain et al. (2024). It provides a mechanistic interpretability perspective on parameter-level modifications during fine-tuning, explaining why shallow updates often suppress rather than fundamentally delete underlying capabilities.
Cover for Machine Unlearning of Pre-trained Large Language Models

Abstract

This study investigates the concept of the ‘right to be forgotten’ within the context of large language models (LLMs). We explore machine unlearning as a pivotal solution, with a focus on pre-trained models–a notably under-researched area. Our research delineates a comprehensive framework for machine unlearning in pre-trained LLMs, encompassing a critical analysis of seven diverse unlearning methods. Through rigorous evaluation using curated datasets from arXiv, books, and GitHub, we establish a robust benchmark for unlearning performance, demonstrating that these methods are over 10⁵ times more computationally efficient than retraining. Our results show that integrating gradient ascent with gradient descent on in-distribution data improves hyperparameter robustness. We also provide detailed guidelines for efficient hyperparameter tuning in the unlearning process. Our findings advance the discourse on ethical AI practices, offering substantive insights into the mechanics of machine unlearning for pre-trained LLMs and underscoring the potential for responsible AI development.

Table of Contents

  • 1 Introduction
  • 2 Problem Formulations
  • 3 Unlearning Methods
  • 3.1 Overview
  • 3.2 Approximate Unlearning Methods
  • 3.2.1 Gradient Ascent (or Negative Gradient)
  • 3.2.2 Fine-tuning with Random Labels
  • 3.2.3 Unlearning with Adversarial Samples
  • 3.2.4 Gradient Ascent + Descent or KL Divergence on Retained Set
  • 4 Experiments
  • 4.1 Background
  • 4.2 Evaluation Metrics
  • 4.3 Model and Datasets
  • 4.4 Results
  • 4.5 Computational efficiency analysis
  • 4.6 Ablation studies
  • 5 Related Work
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Additional Ablation Studies
  • B Related Work (Full Version)
  • Second-order methods of Machine Unlearning.

Knowls

  1. Knowl 1 — Unified Optimization Objective for Pre-trained LLM Unlearning

    equation

    Let D={xi}i=1N\mathcal{D} = \{x_i\}_{i=1}^N be a pre-training text corpus containing NN token sequences xi=(w1i,w2i,…,wtii)x_i = (w_1^i, w_2^i, \dots, w_{t_i}^i), and let PM(wt+1∣w1,…,wt)P_M(w_{t+1} \mid w_1, \dots, w_t) denote the conditional next-token prediction probability under a language model MM. To forget a subset of sequences U⊂D\mathcal{U} \subset \mathcal{D} while maintaining language generation performance on a retain subset R⊆D∖U\mathcal{R} \subseteq \mathcal{D} \setminus \mathcal{U}, first-order approximate unlearning updates the model weights using the gradient derived from the objective:

    ∑w∈U∑t=1TEqt∼Qwt[log⁡PM(qt∣w1,w2,…,wt−1)]+∑z∈R∑t=1Tlog⁡PM(zt∣z1,z2,…,zt−1)\sum_{w \in \mathcal{U}} \sum_{t=1}^T \mathbb{E}_{q_t \sim Q_{w_t}} \left[ \log P_M(q_t \mid w_1, w_2, \dots, w_{t-1}) \right] + \sum_{z \in \mathcal{R}} \sum_{t=1}^T \log P_M(z_t \mid z_1, z_2, \dots, z_{t-1})

    where w=(w1,…,wT)w = (w_1, \dots, w_T) and z=(z1,…,zT)z = (z_1, \dots, z_T) represent full token sequences from the forget set U\mathcal{U} and retain set R\mathcal{R}, respectively, and QwtQ_{w_t} is a reference probability distribution over the vocabulary W\mathcal{W} conditioned on or parameterized by the ground-truth token wtw_t.

  2. Knowl 2 — Specializations of First-Order Approximate Unlearning Methods for LLMs

    model/method

    The unified LLM unlearning framework instantiates seven distinct approximate unlearning methods based on the choice of the reference distribution QwtQ_{w_t} over the vocabulary W\mathcal{W} and the regularization term over the retain set R⊆D∖U\mathcal{R} \subseteq \mathcal{D} \setminus \mathcal{U}:

    1. Gradient Ascent (GA): Omits the retain term, sets Qwt=δwtQ_{w_t} = \delta_{w_t} (the Dirac delta placing probability 1 on the true token wtw_t), and multiplies the gradient by −1-1 to maximize the negative log-likelihood on U\mathcal{U}.
    2. Fine-tuning with Random Labels: Omits the retain term and sets Qwt=Uniform(W)Q_{w_t} = \text{Uniform}(\mathcal{W}), penalizing the model toward uniform random prediction over vocabulary W\mathcal{W} on U\mathcal{U}.
    3. Unlearning with Adversarial Samples: Omits the retain term and sets Qwt=δatQ_{w_t} = \delta_{a_t}, where at=arg⁡max⁡a∈W,a≠wtPM(a∣w1,…,wt−1)a_t = \arg\max_{a \in \mathcal{W}, a \neq w_t} P_M(a \mid w_1, \dots, w_{t-1}) is the most probable token under MM excluding the ground truth wtw_t.
    4. Gradient Ascent + Retain Descent (In-Distribution or General): Minimizes the loss on R\mathcal{R} via standard negative log-likelihood next-token prediction while simultaneously performing gradient ascent on U\mathcal{U}. R\mathcal{R} is drawn either from the same domain distribution as U\mathcal{U} (in-distribution) or sampled uniformly from general pre-training data.
    5. Gradient Ascent + Retain KL Divergence (In-Distribution or General): Combines gradient ascent on U\mathcal{U} with a Kullback-Leibler (KL) divergence constraint between the original vanilla model MM and updated model M′M' evaluated on sequences from R\mathcal{R}, where R\mathcal{R} contains either in-distribution data or general pre-training data.
  3. Knowl 3 — Approximate Retraining Evaluation Baseline

    model/method

    Evaluating approximate unlearning against the theoretical gold standard—a model M⋆=A(D∖U)M^\star = \mathcal{A}(\mathcal{D} \setminus \mathcal{U}) retrained from scratch on the retained dataset—is computationally intractable for large language models trained on trillions of tokens.

    To bypass this, the approximate retraining evaluation baseline uses an out-of-training, domain-matched dataset Dapprox\mathcal{D}_{\text{approx}}. Based on membership inference dynamics, a retrained model M⋆M^\star (which has never encountered U\mathcal{U}) will exhibit performance on U\mathcal{U} equivalent to how the vanilla model MM performs on unseen data drawn from the identical domain distribution. Consequently, the vanilla model's perplexity and next-token accuracy measured on Dapprox\mathcal{D}_{\text{approx}} provide the quantitative target performance for an unlearned model on the forget set U\mathcal{U}.

  4. Knowl 4 — Min-K% Prob Membership Inference Attack for Unlearning Verification

    experimental setup

    To evaluate whether specific sequences have been erased from a pre-trained LLM, the Min-K% Prob membership inference attack (MIA) is used. Min-K% Prob relies on the observation that non-member (unseen) sequences tend to contain outlier tokens with high negative log-likelihood, whereas training sequences (members) have more uniform likelihoods.

    Evaluation is conducted using input sequences chunked to 4096 tokens, pairing member sequences from the forget set U\mathcal{U} with an equal number of non-member sequences from an unseen in-distribution approximate set Dapprox\mathcal{D}_{\text{approx}}. The detection efficacy is quantified by the Area Under the Receiver Operating Characteristic Curve (AUC). An MIA AUC score near 1.0 indicates that target sequences remain memorized, whereas an AUC score approaching 0.5 indicates that the model's outputs on the forget set are indistinguishable from unseen data (equivalent to random guessing), reflecting effective unlearning.

  5. Knowl 5 — Computational Cost Comparison: Unlearning vs. Retraining from Scratch

    empirical result

    For a 6-billion parameter language model (Yi-6B) pre-trained on 3 trillion tokens, retraining from scratch on D∖U\mathcal{D} \setminus \mathcal{U} versus executing approximate unlearning on 2,000 sequences of 4,096 tokens (8.192×1068.192 \times 10^6 tokens) reveals a computational efficiency difference exceeding five orders of magnitude (>105×> 10^5\times). Following the standard estimation where training FLOPs equal 6×Tokens×Parameters6 \times \text{Tokens} \times \text{Parameters} and forward FLOPs equal 2×Tokens×Parameters2 \times \text{Tokens} \times \text{Parameters}:

    Method FLOPs
    Retraining 1.08×10231.08 \times 10^{23}
    Gradient Ascent (GA) 2.95×10172.95 \times 10^{17}
    Fine-tuning with Random Label 2.95×10172.95 \times 10^{17}
    Unlearning with Adversarial Samples 3.93×10173.93 \times 10^{17}
    GA + Descent on in-distribution data 5.90×10175.90 \times 10^{17}
    GA + Descent on general data 5.90×10175.90 \times 10^{17}
    GA + KL on in-distribution data 5.90×10175.90 \times 10^{17}
    GA + KL on general data 5.90×10175.90 \times 10^{17}

    The computational requirement for unlearning methods ranges between 2.95×10172.95 \times 10^{17} and 5.90×10175.90 \times 10^{17} FLOPs, representing an efficiency gain of approximately 1.8×1051.8 \times 10^5 to 3.7×1053.7 \times 10^5 times compared to retraining from scratch (1.08×10231.08 \times 10^{23} FLOPs).

  6. Knowl 6 — Benchmark Results of Pre-trained LLM Unlearning Across arXiv, GitHub, and Books Domains

    data/table

    Evaluation of seven unlearning methods applied to the open-source Yi-6B model across three distinct pre-training domains—500 arXiv papers, 2,000 GitHub code repository files, and 100 books (4,096 tokens per chunk)—demonstrates that approximate unlearning shifts forget set accuracy, perplexity, and MIA AUC toward approximate retraining targets while preserving retain set and downstream benchmark accuracy (MMLU, ARC Challenge, HumanEval, GSM8K).

    Domain / Model Forget Set Retain Set Downstream Task Acc (%) ↑\uparrow
    ACC↓\downarrow PPL↑\uparrow MIA↓\downarrow ACC↑\uparrow PPL↓\downarrow MMLU ARC HEval GSM8K Avg
    arXiv (500 papers)
    Vanilla Model 69.02 3.65 50.77 52.68 9.24 63.37 68.49 16.46 33.59 45.48
    Approximate Retrain 68.98 3.69 - - - - - - - -
    Gradient Ascent (GA) 68.79 3.70 50.28 52.66 9.26 63.45 68.77 15.85 34.04 45.53
    FT with Random Labels 68.92 3.69 50.55 52.67 9.25 63.37 68.38 14.02 32.22 44.50
    GA + Descent (In-Dist) 68.87 3.69 50.18 52.66 9.26 63.32 68.52 15.24 33.74 45.21
    GA + KL (In-Dist) 68.82 3.69 50.29 52.65 9.27 63.40 68.57 15.24 33.89 45.28
    GitHub (2K files)
    Vanilla Model 80.65 2.40 81.93 52.68 9.24 63.37 68.49 16.46 33.59 45.48
    Approximate Retrain 72.91 3.42 - - - - - - - -
    Gradient Ascent (GA) 78.19 3.53 74.28 52.60 9.31 63.45 68.40 14.63 35.10 45.40
    FT with Random Labels 78.00 3.12 80.55 52.50 9.47 62.45 67.02 10.98 29.49 42.48
    Adv. Sample Unlearning 75.09 3.40 79.51 52.54 9.41 62.36 67.33 9.76 31.39 42.71
    GA + Descent (In-Dist) 76.88 3.45 76.75 52.48 9.38 62.31 66.77 2.44 31.01 40.63
    GA + KL (In-Dist) 78.78 3.51 76.19 52.61 9.31 63.40 68.21 14.63 34.95 45.30
    Books (100 books)
    Vanilla Model 55.26 7.62 74.03 52.68 9.24 63.37 68.49 16.46 33.59 45.48
    Approximate Retrain 50.65 10.11 - - - - - - - -
    Gradient Ascent (GA) 52.47 9.64 58.47 52.45 9.40 63.32 68.66 16.46 32.90 44.91
    FT with Random Labels 51.90 10.19 63.69 52.56 9.39 63.05 68.01 16.46 29.64 44.29
    Adv. Sample Unlearning 52.07 10.02 63.60 52.59 9.35 63.08 68.18 16.46 31.39 44.78
    GA + Descent (In-Dist) 50.07 10.27 56.39 52.34 9.41 63.08 67.70 17.68 29.80 44.57
    GA + KL (In-Dist) 52.42 10.02 64.02 52.52 9.35 63.50 68.80 16.46 33.59 45.59

    Across all domains, unlearning reduces MIA AUC scores (e.g., from 81.93% down to 74.28% in GitHub code and from 74.03% down to 56.39% in Books), demonstrating measurable mitigation of training data memorization.

  7. Knowl 7 — Hyperparameter Robustness of Gradient Ascent Combined with In-Distribution Descent

    empirical result

    When sweeping learning rates η∈[5×10−6,1×10−4]\eta \in [5 \times 10^{-6}, 1 \times 10^{-4}] (at 4 optimization steps) and optimization steps S∈[1,32]S \in [1, 32] (at η=2×10−5\eta = 2 \times 10^{-5}), most unlearning methods (Gradient Ascent alone, Random Labels, Adversarial Samples, and GA with general retain data) exhibit severe hyperparameter sensitivity. For these methods, increasing η>3×10−5\eta > 3 \times 10^{-5} or steps S>8S > 8 causes an exponential explosion in perplexity on both the forget set and general retain sets (reaching values >104> 10^{4} to 103610^{36}) and severe downstream task degradation.

    In contrast, combining Gradient Ascent on the forget set with Gradient Descent on domain-matched (in-distribution) retain data shows high tolerance: forget set and general set perplexities remain stable and strictly bounded, and average downstream accuracy remains stable across variations in learning rate up to 10−410^{-4} and optimization steps up to 32.

  8. Knowl 8 — Hyperparameter Tuning Protocol for Pre-trained LLM Approximate Unlearning

    model/method

    Because first-order approximate unlearning methods in LLMs are non-convergent (continued optimization damages general utility), hyperparameter selection must follow a bounded search protocol:

    1. Optimization Steps: Set the optimization step count to four (S=4S = 4). A larger step count destabilizes the unlearning trajectory and degrades retain-set utility, while too few steps collapses gradient resolution over large batch sizes.
    2. Two-Stage Learning Rate Search:
      • Coarse Search: Perform a search at coarse granularity (10−510^{-5}) over the range 5×10−65 \times 10^{-6} to 5×10−55 \times 10^{-5}.
      • Target Matching: Compare the resulting forget set perplexity against the target perplexity established by the approximate retraining baseline (evaluated on held-out in-distribution data Dapprox\mathcal{D}_{\text{approx}}).
      • Fine Search: Refine the search within the identified interval to select the learning rate that matches the approximate retraining target.
    3. Batch Size Scaling: When scaling up the forget set size (from 512 to 8192 sequences), larger batch sizes must be used to prevent exponential growth in model perplexity on both forget and retained sets.
  9. Knowl 9 — Task-Specific Capability Degradation Under Pre-trained Code Unlearning

    empirical result

    Unlearning domain-specific data from pre-trained LLMs disproportionately degrades task performance directly aligned with the unlearned domain, even when general reasoning and language understanding metrics remain intact.

    When unlearning 2,000 GitHub code files from Yi-6B, HumanEval pass@1 accuracy drops sharply from the vanilla model baseline of 16.46%:

    • Drops to 2.44% under Gradient Ascent combined with In-Distribution Descent.
    • Drops to 9.76% under Unlearning with Adversarial Samples.
    • Drops to 10.98% under Fine-tuning with Random Labels.
    • Drops by at least 1.83% (to ≤14.63%\le 14.63\%) across all other unlearning methods (pure GA, GA + general descent, GA + KL divergence).

    Conversely, non-code downstream metrics such as MMLU (63.37% vanilla vs. 62.31%–63.45% unlearned) and GSM8K (33.59% vanilla vs. 29.49%–35.10% unlearned) remain largely stable across the unlearned models.

  10. Knowl 10 — Limitations of Pre-trained LLM Machine Unlearning Framework

    limitation

    The methodology and experimental validation of pre-trained LLM unlearning face four primary limitations:

    1. Proprietary Pre-training Data: Most commercial and open-weight LLMs do not release their exact pre-training corpora, restricting empirical validation to models with accessible pre-training data samples (such as Yi-6B). Evaluation on larger models (13B, 70B) and mixture-of-experts architectures remains untested.
    2. Non-Convergence of Approximate Objectives: The first-order approximate unlearning objectives lack formal convergence criteria; extended optimization degrades model utility on retained data, making unlearning strictly dependent on early stopping and hyperparameter tuning.
    3. MIA Detection Constraints: Membership inference attack methods on LLMs often exhibit low baseline separation between member and non-member pre-training text (e.g., initial MIA AUC of 50.77% on arXiv papers), limiting the precision with which complete sample deletion can be audited.
    4. Domain Coverage: Empirical benchmarks are limited to academic papers (arXiv), code repositories (GitHub), and books, leaving broader pre-training sources such as Common Crawl web scrapes and encyclopedic data unverified.

Coverage note — None was omitted; the extraction covers the unified unlearning framework, seven adapted algorithms, approximate retraining evaluation, computational efficiency calculations, empirical benchmark results across three domains, ablation analyses, tuning protocols, task-specific degradation findings, and stated limitations.

References

  1. 1.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv:2204.05862.
  2. 2.Lucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. 2021. Machine unlearning. In S&P, pages 141–159.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. NeurIPS, 33:1877–1901.
  4. 4.Yinzhi Cao and Junfeng Yang. 2015. Towards making systems forget with machine unlearning. In S&P, pages 463–480.
  5. 5.Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramèr. 2022. Membership inference attacks from first principles. In S&P, pages 1897–1914.
  6. 6.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang. 2023. Quantifying memorization across neural language models. In ICLR.
  7. 7.Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019. The secret sharer: Evaluating and testing unintended memorization in neural networks. In USENIX Security, pages 267–284.
  8. 8.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2021. Extracting training data from large language models. In USENIX Security, pages 2633–2650.
  9. 9.Sungmin Cha, Sungjun Cho, Dasol Hwang, Honglak Lee, Taesup Moon, and Moontae Lee. 2023. Learning to unlearn: Instance-wise unlearning for pre-trained classifiers. abs/2301.11578.
  10. 10.Jiaao Chen and Diyi Yang. 2023. Unlearn what you want to forget: Efficient unlearning for llms. arXiv:2310.20150. EMNLP.
  11. 11.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  12. 12.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  13. 13.Eli Chien, Chao Pan, and Olgica Milenkovic. 2023. Efficient model updates for approximate unlearning of graph-structured data. In ICLR.
  14. 14.Rishav Chourasia and Neil Shah. 2023. Forget unlearning: Towards true data-deletion in machine learning. In ICML, pages 6028–6073.
  15. 15.Paul F. Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. In NeurIPS, pages 4299–4307.
  16. 16.Vikram S. Chundawat, Ayush K. Tarun, Murari Mandal, and Mohan S. Kankanhalli. 2023. Can bad teaching induce forgetting? unlearning in deep networks using an incompetent teacher. In AAAI, pages 7210–7217.
  17. 17.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
  18. 18.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  19. 19.Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm.
  20. 20.Minxin Du, Xiang Yue, Sherman SM Chow, and Huan Sun. 2023a. Sanitizing sentence embeddings (and labels) for local differential privacy. In TheWebConf, pages 2349–2359.
  21. 21.Minxin Du, Xiang Yue, Sherman SM Chow, Tianhao Wang, Chenyu Huang, and Huan Sun. 2023b. Dp-forward: Fine-tuning and inference on language models with differential privacy in forward pass. In CCS, pages 2665–2679.
  22. 22.Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. 2024. Do membership inference attacks work on large language models?
  23. 23.Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam D. Smith. 2006. Calibrating noise to sensitivity in private data analysis. In TCC, pages 265–284.
  24. 24.Cynthia Dwork and Aaron Roth. 2014. The algorithmic foundations of differential privacy. Found. Trends Theor. Comput. Sci., 9(3-4):211–407.
  25. 25.Ronen Eldan and Mark Russinovich. 2023. Who’s harry potter? approximate unlearning in llms. arXiv:2310.02238.
  26. 26.Sanjam Garg, Shafi Goldwasser, and Prashant Nalini Vasudevan. 2020. Formalizing data deletion in the context of the right to be forgotten. In EUROCRYPT, pages 373–402.
  27. 27.Antonio Ginart, Melody Y. Guan, Gregory Valiant, and James Zou. 2019. Making AI forget you: Data deletion in machine learning. In NeurIPS, pages 3513–3526.
  28. 28.Aditya Golatkar, Alessandro Achille, and Stefano Soatto. 2020. Eternal sunshine of the spotless net: Selective forgetting in deep networks. In CVPR, pages 9301–9309.
  29. 29.Chen Gong, Kecen Li, Jin Yao, and Tianhao Wang. 2024. Trajdeleter: Enabling trajectory forgetting in offline reinforcement learning agents. arXiv preprint arXiv:2404.12530.
  30. 30.Chuan Guo, Tom Goldstein, Awni Y. Hannun, and Laurens van der Maaten. 2020. Certified data removal from machine learning models. In ICML, pages 3832–3842.
  31. 31.Varun Gupta, Christopher Jung, Seth Neel, Aaron Roth, Saeed Sharifi-Malvajerdi, and Chris Waites. 2021. Adaptive machine unlearning. In NeurIPS, pages 16319–16330.
  32. 32.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring massive multitask language understanding. In ICLR.
  33. 33.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In ICML, pages 2790–2799.
  34. 34.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. Lora: Low-rank adaptation of large language models. In ICLR.
  35. 35.Yiyang Huang and Clément L. Canonne. 2023. Tight bounds for machine unlearning via differential privacy. arXiv:2309.00886.
  36. 36.Zachary Izzo, Mary Anne Smart, Kamalika Chaudhuri, and James Zou. 2021. Approximate data deletion from machine learning models. In AISTATS, pages 2008–2016.
  37. 37.Matthew Jagielski, Om Thakkar, Florian Tramèr, Daphne Ippolito, Katherine Lee, Nicholas Carlini, Eric Wallace, Shuang Song, Abhradeep Guha Thakurta, Nicolas Papernot, and Chiyuan Zhang. 2023. Measuring forgetting of memorized training examples. In ICLR.
  38. 38.Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. 2023. Knowledge unlearning for mitigating privacy risks in language models. In ACL, pages 14389–14408.
  39. 39.Jinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao, Gaowen Liu, Yang Liu, Pranay Sharma, and Sijia Liu. 2023. Model sparsification can simplify machine unlearning. In NeurIPS (Spotlight). ArXiv:2304.04934.
  40. 40.Ronald Kemker, Marc McClure, Angelina Abitino, Tyler L. Hayes, and Christopher Kanan. 2018. Measuring catastrophic forgetting in neural networks. In AAAI, pages 3390–3398.
  41. 41.Hyunjik Kim, George Papamakarios, and Andriy Mnih. 2021. The lipschitz constant of self-attention. In ICML, pages 5562–5571.
  42. 42.Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, Shahul ES, Sameer Suri, David Glushkov, Arnav Dantuluri, Andrew Maguire, Christoph Schuhmann, Huu Nguyen, and Alexander Mattick. 2023. Openassistant conversations - democratizing large language model alignment.
  43. 43.Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah, and Dan Roth. 2022. Privacy adhering machine un-learning in NLP. arXiv:2212.09573.
  44. 44.Meghdad Kurmanji, Peter Triantafillou, and Eleni Triantafillou. 2023. Towards unbounded machine unlearning. arXiv:2302.09880.
  45. 45.Kecen Li, Chen Gong, Zhixiang Li, Yuzhong Zhao, Xinwen Hou, and Tianhao Wang. 2024. Meticulously selecting 1% of the dataset for pre-training! generating differentially private images data with semantics query. USENIX Security.
  46. 46.Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2021. Large language models can be strong differentially private learners. arXiv preprint arXiv:2110.05679.
  47. 47.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  48. 48.Zheyuan Liu, Guangyao Dou, Yijun Tian, Chunhui Zhang, Eli Chien, and Ziwei Zhu. 2024. Breaking the trilemma of privacy, utility, efficiency via controllable machine unlearning.
  49. 49.Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. 2022. Estimating the carbon footprint of bloom, a 176b parameter language model. arXiv:2211.02001.
  50. 50.Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter. 2024. Tofu: A task of fictitious unlearning for llms. arXiv preprint arXiv:2401.06121.
  51. 51.Yunlong Mao, Zexi Xin, Zhenyu Li, Jue Hong, Qingyou Yang, and Sheng Zhong. 2023. Secure split learning against property inference, data reconstruction, and feature space hijacking attacks. arXiv preprint arXiv:2304.09515.
  52. 52.Ilya Mironov. 2017. Rényi differential privacy. In CSF, pages 263–275.
  53. 53.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  54. 54.Alexandra Peste, Dan Alistarh, and Christoph H. Lampert. 2021. SSSE: efficiently erasing samples from trained machine learning models. arXiv:2107.03860.
  55. 55.RealTimeData. 2024. github_latest.
  56. 56.Ayush Sekhari, Jayadev Acharya, Gautam Kamath, and Ananda Theertha Suresh. 2021. Remember what you want to forget: Algorithms for machine unlearning. In NeurIPS, pages 18075–18086.
  57. 57.Chenze Shao and Yang Feng. 2022. Overcoming catastrophic forgetting beyond continual learning: Balanced training for neural machine translation. In ACL, pages 2023–2036.
  58. 58.Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. 2023. Detecting pretraining data from large language models. arXiv preprint arXiv:2310.16789.
  59. 59.Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Membership inference attacks against machine learning models. In S&P, pages 3–18.
  60. 60.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020. Learning to summarize with human feedback. In NeurIPS.
  61. 61.Ayush K. Tarun, Vikram S. Chundawat, Murari Mandal, and Mohan S. Kankanhalli. 2021. Fast yet effective machine unlearning. CoRR, abs/2111.08947.
  62. 62.Kushal Tirumala, Aram H. Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022. Memorization without overfitting: Analyzing the training dynamics of large language models. In NeurIPS.
  63. 63.TogetherComputer. 2023. Redpajama: An open source recipe to reproduce llama training dataset.
  64. 64.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023a. Llama: Open and efficient foundation language models. arXiv:2302.13971.
  65. 65.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton-Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurélien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv:2307.09288.
  66. 66.Lingzhi Wang, Xingshan Zeng, Jinsong Guo, Kam-Fai Wong, and Georg Gottlob. 2024. Selective forgetting: Advancing machine unlearning techniques and evaluation in language models. arXiv preprint arXiv:2402.05813.
  67. 67.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022. Emergent abilities of large language models. TMLR, 2022.
  68. 68.Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023. Bloomberggpt: A large language model for finance. arXiv preprint arXiv:2303.17564.
  69. 69.Heng Xu, Tianqing Zhu, Lefeng Zhang, Wanlei Zhou, and Philip S. Yu. 2024. Machine unlearning: A survey. ACM Comput. Surv., 56(1):9:1–9:36.
  70. 70.Borui Yang, Wei Li, Liyao Xiang, and Bo Li. 2023. Towards code watermarking with dual-channel transformations. arXiv preprint arXiv:2309.00860.
  71. 71.Yuanshun Yao, Xiaojun Xu, and Yang Liu. 2023. Large language model unlearning. arXiv preprint arXiv:2310.10683.
  72. 72.Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, et al. 2024. Yi: Open foundation models by 01. ai. arXiv preprint arXiv:2403.04652.
  73. 73.Xiang Yue, Minxin Du, Tianhao Wang, Yaliang Li, Huan Sun, and Sherman SM Chow. 2021. Differential privacy for text analytics via natural text sanitization. arXiv preprint arXiv:2106.01221.
  74. 74.Xiang Yue, Huseyin A Inan, Xuechen Li, Girish Kumar, Julia McAnallen, Hoda Shajari, Huan Sun, David Levitan, and Robert Sim. 2022. Synthetic text generation with differential privacy: A simple and practical recipe. arXiv preprint arXiv:2210.14348.
  75. 75.Dawen Zhang, Pamela Finckenberg-Broman, Thong Hoang, Shidong Pan, Zhenchang Xing, Mark Staples, and Xiwei Xu. 2023a. Right to be forgotten in the era of large language models: Implications, challenges, and solutions. arXiv:2307.03941.
  76. 76.Xulong Zhang, Jianzong Wang, Ning Cheng, Yifu Sun, Chuanyao Zhang, and Jing Xiao. 2023b. Machine unlearning methodology base on stochastic teacher network. arXiv:2308.14322.

Citation

MLA
Yao, J., et al. “Machine Unlearning of Pre-trained Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 8403–19, https://doi.org/10.18653/v1/2024.acl-long.457.
APA
Yao, J., Chien, E., Du, M., Niu, X., Wang, T., Cheng, Z., & Yue, X. (2024). Machine Unlearning of Pre-trained Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8403–8419. https://doi.org/10.18653/v1/2024.acl-long.457
Chicago
Yao, J., E. Chien, M. Du, et al. 2024. “Machine Unlearning of Pre-trained Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8403–19. https://doi.org/10.18653/v1/2024.acl-long.457.
Harvard
Yao, J. et al. (2024) “Machine Unlearning of Pre-trained Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8403–8419. Available at: https://doi.org/10.18653/v1/2024.acl-long.457.
Vancouver
1. Yao J, Chien E, Du M, Niu X, Wang T, Cheng Z, Yue X (2024) Machine Unlearning of Pre-trained Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8403–8419

BibTeX

@inproceedings{yao-etal-2024-machine,
    title = "Machine Unlearning of Pre-trained Large Language Models",
    author = "Yao, Jin  and
      Chien, Eli  and
      Du, Minxin  and
      Niu, Xinyao  and
      Wang, Tianhao  and
      Cheng, Zezhou  and
      Yue, Xiang",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.457/",
    doi = "10.18653/v1/2024.acl-long.457",
    pages = "8403--8419"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/