MUSE: Machine Unlearning Six-Way Evaluation for Language Models

Weijia ShiJaechan LeeYangsibo HuangSadhika MalladiJieyu ZhaoAri HoltzmanDaogao LiuLuke ZettlemoyerNoah SmithChiyuan Zhang

article2025ICLR303 citations

Introduces MUSE, a six-property evaluation benchmark for machine unlearning in language models, and reveals that popular unlearning methods cause severe privacy leakage and degrade model utility during sequential or large-scale data removal.

Listen

Modern language models are trained on massive text collections that frequently contain copyrighted works and private personal information. Due to rising privacy regulations like the General Data Protection Regulation and ongoing copyright lawsuits, data owners increasingly demand the removal of their data from deployed systems. However, perfectly erasing specific data by retraining models from scratch is computationally intractable, prompting the development of approximate machine unlearning methods that attempt to remove targeted information without full retraining.

The article aims to evaluate the practical effectiveness and viability of existing approximate machine unlearning techniques across six critical dimensions that capture the needs of both data owners and model deployers. Specifically, it introduces a comprehensive evaluation framework called MUSE (Machine Unlearning Six-Way Evaluation) to assess whether these algorithms can reliably remove data without damaging general model performance or leaking private information.

To conduct this evaluation, the researchers tested eight leading unlearning algorithms applied to 7-billion-parameter language models across two realistic datasets: a news dataset comprising up to 3.3 million tokens of recent BBC articles and a books dataset comprising 1.1 million tokens from the Harry Potter series. The evaluation framework measures six specific properties: prevention of verbatim memorization, elimination of factual knowledge retention, prevention of privacy leakage via membership detection, preservation of general model utility on retained data, scalability to larger removal requests, and sustainability over sequential removal requests.

The investigation produced four central findings. First, while most unlearning algorithms successfully suppress verbatim text generation and factual recall—in some cases reducing memorization scores to zero—they fail to prevent privacy leakage. Algorithms consistently either under-unlearn or over-unlearn, creating statistical anomalies that allow membership inference attacks to detect whether the target data was originally in the training set. Second, existing methods cause severe model degradation; every tested algorithm reduced model performance on retained data by 24% to 100%, with several methods completely destroying general model utility. Third, unlearning performance degrades rapidly as the volume of data to be forgotten scales up from 0.8 million to 3.3 million tokens. Fourth, current algorithms lack sustainability, showing steep drops in utility when handling successive, sequential unlearning requests.

These findings demonstrate that current approximate unlearning techniques create a false sense of security and are not ready for practical or regulatory deployment. For organizations facing legal or compliance pressures, relying on these methods introduces significant risks: they fail to ensure true data privacy, carry high risks of operational degradation, and cannot sustain real-world workloads where data removal requests arrive over time.

Decision-makers should exercise caution and avoid deploying current approximate unlearning methods in production environments where compliance or privacy guarantees are strictly required. Instead, organizations must treat data removal as an open technical challenge, using comprehensive multi-metric evaluations like MUSE to validate emerging algorithms. Future research and development should focus on designing balanced unlearning objectives that preserve general model capabilities while preventing statistical signatures that expose member privacy.

The primary limitations of this evaluation include its focus on 7-billion-parameter text models across news and literature domains, leaving multimodal architectures, smaller or larger language models, and specialized domains such as medical records or emails for future study. Nonetheless, the evidence strongly supports high confidence in the conclusion that current unlearning methods fail to meet real-world operational and privacy requirements.

  • Paper: Machine Unlearning, Lucas Bourtoule et al. (2019). Read this foundational SISA work first to understand approximate unlearning, retraining baselines, and the deletion setting that MUSE evaluates at language-model scale.
  • Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). Its demonstration of verbatim training-data extraction establishes the memorization and privacy risks that MUSE measures as unlearning targets.
  • Paper: Machine Unlearning of Pre-trained Large Language Models, Jin Yao et al. (2024). This direct LLM unlearning benchmark introduces methods and evaluation concerns for pre-trained models that MUSE extends with broader criteria, including sequential sustainability and scale.
  • Paper: The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning, Nathaniel Li et al. (2024). Its RMU experiments provide a concrete LLM unlearning method and utility-preservation evaluation that help contextualize MUSE’s comparison of unlearning algorithms.
  • Paper: LLM Dataset Inference: Did you train on my dataset?, Pratyush Maini et al. (2024). Its analysis of dataset-level membership inference clarifies how privacy leakage can persist beyond verbatim recall, a distinction central to MUSE’s evaluation.
Cover for MUSE: Machine Unlearning Six-Way Evaluation for Language Models

Abstract

Language models (LMs) are trained on vast amounts of text data, which may include private and copyrighted content. Data owners may request the removal of their data from a trained model due to privacy or copyright concerns. However, exactly unlearning only these datapoints (i.e., retraining with the data removed) is intractable in modern-day models. This has led to the development of many approximate unlearning algorithms. The evaluation of the efficacy of these algorithms has traditionally been narrow in scope, failing to precisely quantify the success and practicality of the algorithm from the perspectives of both the model deployers and the data owners. We address this issue by proposing MUSE, a comprehensive machine unlearning evaluation benchmark that enumerates six diverse desirable properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. Using these criteria, we benchmark how effectively eight popular unlearning algorithms on 7B-parameter LMs can unlearn Harry Potter books and news articles. Our results demonstrate that most algorithms can prevent verbatim memorization and knowledge memorization to varying degrees, but only one algorithm does not lead to severe privacy leakage. Furthermore, existing algorithms fail to meet deployer's expectations because they often degrade general model utility and also cannot sustainably accommodate successive unlearning requests or large-scale content removal. Our findings identify key issues with the practicality of existing unlearning algorithms on language models, and we release our benchmark to facilitate further evaluations: this http URL

Table of Contents

  • 1 Introduction
  • 2 Machine Unlearning: Preliminaries and Notations
  • 3 The MUSE Evaluation Benchmark
  • 3.1 Evaluation Metrics
  • 3.2 Evaluation Corpus
  • 4 Unlearning Methods
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Results: Data Owner Expectations
  • 5.3 Results: Deployment Considerations
  • 6 Related Work
  • 7 Conclusion
  • 8 Acknowledgements
  • References
  • A Broader Impact
  • B Experimental Details
  • B.1 Compute Configurations
  • B.2 Experimental Setup
  • B.3 Efficiency of Unlearning Methods
  • C More Experimental Results
  • C.1 Confidence Intervals for C1, C2 and C4 in
  • D Dataset Details

Knowls

  1. Knowl 1 — MUSE defines six requirements for practical language-model unlearning

    definition

    MUSE evaluates approximate unlearning from the perspectives of data owners and model deployers. Data-owner criteria are: no verbatim memorization, meaning the model should not reproduce text from the removal-request data; no knowledge memorization, meaning it should not retain the ability to answer questions about that data; and no privacy leakage, meaning an attacker should not be able to infer that the data was used for training. Deployer criteria are: utility preservation on data not requested for removal; scalability as the size of removal requests grows; and sustainability when removal requests arrive sequentially. The benchmark treats these as distinct goals: suppressing targeted content alone does not establish privacy, preserve general capability, or show that unlearning works at larger scales or over repeated requests.

  2. Knowl 2 — MUSE operationalizes memorization and privacy with three data-owner metrics

    equation

    Let DforgetD_{\mathrm{forget}} be the examples removed from training, and let ff be a language model after unlearning. For a text sequence xx, let x[:l]x[:l] be its first ll tokens and x[l+1:]x[l+1:] its remaining tokens; the model is prompted with the prefix and its generated continuation is compared with the true continuation using ROUGE-L F1. For question-answer pairs (q,a)(q,a) derived from the removal data, compare the model’s answer f(q)f(q) with the reference answer aa using ROUGE. Lower scores indicate less measured memorization.

    VerbMem(f,Dforget)=1∣Dforget∣∑x∈DforgetROUGE\mbox−L(f(x[:l]),x[l+1:]),\mathrm{VerbMem}(f,D_{\mathrm{forget}})=\frac{1}{|D_{\mathrm{forget}}|}\sum_{x\in D_{\mathrm{forget}}}\mathrm{ROUGE\mbox{-}L}\bigl(f(x[:l]),x[l+1:]\bigr), KnowMem(f,Dforget)=1∣Dforget∣∑(q,a)∈DforgetROUGE(f(q),a).\mathrm{KnowMem}(f,D_{\mathrm{forget}})=\frac{1}{|D_{\mathrm{forget}}|}\sum_{(q,a)\in D_{\mathrm{forget}}}\mathrm{ROUGE}\bigl(f(q),a\bigr).

    For privacy, a membership-inference attack (MIA) uses Min-K% Prob to distinguish removal-set examples, treated as members, from a disjoint in-distribution holdout set DholdoutD_{\mathrm{holdout}}, treated as non-members. Let fretrainf_{\mathrm{retrain}} be a reference model trained without the removal set, and let AUC(f;Dforget,Dholdout)\mathrm{AUC}(f;D_{\mathrm{forget}},D_{\mathrm{holdout}}) be the attack’s ROC AUC for model ff. MUSE defines relative privacy leakage as

    PrivLeak=AUC(funlearn;Dforget,Dholdout)−AUC(fretrain;Dforget,Dholdout)AUC(fretrain;Dforget,Dholdout).\mathrm{PrivLeak}=\frac{\mathrm{AUC}(f_{\mathrm{unlearn}};D_{\mathrm{forget}},D_{\mathrm{holdout}})-\mathrm{AUC}(f_{\mathrm{retrain}};D_{\mathrm{forget}},D_{\mathrm{holdout}})}{\mathrm{AUC}(f_{\mathrm{retrain}};D_{\mathrm{forget}},D_{\mathrm{holdout}})}.

    A value near zero means the unlearned model’s attack score is near the retrained reference. Large positive or negative values indicate that the removal-set score has diverged from that reference through over- or under-unlearning, respectively. The retrained AUC is generally around 0.50.5, although differences between removal and holdout data can shift the baseline.

  3. Knowl 3 — Utility, scalability, and sustainability are measured on retained knowledge

    definition

    MUSE measures utility preservation as KnowMem(funlearn,Dretain)\mathrm{KnowMem}(f_{\mathrm{unlearn}},D_{\mathrm{retain}}): the unlearned model’s ROUGE performance on question-answer pairs from data that should remain available. Higher values indicate better preservation of the measured capability.

    For scalability, let DucD_u^c be a removal set of size cc and fucf_u^c the model produced after unlearning it. Scalability is assessed by how utility and other relevant evaluation metrics change as cc increases. For sustainability, let fu,kf_{u,k} be the model after processing the kkth sequential removal request; sustainability is assessed by how those metrics change as the number of requests kk grows. These are trend-based criteria: they test whether unlearning remains effective as removal grows or is repeated, rather than only whether a single request can be handled.

  4. Knowl 4 — MUSE supplies real-text news and book corpora with text and knowledge evaluations

    data/table

    MUSE builds two corpora around material that can be the subject of removal requests. NEWS uses BBC articles collected after August 2023, with disjoint news-article removal, retain, and holdout sets. BOOKS uses the Harry Potter series as the removal data and Harry Potter FanWiki material as retained, related domain knowledge; the book holdout is a disjoint portion of the books. Each corpus includes original verbatim text for continuation-based evaluation and question-answer pairs for knowledge evaluation. The knowledge pairs are generated by partitioning text into 2,048-token excerpts, prompting GPT-4 to produce questions answerable only from excerpt-specific details and short answers copied verbatim from the excerpt, and excluding pairs whose answers fail the verbatim check.

    The standard NEWS evaluation uses a 0.8M-token removal set; the extended NEWS removal set used for scale and sequential-request tests contains 3.3M tokens. NEWS also has 1.6M-token retain and 2.0M-token holdout sets. BOOKS has 1.1M tokens of Harry Potter books in the removal set, 0.5M tokens of FanWiki in the evaluation retain set, and 0.6M tokens of books in the holdout set. For regularized unlearning, the benchmark provides a separate 1.6M-token NEWS retain subset and a separate 0.2M-token FanWiki retain subset; these are disjoint from the evaluation retain sets so the regularizers do not directly train on the utility evaluation data.

  5. Knowl 5 — The evaluated unlearning suite combines four methods with two utility regularizers

    model/method

    MUSE compares eight approximate unlearning methods. Gradient Ascent (GA) performs gradient ascent on the cross-entropy loss of the removal data, reducing the likelihood of correct predictions there. Negative Preference Optimization (NPO) treats removal examples as negative preferences and discourages their likelihood while constraining divergence from the starting target model. With target-model parameters θtarget\theta_{\mathrm{target}}, unlearning parameters θ\theta, sequence probability fθ(x)f_\theta(x), and sigmoid σ\sigma, its objective is

    LNPO(θ)=−2β Ex∼Dforget[log⁡σ(−βlog⁡fθ(x)fθtarget(x))],L_{\mathrm{NPO}}(\theta)=-\frac{2}{\beta}\,\mathbb{E}_{x\sim D_{\mathrm{forget}}}\left[\log\sigma\left(-\beta\log\frac{f_\theta(x)}{f_{\theta_{\mathrm{target}}}(x)}\right)\right],

    where β\beta controls the allowed divergence and is set to 0.10.1 in the experiments. Task Vector first fine-tunes the target model on removal data until it overfits, producing reinforced parameters θreinforce\theta_{\mathrm{reinforce}}, then subtracts the resulting weight difference: θunlearn=θtarget−(θreinforce−θtarget)\theta_{\mathrm{unlearn}}=\theta_{\mathrm{target}}-(\theta_{\mathrm{reinforce}}-\theta_{\mathrm{target}}). Who’s Harry Potter (WHP) combines target and reinforced next-token distributions at inference time: for prompt xx, it uses punlearn(⋅∣x)=ptarget(⋅∣x)−α(preinforce(⋅∣x)−ptarget(⋅∣x))p_{\mathrm{unlearn}}(\cdot\mid x)=p_{\mathrm{target}}(\cdot\mid x)-\alpha\bigl(p_{\mathrm{reinforce}}(\cdot\mid x)-p_{\mathrm{target}}(\cdot\mid x)\bigr), with α\alpha controlling the degree of interpolation.

    GA and NPO are each evaluated alone and with either of two utility regularizers, yielding GA, GAGDR, GAKLR, NPO, NPOGDR, and NPOKLR. Gradient Descent on the Retain set (GDR) adds a conventional cross-entropy learning objective on a separate retain-training subset. KL Divergence Minimization on the Retain set (KLR) encourages the unlearned model’s output distribution to match the target model’s distribution on that subset. Task Vector and WHP are evaluated without these regularizers: the former deliberately overfits removal data to construct its weight direction, and the latter performs unlearning at inference time rather than through optimization.

  6. Knowl 6 — Experiments compare unlearned models with target and retrained 7B references

    experimental setup

    For each corpus, the target model is fine-tuned on the union of removal and retain training data, while the retrained reference is fine-tuned on retain data alone. The NEWS starting model is LLaMA-2 7B, released before the BBC articles used; the BOOKS starting model is ICLM-7B, which the authors state did not contain the Harry Potter books in pretraining. Both starting models were fine-tuned for five epochs at learning rate 10−510^{-5} and batch size 32. GA, NPO, and their regularized variants use AdamW at learning rate 10−510^{-5} and batch size 32. Unlearning stops at the first epoch, within a ten-epoch limit, when utility falls below the retrained reference’s utility; if that condition is not met, the tenth-epoch checkpoint is used. Task Vector and WHP use a reinforced model obtained by fine-tuning the target model for ten epochs. The selected settings were: GA, epoch 1 for both corpora; GAGDR, epochs 7 and 1 for NEWS and BOOKS; GAKLR, epochs 10 and 5; NPO, epoch 1 for both; NPOGDR, epochs 10 and 1; NPOKLR, epochs 10 and 4; Task Vector, α=29\alpha=29 for both; and WHP, α=22\alpha=22 for NEWS and α=28\alpha=28 for BOOKS. Experiments used eight NVIDIA A40 GPUs on one node.

  7. Knowl 7 — Most tested methods reduce removal-set memorization but lose retained utility

    data/table

    The table reports mean ROUGE-L F1 scores on a 0–100 scale for verbatim continuation and knowledge-question answering on the removal set, relative privacy leakage in percentage units, and knowledge-question utility on the retain set. Lower removal-set ROUGE scores are desirable, privacy leakage near zero matches the retrained reference, and higher retain-set ROUGE indicates better utility. The comparison shows that GA and NPO without regularization can drive both removal-set scores to zero while also driving retain utility to zero. Other methods often reduce one or both removal-set scores but retain less utility than the retrained reference. Task Vector largely preserves NEWS utility but barely reduces NEWS memorization; on BOOKS it barely changes verbatim memorization. PrivLeak values outside the paper’s stated [−5%,5%][-5\%,5\%] target indicate privacy-score divergence.

    Corpus Model VerbMem KnowMem forget PrivLeak (%) KnowMem retain
    NEWS Target 58.4 63.9 -99.8 55.2
    NEWS Retrained 20.8 33.1 0.0 55.0
    NEWS GA 0.0 0.0 5.2 0.0
    NEWS GAGDR 4.9 31.0 108.1 27.3
    NEWS GAKLR 27.4 50.2 -96.1 44.8
    NEWS NPO 0.0 0.0 24.4 0.0
    NEWS NPOGDR 1.2 54.6 105.8 40.5
    NEWS NPOKLR 26.9 49.0 -95.8 45.4
    NEWS Task Vector 57.2 66.2 -99.8 55.8
    NEWS WHP 19.7 21.2 109.6 28.3
    BOOKS Target 99.8 59.4 -57.5 66.9
    BOOKS Retrained 14.3 28.9 0.0 74.5
    BOOKS GA 0.0 0.0 -25.0 0.0
    BOOKS GAGDR 0.0 0.0 -26.5 10.7
    BOOKS GAKLR 16.0 21.9 -40.2 37.2
    BOOKS NPO 0.0 0.0 -24.3 0.0
    BOOKS NPOGDR 0.0 0.0 -30.8 22.8
    BOOKS NPOKLR 17.0 25.0 -43.5 44.6
    BOOKS Task Vector 99.7 52.4 -57.5 64.7
    BOOKS WHP 18.0 55.7 56.5 63.6
  8. Knowl 8 — Membership inference distinguishes most approximate unlearning outcomes from retraining

    empirical result

    On NEWS, the Min-K% Prob membership-inference ROC AUC was 0.48 for the retrained reference and 0.00 for the target model. The approximate-method AUCs were 0.50 for GA, 0.99 for GAGDR, 0.02 for GAKLR, 0.59 for NPO, 0.98 for NPOGDR, 0.02 for NPOKLR, 1.00 for WHP, and 0.00 for Task Vector. The paper treats the retrained model’s near-diagonal ROC behavior as the desired no-leakage reference; divergence from that behavior indicates that the attack can distinguish removal examples from holdout examples. Approximate methods diverged in different directions: the NEWS GDR-regularized methods showed strong over-unlearning behavior, while KLR-regularized methods showed under-unlearning and gave little improvement over the target on the reported privacy measure. Across the two corpora, the relative leakage scores in the table also show substantial positive or negative divergence for nearly all approximate methods, rather than consistent matching to retraining.

  9. Knowl 9 — Larger and sequential removal requests reduce utility in the tested methods

    empirical result

    The scalability and sustainability experiments track retain-set utility for GA, NPO, and their GDR- and KLR-regularized variants. For scalability, the NEWS removal set is expanded from 0.8M to 3.3M tokens, with measurements at 0.0M, 0.8M, 1.7M, 2.5M, and 3.3M tokens. Utility generally declines as the removal set grows, reaching its lowest level at the largest size. For sustainability, the 3.3M-token NEWS removal set is divided into four disjoint folds of about 0.8M tokens; the methods are applied sequentially to those folds, and utility is tracked from before any request through the fourth request. Utility tends to fall as requests accumulate. The reported trends indicate that these methods do not reliably preserve utility under larger or repeated removal workloads.

  10. Knowl 10 — The benchmark leaves several unlearning requirements and deployment contexts unevaluated

    limitation

    MUSE does not evaluate whether removed information can still be recovered from intermediate model activations, and it does not provide formal guarantees that unlearning succeeded. It also does not test whether unlearning preserves capabilities such as fine-tuning or in-context learning, or systematically assess effects across user groups and fairness outcomes. Its experiments cover books and news with 7B-parameter models, so the findings do not establish performance on other corpora, model sizes, or modalities. The authors identify computational and storage costs, including whether an algorithm must retain a copy of the retain set, as additional deployment considerations not fully captured by the six evaluation criteria.

Coverage note — The appendix’s bootstrap confidence intervals and per-step wall-clock timing measurements are omitted as supporting robustness and implementation details rather than distinct findings of the six-facet benchmark.

References

  1. 1.Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. Leace: Perfect linear concept erasure in closed form. Advances in Neural Information Processing Systems, 36, 2024.
  2. 2.Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. Nuanced metrics for measuring unintended bias with real data for text classification, 2019.
  3. 3.Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. Machine unlearning. In 2021 IEEE Symposium on Security and Privacy (SP), pp. 141–159. IEEE, 2021.
  4. 4.Yinzhi Cao and Junfeng Yang. Towards making systems forget with machine unlearning. In 2015 IEEE symposium on security and privacy, pp. 463–480. IEEE, 2015.
  5. 5.Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21), pp. 2633–2650, 2021.
  6. 6.Gert Cauwenberghs and Tomaso Poggio. Incremental and decremental support vector machine learning. Advances in neural information processing systems, 13, 2000.
  7. 7.Tianshi Che, Yang Zhou, Zijie Zhang, Lingjuan Lyu, Ji Liu, Da Yan, Dejing Dou, and Jun Huan. Fast federated machine unlearning with nonlinear functional theory. In International conference on machine learning, pp. 4241–4268. PMLR, 2023.
  8. 8.Jiali Cheng and Hadi Amiri. Multimodal machine unlearning, 2023.
  9. 9.Rishav Chourasia and Neil Shah. Forget unlearning: Towards true data-deletion in machine learning. In International Conference on Machine Learning, pp. 6028–6073. PMLR, 2023.
  10. 10.Cynthia Dwork, Krishnaram Kenthapadi, Frank McSherry, Ilya Mironov, and Moni Naor. Our data, ourselves: Privacy via distributed noise generation. In Advances in Cryptology-EUROCRYPT 2006: 24th Annual International Conference on the Theory and Applications of Cryptographic Techniques, St. Petersburg, Russia, May 28-June 1, 2006. Proceedings 25, pp. 486–503. Springer, 2006a.
  11. 11.Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. Calibrating noise to sensitivity in private data analysis. In Theory of Cryptography: Third Theory of Cryptography Conference, TCC 2006, New York, NY, USA, March 4-7, 2006. Proceedings 3, pp. 265–284. Springer, 2006b.
  12. 12.Ronen Eldan and Mark Russinovich. Who’s Harry Potter? Approximate Unlearning in LLMs. arXiv preprint arXiv:2310.02238, 2023.
  13. 13.DOE 1 v. GitHub, Inc. 4:22-cv-06823, N.D. Cal. 2022.
  14. 14.Tremblay v. OpenAI, Inc.,. 23-cv-03416-AMO, (N.D. Cal.), 2023.
  15. 15.European Parliament and Council of the European Union. Regulation (EU) 2016/679 of the European Parliament and of the Council. URL https://data.europa.eu/eli/reg/2016/679/oj.
  16. 16.Chongyu Fan, Jiancheng Liu, Yihua Zhang, Dennis Wei, Eric Wong, and Sijia Liu. Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generation. arXiv preprint arXiv:2310.12508, 2023.
  17. 17.Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau. Erasing concepts from diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 2426–2436, October 2023.
  18. 18.Badih Ghazi, Pritish Kamath, Ravi Kumar, Pasin Manurangsi, Ayush Sekhari, and Chiyuan Zhang. Ticketed learning–unlearning schemes. In The Thirty Sixth Annual Conference on Learning Theory, pp. 5110–5139. PMLR, 2023.
  19. 19.Antonio Ginart, Melody Guan, Gregory Valiant, and James Y Zou. Making ai forget you: Data deletion in machine learning. Advances in neural information processing systems, 32, 2019.
  20. 20.Aditya Golatkar, Alessandro Achille, and Stefano Soatto. Eternal sunshine of the spotless net: Selective forgetting in deep networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9304–9312, 2020a.
  21. 21.Aditya Golatkar, Alessandro Achille, and Stefano Soatto. Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIX 16, pp. 383–398. Springer, 2020b.
  22. 22.Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens Van Der Maaten. Certified data removal from machine learning models. In International Conference on Machine Learning, pp. 3832–3842. PMLR, 2020.
  23. 23.Varun Gupta, Christopher Jung, Seth Neel, Aaron Roth, Saeed Sharifi-Malvajerdi, and Chris Waites. Adaptive machine unlearning. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (eds.), Advances in Neural Information Processing Systems, volume 34, pp. 16319–16330. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper_files/paper/2021/file/87f7ee4fdb57bdfd52179947211b7ebb-Paper.pdf.
  24. 24.Anisa Halimi, Swanand Kadhe, Ambrish Rawat, and Nathalie Baracaldo. Federated unlearning: How to efficiently erase a client in fl? arXiv preprint arXiv:2207.05521, 2022.
  25. 25.Jamie Hayes, Ilia Shumailov, Eleni Triantafillou, Amr Khalifa, and Nicolas Papernot. Inexact unlearning needs more careful evaluations to avoid a false sense of privacy. arXiv preprint arXiv:2403.01218, 2024.
  26. 26.Luxi He, Yangsibo Huang, Weijia Shi, Tinghao Xie, Haotian Liu, Yue Wang, Luke Zettlemoyer, Chiyuan Zhang, Danqi Chen, and Peter Henderson. Fantastic copyrighted beasts and how (not) to generate them. arXiv preprint arXiv:2406.14526, 2024.
  27. 27.Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang. Foundation models and fair use. arXiv preprint arXiv:2303.15715, 2023.
  28. 28.Yangsibo Huang, Chun-Yin Huang, Xiaoxiao Li, and Kai Li. A dataset auditing method for collaboratively trained machine learning models. IEEE Transactions on Medical Imaging, 42(7):2081–2090, 2022.
  29. 29.Yangsibo Huang, Samyak Gupta, Zexuan Zhong, Kai Li, and Danqi Chen. Privacy implications of retrieval-based language models. arXiv preprint arXiv:2305.14888, 2023.
  30. 30.Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic, 2023.
  31. 31.Zachary Izzo, Mary Anne Smart, Kamalika Chaudhuri, and James Zou. Approximate data deletion from machine learning models. In International Conference on Artificial Intelligence and Statistics, pp. 2008–2016. PMLR, 2021.
  32. 32.Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. Knowledge unlearning for mitigating privacy risks in language models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 14389–14408, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.805. URL https://aclanthology.org/2023.acl-long.805.
  33. 33.Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. Mimic-iv. PhysioNet. Available online at: https://physionet. org/content/mimiciv/1.0/(accessed August 23, 2021), pp. 49–55, 2020.
  34. 34.Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9, 2016.
  35. 35.Masayuki Karasuyama and Ichiro Takeuchi. Multiple incremental decremental learning of support vector machines. IEEE Transactions on Neural Networks, 21(7):1048–1059, 2010.
  36. 36.Bryan Klimt and Yiming Yang. The enron corpus: A new dataset for email classification research. In European conference on machine learning, pp. 217–226. Springer, 2004.
  37. 37.Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song. Multi-step jailbreaking privacy attacks on chatgpt. arXiv preprint arXiv:2304.05197, 2023a.
  38. 38.Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al. The wmdp benchmark: Measuring and reducing malicious use with unlearning. arXiv preprint arXiv:2403.03218, 2024a.
  39. 39.Yucheng Li, Frank Guerin, and Chenghua Lin. Avoiding data contamination in language model evaluation: Dynamic test construction with latest materials, 2023b.
  40. 40.Yuyuan Li, Chaochao Chen, Xiaolin Zheng, Junlin Liu, and Jun Wang. Making recommender systems forget: Learning and unlearning for erasable recommendation. Knowledge-Based Systems, 283:111124, 2024b.
  41. 41.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.
  42. 42.Bo Liu, Qiang Liu, and Peter Stone. Continual learning and private unlearning, 2022.
  43. 43.Gaoyang Liu, Xiaoqiang Ma, Yang Yang, Chen Wang, and Jiangchuan Liu. Federated unlearning. arXiv preprint arXiv:2012.13891, 2020.
  44. 44.Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Hang Li, Kush R Varshney, et al. Rethinking machine unlearning for large language models. arXiv preprint arXiv:2402.08787, 2024.
  45. 45.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  46. 46.Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi. Quark: Controllable text generation with reinforced unlearning. Advances in neural information processing systems, 35:27591–27609, 2022.
  47. 47.Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell. Eight methods to evaluate robust unlearning in llms. arXiv preprint arXiv:2402.16835, 2024.
  48. 48.Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary Chase Lipton, and J. Zico Kolter. Tofu: A task of fictitious unlearning for llms. ArXiv, abs/2401.06121, 2024. URL https://api.semanticscholar.org/CorpusID:266933371.
  49. 49.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in gpt. Advances in Neural Information Processing Systems, 35:17359–17372, 2022.
  50. 50.Sewon Min, Suchin Gururangan, Eric Wallace, Weijia Shi, Hannaneh Hajishirzi, Noah A Smith, and Luke Zettlemoyer. Silo language models: Isolating legal risk in a nonparametric datastore. arXiv preprint arXiv:2308.04430, 2023.
  51. 51.Sasi Kumar Murakonda, Reza Shokri, and George Theodorakopoulos. Quantifying the privacy risks of learning high-dimensional graphical models. In International Conference on Artificial Intelligence and Statistics, pp. 2287–2295. PMLR, 2021.
  52. 52.Seth Neel, Aaron Roth, and Saeed Sharifi-Malvajerdi. Descent-to-delete: Gradient-based methods for machine unlearning. In Algorithmic Learning Theory, pp. 931–962. PMLR, 2021.
  53. 53.Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen. A survey of machine unlearning. arXiv preprint arXiv:2209.02299, 2022.
  54. 54.Alex Oesterling, Jiaqi Ma, Flavio Calmon, and Himabindu Lakkaraju. Fair machine unlearning: Data removal while mitigating disparities. In International Conference on Artificial Intelligence and Statistics, pp. 3736–3744. PMLR, 2024.
  55. 55.OpenAI. Gpt-4 technical report, 2023.
  56. 56.Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju. In-context unlearning: Language models as few shot unlearners. arXiv preprint arXiv:2310.07579, 2023.
  57. 57.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2023.
  58. 58.Enrique Romero, Ignacio Barrio, and Lluís Belanche. Incremental and decremental learning for linear support vector machines. In International Conference on Artificial Neural Networks, pp. 209–218. Springer, 2007.
  59. 59.Ayush Sekhari, Jayadev Acharya, Gautam Kamath, and Ananda Theertha Suresh. Remember what you want to forget: Algorithms for machine unlearning. Advances in Neural Information Processing Systems, 34:18075–18086, 2021.
  60. 60.Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. Detecting pretraining data from large language models. In The Twelfth International Conference on Learning Representations, 2024a. URL https://openreview.net/forum?id=zWqr3MQuNs.
  61. 61.Weijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Xi Victoria Lin, Noah A. Smith, Luke Zettlemoyer, Wen tau Yih, and Mike Lewis. In-context pretraining: Language modeling beyond document boundaries. In The Twelfth International Conference on Learning Representations, 2024b. URL https://openreview.net/forum?id=LXVswInHOo.
  62. 62.Saurabh Shintre, Kevin A Roundy, and Jasjeet Dhaliwal. Making machine learning forget. In Privacy Technologies and Policy: 7th Annual Privacy Forum, APF 2019, Rome, Italy, June 13–14, 2019, Proceedings 7, pp. 72–83. Springer, 2019.
  63. 63.Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP), pp. 3–18. IEEE, 2017.
  64. 64.Ilia Shumailov, Jamie Hayes, Eleni Triantafillou, Guillermo Ortiz-Jimenez, Nicolas Papernot, Matthew Jagielski, Itay Yona, Heidi Howard, and Eugene Bagdasaryan. Ununlearning: Unlearning is not sufficient for content regulation in advanced generative ai. arXiv preprint arXiv:2407.00106, 2024.
  65. 65.Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang. Knowledge unlearning for llms: Tasks, methods, and challenges. arXiv preprint arXiv:2311.15766, 2023.
  66. 66.Congzheng Song and Ananth Raghunathan. Information leakage in embedding models. In Proceedings of the 2020 ACM SIGSAC conference on computer and communications security, pp. 377–390, 2020.
  67. 67.Anvith Thudi, Gabriel Deza, Varun Chandrasekaran, and Nicolas Papernot. Unrolling sgd: Understanding factors influencing machine unlearning. In 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P), pp. 303–319. IEEE, 2022.
  68. 68.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023.
  69. 69.Eleni Triantafillou, Fabian Pedregosa, Jamie Hayes, Peter Kairouz, Isabelle Guyon, Meghdad Kurmanji, Gintare Karolina Dziugaite, Peter Triantafillou, Kairan Zhao, Lisheng Sun Hosoya, Julio C. S. Jacques Junior, Vincent Dumoulin, Ioannis Mitliagkas, Sergio Escalera, Jun Wan, Sohier Dane, Maggie Demkin, and Walter Reade. Neurips 2023 machine unlearning challenge, 2023. URL https://kaggle.com/competitions/neurips-2023-machine-unlearning.
  70. 70.Amund Tveit, Magnus Lie Hetland, and Håavard Engum. Incremental and decremental proximal support vector classification using decay coefficients. In International Conference on Data Warehousing and Knowledge Discovery, pp. 422–429. Springer, 2003.
  71. 71.Enayat Ullah, Tung Mai, Anup Rao, Ryan A Rossi, and Raman Arora. Machine unlearning via algorithmic stability. In Conference on Learning Theory, pp. 4126–4142. PMLR, 2021.
  72. 72.Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson. Assessing the brittleness of safety alignment via pruning and low-rank modifications. arXiv preprint arXiv:2402.05162, 2024a.
  73. 73.Boyi Wei, Weijia Shi, Yangsibo Huang, Noah A Smith, Chiyuan Zhang, Luke Zettlemoyer, Kai Li, and Peter Henderson. Evaluating copyright takedown methods for language models. arXiv preprint arXiv:2406.18664, 2024b.
  74. 74.Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. Depn: Detecting and editing privacy neurons in pretrained language models. arXiv preprint arXiv:2310.20138, 2023.
  75. 75.Yinjun Wu, Edgar Dobriban, and Susan Davidson. Deltagrad: Rapid retraining of machine learning models. In International Conference on Machine Learning, pp. 10355–10366. PMLR, 2020.
  76. 76.Heng Xu, Tianqing Zhu, Lefeng Zhang, Wanlei Zhou, and Yu Philip. Machine unlearning: A survey. ACM Computing Surveys, 2023.
  77. 77.Yuanshun Yao, Xiaojun Xu, and Yang Liu. Large language model unlearning. arXiv preprint arXiv:2310.10683, 2023.
  78. 78.Jiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler, and Reza Shokri. Enhanced membership inference attacks against machine learning models. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security, pp. 3093–3106, 2022a.
  79. 79.Jingwen Ye, Yifang Fu, Jie Song, Xingyi Yang, Songhua Liu, Xin Jin, Mingli Song, and Xinchao Wang. Learning with recoverable forgetting. In European Conference on Computer Vision, pp. 87–103. Springer, 2022b.
  80. 80.Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji. Unlearning bias in language models by partitioning gradients. In Findings of the Association for Computational Linguistics: ACL 2023, pp. 6032–6048, 2023.
  81. 81.Dawen Zhang, Shidong Pan, Thong Hoang, Zhenchang Xing, Mark Staples, Xiwei Xu, Lina Yao, Qinghua Lu, and Liming Zhu. To be forgotten or to be fair: Unveiling fairness implications of machine unlearning methods. AI and Ethics, pp. 1–11, 2024a.
  82. 82.Eric Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, and Humphrey Shi. Forget-me-not: Learning to forget in text-to-image diffusion models. arXiv preprint arXiv:2303.17591, 2023.
  83. 83.Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei. Negative preference optimization: From catastrophic collapse to effective unlearning, 2024b.
  84. 84.Yihua Zhang, Yimeng Zhang, Yuguang Yao, Jinghan Jia, Jiancheng Liu, Xiaoming Liu, and Sijia Liu. Unlearncanvas: A stylized image dataset to benchmark machine unlearning for diffusion models. arXiv preprint arXiv:2402.11846, 2024c.

Citation

MLA
Shi, W., et al. “MUSE: Machine Unlearning Six-Way Evaluation for Language Models”. arXiv, 2024, http://arxiv.org/abs/2407.06460v2.
APA
Shi, W., Lee, J., Huang, Y., Malladi, S., Zhao, J., Holtzman, A., Liu, D., Zettlemoyer, L., Smith, N. A., & Zhang, C. (2024). MUSE: Machine Unlearning Six-Way Evaluation for Language Models. arXiv. http://arxiv.org/abs/2407.06460v2
Chicago
Shi, W., J. Lee, Y. Huang, et al. 2024. “MUSE: Machine Unlearning Six-Way Evaluation for Language Models”. arXiv. http://arxiv.org/abs/2407.06460v2.
Harvard
Shi, W. et al. (2024) “MUSE: Machine Unlearning Six-Way Evaluation for Language Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2407.06460v2.
Vancouver
1. Shi W, Lee J, Huang Y, Malladi S, Zhao J, Holtzman A, Liu D, Zettlemoyer L, Smith NA, Zhang C (2024) MUSE: Machine Unlearning Six-Way Evaluation for Language Models. arXiv

BibTeX

@article{shi2024muse,
  title = {MUSE: Machine Unlearning Six-Way Evaluation for Language Models},
  author = {Shi, Weijia and Lee, Jaechan and Huang, Yangsibo and Malladi, Sadhika and Zhao, Jieyu and Holtzman, Ari and Liu, Daogao and Zettlemoyer, Luke and Smith, Noah A. and Zhang, Chiyuan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2407.06460v2},
  eprint = {2407.06460}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/