MUSE: Machine Unlearning Six-Way Evaluation for Language Models
Weijia ShiJaechan LeeYangsibo HuangSadhika MalladiJieyu ZhaoAri HoltzmanDaogao LiuLuke ZettlemoyerNoah SmithChiyuan Zhang
Introduces MUSE, a six-property evaluation benchmark for machine unlearning in language models, and reveals that popular unlearning methods cause severe privacy leakage and degrade model utility during sequential or large-scale data removal.
Modern language models are trained on massive text collections that frequently contain copyrighted works and private personal information. Due to rising privacy regulations like the General Data Protection Regulation and ongoing copyright lawsuits, data owners increasingly demand the removal of their data from deployed systems. However, perfectly erasing specific data by retraining models from scratch is computationally intractable, prompting the development of approximate machine unlearning methods that attempt to remove targeted information without full retraining.
The article aims to evaluate the practical effectiveness and viability of existing approximate machine unlearning techniques across six critical dimensions that capture the needs of both data owners and model deployers. Specifically, it introduces a comprehensive evaluation framework called MUSE (Machine Unlearning Six-Way Evaluation) to assess whether these algorithms can reliably remove data without damaging general model performance or leaking private information.
To conduct this evaluation, the researchers tested eight leading unlearning algorithms applied to 7-billion-parameter language models across two realistic datasets: a news dataset comprising up to 3.3 million tokens of recent BBC articles and a books dataset comprising 1.1 million tokens from the Harry Potter series. The evaluation framework measures six specific properties: prevention of verbatim memorization, elimination of factual knowledge retention, prevention of privacy leakage via membership detection, preservation of general model utility on retained data, scalability to larger removal requests, and sustainability over sequential removal requests.
The investigation produced four central findings. First, while most unlearning algorithms successfully suppress verbatim text generation and factual recall—in some cases reducing memorization scores to zero—they fail to prevent privacy leakage. Algorithms consistently either under-unlearn or over-unlearn, creating statistical anomalies that allow membership inference attacks to detect whether the target data was originally in the training set. Second, existing methods cause severe model degradation; every tested algorithm reduced model performance on retained data by 24% to 100%, with several methods completely destroying general model utility. Third, unlearning performance degrades rapidly as the volume of data to be forgotten scales up from 0.8 million to 3.3 million tokens. Fourth, current algorithms lack sustainability, showing steep drops in utility when handling successive, sequential unlearning requests.
These findings demonstrate that current approximate unlearning techniques create a false sense of security and are not ready for practical or regulatory deployment. For organizations facing legal or compliance pressures, relying on these methods introduces significant risks: they fail to ensure true data privacy, carry high risks of operational degradation, and cannot sustain real-world workloads where data removal requests arrive over time.
Decision-makers should exercise caution and avoid deploying current approximate unlearning methods in production environments where compliance or privacy guarantees are strictly required. Instead, organizations must treat data removal as an open technical challenge, using comprehensive multi-metric evaluations like MUSE to validate emerging algorithms. Future research and development should focus on designing balanced unlearning objectives that preserve general model capabilities while preventing statistical signatures that expose member privacy.
The primary limitations of this evaluation include its focus on 7-billion-parameter text models across news and literature domains, leaving multimodal architectures, smaller or larger language models, and specialized domains such as medical records or emails for future study. Nonetheless, the evidence strongly supports high confidence in the conclusion that current unlearning methods fail to meet real-world operational and privacy requirements.
- Paper: Machine Unlearning, Lucas Bourtoule et al. (2019). Read this foundational SISA work first to understand approximate unlearning, retraining baselines, and the deletion setting that MUSE evaluates at language-model scale.
- Paper: Extracting Training Data from Large Language Models, Nicholas Carlini et al. (2020). Its demonstration of verbatim training-data extraction establishes the memorization and privacy risks that MUSE measures as unlearning targets.
- Paper: Machine Unlearning of Pre-trained Large Language Models, Jin Yao et al. (2024). This direct LLM unlearning benchmark introduces methods and evaluation concerns for pre-trained models that MUSE extends with broader criteria, including sequential sustainability and scale.
- Paper: The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning, Nathaniel Li et al. (2024). Its RMU experiments provide a concrete LLM unlearning method and utility-preservation evaluation that help contextualize MUSE’s comparison of unlearning algorithms.
- Paper: LLM Dataset Inference: Did you train on my dataset?, Pratyush Maini et al. (2024). Its analysis of dataset-level membership inference clarifies how privacy leakage can persist beyond verbatim recall, a distinction central to MUSE’s evaluation.
- Paper: Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient, Yongliang Wu et al. (2025). This 2025 diffusion-model study carries the unlearning evaluation problem into generative image models, testing generalization of removal and preservation of unrelated utility.
