The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning
Nathaniel LiAlexander PanAnjali GopalSummer YueDaniel BerriosAlice GattiJustin D. LiAnn-Kathrin DombrowskiShashwat GoelGabriel Mukobi
Establishes a public benchmark to measure hazardous biological, chemical, and cyber knowledge in language models alongside a representation-control method that removes these weapon-related risks without degrading broader model capabilities.
Recent advances in artificial intelligence have raised significant security concerns regarding the dual-use potential of large language models. Policy mandates, including the White House Executive Order on Artificial Intelligence, underscore the critical need to prevent these models from assisting malicious actors in developing biological, cyber, and chemical weapons. Existing safety mechanisms primarily rely on refusal training—teaching models to decline harmful requests—which can be easily bypassed using adversarial attacks or downstream fine-tuning. Furthermore, current evaluations of dangerous capabilities are typically proprietary and narrowly focused, limiting broader scientific inquiry into effective risk mitigation.
The main objective of the article is to publicly introduce the Weapons of Mass Destruction Proxy (WMDP) benchmark to measure hazardous capabilities in language models, and to develop and evaluate Representation Misdirection for Unlearning (RMU), a novel method designed to excise this dangerous knowledge while preserving general model utility.
To construct WMDP, a consortium of academic and technical experts designed comprehensive threat models across biosecurity, cybersecurity, and chemical security. They created a dataset of 3,668 multiple-choice questions targeting precursors, neighbors, and technical components of hazardous workflows while rigorously filtering out sensitive and export-controlled information. To eliminate hazardous capabilities, the authors engineered the RMU technique, which perturbs internal neural activations on dangerous concepts while regularizing activations on benign data. The authors evaluated RMU and several baseline unlearning techniques on leading open-source models (including ZEPHYR-7B, YI-34B, and MIXTRAL-8X7B) across the WMDP dataset, standard knowledge benchmarks like MMLU, and conversational fluency tests like MT-Bench.
The analysis produced several key findings. First, advanced language models inherently possess substantial hazardous knowledge, with baseline models scoring between 44% and 75% on WMDP categories (where random chance is 25%). Second, RMU successfully reduces model performance on biosecurity and cybersecurity subsets to near-random levels (dropping to roughly 28% to 34%) while largely maintaining overall performance on MMLU and conversational fluency on MT-Bench. Third, knowledge excised through RMU proves highly robust against extraction: internal linear probes failed to recover the removed concepts, and intensive adversarial optimization attacks (such as GCG) could not elicit harmful instructions from unlearned models even after 2,500 optimization steps. Finally, evaluation on a held-out private set of high-hazard biology questions confirmed that lower scores on WMDP strongly correlate with the reduction of severe, real-world biological risks.
These findings demonstrate that machine unlearning is a viable, high-impact approach for securing closed-source models against jailbreaks and malicious fine-tuning before serving them to the public. However, the results also show collateral degradation in closely adjacent benign domains, such as introductory virology and foundational computer security, highlighting the need for higher precision to avoid inadvertently disabling defensive or legitimate scientific research.
The article recommends that model developers integrate machine unlearning into a broader safety framework alongside structured access models. Under this paradigm, unlearned models are served to the general public via application programming interfaces (APIs), while unrestricted base models are reserved for verified researchers and defensive security professionals through rigorous vetting procedures. Future technical efforts must focus on improving unlearning precision and developing architectures that prevent models from easily relearning hazardous information if weights are exposed to direct fine-tuning.
Decision-makers should interpret these results with an understanding of certain operational boundaries. WMDP relies on a multiple-choice structure, which serves as a proxy for knowledge retention rather than a direct test of end-to-end operational execution or multi-step reasoning in real-world attacks. Additionally, while RMU provides high confidence and robust protection within closed-source, API-served environments, it does not prevent adversaries from reintroducing hazardous knowledge into open-source models if they have full access to model weights.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). Provides the foundational multi-task multiple-choice evaluation paradigm that WMDP adapts into a proxy benchmark for measuring specialized hazardous capabilities and unlearning efficacy.
- Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models, Andy Zou et al. (2023). Establishes standard adversarial attack methodologies against aligned models, motivating the need for permanent unlearning of hazardous knowledge over surface-level refusal.
- Paper: Jailbroken: How Does LLM Safety Training Fail?, Alexander Wei et al. (2023). Analyzes the fundamental failure modes of standard safety fine-tuning, demonstrating why post-hoc representation engineering and unlearning are necessary to prevent severe misuse.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Introduces automated red-teaming paradigms for auditing language model risks, establishing core motivations for standardized safety evaluations.
- Paper: Ethical and social risks of harm from Language Models, Laura Weidinger et al. (2022). Presents a structured risk taxonomy of language model harms, providing the conceptual foundation for evaluating dangerous chemical, biological, and cyber capabilities.
- Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal, Mantas Mazeika et al. (2024). Complements WMDP's unlearning evaluations by introducing a standardized adversarial red-teaming benchmark across severe operational harms.
- Paper: Modular Pretraining Enables Access Control, Ethan Roland et al. (2026). Extends dual-use safety research by testing architectural access controls and selective module ablation against post-hoc unlearning baselines on dual-use domains like virology and cybersecurity.
- Paper: What Makes and Breaks Safety Fine-tuning? A Mechanistic Study, Samyak Jain et al. (2024). Mechanistically analyzes how safety interventions—including machine unlearning—alter internal representations and weight spaces to suppress dangerous concepts.
- Paper: Machine Unlearning of Pre-trained Large Language Models, Jin Yao et al. (2024). Provides a comprehensive empirical benchmark of first-order machine unlearning algorithms across pre-trained models to evaluate knowledge erasure and capability preservation.
- Paper: SafeArena: Evaluating the Safety of Autonomous Web Agents, Ada Defne Tur et al. (2025). Evaluates how autonomous agent architectures resist malicious cyber and illegal instructions when deployed in dynamic online environments.
- Paper: Position: TrustLLM: Trustworthiness in Large Language Models, Yue Huang et al. (2024). Broadens the evaluation of model safety and trustworthiness by establishing a comprehensive multi-dimensional benchmark beyond dual-use knowledge.
