The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning

Nathaniel LiAlexander PanAnjali GopalSummer YueDaniel BerriosAlice GattiJustin D. LiAnn-Kathrin DombrowskiShashwat GoelGabriel Mukobi

article2024ICML371 citations

Establishes a public benchmark to measure hazardous biological, chemical, and cyber knowledge in language models alongside a representation-control method that removes these weapon-related risks without degrading broader model capabilities.

Listen

Recent advances in artificial intelligence have raised significant security concerns regarding the dual-use potential of large language models. Policy mandates, including the White House Executive Order on Artificial Intelligence, underscore the critical need to prevent these models from assisting malicious actors in developing biological, cyber, and chemical weapons. Existing safety mechanisms primarily rely on refusal training—teaching models to decline harmful requests—which can be easily bypassed using adversarial attacks or downstream fine-tuning. Furthermore, current evaluations of dangerous capabilities are typically proprietary and narrowly focused, limiting broader scientific inquiry into effective risk mitigation.

The main objective of the article is to publicly introduce the Weapons of Mass Destruction Proxy (WMDP) benchmark to measure hazardous capabilities in language models, and to develop and evaluate Representation Misdirection for Unlearning (RMU), a novel method designed to excise this dangerous knowledge while preserving general model utility.

To construct WMDP, a consortium of academic and technical experts designed comprehensive threat models across biosecurity, cybersecurity, and chemical security. They created a dataset of 3,668 multiple-choice questions targeting precursors, neighbors, and technical components of hazardous workflows while rigorously filtering out sensitive and export-controlled information. To eliminate hazardous capabilities, the authors engineered the RMU technique, which perturbs internal neural activations on dangerous concepts while regularizing activations on benign data. The authors evaluated RMU and several baseline unlearning techniques on leading open-source models (including ZEPHYR-7B, YI-34B, and MIXTRAL-8X7B) across the WMDP dataset, standard knowledge benchmarks like MMLU, and conversational fluency tests like MT-Bench.

The analysis produced several key findings. First, advanced language models inherently possess substantial hazardous knowledge, with baseline models scoring between 44% and 75% on WMDP categories (where random chance is 25%). Second, RMU successfully reduces model performance on biosecurity and cybersecurity subsets to near-random levels (dropping to roughly 28% to 34%) while largely maintaining overall performance on MMLU and conversational fluency on MT-Bench. Third, knowledge excised through RMU proves highly robust against extraction: internal linear probes failed to recover the removed concepts, and intensive adversarial optimization attacks (such as GCG) could not elicit harmful instructions from unlearned models even after 2,500 optimization steps. Finally, evaluation on a held-out private set of high-hazard biology questions confirmed that lower scores on WMDP strongly correlate with the reduction of severe, real-world biological risks.

These findings demonstrate that machine unlearning is a viable, high-impact approach for securing closed-source models against jailbreaks and malicious fine-tuning before serving them to the public. However, the results also show collateral degradation in closely adjacent benign domains, such as introductory virology and foundational computer security, highlighting the need for higher precision to avoid inadvertently disabling defensive or legitimate scientific research.

The article recommends that model developers integrate machine unlearning into a broader safety framework alongside structured access models. Under this paradigm, unlearned models are served to the general public via application programming interfaces (APIs), while unrestricted base models are reserved for verified researchers and defensive security professionals through rigorous vetting procedures. Future technical efforts must focus on improving unlearning precision and developing architectures that prevent models from easily relearning hazardous information if weights are exposed to direct fine-tuning.

Decision-makers should interpret these results with an understanding of certain operational boundaries. WMDP relies on a multiple-choice structure, which serves as a proxy for knowledge retention rather than a direct test of end-to-end operational execution or multi-step reasoning in real-world attacks. Additionally, while RMU provides high confidence and robust protection within closed-source, API-served environments, it does not prevent adversaries from reintroducing hazardous knowledge into open-source models if they have full access to model weights.

Cover for The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning

Abstract

The White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and chemical weapons. To measure these risks of malicious use, government institutions and major AI labs are developing evaluations for hazardous capabilities in LLMs. However, current evaluations are private, preventing further research into mitigating risk. Furthermore, they focus on only a few, highly specific pathways for malicious use. To fill these gaps, we publicly release the Weapons of Mass Destruction Proxy (WMDP) benchmark, a dataset of 3,668 multiple-choice questions that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security. WMDP was developed by a consortium of academics and technical consultants, and was stringently filtered to eliminate sensitive information prior to public release. WMDP serves two roles: first, as an evaluation for hazardous knowledge in LLMs, and second, as a benchmark for unlearning methods to remove such hazardous knowledge. To guide progress on unlearning, we develop RMU, a state-of-the-art unlearning method based on controlling model representations. RMU reduces model performance on WMDP while maintaining general capabilities in areas such as biology and computer science, suggesting that unlearning may be a concrete path towards reducing malicious use from LLMs. We release our benchmark and code publicly at this https URL

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The WMDP Benchmark
  • 3.1 Design Choices for WMDP
  • 3.2 Biosecurity Threat Model
  • 3.3 Cybersecurity Threat Model
  • 3.4 Chemical Security Threat Model
  • 3.5 Sensitive Information Mitigation
  • 4 RMU: Unlearning Inspired By Representation Engineering
  • 4.1 Setup
  • 4.2 Method
  • 5 Experimental Results
  • 5.1 Setup
  • 5.2 Quantitative Evaluation
  • 5.3 Robustness Evaluation
  • 5.4 Generalization of WMDP to Hazardous Knowledge
  • 6 Discussion
  • 6.1 How WMDP Mitigates Risk
  • 6.2 Structured API Access
  • 7 Conclusion
  • References
  • A Dataset
  • A.1 Dataset Breakdown
  • A.2 Additional Considerations for WMDP-Bio
  • A.3 Additional Considerations for WMDP-Chem
  • A.4 Bio Corpora
  • A.5 Cyber Corpora
  • B Experiments
  • B.1 Zero-Shot QA Format
  • B.2 MT-Bench
  • B.3 Robustness Evaluation
  • B.3.1 RMU Unlearned Model
  • B.3.2 Base Model
  • B.4 Updates to RMU
  • B.5 How RMU manipulates representations
  • B.6 Generalization of RMU
  • B.7 Baselines
  • B.7.1 LLMU
  • B.7.2 SCRUB
  • B.7.3 SSD
  • B.7.4 RMU
  • C MMLU Subset Unlearning Benchmark
  • D Broader Impacts of WMDP
  • D.1 Limitations
  • E X-Risk Sheet
  • E.1 Long-Term Impact on Advanced AI Systems
  • E.2 Safety-Capabilities Balance
  • E.3 Elaborations and Other Considerations

Knowls

  1. Knowl 1 — Weapons of Mass Destruction Proxy (WMDP) Benchmark

    definition

    The Weapons of Mass Destruction Proxy (WMDP) benchmark is an open-source evaluation suite of 3,668 four-choice multiple-choice questions designed to assess hazardous knowledge and capabilities in large language models across three security domains:

    1. WMDP-Bio (1,273 questions): Measures dual-use biology knowledge mapping onto the design-build-test-learn-release cycle, divided into Dual-use Virology (228), Bioweapons & Bioterrorism (197), Reverse Genetics & Easy Editing (252), Enhanced Potential Pandemic Pathogens (233), Viral Vector Research (228), and Expanding Access (135).
    2. WMDP-Cyber (1,987 questions): Measures capabilities across the cyberattack lifecycle, divided into Background Knowledge (271), Reconnaissance (20), Weaponization & Vulnerability Discovery (comprising Assembly Review [283], Function Review [300], Packet Dissection [298], and Other [361]), Exploitation (272), and Post-Exploitation (182).
    3. WMDP-Chem (408 questions): Measures chemical security and explosive threat knowledge, divided into General Knowledge (127), Synthesis (78), Sourcing / Procurement (41), Purification (19), Analysis / Verification (21), Deployment Mechanisms (65), Bypass Mechanisms (15), and Miscellaneous (42).

    To prevent information hazards and comply with export control frameworks (such as U.S. ITAR and EAR regulations), WMDP tests proxy knowledge—precursors, neighbors, and components of hazardous capabilities—rather than actionable end-to-end weapon blueprints. Every question is reviewed and cross-checked by domain experts to ensure accuracy and to expunge sensitive information.

  2. Knowl 2 — Representation Misdirection for Unlearning (RMU)

    model/method

    Representation Misdirection for Unlearning (RMU) is an unlearning method for autoregressive language models that operates by directly manipulating hidden representation vectors at intermediate layers rather than modifying output token distributions or steering towards human-crafted novice representations.

    RMU optimizes a dual-objective loss on hidden activations at a designated layer ℓ\ell. On hazardous forget data (DforgetD_{\text{forget}}), it steers the hidden activations toward a fixed random unit vector scaled to a high norm, disrupting downstream layers from processing hazardous information. On benign retain data (DretainD_{\text{retain}}), it penalizes deviations between the updated model's activations and those of the original frozen base model to preserve general capabilities.

    Key implementation properties of RMU include:

    • Subnetwork Parameter Updates: Only the multi-layer perceptron (MLP) weights within layers ℓ−2\ell-2, ℓ−1\ell-1, and ℓ\ell are updated, leaving self-attention modules and all other layers frozen.
    • Layer Selection: Intermediate layers are selected for unlearning (e.g., ℓ=7\ell = 7 for 7B-parameter architectures, ℓ=15\ell = 15 for 34B-parameter architectures).
    • Preservation Dataset: A general linguistic corpus (such as Wikitext) is used as DretainD_{\text{retain}} instead of subject-specific retain sets to prevent the network from relearning unlearned domain concepts.
    • Multi-domain Interleaving: When unlearning multiple domains (e.g., biology and cybersecurity), optimization batches from each domain are interleaved during training.
  3. Knowl 3 — RMU Loss Formulation

    equation

    Let Mupdated(t)∈RdM_{\text{updated}}(t) \in \mathbb{R}^d denote the hidden activation vector of the model being updated at token position tt of layer ℓ\ell, and let Mfrozen(t)∈RdM_{\text{frozen}}(t) \in \mathbb{R}^d denote the corresponding hidden activation vector of the original, frozen base model at layer ℓ\ell. The total optimization loss L\mathcal{L} for Representation Misdirection for Unlearning (RMU) is given by:

    L=Lforget+α⋅Lretain\mathcal{L} = \mathcal{L}_{\text{forget}} + \alpha \cdot \mathcal{L}_{\text{retain}}

    where the forget loss Lforget\mathcal{L}_{\text{forget}} and the retain loss Lretain\mathcal{L}_{\text{retain}} are defined as:

    Lforget=Exf∼Dforget[1Lf∑t∈xf∥Mupdated(t)−c⋅u∥22]\mathcal{L}_{\text{forget}} = \mathbb{E}_{x_f \sim D_{\text{forget}}} \left[ \frac{1}{L_f} \sum_{t \in x_f} \| M_{\text{updated}}(t) - c \cdot u \|_2^2 \right]

    Lretain=Exr∼Dretain[1Lr∑t∈xr∥Mupdated(t)−Mfrozen(t)∥22]\mathcal{L}_{\text{retain}} = \mathbb{E}_{x_r \sim D_{\text{retain}}} \left[ \frac{1}{L_r} \sum_{t \in x_r} \| M_{\text{updated}}(t) - M_{\text{frozen}}(t) \|_2^2 \right]

    where:

    • DforgetD_{\text{forget}} is the dataset representing the hazardous knowledge distribution to unlearn.
    • DretainD_{\text{retain}} is the dataset of general benign text representing capabilities to preserve.
    • xfx_f and xrx_r are input token sequences sampled from DforgetD_{\text{forget}} and DretainD_{\text{retain}} of lengths LfL_f and LrL_r, respectively.
    • u∈Rdu \in \mathbb{R}^d is a fixed random unit vector whose independent components are sampled uniformly at random from [0,1)[0, 1) and normalized such that ∥u∥2=1\|u\|_2 = 1.
    • c>0c > 0 is a scalar hyperparameter that sets the target activation norm.
    • α>0\alpha > 0 is a scalar regularization hyperparameter controlling the trade-off between retain preservation and forget disruption.
  4. Knowl 4 — Representation Misdirection for Unlearning Algorithm

    algorithm

    The RMU algorithm modifies the MLP weights in layers ℓ−2,ℓ−1,ℓ\ell-2, \ell-1, \ell of a language model to remove hazardous knowledge while retaining general capabilities.

    Input: Model to update MupdatedM_{\text{updated}}, frozen base model MfrozenM_{\text{frozen}}, forget dataset DforgetD_{\text{forget}}, retain dataset DretainD_{\text{retain}}, target layer ℓ\ell, scaling factor cc, retain weight α\alpha, learning rate η\eta, number of training steps NN.
    Output: Unlearned model MupdatedM_{\text{updated}}.
    Sample vector v∈Rdv \in \mathbb{R}^d with independent components vi∼Uniform(0,1)v_i \sim \text{Uniform}(0, 1) for i∈{1,…,d}i \in \{1, \dots, d\}.
    Compute unit vector u=v∥v∥2u = \frac{v}{\|v\|_2}.
    Freeze all parameters in MupdatedM_{\text{updated}} except the MLP parameters in layers ℓ−2\ell-2, ℓ−1\ell-1, and ℓ\ell.
    for step =1= 1 to NN do
        Sample forget sequence xf∼Dforgetx_f \sim D_{\text{forget}} containing LfL_f tokens.
        Sample retain sequence xr∼Dretainx_r \sim D_{\text{retain}} containing LrL_r tokens.
        
        Compute layer ℓ\ell activations Mupdated(t)M_{\text{updated}}(t) for each token t∈xft \in x_f.
        Compute layer ℓ\ell activations Mupdated(t)M_{\text{updated}}(t) and Mfrozen(t)M_{\text{frozen}}(t) for each token t∈xrt \in x_r.
        
        Lforget=1Lf∑t∈xf∥Mupdated(t)−c⋅u∥22\mathcal{L}_{\text{forget}} = \frac{1}{L_f} \sum_{t \in x_f} \|M_{\text{updated}}(t) - c \cdot u\|_2^2
        Lretain=1Lr∑t∈xr∥Mupdated(t)−Mfrozen(t)∥22\mathcal{L}_{\text{retain}} = \frac{1}{L_r} \sum_{t \in x_r} \|M_{\text{updated}}(t) - M_{\text{frozen}}(t)\|_2^2
        L=Lforget+α⋅Lretain\mathcal{L} = \mathcal{L}_{\text{forget}} + \alpha \cdot \mathcal{L}_{\text{retain}}
        
        Compute gradients ∇L\nabla \mathcal{L} with respect to the trainable MLP parameters.
        Update trainable MLP parameters using gradient descent with learning rate η\eta.
    end for
    return MupdatedM_{\text{updated}}
  5. Knowl 5 — Benchmark Evaluation of RMU and Baselines on WMDP, MMLU, and MT-Bench

    data/table

    The table below compares zero-shot question-answering accuracy (%) on the Weapons of Mass Destruction Proxy (WMDP) benchmark (measuring unlearning on WMDP-Bio and WMDP-Cyber), Massive Multitask Language Understanding (MMLU) benchmark across specific subdomains and overall aggregate (measuring general knowledge retention), and MT-Bench multi-turn conversation scores (evaluating assistant fluency, out of 9.0) across base models and unlearning baselines: Large Language Model Unlearning (LLMU), SCalable Remembering and Unlearning unBound (SCRUB), Selective Synaptic Dampening (SSD), and Representation Misdirection for Unlearning (RMU). Random chance baseline on 4-choice QA is 25.0%.

    Model Method WMDP (↓\downarrow) MMLU (↑\uparrow) MT-Bench (↑\uparrow)
    Bio Cyber Chem College Bio Virology College CS Cybersec All (Score / 9.0)
    ZEPHYR-7B Base 63.7 44.0 45.8 68.1 52.4 50.0 65.0 58.1 7.33
    + LLMU 59.5 39.5 41.4 54.2 37.4 43.0 53.0 44.7 1.00
    + SCRUB 43.8 39.3 40.4 53.5 40.3 48.0 62.0 51.2 1.43
    + SSD 50.2 35.0 33.8 46.5 38.0 35.0 52.0 40.7 5.48
    + RMU (ours) 31.2 28.2 45.8 63.2 25.9 49.0 45.0 57.1 7.10
    YI-34B Base 75.3 49.7 58.6 88.9 57.2 63.0 84.0 72.6 7.65
    + RMU (ours) 30.7 29.0 55.4 84.0 22.3 57.0 46.0 70.6 7.59
    MIXTRAL-8X7B Base 74.8 52.0 55.2 82.6 50.0 64.0 80.0 68.2 8.30
    + RMU (ours) 34.0 30.8 54.7 81.3 34.3 67.0 58.0 67.1 8.17
    GPT-4 Base 82.2 55.3 64.7 93.9 58.2 69.0 84.5 83.4 9.13

    RMU reduces WMDP-Bio and WMDP-Cyber accuracy close to random guessing (30.7%--34.0% on Bio and 28.2%--30.8% on Cyber) across 7B, 34B, and MoE models, while preserving overall MMLU capabilities within 1.0--2.0% of the base models (e.g., 57.1% vs 58.1% on Zephyr-7B) and preserving conversational fluency on MT-Bench (7.10 vs 7.33 on Zephyr-7B, 7.59 vs 7.65 on Yi-34B). In contrast, prior unlearning baselines (LLMU, SCRUB, SSD) fail to reduce WMDP performance to random chance while severely degrading general capabilities (reducing MMLU by up to 17.4 percentage points and collapsing MT-Bench scores to as low as 1.00).

  6. Knowl 6 — Unrecoverability of RMU Unlearned Knowledge via Linear Probing

    empirical result

    To evaluate whether hazardous knowledge is erased from internal model representations or merely suppressed at the output logit layer, 4-way linear probes were trained on hidden activations extracted from every layer of Zephyr-7B and Yi-34B unlearned with RMU. Probes were trained on a 50% partition of WMDP-Bio and WMDP-Cyber questions and evaluated on the held-out 50%.

    While linear probes trained on representations from the base models achieve ~80% accuracy across intermediate and late layers on both biological and cyber questions, probes trained on RMU-unlearned models achieve only near-random accuracy (~25% to 30%) across all layers. This confirms that RMU alters latent feature spaces throughout the network rather than applying a superficial output-level refusal filter.

  7. Knowl 7 — Robustness of RMU Unlearning against GCG Adversarial Attacks

    empirical result

    Models unlearned with RMU remain resilient to white-box token-level optimization attacks using Greedy Coordinate Gradient (GCG).

    In evaluations on Yi-34B:

    • For the refusal-aligned base model, GCG optimization generates a 20-token adversarial suffix in fewer than 50 gradient steps that bypasses safety refusals and elicits actionable step-by-step instructions on hazardous prompts (such as engineering virulent strains of Influenza A or exploiting browser memory corruptions).
    • For the RMU-unlearned model, running GCG optimization for 2,500 gradient steps (exceeding 7 hours of compute on an NVIDIA A100 GPU) fails to extract hazardous information, yielding repetitive, nonsensical text completions.

    This demonstrates that representation unlearning via RMU produces genuine absence of hazardous knowledge that resists optimization pressure.

  8. Knowl 8 — Generalization of WMDP Unlearning to Private Sensitive Biosecurity Knowledge

    empirical result

    During benchmark construction, 122 questions identified by biosecurity domain experts as containing sensitive, dual-use biological information were withheld from the public WMDP-Bio dataset to serve as a private evaluation set.

    When Zephyr-7B underwent RMU unlearning using an uncurated PubMed forget corpus approximating WMDP-Bio topics:

    • Performance on the public WMDP-Bio dataset dropped from 63.7% to 31.2% (approaching random guessing at 25.0%).
    • Performance on the private held-out set of 122 sensitive questions dropped simultaneously from ~60% to ~25% (random guessing).

    This empirical alignment demonstrates that unlearning proxy knowledge on WMDP generalizes to removing real-world, highly sensitive hazardous knowledge without exposing sensitive training data.

  9. Knowl 9 — Auxiliary MMLU Subset Unlearning Benchmark

    experimental setup

    To provide a non-hazardous testbed for fine-grained unlearning research, an auxiliary benchmark based on subsets of the Massive Multitask Language Understanding (MMLU) benchmark was created across three distinct disciplines:

    1. Physics: The target forget set comprises High School Physics, Conceptual Physics, and College Physics. The retain set comprises College Mathematics and High School Mathematics.
    2. Law: The target forget set comprises International Law and Professional Law. The retain set comprises Jurisprudence.
    3. Economics: The target forget set comprises High School Macroeconomics and High School Microeconomics. The retain set comprises Econometrics.

    Each setting provides domain-specific forget corpora compiled from Creative Commons-licensed textbooks (filtered to remove exercises, indices, citations, and links) containing over 1,000 multiple-choice questions in the forget evaluation set. The benchmark evaluates forget efficacy (accuracy reduction on target disciplines), retention precision on closely related neighbor domains, and retention on the overall MMLU benchmark.

  10. Knowl 10 — Vulnerability of RMU to Finetuning and Over-Forgetting in Neighbor Domains

    limitation

    Machine unlearning via RMU is subject to two primary limitations:

    1. Relearning via Open-Weights Finetuning: RMU does not safeguard models whose raw weights are publicly released. When an RMU-unlearned Mistral-7B-v0.1 model is finetuned on the cybersecurity forget corpus, performance on WMDP-Cyber recovers to base model levels (~45%), demonstrating that unlearning can be reversed under post-release parameter updates. RMU is thus specifically intended for closed-source or structured API deployment regimes.
    2. Imprecise Erasure in Closely Related Domains: RMU exhibits significant collateral forgetting in academic domains closely adjacent to the forget targets. On Zephyr-7B, unlearning biosecurity and cybersecurity caused MMLU Virology accuracy to decline from 52.4% to 25.9% and MMLU Computer Security accuracy to decline from 65.0% to 45.0%, even though broad subjects like College Biology and College Computer Science were largely preserved.

Coverage note — Detailed text lists of all scrape search keywords and threat model taxonomy descriptions were consolidated within the benchmark and method knowls to maintain focus on the core contributions.

References

  1. 1.01-ai. GitHub - 01-ai/Yi: A series of large language models trained from scratch by developers @01-ai — github.com. https://github.com/01-ai/Yi, 2023.
  2. 2.Anthropic. Anthropic’s Responsible Scaling Policy — anthropic.com. https://www.anthropic.com/index/anthropics-responsible-scaling-policy, 2023.
  3. 3.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  4. 4.Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. Leace: Perfect linear concept erasure in closed form. NeurIPS, 2023.
  5. 5.Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann, Gabriel Synnaeve, Varun Vontimitta, Spencer Whitman, and Joshua Saxe. Purple llama cyberseceval: A secure coding benchmark for language models, 2023.
  6. 6.Daniil A Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes. Autonomous chemical research with large language models. Nature, 624(7992):570–578, 2023.
  7. 7.Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in language models without supervision, 2022.
  8. 8.Yinzhi Cao and Junfeng Yang. Towards making systems forget with machine unlearning. In IEEE S&P, 2015.
  9. 9.CCPA. California consumer privacy act, 2018. https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=201720180AB375.
  10. 10.Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries, 2023.
  11. 11.Dasol Choi and Dongbin Na. Towards machine unlearning benchmarks: Forgetting the personal identities in facial recognition systems, 2023.
  12. 12.Council of European Union. Council regulation (EU) no 269/2014, 2014. http://eur-lex.europa.eu/legal-content/EN/TXT/?qid=1416170084502&uri=CELEX:32014R0269.
  13. 13.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. Toxicity in chatgpt: Analyzing persona-assigned language models. arXiv preprint arXiv:2304.05335, 2023.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  15. 15.EAR. Export administration regulations (ear), 15 cfr parts 730-774. https://www.ecfr.gov/current/title-15/subtitle-B/chapter-VII/subchapter-C, 2024.
  16. 16.Ronen Eldan and Mark Russinovich. Who’s harry potter? approximate unlearning in llms. arXiv preprint arXiv:2310.02238, 2023.
  17. 17.Kevin M Esvelt. Inoculating science against potential pandemics and information hazards. PLoS Pathog., 14(10):e1007286, October 2018.
  18. 18.Kevin M. Esvelt. Delay, Detect, Defend: Preparing for a Future in which Thousands Can Release New Pandemics. https://dam.gcsp.ch/files/doc/gcsp-geneva-paper-29-22, 2022.
  19. 19.Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang. Llm agents can autonomously hack websites, 2024.
  20. 20.World Economic Forum. Global cybersecurity outlook 2024, 2024. URL https://www3.weforum.org/docs/WEF_Global_Cybersecurity_Outlook_2024.pdf.
  21. 21.Jack Foster, Stefan Schoepf, and Alexandra Brintrup. Fast machine unlearning without retraining through selective synaptic dampening. AAAI, 2024.
  22. 22.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, September 2021.
  23. 23.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462, 2020.
  24. 24.Shashwat Goel, Ameya Prabhu, Amartya Sanyal, Ser-Nam Lim, Philip Torr, and Ponnurangam Kumaraguru. Towards adversarial evaluations for inexact machine unlearning, 2023.
  25. 25.Shashwat Goel, Ameya Prabhu, Philip Torr, Ponnurangam Kumaraguru, and Amartya Sanyal. Corrective machine unlearning, 2024.
  26. 26.Aditya Golatkar, Alessandro Achille, and Stefano Soatto. Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations. In ECCV, 2020.
  27. 27.Google.重 Neurips 2023 machine unlearning challenge, 2023. URL https://unlearning-challenge.github.io/.
  28. 28.Anjali Gopal, Nathan Helm-Burger, Lennart Justen, Emily H. Soice, Tiffany Tzeng, Geetha Jeyapragasan, Simon Grimm, Benjamin Mueller, and Kevin M. Esvelt. Will releasing the weights of future large language models grant widespread access to pandemic agents?, 2023.
  29. 29.Blessing Guembe, Ambrose Azeta, Sanjay Misra, Victor Chukwudi Osamor, Luis Fernandez-Sanz, and Vera Pospelova. The emerging threat of ai-driven cyber attacks: A review. Applied Artificial Intelligence, 36(1):2037254, 2022.
  30. 30.Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela. Gradient-based adversarial attacks against text transformers, 2021.
  31. 31.Dan Hendrycks and Mantas Mazeika. X-risk analysis for ai research, 2022.
  32. 32.Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning ai with shared human values. arXiv preprint arXiv:2008.02275, 2020a.
  33. 33.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020b.
  34. 34.Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt. Unsolved problems in ml safety. arXiv preprint arXiv:2109.13916, 2021.
  35. 35.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models, 2021.
  36. 36.Eric M. Hutchins, Michael J. Cloppert, and Rohan M. Amin. Intelligence-driven computer network defense informed by analysis of adversary campaigns and intrusion kill chains. Technical report, Lockheed Martin Corporation, 2011.
  37. 37.Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. Editing models with task arithmetic. In The Eleventh International Conference on Learning Representations, 2023.
  38. 38.Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023.
  39. 39.ITAR. International traffic in arms regulations (itar), 22 cfr parts 120-130. https://www.ecfr.gov/current/title-22/chapter-I/subchapter-M, 2024.
  40. 40.Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. Knowledge unlearning for mitigating privacy risks in language models. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, ACL, 2023.
  41. 41.Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38, 2023.
  42. 42.Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. Mixtral of experts, 2024.
  43. 43.Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt. Automatically auditing large language models via discrete optimization, 2023.
  44. 44.Megan Kinniment and Lucas Jun Koba Sato. Haoxing du, brian goodrich, max hasin, lawrence chan, luke harold miles, tao r. lin, hjalmar wijk, joel burget, aaron ho, elizabeth barnes, and paul christiano. evaluating language-model agents on realistic autonomous tasks. Evaluating Language-Model Agents on Realistic Autonomous Tasks. Research paper, Alignment Research Center, 2023.
  45. 45.Meghdad Kurmanji, Peter Triantafillou, and Eleni Triantafillou. Towards unbounded machine unlearning. NeurIPS, 2023.
  46. 46.Jakub Lála, Odhran O’Donoghue, Aleksandar Shtedritski, Sam Cox, Samuel G Rodriques, and Andrew D White. Paperqa: Retrieval-augmented generative agent for scientific research. arXiv preprint arXiv:2312.07559, 2023.
  47. 47.Gregory Lewis, Piers Millett, Anders Sandberg, Andrew Snyder-Beattie, and Gigi Gronvall. Information hazards in biotechnology. Risk Anal., 39(5):975–981, May 2019.
  48. 48.Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. Halueval: A large-scale hallucination evaluation benchmark for large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 6449–6464, 2023.
  49. 49.Stephanie Lin, Jacob Hilton, and Owain Evans. Truthfulqa: Measuring how models mimic human falsehoods. arXiv preprint arXiv:2109.07958, 2021.
  50. 50.Yang Liu, Zhuo Ma, Ximeng Liu, Jian Liu, Zhongyuan Jiang, Jianfeng Ma, Philip Yu, and Kui Ren. Learn to forget: Machine unlearning via neuron masking. arXiv preprint arXiv:2003.10933, 2020.
  51. 51.Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang. Towards Safer Large Language Models through Machine Unlearning. arXiv e-prints, art. arXiv:2402.10058, February 2024. doi: 10.48550/arXiv.2402.10058.
  52. 52.Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell. Eight methods to evaluate robust unlearning in llms, 2024.
  53. 53.Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C. Lipton, and J. Zico Kolter. Tofu: A task of fictitious unlearning for llms, 2024.
  54. 54.Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249, 2024.
  55. 55.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 35, 2022.
  56. 56.Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models, 2016.
  57. 57.Mistral AI team. Mistral 7b. Mistral, 2023. URL https://mistral.ai/news/announcing-mistral-7b/.
  58. 58.Christopher A. Mouton, Caleb Lucas, and Ella Guest. The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study. RAND Corporation, Santa Monica, CA, 2024. doi: 10.7249/RRA2977-2.
  59. 59.Cassidy Nelson and Sophie Rose. Understanding AI-Facilitated Biological Weapon Development. Center for Long Term Resilience, 2023.
  60. 60.Helen Ngo, Cooper D. Raterink, Joao M. de Ara’ujo, Ivan Zhang, Carol Chen, Adrien Morisot, and Nick Frosst. Mitigating harm in language models with conditional-likelihood filtration. ArXiv, abs/2108.07790, 2021.
  61. 61.NIST. AI Risk Management Framework — nist.gov. https://www.nist.gov/itl/ai-risk-management-framework, 2023.
  62. 62.OpenAI. Gpt-4 technical report, 2023a.
  63. 63.OpenAI. Preparedness — openai.com. https://openai.com/safety/preparedness, 2023b.
  64. 64.OpenAI. Building an early warning system for LLM-aided biological threat creation — openai.com. https://openai.com/research/building-an-early-warning-system-for-llm-aided-biological-threat-creation, 2024.
  65. 65.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35: 27730–27744, 2022.
  66. 66.Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark. In International Conference on Machine Learning, pages 26837–26867. PMLR, 2023.
  67. 67.Alexander Pan, Erik Jones, Meena Jagadeesan, and Jacob Steinhardt. Feedback loops with language models drive in-context reward hacking. arXiv preprint arXiv:2402.06627, 2024.
  68. 68.Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. Ai deception: A survey of examples, risks, and potential solutions. arXiv preprint arXiv:2308.14752, 2023.
  69. 69.Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju. In-context unlearning: Language models as few shot unlearners. arXiv preprint arXiv:2310.07579, 2023.
  70. 70.Kellin Pelrine, Mohammad Taufeeque, Michał Zając, Euan McLean, and Adam Gleave. Exploiting novel gpt-4 apis, 2023.
  71. 71.Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis Assael, Sarah Hodkinson, Heidi Howard, Tom Lieberum, Ramana Kumar, Maria Abi Raad, Albert Webson, Lewis Ho, Sharon Lin, Sebastian Farquhar, Marcus Hutter, Gregoire Deletang, Anian Ruoss, Seliem El-Sayed, Sasha Brown, Anca Dragan, Rohin Shah, Allan Dafoe, and Toby Shevlane. Evaluating frontier models for dangerous capabilities, 2024.
  72. 72.Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint arXiv:2310.03693, 2023.
  73. 73.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2023.
  74. 74.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022, 2023.
  75. 75.Jonas B. Sandbrink. Artificial intelligence and biological misuse: Differentiating risks of language models and biological design tools, 2023.
  76. 76.Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. Technical report: Large language models can strategically deceive their users when put under pressure. arXiv preprint arXiv:2311.07590, 2023.
  77. 77.Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, and Stephan Gunnemann. Soft prompt threats: Attacking safety alignment and unlearning in open-source llms through the embedding space, 2024.
  78. 78.Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting, 2023.
  79. 79.Toby Shevlane. Structured access: an emerging paradigm for safe ai deployment, 2022.
  80. 80.Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al. Model evaluation for extreme risks. arXiv preprint arXiv:2305.15324, 2023.
  81. 81.Blake E. Strom, Andy Applebaum, Doug P. Miller, Kathryn C. Nickels, Adam G. Pennington, and Cody B. Thomas. Mitre att&ck: Design and philosophy. Technical report, MITRE Corporation, 2020.
  82. 82.Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf. Zephyr: Direct distillation of lm alignment, 2023.
  83. 83.Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. Activation addition: Steering language models without optimization, 2023.
  84. 84.UK AI Safety Summit. The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023 — gov.uk. https://www.gov.uk/government/publications/ai-safety-summit-2023-the-bletchley-declaration/the-bletchley-declaration-by-countries-attending-the-ai-safety-summit-1-2-november-2023, 2023.
  85. 85.UK Cabinet Office. National risk register. Technical report, UK Cabinet Office, 2023.
  86. 86.Fabio Urbina, Filippa Lentzos, Cédric Invernizzi, and Sean Ekins. Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence, 4(3):189–191, 2022.
  87. 87.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  88. 88.Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. Universal adversarial triggers for attacking and analyzing nlp. arXiv preprint arXiv:1908.07125, 2019.
  89. 89.Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. arXiv preprint arXiv:2211.00593, 2022.
  90. 90.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? arXiv preprint arXiv:2307.02483, 2023.
  91. 91.The White House. Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence. https://www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/, 2023.
  92. 92.Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. Shadow alignment: The ease of subverting safely-aligned language models, 2023.
  93. 93.Dongyu Yao, Jianshu Zhang, Ian G. Harris, and Marcel Carlsson. Fuzzllm: A novel and universal fuzzing framework for proactively discovering jailbreak vulnerabilities in large language models, 2023a.
  94. 94.Yuanshun Yao, Xiaojun Xu, and Yang Liu. Large language model unlearning. arXiv preprint arXiv:2310.10683, 2023b.
  95. 95.Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu. Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher, 2023.
  96. 96.Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang. Removing rlhf protections in gpt-4 via fine-tuning, 2023.
  97. 97.Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. Siren’s song in the ai ocean: A survey on hallucination in large language models. arXiv preprint arXiv:2309.01219, 2023.
  98. 98.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. arXiv preprint arXiv:2306.05685, 2023a.
  99. 99.Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023b.
  100. 100.Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. Fine-tuning language models from human preferences, 2020.
  101. 101.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405, 2023a.
  102. 102.Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023b.

Citation

MLA
Li, N., et al. “The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning”. arXiv, 2024, http://arxiv.org/abs/2403.03218v7.
APA
Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A. B., Chen, M., Barrass, I., Zhang, O., Zhu, X., … Hendrycks, D. (2024). The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. arXiv. http://arxiv.org/abs/2403.03218v7
Chicago
Li, N., A. Pan, A. Gopal, et al. 2024. “The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning”. arXiv. http://arxiv.org/abs/2403.03218v7.
Harvard
Li, N. et al. (2024) “The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.03218v7.
Vancouver
1. Li N, Pan A, Gopal A, et al (2024) The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning. arXiv

BibTeX

@article{li2024the,
  title = {The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning},
  author = {Li, Nathaniel and Pan, Alexander and Gopal, Anjali and Yue, Summer and Berrios, Daniel and Gatti, Alice and Li, Justin D. and Dombrowski, Ann-Kathrin and Goel, Shashwat and Phan, Long and Mukobi, Gabriel and Helm-Burger, Nathan and Lababidi, Rassin and Justen, Lennart and Liu, Andrew B. and Chen, Michael and Barrass, Isabelle and Zhang, Oliver and Zhu, Xiaoyuan and Tamirisa, Rishub and Bharathi, Bhrugu and Khoja, Adam and Zhao, Zhenqi and Herbert-Voss, Ariel and Breuer, Cort B. and Marks, Samuel and Patel, Oam and Zou, Andy and Mazeika, Mantas and Wang, Zifan and Oswal, Palash and Lin, Weiran and Hunt, Adam A. and Tienken-Harder, Justin and Shih, Kevin Y. and Talley, Kemper and Guan, John and Kaplan, Russell and Steneker, Ian and Campbell, David and Jokubaitis, Brad and Levinson, Alex and Wang, Jean and Qian, William and Karmakar, Kallol Krishna and Basart, Steven and Fitz, Stephen and Levine, Mindy and Kumaraguru, Ponnurangam and Tupakula, Uday and Varadharajan, Vijay and Wang, Ruoyu and Shoshitaishvili, Yan and Ba, Jimmy and Esvelt, Kevin M. and Wang, Alexandr and Hendrycks, Dan},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.03218v7},
  eprint = {2403.03218}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/