Detoxifying Large Language Models via Knowledge Editing
Mengru WangNingyu ZhangZiwen XuZekun XiShumin DengYunzhi YaoQishen ZhangLinyi YangJindong WangHuajun Chen
Presents the SafeEdit benchmark and a single-instance knowledge editing method, DINM, to directly modify toxic model parameters rather than merely suppressing their activations, effectively detoxifying large language models without degrading their general performance.
Large Language Models often remain vulnerable to adversarial attack prompts and jailbreaks despite safety alignment. While conventional safety interventions such as Supervised Fine-Tuning and Direct Preference Optimization improve defense rates, they frequently fail against novel or out-of-domain attacks because they primarily suppress toxic activations rather than modifying the underlying parameters responsible for generating harmful content.
The article aims to evaluate whether post-training knowledge editing can directly modify toxic internal model parameters to achieve robust detoxification with minimal impact on general capabilities. To support this evaluation, the authors develop a comprehensive benchmarking framework and introduce a targeted editing baseline designed to locate and diminish toxic model regions.
To establish a rigorous evaluation, the authors constructed SafeEdit, a benchmark covering nine distinct unsafe categories (including illegal acts, privacy violations, offensiveness, and bias) combined with 48 attack templates, comprising over 8,000 total instances. Using this dataset, the authors evaluated standard alignment baselines and knowledge editing techniques on two open-source foundation models: LLaMA2-7B-Chat and Mistral-7B-v0.1. They introduced a new method named Detoxifying with Intraoperative Neural Monitoring (DINM), which identifies the specific neural network layer displaying the largest semantic disparity between safe and unsafe hidden states, then directly updates the toxic parameters using a single data instance across 10 optimization steps while constraining changes with general knowledge pairs.
The experimental findings show that DINM significantly enhances defense performance and out-of-domain robustness. On LLaMA2-7B-Chat, generalized defense success improved from 43.51% in the unedited model to 86.74% after editing, while on Mistral-7B-v0.1, it increased from 47.30% to 96.84%. Mechanistic probing revealed that traditional methods like Supervised Fine-Tuning and Direct Preference Optimization left the intrinsic toxicity of model parameters virtually unchanged (reducing toxicity by less than 1%) while relying on large shifts in intermediate activations to bypass toxic areas. In contrast, DINM reduced parameter toxicity by 2.72% without altering activation pathways, enabling single-instance edits in one category (such as offensiveness) to generalize defense rates above 70% to 95% across entirely different harmful categories.
These results demonstrate that direct parameter editing offers an efficient, low-overhead alternative to full-model retraining, requiring substantially less GPU memory and compute time while offering stronger resilience against unseen jailbreaks. However, the evaluation also revealed trade-offs: aggressively editing parameters occasionally degraded performance on general tasks such as question answering and summarization, and frequently led the model to produce repetitive phrasing when generating safe refusals.
For technical leaders and AI practitioners, targeted parameter editing represents a promising approach to supplement existing safety pipelines, particularly for rapid patching of emerging vulnerabilities. Organizations should consider piloting localized editing strategies while implementing safeguards against over-refusal and text repetition. Further development is necessary to refine toxic neuron localization at finer granularities and develop robust sequential batch-editing pipelines before deploying knowledge editing as a standalone production safeguard.
- Paper: Editing Large Language Models: Problems, Methods, and Opportunities, Yunzhi Yao et al. (2023). This survey establishes the model-editing methods and evaluation criteria that the source applies to detoxification.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Its toxicity benchmark and mitigation experiments provide essential context for the source’s detoxification goals and evaluation.
- Paper: A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity, Andrew Lee et al. (2024). Its mechanistic analysis of how DPO suppresses toxicity directly informs the source’s comparison of temporary suppression with lasting parameter changes.
- Paper: ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection, Thomas Hartvigsen et al. (2022). The source uses ToxiGen-related toxicity evaluation, so this paper explains a key benchmark for measuring toxic outputs.
- Paper: On the Exploitability of Instruction Tuning, Manli Shu et al. (2023). Its account of instruction-tuning vulnerabilities frames the safety risks that motivate methods for changing model behavior.
- Paper: PMET: Precise Model Editing in a Transformer, Xiaopeng Li et al. (2024). Building on the source’s targeted weight editing, PMET extends precise internal parameter updates to factual knowledge edits while limiting collateral effects.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). This later work extends intervention on model internals from detoxification edits to controllable selection between parametric knowledge and retrieved context.
