Detoxifying Large Language Models via Knowledge Editing

Mengru WangNingyu ZhangZiwen XuZekun XiShumin DengYunzhi YaoQishen ZhangLinyi YangJindong WangHuajun Chen

article2024ACL126 citations

Presents the SafeEdit benchmark and a single-instance knowledge editing method, DINM, to directly modify toxic model parameters rather than merely suppressing their activations, effectively detoxifying large language models without degrading their general performance.

Listen

Large Language Models often remain vulnerable to adversarial attack prompts and jailbreaks despite safety alignment. While conventional safety interventions such as Supervised Fine-Tuning and Direct Preference Optimization improve defense rates, they frequently fail against novel or out-of-domain attacks because they primarily suppress toxic activations rather than modifying the underlying parameters responsible for generating harmful content.

The article aims to evaluate whether post-training knowledge editing can directly modify toxic internal model parameters to achieve robust detoxification with minimal impact on general capabilities. To support this evaluation, the authors develop a comprehensive benchmarking framework and introduce a targeted editing baseline designed to locate and diminish toxic model regions.

To establish a rigorous evaluation, the authors constructed SafeEdit, a benchmark covering nine distinct unsafe categories (including illegal acts, privacy violations, offensiveness, and bias) combined with 48 attack templates, comprising over 8,000 total instances. Using this dataset, the authors evaluated standard alignment baselines and knowledge editing techniques on two open-source foundation models: LLaMA2-7B-Chat and Mistral-7B-v0.1. They introduced a new method named Detoxifying with Intraoperative Neural Monitoring (DINM), which identifies the specific neural network layer displaying the largest semantic disparity between safe and unsafe hidden states, then directly updates the toxic parameters using a single data instance across 10 optimization steps while constraining changes with general knowledge pairs.

The experimental findings show that DINM significantly enhances defense performance and out-of-domain robustness. On LLaMA2-7B-Chat, generalized defense success improved from 43.51% in the unedited model to 86.74% after editing, while on Mistral-7B-v0.1, it increased from 47.30% to 96.84%. Mechanistic probing revealed that traditional methods like Supervised Fine-Tuning and Direct Preference Optimization left the intrinsic toxicity of model parameters virtually unchanged (reducing toxicity by less than 1%) while relying on large shifts in intermediate activations to bypass toxic areas. In contrast, DINM reduced parameter toxicity by 2.72% without altering activation pathways, enabling single-instance edits in one category (such as offensiveness) to generalize defense rates above 70% to 95% across entirely different harmful categories.

These results demonstrate that direct parameter editing offers an efficient, low-overhead alternative to full-model retraining, requiring substantially less GPU memory and compute time while offering stronger resilience against unseen jailbreaks. However, the evaluation also revealed trade-offs: aggressively editing parameters occasionally degraded performance on general tasks such as question answering and summarization, and frequently led the model to produce repetitive phrasing when generating safe refusals.

For technical leaders and AI practitioners, targeted parameter editing represents a promising approach to supplement existing safety pipelines, particularly for rapid patching of emerging vulnerabilities. Organizations should consider piloting localized editing strategies while implementing safeguards against over-refusal and text repetition. Further development is necessary to refine toxic neuron localization at finer granularities and develop robust sequential batch-editing pipelines before deploying knowledge editing as a standalone production safeguard.

arXiv: 2403.14472zjunlp/EasyEdit
Cover for Detoxifying Large Language Models via Knowledge Editing

Abstract

This paper investigates using knowledge editing techniques to detoxify Large Language Models (LLMs). We construct a benchmark, SafeEdit, which covers nine unsafe categories with various powerful attack prompts and equips comprehensive metrics for systematic evaluation. We conduct experiments with several knowledge editing approaches, indicating that knowledge editing has the potential to detoxify LLMs with a limited impact on general performance efficiently. Then, we propose a simple yet effective baseline, dubbed Detoxifying with Intraoperative Neural Monitoring (DINM), to diminish the toxicity of LLMs within a few tuning steps via only one instance. We further provide an in-depth analysis of the internal mechanism for various detoxifying approaches, demonstrating that previous methods like SFT and DPO may merely suppress the activations of toxic parameters, while DINM mitigates the toxicity of the toxic parameters to a certain extent, making permanent adjustments. We hope that these insights could shed light on future work of developing detoxifying approaches and the underlying knowledge mechanisms of LLMs¹.

Table of Contents

  • 1 Introduction
  • 2 Benchmark Construction
  • 2.1 Task Definition
  • 2.2 Dataset
  • 2.2.1 Harmful Question
  • 2.2.2 Attack Prompt
  • 2.2.3 Response Generation
  • 2.2.4 General Knowledge
  • 2.2.5 Quality Control
  • 2.3 Evaluation Metrics
  • 2.3.1 Defense Success
  • 2.3.2 Defense Generalization
  • 2.3.3 General Performance
  • 3 The Proposed Baseline: DINM
  • 3.1 Toxic Regions Location
  • 3.2 Detoxifying Editor
  • 4 Experiment
  • 4.1 Settings
  • 4.2 Results
  • 4.3 Analysis
  • 5 Related Work
  • 5.1 Traditional Detoxifying Method
  • 5.2 Knowledge Editing
  • 6 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Dataset
  • A.1 Harmful Question
  • A.2 Attack Prompt
  • A.3 Data Samples
  • A.4 Data Split
  • A.5 Additional Test Dataset
  • A.6 The Difference Between SafeEdit and Existing Dataset
  • B Metrics
  • B.1 Defense Success
  • B.2 Details of Defense Generalization
  • C Experiment Details
  • C.1 Baselines
  • C.2 The Differences in Data Utilization Among Different Paradigmatic Methods
  • C.3 Safety Classifier C
  • C.4 General Performance for Knowledge Editing
  • C.5 DINM for LLaMA2-7B-Chat
  • C.6 DINM for Mistral-7B-v0.1
  • D More Experimental Analysis
  • D.1 Generalization Among Different Categories
  • D.2 Memory Usage Consumption
  • D.3 Different Suffix System Prompts
  • D.4 Different Layers As The Toxic Region
  • D.5 Case Study
  • E Detoxification Mechanism
  • E.1 Instance performance
  • E.2 Toxic Probe
  • E.3 Toxicity Quantification
  • E.4 The Shift of Information Flowing into Toxic Region

Knowls

  1. Knowl 1 — DINM locates toxic regions by safe–unsafe hidden-state separation

    model/method

    Detoxifying with Intraoperative Neural Monitoring (DINM) identifies a transformer layer whose representations distinguish a harmful response from its safe alternative for the same adversarial query. For a model with LL transformer layers, let hℓsafeh^{\mathrm{safe}}_\ell and hℓunsafeh^{\mathrm{unsafe}}_\ell denote the respective hidden-state representations after layer ℓ\ell. DINM selects

    ℓtoxic=arg⁡max⁡ℓ∈{1,…,L}∥hℓsafe−hℓunsafe∥2.\ell_{\mathrm{toxic}} = \arg\max_{\ell\in\{1,\ldots,L\}} \left\|h^{\mathrm{safe}}_\ell-h^{\mathrm{unsafe}}_\ell\right\|_2.

    Within the selected layer, DINM defines the toxic region as the weight matrix WℓtoxicVW^V_{\ell_{\mathrm{toxic}}} in the second linear transformation of the feed-forward network (MLP). The method therefore localizes a layer using response-level contextual semantics, rather than requiring a harmful subject token or phrase. In the authors’ test-set analysis, the selected layer was layer 32 for all 1,350 Mistral-7B-v0.1 instances; for LLaMA2-7B-Chat it was layer 29 for 1,147 instances, layer 30 for 182, and layer 32 for 21.

  2. Knowl 2 — DINM edits only the located MLP weights while constraining ordinary responses

    model/method

    Given one adversarial input XX, its safe target response YsafeY_{\mathrm{safe}}, and a normal user request qconsq_{\mathrm{cons}}, DINM updates only the matrix WℓtoxicVW^V_{\ell_{\mathrm{toxic}}} identified by its safe–unsafe hidden-state comparison; all other model parameters remain frozen. The model is trained to generate the safe response after appending the suffix system prompt SS: “If the user’s request is reasonable, give the correct answer. Please give the right response:”. Its losses are

    Le=−log⁡PWt(Ysafe∣[X;S]),Lc=KL ⁣(PWt(⋅∣[qcons;S]) ∥ PW(⋅∣[qcons;S])),L_e=-\log P_{W_t}(Y_{\mathrm{safe}}\mid[X;S]),\qquad L_c=\mathrm{KL}\!\left(P_{W_t}(\cdot\mid[q_{\mathrm{cons}};S])\,\middle\|\,P_W(\cdot\mid[q_{\mathrm{cons}};S])\right), Ltotal=ceditLe+Lc.L_{\mathrm{total}}=c_{\mathrm{edit}}L_e+L_c.

    Here WW denotes the original model parameters, WtW_t the parameters at tuning step tt, P(⋅∣⋅)P(\cdot\mid\cdot) the model’s conditional output distribution, and ceditc_{\mathrm{edit}} the weight on the detoxification loss. The constraint request qconsq_{\mathrm{cons}} is drawn from general instruction-following data. DINM takes gradient steps on LtotalL_{\mathrm{total}} with respect to the selected MLP matrix alone. The reported setup uses 10 steps, batch size 1, cedit=0.1c_{\mathrm{edit}}=0.1, zero weight decay, maximum input length 1,000, and maximum output length 600. The learning rate is 5×10−45\times10^{-4} for LLaMA2-7B-Chat and 10−510^{-5} for Mistral-7B-v0.1. DINM uses one edit instance and no separate training stage.

  3. Knowl 3 — SafeEdit combines nine harmful categories, attack prompts, and safety and utility evaluation data

    data/table

    SafeEdit is a benchmark for detoxifying language models against adversarially prompted harmful requests while measuring effects on ordinary capabilities. It covers nine categories—offensiveness, bias, physical harm, mental harm, illegal activity, ethics, privacy, pornography, and political content—and contains 60 harmful questions per category, generated with GPT-4. The authors collected 48 attack templates from public sources and hand-written prompts; 45 templates are divided into training, validation, and test groups of 15 each, while the remaining three are reserved as out-of-domain attack prompts. Combining the 60 questions per category with the 15 attack prompts for each split yields 4,050 training, 2,700 validation, and 1,350 test adversarial instances. Each editing instance includes an adversarial input, an unsafe response, a safe response, and a general-knowledge constraint request. Safe responses were generated by GPT-4 with a refusal instruction; unsafe responses were generated by text-davinci-003 and manually checked. General instruction-following examples from the Alpaca evaluation set provide constraint requests. For quality control, a RoBERTa-large safety classifier was trained on manually annotated data and reported to achieve about 97% accuracy; detected unsafe target responses were manually corrected.

  4. Knowl 4 — SafeEdit separates success on the edited query from generalization to new malicious inputs

    definition

    SafeEdit evaluates detoxification with a safety classifier CC, whose safe label is η\eta. Defense Success (DS) is the fraction of test harmful questions qq and attack prompts aa for which the edited model fW′f_{W'} produces a response classified as safe for the edited combination [q,a][q,a]:

    DS=Eq∼Q, a∼A[1{C(fW′([q,a]))=η}].\mathrm{DS}=\mathbb{E}_{q\sim Q,\,a\sim A}\left[\mathbf{1}\{C(f_{W'}([q,a]))=\eta\}\right].

    Defense Generalization (DG) measures the same safe-classification rate on variations of the edited input: harmful questions without an attack prompt (DGonlyQ_{\mathrm{onlyQ}}), the edited question paired with another attack prompt (DGotherA_{\mathrm{otherA}}), another harmful question paired with the edited attack prompt (DGotherQ_{\mathrm{otherQ}}), and a different question paired with a different attack prompt (DGotherAQ_{\mathrm{otherAQ}}). In these definitions, QQ and AA are the question and attack-prompt sets, while q′q' and a′a' denote alternatives to the edited question qq and prompt aa. The benchmark also assesses side effects: fluency by an n-gram measure on malicious-input responses, knowledge-question-answering success on TriviaQA, and summarization with ROUGE-1 on XSum.

  5. Knowl 5 — DINM improves generalized detoxification on both evaluated chat models

    data/table

    The main benchmark experiment compares vanilla models with four editing methods on LLaMA2-7B-Chat and Mistral-7B-v0.1. Detoxification columns are success rates multiplied by 100; DG-Avg is the mean of the four generalization rates. Fluency, KQA, CSum, and their reported average measure side effects. DINM attains the strongest DG-Avg for both models: 86.74 versus the vanilla 43.51 on LLaMA2-7B-Chat, and 96.84 versus 47.30 on Mistral-7B-v0.1. Its general-task performance is not uniformly preserved, particularly for Mistral, where KQA and summarization decline.

    Model Method DS DGonlyQ_{\rm onlyQ} DGotherA_{\rm otherA} DGotherQ_{\rm otherQ} DGotherAQ_{\rm otherAQ} DG-Avg Fluency KQA CSum Avg
    LLaMA2-7B-Chat Vanilla 44.44 84.30 22.00 46.59 21.15 43.51 6.66 55.15 22.29 28.03
    LLaMA2-7B-Chat FT-L 97.70 89.67 47.48 96.53 38.81 74.04 6.44 55.71 22.42 28.19
    LLaMA2-7B-Chat Ext-Sub – 85.70 43.96 59.22 46.81 58.92 4.14 55.37 23.55 27.69
    LLaMA2-7B-Chat MEND 92.88 87.05 42.92 88.99 30.93 62.47 5.80 55.27 22.39 27.82
    LLaMA2-7B-Chat DINM 96.02 95.58 77.28 96.55 77.54 86.74 5.28 53.37 20.22 26.29
    Mistral-7B-v0.1 Vanilla 41.33 50.00 47.22 43.26 48.70 47.30 5.34 51.24 16.43 24.34
    Mistral-7B-v0.1 FT-L 69.85 54.44 50.93 59.89 51.81 57.38 5.20 56.34 16.80 26.11
    Mistral-7B-v0.1 Ext-Sub – 54.22 42.11 74.33 41.81 53.12 4.29 49.72 18.41 24.14
    Mistral-7B-v0.1 MEND 88.74 70.66 56.41 80.96 56.44 66.12 4.42 54.78 17.74 25.65
    Mistral-7B-v0.1 DINM 95.41 99.19 95.00 99.56 93.59 96.84 4.58 47.53 13.01 21.71

    The dash for Ext-Sub DS reflects that Ext-Sub edits using the full training set rather than the current individual test instance, so the paper treats the per-instance DS measure as inapplicable.

  6. Knowl 6 — On a common test set, one-instance DINM matches or exceeds SFT and DPO on detoxification

    empirical result

    For a comparison with methods trained on datasets, the authors evaluated SFT, DPO, Self-Reminder, and DINM on SafeEdit_test_ALL, which has no overlap with DINM’s edit instance or the SFT/DPO training data. DINM’s results are means over three edit instances, with the reported standard deviations shown below. The detoxification average in this comparison averages DGonlyQ_{\mathrm{onlyQ}} and DGotherAQ_{\mathrm{otherAQ}}. On LLaMA2-7B-Chat, DINM scored 92.20% (standard deviation 2.33), compared with 84.20% for DPO and 81.30% for SFT. On Mistral-7B-v0.1, DINM scored 97.12% (0.35), compared with 93.70% for DPO and 87.53% for SFT. DINM uses a single edit instance without a separate training process, whereas SFT and DPO use training data. Utility results are mixed: on LLaMA2-7B-Chat, DINM’s reported general-performance average was 25.85 (0.57), close to DPO’s 25.94; on Mistral-7B-v0.1, it was 20.79 (0.51), above DPO’s 9.66 and SFT’s 11.91. The authors also report substantial variability between different DINM edit instances, so the three-instance mean does not imply that every single edit performs equally.

  7. Knowl 7 — Ablations identify toxic-layer localization and parameter tuning as important components

    empirical result

    DINM’s ablation study finds that randomly choosing a layer instead of locating the toxic layer reduces the detoxification average (the mean of DS and the four DG metrics) from 88.59% to 80.26% for LLaMA2-7B-Chat, and from 96.55% to 67.88% for Mistral-7B-v0.1. Removing parameter tuning and using only the suffix system prompt lowers that average to 64.75% and 71.64%, respectively. These results support the authors’ claim that locating and editing a specific region contributes more than relying on a reminder prompt alone. Component effects are not uniformly beneficial across every measure: for example, removing the general-knowledge constraint can slightly raise detoxification averages while worsening some general-performance scores. The paper’s separate random-layer experiment also reports better detoxification when the chosen layer is closer to Mistral’s identified toxic layer 32: editing layer 31 yields a 77.57% detoxification average, compared with 76.73% for layer 15 and 67.88% for layer 1.

  8. Knowl 8 — Mechanistic measurements suggest DINM changes toxic weights while SFT and DPO shift their inputs

    empirical result

    For Mistral-7B-v0.1, the authors use a probe trained on Jigsaw toxic-comment data to estimate toxicity associated with the selected MLP weights, and compare that estimate and the activations entering those weights before and after detoxification. The reported toxicity-reduction rates are 0.49% for SFT, 0.60% for DPO, and 2.72% for DINM; the corresponding activation-shift rates are 6.57%, 36.67%, and 0.0%. Thus, in this analysis, SFT and DPO leave the measured toxicity of the selected region nearly unchanged while shifting its input activations, whereas DINM reduces the measured toxicity without a measured activation shift. The authors interpret this as evidence that SFT and DPO may bypass toxic regions and that DINM may partially erase their toxicity, but explicitly characterize this mechanism as preliminary and hypothetical, not as a proven causal account.

  9. Knowl 9 — A single-category DINM edit generalizes across the benchmark’s unsafe categories

    empirical result

    The authors tested whether editing an instance from one unsafe category improves defense against malicious inputs from other categories. On the reported cross-category evaluation, generalization rates exceed 70% for LLaMA2-7B-Chat and 95% for Mistral-7B-v0.1. For example, an offensiveness-category edit improves defense across all nine categories for both models; the reported Mistral rates for that edit are 100.00% for offensiveness, bias, physical, mental, illegal, ethics, and privacy, 97.77% for pornography, and 98.14% for political content. The authors hypothesize that different unsafe categories can activate overlapping toxic regions; this explanation is not established as a causal mechanism.

  10. Knowl 10 — The evaluation is limited to two models, layer-level localization, and preliminary mechanism analysis

    limitation

    The experiments cover only LLaMA2-7B-Chat and Mistral-7B-v0.1, and the paper does not establish that the results transfer to other models, multilingual or multimodal settings, or a broader set of editing baselines. DINM’s localization selects a transformer layer rather than a more precise neuron-level region, and its authors describe that localization procedure as simple. The work evaluates edits made from one instance and does not test sustained batch or sequential editing under repeated attacks. The authors also report that DINM can produce repetitive responses and reduce fluency. Finally, the proposed account of how SFT, DPO, and DINM alter toxic regions is preliminary and may be limited by the probe, data, and analysis used.

Coverage note — The paper’s implementation details for the safety classifier and the full per-category cross-generalization matrix are not given separate knowls: the classifier’s construction and accuracy are included with SafeEdit, while the matrix is summarized by its main cross-category result rather than reproducing every cell.

References

  1. 1.Afra Feyza Akyürek, Eric Pan, Garry Kuwanto, and Derry Wijaya. 2023. Dune: Dataset for unified editing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 1847–1861. Association for Computational Linguistics.
  2. 2.Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. CoRR, abs/2204.05862.
  3. 3.Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023. LEACE: perfect linear concept erasure in closed form. CoRR, abs/2306.03819.
  4. 4.Bochuan Cao, Yuanpu Cao, Lu Lin, and Jinghui Chen. 2023. Defending against alignment-breaking attacks via robustly aligned LLM. CoRR, abs/2309.14348.
  5. 5.Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2023. Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons. CoRR, abs/2308.13198.
  6. 6.Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2023. Evaluating the ripple effects of knowledge editing in language models. CoRR, abs/2307.12976.
  7. 7.OpenCompass Contributors. 2023. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass.
  8. 8.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 8493–8502. Association for Computational Linguistics.
  9. 9.Jiawen Deng, Hao Sun, Zhexin Zhang, Jiale Cheng, and Minlie Huang. 2023. Recent advances towards safe, responsible, and moral dialogue systems: A survey. CoRR, abs/2302.09270.
  10. 10.Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023. Toxicity in chatgpt: Analyzing persona-assigned language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 1236–1270. Association for Computational Linguistics.
  11. 11.Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang. 2023. A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily. CoRR, abs/2311.08268.
  12. 12.Zhangyin Feng, Weitao Ma, Weijiang Yu, Lei Huang, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. Trends in integration of knowledge and large language models: A survey and taxonomy of methods, benchmarks, and applications. CoRR, abs/2311.05876.
  13. 13.Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 30–45, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  14. 14.Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng. 2024. Model editing can hurt general abilities of large language models. CoRR, abs/2401.04700.
  15. 15.Akshat Gupta, Anurag Rao, and Gopala Anumanchipalli. 2024. Model editing at scale leads to gradual and catastrophic forgetting. CoRR, abs/2401.07453.
  16. 16.Anshita Gupta, Debanjan Mondal, Akshay Krishna Sheshadri, Wenlong Zhao, Xiang Lorraine Li, Sarah Wiegreffe, and Niket Tandon. 2023. Editing commonsense knowledge in GPT. CoRR, abs/2305.14956.
  17. 17.Skyler Hallinan, Alisa Liu, Yejin Choi, and Maarten Sap. 2023. Detoxifying text with marco: Controllable revision with experts and anti-experts. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 228–242. Association for Computational Linguistics.
  18. 18.Thomas Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim, and Marzyeh Ghassemi. 2022. Aging with GRACE: lifelong model editing with discrete key-value adaptors. CoRR, abs/2211.11031.
  19. 19.Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models. CoRR, abs/2301.04213.
  20. 20.Rima Hazra, Sayan Layek, Somnath Banerjee, and Soujanya Poria. 2024. Sowing the wind, reaping the whirlwind: The impact of editing language models. CoRR, abs/2401.10647.
  21. 21.Xinshuo Hu, Dongfang Li, Zihao Zheng, Zhenyu Liu, Baotian Hu, and Min Zhang. 2023. Separate the wheat from the chaff: Model deficiency unlearning via parameter-efficient module operation. CoRR, abs/2308.08090.
  22. 22.Wenyue Hua, Jiang Guo, Mingwen Dong, Henghui Zhu, Patrick Ng, and Zhiguo Wang. 2024. Propagation and pitfalls: Reasoning-based assessment of knowledge editing through counterfactual tasks. CoRR, abs/2401.17585.
  23. 23.Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, Kaiwen Cai, Yanghao Zhang, Sihao Wu, Peipei Xu, Dengyu Wu, André Freitas, and Mustafa A. Mustafa. 2023a. A survey of safety and trustworthiness of large language models through the lens of verification and validation. CoRR, abs/2305.11391.
  24. 24.Yangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li, and Danqi Chen. 2023b. Catastrophic jailbreak of open-source llms via exploiting generation. CoRR, abs/2310.06987.
  25. 25.Youcheng Huang, Wenqiang Lei, Zheng Zhang, Jiancheng Lv, and Shuicheng Yan. 2024. See the unseen: Better context-consistent knowledge-editing by noises. CoRR, abs/2401.07544.
  26. 26.Zeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou, Wenge Rong, and Zhang Xiong. 2023c. Transformer-patcher: One mistake worth one neuron. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  27. 27.Yoichi Ishibashi and Hidetoshi Shimodaira. 2023. Knowledge sanitization of large language models. CoRR, abs/2309.11852.
  28. 28.Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023. Mistral 7b. CoRR, abs/2310.06825.
  29. 29.Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. CoRR, abs/1705.03551.
  30. 30.Abhinav Kumar, Chenhao Tan, and Amit Sharma. 2022. Probing classifiers are unreliable for concept removal and detection. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  31. 31.Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K. Kummerfeld, and Rada Mihalcea. 2024. A mechanistic understanding of alignment algorithms: A case study on DPO and toxicity. CoRR, abs/2401.01967.
  32. 32.Chak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang, and Wenjie Li. 2023. Self-detoxifying language models via toxification reversal. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 4433–4449. Association for Computational Linguistics.
  33. 33.Xiaopeng Li, Shasha Li, Bin Ji, Shezheng Song, Xi Wang, Jun Ma, Jie Yu, Xiaodong Liu, Jing Wang, and Weimin Zhang. 2024. SWEA: changing factual knowledge in large language models via subject word embedding altering. CoRR, abs/2401.17809.
  34. 34.Xiaopeng Li, Shasha Li, Shezheng Song, Jing Yang, Jun Ma, and Jie Yu. 2023a. PMET: precise model editing in a transformer. CoRR, abs/2308.08742.
  35. 35.Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023b. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval.
  36. 36.Zichao Li, Ines Arous, Siva Reddy, and Jackie Chi Kit Cheung. 2023c. Evaluating dependencies in fact editing for language models: Specificity and implication awareness. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 7623–7636. Association for Computational Linguistics.
  37. 37.Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023. Jailbreaking chatgpt via prompt engineering: An empirical study. CoRR, abs/2305.13860.
  38. 38.Michelle Lo, Shay B. Cohen, and Fazl Barez. 2024. Large language models relearn removed concepts. CoRR, abs/2401.01814.
  39. 39.Jaime R Lopez. 1996. Intraoperative neurophysiological monitoring. International anesthesiology clinics, 34(4):33–54.
  40. 40.Xinbei Ma, Tianjie Ju, Jiyang Qiu, Zhuosheng Zhang, Hai Zhao, Lifeng Liu, and Yulong Wang. 2024. Is it possible to edit large language models robustly? CoRR, abs/2402.05827.
  41. 41.Vittorio Mazzia, Alessandro Pedrani, Andrea Caciolai, Kay Rottmann, and Davide Bernardi. 2023. A survey on knowledge editing of neural networks. CoRR, abs/2310.19704.
  42. 42.Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tur. 2023. Using in-context learning to improve dialogue safety. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 11882–11910. Association for Computational Linguistics.
  43. 43.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022.
  44. 44.Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau. 2023. Mass-editing memory in a transformer. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  45. 45.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. 2022a. Fast model editing at scale. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  46. 46.Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, and Chelsea Finn. 2022b. Memory-based model editing at scale. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 15817–15831. PMLR.
  47. 47.Silen Naihin, David Atkinson, Marc Green, Merwane Hamadi, Craig Swift, Douglas Schonholtz, Adam Tauman Kalai, and David Bau. 2023. Testing language model agents safely in the wild. CoRR, abs/2311.10538.
  48. 48.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. CoRR, abs/1808.08745.
  49. 49.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  50. 50.Haowen Pan, Yixin Cao, Xiaozhi Wang, and Xun Yang. 2023. Finding and editing multi-modal neurons in pre-trained transformer. CoRR, abs/2311.07470.
  51. 51.Yuval Pinter and Michael Elhadad. 2023. Emptying the ocean with a spoon: Should we edit models? In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 15164–15172. Association for Computational Linguistics.
  52. 52.Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2023. Adding instructions during pretraining: Effective way of controlling toxicity in language models. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2023, Dubrovnik, Croatia, May 2-6, 2023, pages 2628–2643. Association for Computational Linguistics.
  53. 53.Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena D. Hwang, Ronan Le Bras, Antoine Bosselut, and Yejin Choi. 2020. Back to the future: Unsupervised backprop-based decoding for counterfactual and abductive commonsense reasoning. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 794–805. Association for Computational Linguistics.
  54. 54.Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. CoRR, abs/2305.18290.
  55. 55.Alexander Robey, Eric Wong, Hamed Hassani, and George J. Pappas. 2023. Smoothllm: Defending large language models against jailbreaking attacks. CoRR, abs/2310.03684.
  56. 56.Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023. "do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models. CoRR, abs/2308.03825.
  57. 57.Nianwen Si, Hao Zhang, and Weiqiang Zhang. 2024. MPN: leveraging multilingual patch neuron for cross-lingual model editing. CoRR, abs/2401.03190.
  58. 58.Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. 2023. Safety assessment of chinese large language models. CoRR, abs/2304.10436.
  59. 59.Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P. Xing, Furong Huang, Hao Liu, Heng Ji, Hongyi Wang, Huan Zhang, Huaxiu Yao, Manolis Kellis, Marinka Zitnik, Meng Jiang, Mohit Bansal, James Zou, Jian Pei, Jian Liu, Jianfeng Gao, Jiawei Han, Jieyu Zhao, Jiliang Tang, Jindong Wang, John Mitchell, Kai Shu, Kaidi Xu, Kai-Wei Chang, Lifang He, Lifu Huang, Michael Backes, Neil Zhenqiang Gong, Philip S. Yu, Pin-Yu Chen, Quanquan Gu, Ran Xu, Rex Ying, Shuiwang Ji, Suman Jana, Tianlong Chen, Tianming Liu, Tianyi Zhou, William Wang, Xiang Li, Xiangliang Zhang, Xiao Wang, Xing Xie, Xun Chen, Xuyu Wang, Yan Liu, Yanfang Ye, Yinzhi Cao, and Yue Zhao. 2024. Trustllm: Trustworthiness in large language models. CoRR, abs/2401.05561.
  60. 60.Zecheng Tang, Keyan Zhou, Pinzheng Wang, Yuyang Ding, Juntao Li, and Min Zhang. 2023. Detoxify language model step-by-step. CoRR, abs/2308.08295.
  61. 61.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  62. 62.Binghai Wang, Rui Zheng, Lu Chen, Yan Liu, Shihan Dou, Caishuang Huang, Wei Shen, Senjie Jin, Enyu Zhou, Chenyu Shi, Songyang Gao, Nuo Xu, Yuhao Zhou, Xiaoran Fan, Zhiheng Xi, Jun Zhao, Xiao Wang, Tao Ji, Hang Yan, Lixing Shen, Zhan Chen, Tao Gui, Qi Zhang, Xipeng Qiu, Xuanjing Huang, Zuxuan Wu, and Yu-Gang Jiang. 2024a. Secrets of RLHF in large language models part II: reward modeling. CoRR, abs/2401.06080.
  63. 63.Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023a. Decodingtrust: A comprehensive assessment of trustworthiness in GPT models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023.
  64. 64.Jiaan Wang, Yunlong Liang, Zengkui Sun, Yuxuan Cao, and Jiarong Xu. 2023b. Cross-lingual knowledge editing in large language models. CoRR, abs/2309.08952.
  65. 65.Peng Wang, Ningyu Zhang, Xin Xie, Yunzhi Yao, Bozhong Tian, Mengru Wang, Zekun Xi, Siyuan Cheng, Kangwei Liu, Guozhou Zheng, et al. 2023c. Easyedit: An easy-to-use knowledge editing framework for large language models. arXiv preprint arXiv:2308.07269.
  66. 66.Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, and Jundong Li. 2023d. Knowledge editing for large language models: A survey. CoRR, abs/2310.16218.
  67. 67.Tiannan Wang, Jiamin Chen, Qingrui Jia, Shuai Wang, Ruoyu Fang, Huilin Wang, Zhaowei Gao, Chunzhao Xie, Chuou Xu, Jihong Dai, Yibin Liu, Jialong Wu, Shengwei Ding, Long Li, Zhiwei Huang, Xinle Deng, Teng Yu, Gangan Ma, Han Xiao, Zixin Chen, Danjun Xiang, Yunxia Wang, Yuanyuan Zhu, Yi Xiao, Jing Wang, Yiru Wang, Siran Ding, Jiayang Huang, Jiayi Xu, Yilihamu Tayier, Zhenyu Hu, Yuan Gao, Chengfeng Zheng, Yueshu Ye, Yihang Li, Lei Wan, Xinyue Jiang, Yujie Wang, Siyu Cheng, Zhule Song, Xiangru Tang, Xiaohua Xu, Ningyu Zhang, Huajun Chen, Yuchen Eleanor Jiang, and Wangchunshu Zhou. 2024b. Weaver: Foundation models for creative writing. CoRR, abs/2401.17268.
  68. 68.Weixuan Wang, Barry Haddow, and Alexandra Birch. 2023e. Retrieval-augmented multilingual knowledge editing. CoRR, abs/2312.13040.
  69. 69.Yiwei Wang, Muhao Chen, Nanyun Peng, and Kai-Wei Chang. 2024c. Deepedit: Knowledge editing as decoding with constraints. CoRR, abs/2401.10471.
  70. 70.Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. 2023a. Jailbroken: How does LLM safety training fail? CoRR, abs/2307.02483.
  71. 71.Yifan Wei, Xiaoyan Yu, Huanhuan Ma, Fangyu Lei, Yixuan Weng, Ran Song, and Kang Liu. 2023b. Assessing knowledge editing in language models via relation perspective. CoRR, abs/2311.09053.
  72. 72.Zeming Wei, Yifei Wang, and Yisen Wang. 2023c. Jailbreak and guard aligned language models with only few in-context demonstrations. CoRR, abs/2310.06387.
  73. 73.Svante Wold, Kim Esbensen, and Paul Geladi. 1987. Principal component analysis. Chemometrics and intelligent laboratory systems, 2(1-3):37–52.
  74. 74.Suhang Wu, Minlong Peng, Yue Chen, Jinsong Su, and Mingming Sun. 2023a. Eva-kellm: A new benchmark for evaluating knowledge editing of llms. CoRR, abs/2308.09954.
  75. 75.Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. 2023b. DEPN: detecting and editing privacy neurons in pre-trained language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 2875–2886. Association for Computational Linguistics.
  76. 76.Yueqi Xie, Jingwei Yi, Jiawei Shao, Justin Curl, Lingjuan Lyu, Qifeng Chen, Xing Xie, and Fangzhao Wu. 2023. Defending chatgpt against jailbreak attack via self-reminders. Nat. Mac. Intell., 5(12):1486–1496.
  77. 77.Yang Xu, Yutai Hou, Wanxiang Che, and Min Zhang. 2023. Language anisotropic cross-lingual model editing. In Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, pages 5554–5569. Association for Computational Linguistics.
  78. 78.Jianhao Yan, Futing Wang, Yafu Li, and Yue Zhang. 2024. Potential and challenges of model editing for social debiasing. CoRR.
  79. 79.Jing Yao, Xiaoyuan Yi, Xiting Wang, Jindong Wang, and Xing Xie. 2023a. From instructions to intrinsic human values - A survey of alignment goals for big models. CoRR, abs/2308.12014.
  80. 80.Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Eric Sun, and Yue Zhang. 2023b. A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly. CoRR, abs/2312.02003.
  81. 81.Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang. 2023c. Editing large language models: Problems, methods, and opportunities. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 10222–10240. Association for Computational Linguistics.
  82. 82.Xiaoyuan Yi, Jing Yao, Xiting Wang, and Xing Xie. 2023. Unpacking the ethical value alignment in big models. CoRR, abs/2310.17551.
  83. 83.Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2023a. GPTFUZZER: red teaming large language models with auto-generated jailbreak prompts. CoRR, abs/2309.10253.
  84. 84.Lang Yu, Qin Chen, Jie Zhou, and Liang He. 2023b. MELO: enhancing model editing with neuron-indexed dynamic lora. CoRR, abs/2312.11795.
  85. 85.Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, Siyuan Cheng, Ziwen Xu, Xin Xu, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Lei Liang, Zhiqiang Zhang, Xiaowei Zhu, Jun Zhou, and Huajun Chen. 2024. A comprehensive study of knowledge editing for large language models. CoRR, abs/2401.01286.
  86. 86.Xu Zhang and Xiaojun Wan. 2023. Mil-decoding: Detoxifying language models at token-level via multiple instance learning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 190–202. Association for Computational Linguistics.
  87. 87.Zhexin Zhang, Jiale Cheng, Hao Sun, Jiawen Deng, and Minlie Huang. 2023a. Instructsafety: A unified framework for building multidimensional and explainable safety detector through instruction tuning. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pages 10421–10436. Association for Computational Linguistics.
  88. 88.Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. 2023b. Safetybench: Evaluating the safety of large language models with multiple choice questions. CoRR, abs/2309.07045.
  89. 89.Zhexin Zhang, Junxiao Yang, Pei Ke, and Minlie Huang. 2023c. Defending large language models against jailbreaking attacks through goal prioritization. CoRR, abs/2311.09096.
  90. 90.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.
  91. 91.Ce Zheng, Lei Li, Qingxiu Dong, Yuxuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. 2023. Can we edit factual knowledge by in-context learning? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 4862–4876. Association for Computational Linguistics.
  92. 92.Chujie Zheng, Fan Yin, Hao Zhou, Fandong Meng, Jie Zhou, Kai-Wei Chang, Minlie Huang, and Nanyun Peng. 2024. Prompt-driven LLM safeguarding via directed representation optimization. CoRR, abs/2401.18018.
  93. 93.Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, and Danqi Chen. 2023. Mquake: Assessing knowledge editing in language models via multi-hop questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, pages 15686–15702. Association for Computational Linguistics.
  94. 94.Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, Shashwat Goel, Nathaniel Li, Michael J. Byun, Zifan Wang, Alex Mallen, Steven Basart, Sanmi Koyejo, Dawn Song, Matt Fredrikson, J. Zico Kolter, and Dan Hendrycks. 2023. Representation engineering: A top-down approach to AI transparency. CoRR, abs/2310.01405.

Citation

MLA
Wang, M., et al. “Detoxifying Large Language Models via Knowledge Editing”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 3093–118, https://doi.org/10.18653/v1/2024.acl-long.171.
APA
Wang, M., Zhang, N., Xu, Z., Xi, Z., Deng, S., Yao, Y., Zhang, Q., Yang, L., Wang, J., & Chen, H. (2024). Detoxifying Large Language Models via Knowledge Editing. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3093–3118. https://doi.org/10.18653/v1/2024.acl-long.171
Chicago
Wang, M., N. Zhang, Z. Xu, et al. 2024. “Detoxifying Large Language Models via Knowledge Editing”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3093–3118. https://doi.org/10.18653/v1/2024.acl-long.171.
Harvard
Wang, M. et al. (2024) “Detoxifying Large Language Models via Knowledge Editing”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3093–3118. Available at: https://doi.org/10.18653/v1/2024.acl-long.171.
Vancouver
1. Wang M, Zhang N, Xu Z, Xi Z, Deng S, Yao Y, Zhang Q, Yang L, Wang J, Chen H (2024) Detoxifying Large Language Models via Knowledge Editing. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3093–3118

BibTeX

@inproceedings{wang-etal-2024-detoxifying,
    title = "Detoxifying Large Language Models via Knowledge Editing",
    author = "Wang, Mengru  and
      Zhang, Ningyu  and
      Xu, Ziwen  and
      Xi, Zekun  and
      Deng, Shumin  and
      Yao, Yunzhi  and
      Zhang, Qishen  and
      Yang, Linyi  and
      Wang, Jindong  and
      Chen, Huajun",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.171/",
    doi = "10.18653/v1/2024.acl-long.171",
    pages = "3093--3118"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/