Editing Large Language Models: Problems, Methods, and Opportunities

Yunzhi YaoPeng WangBozhong TianSiyuan ChengZhoubo LiShumin DengHuajun ChenNingyu Zhang

article2023EMNLP498 citations

Presents a unified taxonomy and empirical evaluation of large language model editing techniques across multiple architectures and settings, providing a standardized benchmark to assess reliability, generalization, locality, and computational efficiency.

Listen

Large language models frequently require updates to correct factual errors, update obsolete information, and remove sensitive content. However, full model retraining is prohibitively expensive in terms of computational cost and time. Model editing has emerged as an efficient alternative that aims to modify specific model behaviors within a targeted scope while preserving overall performance on unrelated inputs.

The main objective of the article is to establish a standardized framework for evaluating model editing methods and to provide a comprehensive, direct comparative assessment of leading techniques across various model architectures and operational settings.

The authors conducted rigorous, controlled experiments evaluating two primary paradigms: parameter-preserving methods (such as retrieval-based memory systems and in-context learning prompts) and parameter-modifying methods (such as direct weight updates and meta-learning hypernetworks). Using standard benchmarks (ZsRE and COUNTERFACT) alongside structural baselines (including T5-XL, GPT-J, OPT-13B, and GPT-NEOX-20B), the study measured reliability, generalization to paraphrased phrasing, and locality to ensure unrelated data remains unaffected. Additionally, the authors developed a benchmark to evaluate "portability" (reasoning across related concepts), susceptibility to distracting context, side effects on downstream reasoning tasks, sequential update stability, and computational efficiency.

The study produced several key findings. First, while methods like SERAC and ROME excel on standard benchmarks (often exceeding 90% accuracy in reliability and locality for single edits), their capability degrades significantly under rigorous downstream reasoning; SERAC scored below 20% on all portability metrics, whereas in-context prompting (IKE) demonstrated superior cross-fact generalization (often exceeding 85% to 90%). Second, parameter-modifying approaches degrade sharply during sequential updates; methods like ROME experience substantial performance declines after 10 to 100 consecutive edits as parameter drift compounds. Third, scalability across model architectures is uneven: ROME and MEMIT perform well on GPT-NEOX-20B but fail completely on OPT-13B due to mathematical matrix degeneracies. Fourth, batched editing exhibits major operational trade-offs; MEMIT successfully scales to 1,000 simultaneous edits with minimal overhead, whereas retrieval and meta-learning approaches hit memory barriers at roughly 100 edits. Finally, training-heavy editors (such as SERAC and MEND) require 7 to 36 hours of pre-training and consume more than 60 gigabytes of memory, whereas direct locate-and-edit approaches execute quickly but require substantial offline statistic gathering.

These findings indicate that existing model editing techniques are not yet plug-and-play solutions for production environments that require continuous, complex factual updates. Incomplete knowledge propagation and parameter degradation pose significant operational risks for applications requiring multi-step reasoning or high-frequency maintenance. Decision-makers must weigh clear trade-offs: retrieval and in-context prompting techniques avoid parameter corruption and generalize well to entity aliases, but they increase prompt overhead and offer weak locality in standard configurations. Conversely, direct weight-editing techniques enable high-throughput batch updates but risk cumulative degradation and architecture-specific failures.

Organizations considering model editing should select approaches tailored strictly to their operational constraints: deploying prompt-based or retrieval-augmented editing when stability and reasoning portability are paramount, and utilizing batch weight editors like MEMIT only for discrete, large-scale, one-time factual updates. Before adopting these methods for production workflows, practitioners must establish validation pipelines that test for multi-hop reasoning propagation, consecutive edit drift, and collateral degradation on core reasoning capabilities.

The study's primary limitations include restricting evaluations to models up to 20 billion parameters, focusing predominantly on open-source transformer architectures rather than proprietary black-box APIs, and centering on single-edit, factual question-answering scenarios. Consequently, caution is advised when extrapolating these findings to larger closed models, non-factual attributes (such as tone, safety alignment, or multilingual processing), or complex multi-hop editing pipelines.

  • Paper: Locating and Editing Factual Associations in GPT, Kevin Meng et al. (2022). This seminal paper introduces causal mediation analysis and the ROME method for locating and editing factual associations in transformers, establishing the foundational locate-then-edit paradigm evaluated by the survey.
  • Paper: MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions, Zexuan Zhong et al. (2023). This work establishes the MQuAKE benchmark to reveal how knowledge-editing techniques fail at multi-hop reasoning, forming a critical evaluation perspective contextualized in the survey.
  • Paper: Can We Edit Factual Knowledge by In-Context Learning?, Ce Zheng et al. (2023). This paper establishes In-Context Knowledge Editing (IKE) as a primary black-box editing baseline, providing essential context for the survey's comparative analysis of parameter-modifying versus prompt-based editing.
Cover for Editing Large Language Models: Problems, Methods, and Opportunities

Abstract

Despite the ability to train capable LLMs, the methodology for maintaining their relevancy and rectifying errors remains elusive. To this end, the past few years have witnessed a surge in techniques for editing LLMs, the objective of which is to efficiently alter the behavior of LLMs within a specific domain without negatively impacting performance across other inputs. This paper embarks on a deep exploration of the problems, methods, and opportunities related to model editing for LLMs. In particular, we provide an exhaustive overview of the task definition and challenges associated with model editing, along with an in-depth empirical analysis of the most progressive methods currently at our disposal. We also build a new benchmark dataset to facilitate a more robust evaluation and pinpoint enduring issues intrinsic to existing techniques. Our objective is to provide valuable insights into the effectiveness and feasibility of each editing technique, thereby assisting the community in making informed decisions on the selection of the most appropriate method for a specific task or context¹.

Table of Contents

  • 1 Introduction
  • 2 Problems Definition
  • 3 Current Methods
  • 3.1 Methods for Preserving LLMs’ Parameters
  • 3.2 Methods for Modifying LLMs’ Parameters
  • 4 Preliminary Experiments
  • 4.1 Experiment Setting
  • 4.2 Experiment Results
  • 5 Comprehensive Study
  • 5.1 Portability - Robust Generalization
  • 5.2 Locality - Side Effect of Model Editing
  • 5.3 Efficiency
  • 6 Relationship with Relevant Works
  • 6.1 Knowledge in LLMs
  • 6.2 Lifelong Learning and Unlearning
  • 6.3 Security and Privacy for LLMs
  • 7 Conclusion
  • Acknowledgment
  • Limitations
  • Ethic Consideration
  • References
  • A Implementing Details
  • B Dataset Details
  • B.1 Basic DataSet
  • B.2 Dataset Construction for Portability Evaluation
  • B.2.1 One hop
  • B.2.2 Subject Replace
  • B.2.3 Reversed Relation
  • B.3 Dataset Construction for Locality Evaluation
  • B.3.1 Other Attribution
  • B.3.2 Distract Neighbor
  • B.3.3 Other Task

Knowls

  1. Knowl 1 — Problem Formulation and Standard Evaluation Metrics for LLM Model Editing

    definition

    Model editing aims to alter the behavior of a pre-trained base model fθ:X→Yf_\theta: \mathcal{X} \to \mathcal{Y} parameterized by θ\theta on a specific edit descriptor (xe,ye)(x_e, y_e) where the original model fails (fθ(xe)≠yef_\theta(x_e) \neq y_e), producing an edited model fθef_{\theta_e} without damaging model performance on unrelated inputs.

    The ideal edited model satisfies: fθe(x)={yeif x∈I(xe,ye)fθ(x)if x∈O(xe,ye)f_{\theta_e}(x) = \begin{cases} y_e & \text{if } x \in I(x_e, y_e) \\ f_\theta(x) & \text{if } x \in O(x_e, y_e) \end{cases} where I(xe,ye)I(x_e, y_e) is the in-scope edit space (comprising (xe,ye)(x_e, y_e) and its equivalence neighborhood N(xe,ye)N(x_e, y_e), such as rephrased descriptions) and O(xe,ye)O(x_e, y_e) is the out-of-scope space of unrelated examples.

    The three foundational properties of model editing are evaluated as follows:

    1. Reliability: The accuracy of the edited model on the exact target edit instance: E(xe′,ye′)∼{(xe,ye)}I(argmax⁡yfθe(y∣xe′)=ye′)\mathbb{E}_{(x'_e, y'_e) \sim \{(x_e, y_e)\}} \mathbb{I}\left(\operatorname{argmax}_y f_{\theta_e}(y \mid x'_e) = y'_e\right)

    2. Generalization: The accuracy of the edited model across equivalent rephrasings drawn uniformly from the equivalence neighborhood N(xe,ye)N(x_e, y_e): E(xe′,ye′)∼N(xe,ye)I(argmax⁡yfθe(y∣xe′)=ye′)\mathbb{E}_{(x'_e, y'_e) \sim N(x_e, y_e)} \mathbb{I}\left(\operatorname{argmax}_y f_{\theta_e}(y \mid x'_e) = y'_e\right)

    3. Locality (Specificity): The rate at which the post-edit model preserves the original pre-edit outputs on out-of-scope examples O(xe,ye)O(x_e, y_e): E(xe′,ye′)∼O(xe,ye)I(fθe(y∣xe′)=fθ(y∣xe′))\mathbb{E}_{(x'_e, y'_e) \sim O(x_e, y_e)} \mathbb{I}\left(f_{\theta_e}(y \mid x'_e) = f_\theta(y \mid x'_e)\right) where I(⋅)\mathbb{I}(\cdot) denotes the indicator function.

  2. Knowl 2 — Taxonomy and Architectural Comparison of LLM Editing Paradigms

    model/method

    Existing model editing approaches for large language models are classified into two major paradigms based on whether they modify or preserve base model parameters:

    1. Preserving Parameters:

      • Memory-based Methods: Store edits explicitly in memory and retrieve relevant items at inference. SERAC uses a scope classifier to route in-scope queries to an auxiliary counterfactual model while passing out-of-scope queries to the frozen base model. In-context editing methods (e.g., IKE, MemPrompt, MeLLo) prepend retrieved factual demonstrations directly into the prompt without introducing extra trainable modules.
      • Additional Parameters: Add extra trainable modules while freezing base weights. T-Patcher adds an individual neuron patch to the final Feed-Forward Network (FFN) layer per mistake; CaliNet integrates calibration memory slots into FFN layers; GRACE utilizes discrete key-value adapters.
    2. Modifying Parameters:

      • Locate-Then-Edit: First localize the weights storing target knowledge, then directly update those parameters. Knowledge Neuron (KN) locates key neurons in FFNs using attribution gradients and updates corresponding rows. ROME views the MLP as a linear key-value store and applies a closed-form rank-one update to the second linear projection layer (mlpprojmlp_{proj}) via Lagrange multipliers. MEMIT extends ROME to multi-layer, synchronous batch updates across layers.
      • Meta-learning: Employ an auxiliary hypernetwork to predict the required parameter update Δθ\Delta\theta. Knowledge Editor (KE) trains a bidirectional LSTM to predict rank-1 weight updates. Model Editor Networks with Gradient Decomposition (MEND) uses a low-rank decomposition of fine-tuning gradients to transform gradients into weight updates.
  3. Knowl 3 — Single-Edit Benchmark Results on Decoder and Encoder-Decoder LLMs

    data/table

    A performance comparison of model editing methods across T5-XL (3B, encoder-decoder) and GPT-J (6B, decoder-only) evaluated on Zero-Shot Relation Extraction (ZsRE) and COUNTERFACT benchmarks. Metrics are reported in percentages.

    Dataset Metric FT-L SERAC IKE CaliNet T-Patcher KE MEND KN ROME MEMIT
    ZsRE
    T5-XL Reliability 20.71 99.80 67.00 5.17 30.52 3.00 78.80 22.51 – –
    Generalization 19.68 99.66 67.11 4.81 30.53 5.40 89.80 22.70 – –
    Locality 89.01 98.13 63.60 72.47 77.10 96.43 98.45 16.43 – –
    GPT-J Reliability 54.70 90.16 99.96 22.72 97.12 6.60 98.15 11.34 99.18 99.23
    Generalization 49.20 89.96 99.87 0.12 94.95 7.80 97.66 9.40 94.90 87.16
    Locality 37.24 99.90 59.21 12.03 96.24 94.18 97.39 90.03 99.19 99.62
    COUNTERFACT
    T5-XL Reliability 33.57 99.89 97.77 7.76 80.26 1.00 81.40 47.86 – –
    Generalization 23.54 98.71 82.99 7.57 21.73 1.40 93.40 46.78 – –
    Locality 72.72 99.93 37.76 27.75 85.09 96.28 91.58 57.10 – –
    GPT-J Reliability 99.90 99.78 99.61 43.58 100.00 13.40 73.80 1.66 99.80 99.90
    Generalization 97.53 99.41 72.67 0.66 83.98 11.00 74.20 1.38 86.63 73.13
    Locality 1.02 98.89 35.57 2.69 8.37 94.38 93.75 58.28 93.61 97.17

    Key findings include:

    • SERAC, ROME, and MEMIT achieve strong single-edit accuracy across reliability, generalization, and locality on GPT-J.
    • KE, CaliNet, and KN perform poorly on larger models despite prior success on sub-1B architectures.
    • Constrained fine-tuning (FT-L) yields severe locality degradation (e.g., 1.02% on COUNTERFACT for GPT-J).
    • T-Patcher displays architectural sensitivity: patching the final decoder FFN of T5-XL is insufficient because the encoder continues to retain pre-edit knowledge.
  4. Knowl 4 — Scaling Breakdown and Non-Invertible Matrix Failure in Locate-Then-Edit Methods

    empirical result

    When scaling model editing methods to larger models (OPT-13B and GPT-NEOX-20B), Locate-Then-Edit methods (ROME and MEMIT) succeed on GPT-NEOX-20B but fail completely on OPT-13B.

    ZsRE COUNTERFACT
    Model Method Reliability Generalization Locality Reliability Generalization Locality
    OPT-13B ROME 22.23 6.08 99.74 36.85 2.86 95.46
    MEMIT 7.95 2.87 92.61 4.95 0.36 93.28
    IKE 69.97 69.93 64.83 49.71 34.98 53.08
    GPT-NEOX-20B ROME 99.34 95.49 99.79 99.80 85.45 94.54
    MEMIT 77.30 71.44 99.67 87.22 70.26 96.48
    IKE 100.00 99.95 59.69 98.64 67.67 43.03

    The catastrophic failure of ROME and MEMIT on OPT-13B stems from their mathematical formulation, which relies on inverting the uncentered covariance matrix of key representations collected over reference text. In OPT-13B, this matrix is degenerate (non-invertible). Attempting to approximate the solution via least squares yields unsatisfactory editing performance, exposing a critical architectural assumption in current weight-modification algorithms.

  5. Knowl 5 — Batch and Sequential Editing Behaviors in LLMs

    empirical result

    Evaluating editing methods beyond single-instance edits reveals distinct behaviors under batch editing (simultaneous multi-fact edits) and sequential editing (continuous updates without rollback):

    • Batch Editing: MEMIT scales effectively to massive knowledge editing, supporting 10 to 1,000 simultaneous edits with minimal time and memory overhead. Its reliability and generalization remain robust up to 1,000 edits, though locality declines at batch size 1,000. SERAC performs robust batch edits up to 100 cases but cannot scale to 1,000 due to memory limits. MEND and FT-L degrade rapidly as batch size increases.

    • Sequential Editing: Parameter-preserving approaches (SERAC, T-Patcher) demonstrate flat, stable performance over continuous edit streams (n=1n = 1 to n=1000n = 1000). In contrast, parameter-modifying approaches degrade severely over time due to cumulative parameter deviation from the base model: MEND degrades at n=10n = 10, ROME degrades sharply at n=100n = 100, and MEMIT's performance declines steadily over successive editing iterations.

  6. Knowl 6 — Portability Evaluation Framework for Robust Generalization

    definition

    Standard generalization metrics evaluate only superficial surface rephrasings (e.g., via back-translation). Portability measures whether an edit transfers logically to related downstream reasoning contexts.

    Portability is computed as the average accuracy of the edited model fhetaef_{ heta_e} on a reasoning dataset P(xe,ye)P(x_e, y_e): Portability=E(xe′,ye′)∼P(xe,ye)I(argmax⁡yfhetae(y∣xe′)=ye′)\text{Portability} = \mathbb{E}_{(x'_e, y'_e) \sim P(x_e, y_e)} \mathbb{I}\left(\operatorname{argmax}_y f_{ heta_e}(y \mid x'_e) = y'_e\right)

    The portability dataset is structured across three distinct reasoning categories:

    1. Subject Replace: Evaluates whether the model generalizes the edited attribute to other expressions of the subject by replacing the subject entity with an alias from Wikidata or a synonym generated by GPT-4.
    2. Reversed Relation: Evaluates bidirectional consistency by querying the inverse relation of symmetric or one-to-one facts (e.g., querying the child entity after editing the father entity).
    3. One-hop Reasoning: Evaluates whether the model can compose the edited fact (s,r,o∗)(s, r, o^*) with an existing known second-hop triple (o∗,r∗,o′∗)(o^*, r^*, o'^*) to answer questions about (s,r∘r∗,o′∗)(s, r \circ r^*, o'^*). Candidates are filtered to guarantee that the base model possessed prior knowledge of (o∗,r∗,o′∗)(o^*, r^*, o'^*) via top-10 logit link prediction.
  7. Knowl 7 — Portability Performance of Model Editing Approaches

    data/table

    Evaluation of model editing methods on GPT-J-6B and GPT-NEOX-20B across Subject Replace, Reversed Relation, and One-hop reasoning portability benchmarks (accuracies reported in percentages).

    Method Subject-Replace Reverse-Relation One-hop
    GPT-J-6B
    FT-L 72.96 8.05 1.34
    SERAC 17.79 1.30 5.53
    T-Patcher 96.65 33.62 3.10
    MEND 42.45 0.00 11.34
    ROME 37.42 46.42 50.91
    MEMIT 27.73 47.67 52.74
    IKE 88.77 92.96 55.38
    GPT-NEOX-20B
    ROME 44.57 48.99 51.03
    MEMIT 30.98 49.19 49.58
    IKE 85.54 96.46 58.97

    Current editing techniques exhibit substantial portability deficiencies:

    • SERAC scores below 20% across all portability metrics because its auxiliary scope classifier fails to recognize rephrased subjects or downstream multi-hop prompts as within the editing scope.
    • Most methods fail to propagate changes to reversed relations (e.g., MEND achieves 0.00% and FT-L achieves 8.05%), indicating that edits are strictly unidirectional.
    • In-context editing (IKE) achieves the strongest overall portability, maintaining >85% on Subject-Replace and >92% on Reverse-Relation.
  8. Knowl 8 — Multi-Level Locality Evaluation Framework for Side Effects

    experimental setup

    To comprehensively capture side effects induced by model editing, locality is evaluated across three challenging dimensions:

    1. Other Attribution: Assesses whether unedited attributes of the target entity remain preserved. Built by using the Wikidata API to extract unrelated relations (s,rother,oother)(s, r_{other}, o_{other}) for subject ss, verifying that modifying (s,rtarget,ye)(s, r_{target}, y_e) does not distort (s,rother)(s, r_{other}).
    2. Distract Neighbor: Tests whether prepending the newly edited factual statement directly before an unrelated neighborhood prompt causes the model to over-generalize and produce corrupted outputs.
    3. Other Downstream Tasks: Measures degradation on general task performance. Evaluated by sequentially performing 100 edits from COUNTERFACT on GPT-J and measuring zero-shot physical commonsense multiple-choice accuracy on Physical Interaction QA (PIQA): Accuracy=1N∑k=1NI(ak,p=ak,t)\text{Accuracy} = \frac{1}{N} \sum_{k=1}^N \mathbb{I}\left(a_{k, p} = a_{k, t}\right) where ak,pa_{k, p} is the option minimizing the post-edit model perplexity given question qkq_k and context ckc_k, and ak,ta_{k, t} is the ground-truth target option.
  9. Knowl 9 — Side-Effect and General Task Retention Results on GPT-J

    data/table

    Locality and side-effect evaluation on GPT-J across Other Attribution, Distract Neighbor, and downstream Physical Interaction QA (PIQA) performance following 100 sequential edits.

    Method Other-Attribution Distract-Neighbor Other-Task (PIQA)
    FT-L 12.88 9.48 49.56
    MEND 73.50 32.96 48.86
    SERAC 99.50 39.18 74.84
    T-Patcher 91.51 17.56 75.03
    ROME 78.94 50.35 52.12
    MEMIT 86.78 60.47 74.62
    IKE 84.13 66.04 75.33

    Key observations:

    • Most methods preserve unrelated subject attributes (Other-Attribution), led by SERAC (99.50%) and T-Patcher (91.51%), while fine-tuning (FT-L) damages other attributes (12.88%).
    • In the Distract-Neighbor setting, almost all methods experience major degradation, showing vulnerability to prompt distraction. IKE (66.04%) and MEMIT (60.47%) retain the highest robustness.
    • Parameter-modifying methods (FT-L at 49.56%, MEND at 48.86%, ROME at 52.12%) significantly degrade downstream commonsense reasoning (PIQA), whereas parameter-preserving methods (SERAC, T-Patcher, IKE) and MEMIT maintain robust baseline downstream accuracy (~74--75%).
  10. Knowl 10 — Time and GPU Memory Efficiency of Model Editing Algorithms

    data/table

    Wall-clock execution time for conducting 10 edits on GPT-J using a 2×V1002 \times \text{V100} (32GB) setup, alongside GPU VRAM consumption during editing and auxiliary training phases.

    Method Time (COUNTERFACT) Time (ZsRE) Edit VRAM (GB) Training VRAM (GB)
    FT-L 35.94s 58.86s 27.0 –
    SERAC 5.31s 6.51s 33.2 37.3
    CaliNet 1.88s 1.93s 24.9 –
    T-Patcher 1864.74s 1825.15s 34.0 –
    KE 2.20s 2.21s 42.3 45.5
    MEND 0.51s 0.52s 53.4 65.5
    KN 225.43s 173.57s 30.6 –
    ROME 147.2s 183.0s 36.3 –
    MEMIT 143.2s 145.6s 40.1 –

    Efficiency trade-offs include:

    • Meta-learning methods (MEND, KE) and SERAC provide fast test-time execution (<7<7 seconds for 10 edits) once trained, but require hours-to-days of upfront auxiliary training (e.g., >36 hours for SERAC on 3×V1003 \times \text{V100}, >7 hours for MEND) and high VRAM overhead (>65>65 GB for MEND training).
    • ROME and MEMIT require substantial upfront computation to collect key covariance statistics over tens of thousands of Wikitext samples.
    • T-Patcher is by far the slowest at test time (>1800>1800 seconds for 10 edits) because each mistake requires individual backpropagation to train its dedicated neuron patch.

Coverage note — None was omitted; all primary problem formulations, taxonomy comparisons, preliminary benchmark results, scaling analyses, portability evaluations, locality side-effect analyses, and efficiency evaluations were formalized into standalone knowls.

References

  1. 1.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  2. 2.Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman. 2023. Leace: Perfect linear concept erasure in closed form.
  3. 3.Magdalena Biesialska, Katarzyna Biesialska, and Marta Ruiz Costa-jussà. 2020. Continual lifelong learning in natural language processing: A survey. ArXiv, abs/2012.09823.
  4. 4.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. PIQA: reasoning about physical commonsense in natural language. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 7432–7439. AAAI Press.
  5. 5.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. Gpt-neox-20b: An opensource autoregressive language model.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  7. 7.Boxi Cao, Qiaoyu Tang, Hongyu Lin, Xianpei Han, Jiawei Chen, Tianshu Wang, and Le Sun. 2023. Retentive or forgetful? diving into the knowledge memorizing mechanism of language models. arXiv preprint arXiv:2305.09144.
  8. 8.Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2020. Extracting training data from large language models. In USENIX Security Symposium.
  9. 9.Yuheng Chen, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. 2023. Journey to the center of the knowledge neurons: Discoveries of language-independent knowledge neurons and degenerate knowledge neurons. CoRR, abs/2308.13198.
  10. 10.Siyuan Cheng, Ningyu Zhang, Bozhong Tian, Zelin Dai, Feiyu Xiong, Wei Guo, and Huajun Chen. 2023. Editing language model-based knowledge graph embeddings. CoRR, abs/2301.10405.
  11. 11.Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2023. Evaluating the ripple effects of knowledge editing in language models.
  12. 12.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493–8502, Dublin, Ireland. Association for Computational Linguistics.
  13. 13.Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. Editing factual knowledge in language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6491–6506, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  14. 14.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  15. 15.Qingxiu Dong, Damai Dai, Yifan Song, Jingjing Xu, Zhifang Sui, and Lei Li. 2022. Calibrating factual knowledge in pretrained language models. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 5937–5947, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  16. 16.Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau. 2023. Erasing concepts from diffusion models. CoRR, abs/2303.07345.
  17. 17.Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 30–45, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  18. 18.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  19. 19.Y. Hao, Li Dong, Furu Wei, and Ke Xu. 2021. Self-attention attribution: Interpreting information interactions inside transformer. In Proc. of AAAI.
  20. 20.Thomas Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim, and Marzyeh Ghassemi. 2022. Aging with grace: Lifelong model editing with discrete key-value adaptors. ArXiv, abs/2211.11031.
  21. 21.Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models. ArXiv, abs/2301.04213.
  22. 22.Adi Haviv, Ido Cohen, Jacob Gidron, Roei Schuster, Yoav Goldberg, and Mor Geva. 2023. Understanding transformer memorization recall through idioms. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 248–264, Dubrovnik, Croatia. Association for Computational Linguistics.
  23. 23.Evan Hernandez, Belinda Z. Li, and Jacob Andreas. 2023. Inspecting and editing knowledge representations in language models.
  24. 24.J. Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, and Fazl Barez. 2023a. Detecting edit failures in large language models: An improved specificity benchmark. In ACL Findings.
  25. 25.Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ionnis Konstas, and Fazl Barez. 2023b. Detecting edit failures in large language models: An improved specificity benchmark. In Findings of ACL. Association for Computational Linguistics.
  26. 26.Xinshuo Hu, Dongfang Li, Zihao Zheng, Zhenyu Liu, Baotian Hu, and Min Zhang. 2023. Separate the wheat from the chaff: Model deficiency unlearning via parameter-efficient module operation. CoRR, abs/2308.08090.
  27. 27.Zeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou, Wenge Rong, and Zhang Xiong. 2023. Transformer-patcher: One mistake worth one neuron. In The Eleventh International Conference on Learning Representations.
  28. 28.Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023. Editing models with task arithmetic. In The Eleventh International Conference on Learning Representations.
  29. 29.Yoichi Ishibashi and Hidetoshi Shimodaira. 2023. Knowledge sanitization of large language models. arXiv preprint arXiv:2309.11852.
  30. 30.Jacques Thibodeau. 2022. But is it really in rome? an investigation of the rome model editing technique.
  31. 31.Yiming Ju and Zheng Zhang. 2023. Klob: a benchmark for assessing knowledge locating methods in language models. arXiv preprint arXiv:2309.16535.
  32. 32.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  33. 33.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
  34. 34.Max Lamparth and Anka Reuel. 2023. Analyzing and editing inner mechanisms of backdoored language models. arXiv preprint arXiv:2302.12461.
  35. 35.Omer Levy, Minjoon Seo, Eunsol Choi, and Luke Zettlemoyer. 2017. Zero-shot relation extraction via reading comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 333–342, Vancouver, Canada. Association for Computational Linguistics.
  36. 36.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems, 33:9459–9474.
  37. 37.Xiaopeng Li, Shasha Li, Shezheng Song, Jing Yang, Jun Ma, and Jie Yu. 2023a. Pmet: Precise model editing in a transformer.
  38. 38.Zhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang, Xi Chen, and Huajun Chen. 2023b. Unveiling the pitfalls of knowledge editing for large language models. arXiv preprint arXiv:2310.02129.
  39. 39.Aman Madaan, Niket Tandon, Peter Clark, and Yiming Yang. 2022. Memory-assisted prompt editing to improve GPT-3 after deployment. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2833–2861, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  40. 40.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. Advances in Neural Information Processing Systems, 36.
  41. 41.Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau. 2023. Mass-editing memory in a transformer. In The Eleventh International Conference on Learning Representations.
  42. 42.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2022a. Fast model editing at scale. In International Conference on Learning Representations.
  43. 43.Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, and Chelsea Finn. 2022b. Memory-based model editing at scale. In International Conference on Machine Learning.
  44. 44.Shikhar Murty, Christopher D. Manning, Scott M. Lundberg, and Marco Túlio Ribeiro. 2022. Fixing model bugs with natural language patches. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 11600–11613. Association for Computational Linguistics.
  45. 45.Yasumasa Onoe, Michael J. Q. Zhang, Shankar Padmanabhan, Greg Durrett, and Eunsol Choi. 2023. Can lms learn new entities from descriptions? challenges in propagating injected knowledge. CoRR, abs/2305.01651.
  46. 46.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  47. 47.Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang. 2023. Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies. CoRR, abs/2308.03188.
  48. 48.Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2022. Reasoning with language model prompting: A survey. CoRR, abs/2212.09597.
  49. 49.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020a. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  50. 50.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020b. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  51. 51.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, abs/1910.01108.
  52. 52.Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023. In chatgpt we trust? measuring and characterizing the reliability of chatgpt.
  53. 53.Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitry Pyrkin, Sergei Popov, and Artem Babenko. 2020. Editable neural networks. In International Conference on Learning Representations.
  54. 54.Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. 2023. Safety assessment of chinese large language models. CoRR, abs/2304.10436.
  55. 55.Ayush K Tarun, Vikram S Chundawat, Murari Mandal, and Mohan S. Kankanhalli. 2021. Fast yet effective machine unlearning. IEEE transactions on neural networks and learning systems, PP.
  56. 56.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  57. 57.Ben Wang and Aran Komatsuzaki. 2021a. Gpt-j-6b: A 6 billion parameter autoregressive language model.
  58. 58.Ben Wang and Aran Komatsuzaki. 2021b. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  59. 59.Jiaan Wang, Yunlong Liang, Zengkui Sun, Yuxuan Cao, and Jiarong Xu. 2023a. Cross-lingual knowledge editing in large language models.
  60. 60.Peng Wang, Ningyu Zhang, Xin Xie, Yunzhi Yao, Bozhong Tian, Mengru Wang, Zekun Xi, Siyuan Cheng, Kangwei Liu, Guozhou Zheng, and Huajun Chen. 2023b. Easyedit: An easy-to-use knowledge editing framework for large language models. CoRR, abs/2308.07269.
  61. 61.Xiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou, Zhiyuan Liu, and Juanzi Li. 2022. Finding skill neurons in pre-trained transformer-based language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11132–11152, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  62. 62.Ga Wu, Masoud Hashemi, and Christopher Srinivasa. 2022. Puma: Performance unchanged model augmentation for training data removal. In AAAI Conference on Artificial Intelligence.
  63. 63.Suhang Wu, Minlong Peng, Yue Chen, Jinsong Su, and Mingming Sun. 2023. Eva-kellm: A new benchmark for evaluating knowledge editing of llms.
  64. 64.Yang Xu, Yutai Hou, and Wanxiang Che. 2022. Language anisotropic cross-lingual model editing. ArXiv, abs/2205.12677.
  65. 65.Yunzhi Yao, Shaohan Huang, Ningyu Zhang, Li Dong, Furu Wei, and Huajun Chen. 2022. Kformer: Knowledge injection in transformer feed-forward layers. In Natural Language Processing and Chinese Computing.
  66. 66.Yunzhi Yao, Peng Wang, Shengyu Mao, Chuanqi Tan, Fei Huang, Huajun Chen, and Ningyu Zhang. 2023. Knowledge rumination for pre-trained language models. CoRR, abs/2305.08732.
  67. 67.Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. QA-GNN: Reasoning with language models and knowledge graphs for question answering. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 535–546, Online. Association for Computational Linguistics.
  68. 68.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022a. OPT: open pre-trained transformer language models. CoRR, abs/2205.01068.
  69. 69.Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D Manning, and Jure Leskovec. 2022b. GreaseLM: Graph REASoning enhanced language models. In International Conference on Learning Representations.
  70. 70.Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019. ERNIE: enhanced language representation with informative entities. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 1441–1451. Association for Computational Linguistics.
  71. 71.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen. 2023. A survey of large language models. CoRR, abs/2303.18223.
  72. 72.Ce Zheng, Lei Li, Qingxiu Dong, Yixuan Fan, Zhiyong Wu, Jingjing Xu, and Baobao Chang. 2023. Can we edit factual knowledge by in-context learning? ArXiv, abs/2305.12740.
  73. 73.Zexuan Zhong, Zhengxuan Wu, Christopher D. Manning, Christopher Potts, and Danqi Chen. 2023. Mquake: Assessing knowledge editing in language models via multi-hop questions.
  74. 74.Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix X. Yu, and Sanjiv Kumar. 2020. Modifying memories in transformer models. ArXiv, abs/2012.00363.

Citation

MLA
Yao, Y., et al. “Editing Large Language Models: Problems, Methods, and Opportunities”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 10222–40, https://doi.org/10.18653/v1/2023.emnlp-main.632.
APA
Yao, Y., Wang, P., Tian, B., Cheng, S., Li, Z., Deng, S., Chen, H., & Zhang, N. (2023). Editing Large Language Models: Problems, Methods, and Opportunities. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10222–10240. https://doi.org/10.18653/v1/2023.emnlp-main.632
Chicago
Yao, Y., P. Wang, B. Tian, et al. 2023. “Editing Large Language Models: Problems, Methods, and Opportunities”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 10222–40. https://doi.org/10.18653/v1/2023.emnlp-main.632.
Harvard
Yao, Y. et al. (2023) “Editing Large Language Models: Problems, Methods, and Opportunities”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 10222–10240. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.632.
Vancouver
1. Yao Y, Wang P, Tian B, Cheng S, Li Z, Deng S, Chen H, Zhang N (2023) Editing Large Language Models: Problems, Methods, and Opportunities. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 10222–10240

BibTeX

@inproceedings{yao-etal-2023-editing,
    title = "Editing Large Language Models: Problems, Methods, and Opportunities",
    author = "Yao, Yunzhi  and
      Wang, Peng  and
      Tian, Bozhong  and
      Cheng, Siyuan  and
      Li, Zhoubo  and
      Deng, Shumin  and
      Chen, Huajun  and
      Zhang, Ningyu",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.632/",
    doi = "10.18653/v1/2023.emnlp-main.632",
    pages = "10222--10240"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/