MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions

Zexuan ZhongZhengxuan WuChristopher D. ManningChristopher PottsDanqi Chen

article2023EMNLP375 citations

Introduces the MQuAKE benchmark to evaluate whether knowledge-edited language models can propagate updates through multi-hop reasoning, revealing that existing weight-editing methods fail catastrophically and proposing a memory-based prompting alternative, MeLLo, that outperforms them by a large margin.

Listen

Large language models store massive amounts of factual information that quickly becomes outdated, but retraining these models from scratch requires prohibitive computational and financial resources. Consequently, organizations and researchers have turned to parameter-editing methods that directly alter internal model weights to inject new facts. However, current evaluations only check whether a model can recall the single edited fact, ignoring whether the model updates subsequent conclusions that logically depend on that fact. For example, if a model is updated to reflect a new national leader, it should also change its answer when asked who is married to that leader.

The article's main objective is to evaluate whether current knowledge-editing methods enable language models to handle multi-step reasoning over updated facts, and to propose a practical alternative for maintaining accurate, consistent knowledge. To do this, the authors introduce MQUAKE (Multi-hop Question Answering for Knowledge Editing), a benchmark consisting of multi-hop questions across synthetic counterfactual scenarios (MQUAKE-CF) and real-world temporal updates (MQUAKE-T) using facts sourced from Wikidata.

The evaluation tested leading parameter-editing techniques—including Fine-tuning, MEND, ROME, and MEMIT—on models such as GPT-J and Vicuna-7B. The findings reveal a critical failure mode: while weight-editing methods reliably recalled direct single-hop facts (often achieving over 90% accuracy), they failed drastically when answering multi-hop questions requiring logical propagation of those facts. For instance, after applying ROME to GPT-J, multi-hop question accuracy plummeted from an initial 43.4% down to 7.6%. Even with structured reasoning prompts, performance remained poor, and accuracy collapsed further toward near-zero levels as the number of simultaneous edits scaled into the thousands.

To overcome these limitations, the article proposes MeLLo (Memory-based Editing for Large Language Models). Instead of modifying model parameters, MeLLo leaves the base model frozen and stores edited facts in an external memory. During inference, it breaks complex questions into sub-questions, generates tentative answers, retrieves relevant edits from memory, and self-checks for contradictions before finalizing the response. When tested on GPT-3, MeLLo achieved between 41.2% and 68.7% accuracy on counterfactual multi-hop questions and over 85% on real-world updates, outperforming direct weight-editing baselines by substantial margins.

These results demonstrate that direct weight-updating techniques merely hardcode isolated facts rather than integrating them into a coherent internal belief system, creating severe risks of factual inconsistency in downstream reasoning tasks. Leaders deploying language models should avoid relying on direct parameter-editing for critical knowledge updates. Instead, organizations should prioritize external memory-based architectures like MeLLo, which eliminate retraining costs, avoid model degradation, and allow seamless addition or deletion of facts. Next steps should focus on improving dense retrieval components to maintain high accuracy when scaling external knowledge bases, and expanding evaluations across broader language models and human-authored question sets.

arXiv: 2305.14795
Cover for MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions

Abstract

The information stored in large language models (LLMs) falls out of date quickly, and retraining from scratch is often not an option. This has recently given rise to a range of techniques for injecting new facts through updating model weights. Current evaluation paradigms are extremely limited, mainly validating the recall of edited facts, but changing one fact should cause rippling changes to the model's related beliefs. If we edit the UK Prime Minister to now be Rishi Sunak, then we should get a different answer to Who is married to the British Prime Minister? In this work, we present a benchmark, MQuAKE (Multi-hop Question Answering for Knowledge Editing), comprising multi-hop questions that assess whether edited models correctly answer questions where the answer should change as an entailed consequence of edited facts. While we find that current knowledge-editing approaches can recall edited facts accurately, they fail catastrophically on the constructed multi-hop questions. We thus propose a simple memory-based approach, MeLLo, which stores all edited facts externally while prompting the language model iteratively to generate answers that are consistent with the edited facts. While MQuAKE remains challenging, we show that MeLLo scales well with LLMs (up to 175B) and outperforms previous model editors by a large margin.¹

Table of Contents

  • 1 Introduction
  • 2 Problem Definition
  • 2.1 Querying Factual Knowledge in LLMs
  • 2.2 Knowledge Editing
  • 2.3 Evaluation of Multi-hop Questions
  • 3 MQUAKE: Multi-hop Question Answering for Knowledge Editing
  • 3.1 Data Construction of MQUAKE-CF
  • 3.2 Data Construction of MQUAKE-T
  • 3.3 Dataset Summary
  • 3.4 Evaluation Metrics
  • 4 MQUAKE Challenges Model Editors
  • 4.1 Experimental Setup
  • 4.2 Results on MQUAKE-CF
  • 4.3 Results on MQUAKE-T
  • 4.4 Evaluation with Edits at Scale
  • 5 MeLLo: A Proposal for Editing Large Language Models
  • 5.1 Method
  • 5.2 Evaluation Results
  • 6 Related Work
  • 7 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Details of Dataset Construction
  • A.1 Sampling Fact Chains from Wikidata
  • A.2 Filtering Unrecallable Facts with GPT-J
  • A.3 Generating Questions using ChatGPT
  • B Evaluation Metrics
  • C Implementation Details for Knowledge Editing Methods
  • C.1 Fine-tuning
  • C.2 MEND
  • C.4 MEMIT
  • D Chain-of-thought Prompting for Multi-hop Questions
  • E Extended Golden Labels for MQUAKE-T
  • F Prompts used in MeLLo
  • G Breakdown Results on MQUAKE-CF
  • H Impact of Retrieval Performance
  • I Question/Cloze Statement Templates used in MQUAKE

Knowls

  1. Knowl 1 — Faithful knowledge editing requires propagating edits through multi-hop consequences

    definition

    MQUAKE evaluates whether a language model updates the consequences of an edited fact, rather than merely recalling the edited fact itself. A factual statement is represented as a triple (s,r,o)(s,r,o) consisting of a subject entity ss, a relation rr, and an object entity oo. A fact edit is e=(s,r,o→o∗)e=(s,r,o\rightarrow o^\ast), meaning that the object of the fact should change from oo to o∗o^\ast.

    A multi-hop chain is an ordered sequence C=⟨(s1,r1,o1),…,(sn,rn,on)⟩C=\langle(s_1,r_1,o_1),\ldots,(s_n,r_n,o_n)\rangle satisfying oi=si+1o_i=s_{i+1} for every i<ni<n. The chain defines a question about the head entity s1s_1 whose answer is the tail entity ono_n. After one or more edits, the chain becomes C∗C^\ast with answer a∗a^\ast instead of the original answer aa. For example, changing the United Kingdom’s head of government from Boris Johnson to Rishi Sunak should also change the answer to a question about the spouse of the British prime minister. An editing method is therefore faithful only if its edited model answers such entailed multi-hop questions with the updated answer, not merely if it assigns high probability to the directly edited object.

  2. Knowl 2 — MQUAKE benchmark construction and dataset composition

    data/table

    MQUAKE contains two datasets designed to test whether language models can use edited facts compositionally.

    MQUAKE-CF is a counterfactual diagnostic dataset built from Wikidata. The authors select 37 common relations and retain the top 20% of entities ranked by Wikipedia hyperlink counts. They sample coherent chains of 2, 3, or 4 triples, discard chains containing any single-hop fact that GPT-J cannot recall using relation-specific prompts with eight in-context demonstrations, and use ChatGPT, specifically gpt-3.5-turbo, to generate three natural-language questions per chain. For each chain, between 1 and NN facts are randomly counterfactually edited, where NN is the chain length; replacement objects are chosen from valid objects of the same relation, and edits are retained only when the resulting chain remains valid and has an answer different from the original.

    MQUAKE-T targets real temporal updates. It compares the April 2021 and April 2023 Wikidata dumps, manually retains six relations whose changes correspond to real-world updates rather than schema changes, filters chains using the same GPT-J recallability test, and uses only edits appearing in the Wikidata difference set. Each MQUAKE-T instance has exactly one edit because the selected changes mostly concern current positions such as head of government or head of state.

    Each instance has the form d=⟨E,Q,a,a∗,C,C∗⟩d=\langle E,Q,a,a^\ast,C,C^\ast\rangle, where EE is the edit set, QQ contains three generated multi-hop questions, aa and a∗a^\ast are the answers before and after editing, and CC and C∗C^\ast are the corresponding fact chains. The full dataset contains the following instances:

    Dataset / edits 2-hop 3-hop 4-hop Total
    MQUAKE-CF, 1 edit 2,454 855 446 3,755
    MQUAKE-CF, 2 edits 2,425 853 467 3,745
    MQUAKE-CF, 3 edits - 827 455 1,282
    MQUAKE-CF, 4 edits - - 436 436
    MQUAKE-CF, all 4,879 2,535 1,804 9,218
    MQUAKE-T, 1 edit 1,421 445 2 1,868

    Because of compute limits, the model-editor experiments on MQUAKE-CF use a random 3,000-instance subset containing 1,000 instances of each hop length.

  3. Knowl 3 — MeLLo performs memory-based editing with iterative contradiction checking

    model/method

    MeLLo, or Memory-based Editing for Large Language Models, leaves the base language model frozen and stores every edit in an external memory. Each edited triple is converted into a natural-language statement using a manually defined relation template. The statements are embedded with the pretrained Contriever retrieval model and stored in a retrieval index.

    To answer a multi-hop question, MeLLo repeatedly performs four operations. First, the language model decomposes the question into the next simple subquestion. Second, the frozen model generates a tentative answer using its original knowledge. Third, the subquestion retrieves the most relevant edited statement from the external index. Fourth, the language model checks whether that statement contradicts the tentative answer; if it does, the tentative answer is replaced with the answer implied by the retrieved edit, and otherwise it is retained. The corrected intermediate answer is then used to generate the next subquestion. The process continues until the model produces the final answer.

    This design does not require gradient updates, additional training, or white-box access to model weights. It also allows edits to be added or removed by changing the external memory while leaving the original model parameters and non-edited knowledge intact.

  4. Knowl 4 — Existing weight-editing methods recall edits but fail on counterfactual multi-hop questions

    empirical result

    On MQUAKE-CF, the authors edit each instance independently, with at most four edited facts, using GPT-J-6B and Vicuna-7B. Fine-tuning (FT), MEND, ROME, and MEMIT generally achieve much higher edit-wise and instance-wise scores than their multi-hop scores. The multi-hop metric is the percentage of instances for which at least one of the three generated questions is answered with the post-edit answer.

    Base model Method Edit-wise Instance-wise Multi-hop Multi-hop (CoT)
    GPT-J Base - 100.0 43.4 42.1
    FT 44.1 24.1 1.6 41.8) 1.9 40.2)
    MEND 72.8 59.6 9.2 34.2) 11.5 30.6)
    ROME 90.8 86.7 7.6 35.8) 18.1 24.0)
    MEMIT 97.4 94.0 8.1 35.3) 12.3 29.8)
    Vicuna-7B Base - 61.0 30.0 36.6
    FT 20.2 7.8 0.7 29.3) 0.2 36.4)
    MEND 65.2 47.6 7.4 22.6) 8.4 28.2)
    ROME 99.8 89.6 8.4 21.6) 12.2 24.4)
    MEMIT 96.6 84.0 7.6 22.4) 9.0 27.6)

    The strongest direct-editing methods, ROME and MEMIT, recall more than 90% of edited facts for both base models, yet GPT-J multi-hop accuracy falls from 43.4% before editing to 8.1% after MEMIT editing, and Vicuna-7B accuracy falls from 30.0% to 7.6%. Chain-of-thought prompting improves some scores but does not remove the failure. These results support the paper’s conclusion that local parameter updates can encode direct edits without integrating them into the model’s broader relational knowledge.

  5. Knowl 5 — Temporal edits and larger edit batches further expose the compositionality failure

    empirical result

    On MQUAKE-T, the authors evaluate GPT-J because Vicuna-7B was trained more recently and could have already seen the temporal facts. The real-world edits are usually recalled accurately by MEND, ROME, and MEMIT, but multi-hop performance remains substantially below the unedited model.

    Method Edit-wise Instance-wise Multi-hop Multi-hop (CoT)
    Base - 100.0 34.3 46.8
    FT 19.5 19.0 0.0 34.3) 0.2 46.6)
    MEND 99.0 98.5 16.0 18.3) 38.2 8.6)
    ROME 100.0 97.7 0.3 34.0) 11.3 35.5)
    MEMIT 100.0 98.9 0.3 34.0) 4.8 42.0)

    MEND performs comparatively well with chain-of-thought prompting on MQUAKE-T, reaching 38.2% versus the pre-edit score of 46.8%, but the other editors remain near zero to 11.3%. When edits from 100, 500, 1,000, or 3,000 instances are injected simultaneously, multi-hop performance decreases further for every evaluated weight-editing method on both MQUAKE-CF and MQUAKE-T. Thus, scaling the number of edits intensifies rather than solves the failure to use edited facts compositionally.

  6. Knowl 6 — MQUAKE evaluates edit recall, complete chain recall, and answer-level propagation separately

    equation

    Let an autoregressive language model be ff, let f∗f^\ast be its edited version, and let tr(s)t_r(s) be the manually defined prompt for relation rr applied to subject ss. Let an instance be d=⟨E,Q,a,a∗,C,C∗⟩d=\langle E,Q,a,a^\ast,C,C^\ast\rangle, where EE is a set of edits, QQ is its set of three multi-hop questions, aa and a∗a^\ast are the pre-edit and post-edit answers, and CC and C∗C^\ast are the pre-edit and post-edit chains. The indicator 1[P]\mathbf{1}[P] equals 1 when proposition PP is true and 0 otherwise.

    Edit-wise success rate measures the fraction of individual edits recalled correctly:

    EditSR=1∣E∣∑e=(s,r,o→o∗)∈E1[f∗(tr(s))=o∗].\mathrm{EditSR}=\frac{1}{|E|}\sum_{e=(s,r,o\rightarrow o^\ast)\in E}\mathbf{1}[f^\ast(t_r(s))=o^\ast].

    Instance-wise accuracy measures whether every single-hop fact in an instance is recalled. Before editing and after editing, respectively, it is

    InstAccpre=1[⋀(s,r,o)∈Cf(tr(s))=o],\mathrm{InstAcc}_{\mathrm{pre}}=\mathbf{1}\left[\bigwedge_{(s,r,o)\in C} f(t_r(s))=o\right], InstAccpost=1[⋀(s,r,o)∈C∗f∗(tr(s))=o].\mathrm{InstAcc}_{\mathrm{post}}=\mathbf{1}\left[\bigwedge_{(s,r,o)\in C^\ast} f^\ast(t_r(s))=o\right].

    Multi-hop accuracy measures whether the model answers at least one of the three generated questions correctly:

    MHAccpre=1[⋁q∈Qf(q)=a],\mathrm{MHAcc}_{\mathrm{pre}}=\mathbf{1}\left[\bigvee_{q\in Q} f(q)=a\right], MHAccpost=1[⋁q∈Qf∗(q)=a∗].\mathrm{MHAcc}_{\mathrm{post}}=\mathbf{1}\left[\bigvee_{q\in Q} f^\ast(q)=a^\ast\right].

    The reported scores are averages of these instance-level indicators over the evaluated dataset. Edit-wise and instance-wise metrics test whether facts are stored, whereas multi-hop accuracy tests whether the stored or edited facts are used consistently to produce an entailed answer.

  7. Knowl 7 — Evaluation compares four weight editors across single-instance and mass-edit settings

    experimental setup

    The primary base models are GPT-J-6B and Vicuna-7B. The compared editors are: FT, which performs gradient descent on an edit while constraining weight changes; MEND, which uses a learned hypernetwork to transform fine-tuning gradients into weight updates; ROME, which locates factual associations at a Transformer layer and updates a feed-forward network; and MEMIT, which extends ROME to encode many facts across multiple layers.

    For an edit (s,r,o→o∗)(s,r,o\rightarrow o^\ast), each editor receives the cloze prompt tr(s)t_r(s) and is trained or optimized to predict o∗o^\ast. The evaluation queries single-hop facts with manually written templates and multi-hop chains with ChatGPT-generated questions. In-context demonstrations are used to encourage a consistent answer format, and a separate chain-of-thought prompting condition tests whether explicit reasoning improves propagation.

    Two edit-load regimes are evaluated. In the instance-wise regime, one MQUAKE instance is edited at a time and can contain up to four edits. In the mass-edit regime, all edits from a randomly selected group are injected together: k∈{1,100,1000,3000}k\in\{1,100,1000,3000\} for MQUAKE-CF and k∈{1,100,500,1868}k\in\{1,100,500,1868\} for MQUAKE-T. The FT implementation updates layer 21 of GPT-J or layer 31 of Vicuna-7B with a weight-change norm coefficient of 5×10−55\times10^{-5}; ROME and MEMIT use their standard settings, with Vicuna-7B updates at layer 9 for ROME and layers {5,6,7,8,9}\{5,6,7,8,9\} for MEMIT.

  8. Knowl 8 — MeLLo substantially improves multi-hop accuracy and scales to larger language models

    data/table

    MeLLo is evaluated on batches of edited instances, and its reported values are multi-hop accuracy percentages. The comparison rows are the strongest GPT-J baselines from the weight-editing experiments: MEMIT for MQUAKE-CF and MEND for MQUAKE-T.

    MQUAKE-CF: number of edited instances MQUAKE-T: number of edited instances
    Base Method 1 100 1000 3000 1 100 500 1868
    GPT-J MEMIT 12.3 9.8 8.1 1.8 4.8 1.0 0.2 0.0
    GPT-J MEND 11.5 9.1 4.3 3.5 38.2 17.4 12.7 4.6
    GPT-J MeLLo 20.3 12.5 10.4 9.8 85.9 45.7 33.8 30.7
    Vicuna-7B MeLLo 20.3 11.9 11.0 10.2 84.4 56.3 52.6 51.3
    GPT-3 MeLLo 68.7 50.5 43.6 41.2 91.1 87.4 86.2 85.5

    With GPT-J, MeLLo outperforms MEMIT on every MQUAKE-CF batch size and MEND on every MQUAKE-T batch size. It also remains effective with Vicuna-7B, and the black-box text-davinci-003 model, denoted GPT-3, reaches 68.7% on the smallest MQUAKE-CF batch and 91.1% on the smallest MQUAKE-T batch. Although accuracy decreases as more edits populate memory, MeLLo retains a large advantage over the weight-editing baselines and can be applied to models up to 175 billion parameters without updating their weights.

  9. Knowl 9 — MeLLo performance is constrained by retrieval of all relevant edited facts

    empirical result

    For a multi-hop question associated with one to four edited facts, MeLLo must retrieve every relevant edit from its external memory. Using GPT-3 as the base model on MQUAKE-CF, retrieval accuracy is defined as the percentage of instances for which all associated edited facts are retrieved correctly.

    Number of edited instances 1 100 1000 3000
    Retrieval accuracy 93.6 67.7 59.4 58.7
    MeLLo multi-hop accuracy 68.7 50.5 43.6 41.2

    When all associated edited facts are successfully retrieved, MeLLo answers 73.1% of the corresponding questions correctly. Increasing the number of unrelated edits makes retrieval harder and lowers both retrieval accuracy and overall multi-hop performance, showing that the external memory improves faithfulness but introduces a dependence on retrieval quality.

  10. Knowl 10 — Stated limitations concern model coverage, prompt dependence, and question authorship

    limitation

    The evidence for existing weight-editing methods is concentrated on GPT-J-6B and Vicuna-7B because these methods are computationally expensive; their effectiveness on other language models is not established by the study. MeLLo is demonstrated on models larger than 6 billion parameters, and its reliance on language-model question decomposition and self-checking leaves its behavior on smaller models untested.

    MeLLo also depends on manually designed prompts and relation-specific templates, so its transfer to tasks beyond MQUAKE is not demonstrated. Finally, MQUAKE’s multi-hop questions are generated automatically by ChatGPT rather than written by human annotators. Although MQUAKE-T uses real temporal fact changes, human-authored questions could provide a more realistic test of knowledge-editing systems.

Coverage note — The complete relation-template inventory, qualitative question examples, detailed editor covariance statistics, and auxiliary prompt demonstrations were omitted because they are implementation-level details rather than additional load-bearing contributions.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems (NeurIPS).
  2. 2.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality.
  3. 3.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022a. Knowledge neurons in pretrained transformers. In Association for Computational Linguistics (ACL).
  4. 4.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022b. Knowledge neurons in pretrained transformers. In Association for Computational Linguistics (ACL).
  5. 5.Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. Editing factual knowledge in language models. In Empirical Methods in Natural Language Processing (EMNLP).
  6. 6.Qingxiu Dong, Damai Dai, Yifan Song, Jingjing Xu, Zhifang Sui, and Lei Li. 2022. Calibrating factual knowledge in pretrained language models. In Findings of Empirical Methods in Natural Language Processing (EMNLP).
  7. 7.Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. 2023. Do language models have beliefs? Methods for detecting, updating, and visualizing model beliefs. In European Chapter of the Association for Computational Linguistics (EACL).
  8. 8.Evan Hernandez, Belinda Z Li, and Jacob Andreas. 2023. Measuring and manipulating knowledge representations in language models. arXiv preprint arXiv:2304.00740.
  9. 9.Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, and Fazl Barez. 2023. Detecting edit failures in large language models: An improved specificity benchmark. arXiv preprint arXiv:2305.17553.
  10. 10.Zeyu Huang, Yikang Shen, Xiaofeng Zhang, Jie Zhou, Wenge Rong, and Zhang Xiong. 2023. Transformer-patcher: One mistake worth one neuron. In International Conference on Learning Representations (ICLR).
  11. 11.Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bojanowski, Armand Joulin, and Edouard Grave. 2021. Towards unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research (TMLR).
  12. 12.Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? In Transactions of the Association of Computational Linguistics (TACL).
  13. 13.Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP. arXiv preprint arXiv:2212.14024.
  14. 14.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2023. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. In Association for Computational Linguistics (ACL).
  15. 15.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems (NeurIPS).
  16. 16.Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass-editing memory in a transformer. In International Conference on Learning Representations (ICLR).
  17. 17.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2022a. Fast model editing at scale. In International Conference on Learning Representations (ICLR).
  18. 18.Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn. 2022b. Memory-based model editing at scale. In International Conference on Machine Learning (ICML).
  19. 19.Eric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong, Ananth Agarwal, Patrick Liu, Chelsea Finn, and Christopher D. Manning. 2022c. Enhancing self-consistency and performance of pretrained language models with NLI. In Empirical Methods in Natural Language Processing (EMNLP).
  20. 20.Yasumasa Onoe, Michael JQ Zhang, Shankar Padmanabhan, Greg Durrett, and Eunsol Choi. 2023. Can lms learn new entities from descriptions? Challenges in propagating injected knowledge. In Association for Computational Linguistics (ACL).
  21. 21.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems (NeurIPS).
  22. 22.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, An-Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? In Empirical Methods in Natural Language Processing (EMNLP).
  23. 23.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350.
  24. 24.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  25. 25.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235.
  26. 26.Anton Sinitsin, Vsevolod Plokhotnyuk, Dmitry Pyrkin, Sergei Popov, and Artem Babenko. 2020. Editable neural networks. In International Conference on Machine Learning (ICML).
  27. 27.Matthew Sotoudeh and Aditya V Thakur. 2019. Correcting deep neural networks with small, generalizing patches. In Workshop on Safety and Robustness in Decision Making.
  28. 28.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Timothée Martin, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  29. 29.Denny Vrandeciˇ c and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10):78–85.
  30. 30.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Lan￾guage Model. https://github.com/kingoflolz/mesh-transformer-jax.
  31. 31.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS).
  32. 32.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR).
  33. 33.Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. Factual probing is [MASK]: Learning vs. learning to recall. In North American Chapter of the Association for Computational Linguistics (NAACL).
  34. 34.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2023a. Least-to-most prompting enables complex reasoning in large language models. In International Conference on Learning Representations (ICLR).
  35. 35.Wenxuan Zhou, Sheng Zhang, Hoifung Poon, and Muhao Chen. 2023b. Context-faithful prompting for large language models. arXiv preprint arXiv:2303.11315.
  36. 36.Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar. 2021. Modifying memories in transformer models. In International Conference on Machine Learning (ICML).

Citation

MLA
Zhong, Z., et al. “MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 15686–702, https://doi.org/10.18653/v1/2023.emnlp-main.971.
APA
Zhong, Z., Wu, Z., Manning, C. D., Potts, C., & Chen, D. (2023). MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 15686–15702. https://doi.org/10.18653/v1/2023.emnlp-main.971
Chicago
Zhong, Z., Z. Wu, C. D. Manning, C. Potts, and D. Chen. 2023. “MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 15686–702. https://doi.org/10.18653/v1/2023.emnlp-main.971.
Harvard
Zhong, Z. et al. (2023) “MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 15686–15702. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.971.
Vancouver
1. Zhong Z, Wu Z, Manning CD, Potts C, Chen D (2023) MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 15686–15702

BibTeX

@inproceedings{zhong-etal-2023-mquake,
    title = "{MQ}u{AKE}: Assessing Knowledge Editing in Language Models via Multi-Hop Questions",
    author = "Zhong, Zexuan  and
      Wu, Zhengxuan  and
      Manning, Christopher  and
      Potts, Christopher  and
      Chen, Danqi",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.971/",
    doi = "10.18653/v1/2023.emnlp-main.971",
    pages = "15686--15702"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/