MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
Zexuan ZhongZhengxuan WuChristopher D. ManningChristopher PottsDanqi Chen
Introduces the MQuAKE benchmark to evaluate whether knowledge-edited language models can propagate updates through multi-hop reasoning, revealing that existing weight-editing methods fail catastrophically and proposing a memory-based prompting alternative, MeLLo, that outperforms them by a large margin.
Large language models store massive amounts of factual information that quickly becomes outdated, but retraining these models from scratch requires prohibitive computational and financial resources. Consequently, organizations and researchers have turned to parameter-editing methods that directly alter internal model weights to inject new facts. However, current evaluations only check whether a model can recall the single edited fact, ignoring whether the model updates subsequent conclusions that logically depend on that fact. For example, if a model is updated to reflect a new national leader, it should also change its answer when asked who is married to that leader.
The article's main objective is to evaluate whether current knowledge-editing methods enable language models to handle multi-step reasoning over updated facts, and to propose a practical alternative for maintaining accurate, consistent knowledge. To do this, the authors introduce MQUAKE (Multi-hop Question Answering for Knowledge Editing), a benchmark consisting of multi-hop questions across synthetic counterfactual scenarios (MQUAKE-CF) and real-world temporal updates (MQUAKE-T) using facts sourced from Wikidata.
The evaluation tested leading parameter-editing techniques—including Fine-tuning, MEND, ROME, and MEMIT—on models such as GPT-J and Vicuna-7B. The findings reveal a critical failure mode: while weight-editing methods reliably recalled direct single-hop facts (often achieving over 90% accuracy), they failed drastically when answering multi-hop questions requiring logical propagation of those facts. For instance, after applying ROME to GPT-J, multi-hop question accuracy plummeted from an initial 43.4% down to 7.6%. Even with structured reasoning prompts, performance remained poor, and accuracy collapsed further toward near-zero levels as the number of simultaneous edits scaled into the thousands.
To overcome these limitations, the article proposes MeLLo (Memory-based Editing for Large Language Models). Instead of modifying model parameters, MeLLo leaves the base model frozen and stores edited facts in an external memory. During inference, it breaks complex questions into sub-questions, generates tentative answers, retrieves relevant edits from memory, and self-checks for contradictions before finalizing the response. When tested on GPT-3, MeLLo achieved between 41.2% and 68.7% accuracy on counterfactual multi-hop questions and over 85% on real-world updates, outperforming direct weight-editing baselines by substantial margins.
These results demonstrate that direct weight-updating techniques merely hardcode isolated facts rather than integrating them into a coherent internal belief system, creating severe risks of factual inconsistency in downstream reasoning tasks. Leaders deploying language models should avoid relying on direct parameter-editing for critical knowledge updates. Instead, organizations should prioritize external memory-based architectures like MeLLo, which eliminate retraining costs, avoid model degradation, and allow seamless addition or deletion of facts. Next steps should focus on improving dense retrieval components to maintain high accuracy when scaling external knowledge bases, and expanding evaluations across broader language models and human-authored question sets.
- Paper: Locating and Editing Factual Associations in GPT, Kevin Meng et al. (2022). Introduces foundational parameter-updating techniques for factual knowledge editing (ROME) whose limitations on ripple effects and multi-hop reasoning motivated the creation of MQuAKE.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). Formalizes the compositionality gap and multi-step factual querying in language models that MQuAKE extends to the knowledge-editing regime.
- Paper: Memory-assisted prompt editing to improve GPT-3 after deployment, Aman Madaan et al. (2022). Demonstrates external memory-assisted prompt editing to update model behavior without retraining, anticipating the memory-based prompting design in MeLLo.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). Establishes the standard paradigm of probing language models as factual knowledge bases using fill-in-the-blank queries, providing context for why editing stored facts is necessary.
- Paper: Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps, Xanh Ho et al. (2020). Provides the foundational multi-hop reasoning construction methodology using Wikipedia and Wikidata relations that informs multi-hop evaluation of edited facts.
- Paper: PMET: Precise Model Editing in a Transformer, Xiaopeng Li et al. (2024). Refines direct parameter editing by isolating attention and feed-forward hidden states to overcome the imprecision and failure modes highlighted by multi-hop benchmarks like MQuAKE.
- Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). Develops a self-updatable transformer with an integrated latent memory pool, addressing the need for scalable and continuous factual updates exposed by MQuAKE's findings.
- Paper: Can We Edit Factual Knowledge by In-Context Learning?, Ce Zheng et al. (2023). Extends the exploration of training-free, prompt-based knowledge updating by evaluating in-context learning as an alternative to parameter-modifying editors.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Employs dynamic memory and adaptive planning over external knowledge graphs to resolve multi-hop factual consistency and hallucination issues in language models.
