Can We Edit Factual Knowledge by In-Context Learning?
Ce ZhengLei LiQingxiu DongYuxuan FanZhiyong WuJingjing XuBaobao Chang
Proposes an in-context knowledge editing framework that updates facts in large language models using structured demonstration prompts, achieving competitive editing performance without costly parameter updates or catastrophic forgetting.
Large language models store immense amounts of factual knowledge in their parameters, but this information frequently becomes outdated, incorrect, or biased. Correcting these errors traditionally requires parameter-tuning or gradient-based methods, which identify and adjust internal model weights. As models scale up and are increasingly delivered as closed, black-box cloud services, these traditional editing approaches become computationally expensive, hard to scale, and risky due to parameter damage that alters unrelated behaviors.
The article evaluates whether factual knowledge can be effectively updated without changing model weights by using in-context learning, a method called In-Context Knowledge Editing (IKE). The primary objective is to demonstrate that structured demonstration prompts can balance generalization—applying the new fact to paraphrased questions—with specificity, which ensures unrelated facts remain unchanged.
The authors tested this approach across multiple autoregressive language models spanning 1.5 billion to 175 billion parameters, including GPT-J, GPT-NeoX, and OPT-175B, as well as instruction-tuned models like LLaMA and Vicuna. The core evaluation used the COUNTERFACT benchmark consisting of over 21,000 records designed to test difficult counterfactual edits and evaluate side effects. The method constructs prompts containing three distinct demonstration types: copy examples to introduce the target fact, update examples to generalize across different phrasings, and retain examples to preserve unrelated knowledge retrieved via sentence similarity.
The investigation produced four central findings. First, IKE achieved competitive editing efficacy without altering any model parameters, matching or closely trailing complex parameter-modifying methods like ROME and outperforming baseline hyper-networks by roughly 10% in editing success on GPT-J. Second, the method dramatically reduced side effects: it prevented catastrophic forgetting of historical facts, retaining an 88% memorization ratio on sequential time-aware edits compared to less than 0.1% for ROME. Third, contrastive evaluations showed that IKE substantially lowered over-editing on similar, unrelated relations. Finally, the approach scaled positively with model size, delivering its highest accuracy, generalization (98.8%), and specificity (85.1%) on the largest 175-billion-parameter model.
These results indicate that organizations deploying large language models can correct facts with lower computational expense, reduced deployment risk, and greater operational transparency. Because the underlying model remains untouched, edits are fully reversible and easily calibrated through natural language. Furthermore, this method bypasses the access restrictions of black-box model-as-a-service APIs where parameter modification is impossible.
For practical implementation, technical teams should consider pairing this in-context approach with external retrieval memory systems, allowing systems to dynamically fetch relevant edit demonstrations per query rather than overloading input windows. Further engineering work and pilot testing are recommended to refine demonstration retrieval and validate performance across diverse query formats before full production rollout.
The primary limitations include increased per-query inference costs caused by longer input context lengths and the current restriction of the evaluation to factual knowledge rather than broader commonsense reasoning. Nevertheless, the evidence provides high confidence that demonstration-based in-context editing offers a viable, scalable alternative to direct parameter manipulation for large language models.
- Paper: Locating and Editing Factual Associations in GPT, Kevin Meng et al. (2022). This seminal paper introduces Rank-One Model Editing (ROME) and establishes the standard paradigms, datasets (zsRE, CounterFact), and evaluation metrics for factual knowledge editing that the source paper directly builds upon and compares against.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). This foundational work formalizes the premise that pretrained language models store massive factual associations in their parameters, motivating subsequent efforts to evaluate and modify those stored facts.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). This study analyzes how demonstration components drive in-context learning, providing essential theoretical and empirical background for using demonstration context prompts to steer model predictions.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). This paper establishes similarity-based demonstration retrieval for in-context learning, which directly informs how in-context knowledge editing retrieves relevant demonstration examples.
- Paper: Memory-assisted prompt editing to improve GPT-3 after deployment, Aman Madaan et al. (2022). This work demonstrates memory-assisted prompt editing to correct deployed black-box language models dynamically via external memory rather than retraining, prefiguring tuning-free knowledge editing.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This empirical study investigates the boundaries between parametric memory and external non-parametric retrieval in language models, framing why parameter-free editing methods are needed.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This paper introduces retrieval-augmented generation to ground parametric language models with retrieved knowledge, establishing key concepts for injecting non-parametric factual context.
- Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). This paper extends parameter-efficient knowledge updating by introducing a self-updatable latent memory pool that allows language models to continually integrate text knowledge without full retraining.
- Paper: MeMo: Memory as a Model, Ryan Wei Heng Quek et al. (2026). This work continues the exploration of parameter-free factual updating for frozen executive models by offloading ongoing knowledge updates into a small, modular parametric memory model.
- Paper: Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models, Qizheng Zhang et al. (2026). This work develops an agentic framework to evolve and update context playbooks dynamically over time, advancing beyond static in-context demonstration engineering.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). This paper investigates an alternative lightweight inference-time approach to eliciting truthful factual knowledge by steering attention head activations rather than prompting demonstrations.
- Paper: Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning, Shuai Zhao et al. (2024). This paper explores the security implications and vulnerabilities of steering language model behavior through in-context learning demonstration prompts.
