Can LMs Learn New Entities from Descriptions? Challenges in Propagating Injected Knowledge
Yasumasa OnoeMichael J. Q. ZhangShankar PadmanabhanGreg DurrettEunsol Choi
Demonstrates that current model-editing techniques struggle to make downstream inferences about newly injected entities unless there is direct lexical overlap, highlighting a critical limitation compared to simple in-context prompting.
Pre-trained language models quickly become outdated as real-world information changes, which degrades their performance on downstream knowledge tasks. While existing model-editing techniques can update isolated facts, deployed systems require models to perform reasoning and draw valid inferences about newly introduced entities rather than merely memorizing updated statements.
The article evaluates whether current parameter-updating methods allow language models to effectively learn new entities from short definitions and propagate that knowledge to downstream inferences. To test this, the authors benchmarked several model-editing and fine-tuning approaches across three base language models using real-world Wikipedia data from the Entity Cloze By Date dataset and a new controlled evaluation set containing explicit and commonsense implicit reasoning probes.
The investigation produced four central findings. First, parameter-updating techniques struggle significantly on real-world inference tasks: on the standard Entity Cloze By Date benchmark, fine-tuning and specialized editing methods failed to improve perplexity over base models. Second, existing methods succeeded only when target answers directly overlapped verbatim with the injected definition sentences, reducing perplexity by 8.5 to 9.0 points in simpler, high-overlap subsets. Third, while full-model fine-tuning improved accuracy by 21 to 32 percentage points on the controlled inference dataset, it degraded model specificity on unrelated facts by up to 16 points. Fourth, simply prepending the entity definition directly into the model context consistently outperformed all parameter-updating methods, improving accuracy by 25 to 31 points and cutting perplexity substantially without harming specificity.
These findings indicate that current parameter-editing techniques are largely restricted to shallow factual recall rather than true knowledge propagation and reasoning. For organizational decision-makers, relying on parameter-editing methods to maintain up-to-date language models introduces operational risks of incorrect inferences and unintended model degradation on unrelated tasks. Although input augmentation via in-context prompt updates carries higher inference-time computational costs, it remains far more dependable than model weight updates for incorporating novel knowledge.
Organizations should treat in-context knowledge injection and retrieval-augmented methods as the primary production strategy for emerging entities until more robust model-editing architectures are developed. Future research should prioritize designing parameter-updating algorithms capable of integrating complex, multi-hop reasoning without sacrificing specificity. Decision-makers should exercise caution when reviewing these results, as the study focused exclusively on English-language models up to 1.5 billion parameters and examined single-run updates for newly emerging entities rather than modifications to existing entities.
- Paper: Locating and Editing Factual Associations in GPT, Kevin Meng et al. (2022). ROME establishes a foundational locate-and-edit method whose limits in generalizing beyond an inserted fact motivate the source’s evaluation of knowledge propagation.
- Paper: MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions, Zexuan Zhong et al. (2023). MQuAKE shows why single-fact edit success is not enough, providing a direct benchmark precedent for the source’s focus on downstream multi-hop inference.
- Paper: Can We Edit Factual Knowledge by In-Context Learning?, Ce Zheng et al. (2023). IKE provides the in-context knowledge-editing alternative that helps frame the source’s comparison between prompt-based updates and parameter changes.
- Paper: Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?, Zorik Gekhman et al. (2024). This study directly tests whether fine-tuning on unfamiliar facts causes hallucinations, extending the source’s concern that updates can impair unrelated knowledge.
- Paper: From RAG to Memory: Non-Parametric Continual Learning for Large Language Models, Bernal Jimnez Gutirrez et al. (2025). HippoRAG 2 develops a non-parametric continual-learning approach for factual and multi-hop retrieval, advancing the source’s case for external knowledge over fragile weight updates.
- Paper: MEMORYLLM: Towards Self-Updatable Large Language Models, Yu Wang et al. (2024). MemoryLLM explores a self-updatable memory architecture as a later alternative for incorporating new knowledge while limiting the drawbacks of conventional weight editing.
