Dissecting Recall of Factual Associations in Auto-Regressive Language Models
Mor GevaJasmijn BastingsKatja FilippovaAmir Globerson
Reveals the step-by-step internal mechanism auto-regressive language models use to recall facts, demonstrating that early feed-forward layers enrich subject representations while upper attention heads directly extract the correct attributes.
Transformer-based language models capture vast amounts of factual knowledge in their internal parameters, yet how they retrieve and assemble this information during inference has remained poorly understood. As enterprises increasingly deploy language models for critical tasks, diagnosing errors and updating outdated or incorrect facts require a precise mechanistic understanding of internal information flow. The article investigates how autoregressive language models recall factual associations when answering subject-relation queries, demonstrating the exact internal pathway through which knowledge is assembled and predicted.
To analyze this retrieval mechanism, the authors applied a genetic-inspired knockout approach, systematically blocking attention edges across different layers to observe performance drops in GPT-2 and GPT-J across approximately 1,200 factual queries per model. They coupled these interventions with vocabulary projections and representation patching to trace how subject and relation representations evolve across network depths.
The investigation revealed a distinct three-stage retrieval mechanism across both architectures. First, early feed-forward sublayers actively enrich the final subject token with a broad pool of related concepts, increasing the proportion of relevant subject attributes in the representation to nearly 50%, whereas static input embeddings contain only a fraction of this information. Second, the representation of the relation propagates to the final prompt position in the early-to-intermediate layers, preparing the model to query the subject. Third, upper attention sublayers extract the specific target attribute, directly driving the correct prediction in 68% to 77% of evaluated cases. Crucially, the authors found that 30% to 39% of extraction events rely on specific attention heads that encode direct subject-attribute mappings within their own parameters, effectively functioning as knowledge hubs across the network.
These findings challenge the prevailing assumption that factual knowledge resides exclusively in intermediate feed-forward layers. Instead, factual recall relies on early feed-forward enrichment followed by extraction via upper-layer attention mechanisms. For organizational leaders and technical teams, this insight significantly alters the technical strategy for model editing, safety patching, and knowledge localization; attempting to edit facts solely by modifying feed-forward weights risks failure if the corresponding attention parameters and early enrichment stages are ignored.
Practitioners developing model-editing workflows should expand their diagnostic and editing frameworks to account for attention head parameters and multi-stage information pipelines rather than relying solely on localized feed-forward updates. Before committing to production-scale model editing pipelines, teams should run targeted pilot tests that monitor both feed-forward layers and upper attention heads to verify that modified associations propagate correctly without unintended side effects.
The findings are established with high confidence across standard autoregressive model architectures. However, decision-makers should note that vocabulary projection techniques provide an approximation of early-layer semantic content, and models exhibit a general structural bias toward their initial prompt token. Further validation is recommended when evaluating non-autoregressive or significantly larger frontier architectures.
- Paper: Locating and Editing Factual Associations in GPT, Kevin Meng et al. (2022). Introduces causal mediation analysis to locate factual recall in mid-layer MLPs at the subject token, providing the foundational factual localization framework that this paper dissects mechanistically.
- Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). Establishes that transformer feed-forward layers act as key-value associative memories, an essential prerequisite for understanding how subject representations become enriched with attributes in early MLP sublayers.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). Develops methods for tracking attention-based information flow and residual connections through layers, directly informing the attention intervention methodology used to trace factual extraction pathways.
- Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). Frames the fundamental premise of probing autoregressive language models as parametric knowledge bases using cloze-style subject-relation queries.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Demonstrates the behavior and sensitivity of language models when queried for factual relations via prompt variations, contextualizing how query representations extract factual associations.
- Paper: Do Large Language Models Latently Perform Multi-Hop Reasoning?, Sohee Yang et al. (2024). Extends single-hop factual retrieval mechanisms by using causal interventions to investigate whether transformers internally execute multi-hop reasoning by chaining latent entity recall.
- Paper: RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations, Jing Huang et al. (2024). Builds upon causal intervention and attribute extraction findings by benchmarking interpretability methods on their capacity to disentangle and isolate specific entity attributes within internal representations.
- Paper: Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering, Yu Zhao 0043 et al. (2025). Applies mechanistic understanding of internal factual representations to steer model reliance between internal parametric memory and external retrieved context during knowledge conflicts.
- Paper: Inference-Time Intervention: Eliciting Truthful Answers from a Language Model, Kenneth Li et al. (2023). Leverages the localization of factual truthfulness in specific attention heads to perform lightweight inference-time steering of model outputs toward accurate factual generation.
