Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Mor GevaJasmijn BastingsKatja FilippovaAmir Globerson

article2023EMNLP588 citations

Reveals the step-by-step internal mechanism auto-regressive language models use to recall facts, demonstrating that early feed-forward layers enrich subject representations while upper attention heads directly extract the correct attributes.

Listen

Transformer-based language models capture vast amounts of factual knowledge in their internal parameters, yet how they retrieve and assemble this information during inference has remained poorly understood. As enterprises increasingly deploy language models for critical tasks, diagnosing errors and updating outdated or incorrect facts require a precise mechanistic understanding of internal information flow. The article investigates how autoregressive language models recall factual associations when answering subject-relation queries, demonstrating the exact internal pathway through which knowledge is assembled and predicted.

To analyze this retrieval mechanism, the authors applied a genetic-inspired knockout approach, systematically blocking attention edges across different layers to observe performance drops in GPT-2 and GPT-J across approximately 1,200 factual queries per model. They coupled these interventions with vocabulary projections and representation patching to trace how subject and relation representations evolve across network depths.

The investigation revealed a distinct three-stage retrieval mechanism across both architectures. First, early feed-forward sublayers actively enrich the final subject token with a broad pool of related concepts, increasing the proportion of relevant subject attributes in the representation to nearly 50%, whereas static input embeddings contain only a fraction of this information. Second, the representation of the relation propagates to the final prompt position in the early-to-intermediate layers, preparing the model to query the subject. Third, upper attention sublayers extract the specific target attribute, directly driving the correct prediction in 68% to 77% of evaluated cases. Crucially, the authors found that 30% to 39% of extraction events rely on specific attention heads that encode direct subject-attribute mappings within their own parameters, effectively functioning as knowledge hubs across the network.

These findings challenge the prevailing assumption that factual knowledge resides exclusively in intermediate feed-forward layers. Instead, factual recall relies on early feed-forward enrichment followed by extraction via upper-layer attention mechanisms. For organizational leaders and technical teams, this insight significantly alters the technical strategy for model editing, safety patching, and knowledge localization; attempting to edit facts solely by modifying feed-forward weights risks failure if the corresponding attention parameters and early enrichment stages are ignored.

Practitioners developing model-editing workflows should expand their diagnostic and editing frameworks to account for attention head parameters and multi-stage information pipelines rather than relying solely on localized feed-forward updates. Before committing to production-scale model editing pipelines, teams should run targeted pilot tests that monitor both feed-forward layers and upper attention heads to verify that modified associations propagate correctly without unintended side effects.

The findings are established with high confidence across standard autoregressive model architectures. However, decision-makers should note that vocabulary projection techniques provide an approximation of early-layer semantic content, and models exhibit a general structural bias toward their initial prompt token. Further validation is recommended when evaluating non-autoregressive or significantly larger frontier architectures.

  • Paper: Locating and Editing Factual Associations in GPT, Kevin Meng et al. (2022). Introduces causal mediation analysis to locate factual recall in mid-layer MLPs at the subject token, providing the foundational factual localization framework that this paper dissects mechanistically.
  • Paper: Transformer Feed-Forward Layers Are Key-Value Memories, Mor Geva et al. (2020). Establishes that transformer feed-forward layers act as key-value associative memories, an essential prerequisite for understanding how subject representations become enriched with attributes in early MLP sublayers.
  • Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). Develops methods for tracking attention-based information flow and residual connections through layers, directly informing the attention intervention methodology used to trace factual extraction pathways.
  • Paper: Language Models as Knowledge Bases?, Fabio Petroni et al. (2019). Frames the fundamental premise of probing autoregressive language models as parametric knowledge bases using cloze-style subject-relation queries.
  • Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Demonstrates the behavior and sensitivity of language models when queried for factual relations via prompt variations, contextualizing how query representations extract factual associations.
Cover for Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Abstract

Transformer-based language models (LMs) are known to capture factual knowledge in their parameters. While previous work looked into where factual associations are stored, only little is known about how they are retrieved internally during inference. We investigate this question through the lens of information flow. Given a subject-relation query, we study how the model aggregates information about the subject and relation to predict the correct attribute. With interventions on attention edges, we first identify two critical points where information propagates to the prediction: one from the relation positions followed by another from the subject positions. Next, by analyzing the information at these points, we unveil a three-step internal mechanism for attribute extraction. First, the representation at the last-subject position goes through an enrichment process, driven by the early MLP sublayers, to encode many subject-related attributes. Second, information from the relation propagates to the prediction. Third, the prediction representation “queries” the enriched subject to extract the attribute. Perhaps surprisingly, this extraction is typically done via attention heads, which often encode subject-attribute mappings in their parameters. Overall, our findings introduce a comprehensive view of how factual associations are stored and extracted internally in LMs, facilitating future research on knowledge localization and editing.¹

Table of Contents

  • 1 Introduction
  • 2 Background and Notation
  • 3 Experimental Setup
  • 4 Overview: Experiments & Findings
  • 5 Localizing Information Flow via Attention Knockout
  • 6 Intermediate Subject Representations
  • 6.1 Inspection of Subject Representations
  • 6.2 Attribute Rate in Token Embeddings
  • 6.3 Subject Representation Enrichment
  • 7 Attribute Extraction via Attention
  • 7.1 Attribute Extraction to the Last Position
  • 7.2 Extraction Significance
  • 7.3 Importance of Subject Enrichment
  • 7.4 'Knowledge' Attention Heads
  • 8 Related Work
  • 9 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Additional Information Flow Analysis
  • A.1 Subject-Relation Order
  • A.2 First-position Bias
  • A.3 Information Flow from Subject Positions
  • A.4 Window Size
  • B Gradient-based Analysis
  • C Attributes Rate Evaluation
  • C.1 Evaluation Details
  • C.2 Additional Sublayer Knockout Results
  • D Projection of Subject Representations
  • E Additional Results for GPT-J
  • F Analysis of MLP Outputs
  • G Subject-Attribute Mappings in Attention Heads
  • H Example Interventions on Information Flow

Knowls

  1. Knowl 1 — Three-Step Mechanism of Factual Association Recall in Auto-Regressive Transformers

    theoretical result

    In auto-regressive decoder-only transformer language models (such as GPT-2 and GPT-J), the internal process of recalling and predicting a factual attribute aa for a given subject-relation input query (s,r)(s, r) (e.g., "Beats Music is owned by [Apple]") operates via a three-step mechanism:

    1. Subject Enrichment: Across early transformer layers, multi-layer perceptron (MLP) sublayers enrich the representation at the last subject token position with numerous concepts and attributes semantically related to the subject ss.
    2. Relation Propagation: Information from the non-subject (relation) tokens propagates to the final token position in early-to-intermediate layers, constructing a relation query representation at the prediction position.
    3. Attribute Extraction: In the upper layers, the multi-head self-attention (MHSA) sublayer at the final position attends to the enriched last-subject position representation and extracts the target attribute aa via attention head parameters, directly projecting it into the residual stream of the final position to form the next-token prediction.
  2. Knowl 2 — Attention Knockout Intervention Method

    model/method

    Attention Knockout is a causal intervention technique designed to localize where and when critical information flows between token positions across transformer layers.

    In an auto-regressive transformer with LL layers and HH attention heads per layer, the pre-softmax attention matrix for head j∈[1,H]j \in [1, H] at layer ℓ\ell is computed as: Aℓ,j=softmax((Xℓ−1WQℓ,j)(Xℓ−1WKℓ,j)Td/H+Mℓ,j)A^{\ell,j} = \text{softmax}\left(\frac{(X^{\ell-1}W_Q^{\ell,j})(X^{\ell-1}W_K^{\ell,j})^T}{\sqrt{d/H}} + M^{\ell,j}\right) where Xℓ−1∈RN×dX^{\ell-1} \in \mathbb{R}^{N \times d} is the matrix of token representations at layer ℓ−1\ell-1, WQℓ,j,WKℓ,j∈Rd×(d/H)W_Q^{\ell,j}, W_K^{\ell,j} \in \mathbb{R}^{d \times (d/H)} are query and key projection matrices, and Mℓ,j∈RN×NM^{\ell,j} \in \mathbb{R}^{N \times N} is the attention mask.

    To block information from target position cc to source position rr (r≥cr \ge c) at layer ℓ\ell, the attention weights are masked for all heads j∈[1,H]j \in [1, H] over a symmetric window of kk layers centered at layer ℓ\ell: Mrcℓ′,j=−∞∀j∈[1,H],∀ℓ′∈[max⁡(1,ℓ−⌊k/2⌋),min⁡(L,ℓ+⌊k/2⌋)]M^{\ell', j}_{rc} = -\infty \quad \forall j \in [1, H], \quad \forall \ell' \in \left[\max(1, \ell - \lfloor k/2 \rfloor), \min(L, \ell + \lfloor k/2 \rfloor)\right]

    The importance of the edge (c→r)(c \to r) across those layers is measured by the relative change in the predicted token probability pintervened−pbasepbase\frac{p_{\text{intervened}} - p_{\text{base}}}{p_{\text{base}}}.

  3. Knowl 3 — Sequential Information Flow from Relation and Subject Positions

    empirical result

    Applying Attention Knockout to GPT-2 XL (L=48L = 48, intervention window k=9k = 9) and GPT-J (L=28L = 28, intervention window k=5k = 5) on the CounterFact dataset reveals two distinct, sequential stages of critical information propagation to the final input position:

    1. Early-to-Intermediate Layers (Relation Flow): Blocking attention from the last position to non-subject (relation) token positions causes a substantial decrease in the target attribute prediction probability of 35%35\% to 45%45\%.
    2. Middle-to-Upper Layers (Subject Flow): Blocking attention from the last position to subject token positions causes a sharp degradation of up to 60%60\% in prediction probability (concentrated in layers 25–40 for GPT-2 and layers 15–23 for GPT-J).
    3. Last-Subject Position Primacy: Refining the intervention to block attention to all subject positions except one reveals that critical subject information originates almost entirely from the last subject token position. Blocking the last subject position degrades prediction probability by 50%−100%50\% - 100\%, whereas preserving the edge to the last subject position leaves prediction probability largely intact.
  4. Knowl 4 — Attributes Rate Metric for Subject Representation Evaluation

    definition

    The attributes rate quantitatively measures the degree to which an intermediate token representation htℓ∈Rdh_t^\ell \in \mathbb{R}^d encodes concepts semantically related to a subject ss.

    1. Candidate Attribute Set (AsA_s): For a subject ss, 100 paragraphs are retrieved from English Wikipedia using BM25. Only paragraphs containing ss verbatim in their text or section/page title are retained (averaging 58.158.1 paragraphs per subject for GPT-2 evaluation and 54.354.3 for GPT-J). The text is tokenized, stripped of stopwords and sub-words shorter than 3 characters, yielding a set AsA_s of candidate attribute tokens (averaging 1154.41154.4 tokens for GPT-2 and 1073.11073.1 for GPT-J).
    2. Metric Definition: Let ptℓ=softmax(δ(htℓ))p_t^\ell = \text{softmax}(\delta(h_t^\ell)) be the probability distribution over vocabulary VV obtained by projecting htℓh_t^\ell using the model's vocabulary prediction head δ:Rd→R∣V∣\delta: \mathbb{R}^d \to \mathbb{R}^{|V|}. Let Tk(htℓ)T_k(h_t^\ell) denote the top-kk tokens with the highest probabilities in ptℓp_t^\ell (with k=50k=50). The attributes rate is defined as the fraction of tokens in Tk(htℓ)T_k(h_t^\ell) that belong to AsA_s: AttributesRate(htℓ,As)=∣Tk(htℓ)∩As∣k\text{AttributesRate}(h_t^\ell, A_s) = \frac{|T_k(h_t^\ell) \cap A_s|}{k}
  5. Knowl 5 — Subject Enrichment Driven by Lower-to-Middle MLP Sublayers

    empirical result

    During inference, the hidden representation at the last subject position undergoes enrichment, reaching an attributes rate of nearly 50%50\% in middle-upper layers (compared to <20%<20\% in static token embeddings EtiE_{t_i} and 4.1%−11.5%4.1\% - 11.5\% in the mean token embedding eˉ=1∣s∣∑i=1∣s∣Eti\bar{e} = \frac{1}{|s|}\sum_{i=1}^{|s|} E_{t_i}).

    Causal sublayer knockout experiments—zeroing out 10 consecutive sublayer updates (miℓ′=0m_i^{\ell'} = 0 or aiℓ′=0a_i^{\ell'} = 0 for ℓ′∈[ℓ,ℓ+9]\ell' \in [\ell, \ell+9]) at the last subject position—show:

    • Canceling updates from early MLP sublayers (e.g., layers 1–10) reduces the attributes rate of the subject representation at layer 40 in GPT-2 by approximately 88%88\%.
    • Canceling updates from early MHSA sublayers over the same span causes a much smaller reduction of <30%<30\%.

    This demonstrates that lower-to-middle MLP sublayers are the primary mechanism that injects subject-related attributes into the last-subject token representation.

  6. Knowl 6 — Attribute Extraction Events and Extraction Rate

    definition

    For an auto-regressive transformer generating a prediction for an input sequence of length NN, let t∗=arg⁡max⁡(pNL)t^* = \arg\max(p_N^L) be the token predicted by the full network at final layer LL, where pNL=softmax(δ(xNL))∈R∣V∣p_N^L = \text{softmax}(\delta(x_N^L)) \in \mathbb{R}^{|V|}.

    An attribute extraction event occurs at layer ℓ∈[1,L]\ell \in [1, L] within a multi-head self-attention (MHSA) sublayer if the top token obtained by projecting the MHSA update vector aNℓ∈Rda_N^\ell \in \mathbb{R}^d to the vocabulary matches the final predicted attribute t∗t^*: t∗=arg⁡max⁡(EaNℓ)t^* = \arg\max(E a_N^\ell) where E∈R∣V∣×dE \in \mathbb{R}^{|V| \times d} is the vocabulary embedding matrix.

    Similarly, an extraction event occurs for an MLP sublayer at layer ℓ\ell if t∗=arg⁡max⁡(EmNℓ)t^* = \arg\max(E m_N^\ell), where mNℓ∈Rdm_N^\ell \in \mathbb{R}^d is the MLP update vector at position NN.

    The attribute extraction rate is defined as the fraction of evaluation queries for which at least one extraction event occurs across layers ℓ∈[1,L]\ell \in [1, L].

  7. Knowl 7 — Attribute Extraction by Upper MHSA Sublayers and Extraction Prominence

    empirical result

    Attribute extraction for factual recall queries is primarily executed by upper MHSA sublayers rather than MLP sublayers:

    1. Extraction Frequency: In GPT-2 XL, 68.2%68.2\% of queries exhibit MHSA extraction events (averaging 2.152.15 extracting layers per query, peaking in layers 25–45). In GPT-J, MHSA extraction events occur in 76.7%76.7\% of queries (averaging 1.91.9 extracting layers per query).
    2. MLP Extraction Role: MLP sublayers at the last position trigger extraction events in only 31.3%31.3\% of queries in GPT-2 (41.0%41.0\% in GPT-J). For 17.4%17.4\% of all queries, MLP extraction is directly preceded by an MHSA extraction event, indicating MHSA is the primary extraction driver.
    3. Extraction Prominence (Non-Identity Operation): At the point where an MHSA extraction event occurs at layer ℓ\ell, the target attribute token has an average rank of 999.5999.5 in the vocabulary projection of the subject representation δ(hsℓ)\delta(h_s^\ell). The MHSA sublayer promotes the attribute from this lower rank directly to rank 1 in its output update aNℓa_N^\ell, showing that the MHSA sublayer does not merely copy the top token of the subject representation.
  8. Knowl 8 — Per-Example Attribute Extraction Rates Across Sublayers and Knockout Interventions

    data/table

    The table below details the attribute extraction rate (fraction of queries with an extraction event in at least one layer) and the mean number of extracting layers per query for GPT-2 XL (L=48L=48) and GPT-J (L=28L=28) across sublayers and positional attention restrictions from the last position:

    Module / Intervention GPT-2 XL GPT-J
    Extraction Rate (%) # Layers Extraction Rate (%) # Layers
    MHSA (no knockout) 68.2 2.15 76.7 1.90
    – all but subj. last + last 44.4 1.03 34.8 0.50
    – all non-subj. but last 42.1 1.04 44.4 0.90
    – last 39.4 0.97 46.7 0.80
    – subj. last 37.7 0.82 41.1 0.70
    – all but last 32.9 0.55 24.8 0.30
    – subj. last + last 32.8 0.70 40.3 0.70
    – non-subj. 31.5 0.71 45.6 0.80
    – subj. 30.2 0.51 27.8 0.30
    – all but subj. last 23.5 0.44 31.3 0.50
    – all but first 0.0 0.01 0.0 0.00
    MLP (no knockout) 31.3 0.38 41.0 0.60

    These measurements demonstrate that blocking attention to either subject positions (30.2%30.2\%) or non-subject positions (31.5%31.5\%) severely degrades MHSA extraction compared to the unconstrained baseline (68.2%68.2\%), whereas allowing the last position to attend only to itself and the last subject position preserves 44.4%44.4\% extraction rate.

  9. Knowl 9 — Encoding of Subject-Attribute Mappings in Knowledge Attention Heads

    empirical result

    The parameter matrices of individual attention heads directly encode subject-to-attribute factual associations.

    For head jj at layer ℓ\ell with value projection WVℓ,j∈Rd×(d/H)W_V^{\ell,j} \in \mathbb{R}^{d \times (d/H)} and output projection WOℓ,j∈R(d/H)×dW_O^{\ell,j} \in \mathbb{R}^{(d/H) \times d}, their product is WVOℓ,j=WVℓ,jWOℓ,j∈Rd×dW_{VO}^{\ell,j} = W_V^{\ell,j} W_O^{\ell,j} \in \mathbb{R}^{d \times d}. Projecting WVOℓ,jW_{VO}^{\ell,j} into vocabulary space yields the transition matrix: Gℓ,j=ETWVOℓ,jE∈R∣V∣×∣V∣G^{\ell,j} = E^T W_{VO}^{\ell,j} E \in \mathbb{R}^{|V| \times |V|} where row tt of Gℓ,jG^{\ell,j} specifies the output vocabulary token preferences for input vocabulary token tt.

    Empirical evaluation reveals:

    • In 30.2%30.2\% of MHSA extraction events in GPT-2 (39.3%39.3\% in GPT-J), the predicted attribute token aa is among the top-10 scoring tokens in the Gℓ,jG^{\ell,j} rows corresponding to the input subject's tokens.
    • In GPT-2, these factual mappings are distributed across approximately 150 attention heads, primarily in layers 24–45.
    • A small subset of 7 frequent attention heads act as "knowledge hubs" (each extracting the correct attribute for ≥10%\ge 10\% of all queries), whose WVOW_{VO} matrices encode hundreds of distinct factual subject-attribute pairs across various relations.
  10. Knowl 10 — Necessity of Subject Enrichment for Extraction Demonstrated via Representation Patching

    empirical result

    Layer patching experiments confirm that multi-layer enrichment of subject representations is required for successful attribute extraction by upper MHSA sublayers.

    When hidden representations at subject positions from early layer ℓ∈{0,1,5,10,20}\ell \in \{0, 1, 5, 10, 20\} (where ℓ=0\ell=0 represents input embeddings) are patched as input to MHSA sublayers at all subsequent layers ℓ′>ℓ\ell' > \ell:

    • Patching subject representations from layer 0 reduces the attribute extraction rate by up to 50%50\%, demonstrating that static token embeddings lack the attribute information accumulated during layer-by-layer enrichment.
    • In contrast, patching non-subject (relation) representations exhibits a steep jump from layer 0 to layer 1 (extraction rate increases from 0.050.05 at ℓ=0\ell=0 to 0.590.59 at ℓ=1\ell=1), indicating that relation representations are contextualized in the very first transformer layer.
    • Layer-wise Gradient-times-Input saliency attributions, defined as ∇xiℓfcℓ(x1ℓ,…,xNℓ)⊙xiℓ\nabla_{x_i^\ell} f_c^\ell(x_1^\ell, \dots, x_N^\ell) \odot x_i^\ell, show that subject tokens maintain high importance through roughly the first two-thirds of network depth, while relation tokens peak in the earliest layers before their relevance shifts into the final position.
  11. Knowl 11 — Limitations of Vocabulary Projection and Attention Knockout

    limitation

    The methodology has two principal limitations acknowledged by the authors:

    1. Vocabulary Space Projection Approximation: Projecting intermediate representations htℓh_t^\ell and parameter matrices WVOℓ,jW_{VO}^{\ell,j} into vocabulary space using the unembedding matrix EE or prediction head δ(⋅)\delta(\cdot) provides an approximation that is less exact in lower layers, where representations have not yet aligned with output vocabulary logits.
    2. Information Leakage in Attention Knockout: Knocking out attention edges between positions cc and rr across a window of layers ℓ…ℓ+k\ell \dots \ell+k does not prevent information from having transferred between those positions at layers prior to the window (<ℓ< \ell). While multi-layer intervention windows alleviate single-layer leakage, indirect propagation through intermediate residual pathways cannot be entirely ruled out.

Coverage note — Qualitative token listings for individual subject case studies (Mark Messier, iPod Classic, Sukarno) and individual prompt intervention heatmaps were omitted as they serve as concrete examples illustrating the general empirical results.

References

  1. 1.Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross. 2019. Gradient-Based Attribution Methods, pages 169–191. Springer International Publishing, Cham.
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  3. 3.Jasmijn Bastings and Katja Filippova. 2020. The elephant in the interpretability room: Why use attention as explanation when we have saliency methods? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 149–155, Online. Association for Computational Linguistics.
  4. 4.Roi Cohen, Mor Geva, Jonathan Berant, and Amir Globerson. 2023. Crawling the internal knowledgebase of language models. In Findings of the Association for Computational Linguistics: EACL 2023, pages 1856–1869, Dubrovnik, Croatia. Association for Computational Linguistics.
  5. 5.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493–8502, Dublin, Ireland. Association for Computational Linguistics.
  6. 6.Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2022. Analyzing transformers in embedding space. arXiv preprint arXiv:2209.02535.
  7. 7.Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. Editing factual knowledge in language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6491–6506, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  8. 8.Misha Denil, Alban Demiraj, and Nando De Freitas. 2014. Extraction of salient sentences from labelled documents. arXiv preprint arXiv:1412.6815.
  9. 9.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread. Https://transformer-circuits.pub/2021/framework/index.html.
  10. 10.Yue Feng, Ebrahim Bagheri, Faezeh Ensan, and Jelena Jovanovic. 2017. The state of the art in semantic relatedness: a framework for comparison. The Knowledge Engineering Review, 32:e10.
  11. 11.Mor Geva, Avi Caciularu, Guy Dar, Paul Roit, Shoval Sadde, Micah Shlain, Bar Tamir, and Yoav Goldberg. 2022a. LM-debugger: An interactive tool for inspection and intervention in transformer-based language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 12–21, Abu Dhabi, UAE. Association for Computational Linguistics.
  12. 12.Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022b. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 30–45, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  13. 13.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  14. 14.John F Griffiths, Anthony JF Griffiths, Susan R Wessler, Richard C Lewontin, William M Gelbart, David T Suzuki, Jeffrey H Miller, et al. 2005. An introduction to genetic analysis. Macmillan.
  15. 15.Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. 2023. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models. arXiv preprint arXiv:2301.04213.
  16. 16.Adi Haviv, Ido Cohen, Jacob Gidron, Roei Schuster, Yoav Goldberg, and Mor Geva. 2023. Understanding transformer memorization recall through idioms. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 248–264, Dubrovnik, Croatia. Association for Computational Linguistics.
  17. 17.Evan Hernandez, Belinda Z Li, and Jacob Andreas. 2023. Measuring and manipulating knowledge representations in language models. arXiv preprint arXiv:2304.00740.
  18. 18.John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4129–4138, Minneapolis, Minnesota. Association for Computational Linguistics.
  19. 19.Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423–438.
  20. 20.Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016. Visualizing and understanding neural models in NLP. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 681–691, San Diego, California. Association for Computational Linguistics.
  21. 21.Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems.
  22. 22.Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass-editing memory in a transformer.
  23. 23.Timothee Mickus, Denis Paperno, and Mathieu Constant. 2022. How to dissect a Muppet: The structure of transformer embedding spaces. Transactions of the Association for Computational Linguistics, 10:981–996.
  24. 24.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2022. Fast model editing at scale. In International Conference on Learning Representations.
  25. 25.Hosein Mohebbi, Willem Zuidema, Grzegorz Chrupała, and Afra Alishahi. 2023. Quantifying context mixing in transformers. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 3378–3400, Dubrovnik, Croatia. Association for Computational Linguistics.
  26. 26.Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt. 2023. Progress measures for grokking via mechanistic interpretability. arXiv preprint arXiv:2301.05217.
  27. 27.Nostalgebraist. 2020. interpreting GPT: the logit lens.
  28. 28.Chris Olah. 2022. Mechanistic interpretability, variables, and the importance of interpretable bases. Transformer Circuits Thread(June 27). http://www.transformer-circuits.pub/2022/mech-interp-essay/index.html.
  29. 29.Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  30. 30.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. .
  31. 31.Ori Ram, Liat Bezalel, Adi Zicher, Yonatan Belinkov, Jonathan Berant, and Amir Globerson. 2022. What are you token about? dense retrieval as distributions over the vocabulary. arXiv preprint arXiv:2212.10380.
  32. 32.Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019. Visualizing and measuring the geometry of bert. Advances in Neural Information Processing Systems, 32.
  33. 33.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  34. 34.Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995. Okapi at trec-3. Nist Special Publication Sp, 109:109.
  35. 35.Gabriele Sarti, Nils Feldhus, Ludwig Sickert, and Oskar van der Wal. 2023. Inseq: An interpretability toolkit for sequence generation models. arXiv preprint arXiv:2302.13942.
  36. 36.Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593–4601, Florence, Italy. Association for Computational Linguistics.
  37. 37.Martin J Tymms and Ismail Kola. 2008. Gene knockout protocols, volume 158. Springer Science & Business Media.
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  39. 39.Elena Voita, Rico Sennrich, and Ivan Titov. 2019. The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4396–4406, Hong Kong, China. Association for Computational Linguistics.
  40. 40.Jonas Wallat, Jaspreet Singh, and Avishek Anand. 2020. BERTnesia: Investigating the capture and forgetting of knowledge in BERT. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 174–183, Online. Association for Computational Linguistics.
  41. 41.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  42. 42.Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. arXiv preprint arXiv:2211.00593.
  43. 43.Kayo Yin and Graham Neubig. 2022. Interpreting language models with contrastive explanations. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 184–198, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  44. 44.Torsten Zesch and Iryna Gurevych. 2010. Wisdom of crowds versus wisdom of linguists–measuring the semantic relatedness of words. Natural Language Engineering, 16(1):25–59.

Citation

MLA
Geva, M., et al. “Dissecting Recall of Factual Associations in Auto-Regressive Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 12216–35, https://doi.org/10.18653/v1/2023.emnlp-main.751.
APA
Geva, M., Bastings, J., Filippova, K., & Globerson, A. (2023). Dissecting Recall of Factual Associations in Auto-Regressive Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12216–12235. https://doi.org/10.18653/v1/2023.emnlp-main.751
Chicago
Geva, M., J. Bastings, K. Filippova, and A. Globerson. 2023. “Dissecting Recall of Factual Associations in Auto-Regressive Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12216–35. https://doi.org/10.18653/v1/2023.emnlp-main.751.
Harvard
Geva, M. et al. (2023) “Dissecting Recall of Factual Associations in Auto-Regressive Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 12216–12235. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.751.
Vancouver
1. Geva M, Bastings J, Filippova K, Globerson A (2023) Dissecting Recall of Factual Associations in Auto-Regressive Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 12216–12235

BibTeX

@inproceedings{geva-etal-2023-dissecting,
    title = "Dissecting Recall of Factual Associations in Auto-Regressive Language Models",
    author = "Geva, Mor  and
      Bastings, Jasmijn  and
      Filippova, Katja  and
      Globerson, Amir",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.751/",
    doi = "10.18653/v1/2023.emnlp-main.751",
    pages = "12216--12235"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/