Do Large Language Models Latently Perform Multi-Hop Reasoning?

Sohee YangElena GribovskayaNora KassnerMor GevaSebastian Riedel

article2024ACL175 citations

Investigates whether large language models internally connect factual knowledge across multi-hop prompts, revealing that while models frequently recall intermediate bridge entities, their ability to utilize that recalled information for the final reasoning step remains context-dependent and fails to scale with model size.

Listen

Large Language Models often complete complex prompts that require combining multiple facts, such as answering "The mother of the singer of 'Superstition' is" without explicit intermediate steps. However, it remains unclear whether models solve these queries by latently traversing intermediate reasoning steps—first recalling the intermediate "bridge" entity Stevie Wonder and then retrieving his mother—or merely by relying on memorized text patterns. Understanding how internal multi-step reasoning operates is essential for evaluating whether model scaling improves real reasoning capabilities and determining whether targeted updates to base facts will reliably propagate across dependent knowledge.

The article evaluates whether and how frequently large language models internally execute latent two-hop reasoning during inference. It specifically examines the degree to which models internally recall intermediate entities and subsequently utilize knowledge about those entities to determine final prompt completions.

To assess these mechanisms, the researchers introduced a benchmark dataset consisting of 45,595 two-hop prompts covering 52 distinct composition types derived from factual relational data. They evaluated three open foundation models ranging from 7 billion to 70 billion parameters (LLaMA-2 7B, 13B, and 70B). The analysis measured two core components across individual model layers: an internal entity recall score that tracks hidden activation projections to intermediate entities, and a consistency score that measures how closely the output matches a direct single-hop query about the intermediate entity. The authors applied intervention and causal patching techniques to determine whether enhancing intermediate entity recall directly improved final output consistency.

The investigation produced four central findings. First, models demonstrate strong internal recognition of the first reasoning step: in approximately 70% to 78% of test cases, models successfully identified the intermediate bridge entity, with recall improving markedly as model size scaled from 7B to 70B parameters. Second, the execution of the second reasoning step was substantially weaker, with increased intermediate recall improving output consistency in roughly 60% to 65% of cases. Third, unlike the first step, this second step showed no positive scaling trend, remaining flat across the 7B, 13B, and 70B models. Finally, complete end-to-end multi-hop reasoning occurred in about 38% to 46% of cases overall, although performance was highly contextual, exceeding 80% success rates in up to 23% of specific relation categories.

These findings suggest that while modern language models can successfully retrieve intermediate entities internally, they frequently struggle to route that retrieved information into the next step of inference. Simply increasing model size improves initial factual recognition but fails to resolve the bottleneck in chaining knowledge together. This explains why standard parameter scaling alone has not resolved compositional reasoning failures. Furthermore, this structural limitation indicates substantial risk for knowledge editing and factual maintenance strategies, as updating a fundamental fact in a model will generally fail to automatically update multi-step dependent outputs.

Organizations developing or deploying large language models should not rely solely on parameter scaling to achieve reliable implicit multi-step reasoning. Instead, technical roadmaps should focus on architectures with explicit multi-step reasoning mechanisms, improved pretraining data structures, and targeted objective functions that encourage internal knowledge routing. Furthermore, for mission-critical applications requiring multi-step factual integrity, leaders should implement explicit prompt-based reasoning workflows (such as chain-of-thought methods) or external retrieval rather than expecting implicit parameter-based inference to remain consistent.

These conclusions are bounded by specific methodological conditions. The analysis focused on two-hop factual associations within the LLaMA-2 model family and tracked single-layer latent pathways, which may represent a conservative lower bound of more distributed internal reasoning. Despite these boundaries, the large sample size and consistent empirical trends across model scales provide high confidence that current standard architectures face a persistent bottleneck in executing multi-step latent reasoning.

Cover for Do Large Language Models Latently Perform Multi-Hop Reasoning?

Abstract

We study whether Large Language Models (LLMs) latently perform multi-hop reasoning with complex prompts such as “The mother of the singer of ‘Superstition’ is”. We look for evidence of a latent reasoning pathway where an LLM (1) latently identifies “the singer of ‘Superstition’” as Stevie Wonder, the bridge entity, and (2) uses its knowledge of Stevie Wonder’s mother to complete the prompt. We analyze these two hops individually and consider their co-occurrence as indicative of latent multi-hop reasoning. For the first hop, we test if changing the prompt to indirectly mention the bridge entity instead of any other entity increases the LLM’s internal recall of the bridge entity. For the second hop, we test if increasing this recall causes the LLM to better utilize what it knows about the bridge entity. We find strong evidence of latent multi-hop reasoning for the prompts of certain relation types, with the reasoning pathway used in more than 80% of the prompts. However, the utilization is highly contextual, varying across different types of prompts. Also, on average, the evidence for the second hop and the full multi-hop traversal is rather moderate and only substantial for the first hop. Moreover, we find a clear scaling trend with increasing model size for the first hop of reasoning but not for the second hop. Our experimental findings suggest potential challenges and opportunities for future development and applications of LLMs.^1

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Problem Formulation
  • 3.1 Preliminaries
  • 3.2 Latent Multi-Hop Reasoning in LLMs
  • 4 TWOHOPFACT Dataset
  • 5 First Hop of Multi-Hop Reasoning
  • 5.1 Internal Entity Recall Score
  • 5.2 Experiment
  • 5.3 Results
  • 6 Second Hop of Multi-Hop Reasoning
  • 6.1 Consistency Score
  • 6.2 Experiment
  • 6.3 Results
  • 7 Latent Multi-Hop Reasoning
  • 8 Discussion and Conclusion
  • 9 Limitations
  • Acknowledgements
  • References
  • A Dataset construction
  • A.1 Data Selection
  • A.2 Natural Language Templates
  • B Dataset Statistics
  • C Justification of Internal Entity Recall Score: Appositive Generation Experiment
  • D Justification of Consistency Score: Comparative Experiment with Chain-of-Thought Cases
  • E Technical Details
  • F Accuracy-based Analysis for the Second Hop of Multi-Hop Reasoning

Knowls

  1. Knowl 1 — Internal Entity Recall Score (EntRec)

    definition

    The internal entity recall score (EntRec\text{EntRec}) measures the degree to which a transformer-based language model internally recalls a bridge entity e2e_2 at layer ll while processing an indirect descriptive mention μ(r1(e1))\mu(r_1(e_1)) (e.g., "the singer of 'Superstition'" for Stevie Wonder) within a multi-hop prompt τ2H\tau_{2H}.

    For layer l∈[0,L−1]l \in [0, L-1], let xl∈Rhx^l \in \mathbb{R}^h denote the hidden representation output from the ll-th transformer layer at the final token position of the descriptive mention μ(r1(e1))\mu(r_1(e_1)) in prompt τ2H\tau_{2H}. Let e2(0)e_2^{(0)} denote the first token of the entity name e2e_2, and let index(e2(0))∈[0,V−1]\text{index}(e_2^{(0)}) \in [0, V-1] denote its index in the vocabulary of size VV. Let WU∈Rh×VW_U \in \mathbb{R}^{h \times V} denote the unembedding projection matrix, and LayerNorm\text{LayerNorm} denote the final layer normalization applied before projection to vocabulary space. The internal entity recall score at layer ll is defined as:

    EntRecl(e2,τ2H)=log⁡softmax(LayerNorm(xl)WU)index(e2(0))\text{EntRec}^l(e_2, \tau_{2H}) = \log \text{softmax}\left(\text{LayerNorm}(x^l) W_U\right)_{\text{index}(e_2^{(0)})}

    Higher values of EntRecl(e2,τ2H)\text{EntRec}^l(e_2, \tau_{2H}) reflect stronger internal activation of the bridge entity e2e_2 at layer ll.

  2. Knowl 2 — First-Hop Latent Reasoning Evaluation via Input Substitution

    model/method

    To evaluate whether a language model executes the first hop of multi-hop reasoning by retrieving the bridge entity e2e_2 upon encountering its descriptive mention μ(r1(e1))\mu(r_1(e_1)), the model's internal entity recall score EntRecl(e2,τ2H)\text{EntRec}^l(e_2, \tau_{2H}) on a two-hop prompt τ2H\tau_{2H} is compared against its recall on an altered prompt τ2H′\tau'_{2H} that does not refer to e2e_2.

    Two forms of substitution are used to generate τ2H′\tau'_{2H}:

    1. Entity Substitution: The subject entity e1e_1 is replaced by e1′e'_1 from the same fact composition type such that μ(r1(e1′))≠e2\mu(r_1(e'_1)) \neq e_2 (e.g., replacing "the singer of 'Superstition'" with "the singer of 'Thriller'").
    2. Relation Substitution: The relation r1r_1 is replaced by an alternative relation r1′r'_1 from a candidate set such that μ(r1′(e1))≠e2\mu(r'_1(e_1)) \neq e_2 (e.g., replacing "the singer of 'Superstition'" with "a plagiarist of 'Superstition'").

    The evaluation metric is the relative frequency across prompts where the internal recall of e2e_2 strictly increases upon restoring the true descriptive mention:

    P(EntRecl(e2,τ2H)>EntRecl(e2,τ2H′))P\left(\text{EntRec}^l(e_2, \tau_{2H}) > \text{EntRec}^l(e_2, \tau'_{2H})\right)

    A relative frequency exceeding the random chance baseline of 0.50.5 indicates that the model systematically performs latent first-hop recall.

  3. Knowl 3 — Symmetric Consistency Score (CnstScore)

    definition

    The consistency score (CnstScore\text{CnstScore}) evaluates the distributional similarity between a model's next-token generation for a two-hop prompt τ2H\tau_{2H} (e.g., "The mother of the singer of 'Superstition' is") and its corresponding direct one-hop prompt τ1H\tau_{1H} (e.g., "The mother of Stevie Wonder is"), independent of ground-truth label correctness.

    Let pτ2H,pτ1H∈RVp_{\tau_{2H}}, p_{\tau_{1H}} \in \mathbb{R}^V denote the model's output probability distributions over vocabulary VV for τ2H\tau_{2H} and τ1H\tau_{1H}, respectively. Let H(P,Q)=−∑i=0V−1Pilog⁡QiH(P, Q) = -\sum_{i=0}^{V-1} P_i \log Q_i denote the cross-entropy from distribution PP to distribution QQ. The symmetric consistency score is defined as:

    CnstScore(τ2H,τ1H)=−0.5 H(pτ2H,pτ1H)−0.5 H(pτ1H,pτ2H)\text{CnstScore}(\tau_{2H}, \tau_{1H}) = -0.5 \, H(p_{\tau_{2H}}, p_{\tau_{1H}}) - 0.5 \, H(p_{\tau_{1H}}, p_{\tau_{2H}})

    A higher CnstScore\text{CnstScore} indicates greater alignment between how the model answers the two-hop composition and how it answers the direct one-hop query about the bridge entity's attribute.

  4. Knowl 4 — Second-Hop Utilization Evaluation via Directional Gradient Intervention

    model/method

    To evaluate whether an autoregressive language model utilizes its recalled bridge entity representation to solve the second hop of reasoning, causal intervention via activation patching along the gradient of entity recall is performed.

    Let xl∈Rhx^l \in \mathbb{R}^h be the intermediate representation at layer l∈[0,L−1)l \in [0, L-1) at the final token position of the descriptive mention in two-hop prompt τ2H\tau_{2H}. The internal recall EntRec(xl)\text{EntRec}(x^l) and consistency CnstScore(xl)\text{CnstScore}(x^l) are treated as functions of xlx^l. The representation is intervened upon by shifting it in the direction of steepest increase in bridge entity recall by magnitude α\alpha:

    x^l(α)=xl+α∇xlEntRec(xl)\hat{x}^l(\alpha) = x^l + \alpha \nabla_{x^l} \text{EntRec}(x^l)

    Using activation patching to replace xlx^l with x^l(α)\hat{x}^l(\alpha), the prompt consistency is evaluated as a parameterized function CnstScore(α)\text{CnstScore}(\alpha). The directional derivative at α=0\alpha = 0 is computed:

    ddαCnstScore(α)∣α=0\left. \frac{d}{d\alpha} \text{CnstScore}(\alpha) \right|_{\alpha=0}

    A positive derivative demonstrates that increasing internal bridge entity recall at layer ll causally increases output consistency with the one-hop prompt τ1H\tau_{1H}. The proportion of dataset prompts exhibiting a positive derivative quantifies second-hop utilization against a random chance baseline of 0.50.5.

  5. Knowl 5 — Empirical Scaling of First-Hop Latent Recall

    empirical result

    Experiments on LLaMA-2 models (7B, 13B, and 70B) across 52 fact composition types in the TwoHopFact dataset show substantial evidence of latent first-hop entity recall that scales positively with model parameter size:

    • Entity Substitution: The peak relative frequency of increased bridge entity recall across layers rises from 0.710.71 (at layer 31) in 7B to 0.720.72 (at layer 39) in 13B and 0.780.78 (at layer 79) in 70B.
    • Relation Substitution: The peak relative frequency rises from 0.630.63 (at layer 20) in 7B to 0.640.64 (at layer 36) in 13B and 0.760.76 (at layer 75) in 70B.
    • Composition Type Prevalence: The number of fact composition types (out of 52) achieving a maximum relative frequency above 0.800.80 increases from 1818 (7B) to 2525 (13B) and 3434 (70B) for entity substitution, and from 2121 (7B) to 2727 (13B) and 3838 (70B) for relation substitution. Eleven composition types consistently maintain peak relative frequencies >0.80>0.80 across all three model scales and both substitution types.
  6. Knowl 6 — Empirical Non-Scaling Behavior in Second-Hop Latent Utilization

    empirical result

    Causal gradient intervention experiments on LLaMA-2 models measuring the rate at which increasing bridge entity recall increases consistency (CnstScore\text{CnstScore}) reveal moderate second-hop reasoning that does not scale with model parameter count:

    • Aggregated Performance: The maximum relative frequency of positive consistency gradients across layers remains essentially flat as model size increases: 0.640.64 at layer 30 in LLaMA-2 7B, 0.650.65 at layer 14 in LLaMA-2 13B, and 0.610.61 at layer 40 in LLaMA-2 70B (compared to a randomly initialized 7B baseline of 0.500.50).
    • Ground-Truth Log-Probability Metric: When replacing CnstScore\text{CnstScore} with the log-probability of the ground-truth target token, peak relative frequencies remain similarly flat: 0.600.60 at layer 30 (7B), 0.620.62 at layer 13 (13B), and 0.590.59 at layer 40 (70B).
    • One-Hop Accuracy Invariance: Filtering prompts into subsets where the direct one-hop prompt is answered correctly versus incorrectly yields comparable peak frequencies (0.620.62–0.650.65 vs. 0.610.61–0.650.65) across all model sizes.
    • Contextual Outliers: Up to 19%19\% of fact composition types (10 for 7B, 7 for 13B, and 5 for 70B) exhibit peak relative frequencies >0.80>0.80, with types such as "founder of person's undergrad university" and "president of anthem's country" exceeding 0.800.80 across all model sizes.
  7. Knowl 7 — Joint Latent Multi-Hop Traversal and Contextual Specificity

    empirical result

    The joint execution of full two-hop latent reasoning is defined as the simultaneous occurrence of first-hop success (recall increases under prompt substitution) and second-hop success (consistency increases under recall gradient ascent):

    • Aggregate Traversal Rate: In LLaMA-2 7B, joint success occurs with a peak frequency of 0.460.46 under entity substitution and 0.380.38 under relation substitution, exceeding the random chance baseline of 0.250.25 (0.5×0.50.5 \times 0.5).
    • Scaling Characteristics: Under entity substitution, joint success remains flat across scales (0.460.46 in 7B, 0.430.43 in 13B, 0.450.45 in 70B). Under relation substitution, joint success scales modestly from 0.380.38 (7B) to 0.370.37 (13B) and 0.430.43 (70B).
    • High Contextuality: Latent multi-hop traversal is concentrated in specific relation types. Using a joint success threshold of 0.640.64 (0.820.8^2), 7/527/52 types in 7B, 3/523/52 in 13B, and 12/5212/52 in 70B exceed the threshold under entity substitution (up to 23%23\% of composition types). For example, "anthem of capital's country" achieves joint success rates of 0.680.68 (7B), 0.820.82 (13B), and 0.660.66 (70B) under entity substitution.
  8. Knowl 8 — TwoHopFact Benchmark Dataset

    experimental setup

    The TwoHopFact dataset contains 45,595 unique pairs of one-hop (τ1H\tau_{1H}) and two-hop (τ2H\tau_{2H}) prompts covering 52 fact composition types derived from Wikidata.

    Each sample is built from pairs of factual triplets ((e1,r1,e2),(e2,r2,e3))((e_1, r_1, e_2), (e_2, r_2, e_3)) satisfying r1(e1)=e2r_1(e_1) = e_2 and r2(e2)=e3r_2(e_2) = e_3, where e2e_2 is unique among facts within the same composition type and the most frequent majority bridge entity within any single type covers no more than 15%15\% of prompts.

    Prompts are created via manual templates:

    • Descriptive mention template mr1(⋅)m_{r_1}(\cdot) yields noun phrase μ(r1(e1))\mu(r_1(e_1)) (e.g., "the singer of 'Superstition'").
    • Relation prompt template tr2(⋅)t_{r_2}(\cdot) produces two-hop prompt τ2H=tr2(μ(r1(e1)))\tau_{2H} = t_{r_2}(\mu(r_1(e_1))) (e.g., "The mother of the singer of 'Superstition' is") and one-hop prompt τ1H=tr2(e2)\tau_{1H} = t_{r_2}(e_2) (e.g., "The mother of Stevie Wonder is").

    The 52 composition types span categories including countries, organizations, real persons, fictional characters, movies, novels, and cities, with every composition type containing at least 30 prompt pairs.

  9. Knowl 9 — Validation of Internal Entity Recall via Appositive Generation

    model/method

    To validate that the metric EntRecl(e2,τ2H)\text{EntRec}^l(e_2, \tau_{2H}) reflects the model's internal representation of the bridge entity e2e_2, an appositive generation control probe tests whether increasing EntRec\text{EntRec} increases the probability of generating e2e_2 as an appositive following a comma.

    In a two-hop prompt prefix ending at the descriptive mention μ(r1(e1))\mu(r_1(e_1)), appending a comma makes the bridge entity name e2e_2 a grammatically natural appositive continuation (e.g., "The mother of the singer of 'Superstition'," followed by "Stevie Wonder").

    The hidden state representation xlx^l at layer ll is perturbed along the gradient of EntRecl(e2,τ2H)\text{EntRec}^l(e_2, \tau_{2H}) via activation patching: x^l(α)=xl+α∇xlEntRec(xl)\hat{x}^l(\alpha) = x^l + \alpha \nabla_{x^l} \text{EntRec}(x^l). Evaluating the directional derivative of the model's output log-probability for the first token e2(0)e_2^{(0)} of the entity name shows that positive gradients occur significantly above the 0.50.5 random baseline across middle and late layers in LLaMA-2 7B, confirming that EntRec\text{EntRec} controls the model's generation of the bridge entity.

  10. Knowl 10 — Limitations of Single-Layer Latent Multi-Hop Reasoning Analysis

    limitation

    The methodology for measuring latent multi-hop reasoning in language models has several inherent structural limitations:

    1. Single-Layer Interventions: The analysis evaluates representation changes and interventions within isolated individual layers rather than tracking end-to-end multi-layer causal propagation across depth.
    2. First-Token Approximation: Internal entity recall is approximated via the log-probability of only the first subword token e2(0)e_2^{(0)} of the bridge entity rather than its full multi-token representation.
    3. Logit Lens Vulnerabilities: Projecting intermediate hidden states with the unembedding matrix inherits known logit-lens biases, representation drift, and vocabulary calibration noise.
    4. Inference Pathway Redundancy: LLMs can rely on distributed or redundant inference pathways; measuring a single mechanistic pathway provides a lower bound on the model's latent multi-hop reasoning capacity.

Coverage note — No substantial contributed material was omitted; all definitions, experimental metrics, empirical findings across RQ1/RQ2/joint analysis, dataset specifications, validation controls, and limitations are represented.

References

  1. 1.Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. 2023. What learning algorithm is in-context learning? investigations with linear models. In ICLR.
  2. 2.Zeyuan Allen-Zhu and Yuanzhi Li. 2023a. Physics of language models: Part 3.1, knowledge storage and extraction. arXiv.
  3. 3.Zeyuan Allen-Zhu and Yuanzhi Li. 2023b. Physics of language models: Part 3.2, knowledge manipulation. arXiv.
  4. 4.Akari Asai and Hannaneh Hajishirzi. 2020. Logic-guided data augmentation and regularization for consistent question answering. In ACL.
  5. 5.Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. 2023. Eliciting latent predictions from transformers with the tuned lens. arXiv.
  6. 6.Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. 2023. Taken out of context: On measuring situational awareness in llms. arXiv.
  7. 7.Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2024. The reversal curse: LLMs trained on “a is b” fail to learn “b is a”. In ICLR.
  8. 8.Jannik Brinkmann, Abhay Sheshadri, Victor Levoso, Paul Swoboda, and Christian Bartelt. 2023. A mechanistic analysis of a transformer trained on a symbolic multi-step reasoning task.
  9. 9.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In NeurIPS.
  10. 10.Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill. 2022. Data distributional properties drive emergent in-context learning in transformers. In NeurIPS.
  11. 11.David Chanin, Anthony Hunter, and Oana-Maria Camburu. 2023. Identifying linear relational concepts in large language models. arXiv.
  12. 12.Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2023. Evaluating the ripple effects of knowledge editing in language models. arXiv.
  13. 13.Arthur Conmy, Augustine N Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023. Towards automated circuit discovery for mechanistic interpretability. In NeurIPS.
  14. 14.Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2023. Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers. In Findings of ACL.
  15. 15.Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. Editing factual knowledge in language models. In EMNLP.
  16. 16.Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. 2023. Jump to conclusions: Shortcutting transformers with linear transformations. arXiv.
  17. 17.Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D Hwang, Soumya Sanyal, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi. 2023. Faith and fate: Limits of transformers on compositionality. In NeurIPS.
  18. 18.Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. TACL.
  19. 19.Jiahai Feng and Jacob Steinhardt. 2024. How do language models bind entities in context? In ICLR.
  20. 20.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. Dissecting recall of factual associations in auto-regressive language models. In EMNLP.
  21. 21.Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In EMNLP.
  22. 22.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In EMNLP.
  23. 23.Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. 2024. Linearity of relation decoding in transformer language models. In ICLR.
  24. 24.Yifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosse lut, and Mrinmaya Sachan. 2023. Towards a mechanistic interpretation of multi-step reasoning capabilities of language models. In ACL.
  25. 25.Myeongjun Jang, Bodhisattwa Prasad Majumder, Julian McAuley, Thomas Lukasiewicz, and Oana-Maria Camburu. 2023. Know how to make up your mind! adversarially detecting and alleviating inconsistencies in natural language explanations. In ACL.
  26. 26.Nora Kassner, Oyvind Tafjord, Ashish Sabharwal, Kyle Richardson, Hinrich Schuetze, and Peter Clark. 2023. Language models with rationality. In EMNLP.
  27. 27.Nora Kassner, Oyvind Tafjord, Hinrich Schütze, and Peter Clark. 2021. BeliefBank: Adding memory to a pre-trained language model for a systematic notion of belief. In EMNLP.
  28. 28.Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2023. Transformer language models handle word frequency in prediction head. In ACL.
  29. 29.Tao Li, Vivek Gupta, Maitrey Mehta, and Vivek Srikumar. 2019. A logic-driven framework for consistency of neural models. In EMNLP.
  30. 30.Zhaoyi Li, Gangwei Jiang, Hong Xie, Linqi Song, Defu Lian, and Ying Wei. 2024. Understanding and patching compositional reasoning in llms. arXiv.
  31. 31.Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik. 2023. Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla. arXiv.
  32. 32.Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. 2023. The hydra effect: Emergent self-repair in language model computations. arXiv.
  33. 33.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in GPT. In NeurIPS.
  34. 34.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2022. Fast model editing at scale. In ICLR.
  35. 35.Neel Nanda and Joseph Bloom. 2022. Transformer-lens. https://github.com/neelnanda-io/TransformerLens.
  36. 36.Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2022. Progress measures for grokking via mechanistic interpretability. In ICLR.
  37. 37.nostalgebraist. 2020. interpreting gpt: the logit lens.
  38. 38.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. In Findings of EMNLP.
  39. 39.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022. In-context learning and induction heads. arXiv.
  40. 40.Yasumasa Onoe, Michael J Q Zhang, Shankar Padmanabhan, Greg Durrett, and Eunsol Choi. 2023. Can LMs learn new entities from descriptions? challenges in propagating injected knowledge. In ACL.
  41. 41.OpenAI, :, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jason Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O’Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang Song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, and Barret Zoph. 2023. Gpt-4 technical report. arXiv.
  42. 42.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? In EMNLP.
  43. 43.Ben Prystawski and Noah D Goodman. 2023. Why think step-by-step? reasoning emerges from the locality of experience. In NeurIPS.
  44. 44.Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh. 2019. Are red roses red? evaluating consistency of question-answering models. In ACL.
  45. 45.Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Daniel Grzenda, Nathaniel Hudson, André Bauer, Kyle Chard, and Ian Foster. 2023. Memory injections: Correcting multi-hop reasoning failures during inference in transformer-based language models. arXiv.
  46. 46.Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Seyed Mehran Kazemi, Najoung Kim, and He He. 2023. Testing the general deductive reasoning capacity of large language models using OOD examples. In NeurIPS.
  47. 47.William Timkey and Marten van Schijndel. 2021. All bard and no bite: Rogue dimensions in transformer language models obscure representational quality. In EMNLP.
  48. 48.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv.
  49. 49.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS.
  50. 50.Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2023. Transformers learn in-context by gradient descent. In ICML.
  51. 51.Denny Vrandečic and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM.
  52. 52.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In ICLR.
  53. 53.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022a. Emergent abilities of large language models. TMLR.
  54. 54.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In NeurIPS.
  55. 55.Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel. 2018. Constructing datasets for multi-hop reading comprehension across documents. TACL.
  56. 56.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Huggingface’s transformers: State-of-the-art natural language processing. arXiv.
  57. 57.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In EMNLP.
  58. 58.Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, Siyuan Cheng, Ziwen Xu, Xin Xu, Jia-Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Lei Liang, Zhiqiang Zhang, Xiaowei Zhu, Jun Zhou, and Huajun Chen. 2024. A comprehensive study of knowledge editing for large language models. arXiv.
  59. 59.Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen. 2023. MQAKE: Assessing knowledge editing in language models via multi-hop questions. In EMNLP.
  60. 60.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. In ICLR.

Citation

MLA
Yang, S., et al. “Do Large Language Models Latently Perform Multi-Hop Reasoning?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10210–29, https://doi.org/10.18653/v1/2024.acl-long.550.
APA
Yang, S., Gribovskaya, E., Kassner, N., Geva, M., & Riedel, S. (2024). Do Large Language Models Latently Perform Multi-Hop Reasoning?. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10210–10229. https://doi.org/10.18653/v1/2024.acl-long.550
Chicago
Yang, S., E. Gribovskaya, N. Kassner, M. Geva, and S. Riedel. 2024. “Do Large Language Models Latently Perform Multi-Hop Reasoning?”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10210–29. https://doi.org/10.18653/v1/2024.acl-long.550.
Harvard
Yang, S. et al. (2024) “Do Large Language Models Latently Perform Multi-Hop Reasoning?”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10210–10229. Available at: https://doi.org/10.18653/v1/2024.acl-long.550.
Vancouver
1. Yang S, Gribovskaya E, Kassner N, Geva M, Riedel S (2024) Do Large Language Models Latently Perform Multi-Hop Reasoning?. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10210–10229

BibTeX

@inproceedings{yang-etal-2024-large-language-models,
    title = "Do Large Language Models Latently Perform Multi-Hop Reasoning?",
    author = "Yang, Sohee  and
      Gribovskaya, Elena  and
      Kassner, Nora  and
      Geva, Mor  and
      Riedel, Sebastian",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.550/",
    doi = "10.18653/v1/2024.acl-long.550",
    pages = "10210--10229"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/