Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries

Eden BiranDaniela GottesmanSohee YangMor GevaAmir Globerson

article2024EMNLP109 citations

Reveals that large language models often fail at multi-hop reasoning because the intermediate entity resolves too late for subsequent layers to extract the final answer, and shows that patching hidden representations back to earlier layers recovers correct predictions in up to 66% of failed cases.

Listen

Large language models often struggle to answer multi-hop factual queries that require chaining pieces of knowledge together, such as identifying the spouse of a song's performer. Even when a model knows each individual fact in isolation, it frequently fails when combining them in a single prompt. Understanding how models process this multi-step reasoning internally is crucial for diagnosing errors, improving factual accuracy, and advancing reliable automated reasoning.

The article demonstrates the internal mechanics of latent multi-hop reasoning in large language models and evaluates why these reasoning processes fail. Specifically, it tracks how and where models extract intermediary information across their internal computational layers during multi-step tasks.

To investigate this, the researchers created a dataset of 82,020 two-hop queries using factual data from Wikidata. They filtered out shortcuts to isolate genuine multi-hop reasoning and evaluated six prominent open-source language models across different families (LLaMA 2, LLaMA 3, and Pythia) ranging from 6.9 billion to 70 billion parameters. Using representation probing techniques (primarily Patchscopes) alongside attention knockout and sublayer projections, the team tracked the flow of information across network layers and token positions. They also introduced an experimental analysis technique called back-patching, which copies intermediate representations from later layers back into earlier layers to provide more computational depth.

The analysis revealed a clear, sequential four-stage reasoning pathway: the intermediary "bridge" entity is first resolved in the early layers at the end of the first-hop phrase, this information propagates across middle layers to the prompt's final token, and the ultimate target entity is resolved in the later layers, where feed-forward sublayers heavily promote the final output. Crucially, failures predominantly occur when the first hop takes too long to resolve; in incorrect cases, the bridge entity emerged significantly later in the network, leaving insufficient layers to retrieve the final answer. Testing back-patching confirmed this limitation, successfully recovering the correct answer in 32% to 66% of previously failed cases without requiring any parameter updates or retraining.

These findings indicate that transformer models face an inherent architectural bottleneck: because knowledge retrieval is distributed across a fixed depth, a delayed initial step leaves the model with too few remaining layers to perform subsequent lookups. For practitioners and decision-makers, this highlights why standard language models struggle with complex, chained tasks and underscores the risks of relying on direct generation for multi-step reasoning. It also explains why explicit reasoning strategies, such as chain-of-thought prompting that forces the model to output intermediate steps into the text, remain substantially more reliable than relying on hidden internal computation.

Organizations deploying language models for complex knowledge retrieval should avoid expecting models to reliably perform multi-step latent reasoning in a single pass. Instead, leaders should prioritize structured prompting workflows or chain-of-thought methods when high factual accuracy is required. Future technical research should focus on methods to predict optimal source-target layer pairs for back-patching during inference or design architectures that dynamically allocate computational depth.

The analysis is subject to certain limitations. Mechanistic probing methods approximate internal states rather than providing perfect readouts, and the empirical study focused exclusively on two-hop factual queries rather than arbitrary multi-step or non-factual reasoning tasks. Furthermore, while back-patching proves the root cause of these reasoning failures, it is currently an analytical diagnostic rather than a real-time production inference technique. Confidence in the underlying conclusion—that models fail multi-hop queries due to running out of network depth—remains high due to consistent results across all tested model sizes and families.

Biran et al (2024).pdf
Cover for Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries

Abstract

Large language models (LLMs) can solve complex multi-step problems, but little is known about how these computations are implemented internally. Motivated by this, we study how LLMs answer multi-hop queries such as “The spouse of the performer of Imagine is”. These queries require two information extraction steps: a latent one for resolving the first hop (“the performer of Imagine”) into the bridge entity (John Lennon), and another for resolving the second hop (“the spouse of John Lennon”) into the target entity (Yoko Ono). Understanding how the latent step is computed internally is key to understanding the overall computation. By carefully analyzing the internal computations of transformer-based LLMs, we discover that the bridge entity is resolved in the early layers of the model. Then, only after this resolution, the two-hop query is solved in the later layers. Because the second hop commences in later layers, there could be cases where these layers no longer encode the necessary knowledge for correctly predicting the answer. Motivated by this, we propose a novel “back-patching” analysis method whereby a hidden representation from a later layer is patched back to an earlier layer. We find that in up to 66% of previously incorrect cases there exists a back-patch that results in the correct generation of the answer, showing that the later layers indeed sometimes lack the needed functionality. Overall, our methods and findings open further opportunities for understanding and improving latent reasoning in transformer-based LLMs.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Experimental Setup
  • 2.1 Two-Hop Queries
  • 2.2 Dataset
  • 2.3 Models
  • 3 Localizing First Hop Resolution
  • 3.1 Interpreting Hidden Representations
  • 3.2 Experiment
  • 3.3 Results
  • 4 Second Hop is Resolved at Last Position
  • 4.1 Method
  • 4.2 Experiment
  • 4.3 Results
  • 5 Information Propagates to Last Token
  • 5.1 Method
  • 5.2 Experiment
  • 5.3 Results
  • 6 Back-patching Improves Two-Hop Performance
  • 6.1 Method
  • 6.2 Experiment
  • 6.3 Results
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Technical Details
  • B Patchscopes First Decoded Layers
  • C Patchscopes Heat-maps
  • D First Observed Layers of Pathway Stages
  • E Back-patching Heat-maps

Knowls

  1. Knowl 1 — Two-hop answers emerge through ordered hidden-state resolutions

    empirical result

    For a query formed by composing facts (e1,r1,e2)(e_1,r_1,e_2) and (e2,r2,e3)(e_2,r_2,e_3), e2e_2 is the bridge entity and e3e_3 is the final answer. Let t1t_1 be the final token of the first-hop clause and t2t_2 the final token of the full prompt. Patchscopes decoding found that the bridge entity often appears in the representation at t1t_1 in early layers; it appears at t2t_2 later, and the final answer appears at t2t_2 later still. Each percentage below is the share of tested examples in which Patchscopes decoded the specified entity; the layer value is the mean source layer of its first successful decoding. A decoding counted as successful if at least one of three sampled generations contained the entity.

    Model Outcome e2e_2 at t1t_1 (%; layer) e2e_2 at t2t_2 (%; layer) e3e_3 at t2t_2 (%; layer)
    LLaMA 2 7B Correct 56%;8.956\%; 8.9 41%;15.341\%; 15.3 73%;16.273\%; 16.2
    LLaMA 2 7B Incorrect 41%;9.141\%; 9.1 36%;16.436\%; 16.4 47%;17.547\%; 17.5
    LLaMA 2 13B Correct 49%;7.749\%; 7.7 34%;18.234\%; 18.2 71%;16.971\%; 16.9
    LLaMA 2 13B Incorrect 48%;7.448\%; 7.4 37%;17.937\%; 17.9 33%;16.933\%; 16.9
    LLaMA 3 8B Correct 46%;8.946\%; 8.9 37%;14.337\%; 14.3 74%;13.574\%; 13.5
    LLaMA 3 8B Incorrect 46%;11.946\%; 11.9 25%;15.325\%; 15.3 40%;17.240\%; 17.2
    LLaMA 3 70B Correct 63%;27.163\%; 27.1 46%;35.246\%; 35.2 86%;30.786\%; 30.7
    LLaMA 3 70B Incorrect 57%;29.257\%; 29.2 46%;34.546\%; 34.5 50%;35.950\%; 35.9
    Pythia 6.9B Correct 75%;5.475\%; 5.4 75%;11.475\%; 11.4 80%;14.180\%; 14.1
    Pythia 6.9B Incorrect 78%;4.978\%; 4.9 67%;9.267\%; 9.2 65%;13.665\%; 13.6
    Pythia 12B Correct 73%;5.073\%; 5.0 66%;12.266\%; 12.2 77%;13.577\%; 13.5
    Pythia 12B Incorrect 61%;6.361\%; 6.3 42%;12.942\%; 12.9 52%;17.252\%; 17.2

    The results support a sequential pathway: the first hop is often resolved at t1t_1 before the bridge entity becomes detectable at the answer-producing position t2t_2, where the second hop yields e3e_3. Final-answer decoding at t2t_2 was more frequent for correct than incorrect cases in all six models (71–86% versus 33–65%). These rates are approximate indicators of what a representation can expose through the decoding probe, not direct measurements of every computation performed by the model.

  2. Knowl 2 — Bridge-entity information is detected propagating to the answer position

    empirical result

    In a two-hop prompt, let t1t_1 denote the final token of the first-hop clause and t2t_2 the final token of the full prompt. The study tested whether information about the bridge entity e2e_2 travels from t1t_1 to t2t_2 using three signals: attention knockout, sublayer vocabulary projections, and Patchscopes decoding. Attention knockout blocked the t1t_1-to-t2t_2 attention edge over a window of seven layers; the other tests checked for evidence of e2e_2 in updates or decoded representations at t2t_2. A case counted if any method detected relevant information. The reported layer is the mean layer of first detection.

    Model Outcome Detected cases (%) Mean first-detection layer
    LLaMA 2 7B Correct 85.82 12.83
    LLaMA 2 7B Incorrect 95.44 8.10
    LLaMA 2 13B Correct 81.94 15.45
    LLaMA 2 13B Incorrect 71.11 16.99
    LLaMA 3 8B Correct 78.36 10.15
    LLaMA 3 8B Incorrect 82.74 7.45
    LLaMA 3 70B Correct 90.39 29.14
    LLaMA 3 70B Incorrect 94.28 20.06
    Pythia 6.9B Correct 96.53 11.94
    Pythia 6.9B Incorrect 95.04 8.06
    Pythia 12B Correct 94.18 11.23
    Pythia 12B Incorrect 93.81 9.28

    At least one signal detected propagation in 71.11–96.53% of cases. The mean detection layers generally fall between the early first-hop resolution and the later second-hop resolution, consistent with the bridge entity being conveyed to the position from which the model generates its answer. Detection is evidence of information flow under these analyses, not proof that a single fixed pathway is used in every example.

  3. Knowl 3 — Back-patching tests whether later representations can use earlier-layer computation

    model/method

    Back-patching is an intervention for testing whether a representation computed later in a transformer can produce a better answer when processed through earlier layers. Given a prompt, a token position, a source layer ℓs\ell_s, and an earlier target layer ℓt\ell_t with ℓt<ℓs\ell_t<\ell_s, the procedure records the token's hidden representation at ℓs\ell_s. It then reruns the same prompt, replaces the representation at the same token position at ℓt\ell_t with the recorded vector, continues the forward pass, and generates an answer by greedy decoding. The study applies this procedure to both t1t_1 (the first-hop clause's final token) and t2t_2 (the full prompt's final token), testing source–target layer pairs and counting a pair as successful when the generated answer matches the known target entity. This injects later-computed information into an earlier point in the computation without changing model parameters.

  4. Knowl 4 — Back-patching recovers many otherwise incorrect two-hop answers

    empirical result

    The back-patching evaluation searched source–target layer pairs with the source layer later than the target layer, and counted a prompt as successful if at least one pair produced the correct target entity under greedy decoding. The table reports the share of examples with a successful pair for each token position. The 100% rates for originally correct examples mean a successful pair could be found in every tested case; the incorrect-case rates measure recovery potential under this layer-pair search, not performance from a single preset intervention.

    Model Original outcome Patch at t1t_1 (%) Patch at t2t_2 (%)
    LLaMA 2 7B Correct 100 100
    LLaMA 2 7B Incorrect 41.02 42.45
    LLaMA 2 13B Correct 100 100
    LLaMA 2 13B Incorrect 32.44 36.07
    LLaMA 3 8B Correct 100 100
    LLaMA 3 8B Incorrect 38.81 47.16
    LLaMA 3 70B Correct 100 100
    LLaMA 3 70B Incorrect 57.31 57.81
    Pythia 6.9B Correct 100 100
    Pythia 6.9B Incorrect 66.33 56.43
    Pythia 12B Correct 100 100
    Pythia 12B Incorrect 63.17 61.82

    Across the six models, a successful back-patch was found for 32.44–66.33% of initially incorrect cases at t1t_1 and 36.07–61.82% at t2t_2. This demonstrates that changing where a hidden representation enters the remaining computation can rescue many failures, supporting the hypothesis that some failures reflect a mismatch between when information becomes available and which layers can use it.

  5. Knowl 5 — MLP updates usually contribute more than attention updates to answer-token promotion

    empirical result

    For each attention or MLP sublayer update at the final prompt position t2t_2, the study projected the update into vocabulary space using the layer normalization and output embedding matrix. An update counted as promoting the answer token when the highest-probability vocabulary token in its projection matched the first token of the model's generated answer. The table gives the percentage of cases in which this first occurred for each sublayer type and the mean layer of first occurrence.

    Model and outcome Attention (cases %; layer) MLP (cases %; layer)
    LLaMA 2 7B, correct 25.5; 24.7 33.2; 25.2
    LLaMA 2 7B, incorrect 14.8; 23.4 28.4; 26.0
    LLaMA 2 13B, correct 28.8; 32.0 68.5; 33.1
    LLaMA 2 13B, incorrect 14.2; 27.0 51.5; 31.7
    LLaMA 3 8B, correct 9.7; 26.7 17.6; 27.8
    LLaMA 3 8B, incorrect 25.3; 27.4 11.8; 24.6
    LLaMA 3 70B, correct 21.9; 56.3 48.1; 68.0
    LLaMA 3 70B, incorrect 21.8; 65.7 36.1; 67.0
    Pythia 6.9B, correct 11.5; 21.8 36.4; 20.5
    Pythia 6.9B, incorrect 28.2; 23.4 33.6; 21.2
    Pythia 12B, correct 27.9; 22.2 38.9; 22.4
    Pythia 12B, incorrect 43.2; 23.6 47.0; 23.3

    The MLP update has the higher detection rate in most model/outcome combinations, especially in LLaMA 2 13B, while attention updates also promote the predicted token in a non-negligible share of cases. The evidence therefore favors a larger MLP role without establishing an exclusive MLP pathway. The first promotion generally occurs in upper layers, after the answer entity is often decodable from the representation at t2t_2.

  6. Knowl 6 — The study isolates factual two-hop reasoning cases from shortcut successes

    experimental setup

    The dataset contains 82,020 two-hop queries constructed from Wikidata facts. Each query composes (e1,r1,e2)(e_1,r_1,e_2) with (e2,r2,e3)(e_2,r_2,e_3), where e2e_2 is the bridge entity; manually written relation templates convert the facts into natural-language prompts. To filter likely shortcuts, each model was tested with two altered prompts: one omitted the first-hop source entity, and the other omitted the first-hop relation. Examples where either altered prompt still elicited the final answer were excluded using greedy decoding. The researchers then formed model-specific subsets: correct cases had to answer both the complete query and the first hop correctly; incorrect cases had to answer each fact correctly when asked separately but fail on the composed query. They sampled up to 100 correct and 50 incorrect cases per bridge-entity type.

    Model After shortcut filtering Correct cases tested Incorrect cases tested
    LLaMA 2 7B 70,625 388 351
    LLaMA 2 13B 70,972 554 413
    LLaMA 3 8B 71,569 379 371
    LLaMA 3 70B 70,334 656 595
    Pythia 6.9B 73,058 173 202
    Pythia 12B 74,056 172 372

    The analyzed models were LLaMA 2 7B and 13B, LLaMA 3 8B and 70B, and Pythia 6.9B and 12B. They have 32 layers for the 6.9B, 7B, and 8B models; 36 for Pythia 12B; 40 for LLaMA 2 13B; and 80 for LLaMA 3 70B. Because filtering and case selection were model-specific, the tested subsets differ across models.

  7. Knowl 7 — Patchscopes decodes entity information from hidden representations

    model/method

    Patchscopes was used to test whether a hidden representation at a chosen token and layer contains information about a particular entity. The model first processes the two-hop query, and the representation at the source token and source layer is recorded. The same model then processes a separate natural-language decoding prompt containing examples such as “Syria: Syria is a country in the Middle East” and “Leonardo DiCaprio: Leonardo DiCaprio is an American actor,” followed by an unfinished entry for “x.” The recorded representation replaces the representation of “x” at a chosen target layer; the forward pass then continues and generates text describing the represented content. The experiment repeats this across source and target layers. For each source layer, it samples three generations and counts the entity as decoded if any generation contains the entity name. This method can reveal entities that are not the top vocabulary-projection token at the source layer, but its detection rates approximate what a representation makes accessible to this decoding procedure.

  8. Knowl 8 — Incorrect answers show a general tendency toward delayed entity resolution

    empirical result

    Comparisons of the first detected layers for pathway stages—first-hop resolution, bridge-information propagation, second-hop resolution, and prediction extraction—show a general, not universal, difference between successful and failed two-hop queries. In incorrect cases, entity resolutions tend to be detected later, while propagation and prediction extraction tend to be detected earlier than in correct cases. This pattern is consistent with the first hop becoming available too late for the later computation to use it effectively. The layer distributions vary across models and examples, so the comparison supports a tendency rather than a deterministic failure rule.

  9. Knowl 9 — The mechanistic conclusions are bounded by probe coverage and back-patching feasibility

    limitation

    The pathway evidence relies on approximate analyses: Patchscopes decodes representations through a separate prompt, and vocabulary projections interpret residual updates in token space. Using multiple probes reduces dependence on any one measure but does not make the interpretations exhaustive. The study traces one prominent two-hop pathway rather than all possible pathways, and does not resolve every component, including how relation information participates. Its experiments cover two-hop queries only. Back-patching demonstrates recovery potential when source and target layers are searched, but only a subset of tested patches succeeds; choosing useful layers in advance remains unresolved, so the method is not presented as a practical inference procedure. The authors also note that if maximizing question-answering performance is the sole goal, chain-of-thought prompting is likely more effective.

Coverage note — The per-model layer heatmaps and implementation hardware/runtime details are omitted because the aggregate results capture their substantive findings, while the resource details do not add a distinct scientific result.

References

  1. 1.David Bau. 2024. Baukit.
  2. 2.Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219.
  3. 3.Yonatan Belinkov and James Glass. 2019. Analysis methods in neural language processing: A survey. Transactions of the Association for Computational Linguistics, 7:49–72.
  4. 4.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, pages 2397–2430. PMLR.
  5. 5.Jannik Brinkmann, Abhay Sheshadri, Victor Levoso, Paul Swoboda, and Christian Bartelt. 2024. A mechanistic analysis of a transformer trained on a symbolic multi-step reasoning task. In Findings of the Association for Computational Linguistics ACL 2024, pages 4082–4102, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  6. 6.Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva. 2024. Evaluating the ripple effects of knowledge editing in language models. Transactions of the Association for Computational Linguistics, 12:283–298.
  7. 7.Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023. Towards automated circuit discovery for mechanistic interpretability. In Advances in Neural Information Processing Systems, volume 36, pages 16318–16352. Curran Associates, Inc.
  8. 8.Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8493–8502, Dublin, Ireland. Association for Computational Linguistics.
  9. 9.Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant. 2023. Analyzing transformers in embedding space. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16124–16170, Toronto, Canada. Association for Computational Linguistics.
  10. 10.Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. Editing factual knowledge in language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6491–6506, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  11. 11.Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang (Lorraine) Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena Hwang, Soumya Sanyal, Xiang Ren, Allyson Ettinger, Zaid Harchaoui, and Yejin Choi. 2023. Faith and fate: Limits of transformers on compositionality. In Advances in Neural Information Processing Systems, volume 36, pages 70293–70332. Curran Associates, Inc.
  12. 12.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. Dissecting recall of factual associations in auto-regressive language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12216–12235, Singapore. Association for Computational Linguistics.
  13. 13.Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 30–45, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  14. 14.Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 5484–5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  15. 15.Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024. Patchscopes: A unifying framework for inspecting hidden representations of language models. In Forty-first International Conference on Machine Learning.
  16. 16.Yifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, and Mrinmaya Sachan. 2023. Towards a mechanistic interpretation of multi-step reasoning capabilities of language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4902–4919, Singapore. Association for Computational Linguistics.
  17. 17.Cheng-Hsun Hsueh, Paul Kuo-Ming Huang, Tzu-Han Lin, Che-Wei Liao, Hung-Chieh Fang, Chao-Wei Huang, and Yun-Nung Chen. 2024. Editing the mind of giants: An in-depth exploration of pitfalls of knowledge editing in large language models. arXiv preprint arXiv:2406.01436.
  18. 18.Tianjie Ju, Yijin Chen, Xinwei Yuan, Zhuosheng Zhang, Wei Du, Yubin Zheng, and Gongshen Liu. 2024. Investigating multi-hop factual shortcuts in knowledge editing of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8987–9001, Bangkok, Thailand. Association for Computational Linguistics.
  19. 19.Brenden Lake and Marco Baroni. 2018. Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2873–2882. PMLR.
  20. 20.Zhaoyi Li, Gangwei Jiang, Hong Xie, Linqi Song, Defu Lian, and Ying Wei. 2024a. Understanding and patching compositional reasoning in LLMs. In Findings of the Association for Computational Linguistics ACL 2024, pages 9668–9688, Bangkok, Thailand and virtual meeting. Association for Computational Linguistics.
  21. 21.Zhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang, Xi Chen, and Huajun Chen. 2024b. Unveiling the pitfalls of knowledge editing for large language models. In The Twelfth International Conference on Learning Representations.
  22. 22.Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. 2023. The hydra effect: Emergent self-repair in language model computations. arXiv preprint arXiv:2307.15771.
  23. 23.Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in gpt. In Advances in Neural Information Processing Systems, volume 35, pages 17359–17372. Curran Associates, Inc.
  24. 24.Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2024. Language models implement simple Word2Vec-style vector arithmetic. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 5030–5047, Mexico City, Mexico. Association for Computational Linguistics.
  25. 25.Meta AI. 2024. Introducing meta llama 3: The most capable openly available llm to date.
  26. 26.Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2022. Fast model editing at scale. In International Conference on Learning Representations.
  27. 27.Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023. Progress measures for grokking via mechanistic interpretability. In The Eleventh International Conference on Learning Representations.
  28. 28.nostalgebraist. 2020. interpreting gpt: the logit lens.
  29. 29.Yasumasa Onoe, Michael Zhang, Shankar Padmanabhan, Greg Durrett, and Eunsol Choi. 2023. Can LMs learn new entities from descriptions? challenges in propagating injected knowledge. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5469–5485, Toronto, Canada. Association for Computational Linguistics.
  30. 30.Jackson Petty, Sjoerd Steenkiste, Ishita Dasgupta, Fei Sha, Dan Garrette, and Tal Linzen. 2024. The impact of depth on compositional generalization in transformer language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 7239–7252, Mexico City, Mexico. Association for Computational Linguistics.
  31. 31.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023. Measuring and narrowing the compositionality gap in language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 5687–5711, Singapore. Association for Computational Linguistics.
  32. 32.Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Daniel Grzenda, Nathaniel Hudson, André Bauer, Kyle Chard, and Ian Foster. 2023. Memory injections: Correcting multi-hop reasoning failures during inference in transformer-based language models. In Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 342–356, Singapore. Association for Computational Linguistics.
  33. 33.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  34. 34.Denny Vrandecic and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM, 57(10):78–85.
  35. 35.Boshi Wang, Xiang Yue, Yu Su, and Huan Sun. 2024. Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization. arXiv preprint arXiv:2405.15071.
  36. 36.Fei Wang, Wenjie Mo, Yiwei Wang, Wenxuan Zhou, and Muhao Chen. 2023a. A causal view of entity bias in (large) language models. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 15173–15184, Singapore. Association for Computational Linguistics.
  37. 37.Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023b. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. In The Eleventh International Conference on Learning Representations.
  38. 38.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems.
  39. 39.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  40. 40.Nan Xu, Fei Wang, Bangzheng Li, Mingtao Dong, and Muhao Chen. 2022. Does your model classify entities reasonably? diagnosing and mitigating spurious correlations in entity typing. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 8642–8658, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  41. 41.Sohee Yang, Elena Gribovskaya, Nora Kassner, Mor Geva, and Sebastian Riedel. 2024. Do large language models latently perform multi-hop reasoning? In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10210–10229, Bangkok, Thailand. Association for Computational Linguistics.
  42. 42.Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, et al. 2024. A comprehensive study of knowledge editing for large language models. arXiv preprint arXiv:2401.01286.
  43. 43.Zexuan Zhong, Zhengxuan Wu, Christopher Manning, Christopher Potts, and Danqi Chen. 2023. MQuAKE: Assessing knowledge editing in language models via multi-hop questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15686–15702, Singapore. Association for Computational Linguistics.

Citation

MLA
Biran, E., et al. “Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 14113–30, https://doi.org/10.18653/v1/2024.emnlp-main.781.
APA
Biran, E., Gottesman, D., Yang, S., Geva, M., & Globerson, A. (2024). Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 14113–14130. https://doi.org/10.18653/v1/2024.emnlp-main.781
Chicago
Biran, E., D. Gottesman, S. Yang, M. Geva, and A. Globerson. 2024. “Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 14113–30. https://doi.org/10.18653/v1/2024.emnlp-main.781.
Harvard
Biran, E. et al. (2024) “Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 14113–14130. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.781.
Vancouver
1. Biran E, Gottesman D, Yang S, Geva M, Globerson A (2024) Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 14113–14130

BibTeX

@inproceedings{biran-etal-2024-hopping,
    title = "Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries",
    author = "Biran, Eden  and
      Gottesman, Daniela  and
      Yang, Sohee  and
      Geva, Mor  and
      Globerson, Amir",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.781/",
    doi = "10.18653/v1/2024.emnlp-main.781",
    pages = "14113--14130"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/