Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
Eden BiranDaniela GottesmanSohee YangMor GevaAmir Globerson
Reveals that large language models often fail at multi-hop reasoning because the intermediate entity resolves too late for subsequent layers to extract the final answer, and shows that patching hidden representations back to earlier layers recovers correct predictions in up to 66% of failed cases.
Large language models often struggle to answer multi-hop factual queries that require chaining pieces of knowledge together, such as identifying the spouse of a song's performer. Even when a model knows each individual fact in isolation, it frequently fails when combining them in a single prompt. Understanding how models process this multi-step reasoning internally is crucial for diagnosing errors, improving factual accuracy, and advancing reliable automated reasoning.
The article demonstrates the internal mechanics of latent multi-hop reasoning in large language models and evaluates why these reasoning processes fail. Specifically, it tracks how and where models extract intermediary information across their internal computational layers during multi-step tasks.
To investigate this, the researchers created a dataset of 82,020 two-hop queries using factual data from Wikidata. They filtered out shortcuts to isolate genuine multi-hop reasoning and evaluated six prominent open-source language models across different families (LLaMA 2, LLaMA 3, and Pythia) ranging from 6.9 billion to 70 billion parameters. Using representation probing techniques (primarily Patchscopes) alongside attention knockout and sublayer projections, the team tracked the flow of information across network layers and token positions. They also introduced an experimental analysis technique called back-patching, which copies intermediate representations from later layers back into earlier layers to provide more computational depth.
The analysis revealed a clear, sequential four-stage reasoning pathway: the intermediary "bridge" entity is first resolved in the early layers at the end of the first-hop phrase, this information propagates across middle layers to the prompt's final token, and the ultimate target entity is resolved in the later layers, where feed-forward sublayers heavily promote the final output. Crucially, failures predominantly occur when the first hop takes too long to resolve; in incorrect cases, the bridge entity emerged significantly later in the network, leaving insufficient layers to retrieve the final answer. Testing back-patching confirmed this limitation, successfully recovering the correct answer in 32% to 66% of previously failed cases without requiring any parameter updates or retraining.
These findings indicate that transformer models face an inherent architectural bottleneck: because knowledge retrieval is distributed across a fixed depth, a delayed initial step leaves the model with too few remaining layers to perform subsequent lookups. For practitioners and decision-makers, this highlights why standard language models struggle with complex, chained tasks and underscores the risks of relying on direct generation for multi-step reasoning. It also explains why explicit reasoning strategies, such as chain-of-thought prompting that forces the model to output intermediate steps into the text, remain substantially more reliable than relying on hidden internal computation.
Organizations deploying language models for complex knowledge retrieval should avoid expecting models to reliably perform multi-step latent reasoning in a single pass. Instead, leaders should prioritize structured prompting workflows or chain-of-thought methods when high factual accuracy is required. Future technical research should focus on methods to predict optimal source-target layer pairs for back-patching during inference or design architectures that dynamically allocate computational depth.
The analysis is subject to certain limitations. Mechanistic probing methods approximate internal states rather than providing perfect readouts, and the empirical study focused exclusively on two-hop factual queries rather than arbitrary multi-step or non-factual reasoning tasks. Furthermore, while back-patching proves the root cause of these reasoning failures, it is currently an analytical diagnostic rather than a real-time production inference technique. Confidence in the underlying conclusion—that models fail multi-hop queries due to running out of network depth—remains high due to consistent results across all tested model sizes and families.
- Paper: Do Large Language Models Latently Perform Multi-Hop Reasoning?, Sohee Yang et al. (2024). It establishes whether models retrieve and use intermediate entities during latent two-hop reasoning, providing the direct mechanistic foundation for this paper’s analysis of where that process breaks down.
- Paper: Dissecting Recall of Factual Associations in Auto-Regressive Language Models, Mor Geva et al. (2023). Its account of how factual associations move through attention and feed-forward layers prepares you to follow this paper’s tracing of bridge and target information across model depth.
- Paper: Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task, Alicia Curth et al. (2026). It carries the question of reasoning depth forward by testing whether transformers adapt their layer use to relational task difficulty, extending the concern that fixed depth can constrain multi-hop reasoning.
