Do Large Language Models Latently Perform Multi-Hop Reasoning?
Sohee YangElena GribovskayaNora KassnerMor GevaSebastian Riedel
Investigates whether large language models internally connect factual knowledge across multi-hop prompts, revealing that while models frequently recall intermediate bridge entities, their ability to utilize that recalled information for the final reasoning step remains context-dependent and fails to scale with model size.
Large Language Models often complete complex prompts that require combining multiple facts, such as answering "The mother of the singer of 'Superstition' is" without explicit intermediate steps. However, it remains unclear whether models solve these queries by latently traversing intermediate reasoning steps—first recalling the intermediate "bridge" entity Stevie Wonder and then retrieving his mother—or merely by relying on memorized text patterns. Understanding how internal multi-step reasoning operates is essential for evaluating whether model scaling improves real reasoning capabilities and determining whether targeted updates to base facts will reliably propagate across dependent knowledge.
The article evaluates whether and how frequently large language models internally execute latent two-hop reasoning during inference. It specifically examines the degree to which models internally recall intermediate entities and subsequently utilize knowledge about those entities to determine final prompt completions.
To assess these mechanisms, the researchers introduced a benchmark dataset consisting of 45,595 two-hop prompts covering 52 distinct composition types derived from factual relational data. They evaluated three open foundation models ranging from 7 billion to 70 billion parameters (LLaMA-2 7B, 13B, and 70B). The analysis measured two core components across individual model layers: an internal entity recall score that tracks hidden activation projections to intermediate entities, and a consistency score that measures how closely the output matches a direct single-hop query about the intermediate entity. The authors applied intervention and causal patching techniques to determine whether enhancing intermediate entity recall directly improved final output consistency.
The investigation produced four central findings. First, models demonstrate strong internal recognition of the first reasoning step: in approximately 70% to 78% of test cases, models successfully identified the intermediate bridge entity, with recall improving markedly as model size scaled from 7B to 70B parameters. Second, the execution of the second reasoning step was substantially weaker, with increased intermediate recall improving output consistency in roughly 60% to 65% of cases. Third, unlike the first step, this second step showed no positive scaling trend, remaining flat across the 7B, 13B, and 70B models. Finally, complete end-to-end multi-hop reasoning occurred in about 38% to 46% of cases overall, although performance was highly contextual, exceeding 80% success rates in up to 23% of specific relation categories.
These findings suggest that while modern language models can successfully retrieve intermediate entities internally, they frequently struggle to route that retrieved information into the next step of inference. Simply increasing model size improves initial factual recognition but fails to resolve the bottleneck in chaining knowledge together. This explains why standard parameter scaling alone has not resolved compositional reasoning failures. Furthermore, this structural limitation indicates substantial risk for knowledge editing and factual maintenance strategies, as updating a fundamental fact in a model will generally fail to automatically update multi-step dependent outputs.
Organizations developing or deploying large language models should not rely solely on parameter scaling to achieve reliable implicit multi-step reasoning. Instead, technical roadmaps should focus on architectures with explicit multi-step reasoning mechanisms, improved pretraining data structures, and targeted objective functions that encourage internal knowledge routing. Furthermore, for mission-critical applications requiring multi-step factual integrity, leaders should implement explicit prompt-based reasoning workflows (such as chain-of-thought methods) or external retrieval rather than expecting implicit parameter-based inference to remain consistent.
These conclusions are bounded by specific methodological conditions. The analysis focused on two-hop factual associations within the LLaMA-2 model family and tracked single-layer latent pathways, which may represent a conservative lower bound of more distributed internal reasoning. Despite these boundaries, the large sample size and consistent empirical trends across model scales provide high confidence that current standard architectures face a persistent bottleneck in executing multi-step latent reasoning.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). It conceptualizes and measures the compositionality gap across model scales, establishing the foundational behavioral phenomenon that the source investigates mechanistically at the latent layer level.
- Paper: MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions, Zexuan Zhong et al. (2023). It reveals that model editing fails to propagate across multi-hop factual chains, providing essential empirical motivation for the source's study of internal knowledge routing.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It introduces explicit step-by-step chain-of-thought reasoning, providing the core baseline against which implicit, latent multi-step reasoning is framed and evaluated.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). It directly addresses the limitations of standard autoregressive generation by training language models to explicitly execute multi-step reasoning within continuous latent hidden states.
- Paper: Chain-of-Thought Reasoning Without Prompting, Xuezhi Wang et al. (2024). It investigates how latent intermediate reasoning trajectories can be decoded directly from base model representations without explicit prompt interventions.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). It builds on the structural bottlenecks of implicit reasoning by formalizing methods to teach language models deliberate internal search and reflection dynamics.
