Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models
Yifan HouJiaoda LiYu FeiAlessandro StolfoWangchunshu ZhouGuangtao ZengAntoine BosselutMrinmaya Sachan
Reveals that language models execute genuine multi-step reasoning rather than simple memorization by introducing MechanisticProbe, a framework that extracts the underlying procedural reasoning trees directly from internal attention patterns.
Large language models have demonstrated strong capabilities across multi-step reasoning tasks. However, it remains uncertain whether these systems genuinely perform structured, procedural reasoning or simply rely on memorized shortcuts from pretraining data. To address this question, the article evaluates whether language models internally embed a step-by-step reasoning tree corresponding to the ground-truth problem-solving process.
The authors develop MechanisticProbe, a framework that recovers reasoning trees from internal attention patterns using non-parametric nearest-neighbor classifiers. The methodology evaluates two sub-tasks: selecting necessary statements from an input set and determining the hierarchical height of those statements in the reasoning tree. The approach was evaluated on a synthetic numerical ordering task using GPT-2 (across a dataset of roughly one million generated sequences) and two natural language logical reasoning benchmarks—ProofWriter and the AI2 Reasoning Challenge—using the 7-billion-parameter LLaMA model under both few-shot in-context and fine-tuned settings.
The analysis produced several primary findings. First, internal attention patterns clearly encode the underlying reasoning trees, achieving normalized probe scores above 90% on structured tasks when models are fine-tuned. Second, layer-by-layer probing demonstrates that models process reasoning sequentially across their depth: bottom layers immediately filter and isolate relevant statements, while middle and higher layers execute the subsequent reasoning steps. Third, causal pruning experiments revealed that attention heads focused on rank and logical size are critical to performance, where removing just 10% caused severe accuracy drops, whereas removing up to 40% of position-focused heads had little impact. Fourth, strong probe scores strongly correlate with end-to-end task accuracy (with a Pearson correlation of 0.71) and noise tolerance; models showing higher alignment with the reasoning tree maintained robust performance even when presented with corrupted input statements.
These results imply that language models can perform authentic mechanistic reasoning internally rather than merely surface-level pattern matching. This internal procedural execution directly enhances model reliability and robustness against distractors, which is crucial for deploying models in high-stakes compliance, logical deduction, and decision-support applications. Furthermore, fine-tuning substantially improved tree-following fidelity compared to few-shot prompting, which showed vulnerability as the number of irrelevant statements increased.
Organizations developing or deploying language models for complex logical tasks should prioritize targeted fine-tuning and task decomposition to improve internal reasoning reliability. Additionally, internal attention probing can serve as a diagnostic auditing tool to assess whether a model's output is supported by valid procedural logic before putting it into production.
These findings should be interpreted within certain limitations. The evaluation focused primarily on classification-based, single-token outputs with shallow reasoning trees of depth up to one, rather than long-chain autoregressive generation. While confidence in the internal tree structure for shallow procedural tasks is high, further validation on deeper, more complex reasoning chains is recommended before applying these diagnostic methods to larger real-world workflows.
- Paper: A Structural Probe for Finding Syntax in Word Representations, John Hewitt et al. (2019). Its structural-probing approach provides a useful methodological precedent for asking whether hierarchical structure can be recovered from language-model representations.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Its chain-of-thought results establish the step-by-step reasoning paradigm that this paper investigates inside model attention patterns.
- Paper: Do Large Language Models Latently Perform Multi-Hop Reasoning?, Sohee Yang et al. (2024). It extends mechanistic analysis of multi-step reasoning by tracing whether models recall and use intermediate entities during latent two-hop inference.
- Paper: Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries, Eden Biran et al. (2024). It carries internal reasoning analysis into multi-hop factual queries, tracing layer-wise information flow and identifying where sequential reasoning fails.
