Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models

Yifan HouJiaoda LiYu FeiAlessandro StolfoWangchunshu ZhouGuangtao ZengAntoine BosselutMrinmaya Sachan

article2023EMNLP76 citations

Reveals that language models execute genuine multi-step reasoning rather than simple memorization by introducing MechanisticProbe, a framework that extracts the underlying procedural reasoning trees directly from internal attention patterns.

Listen

Large language models have demonstrated strong capabilities across multi-step reasoning tasks. However, it remains uncertain whether these systems genuinely perform structured, procedural reasoning or simply rely on memorized shortcuts from pretraining data. To address this question, the article evaluates whether language models internally embed a step-by-step reasoning tree corresponding to the ground-truth problem-solving process.

The authors develop MechanisticProbe, a framework that recovers reasoning trees from internal attention patterns using non-parametric nearest-neighbor classifiers. The methodology evaluates two sub-tasks: selecting necessary statements from an input set and determining the hierarchical height of those statements in the reasoning tree. The approach was evaluated on a synthetic numerical ordering task using GPT-2 (across a dataset of roughly one million generated sequences) and two natural language logical reasoning benchmarks—ProofWriter and the AI2 Reasoning Challenge—using the 7-billion-parameter LLaMA model under both few-shot in-context and fine-tuned settings.

The analysis produced several primary findings. First, internal attention patterns clearly encode the underlying reasoning trees, achieving normalized probe scores above 90% on structured tasks when models are fine-tuned. Second, layer-by-layer probing demonstrates that models process reasoning sequentially across their depth: bottom layers immediately filter and isolate relevant statements, while middle and higher layers execute the subsequent reasoning steps. Third, causal pruning experiments revealed that attention heads focused on rank and logical size are critical to performance, where removing just 10% caused severe accuracy drops, whereas removing up to 40% of position-focused heads had little impact. Fourth, strong probe scores strongly correlate with end-to-end task accuracy (with a Pearson correlation of 0.71) and noise tolerance; models showing higher alignment with the reasoning tree maintained robust performance even when presented with corrupted input statements.

These results imply that language models can perform authentic mechanistic reasoning internally rather than merely surface-level pattern matching. This internal procedural execution directly enhances model reliability and robustness against distractors, which is crucial for deploying models in high-stakes compliance, logical deduction, and decision-support applications. Furthermore, fine-tuning substantially improved tree-following fidelity compared to few-shot prompting, which showed vulnerability as the number of irrelevant statements increased.

Organizations developing or deploying language models for complex logical tasks should prioritize targeted fine-tuning and task decomposition to improve internal reasoning reliability. Additionally, internal attention probing can serve as a diagnostic auditing tool to assess whether a model's output is supported by valid procedural logic before putting it into production.

These findings should be interpreted within certain limitations. The evaluation focused primarily on classification-based, single-token outputs with shallow reasoning trees of depth up to one, rather than long-chain autoregressive generation. While confidence in the internal tree structure for shallow procedural tasks is high, further validation on deeper, more complex reasoning chains is recommended before applying these diagnostic methods to larger real-world workflows.

Cover for Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models

Abstract

Recent work has shown that language models (LMs) have strong multi-step (i.e., procedural) reasoning capabilities. However, it is unclear whether LMs perform these tasks by cheating with answers memorized from pretraining corpus, or, via a multi-step reasoning mechanism. In this paper, we try to answer this question by exploring a mechanistic interpretation of LMs for multi-step reasoning tasks. Concretely, we hypothesize that the LM implicitly embeds a reasoning tree resembling the correct reasoning process within it. We test this hypothesis by introducing a new probing approach (called MechanisticProbe) that recovers the reasoning tree from the model’s attention patterns. We use our probe to analyze two LMs: GPT-2 on a synthetic task (k-th smallest element), and LLaMA on two simple language-based reasoning tasks (ProofWriter & AI2 Reasoning Challenge). We show that MechanisticProbe is able to detect the information of the reasoning tree from the model’s attentions for most examples, suggesting that the LM indeed is going through a process of multi-step reasoning within its architecture in many cases.¹

Table of Contents

  • 1 Introduction
  • 2 Reasoning with LM
  • 2.1 Reasoning Formulation
  • 2.2 Reasoning Tasks
  • 3 Mechanistic Probe
  • 3.1 Problem Formulation
  • 3.2 Simplification of A_simp
  • 3.3 Simplification of the Probing Task
  • 3.4 Probing Score
  • 4 Mechanistic Probing of LMs
  • 4.1 Attention Visualization
  • 4.2 Probing Scores
  • 4.3 Layer-wise Probing
  • 5 Do LMs Reason Using A_simp?
  • 6 Correlating Probe Scores with Model Accuracy and Robustness
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Proof of The First-Token Domination
  • B Supplementary about GPT-2
  • B.1 GPT-2 Finetuning
  • B.2 Original Probing Scores
  • B.3 Visualization of E[π(A_simp)] for GPT-2
  • B.4 Visualization of A_simp
  • B.5 Attention Head Entropy
  • B.6 Do Finetuning Methods Matter?
  • B.7 Task Difficulty & Model Capacity
  • B.8 What if Reasoning Tasks Become Harder?
  • C Supplementary about LLaMA
  • C.1 Settings for 4-shot and Finetuned LLaMA
  • C.2 Layer (Attention) Pruning
  • C.3 Statistics of Cleaned ProofWriter and Annotated ARC
  • C.4 Reasoning Tree Ambiguity Example
  • C.5 Original Probing Scores

Knowls

  1. Knowl 1 — MechanisticProbe reconstructs reasoning trees from attention

    model/method

    MechanisticProbe tests whether a language model's attention contains information about the reasoning tree GG used to answer a question. It represents the input as statements S1,…,SnS_1,\ldots,S_n and a question QQ, and reduces the attention information before probing it. For a causal language model, it first uses only attention directed to the final input token, reducing the attention data from all token pairs to one attention value per input token, layer, and head; it then averages across heads. For multi-token statements in LLaMA, it treats each statement as a hypernode, averages attention over the tokens in each statement, takes a maximum over question tokens, and ignores attention within a statement. If A(l,h)[i,j]A^{(l,h)}[i,j] is the attention from query token ii to key token jj at layer ll and head hh, the cross-statement simplification is

    Asimpcross(l,h)[i]=max⁡tj′∈Q(1∣Si∣∑ti′∈SiA(l,h)[i′,j′]),A^{\mathrm{cross}}_{\mathrm{simp}}(l,h)[i] =\max_{t_{j'}\in Q}\left(\frac{1}{|S_i|}\sum_{t_{i'}\in S_i} A^{(l,h)}[i',j']\right),

    where SiS_i is statement ii, ∣Si∣|S_i| is its number of tokens, and tj′t_{j'} ranges over question tokens. After simplification, two kk-nearest-neighbor classifiers probe the attention features: one predicts whether each input statement is useful to the reasoning tree, and the other predicts the height of each useful statement in that tree. The second prediction is conditioned on the gold set of useful statements, allowing the method to separately assess statement selection and reasoning-step identification.

  2. Knowl 2 — Random-baseline-normalized probing scores quantify tree information

    equation

    MechanisticProbe reports separate scores for identifying useful statements and for assigning reasoning-tree heights. Let VV be the set of useful statement nodes in the gold tree GG, AsimpA_{\mathrm{simp}} the simplified attention features, and ArandA_{\mathrm{rand}} the corresponding features from a randomly initialized language model. Let F1(⋅)F_1(\cdot) denote macro-F1 for the specified classification task. The scores are

    SP1=F1(V∣Asimp)−F1(V∣Arand)1−F1(V∣Arand),SP2=F1(G∣V,Asimp)−F1(G∣V,Arand)1−F1(G∣V,Arand).S_{P1}=\frac{F_1(V\mid A_{\mathrm{simp}})-F_1(V\mid A_{\mathrm{rand}})}{1-F_1(V\mid A_{\mathrm{rand}})},\qquad S_{P2}=\frac{F_1(G\mid V,A_{\mathrm{simp}})-F_1(G\mid V,A_{\mathrm{rand}})}{1-F_1(G\mid V,A_{\mathrm{rand}})}.

    The first score measures attention-based classification of useful versus irrelevant statements; the second measures classification of node heights when the useful statements are supplied. The random-attention score is the control baseline, and normalization places each score in [0,1][0,1]: values near zero indicate little tree information beyond the control, while larger values indicate more detectable information.

  3. Knowl 3 — Evaluation covers synthetic selection and two language reasoning tasks

    experimental setup

    The study evaluates GPT-2 on a synthetic task and 7B LLaMA on ProofWriter and the AI2 Reasoning Challenge (ARC). In the synthetic task, an input list contains m=16m=16 randomly selected numbers, each representable as one GPT-2 token; the model predicts the kk-th smallest number. A separate GPT-2 is fine-tuned for each kk. The models are trained on 0.98 million generated sequences for two epochs, with batch size 256, AdamW, and learning rate 10−610^{-6}; 10,000 examples are used for validation and testing. The analysis compares pretrained GPT-2 with its fine-tuned counterpart.

    For ProofWriter, the model classifies the question as true or false. The analysis excludes examples with multiple annotated reasoning trees and limits the main analysis to depths 0 and 1, representing about 70% of the data. ARC consists of multiple-choice science questions; because reasoning-tree annotations are limited, the analysis uses the annotated subset and reports the four-shot in-context setting rather than fine-tuned LLaMA. For the language tasks, the paper also evaluates four-shot LLaMA and a supervised LLaMA variant partially fine-tuned on attention parameters. The reasoning trees in the analyzed tasks have multiple leaves but at most one node at each nonzero height.

  4. Knowl 4 — Fine-tuned GPT-2 exposes the synthetic reasoning tree

    data/table

    On the m=16m=16 kk-th-smallest task, pretrained GPT-2 has 0.00% test accuracy across the reported kk values and low tree-probing scores. Fine-tuned GPT-2 achieves high task accuracy and high normalized probe scores for both selecting the top-kk statements and identifying the kk-th smallest among them. The scores below are percentages; the k=1k=1 second-stage scores are omitted because the tree has depth zero and that classification is trivial.

    Could not parse LaTeX table

    The pattern indicates that successful fine-tuning is associated with attention features from which both parts of the stipulated tree can be recovered. The attention heatmaps on page 6 provide a visual counterpart: lower layers emphasize the top-kk numbers, while later layers emphasize the answer number.

  5. Knowl 5 — LLaMA attention contains useful-statement and reasoning-step information

    data/table

    For ProofWriter, normalized probing scores show that statement selection becomes harder as the number of input statements increases, whereas reasoning-height prediction remains strong. The table reports scores as percentages for four-shot LLaMA and the attention-parameter-fine-tuned variant (LLaMAFT). For depth 0 there is no height-prediction task; for depth 1 with two statements, all statements are useful, so the useful-statement task is trivial. ARC results are for four-shot LLaMA only. Dashes mark tasks that are absent or trivial.

    Could not parse LaTeX table

    The scores exceed the random-baseline control, showing detectable reasoning-tree information in attention. On ProofWriter, LLaMAFT has higher height-prediction scores than four-shot LLaMA at every reported nontrivial statement count; four-shot statement-selection scores weaken notably as irrelevant statements accumulate. ARC statement-selection scores are high, while its height-prediction scores are lower than ProofWriter's.

  6. Knowl 6 — Layer-wise probes indicate statement selection precedes later reasoning

    empirical result

    Cumulative layer-wise probing shows a staged pattern in both models: attention features identify useful statements early, while features for the next reasoning step become informative later. For fine-tuned GPT-2 on the kk-th-smallest task, the useful-statement score rises quickly in the bottom layers, whereas the score for selecting the answer from those statements does not become high until about layer 10 of the 12-layer model. For four-shot LLaMA on ProofWriter and ARC, useful-statement selection reaches a plateau around layer 2, while reasoning-height scores continue to rise through middle layers. On ProofWriter, height-specific probes likewise identify height-0 statements in bottom layers and height-1 statements later. The layer-wise plots on pages 7–8 visualize these separate score trajectories.

  7. Knowl 7 — Pruning size-sensitive heads reduces task accuracy

    empirical result

    To test whether attention heads identified by the probe contribute to solving the synthetic task, the authors group GPT-2 heads by attention-distribution entropy over number rank (size entropy) or input position (position entropy), then prune heads in ascending entropy order. Low size entropy means a head concentrates on particular number ranks; low position entropy means it concentrates on particular input locations. Removing 10% of the low-size-entropy heads causes a significant accuracy drop. By contrast, removing 40% of the low-position-entropy heads has little effect, and for k=1k=1 the model retains high accuracy even after 90% of those position-focused heads are removed. The pruning curves on page 8 show the distinction. This intervention supports the interpretation that rank-sensitive heads contribute to the computation, while position-sensitive heads are comparatively redundant.

  8. Knowl 8 — Reasoning-step probe scores track accuracy and noise tolerance

    empirical result

    Across 2,048 repeated samples of four-shot LLaMA on ProofWriter, test accuracy has a stronger Pearson correlation with the reasoning-height score SP2S_{P2} than with the useful-statement score SP1S_{P1}. The reported correlations, expressed as percentages of the coefficient, are ρ(accuracy,SP2)=71.13%\rho(\text{accuracy},S_{P2})=71.13\%, ρ(accuracy,SP1)=27.42%\rho(\text{accuracy},S_{P1})=27.42\%, and ρ(SP1,SP2)=0.01%\rho(S_{P1},S_{P2})=0.01\%. This suggests that accurate identification of reasoning steps is more closely associated with correct answers in this analysis.

    For a robustness analysis, one irrelevant statement per example is corrupted by appending a negation, and the change in accuracy is measured. The plotted results on page 9 show that examples with SP2<0.7S_{P2}<0.7 lose around 10% accuracy after corruption, whereas examples with higher SP2S_{P2} show an accuracy increase of around 4% in the reported change measure. The authors interpret higher reasoning-step scores as associated with greater tolerance to irrelevant-input noise; the correlations are observational and do not by themselves establish that the score causes robustness.

  9. Knowl 9 — Capacity and task difficulty affect multi-step synthetic performance

    empirical result

    In additional GPT-2 experiments, the number list is extended to m=64m=64 and kk ranges up to 32. Larger GPT-2 variants perform better as the number of leaves in the reasoning tree grows, although smaller models can also handle larger trees when the task is made easier by drawing the 64 numbers from a larger pool (256, 384, or 512 possible numbers). On the harder m=64m=64 tasks, test accuracy falls from nearly 100% toward roughly 15% as kk increases. The useful-statement score SP1S_{P1} remains above 30% even at low accuracy, while the reasoning-step score SP2S_{P2} falls to approximately zero around k=20k=20. These results are consistent with useful-number selection persisting after the later step—choosing the requested number among them—has become unreliable. The capacity and difficulty curves are plotted on pages 16–17.

  10. Knowl 10 — Evidence is limited to simple trees, attention pooling, and classification tasks

    limitation

    The analysis covers relatively simple reasoning tasks and tree structures; ProofWriter examples with depth greater than 1 are excluded from the main analysis because alternative valid reasoning trees can make the annotated path ambiguous. The probe averages attention across heads, which may conceal specialized head functions, especially in shallow, wide models. The evaluated tasks are cast as single-token classification or prediction, so the method does not establish how language models reason in autoregressive chain-of-thought settings that require generating multi-token reasoning traces. Consequently, the reported attention patterns support mechanistic reasoning in the studied settings but do not establish that the same process generalizes to more complex reasoning tasks or architectures.

Coverage note — Supplementary visualizations of individual head entropies and alternative GPT-2 fine-tuning methods are omitted because they do not add a distinct central result beyond the reported pruning evidence and the probe's observed robustness.

References

  1. 1.Samira Abnar and Willem Zuidema. 2020. Quantifying attention flow in transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197, Online. Association for Computational Linguistics.
  2. 2.Deniz Bayazit, Negar Foroutan, Zeming Chen, Gail Weiss, and Antoine Bosselut. 2023. Discovering knowledge-critical subnetworks in pretrained language models.
  3. 3.Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219.
  4. 4.Adrien Bibal, Rémi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas François, and Patrick Watrin. 2022. Is attention explanation? an introduction to the debate. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3889–3900, Dublin, Ireland. Association for Computational Linguistics.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. CoRR, abs/2005.14165.
  6. 6.Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020. On identifiability in transformers. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  7. 7.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang. 2022. Quantifying memorization across neural language models. CoRR, abs/2202.07646.
  8. 8.Hila Chefer, Shir Gur, and Lior Wolf. 2021. Generic attention-model explainability for interpreting bimodal and encoder-decoder transformers. In 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021, pages 387–396. IEEE.
  9. 9.Zeming Chen, Gail Weiss, Eric Mitchell, Asli Celikyilmaz, and Antoine Bosselut. 2023. RECKONING: reasoning through dynamic knowledge encoding. CoRR, abs/2305.06349.
  10. 10.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the AI2 reasoning challenge. CoRR, abs/1803.05457.
  11. 11.Antonia Creswell and Murray Shanahan. 2022. Faithful reasoning using large language models. CoRR, abs/2208.14271.
  12. 12.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2022. Selection-inference: Exploiting large language models for interpretable logical reasoning. CoRR, abs/2205.09712.
  13. 13.Joseph F. DeRose, Jiayao Wang, and Matthew Berger. 2021. Attention flows: Analyzing and comparing attention mechanisms in language models. IEEE Trans. Vis. Comput. Graph., 27(2):1160–1170.
  14. 14.Yue Dong, Chandra Bhagavatula, Ximing Lu, Jena D. Hwang, Antoine Bosselut, Jackie Chi Kit Cheung, and Yejin Choi. 2021. On-the-fly attention modulation for neural generation. In Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pages 1261–1274. Association for Computational Linguistics.
  15. 15.Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Peter West, Chandra Bhagavatula, Ronan Le Bras, Jena D. Hwang, Soumya Sanyal, Sean Welleck, Xiang Ren, Allyson Ettinger, Zaïd Harchaoui, and Yejin Choi. 2023. Faith and fate: Limits of transformers on compositionality. CoRR, abs/2305.18654.
  16. 16.Oliver Eberle, Stephanie Brandl, Jonas Pilot, and Anders Søgaard. 2022. Do transformer models show similar attention patterns to task-specific human gaze? In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4295–4309, Dublin, Ireland. Association for Computational Linguistics.
  17. 17.Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175.
  18. 18.Kawin Ethayarajh and Dan Jurafsky. 2021. Attention flows are shapley value explanations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 49–54, Online. Association for Computational Linguistics.
  19. 19.Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023. Dissecting recall of factual associations in auto-regressive language models. CoRR, abs/2304.14767.
  20. 20.John Hewitt and Percy Liang. 2019. Designing and interpreting probes with control tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2733–2743, Hong Kong, China. Association for Computational Linguistics.
  21. 21.Yifan Hou and Mrinmaya Sachan. 2021. Bird’s eye: Probing for linguistic graph structures with a simple information-theoretic approach. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1844–1859, Online. Association for Computational Linguistics.
  22. 22.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. In NeurIPS.
  23. 23.Karim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau, and Ryan Cotterell. 2022. Probing for the usage of grammatical number. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8818–8831, Dublin, Ireland. Association for Computational Linguistics.
  24. 24.Yibing Liu, Haoliang Li, Yangyang Guo, Chenqi Kong, Jing Li, and Shiqi Wang. 2022. Rethinking attention-model explainability through faithfulness violation test. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pages 13807–13824. PMLR.
  25. 25.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  26. 26.Donald W. Loveland. 1980. Automated theorem proving. a logical basis. Journal of Symbolic Logic, 45(3):629–630.
  27. 27.Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020. Emergent linguistic structure in artificial neural networks trained by self-supervision. Proc. Natl. Acad. Sci. USA, 117(48):30046–30054.
  28. 28.Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2023. Language models implement simple word2vec-style vector arithmetic.
  29. 29.Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher Manning. 2023. Grokking of hierarchical structure in vanilla transformers. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 439–448, Toronto, Canada. Association for Computational Linguistics.
  30. 30.Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023. Progress measures for grokking via mechanistic interpretability. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  31. 31.Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020. Zoom in: An introduction to circuits. Distill, 5(3):e00024–001.
  32. 32.Chris Olah, Alexander Mordvintsev, and Ludwig Schubert. 2017. Feature visualization. Distill, 2(11):e7.
  33. 33.Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev. 2018. The building blocks of interpretability. Distill, 3(3):e10.
  34. 34.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022. In-context learning and induction heads. arXiv preprint arXiv:2209.11895.
  35. 35.Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi, Nouha Dziri, et al. 2023. Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with hypothesis refinement. arXiv preprint arXiv:2310.08559.
  36. 36.Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog.
  37. 37.Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2022. Toward transparent AI: A survey on interpreting the inner structures of deep neural networks. CoRR, abs/2207.13243.
  38. 38.Abhilasha Ravichander, Yonatan Belinkov, and Eduard Hovy. 2021. Probing the probing paradigm: Does probing accuracy entail task relevance? In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 3363–3377, Online. Association for Computational Linguistics.
  39. 39.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot numerical reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 840–854, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  40. 40.Danilo Neves Ribeiro, Shen Wang, Xiaofei Ma, Henghui Zhu, Rui Dong, Deguang Kong, Juliette Burger, Anjelica Ramos, Zhiheng Huang, William Yang Wang, George Karypis, Bing Xiang, and Dan Roth. 2023. STREET: A multi-task structured reasoning and explanation benchmark. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net.
  41. 41.Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020. A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8:842–866.
  42. 42.Sebastian Ruder, Jonas Pfeiffer, and Ivan Vulic. 2022. Modular and parameter-efficient fine-tuning for NLP models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts, pages 23–29, Abu Dubai, UAE. Association for Computational Linguistics.
  43. 43.Jules Samaran, Noa Garcia, Mayu Otani, Chenhui Chu, and Yuta Nakashima. 2021. Attending self-attention: A case study of visually grounded supervision in vision-and-language transformers. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: Student Research Workshop, pages 81–86, Online. Association for Computational Linguistics.
  44. 44.Alessandro Stolfo, Yonatan Belinkov, and Mrinmaya Sachan. 2023. Understanding arithmetic reasoning in language models using causal mediation analysis. arXiv preprint arXiv:2305.15054.
  45. 45.Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. 2021. ProofWriter: Generating implications, proofs, and abductive statements over natural language. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021, pages 3621–3634, Online. Association for Computational Linguistics.
  46. 46.Xiaojuan Tang, Zilong Zheng, Jiaqi Li, Fanxu Meng, Song-Chun Zhu, Yitao Liang, and Muhan Zhang. 2023. Large language models are in-context semantic reasoners rather than symbolic reasoners. CoRR, abs/2305.14825.
  47. 47.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  48. 48.Jesse Vig and Yonatan Belinkov. 2019. Analyzing the structure of attention in a transformer language model. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 63–76, Florence, Italy. Association for Computational Linguistics.
  49. 49.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797–5808, Florence, Italy. Association for Computational Linguistics.
  50. 50.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS.
  51. 51.Zhengxuan Wu, Atticus Geiger, Christopher Potts, and Noah D. Goodman. 2023. Interpretability at scale: Identifying causal mechanisms in alpaca. CoRR, abs/2305.08809.
  52. 52.Shizhuo Dylan Zhang, Curt Tigges, Stella Biderman, Maxim Raginsky, and Talia Ringer. 2023. Can transformers learn to solve problems recursively? CoRR, abs/2305.14699.

Citation

MLA
Hou, Y., et al. “Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 4902–19, https://doi.org/10.18653/v1/2023.emnlp-main.299.
APA
Hou, Y., Li, J., Fei, Y., Stolfo, A., Zhou, W., Zeng, G., Bosselut, A., & Sachan, M. (2023). Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4902–4919. https://doi.org/10.18653/v1/2023.emnlp-main.299
Chicago
Hou, Y., J. Li, Y. Fei, et al. 2023. “Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 4902–19. https://doi.org/10.18653/v1/2023.emnlp-main.299.
Harvard
Hou, Y. et al. (2023) “Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4902–4919. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.299.
Vancouver
1. Hou Y, Li J, Fei Y, Stolfo A, Zhou W, Zeng G, Bosselut A, Sachan M (2023) Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4902–4919

BibTeX

@inproceedings{hou-etal-2023-towards,
    title = "Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models",
    author = "Hou, Yifan  and
      Li, Jiaoda  and
      Fei, Yu  and
      Stolfo, Alessandro  and
      Zhou, Wangchunshu  and
      Zeng, Guangtao  and
      Bosselut, Antoine  and
      Sachan, Mrinmaya",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.299/",
    doi = "10.18653/v1/2023.emnlp-main.299",
    pages = "4902--4919"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/