Why think step by step? Reasoning emerges from the locality of experience

Ben PrystawskiMichael LiNoah D. Goodman

article2023NeurIPS168 citationsBest Paper Award

Explains why chain-of-thought reasoning aids language models by proving mathematically and demonstrating empirically that step-by-step generation bridges the gap between locally structured training observations to accurately estimate unseen dependencies between distant variables.

Listen

Large language models frequently perform better on complex tasks when prompted to generate step-by-step reasoning before delivering an answer. However, because generating intermediate steps introduces no new data, it has remained unclear why this approach works and what underlying statistical properties make it effective. The article investigates this question to determine whether the advantage of step-by-step reasoning emerges because training data is naturally organized into local, overlapping clusters of closely related concepts.

The main objective of the article is to demonstrate theoretically and experimentally why chain-of-thought reasoning improves inference over direct prediction in autoregressive models. The authors evaluate whether reasoning through intermediate variables reduces statistical bias when models estimate relationships between variables that were rarely or never seen together during training.

To evaluate this hypothesis, the authors first developed a mathematical proof using a risk-minimizing sequence model on a chain-structured network. They then conducted empirical experiments by generating synthetic datasets from 10 distinct 100-variable probabilistic networks featuring strong dependencies. They trained transformer language models from scratch across varying training conditions, including local neighborhoods where data co-occurred based on true dependencies, mismatched controls with incorrect locality, and fully observed datasets. The models were tested on their ability to estimate conditional probabilities for pairs of variables deliberately held out during training, comparing direct prediction against scaffolded and model-generated intermediate reasoning.

The findings reveal that step-by-step reasoning significantly outperforms direct prediction, but only when the training data is structured locally around true statistical dependencies. Under these local conditions, allowing the model to freely generate intermediate variables achieved nearly the same high accuracy as providing an optimal path of intermediate steps, while direct prediction defaulted toward marginal baseline probabilities. Generating irrelevant intermediate variables provided no benefit, showing that reasoning requires valid dependency chains. Furthermore, training on local neighborhoods combined with intermediate reasoning was far more data-efficient than training on complete datasets: the local reasoning approach reached strong accuracy in roughly 200 million training tokens, whereas a model trained directly on fully observed data required about 650 million tokens—over three times as much training—to match that performance. The authors also found that taking multiple samples of generated reasoning paths reduced variance and improved overall accuracy.

These results imply that step-by-step reasoning works because it chains together familiar, highly accurate local relationships to bridge gaps between concepts that were never observed together. For practitioners, this indicates that model performance and training costs can be optimized simultaneously. Rather than attempting to train models on massive, fully exhaustive datasets, curating training data into tightly connected, overlapping thematic clusters can substantially reduce compute requirements while improving complex multi-step inference at runtime.

Based on these findings, teams designing training pipelines should prioritize data curation that preserves dense local topic structure rather than relying solely on raw scale. For deployment, implementing sampling methods that average over multiple reasoning paths is recommended to minimize estimation variance. Moving forward, additional research is needed to examine how these principles apply to more abstract concepts, richer natural language settings, and diverse prompting methods.

Readers should note that the experimental validation relies on synthetic, propositional networks rather than open-ended natural text corpora, and the study focuses primarily on zero-shot reasoning rather than few-shot demonstration prompting. Nevertheless, because the findings are backed by formal mathematical proofs and remain consistent across multiple transformer model sizes, there is high confidence in the core mechanism connecting local training structure to the emergence of effective reasoning.

Cover for Why think step by step? Reasoning emerges from the locality of experience

Abstract

Humans have a powerful and mysterious capacity to reason. Working through a set of mental steps enables us to make inferences we would not be capable of making directly even though we get no additional data from the world. Similarly, when large language models generate intermediate steps (a chain of thought) before answering a question, they often produce better answers than they would directly. We investigate why and how chain-of-thought reasoning is useful in language models, testing the hypothesis that reasoning is effective when training data consists of overlapping local clusters of variables that influence each other strongly. These training conditions enable the chaining of accurate local inferences to estimate relationships between variables that were not seen together in training. We prove that there will exist a “reasoning gap”, where reasoning through intermediate variables reduces bias, for the simple case of an autoregressive density estimator trained on local samples from a chain-structured probabilistic model. We then test our hypothesis experimentally in more complex models, training an autoregressive language model on samples from Bayes nets but only including a subset of variables in each sample. We test language models’ ability to match conditional probabilities with and without intermediate reasoning steps, finding that intermediate steps are only helpful when the training data is locally structured with respect to dependencies between variables. The combination of locally structured observations and reasoning is much more data-efficient than training on all variables. Our results illustrate how the effectiveness of reasoning step by step is rooted in the local statistical structure of the training data.

Table of Contents

  • 1 Introduction
  • 2 Task setup
  • 2.1 Observation distribution
  • 2.2 Estimators
  • 3 Theoretical results
  • 4 Experimental methods
  • 4.1 Training data
  • 4.2 Estimation
  • 4.3 Model architecture
  • 5 Results
  • 5.1 When reasoning helps
  • 5.2 Data complexity and reasoning
  • 5.3 When reasoning is unnecessary
  • 5.4 When reasoning fails
  • 6 Discussion
  • Acknowledgments and Disclosure of Funding
  • References
  • A Theoretical analysis
  • A.1 Problem Setup
  • A.2 Preliminaries
  • A.3 Main Theorem
  • B Pseudocode for data generation
  • C Full sample of training data
  • D Formatting estimators as prompts
  • D.1 Direct prediction
  • D.2 Scaffolded generation
  • D.3 Free generation
  • E Training details
  • F Comparison of reasoning gaps across different architectures
  • G Mean squared error by number of samples
  • H Data efficiency of fully observed training with no held-out pairs

Knowls

  1. Knowl 1 — Bias Reduction via Scaffolded Intermediate Marginalization in Locally Observed Chains

    theoretical result

    Let a sequence model qq be trained to predict alternating variable indices it∈{1,…,N}i_t \in \{1, \dots, N\} and discrete variable values vt∈Xv_t \in \mathcal{X} drawn from a joint distribution pd(Y1,…,YN)=pd(Y1)∏j=1N−1pd(Yj+1∣Yj)p_d(Y_1, \dots, Y_N) = p_d(Y_1) \prod_{j=1}^{N-1} p_d(Y_{j+1} \mid Y_j) over a directed chain. Suppose training sequences are generated according to an observation distribution pobsp_{\text{obs}} that only assigns non-zero probability to adjacent variable pairs (pobs({i,j})=0p_{\text{obs}}(\{i, j\}) = 0 if ∣i−j∣>1|i-j| > 1), yielding observed sequence distribution pp.

    Let uu be the uniform distribution over valid index-value sequences, and define the regularized risk R(q)=H(p,q)+H(u,q)R(q) = H(p, q) + H(u, q), where H(⋅,⋅)H(\cdot, \cdot) denotes cross entropy. The risk minimizer q∗=arg⁡min⁡qR(q)q^* = \arg\min_q R(q) satisfies:

    1. For adjacent pairs Yi,Yi+1Y_i, Y_{i+1}, q∗(Yi+1∣Yi)=λpd(Yi+1∣Yi)+(1−λ)1∣X∣q^*(Y_{i+1} \mid Y_i) = \lambda p_d(Y_{i+1} \mid Y_i) + (1-\lambda)\frac{1}{|\mathcal{X}|} for some λ∈(0,1)\lambda \in (0, 1).
    2. For non-adjacent pairs Yi,YjY_i, Y_j (∣i−j∣>1|i-j| > 1), q∗(Yi∣Yj)=1∣X∣q^*(Y_i \mid Y_j) = \frac{1}{|\mathcal{X}|}.

    Under the assumption that true conditional probabilities satisfy a doubly stochastic condition ($\

Coverage note — ...

References

  1. 1.R. N. Shepard, “The step to rationality: The efficacy of thought experiments in science, ethics, and free will,” Cognitive Science, vol. 32, no. 1, pp. 3–35, 2008.
  2. 2.D. C. Dennett, Intuition pumps and other tools for thinking. WW Norton & Company, 2013.
  3. 3.T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020.
  4. 4.A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al., “Palm: Scaling language modeling with pathways,” arXiv preprint arXiv:2204.02311, 2022.
  5. 5.H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al., “Llama: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.
  6. 6.B. bench authors, “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” Transactions on Machine Learning Research, 2023.
  7. 7.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al., “Training verifiers to solve math word problems,” arXiv preprint arXiv:2110.14168, 2021.
  8. 8.M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, et al., “Show your work: Scratchpads for intermediate computation with language models,” arXiv preprint arXiv:2112.00114, 2021.
  9. 9.T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” arXiv preprint arXiv:2205.11916, 2022.
  10. 10.J. Wei, X. Wang, D. Schuurmans, M. Bosma, brian ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain of thought prompting elicits reasoning in large language models,” in Advances in Neural Information Processing Systems (A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, eds.), 2022.
  11. 11.D. Dohan, W. Xu, A. Lewkowycz, J. Austin, D. Bieber, R. G. Lopes, Y. Wu, H. Michalewski, R. A. Saurous, J. Sohl-dickstein, et al., “Language model cascades,” arXiv preprint arXiv:2207.10342, 2022.
  12. 12.A. K. Lampinen, I. Dasgupta, S. C. Chan, K. Matthewson, M. H. Tessler, A. Creswell, J. L. McClelland, J. X. Wang, and F. Hill, “Can language models learn from explanations in context?,” arXiv preprint arXiv:2204.02329, 2022.
  13. 13.E. Zelikman, Y. Wu, J. Mu, and N. Goodman, “STaR: Bootstrapping reasoning with reasoning,” in Advances in Neural Information Processing Systems (A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, eds.), 2022.
  14. 14.X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” in The Eleventh International Conference on Learning Representations, 2022.
  15. 15.D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le, et al., “Least-to-most prompting enables complex reasoning in large language models,” arXiv preprint arXiv:2205.10625, 2022.
  16. 16.S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” arXiv preprint arXiv:2305.10601, 2023.
  17. 17.D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent dirichlet allocation,” Journal of machine Learning research, vol. 3, no. Jan, pp. 993–1022, 2003.
  18. 18.D. M. Blei and J. D. Lafferty, “A correlated topic model of Science,” The Annals of Applied Statistics, vol. 1, no. 1, pp. 17 – 35, 2007.
  19. 19.S. C. Chan, A. Santoro, A. K. Lampinen, J. X. Wang, A. K. Singh, P. H. Richemond, J. McClelland, and F. Hill, “Data distributional properties drive emergent in-context learning in transformers,” in Advances in Neural Information Processing Systems (A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, eds.), 2022.
  20. 20.A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019.
  21. 21.T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al., “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pp. 38–45, 2020.
  22. 22.R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 1715–1725, 2016.
  23. 23.D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of the 3rd International Conference on Learning Representations (ICLR), 2015.

Citation

MLA
Prystawski, B., et al. “Why Think Step by Step? Reasoning Emerges from the Locality of Experience”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 70926–47, https://proceedings.neurips.cc/paper_files/paper/2023/file/e0af79ad53a336b4c4b4f7e2a68eb609-Paper-Conference.pdf.
APA
Prystawski, B., Li, M., & Goodman, N. (2023). Why think step by step? Reasoning emerges from the locality of experience. Advances in Neural Information Processing Systems, 36, 70926–70947. https://proceedings.neurips.cc/paper_files/paper/2023/file/e0af79ad53a336b4c4b4f7e2a68eb609-Paper-Conference.pdf
Chicago
Prystawski, B., M. Li, and N. Goodman. 2023. “Why Think Step by Step? Reasoning Emerges from the Locality of Experience”. Advances in Neural Information Processing Systems 36: 70926–47. https://proceedings.neurips.cc/paper_files/paper/2023/file/e0af79ad53a336b4c4b4f7e2a68eb609-Paper-Conference.pdf.
Harvard
Prystawski, B., Li, M. and Goodman, N. (2023) “Why think step by step? Reasoning emerges from the locality of experience”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 70926–70947. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/e0af79ad53a336b4c4b4f7e2a68eb609-Paper-Conference.pdf.
Vancouver
1. Prystawski B, Li M, Goodman N (2023) Why think step by step? Reasoning emerges from the locality of experience. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 70926–70947

BibTeX

@inproceedings{prystawski2023why,
  title = {Why think step by step? Reasoning emerges from the locality of experience},
  author = {Prystawski, Ben and Li, Michael and Goodman, Noah},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {70926-70947},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/e0af79ad53a336b4c4b4f7e2a68eb609-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors