Why think step by step? Reasoning emerges from the locality of experience
Ben PrystawskiMichael LiNoah D. Goodman
Explains why chain-of-thought reasoning aids language models by proving mathematically and demonstrating empirically that step-by-step generation bridges the gap between locally structured training observations to accurately estimate unseen dependencies between distant variables.
Large language models frequently perform better on complex tasks when prompted to generate step-by-step reasoning before delivering an answer. However, because generating intermediate steps introduces no new data, it has remained unclear why this approach works and what underlying statistical properties make it effective. The article investigates this question to determine whether the advantage of step-by-step reasoning emerges because training data is naturally organized into local, overlapping clusters of closely related concepts.
The main objective of the article is to demonstrate theoretically and experimentally why chain-of-thought reasoning improves inference over direct prediction in autoregressive models. The authors evaluate whether reasoning through intermediate variables reduces statistical bias when models estimate relationships between variables that were rarely or never seen together during training.
To evaluate this hypothesis, the authors first developed a mathematical proof using a risk-minimizing sequence model on a chain-structured network. They then conducted empirical experiments by generating synthetic datasets from 10 distinct 100-variable probabilistic networks featuring strong dependencies. They trained transformer language models from scratch across varying training conditions, including local neighborhoods where data co-occurred based on true dependencies, mismatched controls with incorrect locality, and fully observed datasets. The models were tested on their ability to estimate conditional probabilities for pairs of variables deliberately held out during training, comparing direct prediction against scaffolded and model-generated intermediate reasoning.
The findings reveal that step-by-step reasoning significantly outperforms direct prediction, but only when the training data is structured locally around true statistical dependencies. Under these local conditions, allowing the model to freely generate intermediate variables achieved nearly the same high accuracy as providing an optimal path of intermediate steps, while direct prediction defaulted toward marginal baseline probabilities. Generating irrelevant intermediate variables provided no benefit, showing that reasoning requires valid dependency chains. Furthermore, training on local neighborhoods combined with intermediate reasoning was far more data-efficient than training on complete datasets: the local reasoning approach reached strong accuracy in roughly 200 million training tokens, whereas a model trained directly on fully observed data required about 650 million tokens—over three times as much training—to match that performance. The authors also found that taking multiple samples of generated reasoning paths reduced variance and improved overall accuracy.
These results imply that step-by-step reasoning works because it chains together familiar, highly accurate local relationships to bridge gaps between concepts that were never observed together. For practitioners, this indicates that model performance and training costs can be optimized simultaneously. Rather than attempting to train models on massive, fully exhaustive datasets, curating training data into tightly connected, overlapping thematic clusters can substantially reduce compute requirements while improving complex multi-step inference at runtime.
Based on these findings, teams designing training pipelines should prioritize data curation that preserves dense local topic structure rather than relying solely on raw scale. For deployment, implementing sampling methods that average over multiple reasoning paths is recommended to minimize estimation variance. Moving forward, additional research is needed to examine how these principles apply to more abstract concepts, richer natural language settings, and diverse prompting methods.
Readers should note that the experimental validation relies on synthetic, propositional networks rather than open-ended natural text corpora, and the study focuses primarily on zero-shot reasoning rather than few-shot demonstration prompting. Nevertheless, because the findings are backed by formal mathematical proofs and remain consistent across multiple transformer model sizes, there is high confidence in the core mechanism connecting local training structure to the emergence of effective reasoning.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This foundational chain-of-thought study establishes the empirical phenomenon that intermediate reasoning improves large-language-model performance, which the source explains through local statistical structure.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Its zero-shot chain-of-thought experiments provide an earlier demonstration of step-by-step reasoning gains that the source seeks to ground in the organization of training data.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Its formal analysis of local deduction and global proof-planning failures supplies a complementary account of why intermediate reasoning can help, preparing the reader for the source’s probabilistic explanation.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). This work extends explicit chain-of-thought reasoning into continuous latent computation, testing how the source’s account of intermediate inference can motivate nonverbal reasoning mechanisms.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). Graph of Thoughts generalizes the source’s chain-structured reasoning picture to aggregation, refinement, and feedback across arbitrary dependency graphs.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Self-consistency continues the source’s account of useful intermediate computation by exploiting multiple sampled reasoning paths and aggregating their answers.
