Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters
Boshi WangSewon MinXiang DengJiaming ShenYou WuLuke ZettlemoyerHuan Sun
Reveals that chain-of-thought prompting remains highly effective even when demonstrations contain invalid reasoning, demonstrating that query relevance and step ordering drive performance gains rather than the logical validity of exemplar rationales.
Large language models have shown remarkable success on complex tasks through Chain-of-Thought prompting, a technique that supplies step-by-step example rationales within the prompt to encourage explicit multi-step reasoning. Despite its widespread adoption, practitioners and researchers have lacked a clear understanding of why this technique works and which components of the demonstrated reasoning steps are truly essential. Understanding these mechanisms is critical for designing cost-effective prompt engineering pipelines and accurately benchmarking model capabilities.
The article evaluates what makes Chain-of-Thought prompting effective by systematically altering different components of the prompt demonstrations. Specifically, the analysis demonstrates how model performance changes when the logical validity, relevance, and ordering of demonstrated rationales are intentionally degraded.
The authors conducted controlled ablation experiments on two representative multi-step reasoning benchmarks: the GSM8K mathematical reasoning dataset (evaluated on a 800-sample test subset) and the Bamboogle multi-hop factual question answering dataset (125 test questions). They tested several high-capacity models, primarily InstructGPT (text-davinci-002 and text-davinci-003), along with PaLM and Flan-PaLM. The evaluation examined both final output correctness and intermediate intrinsic quality by tracking whether models correctly derived key intermediate entities and numbers.
The key findings reveal that logical validity in prompt demonstrations matters far less than previously assumed. First, providing demonstrations with completely invalid, flawed reasoning steps still enabled models to achieve over 80% to 90% of standard Chain-of-Thought performance, and the models continued to generate coherent, logically sound reasoning during inference. Second, prompt relevance is essential; completely removing query relevance from the demonstrated objects and templates caused performance on mathematical reasoning to collapse from an intermediate score of 48.3% down to 11.9%, falling below basic standard prompting (15.4%). Third, the structural coherence and ordering of the language templates are critical, whereas maintaining the exact sequence of intermediate numbers and entities matters substantially less. Finally, models with extensive pretraining or instruction tuning on similar tasks proved highly resilient to flawed prompt rationales.
These findings imply that large language models do not primarily learn how to reason step-by-step from few-shot in-context demonstrations. Instead, the models already possess underlying reasoning capabilities acquired during large-scale pretraining. Demonstrations mainly serve as formatting guides that activate existing capabilities and enforce structured output. This insight alters operational risk assessments: while prompt engineers do not need to spend excessive effort crafting perfectly verified logical steps, models risk over-relying on pretraining priors and ignoring critical instructions or counterfactual context.
Decision-makers and engineering teams should focus prompt design resources on ensuring strict topic relevance and coherent language framing rather than exhaustive step-by-step manual verification of prompt logic. For benchmarking and evaluation, organizations must develop novel test suites where models have minimal prior exposure, ensuring tests measure true in-context skill acquisition rather than the mere recitation of pretraining data.
These conclusions should be interpreted within the context of specific limitations. The empirical analysis focused on arithmetic and factual question answering benchmarks; highly template-based symbolic tasks were not evaluated due to rigid structures. Furthermore, the experiments relied on single-run evaluations without reporting variance across multiple random seeds, and invalid reasoning steps were written manually rather than generated via a formal algorithmic taxonomy. Nonetheless, the consistent trends observed across diverse model architectures provide high confidence that prompt relevance and structural ordering outweigh step-by-step logical validity in eliciting model reasoning.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces standard Chain-of-Thought prompting, establishing the foundational reasoning paradigm that the source empirically dissects to uncover what components matter.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). Demonstrates that ground-truth labels are non-essential for standard in-context learning, directly inspiring the source's investigation into whether valid intermediate reasoning steps are required in Chain-of-Thought prompting.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Establishes zero-shot Chain-of-Thought reasoning by eliciting step-by-step rationales, providing a critical baseline for comparing prompt-driven intermediate steps against few-shot demonstrations.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Explores sequential subproblem decomposition in prompts, laying groundwork for analyzing the necessity of step relevance and ordering in multi-step demonstrations.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). Evaluates Chain-of-Thought performance across difficult tasks in BIG-Bench Hard, providing the empirical rationale-benchmark settings tested in the source's ablation study.
- Paper: Why think step by step? Reasoning emerges from the locality of experience, Ben Prystawski et al. (2023). Provides a theoretical statistical explanation based on local data dependencies to explain why intermediate steps enhance reasoning even when not strictly ground-truth accurate.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). Extends the insight that reasoning chains can be functionally disconnected from true correctness by showing that generated explanations can be systematically unfaithful and post-hoc rationalizations.
- Paper: Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought, Abulhair Saparov et al. (2023). Formally analyzes step validity versus planning failures in generated chains, continuing the empirical examination of how step order and local validity impact final outputs.
- Paper: ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness, Archiki Prasad et al. (2023). Develops a reference-free framework to evaluate reasoning correctness and step informativeness separately, directly addressing the finding that relevance and ordering matter more than perfect exemplar validity.
- Paper: Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future, Zheng Chu et al. (2024). Synthesizes modern developments in Chain-of-Thought methodologies, placing the source's findings on rationale mechanics into a comprehensive broader taxonomy.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). Introduces metrics for assessing the validity and redundancy of intermediate reasoning steps, moving beyond final-answer accuracy in evaluating rationales.
- Paper: Chain-of-Thought Reasoning Without Prompting, Xuezhi Wang et al. (2024). Investigates how models elicit reasoning paths intrinsically from top alternative tokens without relying on explicit prompt demonstrations.
