Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought
Abulhair SaparovHe He
Introduces the PrOntoQA benchmark to convert chain-of-thought reasoning into formal proofs, demonstrating that while large language models reliably execute single deduction steps, they fail at systematic proof planning when faced with multiple reasoning paths.
Large language models have shown strong performance on reasoning tasks when provided with step-by-step demonstrations, known as chain-of-thought prompting. However, real-world benchmarks make it difficult to determine whether these models genuinely reason or merely retrieve memorized facts and exploit superficial shortcuts. This distinction is critical today as organizations increasingly rely on language models for automated decision-making and complex analytical workflows.
The article investigates whether large language models genuinely reason step by step and pinpoints where their reasoning breaks down. To achieve this, the authors created PRONTOQA, a synthetic question-answering dataset where each problem is generated from a formal logical hierarchy with a single correct proof. By converting each predicted reasoning step into symbolic logic, the authors evaluated multiple versions of GPT-3 and InstructGPT across 48 experimental settings, testing proof lengths from one to five steps, varied sentence orders, and fictional, true, and counterfactual contexts.
The analysis yielded several key findings. First, advanced models execute individual deduction steps with high local validity, with over 93% of generated steps being strictly valid even in fictional settings. Second, models struggle significantly with global proof planning; when faced with multiple valid logical branches, they often follow misleading paths that lead to incomplete proofs and incorrect conclusions. Third, model accuracy degrades as proof length increases, dropping to near random chance on five-step proofs when premise order is inverted. Fourth, models perform substantially better when reasoning with real-world, familiar facts than with fictional or false concepts, showing heavy reliance on pretraining memory. Finally, prompting techniques like self-consistency and depth-first search demonstrations fail to resolve these planning errors, as models often assign higher probability to flawed proofs than to valid ones.
These findings indicate that large language models act as greedy reasoners: they generate locally sensible deductions but lack lookahead capability and systematic search. Deploying standard prompting for complex, multi-step tasks introduces serious operational risks of undetected logical errors. Organizations should avoid relying solely on pure language generation for mission-critical deductions. Instead, decision-makers should consider hybrid systems that combine language models with external symbolic verifiers or formal planning frameworks.
Readers should note that the evaluation is limited to a single deduction rule (modus ponens), short proof lengths up to five steps, and simple sentence structures. While confidence in the finding that language models lack autonomous proof planning is high, further research is required to assess performance across more diverse deduction rules, longer reasoning chains, and richer real-world domains.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It introduces few-shot chain-of-thought prompting in large language models, providing the core reasoning mechanism that the source formally analyzes.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). It demonstrates zero-shot step-by-step reasoning via greedy decoding, establishing the greedy reasoning paradigm scrutinized by the source.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). It establishes multi-path sampling over single greedy decoding paths, highlighting the exact exploration limitations of greedy reasoning analyzed in the source.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). It introduces problem decomposition strategies to address the structural reasoning and generalization failures that the source formalizes.
- Paper: Measuring and Narrowing the Compositionality Gap in Language Models, Ofir Press et al. (2022). It characterizes multi-step compositional failures in language models, setting up the empirical problem of step composition that the source models in first-order logic.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). It benchmarks the multi-step reasoning capabilities and limitations of large language models across complex tasks using chain-of-thought prompting.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). It examines how language models generate reasoning rationales and identifies their tendency to settle into greedy or biased reasoning paths.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). It utilizes the source's PrOntoQA benchmark and directly addresses the source's findings on proof-planning limitations by enabling non-greedy, continuous latent search.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). It overcomes the greedy proof-planning bottleneck identified in the source by implementing tree-based lookahead and backtracking over intermediate thought steps.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). It extends non-linear reasoning beyond linear and tree structures into arbitrary graph networks to remedy the search and planning deficiencies shown in the source.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It builds upon the source's formal inspection of step-by-step rationales to demonstrate how chain-of-thought outputs can systematically unfaithfully reflect internal decision-making.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). It develops meta-reasoning and trial-and-error search patterns to address the core inability of standard greedy chain-of-thought models to explore alternative deduction paths.
- Paper: Chain-of-Thought Reasoning Without Prompting, Xuezhi Wang et al. (2024). It investigates how to uncover alternative valid reasoning paths from base models without explicit prompting, extending the source's analysis of greedy generation dynamics.
- Paper: GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models, Iman Mirzadeh et al. (2025). It applies symbolic templates to expose the fragility of language model chain-of-thought steps, reinforcing the source's formal evaluation methodology.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). It leverages reinforcement learning to incentivize self-verification and autonomous proof search, overcoming the myopic deduction habits documented in the source.
