Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Abulhair SaparovHe He

article2023ICLR461 citations

Introduces the PrOntoQA benchmark to convert chain-of-thought reasoning into formal proofs, demonstrating that while large language models reliably execute single deduction steps, they fail at systematic proof planning when faced with multiple reasoning paths.

Listen

Large language models have shown strong performance on reasoning tasks when provided with step-by-step demonstrations, known as chain-of-thought prompting. However, real-world benchmarks make it difficult to determine whether these models genuinely reason or merely retrieve memorized facts and exploit superficial shortcuts. This distinction is critical today as organizations increasingly rely on language models for automated decision-making and complex analytical workflows.

The article investigates whether large language models genuinely reason step by step and pinpoints where their reasoning breaks down. To achieve this, the authors created PRONTOQA, a synthetic question-answering dataset where each problem is generated from a formal logical hierarchy with a single correct proof. By converting each predicted reasoning step into symbolic logic, the authors evaluated multiple versions of GPT-3 and InstructGPT across 48 experimental settings, testing proof lengths from one to five steps, varied sentence orders, and fictional, true, and counterfactual contexts.

The analysis yielded several key findings. First, advanced models execute individual deduction steps with high local validity, with over 93% of generated steps being strictly valid even in fictional settings. Second, models struggle significantly with global proof planning; when faced with multiple valid logical branches, they often follow misleading paths that lead to incomplete proofs and incorrect conclusions. Third, model accuracy degrades as proof length increases, dropping to near random chance on five-step proofs when premise order is inverted. Fourth, models perform substantially better when reasoning with real-world, familiar facts than with fictional or false concepts, showing heavy reliance on pretraining memory. Finally, prompting techniques like self-consistency and depth-first search demonstrations fail to resolve these planning errors, as models often assign higher probability to flawed proofs than to valid ones.

These findings indicate that large language models act as greedy reasoners: they generate locally sensible deductions but lack lookahead capability and systematic search. Deploying standard prompting for complex, multi-step tasks introduces serious operational risks of undetected logical errors. Organizations should avoid relying solely on pure language generation for mission-critical deductions. Instead, decision-makers should consider hybrid systems that combine language models with external symbolic verifiers or formal planning frameworks.

Readers should note that the evaluation is limited to a single deduction rule (modus ponens), short proof lengths up to five steps, and simple sentence structures. While confidence in the finding that language models lack autonomous proof planning is high, further research is required to assess performance across more diverse deduction rules, longer reasoning chains, and richer real-world domains.

Cover for Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought

Abstract

Large language models (LLMs) have shown remarkable reasoning capabilities given chain-of-thought prompts (examples with intermediate reasoning steps). Existing benchmarks measure reasoning ability indirectly, by evaluating accuracy on downstream tasks such as mathematical reasoning. However, it is unclear how these models obtain the answers and whether they rely on simple heuristics rather than the generated chain-of-thought. To enable systematic exploration of the reasoning ability of LLMs, we present a new synthetic question-answering dataset called PrOntoQA, where each example is generated from a synthetic world model represented in first-order logic. This allows us to parse the generated chain-of-thought into symbolic proofs for formal analysis. Our analysis on InstructGPT and GPT-3 shows that LLMs are quite capable of making correct individual deduction steps, and so are generally capable of reasoning, even in fictional contexts. However, they have difficulty with proof planning: When multiple valid deduction steps are available, they are not able to systematically explore the different options.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 PrOntoQA: A synthetic dataset for logical reasoning
  • 4 Formal analysis of predicted proofs
  • 5 Results
  • 5.1 Experimental setup
  • 5.2 Do correct answers imply correct reasoning?
  • 5.3 Proof analysis results
  • 5.4 What leads to a mistake?
  • 6 Conclusion and future work
  • References
  • A Appendix
  • A.1 Deduction rules
  • A.2 Avoiding shortcuts
  • A.3 Example InstructGPT misprediction
  • A.4 How we evaluate the chain-of-thought
  • A.5 Proof accuracy vs model size
  • A.6 Additional error analysis
  • A.7 Do other prompting strategies help?
  • A.7.1 Self-consistency
  • A.7.2 Can the model learn to do depth-first search from in-context examples?

Knowls

  1. Knowl 1 — PRONTOQA Synthetic Dataset Generation Framework

    model/method

    PRONTOQA (Proof and Ontology-Generated Question-Answering) is a synthetic framework designed to evaluate chain-of-thought (CoT) reasoning in large language models by generating question-answering examples with verifiable symbolic proofs.

    Each example in PRONTOQA is synthesized through a four-stage process:

    1. Ontology Generation: A hierarchical ontology is randomly generated as a tree structure of concepts (varying in size from 3 to 10 concepts) with subtype implications ∀x(f(x)→g(x))\forall x (f(x) \to g(x)) (e.g., ∀x(cat(x)→carnivore(x))\forall x (\text{cat}(x) \to \text{carnivore}(x))) and node properties (e.g., ∀x(mammal(x)→¬cold_blooded(x))\forall x (\text{mammal}(x) \to \neg \text{cold\_blooded}(x))). The ontology trees are constrained to be linear, where each node has at most one child.

    2. Proof Generation: A starting concept node is chosen uniformly at random, establishing an initial axiom f(a)f(a) asserting that an individual entity aa belongs to type ff (e.g., cat(fae)\text{cat}(\text{fae})). The proof is generated by iteratively applying modus ponens upward along the ontology tree until reaching a target concept or property matching a specified proof depth (1, 3, or 5 hops).

    3. Context Verbalization: Ontology edges and node properties are converted into natural language sentences using fixed template grammars (e.g., "Every cat is a carnivore", "All mammals are not cold-blooded"). The presentation order of context sentences is controlled as either top-down (preorder traversal of the ontology tree) or bottom-up (postorder traversal, mirroring the proof sequence).

    4. Query and Chain-of-Thought Construction: The initial axiom is verbalized (e.g., "Fae is a cat."). A query asks whether the proof conclusion or its negation is true with probability 0.5 (e.g., "True or false: Fae is not herbivorous."). The gold CoT consists of the step-by-step verbalizations of the intermediate modus ponens conclusions leading to the final "True" or "False" label.

    To prevent language models from exploiting surface heuristics—such as predicting truth solely based on whether the queried predicate appears in the context—a distractor sentence is inserted at a random context position. The distractor introduces a disconnected concept assigned the negation of the queried property (e.g., "Every insect is not a vertebrate").

  2. Knowl 2 — Formal Taxonomy of Chain-of-Thought Proof Steps

    definition

    In the formal analysis of natural language chains-of-thought parsed into first-order logic, individual reasoning steps (where each step consists of premises and a derived conclusion) are classified along three orthogonal dimensions:

    1. Validity: Evaluates whether the step's conclusion logically follows from previously established premises or axioms.

      • Strictly-valid: Provable using strictly the rules defined in the base proof calculus (axiom assumption and universal modus ponens: ∀x(f(x)→g(x))\forall x (f(x) \to g(x)) and f(a)⊢g(a)f(a) \vdash g(a)).
      • Broadly-valid: Provable via standard natural deduction rules not present in the base calculus, specifically hypothetical syllogism / transitivity: ∀x(f(x)→g(x))\forall x (f(x) \to g(x)) and ∀x(g(x)→h(x))⊢∀x(f(x)→h(x))\forall x (g(x) \to h(x)) \vdash \forall x (f(x) \to h(x)).
      • Invalid: The conclusion cannot be proven from previous premises or axioms under the system.
    2. Atomicity: Evaluates the granularity of inference in strictly-valid steps.

      • Atomic: The step applies exactly one deduction rule.
      • Non-atomic: The step skips intermediate deduction steps, proving a multi-step conclusion in a single utterance (all broadly-valid steps are non-atomic).
    3. Utility: Evaluates whether the step contributes to proving the target query.

      • Correct: The step's conclusion lies on a valid deductive path from premises to the target query.
      • Misleading: The premises are part of the gold proof, but the conclusion branches into an irrelevant sub-tree of the ontology that does not reach the target query.

    A canonical step is defined as a proof step that is strictly-valid, atomic, and correct.

  3. Knowl 3 — Algorithm for Formal Evaluation of Predicted Chains-of-Thought

    algorithm

    The formal evaluation algorithm parses the context, predicted chain-of-thought (CoT), and gold CoT into first-order logical forms, categorizes each predicted deduction step, and determines overall proof correctness.

    Input: Context sentences Q1,…,QmQ_1, \dots, Q_m, predicted CoT sentences C1,…,CnC_1, \dots, C_n, gold CoT sentences T1,…,TrT_1, \dots, T_r
    Output: Boolean indicating whether the predicted proof is valid, and classification of each proof step
    for i=1i = 1 to mm do
        LiQ=semantic_parse(Qi)L^Q_i = \text{semantic\_parse}(Q_i)
    for i=1i = 1 to rr do
        LiT=semantic_parse(Ti)L^T_i = \text{semantic\_parse}(T_i)
    S=∅S = \emptyset
    A={L1Q,…,LmQ}A = \{L^Q_1, \dots, L^Q_m\}
    Ggold={L1T,…,LrT}G_{\text{gold}} = \{L^T_1, \dots, L^T_r\}
    for i=1i = 1 to nn do
        LiC=semantic_parse(Ci)L^C_i = \text{semantic\_parse}(C_i)
        (P,k)=is_provable(LiC,A,S)(P, k) = \text{is\_provable}(L^C_i, A, S)
        if k≥0k \ge 0 then
            S=S∪{LiC}S = S \cup \{L^C_i\}
            if P⊆GgoldP \subseteq G_{\text{gold}} and LiC∉GgoldL^C_i \notin G_{\text{gold}} then
                mark LiCL^C_i as a misleading step
        else
            mark LiCL^C_i as an invalid step
    return LrT∈SL^T_r \in S
    function is_provable(ϕ,A,S\phi, A, S):
        if ϕ∈A\phi \in A then
            return ({ϕ},1)(\{\phi\}, 1)
        else if ϕ∈S\phi \in S then
            return ({ϕ},0)(\{\phi\}, 0)
        else if ϕ\phi is of the form g(c)g(c) or ¬g(c)\neg g(c) then
            for each a∈A∪Sa \in A \cup S do
                if aa is of the form ∀x(ψ→γ)\forall x (\psi \to \gamma) where γ[x↦c]=ϕ\gamma[x \mapsto c] = \phi then
                    (P,k)=is_provable(ψ[x↦c],A,S)(P, k) = \text{is\_provable}(\psi[x \mapsto c], A, S)
                    if k≥0k \ge 0 then
                        return (P∪{a},k+I[a∈A])(P \cup \{a\}, k + \mathbb{I}[a \in A])
        else if ϕ\phi is of the form ∀x(ψ→γ)\forall x (\psi \to \gamma) then
            Let GG be a directed graph with edges α→β\alpha \to \beta for each axiom ∀x(α→β)∈A\forall x (\alpha \to \beta) \in A
            if there is a directed path in GG from ψ\psi to γ\gamma then
                return (axioms corresponding to path edges, path length)
        return (∅,−1)(\emptyset, -1)

    A proof is marked as correct under the most permissive metric (valid proof accuracy) if there exists a valid chain of deductions yielding the gold query target LrT∈SL^T_r \in S.

  4. Knowl 4 — Correlation Between Proof Accuracy Metrics and Label Accuracy

    empirical result

    When evaluating large language models on PRONTOQA across 48 experimental configurations (varying hops ∈{1,3,5}\in \{1, 3, 5\}, ontologies ∈{fictional,false,true}\in \{\text{fictional}, \text{false}, \text{true}\}, and presentation order), label accuracy (the rate of predicting "True" vs "False" correctly) correlates poorly with strict proof accuracy (which requires every step to be strictly-valid, atomic, and correct). Strict proof accuracy substantially underestimates the model's actual reasoning capability because models often skip intermediate atomic steps in their generated text, producing valid non-atomic inferences.

    Label accuracy exhibits the highest agreement with valid proof accuracy, the most permissive metric that counts any continuous derivation path consisting of strictly-valid or broadly-valid steps (including non-atomic step-skipping). This demonstrates that downstream label accuracy is reflective of true logical reasoning paths rather than purely heuristic guesses, provided step-skipping is accommodated.

  5. Knowl 5 — LLM Failure in Global Proof Planning Despite Strong Local Deduction

    empirical result

    On PRONTOQA multi-hop deduction tasks, large language models (specifically text-davinci-002 under 8-shot chain-of-thought prompting) demonstrate high local deduction proficiency but fail at global proof planning.

    In 5-hop experiments with fictional ontologies:

    • 93.2%93.2\% of all generated proof steps are strictly-valid.
    • 2.4%2.4\% are broadly-valid (natural deduction transitive shortcuts).
    • Only 5.9%5.9\% are invalid steps.

    Despite making valid local deduction steps over 95%95\% of the time, the model's overall proof accuracy drops sharply as proof depth increases (falling to near chance on 5-hop top-down fictional ontologies). When the model encounters a branch point in the ontology where multiple valid deduction steps can be applied to current facts, it acts as a greedy reasoner: it selects one valid step without global planning. When it selects a strictly-valid misleading step (a step that is logically sound but leads away from the query target), it rarely recovers, leading to incomplete proofs and invalid final conclusions.

  6. Knowl 6 — Impact of World Knowledge and Ontology Factivity on Multi-Hop Reasoning

    empirical result

    Language model reasoning accuracy is strongly influenced by whether the ontology aligns with real-world knowledge acquired during pretraining:

    1. True Ontologies: When reasoning over ontologies containing real-world facts that are factually true, text-davinci-002 achieves near-perfect proof accuracy across 1, 3, and 5 hops, exhibiting virtually no performance drop as hop count increases.
    2. Fictional and False Ontologies: When reasoning over fictional concepts (e.g., "wumpus", "zumpus") or counterfactual "false" ontologies (real concept names arranged into false assertions like "All mammals are cats"), performance degrades steeply with proof depth. Performance between fictional and false ontologies is comparable, with false ontologies performing slightly worse.

    Analysis reveals that for true ontologies, the model uses its pretrained world knowledge to skip deduction hops, insulating it from multi-hop planning degradation. For fictional and counterfactual settings, where background retrieval cannot substitute for deduction, the model must rely entirely on in-context multi-hop planning.

  7. Knowl 7 — Sensitivity of Deductive Reasoning to Context Presentation Order

    empirical result

    The performance of chain-of-thought reasoning in text-davinci-002 is sensitive to the structural order in which premises are listed in the context:

    • Bottom-Up Traversal (Postorder): Context premises are generated starting from the leaf concepts up to the root, matching the chronological order of forward modus ponens deduction steps in the gold proof. Under this ordering, the model maintains higher proof accuracy across all hop depths.
    • Top-Down Traversal (Preorder): Context premises are generated in reverse order relative to the execution of the proof. As the number of hops increases from 1 to 5, top-down ordering causes a severe performance drop compared to bottom-up ordering, reducing accuracy on 5-hop fictional ontologies to chance levels.

    This discrepancy demonstrates that autoregressive LLM reasoning is heavily reliant on context ordering mirroring the forward derivation sequence.

  8. Knowl 8 — Dominance of Misleading Steps in LLM Reasoning Failures

    empirical result

    Analysis of the first non-canonical proof step in all incorrect proofs generated by text-davinci-002 reveals that strictly-valid atomic misleading steps account for the vast majority of initial errors across fictional, false, and true ontologies.

    Rather than hallucinating invalid logic at the point of failure, the model makes a valid deduction along a distractor branch. Furthermore, the probability of the model returning to the gold proof path decays rapidly as a function of the number of steps taken outside the gold proof graph. The longer the model continues along a misleading deduction path, the less likely it is to backtrack or reach the target conclusion, eventually producing an invalid terminal step to force an answer.

  9. Knowl 9 — Effect of Model Scale on Proof Step Validity and Error Distribution

    empirical result

    Across the GPT-3 and InstructGPT model series (text-ada-001 [350M], text-babbage-001 [1.3B], text-curie-001 [6.7B], davinci [175B], text-davinci-001, and text-davinci-002):

    1. Reasoning Threshold: Only the largest RLHF model (text-davinci-002) performs significantly better than random chance on multi-hop PRONTOQA tasks. Older 175B models (davinci and text-davinci-001) perform comparably to one another but substantially worse than text-davinci-002.
    2. Error Shift with Scale: For smaller models (text-ada-001, text-babbage-001), the first error in an incorrect proof is predominantly an invalid step or an unauthorized non-atomic step. As parameter scale increases, the occurrence of invalid deduction steps drops monotonically. For text-davinci-002, invalid steps are rare at the outset, and the dominant first failure mode becomes strictly-valid misleading steps.
  10. Knowl 10 — Ineffectiveness of Self-Consistency and DFS In-Context Prompting for Proof Planning

    empirical result

    Advanced prompting and decoding strategies fail to resolve the proof planning limitation of text-davinci-002 on 5-hop fictional top-down PRONTOQA:

    1. Self-Consistency Prompting: Sampling 40 chains-of-thought at temperature T=0.7T = 0.7 and aggregating probabilities over semantically identical parsed logical form sequences according to the normalized log probability

    exp⁡(1∣si∣∑j=1∣si∣log⁡p(si,j∣si,1,…,si,j−1))\exp \left( \frac{1}{|s_i|} \sum_{j=1}^{|s_i|} \log p(s_{i,j} \mid s_{i,1}, \ldots, s_{i,j-1}) \right)

    yields a valid proof accuracy of 0.560.56 (on 100 test examples), compared to 0.5450.545 for standard greedy decoding (a statistically insignificant difference). The model assigns higher joint sequence probabilities to globally incorrect proofs than to the gold proof.

    1. Depth-First Search (DFS) In-Context Prompting: Providing few-shot prompt examples that explicitly demonstrate depth-first search search traces (exploring a branch, encountering a dead-end, backtracking, and recovering to the correct path) yields a valid proof accuracy of 0.550.55 compared to 0.5450.545 for standard CoT.

    These results confirm that standard generation objectives favor locally probable but globally erroneous deduction branches, which cannot be fixed simply via voting or DFS prompting.

Coverage note — None was omitted; all key models, algorithms, definitions, mathematical metrics, and empirical findings across ontology types, prompt variations, and model scales are fully covered.

References

  1. 1.Gabor Angeli and Christopher D. Manning. Naturalli: Natural logic inference for common sense reasoning. In Alessandro Moschitti, Bo Pang, and Walter Daelemans (eds.), Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL, pp. 534–545. ACL, 2014. doi: 10.3115/v1/d14-1059. URL https://doi.org/10.3115/v1/d14-1059.
  2. 2.Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay V. Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. Exploring length generalization in large language models. CoRR, abs/2207.04901, 2022. doi: 10.48550/arXiv.2207.04901. URL https://doi.org/10.48550/arXiv.2207.04901.
  3. 3.Gregor Betz. Critical thinking for language models. CoRR, abs/2009.07185, 2020. URL https://arxiv.org/abs/2009.07185.
  4. 4.Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen-tau Yih, and Yejin Choi. Abductive commonsense reasoning. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URL https://openreview.net/forum?id=Byg1v1HKDB.
  5. 5.Kaj Bostrom, Xinyu Zhao, Swarat Chaudhuri, and Greg Durrett. Flexible generation of natural language deductions. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pp. 6266–6278. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.emnlp-main.506. URL https://doi.org/10.18653/v1/2021.emnlp-main.506.
  6. 6.Kaj Bostrom, Zayne Sprague, Swarat Chaudhuri, and Greg Durrett. Natural language deduction through search over statement compositions. CoRR, abs/2201.06028, 2022. URL https://arxiv.org/abs/2201.06028.
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Hugo Larochelle, Marc'Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  8. 8.Jifan Chen, Eunsol Choi, and Greg Durrett. Can NLI models verify QA systems' predictions? In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pp. 3841–3854. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.findings-emnlp.324. URL https://doi.org/10.18653/v1/2021.findings-emnlp.324.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. Palm: Scaling language modeling with pathways. CoRR, abs/2204.02311, 2022. doi: 10.48550/arXiv.2204.02311. URL https://doi.org/10.48550/arXiv.2204.02311.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.
  11. 11.Antonia Creswell and Murray Shanahan. Faithful reasoning using large language models. CoRR, abs/2208.14271, 2022. doi: 10.48550/arXiv.2208.14271. URL https://doi.org/10.48550/arXiv.2208.14271.
  12. 12.Antonia Creswell, Murray Shanahan, and Irina Higgins. Selection-inference: Exploiting large language models for interpretable logical reasoning. CoRR, abs/2205.09712, 2022. doi: 10.48550/arXiv.2205.09712. URL https://doi.org/10.48550/arXiv.2205.09712.
  13. 13.Ishita Dasgupta, Andrew K. Lampinen, Stephanie C. Y. Chan, Antonia Creswell, Dharshan Kumaran, James L. McClelland, and Felix Hill. Language models show human-like content effects on reasoning. CoRR, abs/2207.07051, 2022. doi: 10.48550/arXiv.2207.07051. URL https://doi.org/10.48550/arXiv.2207.07051.
  14. 14.David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A. Saurous, Jascha Sohl-Dickstein, Kevin Murphy, and Charles Sutton. Language model cascades. CoRR, abs/2207.10342, 2022. doi: 10.48550/arXiv.2207.10342. URL https://doi.org/10.48550/arXiv.2207.10342.
  15. 15.Honghua Dong, Jiayuan Mao, Tian Lin, Chong Wang, Lihong Li, and Denny Zhou. Neural logic machines. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019. URL https://openreview.net/forum?id=B1xY-hRctX.
  16. 16.Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, David Peng, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Shafiq R. Joty, Alexander R. Fabbri, Wojciech Kryscinski, Xi Victoria Lin, Caiming Xiong, and Dragomir Radev. FOLIO: natural language reasoning with first-order logic. CoRR, abs/2209.00840, 2022. doi: 10.48550/arXiv.2209.00840. URL https://doi.org/10.48550/arXiv.2209.00840.
  17. 17.Pavan Kapanipathi, Ibrahim Abdelaziz, Srinivas Ravishankar, Salim Roukos, Alexander G. Gray, Ramón Fernandez Astudillo, Maria Chang, Cristina Cornelio, Saswati Dana, Achille Fokoue, Dinesh Garg, Alfio Gliozzo, Sairam Gurajada, Hima Karanam, Naweed Khan, Dinesh Khandelwal, Young-Suk Lee, Yunyao Li, Francois P. S. Luus, Ndivhuwo Makondo, Nandana Mihindukulasooriya, Tahira Naseem, Sumit Neelam, Lucian Popa, Revanth Gangi Reddy, Ryan Riegel, Gaetano Rossiello, Udit Sharma, G. P. Shrivatsa Bhargav, and Mo Yu. Leveraging abstract meaning representation for knowledge base question answering. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pp. 3884–3894. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.findings-acl.339. URL https://doi.org/10.18653/v1/2021.findings-acl.339.
  18. 18.Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay V. Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. Solving quantitative reasoning problems with language models. CoRR, abs/2206.14858, 2022. doi: 10.48550/arXiv.2206.14858. URL https://doi.org/10.48550/arXiv.2206.14858.
  19. 19.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pp. 8086–8098. Association for Computational Linguistics, 2022. doi: 10.18653/v1/2022.acl-long.556. URL https://doi.org/10.18653/v1/2022.acl-long.556.
  20. 20.Bill MacCartney and Christopher D. Manning. An extended model of natural logic. In Harry Bunt, Volha Petukhova, and Sander Wubben (eds.), Proceedings of the Eight International Conference on Computational Semantics, IWCS 2009, Tilburg, The Netherlands, January 7-9, 2009, pp. 140–156. Association for Computational Linguistics, 2009. URL https://aclanthology.org/W09-3714/.
  21. 21.William Merrill, Ashish Sabharwal, and Noah A. Smith. Saturated Transformers are Constant-Depth Threshold Circuits. Transactions of the Association for Computational Linguistics, 10:843–856, 08 2022. ISSN 2307-387X. doi: 10.1162/tacl_a_00493. URL https://doi.org/10.1162/tacl_a_00493.
  22. 22.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback. CoRR, abs/2203.02155, 2022. doi: 10.48550/arXiv.2203.02155. URL https://doi.org/10.48550/arXiv.2203.02155.
  23. 23.Yasaman Razeghi, Robert L. Logan IV, Matt Gardner, and Sameer Singh. Impact of pretraining term frequencies on few-shot reasoning. CoRR, abs/2202.07206, 2022. URL https://arxiv.org/abs/2202.07206.
  24. 24.Tim Rocktäschel and Sebastian Riedel. End-to-end differentiable proving. In Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (eds.), Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pp. 3788–3800, 2017. URL https://proceedings.neurips.cc/paper/2017/hash/b2ab001909a8a6f04b51920306046ce5-Abstract.html.
  25. 25.Abulhair Saparov and Tom M. Mitchell. Towards general natural language understanding with probabilistic worldbuilding. Trans. Assoc. Comput. Linguistics, 10:325–342, 2022. doi: 10.1162/tacl_a_00463. URL https://doi.org/10.1162/tacl_a_00463.
  26. 26.Oyvind Tafjord, Bhavana Dalvi, and Peter Clark. Proofwriter: Generating implications, proofs, and abductive statements over natural language. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Findings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021, volume ACL/IJCNLP 2021 of Findings of ACL, pp. 3621–3634. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.findings-acl.317. URL https://doi.org/10.18653/v1/2021.findings-acl.317.
  27. 27.Karthik Valmeekam, Alberto Olmo Hernandez, Sarath Sreedharan, and Subbarao Kambhampati. Large language models still can't plan (A benchmark for llms on planning and reasoning about change). CoRR, abs/2206.10498, 2022. doi: 10.48550/arXiv.2206.10498. URL https://doi.org/10.48550/arXiv.2206.10498.
  28. 28.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. CoRR, abs/2203.11171, 2022. doi: 10.48550/arXiv.2203.11171. URL https://doi.org/10.48550/arXiv.2203.11171.
  29. 29.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. CoRR, abs/2201.11903, 2022. URL https://arxiv.org/abs/2201.11903.
  30. 30.Sean Welleck, Jiacheng Liu, Ronan Le Bras, Hanna Hajishirzi, Yejin Choi, and Kyunghyun Cho. Naturalproofs: Mathematical theorem proving in natural language. In Joaquin Vanschoren and Sai-Kit Yeung (eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, 2021. URL https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/d9d4f495e875a2e075a1a4a6e1b9770f-Abstract-round1.html.
  31. 31.Jason Weston, Antoine Bordes, Sumit Chopra, and Tomás Mikolov. Towards ai-complete question answering: A set of prerequisite toy tasks. In Yoshua Bengio and Yann LeCun (eds.), 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, 2016. URL http://arxiv.org/abs/1502.05698.
  32. 32.Edwin B Wilson. Probable inference, the law of succession, and statistical inference. J. Am. Stat. Assoc., 22(158):209, June 1927.
  33. 33.Hanlin Zhang, Yi-Fan Zhang, Li Erran Li, and Eric Xing. The impact of symbolic representations on in-context learning for few-shot reasoning. In Neuro Causal and Symbolic AI Workshop at NeurIPS 2022, Virtual Workshop, December 9, 2022, 2022a. URL https://openreview.net/pdf?id=qLgQpeQX3x1. To appear.
  34. 34.Honghua Zhang, Liunian Harold Li, Tao Meng, Kai-Wei Chang, and Guy Van den Broeck. On the paradox of learning to reason from data. CoRR, abs/2205.11502, 2022b. doi: 10.48550/arXiv.2205.11502. URL https://doi.org/10.48550/arXiv.2205.11502.

Citation

MLA
Saparov, A., and H. He. “Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought”. arXiv, 2022, http://arxiv.org/abs/2210.01240v4.
APA
Saparov, A., & He, H. (2022). Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought. arXiv. http://arxiv.org/abs/2210.01240v4
Chicago
Saparov, A., and H. He. 2022. “Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought”. arXiv. http://arxiv.org/abs/2210.01240v4.
Harvard
Saparov, A. and He, H. (2022) “Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2210.01240v4.
Vancouver
1. Saparov A, He H (2022) Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought. arXiv

BibTeX

@article{saparov2022language,
  title = {Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-Thought},
  author = {Saparov, Abulhair and He, He},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2210.01240v4},
  eprint = {2210.01240}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors