Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Wenhu ChenXueguang MaXinyi WangWilliam W. Cohen

article2022Trans. Mach. Learn. Res.1,479 citationsTMLR 2024 Best Paper Award

Proposes Program of Thoughts prompting to disentangle reasoning from computation by offloading code execution to an external interpreter, consistently outperforming Chain-of-Thought prompting across diverse mathematical and financial reasoning benchmarks.

Listen

Large language models often struggle with complex numerical reasoning and quantitative problem solving. When using conventional methods like Chain of Thoughts, models attempt to perform both linguistic reasoning and arithmetic calculation internally. This leads to frequent calculation mistakes, an inability to solve complex equations, and failure when handling iterative steps.

The article evaluates Program of Thoughts prompting, a strategy designed to decouple natural language reasoning from computation by having the language model generate executable Python code and delegating calculation to an external program interpreter.

The authors tested this approach across eight standard benchmarks comprising math word problems and financial question answering datasets. Testing was conducted across zero-shot and few-shot prompting setups using various language models, primarily OpenAI's Codex backend.

The core findings demonstrate substantial accuracy gains across all tasks. Program of Thoughts outperforms standard Chain of Thoughts by an average of 12% across datasets, with few-shot gains around 8% on math word problems and roughly 15% to 20% on financial datasets. Under zero-shot conditions, it surpasses zero-shot Chain of Thoughts by 12% on average. When paired with self-consistency majority voting, Program of Thoughts achieves state-of-the-art results across all evaluated math benchmarks. Ablation studies reveal that assigning semantically meaningful variable names and breaking logic into multiple programmatic steps are critical contributors to these gains.

These results indicate that offloading computation to an external interpreter significantly mitigates arithmetic errors and improves reliability in complex numerical domains such as finance. By addressing calculation pitfalls without extensive model retraining, this approach offers an effective framework for deploying language models in high-accuracy quantitative workflows.

Organizations implementing language models for numerical and financial analysis should adopt programmatic tool-use frameworks rather than relying solely on text-based internal arithmetic. In deployment, security guardrails must be maintained to prevent unsafe code execution from model-generated programs. Further research should focus on improving value-grounding accuracy and exploring broader symbolic reasoning extensions.

Cover for Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Abstract

Recently, there has been significant progress in teaching language models to perform step-by-step reasoning to solve complex numerical reasoning tasks. Chain-of-thoughts prompting (CoT) is by far the state-of-art method for these tasks. CoT uses language models to perform both reasoning and computation in the multi-step thought' process. To disentangle computation from reasoning, we propose Program of Thoughts' (PoT), which uses language models (mainly Codex) to express the reasoning process as a program. The computation is relegated to an external computer, which executes the generated programs to derive the answer. We evaluate PoT on five math word problem datasets (GSM, AQuA, SVAMP, TabMWP, MultiArith) and three financial-QA datasets (FinQA, ConvFinQA, TATQA) for both few-shot and zero-shot setups. Under both few-shot and zero-shot settings, PoT can show an average performance gain over CoT by around 12% across all the evaluated datasets. By combining PoT with self-consistency decoding, we can achieve SoTA performance on all math problem datasets and near-SoTA performance on financial datasets. All of our data and code are released in Github this https URL

Table of Contents

  • 1 Introduction
  • 2 Program of Thoughts
  • 2.1 Preliminaries
  • 2.2 Program of Thoughts
  • 2.3 PoT as an Intermediate Step
  • 3 Experiments
  • 3.1 Experimental Setup
  • 3.2 Main Results
  • 3.3 Ablation Studies
  • 4 Related Work
  • 4.1 Mathematical Reasoning in NLP
  • 4.2 In-context Learning with LLMs
  • 4.3 Chain of Reasoning with LLMs
  • 4.4 Discussion about Contemporary Work
  • 5 Discussion
  • 6 Conclusions
  • References
  • 7 Appendix
  • 7.1 PoT as intermediate step
  • 7.2 Exemplars for Prompting

Knowls

  1. Knowl 1 — Program of Thoughts (PoT) Prompting

    model/method

    Program of Thoughts (PoT) is a prompting framework for large language models (LLMs) to solve numerical and mathematical reasoning tasks. Unlike Chain of Thoughts (CoT), which delegates both problem breakdown and step-by-step arithmetic computation to the language model, PoT disentangles reasoning from computation:

    1. The LLM generates natural language reasoning expressed as a sequence of programming language statements (e.g., Python code) using semantically meaningful variable names and step-by-step assignments or equations.
    2. Symbolic or mathematical packages (such as SymPy) are used within the generated code to represent unknowns (e.g., via Symbol), set up system equations, or perform multi-step arithmetic and algebraic operations.
    3. The generated program is passed to an external deterministic interpreter (such as a standard Python 3.8 interpreter) which executes the code to obtain the final numerical answer.

    This prevents the LLM from making arithmetic calculation errors (e.g., on large numbers or high-precision floats), enables solving complex equations (such as polynomials or systems of equations), and accurately executes long iterative loops.

  2. Knowl 2 — Zero-Shot Program of Thoughts Prompting via Comment Suppression

    model/method

    In zero-shot settings without exemplar demonstrations, language models prompted to write code for solving math problems frequently fall back to generating natural language reasoning inside Python comments (prefixed with #) rather than constructing runnable programmatic logic.

    To enforce executable code generation in zero-shot PoT:

    1. The model is given a direct instruction prompt instructing it to define a solver() function.
    2. A negative logit bias of −2-2 is applied during generation to suppress the comment token '#'. This reduces the probability of generating text comments, forcing the LLM to express intermediate reasoning directly as executable variable assignments and programmatic statements.

    Unlike zero-shot CoT, which requires an additional prompt step to extract the answer from the generated text, zero-shot PoT directly returns the numerical output via the interpreter upon executing solver().

  3. Knowl 3 — Multi-Stage Hybrid PoT-CoT Procedure

    algorithm

    When a numerical reasoning question requires commonsense formatting, unit conversion, or matching against multiple-choice answer strings that cannot be resolved solely through programmatic execution, PoT can be chained with Chain-of-Thought (CoT) prompting in a two-stage pipeline.

    Input: Natural language question qq, prompt generator PoT, prompt generator Prompt
    Output: Final predicted answer predpred
    program = PoT(q)
    ans = exec(program)
    if is_dict_or_structured(ans) then
        intermediate_val = extract_value(ans)
        extra_context = "according to the program: " + intermediate_val
        pred = Prompt(q + extra_context)
    else
        pred = ans
    end if
    return pred

    For example, if a question asks what time two trains meet given a departure at 11:00 AM and the program returns ans = 2.05 (hours), the second CoT stage translates the 2.05 hours into standard HH:MM time format (1:03 PM) to match multiple-choice options.

  4. Knowl 4 — Few-Shot Reasoning Accuracy of Program of Thoughts Across MWP and Financial Datasets

    data/table

    Few-shot evaluation compares Program of Thoughts (PoT) against Direct output and Chain of Thoughts (CoT) across multiple LLMs (Codex code-davinci-002, GPT-3 text-davinci-002, PaLM 540B, and GPT-4) using greedy decoding and self-consistency (SC, with temperature T=0.4T=0.4 and K=40K=40 completions).

    Model / Method GSM8K AQuA SVAMP TabMWP FinQA ConvFin TATQA Avg
    Published SoTA 78.0 52.0 86.8 68.2 68.0 68.9 73.6 70.7
    Few-shot prompt (Greedy Decoding)
    Codex Direct 19.7 29.5 69.9 59.4 25.6 40.0 55.0 42.7
    Codex CoT 63.1 45.3 76.4 65.2 40.4 45.6 61.4 56.7
    GPT-3 Direct 15.6 24.8 65.7 57.1 14.4 29.1 37.9 34.9
    GPT-3 CoT 46.9 35.8 68.9 62.9 26.1 37.4 42.5 45.7
    PaLM Direct 17.9 25.2 69.4 – – – – –
    PaLM CoT 56.9 35.8 79.0 – – – – –
    Codex CoTcalc\text{CoT}_{\text{calc}} 65.4 45.3 77.0 65.8 – – – –
    GPT-3 CoTcalc\text{CoT}_{\text{calc}} 49.6 35.8 70.3 63.4 – – – –
    PaLM CoTcalc\text{CoT}_{\text{calc}} 58.6 35.8 79.8 – – – – –
    PoT-Codex 71.6 54.1 85.2 73.2 64.5 64.6 69.0 68.9
    Few-shot prompt (Self-Consistency Decoding)
    LaMDA CoT-SC 27.7 26.8 53.5 – – – – –
    Codex CoT-SC 78.0 52.0 86.8 75.4 44.4 47.9 63.2 63.9
    PaLM CoT-SC 74.4 48.3 86.6 – – – – –
    PoT-SC-Codex 80.0 58.6 89.1 81.8 68.1 67.3 70.2 73.6
    Few-shot prompt (GPT-4)
    CoT-GPT4 92.0 72.4 97.0 – 58.2 – – –
    PoT-GPT4 97.2 84.4 97.4 – 74.0 – – –

    Under greedy decoding, PoT-Codex achieves an average accuracy of 68.9%68.9\%, outperforming Codex CoT (56.7%56.7\%) by 12.2%12.2\%, with notable gains on FinQA (+24.1%+24.1\%) and ConvFinQA (+19.0%+19.0\%) due to eliminating LLM calculation errors on large numbers. With self-consistency (PoT-SC-Codex), average accuracy reaches 73.6%73.6\%, establishing new state-of-the-art results across the evaluated MWP benchmarks.

  5. Knowl 5 — Zero-Shot Reasoning Accuracy of Program of Thoughts vs Chain of Thoughts

    data/table

    Zero-shot PoT is evaluated against Zero-shot Direct and Zero-shot CoT methods on five math word problem benchmarks without providing any in-context exemplars.

    Model / Method GSM8K AQuA SVAMP TabMWP MultiArith Avg
    Zero-shot Direct (GPT-3 175B) 12.6 22.4 58.7 38.9 22.7 31.0
    Zero-shot CoT (GPT-3 175B) 40.5 31.9 63.7 53.5 79.3 53.7
    Zero-shot CoT (PaLM 540B) 43.0 – – – 66.1 –
    Zero-shot PoT (Codex 175B) 57.0 43.9 70.8 66.5 92.2 66.1

    Zero-shot PoT achieves an overall average accuracy of 66.1%66.1\%, outperforming Zero-shot CoT (53.7%53.7\%) by 12.412.4 percentage points on average across all datasets. On TabMWP, Zero-shot PoT (66.5%66.5\%) outperforms the few-shot performance of Codex CoT (65.2%65.2\%).

  6. Knowl 6 — Contributions of Semantic Binding and Multi-Step Program Decomposition in PoT

    empirical result

    PoT relies on two distinct structural properties:

    1. Multi-step reasoning: Breaking down the mathematical problem into a sequential step-by-step program rather than directly predicting a single complex formula or equation.
    2. Semantic binding: Assigning meaningful, domain-grounded variable names (e.g., interest_rate, cost_of_repair) rather than arbitrary variable tokens (e.g., a,b,ca, b, c).

    The effects of removing these components evaluated with Codex (code-davinci-002) are:

    Method GSM8K SVAMP FinQA
    PoT (Full) 71.6 85.2 64.5
    PoT without Semantic Binding (PoT - Binding) 60.2 83.8 61.6
    PoT without Multi-Step Breakdown (PoT - MultiStep) 45.8 81.9 58.9

    Removing semantic binding degrades performance by 11.4%11.4\% on GSM8K and 2.9%2.9\% on FinQA. Directly predicting a single final equation (removing multi-step decomposition) leads to a drastic degradation of 25.8%25.8\% on GSM8K and 5.6%5.6\% on FinQA, demonstrating that both semantic grounding and incremental decomposition are critical.

  7. Knowl 7 — Backend Model Performance Comparison for Program of Thoughts Prompting

    empirical result

    Evaluating PoT across different backend language model families shows substantial performance variations based on pre-training objective, instruction tuning, and model scale:

    Model #Params GSM8K SVAMP
    code-davinci-002 175B 71.6 85.2
    text-davinci-002 175B 60.4 80.1
    gpt-3.5-turbo – 76.3 88.2
    codegen-16B-multi 16B 8.2 29.2
    codegen-16B-mono 16B 12.7 41.1
    codeT5+ 16B 12.5 38.5
    xgen 7B 11.0 40.6

    Key observations:

    • gpt-3.5-turbo achieves the highest performance (76.3%76.3\% on GSM8K and 88.2%88.2\% on SVAMP).
    • code-davinci-002 outperforms text-davinci-002 (71.6%71.6\% vs 60.4%60.4\% on GSM8K), indicating that text-only instruction tuning impairs code generation capability.
    • Open-source models (such as codegen-16B, codeT5+ 16B, and xgen 7B) lag substantially behind proprietary large models, scoring 8.2%–12.7%8.2\%\text{--}12.7\% on GSM8K.
  8. Knowl 8 — Experimental Setup and Input Linearization for Heterogeneous Numerical Reasoning

    experimental setup

    PoT is evaluated across eight datasets spanning Math Word Problems (MWP) and Financial Question Answering (QA):

    • MWP datasets: GSM8K (1,318 test examples, numerical output), AQuA (253 test examples, multiple-choice option output), SVAMP (1,000 test examples, numerical output), MultiArith (600 test examples, numerical output), TabMWP (7,861 test examples, table + question, number/text output).
    • Financial datasets: FinQA (1,147 test examples, table + text + question, numerical/binary output), ConvFinQA (421 test examples, table + text + conversation, numerical/binary output), TATQA (1,668 development examples, table + text + question, numerical/text output).

    To represent heterogeneous multi-modal inputs uniformly in the LLM prompt:

    • Tables: Linearized into string format where table columns are separated by |, rows are separated by newline characters (\n), and empty cells are filled with -.
    • Hybrid Text and Table: Concatenated by separating text and tables with \n.
    • Conversational History: Multi-turn dialogue history is concatenated using \n turn separators.

    Few-shot prompting uses 4 to 8 exemplar demonstrations (chosen based on dataset difficulty and tuned on a small validation set) executed with Python 3.8 and SymPy.

  9. Knowl 9 — Performance Breakdown of PoT vs CoT by Mathematical Problem Category

    empirical result

    On the AQuA benchmark, categorizing questions into distinct mathematical domains reveals where programmatic execution provides the largest advantage over natural language reasoning (CoT):

    • Iterative problems: PoT achieves 65%65\% vs CoT 40%40\% (+25%+25\%).
    • Polynomial equations: PoT achieves 50%50\% vs CoT 20%20\% (+30%+30\%).
    • Symbolic equations: PoT achieves 66%66\% vs CoT 56%56\% (+10%+10\%).
    • Linear equations: PoT achieves 72%72\% vs CoT 62%62\% (+10%+10\%).
    • Combinatorics: PoT achieves 50%50\% vs CoT 40%40\% (+10%+10\%).
    • Arithmetic: PoT achieves 86%86\% vs CoT 88%88\% (comparable).
    • Probability: PoT achieves 38%38\% vs CoT 40%40\% (comparable).
    • Geometry: PoT achieves 10%10\% vs CoT 20%20\%.

    The largest gains occur in domains requiring complex algebra, non-linear/symbolic equation solving, and multi-step iterative loops, where manual arithmetic calculation by LLMs is failure-prone and symbolic execution via SymPy succeeds.

  10. Knowl 10 — Error Taxonomy and Distribution of PoT on TAT-QA

    empirical result

    An analysis of 198 failure cases produced by greedy PoT on the TAT-QA development set identifies two primary error types:

    1. Value Grounding Errors (47%47\%): The model generates correct computation and equation logic but assigns incorrect numbers or values extracted from the input text or table to the programmatic variables.
    2. Logic Generation Errors (33%33\%): The model correctly identifies and assigns the relevant numbers to variables but synthesizes an incorrect mathematical formulation or operations sequence.
    3. Compound Errors (15%15\%): Both value grounding errors and logic generation errors occur simultaneously.
    4. Annotation Artifacts / Correct (5%5\%): The generated program's answer is mathematically valid but mismatches reference annotations due to dataset labeling formatting.

    Value grounding is the primary failure mode (47%47\%), indicating that extracting accurate entities from complex hybrid tabular and textual context remains the primary bottleneck rather than code generation logic.

  11. Knowl 11 — Empirical Comparison Between PoT and Program-Aided Language Models (PaL)

    data/table

    PoT is compared against contemporary program-synthesis prompting method PaL across multiple math reasoning benchmarks:

    Model GSM8K GSM8K-Hard SVAMP ASDIV ADDSUB MULTIARITH
    PaL 72.0 61.2 79.4 79.6 92.5 99.2
    PoT 71.6 61.8 85.2 85.2 92.2 99.5

    PoT achieves substantial gains over PaL on SVAMP (85.2%85.2\% vs 79.4%79.4\%, +5.8%+5.8\%) and ASDIV (85.2%85.2\% vs 79.6%79.6\%, +5.6%+5.6\%), while maintaining competitive performance on GSM8K (71.6%71.6\% vs 72.0%72.0\%), GSM8K-Hard (61.8%61.8\% vs 61.2%61.2\%), ADDSUB (92.2%92.2\% vs 92.5%92.5\%), and MultiArith (99.5%99.5\% vs 99.2%99.2\%).

  12. Knowl 12 — Security and Problem Diversity Limitations of Program of Thoughts

    limitation

    PoT exhibits two primary limitations:

    1. Code Execution Safety: Generating and automatically executing code from language models carries security risks (e.g., executing malicious statements such as import os; os.rmdir()). While restricting imports and executing code within a sandboxed environment with pre-defined modules (e.g., sympy, math) mitigates risk for math benchmarks, brutal-force blocking restricts generalization to broader symbolic or system-level tasks.
    2. Generalization to Highly Diversified Algebra: PoT performance on benchmarks with wide problem variance and algebraic questions like AQuA remains limited (58.6%58.6\% with self-consistency), because fixed few-shot exemplar prompts cannot cover the broad distribution of question types in such datasets.

Coverage note — Raw exemplar prompts provided in the appendix were omitted as implementation instances of the documented methods.

References

  1. 1.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. Mathqa: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2357–2367, 2019.
  2. 2.Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  4. 4.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021a.
  5. 5.Wenhu Chen. Large language models are few (1)-shot table reasoners. arXiv preprint arXiv:2210.06710, 2022.
  6. 6.Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan R Routledge, et al. Finqa: A dataset of numerical reasoning over financial data. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 3697–3711, 2021b.
  7. 7.Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. Convfinqa: Exploring the chain of numerical reasoning in conversational finance question answering. arXiv preprint arXiv:2210.03849, 2022.
  8. 8.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, et al. Binding language models in symbolic languages. arXiv preprint arXiv:2210.02875, 2022.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  11. 11.Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Xinyun Chen, Olivier Bousquet, and Denny Zhou. Compositional semantic parsing with large language models. arXiv preprint arXiv:2209.15003, 2022.
  12. 12.Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al. Glam: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pp. 5547–5569. PMLR, 2022.
  13. 13.Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 2368–2378, 2019.
  14. 14.Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. Pal: Program-aided language models. arXiv preprint arXiv:2211.10435, 2022.
  15. 15.Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. Critic: Large language models can self-correct with tool-interactive critiquing. arXiv preprint arXiv:2305.11738, 2023.
  16. 16.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874, 2021.
  17. 17.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. Learning to solve arithmetic word problems with verb categorization. In EMNLP, pp. 523–533, 2014.
  18. 18.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  19. 19.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  20. 20.Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics, 3:585–597, 2015.
  21. 21.Fangyu Lei, Shizhu He, Xiang Li, Jun Zhao, and Kang Liu. Answering numerical reasoning questions in table-text hybrid contents with graph-based encoder and tree-based decoder. In Proceedings of the 29th International Conference on Computational Linguistics, pp. 1379–1390, 2022.
  22. 22.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 158–167, 2017.
  23. 23.Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. arXiv preprint arXiv:2209.14610, 2022.
  24. 24.Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. A diverse corpus for evaluating and developing english math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 975–984, 2020.
  25. 25.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837, 2022.
  26. 26.Swaroop Mishra, Matthew Finlayson, Pan Lu, Leonard Tang, Sean Welleck, Chitta Baral, Tanmay Rajpurohit, Oyvind Tafjord, Ashish Sabharwal, Peter Clark, and Ashwin Kalyan. Lila: A unified benchmark for mathematical reasoning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022.
  27. 27.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474, 2022.
  28. 28.Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  29. 29.OpenAI. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  30. 30.Bhargavi Paranjape, Scott Lundberg, Sameer Singh, Hannaneh Hajishirzi, Luke Zettlemoyer, and Marco Tulio Ribeiro. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014, 2023.
  31. 31.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2080–2094, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.168. URL https://aclanthology.org/2021.naacl-main.168.
  32. 32.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  33. 33.Subhro Roy and Dan Roth. Solving general arithmetic word problems. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pp. 1743–1752, 2015.
  34. 34.Subhro Roy and Dan Roth. Mapping to declarative knowledge for word problem solving. Transactions of the Association for Computational Linguistics, 6:159–172, 2018.
  35. 35.David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. Analysing mathematical reasoning abilities of neural models. arXiv preprint arXiv:1904.01557, 2019.
  36. 36.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761, 2023.
  37. 37.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990, 2022.
  38. 38.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  39. 39.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  40. 40.Bin Wang, Jiangzhou Ju, Yunlin Mao, Xin-Yu Dai, Shujian Huang, and Jiajun Chen. A numerical reasoning question answering system with fine-grained retriever and the ensemble of multiple generators for finqa. arXiv preprint arXiv:2206.08506, 2022a.
  41. 41.Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. arXiv preprint arXiv:2305.04091, 2023a.
  42. 42.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022b.
  43. 43.Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi DQ Bui, Junnan Li, and Steven CH Hoi. Codet5+: Open code large language models for code understanding and generation. arXiv preprint arXiv:2305.07922, 2023b.
  44. 44.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  45. 45.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations, 2021.
  46. 46.Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie. Decomposition enhances reasoning via self-evaluation guided decoding. arXiv preprint arXiv:2305.00633, 2023.
  47. 47.Shunyu Yao, Jeffrey Zhao, Dian Yu, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In NeurIPS 2022 Foundation Models for Decision Making Workshop, 2022.
  48. 48.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022.
  49. 49.Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. Tat-qa: A question answering benchmark on a hybrid of tabular and textual content in finance. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 3277–3287, 2021.

Citation

MLA
Chen, W., et al. “Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks”. arXiv, 2022, http://arxiv.org/abs/2211.12588v4.
APA
Chen, W., Ma, X., Wang, X., & Cohen, W. W. (2022). Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks. arXiv. http://arxiv.org/abs/2211.12588v4
Chicago
Chen, W., X. Ma, X. Wang, and W. W. Cohen. 2022. “Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks”. arXiv. http://arxiv.org/abs/2211.12588v4.
Harvard
Chen, W. et al. (2022) “Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.12588v4.
Vancouver
1. Chen W, Ma X, Wang X, Cohen WW (2022) Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks. arXiv

BibTeX

@article{chen2022program,
  title = {Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks},
  author = {Chen, Wenhu and Ma, Xueguang and Wang, Xinyi and Cohen, William W.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.12588v4},
  eprint = {2211.12588}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF