Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Wenhu ChenXueguang MaXinyi WangWilliam W. Cohen
Proposes Program of Thoughts prompting to disentangle reasoning from computation by offloading code execution to an external interpreter, consistently outperforming Chain-of-Thought prompting across diverse mathematical and financial reasoning benchmarks.
Large language models often struggle with complex numerical reasoning and quantitative problem solving. When using conventional methods like Chain of Thoughts, models attempt to perform both linguistic reasoning and arithmetic calculation internally. This leads to frequent calculation mistakes, an inability to solve complex equations, and failure when handling iterative steps.
The article evaluates Program of Thoughts prompting, a strategy designed to decouple natural language reasoning from computation by having the language model generate executable Python code and delegating calculation to an external program interpreter.
The authors tested this approach across eight standard benchmarks comprising math word problems and financial question answering datasets. Testing was conducted across zero-shot and few-shot prompting setups using various language models, primarily OpenAI's Codex backend.
The core findings demonstrate substantial accuracy gains across all tasks. Program of Thoughts outperforms standard Chain of Thoughts by an average of 12% across datasets, with few-shot gains around 8% on math word problems and roughly 15% to 20% on financial datasets. Under zero-shot conditions, it surpasses zero-shot Chain of Thoughts by 12% on average. When paired with self-consistency majority voting, Program of Thoughts achieves state-of-the-art results across all evaluated math benchmarks. Ablation studies reveal that assigning semantically meaningful variable names and breaking logic into multiple programmatic steps are critical contributors to these gains.
These results indicate that offloading computation to an external interpreter significantly mitigates arithmetic errors and improves reliability in complex numerical domains such as finance. By addressing calculation pitfalls without extensive model retraining, this approach offers an effective framework for deploying language models in high-accuracy quantitative workflows.
Organizations implementing language models for numerical and financial analysis should adopt programmatic tool-use frameworks rather than relying solely on text-based internal arithmetic. In deployment, security guardrails must be maintained to prevent unsafe code execution from model-generated programs. Further research should focus on improving value-grounding accuracy and exploring broader symbolic reasoning extensions.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces the foundational Chain-of-Thought prompting paradigm that Program of Thoughts explicitly builds upon and seeks to improve by offloading computation to an external interpreter.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Establishes self-consistency decoding for multi-step reasoning, which Program of Thoughts directly adopts and combines with program generation to achieve state-of-the-art accuracy.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). Provides early empirical foundations for using large language models to synthesize executable Python code for problem-solving tasks, enabling the program-based reasoning evaluated in this work.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Introduces zero-shot step-by-step prompting in large language models, establishing the zero-shot reasoning baseline that Program of Thoughts evaluates against.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Presents problem decomposition strategies for complex numerical reasoning, providing key conceptual groundwork for separating problem breakdown from execution.
- Paper: Training Verifiers to Solve Math Word Problems, Karl Cobbe et al. (2021). Introduces the GSM8K benchmark and formalizes multi-step arithmetic error tracking, establishing a core evaluation dataset and motivation for separating reasoning from computation.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). Generalizes the concept of offloading execution to external engines by teaching language models to autonomously call diverse external tools and APIs.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). Expands structured prompting beyond linear chains and programs into tree-based exploration and deliberate search over intermediate thoughts.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). Further extends structured thought representations by modeling language model reasoning as arbitrary graphs with dynamic feedback and aggregation loops.
- Paper: DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, Zhihong Shao et al. (2024). Builds on code pre-training and mathematical problem-solving to train open foundation models specialized in both natural language and algorithmic math reasoning.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). Extends program-aided and tool-augmented numerical reasoning evaluations to multimodal environments containing visual charts, plots, and diagrams.
- Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). Analyzes optimal inference-time compute scaling across search and verification mechanisms, providing a theoretical and empirical framework for test-time reasoning strategies.
