PAL: Program-aided Language Models
Luyu GaoAman MadaanShuyan ZhouUri AlonPengfei LiuYiming YangJamie CallanGraham Neubig
Proposes Program-aided Language models (PAL), a method that prompts language models to generate executable code for intermediate reasoning and offloads computation to a Python interpreter, significantly outperforming standard chain-of-thought approaches on complex mathematical and symbolic benchmarks.
Large language models frequently struggle with complex arithmetic and multi-step logical tracking, even when using step-by-step chain-of-thought prompting. Although these models effectively decompose problems into intermediate steps, they often fail to execute the calculations or maintain correct state tracking during the solution phase. The article introduces and evaluates Program-Aided Language Models (PAL), a framework designed to resolve this execution bottleneck by delegating the intermediate calculation steps to an external Python interpreter.
Under the PAL approach, a language model translates a natural language problem into executable Python code with meaningful variable names rather than generating free-form text answers. The article evaluates this method across 13 benchmarks covering mathematical word problems, symbolic reasoning, and algorithmic execution tasks. In experimental setups primarily using Codex, PAL established new few-shot performance benchmarks, consistently outperforming standard prompting methods and much larger language models.
Key findings show significant performance gains across diverse domains. On the GSM8K math benchmark, PAL achieved a 72.0% solve rate, surpassing PaLM-540B with chain-of-thought prompting by 15.1 percentage points, and reached 80.4% when paired with majority voting across 40 samples. On GSM-HARD, an evaluation set containing large multi-digit numbers, chain-of-thought accuracy collapsed from 65.6% to 23.1%, whereas PAL maintained a 61.2% solve rate. PAL also delivered substantial improvements in symbolic and algorithmic tasks, reaching 96.7% on object counting (a 23.7 percentage point gain over chain-of-thought) and 95.1% on colored objects reasoning.
These results indicate that neuro-symbolic integration provides substantial operational benefits. By shifting execution away from probabilistic text generation to deterministic code interpreters, organizations can significantly improve computational reliability without requiring larger model parameters or specialized retraining. Analysis shows that the benefit is not limited to dedicated code models, as sufficiently strong general text models like ChatGPT also gain measurable performance improvements when paired with programmatic execution.
Organizations developing reasoning or calculation workflows should adopt programmatic prompting frameworks to mitigate hallucination and arithmetic errors. Teams should prioritize clear, entity-grounded variable naming in prompts and pair models with secure runtime environments to handle exceptions. While the framework demonstrates robust accuracy across benchmark suites, implementation confidence depends on maintaining a secure execution runtime, ensuring sufficient code-generation capabilities in the base language model, and addressing edge-case handling for runtime errors.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It introduces the fundamental Chain-of-Thought prompting paradigm that PAL directly builds upon and modifies by offloading the execution step to an interpreter.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). It pioneers Program of Thoughts prompting to decouple reasoning from computation via an external Python runtime, representing the exact foundational paradigm behind PAL.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). It curates the BIG-Bench Hard suite and establishes Chain-of-Thought benchmarks across complex symbolic tasks that PAL directly targets and evaluates against.
- Paper: Program Synthesis with Large Language Models, Jacob Austin et al. (2021). It establishes early foundations and benchmarks for synthesizing functional, executable Python code with large language models to solve natural language problems.
- Paper: LILA: A Unified Benchmark for Mathematical Reasoning, Swaroop Mishra et al. (2022). It establishes a comprehensive mathematical benchmark demonstrating that delegating solutions to synthesized Python programs outperforms direct answer generation.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). It explores problem decomposition into subproblems within language models, providing crucial conceptual context for delegating multistep tasks.
- Paper: Language Models of Code are Few-Shot Commonsense Learners, Aman Madaan et al. (2022). It demonstrates that code-pretrained language models excel at structured problem decomposition and logic compared to pure text models.
- Paper: Code as Policies: Language Model Programs for Embodied Control, Jacky Liang et al. (2022). It introduces the concept of generating structured Python programs as intermediate policies for solving external, non-coding reasoning tasks.
- Paper: Toolformer: Language Models Can Teach Themselves to Use Tools, Timo Schick et al. (2023). It generalizes program-assisted execution by enabling language models to teach themselves when and how to call external tools and calculators via simple APIs.
- Paper: ReAct: Synergizing Reasoning and Acting in Language Models, Shunyu Yao et al. (2023). It extends the synergy between reasoning and execution by interleaving free-form verbal thinking traces with external tool and environment actions.
- Paper: SciAgent: Tool-augmented Language Models for Scientific Reasoning, Yubo Ma et al. (2024). It builds upon program-aided execution by retrieving and executing modular domain-specific Python functions to solve complex scientific reasoning tasks.
- Paper: Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models, Ling Yang et al. (2024). It expands upon structured thought templates by maintaining an evolving buffer of high-level problem-solving workflows across math and coding tasks.
- Paper: Recursive Language Models, Alex L. Zhang et al. (2025). It advances the programmatic execution paradigm by enabling models to recursively interact with prompts and call sub-instances inside a REPL environment.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). It evaluates the functional reliability and edge-case correctness of LLM-generated code through rigorous automated test mutation frameworks.
- Paper: LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code, Naman Jain et al. (2025). It provides a contamination-free evaluation suite across live coding competitions to benchmark modern code generation and execution abilities.
