PAL: Program-aided Language Models

Luyu GaoAman MadaanShuyan ZhouUri AlonPengfei LiuYiming YangJamie CallanGraham Neubig

article2023ICML678 citations

Proposes Program-aided Language models (PAL), a method that prompts language models to generate executable code for intermediate reasoning and offloads computation to a Python interpreter, significantly outperforming standard chain-of-thought approaches on complex mathematical and symbolic benchmarks.

Listen

Large language models frequently struggle with complex arithmetic and multi-step logical tracking, even when using step-by-step chain-of-thought prompting. Although these models effectively decompose problems into intermediate steps, they often fail to execute the calculations or maintain correct state tracking during the solution phase. The article introduces and evaluates Program-Aided Language Models (PAL), a framework designed to resolve this execution bottleneck by delegating the intermediate calculation steps to an external Python interpreter.

Under the PAL approach, a language model translates a natural language problem into executable Python code with meaningful variable names rather than generating free-form text answers. The article evaluates this method across 13 benchmarks covering mathematical word problems, symbolic reasoning, and algorithmic execution tasks. In experimental setups primarily using Codex, PAL established new few-shot performance benchmarks, consistently outperforming standard prompting methods and much larger language models.

Key findings show significant performance gains across diverse domains. On the GSM8K math benchmark, PAL achieved a 72.0% solve rate, surpassing PaLM-540B with chain-of-thought prompting by 15.1 percentage points, and reached 80.4% when paired with majority voting across 40 samples. On GSM-HARD, an evaluation set containing large multi-digit numbers, chain-of-thought accuracy collapsed from 65.6% to 23.1%, whereas PAL maintained a 61.2% solve rate. PAL also delivered substantial improvements in symbolic and algorithmic tasks, reaching 96.7% on object counting (a 23.7 percentage point gain over chain-of-thought) and 95.1% on colored objects reasoning.

These results indicate that neuro-symbolic integration provides substantial operational benefits. By shifting execution away from probabilistic text generation to deterministic code interpreters, organizations can significantly improve computational reliability without requiring larger model parameters or specialized retraining. Analysis shows that the benefit is not limited to dedicated code models, as sufficiently strong general text models like ChatGPT also gain measurable performance improvements when paired with programmatic execution.

Organizations developing reasoning or calculation workflows should adopt programmatic prompting frameworks to mitigate hallucination and arithmetic errors. Teams should prioritize clear, entity-grounded variable naming in prompts and pair models with secure runtime environments to handle exceptions. While the framework demonstrates robust accuracy across benchmark suites, implementation confidence depends on maintaining a secure execution runtime, ensuring sufficient code-generation capabilities in the base language model, and addressing edge-case handling for runtime errors.

Cover for PAL: Program-aided Language Models

Abstract

Large language models (LLMs) have demonstrated an impressive ability to perform arithmetic and symbolic reasoning tasks, when provided with a few examples at test time (“few-shot prompting”). Much of this success can be attributed to prompting methods such as “chain-of-thought”, which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the problem. While LLMs seem to be adept at this sort of step-by-step decomposition, LLMs often make logical and arithmetic mistakes in the solution part, even when the problem is decomposed correctly. In this paper, we present Program-Aided Language models (PAL): a novel approach that uses the LLM to read natural language problems and generate programs as the intermediate reasoning steps, but offloads the solution step to a runtime such as a Python interpreter. With PAL, decomposing the natural language problem into runnable steps remains the only learning task for the LLM, while solving is delegated to the interpreter. We demonstrate this synergy between a neural LLM and a symbolic interpreter across 13 mathematical, symbolic, and algorithmic reasoning tasks from BIG-Bench Hard and others. In all these natural language reasoning tasks, generating code using an LLM and reasoning using a Python interpreter leads to more accurate results than much larger models. For example, PAL using CODEX achieves state-of-the-art few-shot accuracy on GSM8K, surpassing PaLM-540B which uses chain-of-thought by absolute 15% top-1.

Table of Contents

  • 1. Introduction
  • 2. Background: Few-shot Prompting
  • 3. Program-aided Language Models
  • 4. Experimental Setup
  • 4.1. Mathematical Reasoning
  • 4.2. Symbolic Reasoning
  • 4.3. Algorithmic Tasks
  • 5. Results
  • 5.1. Math Results
  • 5.2. Symbolic Reasoning & Algorithmic Tasks Results
  • 6. Analysis
  • 7. Related Work
  • 8. Conclusion
  • 9. Acknowledgement
  • References
  • Appendix
  • A. Alternative Prompts without Meaningful Variable Names
  • B. Additional analysis on Arithmetic Reasoning
  • C. Effect of Using Language Models of Code
  • D. Experiments with ChatGPT
  • E. Analyzing the Effect of Increasing Number of Samples on PAL
  • F. Standard Deviations Across Multiple Order of Prompts
  • G. PAL Beyond Benchmarks
  • H. Closer Look into Token-level Behaviors of Different Mechanisms
  • I. Datasets
  • I.1. Creating GSM-HARD
  • I.2. GSM-HARD Analysis
  • J. Generalization of PAL to Least-to-Most Prompting
  • K. Prompts
  • K.1. Reasoning about Colored Objects
  • K.2. Penguins in a Table
  • K.3. Date Understanding
  • K.4. Math
  • K.5. Object Counting
  • K.6. Repeat Copy
  • L. Success and Failure Modes in Symbolic Tasks
  • L.1. Colored Objects
  • L.2. Penguins in a Table
  • L.3. Date Understanding

Knowls

  1. Knowl 1 — Program-Aided Language Models (PAL)

    model/method

    Program-Aided Language Models (PAL) is a neuro-symbolic few-shot prompting framework for solving reasoning tasks in natural language (NL). Unlike standard Chain-of-Thought (COT) prompting where a large language model (LLM) generates both the reasoning rationale and the final calculation/solution in free-form text, PAL uses the LLM to read the problem and generate a program (such as a Python script) containing the reasoning steps, while delegating program execution to an external symbolic runtime (such as a standard Python interpreter).

    In PAL, each few-shot in-context example is a pair ⟨xi,ti⟩\langle x_i, t_i \rangle, where xix_i is a natural language problem and ti=[s1,s2,…,sN]t_i = [s_1, s_2, \dots, s_N] is a sequence of intermediate programmatic statements sj∈NL∪PLs_j \in \text{NL} \cup \text{PL} (interleaved natural language comments and executable programming statements). Importantly, tit_i does not provide hard-coded final textual answers; it terminates in a program expression that produces the answer.

    Given a prompt consisting of kk concatenated examples p≡⟨x1⋅t1⟩∥⟨x2⋅t2⟩∥⋯∥⟨xk⋅tk⟩p \equiv \langle x_1 \cdot t_1 \rangle \parallel \langle x_2 \cdot t_2 \rangle \parallel \dots \parallel \langle x_k \cdot t_k \rangle and a new test query xtestx_{\text{test}}, the prompt p∥xtestp \parallel x_{\text{test}} is passed to the LLM to generate code ttestt_{\text{test}}. The generated code ttestt_{\text{test}} is executed by the interpreter to obtain the final run result ytesty_{\text{test}}.

  2. Knowl 2 — GSM-HARD Benchmark for Large-Number Arithmetic Reasoning

    experimental setup

    GSM-HARD is an arithmetic reasoning evaluation dataset designed to test whether language models can generalize reasoning to large numerical values. In the original GSM8K benchmark, approximately 50%50\% of numbers are integers between 0 and 8. GSM-HARD modifies the 1,319 question-answer pairs of the GSM8K test set by replacing one or more numerical values in each question with large random integers of up to 7 digits using pattern matching.

    Reference answers for GSM-HARD were created using a semi-automated pipeline:

    1. For problems where PAL generated a correct program on the original GSM8K question (71%71\% of the dataset), the program's initial variable assignments were updated with the newly sampled large integers and executed to yield the ground-truth answer.
    2. For the remaining 29%29\%, PAL was run with nucleus sampling (p=0.95p = 0.95, temperature T=0.7T = 0.7) to search for correct programs.
    3. The remaining 50 unresolved instances after 100 sampling iterations were manually annotated.
  3. Knowl 3 — Mathematical Reasoning Accuracy Across Benchmarks

    data/table

    The performance of PAL was evaluated on eight mathematical word problem benchmarks using Codex (code-davinci-002) with greedy decoding (temperature 00). Results reflect problem solve rates (%) averaged over three prompt orderings:

    Method GSM8K GSM-HARD SVAMP ASDIV SINGLEEQ SINGLEOP ADDSUB MULTIARITH
    DIRECT (Codex) 19.7 5.0 69.9 74.0 86.8 93.1 90.9 44.0
    COT (UL2-20B) 4.1 - 12.6 16.9 - - 18.2 10.7
    COT (LaMDA-137B) 17.1 - 39.9 49.0 - - 52.9 51.8
    COT (Codex) 65.6 23.1 74.8 76.9 89.1 91.9 86.0 95.9
    COT (PaLM-540B) 56.9 - 79.0 73.9 92.3 94.1 91.9 94.7
    COT (Minerva 540B) 58.8 - - - - - - -
    PAL (Codex) 72.0 61.2 79.4 79.6 96.1 94.6 92.5 99.2

    PAL with Codex outperforms all chain-of-thought (COT) models across all eight benchmarks, including PaLM-540B and Minerva-540B (which is fine-tuned on 164B tokens of mathematical text). On GSM-HARD, COT accuracy drops from 65.6%65.6\% to 23.1%23.1\% (a relative drop of ∼65%\sim 65\%), whereas PAL accuracy drops only from 72.0%72.0\% to 61.2%61.2\% (a relative drop of 15%15\%), demonstrating robust generalization to large-number arithmetic.

  4. Knowl 4 — Symbolic and Algorithmic Reasoning Performance

    data/table

    PAL was evaluated against DIRECT prompting and Chain-of-Thought (COT) across three symbolic reasoning datasets (COLORED OBJECTS, PENGUINS, DATE) and two algorithmic datasets (REPEAT COPY, OBJECT COUNTING) from BIG-Bench Hard, using Codex (code-davinci-002) with greedy decoding:

    Method COLORED OBJECT PENGUINS DATE REPEAT COPY OBJECT COUNTING
    DIRECT (Codex) 75.7 71.1 49.9 81.3 37.6
    COT (LaMDA-137B) - - 26.8 - -
    COT (PaLM-540B) - 65.1 65.3 - -
    COT (Codex) 86.3 79.2 64.8 68.8 73.0
    PAL (Codex) 95.1 93.3 76.2 90.6 96.7

    PAL achieves superior performance on all five tasks. On OBJECT COUNTING and REPEAT COPY, PAL outperforms COT by absolute margins of 23.7%23.7\% and 21.8%21.8\% respectively. On PENGUINS and COLORED OBJECTS, PAL achieves absolute gains of 14.1%14.1\% and 8.8%8.8\% over COT.

  5. Knowl 5 — Ablation of External Runtime Execution and Step-by-Step Code Decomposition

    empirical result

    Ablation experiments on GSM8K using Codex evaluate whether PAL's advantages stem from its code-formatted prompts or from offloading execution to an interpreter:

    Ablation Setting GSM8K Solve Rate (%)
    DIRECT (no intermediate reasoning) 19.7
    COT (free-form natural language) 65.6
    PAL (step-by-step code + Python runtime) 72.0
    Succinct Code (single-line expression + Python runtime) 47.8
    LLM Simulating Runtime (step-by-step code without Python runtime) 23.2

    When the LLM is prompted to generate the code reasoning chain and then directly predict the answer itself without an interpreter (LLM Simulating Runtime), performance drops from 72.0%72.0\% to 23.2%23.2\%, near DIRECT prompting (19.7%19.7\%). When the reasoning is compressed into a single-line arithmetic expression executed by the interpreter (Succinct Code), accuracy falls to 47.8%47.8\%. This confirms that both step-by-step programmatic decomposition and symbolic runtime execution are necessary for PAL's performance.

  6. Knowl 6 — Role of Meaningful Identifier Names and Natural Language Comments in PAL

    empirical result

    Ablations on PAL prompt formats demonstrate the impact of natural language comments and semantic variable naming across reasoning tasks:

    1. Standard PAL: Prompts contain step-by-step Python statements, meaningful variable names (e.g., num_apples_in_basket), and natural language comments.
    2. PAL−comment\text{PAL}_{-\text{comment}}: Intermediate natural language comments are removed, keeping meaningful variable names.
    3. PAL−var−comment\text{PAL}_{-\text{var}-\text{comment}}: Intermediate natural language comments are removed and all variable names are replaced with uninformative single-letter or random characters.

    On GSM8K, solve rates are:

    • COT: 63.1%63.1\%
    • PAL−var\text{PAL}_{-\text{var}}: 59.0%59.0\%
    • PAL−var+comms\text{PAL}_{-\text{var}+\text{comms}}: 69.0%69.0\%
    • PAL: 71.8%71.8\%

    On COLORED OBJECTS, accuracy is 84.4%84.4\% (COT), 95.2%95.2\% (PAL), 91.1%91.1\% (PAL−comment\text{PAL}_{-\text{comment}}), and 79.9%79.9\% (PAL−var−comment\text{PAL}_{-\text{var}-\text{comment}}). On DATE, accuracy is 64.8%64.8\% (COT), 76.2%76.2\% (PAL), 69.1%69.1\% (PAL−comment\text{PAL}_{-\text{comment}}), and 63.4%63.4\% (PAL−var−comment\text{PAL}_{-\text{var}-\text{comment}}).

    Meaningful variable names ground entities in the problem description; stripping variable semantics degrades performance below standard Chain-of-Thought.

  7. Knowl 7 — Performance Across Different Model Families and Model Scales

    empirical result

    PAL's performance was measured across different base language models on GSM8K:

    1. Weaker Code LLMs:

      • code-cushman-001: COT solves 19.1%19.1\%, PAL solves 21.7%21.7\% (13.6%13.6\% relative improvement).
      • code-davinci-001: COT solves 26.0%26.0\%, PAL solves 31.8%31.8\% (22.3%22.3\% relative improvement).
      • code-davinci-002: COT solves 60.1%60.1\%, PAL solves 72.0%72.0\% (19.8%19.8\% relative improvement).
    2. Natural Language / Text LLMs:

      • text-davinci-001: COT achieves 26.5%26.5\%, PAL achieves 8.6%8.6\% (weak code modeling hurts PAL).
      • text-davinci-002: COT achieves 46.9%46.9\%, PAL achieves 65.8%65.8\%.
      • text-davinci-003: COT achieves 65.3%65.3\%, PAL achieves 69.8%69.8\%.
      • ChatGPT (gpt-3.5-turbo): COT achieves 76.8%76.8\%, PAL achieves 79.6%79.6\% (+2.8%+2.8\% absolute).

    PAL is not restricted to code-dedicated models; text-based models with sufficient code generation capabilities also benefit significantly from PAL.

  8. Knowl 8 — Multi-Sample Generation with Majority Voting for PAL

    empirical result

    When combining PAL with self-consistency multi-sample majority voting (extmajority@k ext{majority}@k) using nucleus sampling (p=0.95p = 0.95, temperature T=0.7T = 0.7, k=40k = 40 sampled solutions), PAL with Codex achieves state-of-the-art accuracy on GSM8K:

    Method GSM8K majority@40 (%)
    COT (UL2-20B) 7.3
    COT (LaMDA-137B) 27.7
    COT (PaLM-540B) 74.4
    COT (Codex) 78.0
    COT (Minerva 540B) 78.5
    PAL (Codex) 80.4

    Multi-sample majority voting increases PAL Codex accuracy from 72.0%72.0\% (greedy decoding) to 80.4%80.4\%, outperforming Minerva-540B (78.5%78.5\%) with the same sample budget.

  9. Knowl 9 — Generalization of PAL to Least-to-Most Prompting

    empirical result

    PAL can be integrated into two-stage Least-to-Most prompting. In Least-to-Most prompting, the problem-reduction stage recursively decomposes a question into sub-questions using natural language, and the problem-solving stage sequentially solves each sub-question. Under the Least-to-Most + PAL setup, the natural language reduction prompt is retained, while the solution steps in the solving scripts are replaced with programmatic statements.

    On 500-example subsets of GSM8K and SVAMP, Least-to-Most + PAL consistently outperforms standard Least-to-Most:

    • GSM8K: Least-to-Most achieves 67.2%67.2\%; Least-to-Most + PAL achieves 72.8%72.8\% (+5.6%+5.6\%).
    • SVAMP: Least-to-Most achieves 75.2%75.2\%; Least-to-Most + PAL achieves 78.2%78.2\% (+3.0%+3.0\%).

    Because variable bindings are preserved across the programmatic script, PAL automatically shares intermediate symbolic values across sub-problems without needing to explicitly restate answers in text.

  10. Knowl 10 — Robustness to Complexity and Token-Level Confidence in Symbolic Reasoning

    empirical result

    On the COLORED OBJECTS task, varying the number of objects from 0 to 26 reveals that while Chain-of-Thought (COT) accuracy degrades and fluctuates significantly as the object count increases, PAL accuracy remains consistently close to 100%100\%.

    Token-level log-likelihood analysis on 20 randomly selected examples from COLORED OBJECTS reveals key differences in model behavior:

    • In COT reasoning chains, the LLM exhibits low log-likelihood (low confidence) on tokens representing numbers and quantitative quantities (7/20 examples), spatial grounding adjectives like 'right-most' (6/20 examples), object colors (2/20 examples), and object nouns (6/20 examples).
    • In PAL reasoning chains, the model uses standard programmatic idioms (e.g., len(objects) and fixed list indexing objects[-1]), generating them with high and uniform token log-likelihood. PAL translates varied linguistic contexts into invariant programmatic structures.

Coverage note — None was omitted. All primary contributions—including the PAL methodology, the GSM-HARD benchmark, full empirical tables across math/symbolic/algorithmic benchmarks, ablations (runtime offloading, succinct code, variable names/comments), multi-sample voting, base LM variations, Least-to-Most integration, and token confidence analyses—are fully captured.

References

  1. 1.Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Ho, D., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jang, E., Ruano, R. J., Jeffrey, K., Jesmonth, S., Joshi, N. J., Julian, R., Kalashnikov, D., Kuang, Y., Lee, K.-H., Levine, S., Lu, Y., Luu, L., Parada, C., Pastor, P., Quiambao, J., Rao, K., Rettinghouse, J., Reyes, D., Sermanet, P., Sievers, N., Tan, C., Toshev, A., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Xu, S., Yan, M., and Zeng, A. Do as I Can, not as I Say: Grounding Language in Robotic Affordances. arXiv preprint arXiv:2204.01691, 2022.
  2. 2.Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H. MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms. In ACL, 2019.
  3. 3.Andor, D., He, L., Lee, K., and Pitler, E. Giving bert a calculator: Finding operations and arguments with reading comprehension. arXiv preprint arXiv:1909.00109, 2019.
  4. 4.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language Models are Few-Shot Learners. In NeurIPS, 2020.
  5. 5.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating Large Language Models Trained on Code. arXiv preprint arXiv:2107.03374, 2021a.
  6. 6.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021b.
  7. 7.Chen, W., Ma, X., Wang, X., and Cohen, W. W. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588, 2022.
  8. 8.Cheng, Z., Xie, T., Shi, P., Li, C., Nadkarni, R., Hu, Y., Xiong, C., Radev, D., Ostendorf, M., Zettlemoyer, L., Smith, N. A., and Yu, T. Binding language models in symbolic languages. arXiv preprint arXiv:2210.02875, 2022.
  9. 9.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N. PaLM: Scaling Language Modeling with Pathways. arXiv preprint arXiv:2204.02311, 2022.
  10. 10.Cobbe, K., Kosaraju, V., Bavarian, M., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168, 2021.
  11. 11.Demeter, D. and Downey, D. Just add functions: A neural-symbolic language model. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 7634–7642, 2020.
  12. 12.Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al. Language model cascades. arXiv preprint arXiv:2207.10342, 2022.
  13. 13.Garcez, A. d. and Lamb, L. C. Neurosymbolic ai: the 3rd wave. arXiv preprint arXiv:2012.05876, 2020.
  14. 14.Gehrmann, S., Adewumi, T., Aggarwal, K., Ammanamanchi, P. S., Anuoluwapo, A., Bosselut, A., Chandu, K. R., Clinciu, M., Das, D., Dhole, K. D., Du, W., Durmus, E., Dusek, O., Emezue, C., Gangal, V., Garbacea, C., Hashimoto, T., Hou, Y., Jernite, Y., Jhamtani, H., Ji, Y., Jolly, S., Kale, M., Kumar, D., Ladhak, F., Madaan, A., Maddela, M., Mahajan, K., Mahamood, S., Majumder, B. P., Martins, P. H., McMillan-Major, A., Mille, S., van Miltenburg, E., Nadeem, M., Narayan, S., Nikolaev, V., Niyongabo, R. A., Osei, S., Parikh, A., Perez-Beltrachini, L., Rao, N. R., Raunak, V., Rodriguez, J. D., Santhanam, S., Sedoc, J., Sellam, T., Shaikh, S., Shimorina, A., Cabezudo, M. A. S., Strobelt, H., Subramani, N., Xu, W., Yang, D., Yerukola, A., and Zhou, J. The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics. arXiv preprint arXiv:2102.01672, 2021.
  15. 15.Gellenbeck, E. M. and Cook, C. R. An investigation of procedure and variable names as beacons during program comprehension. In Empirical studies of programmers: Fourth workshop, pp. 65–81. Ablex Publishing, Norwood, NJ, 1991.
  16. 16.Geva, M., Gupta, A., and Berant, J. Injecting numerical reasoning skills into language models. In ACL, 2020.
  17. 17.Gupta, N., Lin, K., Roth, D., Singh, S., and Gardner, M. Neural module networks for reasoning over text. arXiv preprint arXiv:1912.04971, 2019.
  18. 18.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the MATH dataset, 2021. URL https://openreview.net/forum?id=7Bywt2mQsCe.
  19. 19.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The Curious Case of Neural Text Degeneration. In ICLR, 2019.
  20. 20.Koncel-Kedziorski, R., Roy, S., Amini, A., Kushman, N., and Hajishirzi, H. Mawps: A math word problem repository. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1152–1157, 2016.
  21. 21.Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., Wu, Y., Neyshabur, B., Gur-Ari, G., and Misra, V. Solving quantitative reasoning problems with language models. arXiv preprint arXiv:2206.14858, 2022.
  22. 22.Ling, W., Yogatama, D., Dyer, C., and Blunsom, P. Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems. arXiv preprint arXiv:1705.04146, 2017.
  23. 23.Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. arXiv preprint arXiv:2107.13586, 2021.
  24. 24.Madaan, A. and Yazdanbakhsh, A. Text and patterns: For effective chain of thought, it takes two to tango. arXiv preprint arXiv:2209.07686, 2022.
  25. 25.Madaan, A., Zhou, S., Alon, U., Yang, Y., and Neubig, G. Language models of code are few-shot commonsense learners. arXiv preprint arXiv:2210.07128, 2022.
  26. 26.Marcus, G. Deep learning: A critical appraisal. arXiv preprint arXiv:1801.00631, 2018.
  27. 27.Marcus, G. The next decade in ai: four steps towards robust artificial intelligence. arXiv preprint arXiv:2002.06177, 2020.
  28. 28.Miao, S.-y., Liang, C.-C., and Su, K.-Y. A diverse corpus for evaluating and developing English math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 975–984, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.92. URL https://aclanthology.org/2020.acl-main.92.
  29. 29.Mishra, S., Finlayson, M., Lu, P., Tang, L., Welleck, S., Baral, C., Rajpurohit, T., Tafjord, O., Sabharwal, A., Clark, P., and Kalyan, A. Lila: A unified benchmark for mathematical reasoning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2022.
  30. 30.Nogueira, R., Jiang, Z., and Lin, J. Investigating the limitations of transformers with simple arithmetic tasks. arXiv preprint arXiv:2102.13019, 2021.
  31. 31.Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A. Show your Work: Scratchpads for Intermediate Computation with Language Models. arXiv preprint arXiv:2112.00114, 2021a.
  32. 32.Nye, M., Tessler, M., Tenenbaum, J., and Lake, B. M. Improving coherence and consistency in neural sequence models with dual-system, neuro-symbolic reasoning. Advances in Neural Information Processing Systems, 34: 25192–25204, 2021b.
  33. 33.Patel, A., Bhattamishra, S., and Goyal, N. Are NLP Models Really Able to Solve Simple Math Word Problems? arXiv preprint arXiv:2103.07191, 2021.
  34. 34.Pi, X., Liu, Q., Chen, B., Ziyadi, M., Lin, Z., Gao, Y., Fu, Q., Lou, J.-G., and Chen, W. Reasoning like program executors. arXiv preprint arXiv:2201.11473, 2022.
  35. 35.Qian, J., Wang, H., Li, Z., Li, S., and Yan, X. Limitations of language models in arithmetic and symbolic induction. arXiv preprint arXiv:2208.05051, 2022.
  36. 36.Reif, E., Ippolito, D., Yuan, A., Coenen, A., Callison-Burch, C., and Wei, J. A Recipe for Arbitrary Text Style Transfer with Large Language Models. arXiv preprint arXiv:2109.03910, 2021.
  37. 37.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Biderman, S., Gao, L., Bers, T., Wolf, T., and Rush, A. M. Multitask Prompted Training Enables Zero-Shot Task Generalization, 2021.
  38. 38.Shin, R. and Van Durme, B. Few-shot semantic parsing with language models trained on code. arXiv preprint arXiv:2112.08696, 2021.
  39. 39.Shin, R., Lin, C. H., Thomson, S., Chen, C., Roy, S., Platanios, E. A., Pauls, A., Klein, D., Eisner, J., and Van Durme, B. Constrained language models yield few-shot semantic parsers. arXiv preprint arXiv:2104.08768, 2021.
  40. 40.Suzgun, M., Scales, N., Scharli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E., Zhou, D., and Wei, J. Challenging big-bench tasks and whether chain-of-thought can solve them. ArXiv, abs/2210.09261, 2022.
  41. 41.Takang, A. A., Grubb, P. A., and Macredie, R. D. The effects of comments and identifier names on program comprehensibility: an experimental investigation. J. Prog. Lang., 4(3):143–167, 1996.
  42. 42.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D. Rationale-Augmented Ensembles in Language Models. arXiv preprints arXiv:2207.00747, 2022a.
  43. 43.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv preprint arXiv:2203.11171, 2022b.
  44. 44.Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. Finetuned Language Models are Zero-shot Learners. arXiv preprint arXiv:2109.01652, 2021.
  45. 45.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D. Chain of Thought Prompting Elicits Reasoning in Large Language Models. arXiv preprint arXiv:2201.11903, 2022.
  46. 46.Wu, Y., Jiang, A. Q., Li, W., Rabe, M. N., Staats, C., Jamnik, M., and Szegedy, C. Autoformalization with Large Language Models. arXiv preprint arXiv:2205.12615, 2022.
  47. 47.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022.
  48. 48.Zhou, D., Scharli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Bousquet, O., Le, Q., and Chi, E. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. arXiv preprint arXiv:2205.10625, 2022.

Citation

MLA
Gao, L., et al. “PAL: Program-aided Language Models”. International Conference on Machine Learning, vol. 202, 2023, pp. 10764–99, https://proceedings.mlr.press/v202/gao23f.html.
APA
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., & Neubig, G. (2023). PAL: Program-aided Language Models. International Conference on Machine Learning, 202, 10764–10799. https://proceedings.mlr.press/v202/gao23f.html
Chicago
Gao, L., A. Madaan, S. Zhou, et al. 2023. “PAL: Program-aided Language Models”. International Conference on Machine Learning 202: 10764–99. https://proceedings.mlr.press/v202/gao23f.html.
Harvard
Gao, L. et al. (2023) “PAL: Program-aided Language Models”, International Conference on Machine Learning. PMLR, pp. 10764–10799. Available at: https://proceedings.mlr.press/v202/gao23f.html.
Vancouver
1. Gao L, Madaan A, Zhou S, Alon U, Liu P, Yang Y, Callan J, Neubig G (2023) PAL: Program-aided Language Models. In: International Conference on Machine Learning. PMLR, pp 10764–10799

BibTeX

@InProceedings{pmlr-v202-gao23f,
  title = 	 {{PAL}: Program-aided Language Models},
  author =       {Gao, Luyu and Madaan, Aman and Zhou, Shuyan and Alon, Uri and Liu, Pengfei and Yang, Yiming and Callan, Jamie and Neubig, Graham},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {10764--10799},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/gao23f/gao23f.pdf},
  url = 	 {https://proceedings.mlr.press/v202/gao23f.html},
  abstract = 	 {Large language models (LLMs) have demonstrated an impressive ability to perform arithmetic and symbolic reasoning tasks, when provided with a few examples at test time ("few-shot prompting"). Much of this success can be attributed to prompting methods such as "chain-of-thought", which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the problem. While LLMs seem to be adept at this sort of step-by-step decomposition, LLMs often make logical and arithmetic mistakes in the solution part, even when the problem is decomposed correctly. In this paper, we present Program-Aided Language models (PAL): a novel approach that uses the LLM to read natural language problems and generate programs as the intermediate reasoning steps, but offloads the solution step to a runtime such as a Python interpreter. With PAL, decomposing the natural language problem into runnable steps remains the only learning task for the LLM, while solving is delegated to the interpreter. We demonstrate this synergy between a neural LLM and a symbolic interpreter across 13 mathematical, symbolic, and algorithmic reasoning tasks from BIG-Bench Hard and others. In all these natural language reasoning tasks, generating code using an LLM and reasoning using a Python interpreter leads to more accurate results than much larger models. For example, PAL using Codex achieves state-of-the-art few-shot accuracy on GSM8K, surpassing PaLM which uses chain-of-thought by absolute 15% top-1.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/