Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Lei WangWanyu XuYihuai LanZhiqiang HuYunshi LanRoy Ka-Wei LeeEe-Peng Lim

article2023ACL668 citations

Proposes Plan-and-Solve prompting to guide large language models into breaking multi-step reasoning problems into structured subtasks, substantially reducing calculation and missing-step errors in zero-shot chain-of-thought settings to match few-shot performance.

Listen

Large language models often struggle to solve complex, multi-step reasoning problems accurately. While existing zero-shot techniques prompt models with simple phrases like "Let's think step by step" to avoid manual example creation, these approaches frequently fail due to calculation mistakes and omitted reasoning steps. The article evaluates a new zero-shot strategy called Plan-and-Solve (PS) prompting—along with an enhanced variant, PS+—which guides language models to first devise an explicit plan dividing a problem into subtasks and then execute those steps while carefully extracting variables and intermediate values.

To demonstrate the effectiveness of this method, the authors conducted experiments using OpenAI's GPT-3 model across ten benchmark datasets covering arithmetic, commonsense, and symbolic reasoning. The evaluation compared the proposed approach against standard zero-shot baselines and few-shot techniques that rely on manually crafted demonstration examples.

The findings show that PS+ prompting consistently outperforms standard zero-shot prompting across all ten datasets, improving mathematical reasoning accuracy by roughly 3 to 7 percentage points per benchmark and raising the average math accuracy from 70.4% to 76.7%. Error analysis revealed that PS+ reduced calculation errors from 7% to 5% and dropped missing-step errors from 12% to 7%. Furthermore, without using any manual reference demonstrations, zero-shot PS+ achieved an overall mathematical accuracy comparable to an 8-shot manual prompting baseline (76.7% versus 77.6%) and outperformed manual few-shot prompts on symbolic reasoning tasks such as letter concatenation (75.2% versus 70.6%).

These results demonstrate that organizations can achieve near-few-shot or state-of-the-art zero-shot reasoning performance from existing models without the costly labor of hand-crafting demonstration examples for every new domain. Incorporating structured planning instructions lowers operational error rates and enhances task reliability. However, while planning instructions effectively mitigate mechanical arithmetic and step-omission errors, they do not resolve semantic misunderstandings of complex problem text, which remained flat at around 27% across all methods.

Organizations deploying large language models for analytical or multi-step tasks should adopt structured plan-and-solve prompt templates rather than generic step-by-step instructions. For mission-critical workflows, pairing this prompting method with self-consistency voting will provide further performance gains. Future efforts should focus on refining prompt designs to address semantic misinterpretations and exploring dynamic plan correction.

Cover for Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Abstract

Large language models (LLMs) have recently been shown to deliver impressive performance in various NLP tasks. To tackle multi-step reasoning tasks, few-shot chain-of-thought (CoT) prompting includes a few manually crafted step-by-step reasoning demonstrations which enable LLMs to explicitly generate reasoning steps and improve their reasoning task accuracy. To eliminate the manual effort, Zero-shot-CoT concatenates the target problem statement with “Let’s think step by step” as an input prompt to LLMs. Despite the success of Zero-shot-CoT, it still suffers from three pitfalls: calculation errors, missing-step errors, and semantic misunderstanding errors. To address the missing-step errors, we propose Plan-and-Solve (PS) Prompting. It consists of two components: first, devising a plan to divide the entire task into smaller subtasks, and then carrying out the subtasks according to the plan. To address the calculation errors and improve the quality of generated reasoning steps, we extend PS prompting with more detailed instructions and derive PS+ prompting. We evaluate our proposed prompting strategy on ten datasets across three reasoning problems. The experimental results over GPT-3 show that our proposed zero-shot prompting consistently outperforms zero-shot-CoT across all datasets by a large margin, is comparable to or exceeds Zero-shot-Program-of-Thought Prompting, and has comparable performance with 8-shot CoT prompting on the math reasoning problem. The code can be found at https://github.com/AGI-Edgerunners/Plan-and-Solve-Prompting.

Table of Contents

  • 1 Introduction
  • 2 Plan-and-Solve Prompting
  • 2.1 Step 1: Prompting for Reasoning Generation
  • 2.2 Step 2: Prompting for Answer Extraction
  • 3 Experimental Setup
  • 3.1 Benchmarks
  • 3.2 Zero-shot and Few-shot Baselines
  • 3.3 Implementations
  • 4 Experimental Results
  • 4.1 Main Results
  • 4.2 Analysis
  • 5 Related Work
  • 5.1 Reasoning in NLP
  • 5.2 Prompting Methods
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethics
  • References
  • A Appendix
  • A.1 Results of All Trigger Sentences
  • A.2 Example Outputs by Zero-shot-PS+

Knowls

  1. Knowl 1 — Plan-and-Solve (PS) and Plan-and-Solve+ (PS+) Zero-Shot Prompting

    model/method

    Plan-and-Solve (PS) Prompting is a zero-shot prompting strategy designed to address missing-step errors and calculation errors in large language models (LLMs) performing multi-step reasoning without requiring few-shot demonstration examples.

    The framework operates in two sequential inference steps:

    1. Reasoning Generation Step: The target question XX is converted into an input prompt of the form Q: [X]. A: [Trigger], where [Trigger] instructs the model to explicitly decompose the task and carry out each subtask:

      • PS Prompting Trigger: "Let's first understand the problem and devise a plan to solve the problem. Then, let's carry out the plan and solve the problem step by step."
      • PS+ Prompting Trigger (extended with detailed variable extraction and calculation directives): "Let's first understand the problem, extract relevant variables and their corresponding numerals, and make a plan. Then, let's carry out the plan, calculate intermediate variables (pay attention to correct numerical calculation and commonsense), solve the problem step by step, and show the answer."
    2. Answer Extraction Step: The reasoning text generated by the LLM in Step 1 is appended to the prompt, followed by an extraction prompt to isolate the final answer into a standard evaluation format (e.g., "Therefore, the answer (arabic numerals) is").

    By default, generation in both steps uses greedy decoding (T=0T = 0 with 1 output chain).

  2. Knowl 2 — Arithmetic Reasoning Benchmark Accuracy across Zero-Shot and Few-Shot Methods

    data/table

    Evaluated on six arithmetic reasoning datasets using GPT-3 (text-davinci-003) with greedy decoding (temperature T=0T=0), Plan-and-Solve (PS) and Plan-and-Solve+ (PS+) zero-shot prompting consistently outperform Zero-shot Chain-of-Thought (Zero-shot-CoT) and Program-of-Thought (PoT), achieving performance comparable to few-shot demonstration-based methods.

    Method MultiArith GSM8K AddSub AQuA SingleEq SVAMP Average
    Zero-Shot
    CoT 83.8 56.4 85.3 38.9 88.1 69.9 70.4
    PoT 92.2 57.0 85.1 43.9 91.7 70.8 73.5
    PS 87.2 58.2 88.1 42.5 89.2 72.0 72.9
    PS+ 91.8 59.3 92.2 46.0 94.7 75.7 76.7
    Few-Shot
    Manual-CoT 93.6 58.4 91.6 48.4 93.5 80.3 77.6
    Auto-CoT 95.5 57.1 90.8 41.7 92.1 78.1 75.9

    Zero-shot PS+ achieves an average arithmetic accuracy of 76.7%76.7\%, outperforming Zero-shot-CoT by 6.3%6.3\% absolute margin, surpassing the automatic demonstration method Auto-CoT (75.9%75.9\%), and nearing the performance of 8-shot Manual-CoT (77.6%77.6\%) without utilizing any exemplars.

  3. Knowl 3 — Zero-Shot PS+ Prompting Accuracy on Commonsense and Symbolic Reasoning

    data/table

    Plan-and-Solve+ (PS+) prompting generalizes beyond arithmetic to commonsense reasoning and symbolic reasoning benchmarks when evaluated on GPT-3 (text-davinci-003).

    Method CSQA StrategyQA Last Letter Coin Flip
    Manual-CoT (Few-Shot) 78.3 71.2 70.6 100.0
    Zero-shot-CoT 65.2 63.8 64.8 96.8
    Zero-shot-PS+ 71.9 65.4 75.2 99.6

    On commonsense reasoning (CommonsenseQA, StrategyQA), Zero-shot PS+ outperforms Zero-shot-CoT by 6.7%6.7\% and 1.6%1.6\% respectively. On symbolic reasoning tasks (Last Letter Concatenation, Coin Flip), Zero-shot PS+ improves accuracy over Zero-shot-CoT by 10.4%10.4\% and 2.8%2.8\%, and even surpasses few-shot Manual-CoT on Last Letter Concatenation (75.2%75.2\% vs 70.6%70.6\%).

  4. Knowl 4 — Effect of Trigger Sentence Instructions on Reasoning Performance

    data/table

    Ablating the trigger instructions used in Step 1 of zero-shot prompting on GSM8K and SVAMP using GPT-3 (text-davinci-003) demonstrates the cumulative impact of specific prompt components.

    No. Trigger Sentence GSM8K SVAMP
    1 Let’s think step by step. (Zero-shot-CoT) 56.4 69.9
    2 Python solver code generation prompt (Zero-shot-PoT, code-davinci-002) 57.0 70.8
    3 Extract variables and assign their corresponding numerals to these variables first and then solve the problem step by step. 50.5 69.5
    4 Firstly, extract variables and their corresponding numerals. Then, calculate intermediate variables. Finally, solve the problem step by step. 54.8 70.8
    5 Let’s first understand the problem and devise a plan to solve the problem. Then, let’s carry out the plan and solve the problem step by step. (PS) 58.2 72.0
    6 Let’s first understand the problem, extract relevant variables and their corresponding numerals, and make a plan. Then, let’s carry out the plan, calculate intermediate variables (pay attention to correct numerical calculation and commonsense), solve the problem step by step, and show the answer. (PS+) 59.3 75.7

    Extracting variables without plan instructions (Prompt 3) drops performance below standard Zero-shot-CoT (Prompt 1). Incorporating planning instructions (Prompt 5) and adding variable extraction alongside intermediate calculation constraints (Prompt 6) yield monotonic gains, reaching peak performance on both datasets.

  5. Knowl 5 — Error Distribution of Zero-Shot Prompting Methods on GSM8K

    empirical result

    An evaluation of 100 randomly sampled grade-school math word problems from GSM8K using GPT-3 (text-davinci-003) categorized reasoning failures into three primary error types:

    • Calculation Error: Arithmetic mistakes during intermediate or final computations.
    • Missing-Step Error: Omission of necessary intermediate reasoning subtasks.
    • Semantic Misunderstanding Error: Incomprehension of problem semantics or generating logically incoherent steps.
    Method Calculation Error Missing-Step Error Semantic Misunderstanding Total Incorrect
    Zero-shot-CoT 7% 12% 27% 46%
    Zero-shot-PS 7% 10% 26% 43%
    Zero-shot-PS+ 5% 7% 27% 39%

    Zero-shot-PS reduces missing-step errors from 12%12\% to 10%10\%, and Zero-shot-PS+ reduces missing-step errors to 7%7\% and calculation errors from 7%7\% to 5%5\%. The rate of semantic misunderstanding errors remains unchanged across methods (26%–27%26\%\text{--}27\%), indicating that semantic comprehension errors are largely bounded by model capacity rather than prompt structure.

  6. Knowl 6 — Correlation between Generated Reasoning Components and Reasoning Error Types

    empirical result

    Analysis of 100 sampled GSM8K outputs generated by Zero-shot-PS+ with GPT-3 reveals the correlation between the presence of specific sub-parts in the generated reasoning text and the occurrence of different error modes:

    • Variable Definitions: The presence of explicit variable extractions negatively correlates with calculation errors (r=−0.41r = -0.41) and missing-step errors (r=−0.56r = -0.56), while displaying a positive correlation with semantic misunderstanding errors (r=0.76r = 0.76).
    • Reasoning Plan: The presence of an explicit plan negatively correlates with missing-step errors (r=−0.83r = -0.83) and calculation errors (r=−0.02r = -0.02), with a positive correlation to semantic misunderstanding (r=0.70r = 0.70).
    • Solution Steps: The solution component has a negative correlation with calculation errors (r=−0.42r = -0.42), a weak positive correlation with missing-step errors (r=0.076r = 0.076), and a positive correlation with semantic misunderstanding (r=0.24r = 0.24).

    These negative correlations confirm that explicitly prompting language models to define variables and structure a multi-step plan directly suppresses step-omission and arithmetic calculation errors.

  7. Knowl 7 — Accuracy Gains of Plan-and-Solve+ with Self-Consistency

    empirical result

    When paired with Self-Consistency (SC) decoding (generating N=10N = 10 reasoning chains at temperature T=0.7T = 0.7 and resolving the final answer via majority voting), Zero-shot-PS+ achieves substantial performance improvements over greedy decoding (T=0,N=1T = 0, N = 1):

    • GSM8K:
      • Zero-shot-CoT: 56.4%56.4\% (without SC) →70.7%\to 70.7\% (with SC)
      • Zero-shot-PS+: 58.7%58.7\% (without SC) →73.7%\to 73.7\% (with SC)
    • SVAMP:
      • Zero-shot-CoT: 69.9%69.9\% (without SC) →81.7%\to 81.7\% (with SC)
      • Zero-shot-PS+: 75.7%75.7\% (without SC) →84.4%\to 84.4\% (with SC)

    Zero-shot-PS+ with Self-Consistency outperforms Zero-shot-CoT with Self-Consistency by 3.0%3.0\% on GSM8K and 2.7%2.7\% on SVAMP.

  8. Knowl 8 — Frequency of Explicit Plan Formation in Zero-Shot PS Predictions

    empirical result

    In an empirical audit of 100 randomly sampled test instances from GSM8K evaluated under Plan-and-Solve (PS) zero-shot prompting using GPT-3 (text-davinci-003), 90 out of 100 predictions (90%90\%) explicitly generated a step-by-step plan prior to executing calculation and solution steps. This demonstrates that zero-shot instruction prompting effectively elicits latent task decomposition and planning behavior in large language models.

  9. Knowl 9 — Limitations of Plan-and-Solve Prompting

    limitation

    Plan-and-Solve (PS and PS+) prompting has two primary limitations:

    1. Sensitivity to Prompt Phrasing: Large language models are sensitive to exact lexical expressions within prompts. Crafting optimal trigger sentences requires domain-specific wording adjustments across mathematical, commonsense, and symbolic reasoning tasks.
    2. Inability to Resolve Semantic Misunderstandings: While plan-and-solve instructions reduce calculation and step-omission errors, semantic misunderstanding errors persist (affecting approximately 27%27\% of sampled GSM8K failures). Mitigating semantic errors requires underlying improvements in language model comprehension rather than zero-shot prompt formulation alone.

Coverage note — None was omitted; the knowls fully capture the PS and PS+ prompting mechanisms, all benchmark comparisons across arithmetic, commonsense, and symbolic domains, trigger sentence ablations, error and correlation analyses, self-consistency experiments, and stated limitations.

References

  1. 1.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  2. 2.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  3. 3.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  4. 4.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL, pages 4171–4186.
  6. 6.Dheeru Dua, Shivanshu Gupta, Sameer Singh, and Matt Gardner. 2022. Successive prompting for decomposing complex questions. arXiv preprint arXiv:2212.04092.
  7. 7.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  8. 8.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies. TACL, 9:346–361.
  9. 9.Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2021. Towards a unified view of parameter-efficient transfer learning. arXiv preprint arXiv:2110.04366.
  10. 10.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. Measuring mathematical problem solving with the math dataset. arXiv preprint arXiv:2103.03874.
  11. 11.Mohammad Javad Hosseini, Hannaneh Hajishirzi, Oren Etzioni, and Nate Kushman. 2014. Learning to solve arithmetic word problems with verb categorization. In EMNLP, pages 523–533.
  12. 12.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for nlp. In International Conference on Machine Learning, pages 2790–2799. PMLR.
  13. 13.Jie Huang and Kevin Chen-Chuan Chang. 2022. Towards reasoning in large language models: A survey. arXiv preprint arXiv:2212.10403.
  14. 14.Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning, pages 9118–9147. PMLR.
  15. 15.Tushar Khot, Harsh Trivedi, Matthew Finlayson, Yao Fu, Kyle Richardson, Peter Clark, and Ashish Sabharwal. 2022. Decomposed prompting: A modular approach for solving complex tasks. arXiv preprint arXiv:2210.02406.
  16. 16.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  17. 17.Rik Koncel-Kedziorski, Hannaneh Hajishirzi, Ashish Sabharwal, Oren Etzioni, and Siena Dumas Ang. 2015. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics, 3:585–597.
  18. 18.Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In Proceedings of NAACL, pages 1152–1157.
  19. 19.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  20. 20.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 158–167.
  21. 21.Bo Liu, Yuqian Jiang, Xiaohan Zhang, Qiang Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone. 2023. Llm+ p: Empowering large language models with optimal planning proficiency. arXiv preprint arXiv:2304.11477.
  22. 22.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  23. 23.Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and Ashwin Kalyan. 2022. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. arXiv preprint arXiv:2209.14610.
  24. 24.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are NLP models really able to solve simple math word problems? In Proceedings of NAACL, pages 2080–2094.
  25. 25.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350.
  26. 26.Subhro Roy and Dan Roth. 2016. Solving general arithmetic word problems. arXiv preprint arXiv:1608.01413.
  27. 27.Abulhair Saparov and He He. 2022. Language models are greedy reasoners: A systematic formal analysis of chain-of-thought. arXiv preprint arXiv:2210.01240.
  28. 28.Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2022. On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning. arXiv preprint arXiv:2212.08061.
  29. 29.Simeng Sun, Yang Liu, Shuohang Wang, Chenguang Zhu, and Mohit Iyyer. 2023. Pearl: Prompting large language models to plan and execute actions over long documents.
  30. 30.Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu. 2022. Black-box tuning for language-model-as-a-service. arXiv preprint arXiv:2201.03514.
  31. 31.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261.
  32. 32.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of NAACL-HLT, pages 4149–4158.
  33. 33.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  34. 34.Boshi Wang, Sewon Min, Xiang Deng, Jiaming Shen, You Wu, Luke Zettlemoyer, and Huan Sun. 2022a. Towards understanding chain-of-thought prompting: An empirical study of what matters. arXiv preprint arXiv:2212.10001.
  35. 35.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022b. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  36. 36.Zihao Wang, Shaofei Cai, Anji Liu, Xiaojian Ma, and Yitao Liang. 2023. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. arXiv preprint arXiv:2302.01560.
  37. 37.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022a. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  38. 38.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. In Thirty-sixth Conference on Neural Information Processing Systems (NeurIPS 2022).
  39. 39.Yixuan Weng, Minjun Zhu, Shizhu He, Kang Liu, and Jun Zhao. 2022. Large language models are reasoners with self-verification. arXiv preprint arXiv:2212.09561.
  40. 40.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models.
  41. 41.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. ArXiv, abs/2210.03629.
  42. 42.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  43. 43.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.

Citation

MLA
Wang, L., et al. “Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 2609–34, https://doi.org/10.18653/v1/2023.acl-long.147.
APA
Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., & Lim, E.-P. (2023). Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2609–2634. https://doi.org/10.18653/v1/2023.acl-long.147
Chicago
Wang, L., W. Xu, Y. Lan, et al. 2023. “Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2609–34. https://doi.org/10.18653/v1/2023.acl-long.147.
Harvard
Wang, L. et al. (2023) “Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2609–2634. Available at: https://doi.org/10.18653/v1/2023.acl-long.147.
Vancouver
1. Wang L, Xu W, Lan Y, Hu Z, Lan Y, Lee RK-W, Lim E-P (2023) Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2609–2634

BibTeX

@inproceedings{wang-etal-2023-plan,
    title = "Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models",
    author = "Wang, Lei  and
      Xu, Wanyu  and
      Lan, Yihuai  and
      Hu, Zhiqiang  and
      Lan, Yunshi  and
      Lee, Roy Ka-Wei  and
      Lim, Ee-Peng",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.147/",
    doi = "10.18653/v1/2023.acl-long.147",
    pages = "2609--2634"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/