Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models
Lei WangWanyu XuYihuai LanZhiqiang HuYunshi LanRoy Ka-Wei LeeEe-Peng Lim
Proposes Plan-and-Solve prompting to guide large language models into breaking multi-step reasoning problems into structured subtasks, substantially reducing calculation and missing-step errors in zero-shot chain-of-thought settings to match few-shot performance.
Large language models often struggle to solve complex, multi-step reasoning problems accurately. While existing zero-shot techniques prompt models with simple phrases like "Let's think step by step" to avoid manual example creation, these approaches frequently fail due to calculation mistakes and omitted reasoning steps. The article evaluates a new zero-shot strategy called Plan-and-Solve (PS) prompting—along with an enhanced variant, PS+—which guides language models to first devise an explicit plan dividing a problem into subtasks and then execute those steps while carefully extracting variables and intermediate values.
To demonstrate the effectiveness of this method, the authors conducted experiments using OpenAI's GPT-3 model across ten benchmark datasets covering arithmetic, commonsense, and symbolic reasoning. The evaluation compared the proposed approach against standard zero-shot baselines and few-shot techniques that rely on manually crafted demonstration examples.
The findings show that PS+ prompting consistently outperforms standard zero-shot prompting across all ten datasets, improving mathematical reasoning accuracy by roughly 3 to 7 percentage points per benchmark and raising the average math accuracy from 70.4% to 76.7%. Error analysis revealed that PS+ reduced calculation errors from 7% to 5% and dropped missing-step errors from 12% to 7%. Furthermore, without using any manual reference demonstrations, zero-shot PS+ achieved an overall mathematical accuracy comparable to an 8-shot manual prompting baseline (76.7% versus 77.6%) and outperformed manual few-shot prompts on symbolic reasoning tasks such as letter concatenation (75.2% versus 70.6%).
These results demonstrate that organizations can achieve near-few-shot or state-of-the-art zero-shot reasoning performance from existing models without the costly labor of hand-crafting demonstration examples for every new domain. Incorporating structured planning instructions lowers operational error rates and enhances task reliability. However, while planning instructions effectively mitigate mechanical arithmetic and step-omission errors, they do not resolve semantic misunderstandings of complex problem text, which remained flat at around 27% across all methods.
Organizations deploying large language models for analytical or multi-step tasks should adopt structured plan-and-solve prompt templates rather than generic step-by-step instructions. For mission-critical workflows, pairing this prompting method with self-consistency voting will provide further performance gains. Future efforts should focus on refining prompt designs to address semantic misinterpretations and exploring dynamic plan correction.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). This foundational work introduces Zero-shot-CoT via the "Let's think step by step" prompt, providing the direct baseline and setting that Plan-and-Solve Prompting aims to improve.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It establishes the underlying Chain-of-Thought prompting paradigm that demonstrates how eliciting intermediate steps improves multi-step reasoning in large language models.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). This paper presents problem decomposition into sequential subquestions, offering conceptual grounding for dividing reasoning tasks into explicit plans and subtasks.
- Paper: Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Wenhu Chen et al. (2022). It introduces Program-of-Thought prompting to disentangle computation from reasoning, which serves as a primary benchmark and point of comparison in Plan-and-Solve Prompting.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). It demonstrates how sampling diverse reasoning paths mitigates reasoning and calculation errors in chain-of-thought outputs.
- Paper: Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, Mirac Suzgun et al. (2022). It establishes the BIG-Bench Hard suite, providing key benchmark evaluation standards for multi-step reasoning capabilities in language models.
- Paper: Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Shunyu Yao et al. (2023). It extends linear plan-and-solve workflows into deliberate tree-structured exploration with lookahead and backtracking over intermediate thought steps.
- Paper: Graph of Thoughts: Solving Elaborate Problems with Large Language Models, Maciej Besta et al. (2023). It generalizes step-by-step planning beyond trees into arbitrary graph networks to allow thought aggregation and cyclic refinement.
- Paper: Reasoning with Language Model is Planning with World Model, Shibo Hao et al. (2023). It enhances multi-step planning by combining Monte Carlo Tree Search with world-model state simulation to anticipate intermediate reasoning outcomes.
- Paper: Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models, Ling Yang et al. (2024). It builds on multi-step reasoning prompting by maintaining and retrieving distilled high-level thought templates across diverse problem-solving tasks.
- Paper: Let's Verify Step by Step, Hunter Lightman et al. (2023). It advances beyond prompt-level error mitigation by training step-level process reward models to supervise and verify intermediate reasoning stages.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). It automated step-level verification and reinforcement for multi-step math reasoning without requiring human rationale annotations.
- Paper: Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting, Miles Turpin et al. (2023). It critically examines the faithfulness of generated step-by-step reasoning explanations when language models are subjected to biased inputs.
