Active Prompting with Chain-of-Thought for Large Language Models
Shizhe DiaoPengcheng WangYong LinRui PanXiang LiuTong Zhang
Proposes an uncertainty-based active learning framework that identifies and selects the most informative task-specific questions for human chain-of-thought annotation, significantly improving large language model reasoning performance with minimal labeling effort.
Large language models frequently struggle with complex multi-step reasoning tasks. While guiding these models with step-by-step example solutions—known as chain-of-thought prompting—significantly enhances accuracy, current practices rely on a fixed, arbitrarily chosen set of human-annotated examples. Because tasks vary widely in difficulty and structure, arbitrary examples are often suboptimal, leaving organizations to spend excessive time on manual trial-and-error prompt engineering without guaranteed performance improvements.
The article demonstrates an active prompting framework called Active-Prompt, which systematically identifies and annotates the most informative, task-specific example questions using uncertainty metrics. By framing question selection as an active learning problem, the approach determines which few examples provide the highest performance return on human annotation effort.
The proposed method operates across four structured stages: generating multiple candidate answers for a pool of unlabeled task questions, estimating model uncertainty across these outputs, selecting the top uncertain questions for expert step-by-step annotation, and using these newly annotated exemplars to guide final test inference. The article evaluates this framework across eight benchmark datasets covering arithmetic, commonsense, and symbolic reasoning, primarily utilizing language models such as OpenAI's Codex (code-davinci-002) and GPT-3.5 variants, alongside open-source models like Llama 2.
The findings establish that uncertainty-based question selection consistently outperforms conventional baselines. Active-Prompt improved accuracy across all eight benchmarks, outperforming standard self-consistency baselines by an average of 7.0 percentage points on text-davinci-002 and 1.8 percentage points on code-davinci-002, with specific math reasoning gains reaching up to 4.2 percentage points on GSM8K. Analysis confirmed that uncertainty metrics based on answer disagreement and entropy are highly effective, whereas asking the model to evaluate its own confidence failed due to severe overconfidence. Furthermore, a pool size of roughly 1,000 candidate questions and 10 sampled outputs provided robust uncertainty estimation, and questions selected by one model transferred successfully to improve other models.
These results demonstrate that the precision of exemplar selection, rather than prompt length or excessive engineering, is the primary driver of reasoning gains in language models. For technical leaders and operational teams, this provides a cost-effective, reproducible strategy: annotating just 4 to 8 highly uncertain, task-specific examples delivers superior model accuracy while minimizing expensive human labeling labor and compute costs. The finding that smaller, open-source models can identify uncertain examples that successfully transfer to larger commercial models also opens paths to lower API expenses.
Organizations deploying reasoning-intensive language model applications should adopt uncertainty-based active selection instead of arbitrary prompt crafting, utilizing disagreement or entropy metrics rather than model self-confidence scores. Future implementation should explore combining uncertainty-driven selection with prompt diversity metrics and automated step-by-step reasoning generation to further reduce manual annotation requirements.
Confidence in these findings is supported by consistent cross-task validation, though several limitations exist. The main experimental model (code-davinci-002) has been deprecated by OpenAI, and testing did not include top-tier frontier models like GPT-4 due to budget constraints. Additionally, tasks lacking dedicated training splits required transferring prompts across different datasets, indicating that careful evaluation is necessary when deploying the method to entirely new domains without domain-specific data pools.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces chain-of-thought prompting for large language models, providing the foundational reasoning paradigm that Active-Prompt builds upon and optimizes through exemplar selection.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Establishes multi-path sampling and consistency across reasoning paths, which directly motivates the uncertainty and disagreement metrics used to select prompts in Active-Prompt.
- Paper: Large Language Models are Zero-Shot Reasoners, Takeshi Kojima et al. (2022). Demonstrates zero-shot chain-of-thought elicitation, which Active-Prompt employs to initialize uncertainty estimation before selecting exemplars for human annotation.
- Paper: Active Learning Helps Pretrained Models Learn the Intended Task, Alex Tamkin et al. (2022). Provides the uncertainty-based active learning principles that Active-Prompt adapts to the discrete exemplar selection problem for language model reasoning.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). Pioneers instance-level exemplar selection for language model in-context learning, framing the core challenge of choosing optimal demonstrations that Active-Prompt addresses via uncertainty.
- Paper: Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models, Zhihong Shao et al. (2023). Explores selecting and generating diverse chain-of-thought demonstrations, highlighting the exemplar-quality bottleneck resolved by Active-Prompt's active learning approach.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Demonstrates the necessity of task decomposition and specialized multi-step prompting on complex tasks where fixed exemplar sets fall short.
- Paper: Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future, Zheng Chu et al. (2024). Surveys the broader chain-of-thought literature, categorizing advanced exemplar selection and active optimization strategies in relation to the wider landscape.
- Paper: Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models, Ling Yang et al. (2024). Extends dynamic prompt adaptation by distilling and retrieving high-level thought templates from past problem-solving traces rather than annotating individual task-specific exemplars.
- Paper: Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching, Simon A. Aytes et al. (2025). Builds on step-by-step chain-of-thought prompting by introducing adaptive, token-efficient mental sketching to overcome the verbosity and latency of full-sentence reasoning.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). Advances beyond static exemplar prompting toward System 2 search and deliberate exploration processes for hard multi-step reasoning.
- Paper: Chain-of-Thought Reasoning Without Prompting, Xuezhi Wang et al. (2024). Investigates whether the reasoning capabilities elicited by specialized prompting methods like Active-Prompt can be decoded directly from base models without explicit prompt engineering.
