Active Prompting with Chain-of-Thought for Large Language Models

Shizhe DiaoPengcheng WangYong LinRui PanXiang LiuTong Zhang

article2024ACL210 citations

Proposes an uncertainty-based active learning framework that identifies and selects the most informative task-specific questions for human chain-of-thought annotation, significantly improving large language model reasoning performance with minimal labeling effort.

Listen

Large language models frequently struggle with complex multi-step reasoning tasks. While guiding these models with step-by-step example solutions—known as chain-of-thought prompting—significantly enhances accuracy, current practices rely on a fixed, arbitrarily chosen set of human-annotated examples. Because tasks vary widely in difficulty and structure, arbitrary examples are often suboptimal, leaving organizations to spend excessive time on manual trial-and-error prompt engineering without guaranteed performance improvements.

The article demonstrates an active prompting framework called Active-Prompt, which systematically identifies and annotates the most informative, task-specific example questions using uncertainty metrics. By framing question selection as an active learning problem, the approach determines which few examples provide the highest performance return on human annotation effort.

The proposed method operates across four structured stages: generating multiple candidate answers for a pool of unlabeled task questions, estimating model uncertainty across these outputs, selecting the top uncertain questions for expert step-by-step annotation, and using these newly annotated exemplars to guide final test inference. The article evaluates this framework across eight benchmark datasets covering arithmetic, commonsense, and symbolic reasoning, primarily utilizing language models such as OpenAI's Codex (code-davinci-002) and GPT-3.5 variants, alongside open-source models like Llama 2.

The findings establish that uncertainty-based question selection consistently outperforms conventional baselines. Active-Prompt improved accuracy across all eight benchmarks, outperforming standard self-consistency baselines by an average of 7.0 percentage points on text-davinci-002 and 1.8 percentage points on code-davinci-002, with specific math reasoning gains reaching up to 4.2 percentage points on GSM8K. Analysis confirmed that uncertainty metrics based on answer disagreement and entropy are highly effective, whereas asking the model to evaluate its own confidence failed due to severe overconfidence. Furthermore, a pool size of roughly 1,000 candidate questions and 10 sampled outputs provided robust uncertainty estimation, and questions selected by one model transferred successfully to improve other models.

These results demonstrate that the precision of exemplar selection, rather than prompt length or excessive engineering, is the primary driver of reasoning gains in language models. For technical leaders and operational teams, this provides a cost-effective, reproducible strategy: annotating just 4 to 8 highly uncertain, task-specific examples delivers superior model accuracy while minimizing expensive human labeling labor and compute costs. The finding that smaller, open-source models can identify uncertain examples that successfully transfer to larger commercial models also opens paths to lower API expenses.

Organizations deploying reasoning-intensive language model applications should adopt uncertainty-based active selection instead of arbitrary prompt crafting, utilizing disagreement or entropy metrics rather than model self-confidence scores. Future implementation should explore combining uncertainty-driven selection with prompt diversity metrics and automated step-by-step reasoning generation to further reduce manual annotation requirements.

Confidence in these findings is supported by consistent cross-task validation, though several limitations exist. The main experimental model (code-davinci-002) has been deprecated by OpenAI, and testing did not include top-tier frontier models like GPT-4 due to budget constraints. Additionally, tasks lacking dedicated training splits required transferring prompts across different datasets, indicating that careful evaluation is necessary when deploying the method to entirely new domains without domain-specific data pools.

Cover for Active Prompting with Chain-of-Thought for Large Language Models

Abstract

The increasing scale of large language models (LLMs) brings emergent abilities to various complex tasks requiring reasoning, such as arithmetic and commonsense reasoning. It is known that the effective design of task-specific prompts is critical for LLMs' ability to produce high-quality answers. In particular, an effective approach for complex question-and-answering tasks is example-based prompting with chain-of-thought (CoT) reasoning, which significantly improves the performance of LLMs. However, current CoT methods rely on a fixed set of human-annotated exemplars, which are not necessarily the most effective examples for different tasks. This paper proposes a new method, Active-Prompt, to adapt LLMs to different tasks with task-specific example prompts (annotated with human-designed CoT reasoning). For this purpose, we propose a solution to the key problem of determining which questions are the most important and helpful to annotate from a pool of task-specific queries. By borrowing ideas from the related problem of uncertainty-based active learning, we introduce several metrics to characterize the uncertainty so as to select the most uncertain questions for annotation. Experimental results demonstrate the superiority of our proposed method, achieving superior performance on eight complex reasoning tasks. Further analyses of different uncertainty metrics, pool sizes, zero-shot learning, and accuracy-uncertainty relationships demonstrate the effectiveness of our method.

Table of Contents

  • 1 Introduction
  • 2 Active-Prompt
  • 2.1 Uncertainty Estimation
  • 2.2 Selection and Annotation
  • 2.3 Inference
  • 3 Experimental Settings
  • 3.1 Datasets and Evaluation Metrics
  • 3.2 Baselines
  • 3.3 Implementation
  • 4 Experimental Results
  • 5 Analysis
  • 5.1 Ablation Study
  • 5.2 Uncertainty Analysis
  • 5.3 Transferability
  • 5.4 Performance of Weaker Models
  • 5.5 Transferability between GPT and Llama Models
  • 6 Related Work
  • 6.1 Chain-of-thought Prompting
  • 6.2 Active Learning
  • 7 Conclusion
  • Limitations
  • References
  • A Experimental Settings
  • A.1 Datasets and Evaluation Metrics
  • A.2 Baselines
  • A.3 Implementation
  • B Uncertainty Analysis
  • C Variance Analysis
  • D Self-confidence-based Uncertainty Estimation
  • E Logits-based Uncertainty Estimation
  • F Comparison with Diversity-based Methods
  • G Comparison with Complexity-based Methods
  • H Costs of Active-Prompt
  • I Ablation Study of Longer CoT Annotations
  • J Full Exemplars Generated by Active-Prompt

Knowls

  1. Knowl 1 — Active-Prompt: Task-Specific Few-Shot Exemplar Selection Framework

    model/method

    Active-Prompt is a framework designed to adapt Large Language Models (LLMs) to complex reasoning tasks by actively selecting the most informative task-specific questions for human Chain-of-Thought (CoT) annotation.

    The framework operates across four sequential stages:

    1. Uncertainty Estimation: Given an unlabeled training dataset Dtr={q1,q2,…,ql}\mathcal{D}_{tr} = \{q_1, q_2, \dots, q_l\}, a candidate pool of up to 1,000 questions is sampled. For each candidate question qiq_i, the LLM is queried kk times (typically k=10k=10 with temperature T>0T > 0) using either a few initial seed exemplars or a zero-shot prompt ("Let's think step by step.") to generate kk reasoning paths and candidate answers Ai={a1,a2,…,ak}\mathcal{A}_i = \{a_1, a_2, \dots, a_k\}. An uncertainty score u(qi)u(q_i) is then computed over Ai\mathcal{A}_i.

    2. Selection: Questions are ranked in descending order of their uncertainty score u(qi)u(q_i). The top-nn most uncertain questions are selected (with ties broken uniformly at random).

    3. Annotation: A human annotator writes step-by-step reasoning rationales cic_i and correct ground-truth answers aia_i for each of the nn selected questions, producing a task-specific exemplar set E={(q1,c1,a1),…,(qn,cn,an)}\mathcal{E} = \{(q_1, c_1, a_1), \dots, (q_n, c_n, a_n)\}.

    4. Inference: For each test question p∈Dtep \in \mathcal{D}_{te}, the exemplar set E\mathcal{E} is prepended as the few-shot in-context prompt. The LLM then generates the answer, optionally combined with self-consistency majority voting over mm sampled outputs.

  2. Knowl 2 — Active-Prompt Algorithm for Uncertainty-Based Exemplar Selection and Inference

    algorithm

    The Active-Prompt algorithm identifies uncertain instances from an unlabeled question pool, constructs human-annotated chain-of-thought exemplars, and evaluates test queries.

    Input: Unlabeled training question set DtrD_{tr}, test question set DteD_{te}, LLM θ\theta, number of sampled predictions kk, number of exemplars nn, uncertainty function U(⋅)U(\cdot), temperature TT, inference voting paths mm
    Output: Predicted answers for all test questions in DteD_{te}
    if ∣Dtr∣>1000|D_{tr}| > 1000 then
        Dpool←D_{pool} \leftarrow randomly sample 1000 instances from DtrD_{tr}
    else
        Dpool←DtrD_{pool} \leftarrow D_{tr}
    end if
    for each question qi∈Dpoolq_i \in D_{pool} do
        Sample kk responses from θ(qi)\theta(q_i) to extract predicted answers Ai={ai,1,ai,2,…,ai,k}A_i = \{a_{i,1}, a_{i,2}, \dots, a_{i,k}\}
        Compute uncertainty metric ui←U(Ai)u_i \leftarrow U(A_i)
    end for
    Sort DpoolD_{pool} by uiu_i in descending order
    Qselect←Q_{select} \leftarrow select top-nn questions from DpoolD_{pool} (breaking ties uniformly at random)
    Initialize exemplar set E←∅E \leftarrow \emptyset
    for each qj∈Qselectq_j \in Q_{select} do
        Obtain human-annotated reasoning chain cjc_j and ground-truth answer yjy_j
        E←E∪{(qj,cj,yj)}E \leftarrow E \cup \{(q_j, c_j, y_j)\}
    end for
    Initialize predictions dictionary Yte←{}Y_{te} \leftarrow \{\}
    for each test question p∈Dtep \in D_{te} do
        Sample mm outputs from θ(E∘p)\theta(E \circ p) at temperature TT
        Extract predicted answers and select majority consensus answer y^\hat{y}
        Yte[p]←y^Y_{te}[p] \leftarrow \hat{y}
    end for
    return YteY_{te}
  3. Knowl 3 — Uncertainty Metrics for Query Selection in Active-Prompt

    equation

    Active-Prompt measures the uncertainty of a question qiq_i using the distribution of predicted answers Ai={a1,a2,…,ak}\mathcal{A}_i = \{a_1, a_2, \dots, a_k\} obtained across kk stochastic model rollouts. Three primary metrics are defined:

    1. Disagreement Metric (uDu_D): Measures the proportion of unique answer strings among the kk predictions: uD(qi)=hku_D(q_i) = \frac{h}{k} where h=∣{a1,a2,…,ak}∣h = |\{a_1, a_2, \dots, a_k\}| is the count of distinct answers after deduplication.

    2. Entropy Metric (uEu_E): Calculates the Shannon entropy over the empirical discrete probability distribution Pθ(a∣qi)P_\theta(a \mid q_i) of answers generated across the kk predictions: uE(qi)=−∑j=1CPθ(aj∣qi)ln⁡Pθ(aj∣qi)u_E(q_i) = -\sum_{j=1}^{C} P_\theta(a_j \mid q_i) \ln P_\theta(a_j \mid q_i) where CC is the number of unique answers and Pθ(aj∣qi)P_\theta(a_j \mid q_i) denotes the relative frequency of unique answer aja_j in Ai\mathcal{A}_i.

    3. Normalized Variance Metric (uVu_V): Designed for numerical/arithmetic answers, variance is computed over normalized numerical predictions: uV(qi)=1k−1∑j=1k(a~j−a~ˉ)2u_V(q_i) = \frac{1}{k-1} \sum_{j=1}^k (\tilde{a}_j - \bar{\tilde{a}})^2 where a~j=aj∑r=1R∣xr∣\tilde{a}_j = \frac{a_j}{\sum_{r=1}^R |x_r|} normalizes each predicted numeric answer aja_j by the sum of absolute values of all numbers {x1,…,xR}\{x_1, \dots, x_R\} mentioned in question qiq_i, and a~ˉ=1k∑j=1ka~j\bar{\tilde{a}} = \frac{1}{k} \sum_{j=1}^k \tilde{a}_j is the mean of normalized predictions.

  4. Knowl 4 — Benchmark Evaluation of Active-Prompt Across Reasoning Tasks

    data/table

    Active-Prompt was evaluated across eight benchmarks spanning arithmetic reasoning (GSM8K, ASDiv, SVAMP, AQuA, SingleEq), commonsense reasoning (CSQA, StrategyQA), and symbolic reasoning (Letter (4) out-of-distribution). The table reports exact match accuracy (%) comparing standard Chain-of-Thought (CoT), Self-Consistency (SC), Auto-CoT, Random-CoT, Disagreement-based Active-Prompt (D), and Entropy-based Active-Prompt (E).

    Method GSM8K ASDiv SVAMP AQuA SingleEq CSQA Strategy Letter (4) AVG.
    text-davinci-002
    Auto-CoT 47.9 - 69.5 36.5 87.0 74.4 65.4 59.7 -
    CoT 46.9 71.3 68.9 35.8 77.3 73.5 65.4 56.6 61.5
    SC 58.2 76.9 78.2 41.8 87.2 72.9 70.7 57.6 67.9
    Random-CoT 63.9 82.3 81.1 44.1 89.4 74.5 73.3 65.5 71.8
    Active-Prompt (D) 73.2 83.2 82.7 48.4 90.6 76.6 76.9 67.7 74.9
    Active-Prompt (E) 71.1 83.8 81.8 50.3 93.1 78.8 76.9 66.7 75.3
    code-davinci-002
    Auto-CoT 62.8 - - - - - - - -
    CoT 63.1 80.4 76.4 45.3 93.1 77.9 73.2 70.4 72.5
    SC 78.0 87.8 86.8 52.0 93.7 81.5 79.8 73.4 79.1
    Random-CoT 78.6 87.1 88.0 53.1 94.0 82.1 79.4 73.3 79.4
    Active-Prompt (D) 82.2 88.4 88.7 55.1 94.5 83.9 80.6 74.1 80.9
    Active-Prompt (E) 83.4 89.3 87.5 57.0 95.5 82.6 80.6 76.7 81.6
    gpt-3.5-turbo-0613 (without SC)
    CoT 74.2 82.5 83.8 50.0 95.0 79.9 80.5 82.0 78.5
    Active-Prompt (D) 77.1 83.6 85.5 50.0 96.0 81.5 82.1 84.0 80.0
    Active-Prompt (E) 78.2 84.7 86.0 57.3 95.5 80.7 81.3 84.0 81.0

    Active-Prompt (D) and (E) consistently outperform CoT, Self-Consistency, and Auto-CoT baselines across both arithmetic and symbolic reasoning tasks, achieving average accuracy gains of 7.0% on text-davinci-002 and 1.8%–2.5% on code-davinci-002 over Self-Consistency.

  5. Knowl 5 — Exemplar-Free Uncertainty Estimation via Zero-Shot Active-Prompt

    empirical result

    Active-Prompt does not require an initial set of human-crafted few-shot exemplars during the uncertainty estimation stage. In Zero-Shot-Active-Prompt, the initial kk candidate reasoning paths are generated using the zero-shot prompt prefix "Let's think step by step." instead of 4–8 seed exemplars.

    On the code-davinci-002 model, Zero-Shot-Active-Prompt achieves exact match accuracies that match or closely match standard Active-Prompt with seed exemplars:

    • GSM8K: 82.2% (Zero-Shot-Active-Prompt) vs 82.2% (Active-Prompt Disagreement) vs 78.0% (Self-Consistency baseline).
    • ASDiv: 86.7% (Zero-Shot-Active-Prompt) vs 88.4% (Active-Prompt Disagreement) vs 87.8% (Self-Consistency baseline).
    • SingleEq: 94.2% (Zero-Shot-Active-Prompt) vs 94.5% (Active-Prompt Disagreement) vs 93.7% (Self-Consistency baseline).

    This demonstrates that the exemplar-selection pipeline can be executed in a fully autonomous, exemplar-free manner prior to final human annotation.

  6. Knowl 6 — Performance Comparison Between Active Selection and Random Exemplar Selection

    empirical result

    To isolate the performance gain attributable specifically to active uncertainty-based selection versus human annotation of new exemplars, Random-CoT was tested as a control baseline. Random-CoT samples nn questions uniformly at random from the training pool and has the same human annotator write CoT rationales using the identical protocol.

    Results across benchmarks demonstrate that Random-CoT performs essentially at the level of standard Self-Consistency, whereas Active-Prompt significantly exceeds it:

    • On GSM8K with code-davinci-002: Active-Prompt (E) achieves 83.4% and Active-Prompt (D) achieves 82.2%, compared to 78.6% for Random-CoT and 78.0% for Self-Consistency.
    • On text-davinci-002 average across all eight tasks: Active-Prompt achieves 74.9% (D) / 75.3% (E), whereas Random-CoT achieves 71.8% and Self-Consistency achieves 67.9%.

    This confirms that the accuracy improvements of Active-Prompt stem from prioritizing high-uncertainty questions rather than human annotation alone.

  7. Knowl 7 — Cross-Model Transferability of Actively Selected Exemplars

    empirical result

    Exemplars selected by measuring uncertainty on one model transfer effectively to other, distinct LLM architectures and sizes, indicating that uncertainty is predominantly task-inherent rather than specific to a single model's weights.

    Key cross-model evaluation findings include:

    1. Code-to-Text Transfer (code-davinci-002 →\to text-davinci-002 / text-davinci-003):

      • On GSM8K, standard text-davinci-002 (SC) achieves 58.2%. Using exemplars selected by code-davinci-002 improves performance to 73.2%.
      • On CSQA, text-davinci-003 (CoT) achieves 76.2%. Using exemplars selected by code-davinci-002 raises accuracy to 78.9%.
    2. Proprietary-to-Open Transfer (gpt-3.5-turbo ↔\leftrightarrow Llama2-70b-chat):

      • Using gpt-3.5-turbo to select exemplars for Llama2-70b-chat inference achieves 58.9% on GSM8K and 86.0% on SingleEq, outperforming native Llama2-70b-chat selection (56.9% and 83.2%) and baseline CoT (54.8% and 84.6%).
      • Using Llama2-70b-chat to select exemplars for gpt-3.5-turbo inference achieves 78.7% on GSM8K and 85.9% on ASDiv, surpassing native gpt-3.5-turbo baseline CoT (74.2% and 82.5%).

    Selecting exemplars with a stronger model and deploying them on a weaker model yields higher performance gains than selecting with the weaker model directly.

  8. Knowl 8 — Comparison of Uncertainty Estimation via Disagreement, Entropy, Variance, Logits, and Self-Confidence

    empirical result

    Multiple strategies for ranking query uncertainty in Active-Prompt exhibit varying performance and practical limitations:

    1. Disagreement and Entropy: Both metrics yield robust, top-tier performance across reasoning domains (e.g., GSM8K accuracies of 82.2% for Disagreement and 83.4% for Entropy on code-davinci-002). However, Disagreement fails on binary tasks with constrained output spaces (e.g., StrategyQA yes/no), where predictions frequently tie at u=1.0u=1.0, necessitating Entropy.

    2. Normalized Variance: Competitive on arithmetic datasets with continuous numerical answers (86.4% on ASDiv, 94.0% on SingleEq), but achieves lower performance on GSM8K (75.2%) compared to Disagreement (82.2%).

    3. Model Logits: Logit-based uncertainty on gpt-3.5-turbo-0301 yields an average accuracy of 88.2% across GSM8K, ASDiv, and SingleEq, performing comparably to Disagreement (88.1%) and Entropy (88.8%). However, on open-source models like Llama-2-70b, uncalibrated overconfidence in output logits degrades selection quality.

    4. Verbalized Self-Confidence: Prompting the LLM to classify its confidence (choices: very confident, confident, not confident, wrong answer) fails as a selection metric because LLMs exhibit severe overconfidence, routinely classifying erroneous generations as very confident.

  9. Knowl 9 — Disentangling the Impact of Reasoning Chain Length from Question Selection

    empirical result

    To evaluate whether performance gains in Active-Prompt are an artifact of writing longer CoT rationales (which averaged 160 words), an ablation extended the standard CoT exemplars of Wei et al. (2022b) to an average length of 155 words on gpt-3.5-turbo.

    Method GSM8K ASDiv SVAMP SingleEq
    Original CoT 74.2 82.5 83.8 95.0
    Longer CoT 69.4 69.2 70.4 83.2
    Active-Prompt 77.1 83.6 85.5 96.0

    Merely extending the word count of reasoning steps without altering the selected questions caused accuracy to drop substantially (e.g., from 74.2% down to 69.4% on GSM8K, and from 95.0% down to 83.2% on SingleEq). In contrast, Active-Prompt achieved the highest performance across all datasets (77.1% on GSM8K, 96.0% on SingleEq), confirming that the quality and informativeness of the actively chosen questions—not the verbosity of the annotations—drive the method's improvements.

  10. Knowl 10 — Limitations and Computational Trade-offs of Active-Prompt

    limitation

    Active-Prompt involves several practical constraints and computational requirements:

    1. Uncertainty Estimation Query Cost: Evaluating uncertainty requires sampling k=10k=10 model outputs for each candidate in a pool of 1,000 unlabeled questions (totaling 10,000 forward passes per task), introducing inference costs before annotation.

    2. Human Annotation Requirement: While requiring substantially less trial-and-error than arbitrary prompt engineering, the method still requires a knowledgeable human to hand-annotate nn multi-step CoT rationales and correct answers for the selected uncertain questions.

    3. Closed-Source API Deprecation and Reproducibility: The primary baseline experiments relied extensively on OpenAI's code-davinci-002, which was deprecated and shut down, limiting exact replication to researchers with specialized archival access.

    4. Task Transfer for Datasets Without Training Sets: For datasets lacking training splits (such as ASDiv, SVAMP, and SingleEq), Active-Prompt relies on transferring exemplars identified and annotated from related tasks (e.g., GSM8K), which can introduce transfer discrepancies.

Coverage note — Specific textual prompt exemplars for individual benchmark datasets provided in Appendix J (Tables 13-18) were omitted as they represent raw task-specific prompt instances rather than generalizable methods or experimental results.

References

  1. 1.Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi. 2019. MathQA: Towards interpretable math word problem solving with operation-based formalisms. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2357–2367, Minneapolis, Minnesota. Association for Computational Linguistics.
  2. 2.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  3. 3.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  4. 4.Yangyi Chen, Lifan Yuan, Ganqu Cui, Zhiyuan Liu, and Heng Ji. 2022. A close look into the calibration of pre-trained language models. arXiv preprint arXiv:2211.00151.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  7. 7.David A Cohn, Zoubin Ghahramani, and Michael I Jordan. 1996. Active learning with statistical models. Journal of artificial intelligence research, 4:129–145.
  8. 8.Aron Culotta and Andrew McCallum. 2005. Reducing labeling effort for structured prediction tasks. In AAAI, volume 5, pages 746–751.
  9. 9.Andrew Drozdov, Nathanael Scharli, Ekin Akyurek, Nathan Scales, Xinying Song, Xinyun Chen, Olivier Bousquet, and Denny Zhou. 2022. Compositional semantic parsing with large language models. arXiv preprint arXiv:2209.15003.
  10. 10.Yao Fu, Hao Peng, Ashish Sabharwal, Peter Clark, and Tushar Khot. 2022. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720.
  11. 11.Claudio Gentile, Zhilei Wang, and Tong Zhang. 2022. Fast rates in pool-based batch active learning.
  12. 12.Mor Geva, Daniel Khashabi, Elad Segal, Tushar Khot, Dan Roth, and Jonathan Berant. 2021. Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies. Transactions of the Association for Computational Linguistics, 9:346–361.
  13. 13.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017. On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR.
  14. 14.Minghao Hu, Yuxing Peng, Zhen Huang, and Dongsheng Li. 2019. A multi-type multi-span network for reading comprehension that requires discrete reasoning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 1596–1606.
  15. 15.Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2022. Large language models can self-improve. arXiv preprint arXiv:2210.11610.
  16. 16.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916.
  17. 17.Abdullatif Koksal, Timo Schick, and Hinrich Schutze. 2022. Meal: Stable and active learning for few-shot prompting. arXiv preprint arXiv:2211.08358.
  18. 18.Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016. MAWPS: A math word problem repository. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1152–1157, San Diego, California. Association for Computational Linguistics.
  19. 19.Lingkai Kong, Haoming Jiang, Yuchen Zhuang, Jie Lyu, Tuo Zhao, and Chao Zhang. 2020. Calibrated language model fine-tuning for in-and out-of-distribution data. arXiv preprint arXiv:2010.11506.
  20. 20.Yihuai Lan, Lei Wang, Qiyuan Zhang, Yunshi Lan, Bing Tian Dai, Yan Wang, Dongxiang Zhang, and Ee-Peng Lim. 2022. Mwptoolkit: an open-source framework for deep learning-based math word problem solvers. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 13188–13190.
  21. 21.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  22. 22.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2023. Making language models better reasoners with step-aware verifier. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5315–5333.
  23. 23.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. 2022. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
  24. 24.Yong Lin, Chen Liu, Chenlu Ye, Qing Lian, Yuan Yao, and Tong Zhang. 2023. Optimal sample selection through uncertainty estimation and its application in deep learning. arXiv preprint arXiv:2309.02476.
  25. 25.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 158–167, Vancouver, Canada. Association for Computational Linguistics.
  26. 26.Shen-Yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020. A diverse corpus for evaluating and developing english math word problem solvers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 975–984.
  27. 27.Fredrik Olsson. 2009. A literature survey of active machine learning in the context of natural language processing.
  28. 28.OpenAI. 2023. Gpt-4 technical report.
  29. 29.Rui Pan, Shuo Xing, Shizhe Diao, Xiang Liu, Kashun Shum, Jipeng Zhang, and Tong Zhang. 2023. Plum: Prompt learning using metaheuristic. arXiv preprint arXiv:2311.08364.
  30. 30.Shilong Pan, Zhiliang Tian, Liang Ding, Zhen Huang, Zhihua Wen, and Dongsheng Li. 2024. Pomp: Probability-driven meta-graph prompter for llms in low-resource unsupervised neural machine translation.
  31. 31.Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are nlp models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094.
  32. 32.Xinyu Pi, Qian Liu, Bei Chen, Morteza Ziyadi, Zeqi Lin, Yan Gao, Qiang Fu, Jian-Guang Lou, and Weizhu Chen. 2022. Reasoning like program executors. arXiv preprint arXiv:2201.11473.
  33. 33.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  34. 34.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67.
  35. 35.Guy Rotman and Roi Reichart. 2022. Multi-task active learning for pre-trained transformer-based models. Transactions of the Association for Computational Linguistics, 10:1209–1228.
  36. 36.Nicholas Roy and Andrew McCallum. 2001. Toward optimal active learning through sampling estimation of error reduction. int. conf. on machine learning.
  37. 37.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman Castagne, Alexandra Sasha Luccioni, Francois Yvon, Matthias Galle, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  38. 38.Christopher Schroder, Andreas Niekler, and Martin Potthast. 2022. Revisiting uncertainty-based query strategies for active learning with transformers. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2194–2203.
  39. 39.Burr Settles. 2009. Active learning literature survey.
  40. 40.Kashun Shum, Shizhe Diao, and Tong Zhang. 2023. Automatic prompt augmentation and selection with chain-of-thought from labeled data. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12113–12139.
  41. 41.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2022. Prompting gpt-3 to be reliable. arXiv preprint arXiv:2210.09150.
  42. 42.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  43. 43.Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158.
  44. 44.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131.
  45. 45.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothee Lacroix, Baptiste Roziere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023a. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.
  46. 46.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023b. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  47. 47.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  48. 48.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a. Emergent abilities of large language models. Transactions on Machine Learning Research. Survey Certification.
  49. 49.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  50. 50.Xin Xu, Shizhe Diao, Can Yang, and Yang Wang. 2024. Can we verify step by step for incorrect answer detection? arXiv preprint arXiv:2402.10528.
  51. 51.Yichong Xu, Chenguang Zhu, Shuohang Wang, Siqi Sun, Hao Cheng, Xiaodong Liu, Jianfeng Gao, Pengcheng He, Michael Zeng, and Xuedong Huang. 2021. Human parity on commonsenseqa: Augmenting self-attention with external attention. arXiv preprint arXiv:2112.03254.
  52. 52.Eric Zelikman, Jesse Mu, Noah D Goodman, and Yuhuai Tony Wu. 2022. Star: Self-taught reasoner bootstrapping reasoning with reasoning.
  53. 53.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. 2022. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414.
  54. 54.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022a. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  55. 55.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022b. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493.
  56. 56.Denny Zhou, Nathanael Scharli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, Quoc Le, and Ed Chi. 2022. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.

Citation

MLA
Diao, S., et al. “Active Prompting with Chain-of-Thought for Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 1330–50, https://doi.org/10.18653/v1/2024.acl-long.73.
APA
Diao, S., Wang, P., Lin, Y., Pan, R., Liu, X., & Zhang, T. (2024). Active Prompting with Chain-of-Thought for Large Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1330–1350. https://doi.org/10.18653/v1/2024.acl-long.73
Chicago
Diao, S., P. Wang, Y. Lin, R. Pan, X. Liu, and T. Zhang. 2024. “Active Prompting with Chain-of-Thought for Large Language Models”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 1330–50. https://doi.org/10.18653/v1/2024.acl-long.73.
Harvard
Diao, S. et al. (2024) “Active Prompting with Chain-of-Thought for Large Language Models”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 1330–1350. Available at: https://doi.org/10.18653/v1/2024.acl-long.73.
Vancouver
1. Diao S, Wang P, Lin Y, Pan R, Liu X, Zhang T (2024) Active Prompting with Chain-of-Thought for Large Language Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 1330–1350

BibTeX

@inproceedings{diao-etal-2024-active,
    title = "Active Prompting with Chain-of-Thought for Large Language Models",
    author = "Diao, Shizhe  and
      Wang, Pengcheng  and
      Lin, Yong  and
      Pan, Rui  and
      Liu, Xiang  and
      Zhang, Tong",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.73/",
    doi = "10.18653/v1/2024.acl-long.73",
    pages = "1330--1350"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/