Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling

Weijia XuAndrzej BanburskiNebojsa Jojic

article2024ICML62 citations

Proposes an automated method that uses Gibbs sampling to infer effective chain-of-thought prompts without human intervention, outperforming human-written prompts and existing prompt optimization techniques across twenty complex reasoning benchmarks.

Listen

Large language models excel at processing language but struggle with complex logical deduction and mathematical reasoning unless guided by step-by-step demonstrations, commonly known as chain-of-thought prompts. Currently, creating these step-by-step examples requires substantial time and expertise from human engineers who must tailor instructions to each specific problem. This reliance on manual prompt design limits scalability and complicates fair performance comparisons across different language models.

The article introduces and evaluates Reprompting, an automated algorithm designed to discover effective chain-of-thought prompt recipes from a small set of example problems without human intervention. The primary objective is to demonstrate that automated iterative sampling can discover reasoning strategies that match or exceed the quality of human-engineered prompts.

To achieve this, the authors modeled prompt creation as an evolutionary sampling process using Gibbs sampling on twenty example problems per task. The algorithm starts by generating initial zero-shot solutions, uses the best-performing solutions as parent prompts to solve other training problems, and filters out erroneous steps through rejection sampling over thousands of iterations. The researchers evaluated the method across twenty complex reasoning tasks from established benchmarks, testing leading commercial language models including ChatGPT and InstructGPT.

The results establish four major findings. First, prompts generated automatically through Reprompting outperformed expert human-written prompts across the benchmark tasks by an average of 9.4 percentage points in accuracy. Second, the method consistently outperformed existing automated prompt optimization and decoding techniques by margins ranging from 11 to 33 percentage points. Third, combining models—using ChatGPT to propose initial candidate solutions and InstructGPT to iteratively refine them—improved performance by up to 71 percentage points over using InstructGPT alone. Fourth, reasoning strategies did not transfer cleanly across different model architectures; prompts optimized for one model frequently suffered performance drops of 18 to 19 percentage points when applied to another.

These findings indicate that automated prompt discovery eliminates the labor cost and bottleneck of manual prompt engineering while delivering superior model accuracy. The results also demonstrate that comparing language models using identical, static human prompts introduces significant evaluation bias, because different architectures require distinct reasoning patterns to achieve peak performance.

Organizations deploying language models for structured reasoning tasks should adopt automated prompt optimization pipelines instead of relying on manual prompt authoring. Evaluation teams must also tailor prompts to each specific model before benchmarking capabilities. Future initiatives should investigate hybrid approaches, such as combining multiple initialization models or introducing lightweight human feedback into early sampling rounds, to further improve efficiency.

These conclusions are supported by extensive empirical testing across diverse reasoning datasets, though practical adoption requires attention to specific operational boundaries. Reprompting depends on having verified input-answer pairs for the target domain and incurs upfront cloud computing costs from repeated model queries. Nevertheless, confidence in the method is high for standard multi-step reasoning workflows where ground-truth verification is readily available.

No sufficiently relevant recommendations were found.

Cover for Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling

Abstract

We introduce Reprompting, an iterative sampling algorithm that automatically learns the Chain-of-Thought (CoT) recipes for a given task without human intervention. Through Gibbs sampling, Reprompting infers the CoT recipes that work consistently well for a set of training samples by iteratively sampling new recipes using previously sampled recipes as parent prompts to solve other training problems. We conduct extensive experiments on 20 challenging reasoning tasks. Results show that Reprompting outperforms human-written CoT prompts substantially by +9.4 points on average. It also achieves consistently better performance than the state-of-the-art prompt optimization and decoding algorithms.

Table of Contents

  • 1. Introduction
  • 2. Reprompting: Prompt Inference Through Gibbs Sampling
  • 2.1. In-Context Learning
  • 2.2. Prompt Inference Through Gibbs Sampling
  • 3. Experimental Setup
  • 3.1. Reprompting Setup
  • 3.2. Baselines
  • 3.3. Large Language Models (LLMs)
  • 3.4. Evaluation Protocol
  • 4. Results
  • 4.1. Main Results
  • 4.2. Quantitative Analysis
  • 4.3. Qualitative Analysis
  • 5. Related Work
  • 6. Conclusion
  • Acknowledgements
  • Impact Statement
  • References
  • A. Additional Illustrations

Knowls

  1. Knowl 1 — Reprompting iteratively samples and selects task-specific reasoning recipes

    algorithm

    Reprompting takes labeled training examples (xj,yj)(x_j,y_j), where xjx_j is a question and yjy_j its correct answer, and produces a worked reasoning recipe zjz_j for each example. It uses an initialization language model LLM1, a sampling language model LLM2, an instruction message mm, a prompt size KK, a maximum iteration count MM, and a rejection probability prejp_{\mathrm{rej}}.

    Input: Training examples {(x_j, y_j)} for j = 1,...,N; prompt size K; maximum iterations M; rejection probability p_rej; models LLM1 and LLM2; instruction m
    Output: A recipe z_j for each training example
    Initialization:
      For each training example j:
        Generate a candidate recipe z_hat_j and answer y_hat_j from LLM1 given x_j and m, with no demonstrations
        Draw u uniformly from [0,1]
        If y_hat_j equals y_j, or u > p_rej, set z_j = z_hat_j
    Iterative sampling:
      Repeat until convergence or M iterations:
        Select a training index j uniformly at random
        Select K distinct indices S_j uniformly from {1,...,N} excluding j
        Prompt LLM2 with the examples (x_i, z_i, y_i) for i in S_j, plus x_j and m
        Generate a candidate recipe z_hat_j and answer y_hat_j
        Draw u uniformly from [0,1]
        If y_hat_j equals y_j, or u > p_rej, replace z_j with z_hat_j
    Test-time prompt construction:
      Rank the learned training tuples (x_j, z_j, y_j) by the training accuracy obtained when each tuple is used individually in a prompt
      Use the top K tuples as demonstrations for test questions

    An incorrect-answer candidate is therefore still accepted with probability 1−prej1-p_{\mathrm{rej}}; this permits occasionally retaining reasoning fragments that may be useful when later recipes are generated from them. In the reported setup, K=5K=5, MM was at most 20,000, and prej=0.99p_{\mathrm{rej}}=0.99. The experiment also stopped early when average training accuracy did not increase for 1,000 iterations.

  2. Knowl 2 — The recipe set is targeted to make few-shot reasoning insensitive to which demonstrations are selected

    model/method

    Let (xi,yi)(x_i,y_i) for i=1,...,Ni=1,...,N be labeled training questions and answers, ziz_i a text recipe giving step-by-step reasoning for example ii, mm an instruction message, and KK the number of demonstrations in a prompt. Reprompting frames recipe inference as sampling a joint set of recipes conditioned on the training data and the instruction. The intended property is that, for a target question xx, the LLM's distribution over its generated reasoning zz and answer yy should be approximately unchanged by which KK training tuples are used as demonstrations:

    p(z1,…,zN∣{(xi,yi)}i=1N,m),pLLM(z,y∣{(xi,zi,yi)}i∈S,x,m)≈pLLM(z,y∣{(xi,zi,yi)}i=1N,x,m),∣S∣=K.p(z_1,\ldots,z_N\mid \{(x_i,y_i)\}_{i=1}^N,m),\qquad p_{\mathrm{LLM}}(z,y\mid \{(x_i,z_i,y_i)\}_{i\in S},x,m) \approx p_{\mathrm{LLM}}(z,y\mid \{(x_i,z_i,y_i)\}_{i=1}^N,x,m), \quad |S|=K.

    Here SS is a subset of training indices, and pLLMp_{\mathrm{LLM}} is the model's conditional generation distribution. This invariance is the algorithm's target rather than a property established for arbitrary LLMs. Since the joint distribution over recipes is not directly characterized, Reprompting approximates its sampling through iterative Gibbs-style updates and uses answer correctness as a practical proxy for the conditional weight of a candidate recipe.

  3. Knowl 3 — Experimental protocol and model settings

    experimental setup

    The evaluation covers 20 reasoning tasks: 12 Big-Bench Hard (BBH) tasks, GSM8K, and seven MATH task categories. For each task, the authors randomly selected 20 labeled training examples, excluding BBH test examples. For Penguins in a Table, only three non-test examples were available outside the BBH set, so 17 more were sampled from BBH. Reprompting used either one or three clones of each selected training example, choosing the clone count with the higher training accuracy; this yielded N=20kN=20k recipes for k∈{1,3}k\in\{1,3\}. The prompt used K=5K=5 demonstrations, with at most 20,000 iterations and early stopping after 1,000 iterations without an increase in average training accuracy. The tested rejection probabilities were 0.95 and 0.99; 0.99 was selected for the reported results because it gave higher training accuracy on various tasks.

    Experiments used ChatGPT (gpt-3.5-turbo), InstructGPT (text-davinci-003), and a combination using ChatGPT for initialization and InstructGPT for iterative sampling. Generation used a 500-token output limit, top-p=0.5p=0.5, zero frequency and presence penalties, and END as a stop word. Reprompting used temperature 1.0 and test-time generation used temperature 0.0. Accuracy was exact match after extracting the answer between the answer tags; the human-written CoT baseline used the extraction procedure of its source benchmark.

  4. Knowl 4 — Reprompting improves average accuracy across the 20 reasoning tasks

    empirical result

    Across 20 tasks from BBH, GSM8K, and MATH, the paper reports that Reprompting's accuracy averaged 9.4 percentage points higher than human-written chain-of-thought (CoT) prompting. The authors also report that Reprompting outperformed self-consistency decoding and the evaluated state-of-the-art prompt-optimization methods by 11–33 points on average. These comparisons were made using small labeled training sets and without human construction of task-specific reasoning recipes; the numerical task-level results reported in the paper are given in the accompanying benchmark results.

  5. Knowl 5 — BBH comparison shows strong gains over prompting and optimization baselines

    data/table

    The table reports accuracy in percent on five BBH tasks. The SOTA column is the published state-of-the-art result cited by the paper, while the other baseline columns use ChatGPT. The final three columns are Reprompting with ChatGPT, InstructGPT, and ChatGPT initialization followed by InstructGPT sampling (Chat+Ins). ChatGPT Reprompting averages 83.0%, compared with 70.1% for human-written CoT, 71.2% for CoT with self-consistency, and 72.2% for Auto-CoT. It exceeds the CoT result on each of the five tasks; the mixed-model version reaches 99.6% on Object Counting and 99.2% on Temporal Sequences.

    Task SOTA ZS FS CoT CoT+SC APO AutoCoT Reprompt ChatGPT Reprompt InsGPT Chat+Ins
    Logical 60.4 35.1 46.4 63.1 62.7 28.0 53.2 66.3 53.7 60.0
    Geometric 56.0 13.6 20.0 58.0 60.0 52.0 52.4 72.8 40.8 64.4
    ObjectCount 93.2 52.4 46.8 95.6 95.2 74.8 88.8 97.2 42.8 99.6
    Penguins 81.5 50.7 60.3 67.1 71.2 45.2 85.6 85.6 78.1 82.9
    Temporal 96.8 38.4 41.2 66.8 66.8 50.4 80.8 93.2 28.4 99.2
    Average 77.6 38.0 42.9 70.1 71.2 50.1 72.2 83.0 48.8 81.2

    ZS is zero-shot, FS is few-shot, CoT is human-written chain-of-thought, CoT+SC adds self-consistency decoding, and APO is automatic prompt optimization using textual feedback. Values reproduce the paper's reported percentages.

  6. Knowl 6 — Reprompting outperforms CoT on most of the other 15 tasks

    data/table

    This comparison uses ChatGPT on seven additional BBH tasks, GSM8K, and seven MATH categories. Reprompting exceeds CoT on 11 of the 15 tasks and raises the reported mean from 44.9% to 53.0% (+8.1 points). It also improves over zero-shot and few-shot on average by 14.7 and 14.3 points, respectively. Reprompting is not uniformly better than CoT: it is lower on Date Understanding, Colored Objects, Integer Algebra, and Prealgebra. It nonetheless improves substantially over zero-shot on tasks where ordinary CoT performs poorly, including Movie Recommendation, Salient Translation Error Detection, and Word Sorting.

    Task ZS FS CoT Reprompting
    BBH Date 63.6 46.4 76.8 76.4
    BBH Formal 49.2 53.6 48.4 56.8
    BBH Movie 59.2 72.4 25.6 78.4
    BBH ColoredObj 66.8 48.8 76.0 74.0
    BBH Ruin 53.2 66.8 60.8 74.8
    BBH Salient 43.2 53.2 32.8 54.8
    BBH WordSort 58.0 72.0 46.0 73.2
    GSM8K 45.6 26.5 75.6 79.5
    MATH Algebra 37.6 23.7 52.0 53.1
    MATH Counting 17.1 19.8 26.6 32.3
    MATH Geometry 12.4 16.2 28.5 29.2
    MATH IntAlgebra 9.4 12.1 18.0 16.8
    MATH Number 20.8 17.1 32.9 33.3
    MATH Prealgebra 31.4 33.2 54.0 43.8
    MATH Precalculus 7.4 18.4 19.0 19.3
    Average 38.3 38.7 44.9 53.0

    All entries are accuracy percentages. The paper's table uses the abbreviated task labels shown above.

  7. Knowl 7 — ChatGPT initialization substantially boosts InstructGPT sampling

    empirical result

    On the five BBH tasks in the main model comparison, using ChatGPT only to initialize recipes and then InstructGPT for iterative sampling improves over Reprompting with InstructGPT alone by 4.8–70.8 percentage points across tasks (reported in the paper as 5–71 points). The mixed system's scores are 60.0% on Logical Deduction, 64.4% on Geometric Shapes, 99.6% on Object Counting, 82.9% on Penguins in a Table, and 99.2% on Temporal Sequences, compared with 53.7%, 40.8%, 42.8%, 78.1%, and 28.4% for InstructGPT alone. The mixed system also surpasses ChatGPT Reprompting on Object Counting and Temporal Sequences. The authors attribute the benefit to ChatGPT supplying a more diverse set of initial strategies that InstructGPT can subsequently follow, recombine, and refine.

  8. Knowl 8 — Rejection sampling and recipe recombination both contribute to test accuracy

    empirical result

    An ablation on Logical Deduction, Object Counting, and Temporal Sequences compares standard Reprompting (rejection probability prej=0.99p_{\mathrm{rej}}=0.99 and recombination enabled) with three variants. The standard method averages 85.6% accuracy. Accepting every candidate, including incorrect-answer candidates (prej=0p_{\mathrm{rej}}=0), reduces the average to 61.0%, a 24.6-point drop. Rejecting every candidate with an incorrect answer (prej=1p_{\mathrm{rej}}=1) gives 77.8%, a 7.8-point drop. Disabling recombination of previously sampled recipes gives 80.2%, a 5.4-point drop. The results support both occasional retention of incorrect candidates and recombination as useful parts of the method.

    Task p_rej = 0 p_rej = 1 NoRec Standard
    Logical Deduction 56.3 61.9 54.7 66.3
    Object Counting 52.0 97.2 95.6 97.2
    Temporal Sequences 74.8 74.4 90.4 93.2
    Average 61.0 77.8 80.2 85.6

    Entries are test accuracy percentages. NoRec means Reprompting without recombining previously sampled recipes; Standard is the unablated algorithm.

  9. Knowl 9 — Recipes optimized for one LLM can transfer poorly to another

    empirical result

    Testing the best Reprompting recipe on both ChatGPT and InstructGPT shows model-dependent transfer. For Geometric Shapes, the ChatGPT-optimized recipe scores 72.8% on ChatGPT but 53.6% on InstructGPT. For Temporal Sequences, the InstructGPT-optimized recipe scores 99.2% on InstructGPT but 81.6% on ChatGPT. The paper describes these cross-model losses as 18–19 points and reports that optimizing the recipe for the model being evaluated can improve accuracy by 11–12 points over using a recipe optimized for the other model. Transfer is closer on Logical Deduction and Object Counting, so the result is task-dependent rather than universal.

    The reported scores by test model are: Logical Deduction, InstructGPT 65.9% and ChatGPT 66.3%; Geometric Shapes, 53.6% and 72.8%; Object Counting, 99.6% and 96.8%; Penguins in a Table, 82.2% and 85.6%; Temporal Sequences, 99.2% and 81.6%. The asterisk in the source table marks the model used as the iterative sampling model LLM2 when that recipe was inferred.

  10. Knowl 10 — Recipes can improve through iteration even when early reasoning is wrong

    empirical result

    On Logical Deduction, the plotted average training accuracy for Reprompting with InstructGPT, ChatGPT, and the ChatGPT-initialization/InstructGPT-sampling combination rises from relatively low initial values toward a plateau, with occasional fluctuations. The plotted quantity averages training accuracy over the current and all previous iterations. The qualitative examples show how a recipe containing an incorrect deduction can still provide a useful strategy: a later generation may preserve a productive ordering of constraints while correcting or changing how the constraints are handled. With multiple demonstrations, a generated recipe can also combine reasoning fragments that address different cases. These examples illustrate the proposed evolution and recombination mechanism; they do not establish that every erroneous recipe will become useful.

Coverage note — No other substantial contributed result was omitted. The reported API cost estimates and appendix prompt examples are not separate knowls: the cost figures are ancillary resource reporting, and the examples illustrate recipe evolution already captured above.

References

  1. 1.Arora, S., Narayan, A., Chen, M. F., Orr, L. J., Guha, N., Bhatia, K., Chami, I., Sala, F., and Ré, C. Ask me anything: A simple strategy for prompting language models. arXiv preprint arXiv:2210.02441, 2022.
  2. 2.Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al. Program synthesis with large language models. arXiv preprint arXiv:2108.07732, 2021.
  3. 3.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. Neural Information Processing Systems (NeurIPS), 2020.
  4. 4.Casella, G. and George, E. I. Explaining the gibbs sampler. The American Statistician, 46(3):167–174, 1992.
  5. 5.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, M., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W. Evaluating large language models trained on code, 2021. URL https://arxiv.org/abs/2107.03374.
  6. 6.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  7. 7.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  8. 8.Creswell, A., Shanahan, M., and Higgins, I. Selection-inference: Exploiting large language models for interpretable logical reasoning. arXiv preprint arXiv:2205.09712, 2022.
  9. 9.Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548, 2022.
  10. 10.Fu, Y., Peng, H., Sabharwal, A., Clark, P., and Khot, T. Complexity-based prompting for multi-step reasoning. arXiv preprint arXiv:2210.00720, 2022.
  11. 11.Geman, S. and Geman, D. Stochastic relaxation, gibbs distributions, and the bayesian restoration of images. IEEE Transactions on Pattern Analysis and Machine Intelligence, PAMI-6:721–741, 1984.
  12. 12.Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. Measuring mathematical problem solving with the MATH dataset. CoRR, abs/2103.03874, 2021. URL https://arxiv.org/abs/2103.03874.
  13. 13.Jojic, A., Wang, Z., and Jojic, N. Gpt is becoming a turing machine: Here are some ways to program it, 2023.
  14. 14.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  15. 15.Li, Y., Lin, Z., Zhang, S., Fu, Q., Chen, B., Lou, J.-G., and Chen, W. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336, 2022.
  16. 16.Liu, Z., Patwary, M., Prenger, R., Prabhumoye, S., Ping, W., Shoeybi, M., and Catanzaro, B. Multi-stage prompting for knowledgeable dialogue generation. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 1317–1337, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.findings-acl.104. URL https://aclanthology.org/2022.findings-acl.104.
  17. 17.Madaan, A. and Yazdanbakhsh, A. Text and patterns: For effective chain of thought, it takes two to tango. arXiv preprint arXiv:2209.07686, 2022.
  18. 18.Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al. Show your work: Scratchpads for intermediate computation with language models. arXiv preprint arXiv:2112.00114, 2021.
  19. 19.OpenAI. Gpt-4 technical report, 2023.
  20. 20.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  21. 21.Paranjape, B., Lundberg, S., Singh, S., Hajishirzi, H., Zettlemoyer, L., and Ribeiro, M. T. Art: Automatic multi-step reasoning and tool-use for large language models. arXiv preprint arXiv:2303.09014, 2023.
  22. 22.Perez, E., Kiela, D., and Cho, K. True few-shot learning with language models. Advances in neural information processing systems, 34:11054–11070, 2021.
  23. 23.Press, O., Zhang, M., Min, S., Schmidt, L., Smith, N. A., and Lewis, M. Measuring and narrowing the compositionality gap in language models. arXiv preprint arXiv:2210.03350, 2022.
  24. 24.Pryzant, R., Iter, D., Li, J., Lee, Y., Zhu, C., and Zeng, M. Automatic prompt optimization with “gradient descent” and beam search. In Bouamor, H., Pino, J., and Bali, K. (eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 7957–7968, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.494. URL https://aclanthology.org/2023.emnlp-main.494.
  25. 25.Roberts, G. O. and Smith, A. F. Simple conditions for the convergence of the gibbs sampler and metropolis-hastings algorithms. Stochastic processes and their applications, 49(2):207–216, 1994.
  26. 26.Schick, T. and Schütze, H. Exploiting cloze questions for few shot text classification and natural language inference. arXiv preprint arXiv:2001.07676, 2020.
  27. 27.Shwartz, V., West, P., Le Bras, R., Bhagavatula, C., and Choi, Y. Unsupervised commonsense question answering with self-talk. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 4615–4629, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.373. URL https://aclanthology.org/2020.emnlp-main.373.
  28. 28.Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  29. 29.Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  30. 30.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D. Rationale-augmented ensembles in language models. arXiv preprint arXiv:2207.00747, 2022a.
  31. 31.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022b.
  32. 32.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022.
  33. 33.Yoran, O., Wolfson, T., Bogin, B., Katz, U., Deutch, D., and Berant, J. Answering questions by meta-reasoning over multiple chains of thought. arXiv preprint arXiv:2304.13007, 2023.
  34. 34.Zamfirescu-Pereira, J., Wong, R. Y., Hartmann, B., and Yang, Q. Why johnny can’t prompt: How non-ai experts try (and fail) to design llm prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9781450394215. doi: 10.1145/3544548.3581388. URL https://doi.org/10.1145/3544548.3581388.
  35. 35.Zelikman, E., Wu, Y., and Goodman, N. D. STaR: Bootstrapping reasoning with reasoning. arXiv preprint arXiv:2203.14465, 2022.
  36. 36.Zhang, T., Wang, X., Zhou, D., Schuurmans, D., and Gonzalez, J. E. TEMPERA: Test-time prompt editing via reinforcement learning. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=gSHyqBijPFO.
  37. 37.Zhang, Z., Zhang, A., Li, M., and Smola, A. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493, 2022.
  38. 38.Zheng, C., Liu, Z., Xie, E., Li, Z., and Li, Y. Progressive-hint prompting improves reasoning in large language models. arXiv preprint arXiv:2304.09797, 2023.
  39. 39.Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Bousquet, O., Le, Q., and Chi, E. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022.
  40. 40.Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. Large language models are human-level prompt engineers. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=92gvk82DE-.

Citation

MLA
Xu, W., et al. “Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling”. arXiv, 2023, http://arxiv.org/abs/2305.09993v2.
APA
Xu, W., Banburski-Fahey, A., & Jojic, N. (2023). Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling. arXiv. http://arxiv.org/abs/2305.09993v2
Chicago
Xu, W., A. Banburski-Fahey, and N. Jojic. 2023. “Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling”. arXiv. http://arxiv.org/abs/2305.09993v2.
Harvard
Xu, W., Banburski-Fahey, A. and Jojic, N. (2023) “Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.09993v2.
Vancouver
1. Xu W, Banburski-Fahey A, Jojic N (2023) Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling. arXiv

BibTeX

@article{xu2023reprompting,
  title = {Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling},
  author = {Xu, Weijia and Banburski-Fahey, Andrzej and Jojic, Nebojsa},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.09993v2},
  eprint = {2305.09993}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/