Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling
Weijia XuAndrzej BanburskiNebojsa Jojic
Proposes an automated method that uses Gibbs sampling to infer effective chain-of-thought prompts without human intervention, outperforming human-written prompts and existing prompt optimization techniques across twenty complex reasoning benchmarks.
Large language models excel at processing language but struggle with complex logical deduction and mathematical reasoning unless guided by step-by-step demonstrations, commonly known as chain-of-thought prompts. Currently, creating these step-by-step examples requires substantial time and expertise from human engineers who must tailor instructions to each specific problem. This reliance on manual prompt design limits scalability and complicates fair performance comparisons across different language models.
The article introduces and evaluates Reprompting, an automated algorithm designed to discover effective chain-of-thought prompt recipes from a small set of example problems without human intervention. The primary objective is to demonstrate that automated iterative sampling can discover reasoning strategies that match or exceed the quality of human-engineered prompts.
To achieve this, the authors modeled prompt creation as an evolutionary sampling process using Gibbs sampling on twenty example problems per task. The algorithm starts by generating initial zero-shot solutions, uses the best-performing solutions as parent prompts to solve other training problems, and filters out erroneous steps through rejection sampling over thousands of iterations. The researchers evaluated the method across twenty complex reasoning tasks from established benchmarks, testing leading commercial language models including ChatGPT and InstructGPT.
The results establish four major findings. First, prompts generated automatically through Reprompting outperformed expert human-written prompts across the benchmark tasks by an average of 9.4 percentage points in accuracy. Second, the method consistently outperformed existing automated prompt optimization and decoding techniques by margins ranging from 11 to 33 percentage points. Third, combining models—using ChatGPT to propose initial candidate solutions and InstructGPT to iteratively refine them—improved performance by up to 71 percentage points over using InstructGPT alone. Fourth, reasoning strategies did not transfer cleanly across different model architectures; prompts optimized for one model frequently suffered performance drops of 18 to 19 percentage points when applied to another.
These findings indicate that automated prompt discovery eliminates the labor cost and bottleneck of manual prompt engineering while delivering superior model accuracy. The results also demonstrate that comparing language models using identical, static human prompts introduces significant evaluation bias, because different architectures require distinct reasoning patterns to achieve peak performance.
Organizations deploying language models for structured reasoning tasks should adopt automated prompt optimization pipelines instead of relying on manual prompt authoring. Evaluation teams must also tailor prompts to each specific model before benchmarking capabilities. Future initiatives should investigate hybrid approaches, such as combining multiple initialization models or introducing lightweight human feedback into early sampling rounds, to further improve efficiency.
These conclusions are supported by extensive empirical testing across diverse reasoning datasets, though practical adoption requires attention to specific operational boundaries. Reprompting depends on having verified input-answer pairs for the target domain and incurs upfront cloud computing costs from repeated model queries. Nevertheless, confidence in the method is high for standard multi-step reasoning workflows where ground-truth verification is readily available.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). This foundational study establishes chain-of-thought prompting as the reasoning technique Reprompting automates, making its prompt recipes and evaluation easier to follow.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Its Automatic Prompt Engineer frames prompt discovery as generation and performance-based selection, a key precursor to understanding Reprompting’s automated search approach.
No sufficiently relevant recommendations were found.
