Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution
Chrisantha FernandoDylan BanarseHenryk MichalewskiSimon OsinderoTim Rocktäschel
Presents Promptbreeder, an evolutionary framework where large language models self-referentially improve both task prompts and the mutation prompts that modify them, outperforming manual strategies like Chain-of-Thought across arithmetic, commonsense reasoning, and hate-speech classification benchmarks.
The downstream performance and reasoning capabilities of large language models depend heavily on the quality and phrasing of input prompts. While engineered strategies like Chain-of-Thought prompting enhance performance, manually designing prompts is labor-intensive, heuristic, and frequently suboptimal across diverse applications. Earlier attempts to automate prompt discovery faced diminishing returns after a few iterations due to a loss of exploration diversity. The article addresses the challenge of building a fully automated, general-purpose system capable of open-ended, continuous prompt optimization without requiring costly neural network parameter updates.
The main objective of the article is to demonstrate Promptbreeder, an evolutionary mechanism that autonomously adapts task-prompts for specific domains. It evaluates whether self-referential improvement—where the system evolves both task-prompts and the mutation-prompts that modify them—can surpass state-of-the-art hand-crafted prompting techniques.
Promptbreeder operates by maintaining a population of evolutionary units evaluated over multiple generations against training problem sets. Rather than altering model weights, the framework uses the language model itself as an evolutionary mutation operator applied entirely in natural language. The system initializes diverse prompt candidates by combining high-level problem descriptions with distinct thinking styles and mutation instructions. Across successive generations, Promptbreeder applies five classes of mutation operators, including zero-order generation, lineage history analysis, Lamarckian induction from successful reasoning paths, and hyper-mutations that refine the mutation instructions themselves. Binary tournament selection and diversity-preserving embedding filters ensure effective candidates are retained while avoiding stagnation.
The findings show that Promptbreeder consistently outperforms state-of-the-art baseline prompting strategies across arithmetic, commonsense reasoning, and classification benchmarks. In zero-shot mathematical reasoning on the GSM8K dataset, Promptbreeder achieved 83.9% accuracy, exceeding hand-crafted Chain-of-Thought (66.5%) and optimization baselines like OPRO (80.2%). The system outperformed all comparative baselines across eight standard reasoning benchmarks and surpassed prior automated prompt generation methods on 21 of 24 instruction induction tasks. Ablation analyses confirmed that self-referential hyper-mutations and structured initializations were critical to driving performance gains.
These results demonstrate that language models can effectively self-improve in an entirely gradient-free, post-training setup using natural language as the operational substrate. By treating prompts as executable programs that guide model behavior, organizations can significantly elevate output accuracy, reduce reasoning errors, and automate prompt maintenance across complex domains without incurring the infrastructure costs or engineering constraints of full model fine-tuning.
Organizations deploying large language models should consider automated prompt evolution pipelines for complex or domain-specific tasks rather than relying exclusively on manual engineering. When adopting such approaches, practitioners should seed the search space with diverse heuristics and provide representative training samples to maximize exploration quality. Further engineering is warranted to explore dynamic, multi-step prompt topologies, conditional reasoning graphs, and self-evaluating execution frameworks.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Automatic Prompt Engineer (APE) introduces the foundational framework for using LLMs to generate, score, and select discrete prompt candidates, which Promptbreeder directly builds upon and evolves.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Chain-of-thought prompting establishes the core baseline prompting strategy that Promptbreeder aims to outperform through automated evolutionary search.
- Paper: RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning, Mingkai Deng et al. (2022). RLPrompt formalizes the problem of optimizing discrete natural-language prompt tokens for frozen language models via goal-oriented optimization.
- Paper: Large Language Models Can Self-Improve, Jiaxin Huang et al. (2023). This paper establishes the concept of LLM self-improvement via self-generated reasoning and filtering, serving as a conceptual foundation for Promptbreeder's self-referential improvement.
- Paper: Least-to-Most Prompting Enables Complex Reasoning in Large Language Models, Denny Zhou et al. (2022). Least-to-most prompting provides an advanced reasoning prompt decomposition strategy against which Promptbreeder compares and evaluates prompt performance.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Self-consistency establishes multi-path sampling and consensus filtering in LLM reasoning, a key mechanism utilized in prompt evaluation and fitness scoring.
- Paper: Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations, Jaehun Jung et al. (2022). Maieutic prompting illustrates recursive explanation generation and self-corrective reasoning structures that inform recursive and self-referential prompting techniques.
- Paper: Self-Refine: Iterative Refinement with Self-Feedback, Aman Madaan et al. (2023). SELF-REFINE demonstrates how language models can iteratively critique and refine their own outputs in a self-feedback loop, directly preceding self-referential prompt mutation.
- Paper: Learning, Fast and Slow: Towards LLMs That Adapt Continually, Rishabh Tiwari et al. (2026). This work extends evolutionary prompt optimization into a continual learning framework by interleaving prompt evolution with parameter-level reinforcement learning updates.
- Paper: AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Xiaogeng Liu et al. (2024). AutoDAN adapts genetic algorithmic optimization and LLM mutation strategies to iteratively evolve stealthy jailbreak prompts against aligned models.
- Paper: Chain-of-Thought Reasoning Without Prompting, Xuezhi Wang et al. (2024). This paper investigates whether reasoning paths can be elicited directly from model decoding trajectories without relying on prompt engineering or evolution.
- Paper: Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching, Simon A. Aytes et al. (2025). Sketch-of-Thought builds upon prompt optimization paradigms to create concise cognitive shorthand, reducing the token overhead common in evolved chain-of-thought prompts.
