Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution

Chrisantha FernandoDylan BanarseHenryk MichalewskiSimon OsinderoTim Rocktäschel

article2024ICML656 citationsBest Paper Award

Presents Promptbreeder, an evolutionary framework where large language models self-referentially improve both task prompts and the mutation prompts that modify them, outperforming manual strategies like Chain-of-Thought across arithmetic, commonsense reasoning, and hate-speech classification benchmarks.

Listen

The downstream performance and reasoning capabilities of large language models depend heavily on the quality and phrasing of input prompts. While engineered strategies like Chain-of-Thought prompting enhance performance, manually designing prompts is labor-intensive, heuristic, and frequently suboptimal across diverse applications. Earlier attempts to automate prompt discovery faced diminishing returns after a few iterations due to a loss of exploration diversity. The article addresses the challenge of building a fully automated, general-purpose system capable of open-ended, continuous prompt optimization without requiring costly neural network parameter updates.

The main objective of the article is to demonstrate Promptbreeder, an evolutionary mechanism that autonomously adapts task-prompts for specific domains. It evaluates whether self-referential improvement—where the system evolves both task-prompts and the mutation-prompts that modify them—can surpass state-of-the-art hand-crafted prompting techniques.

Promptbreeder operates by maintaining a population of evolutionary units evaluated over multiple generations against training problem sets. Rather than altering model weights, the framework uses the language model itself as an evolutionary mutation operator applied entirely in natural language. The system initializes diverse prompt candidates by combining high-level problem descriptions with distinct thinking styles and mutation instructions. Across successive generations, Promptbreeder applies five classes of mutation operators, including zero-order generation, lineage history analysis, Lamarckian induction from successful reasoning paths, and hyper-mutations that refine the mutation instructions themselves. Binary tournament selection and diversity-preserving embedding filters ensure effective candidates are retained while avoiding stagnation.

The findings show that Promptbreeder consistently outperforms state-of-the-art baseline prompting strategies across arithmetic, commonsense reasoning, and classification benchmarks. In zero-shot mathematical reasoning on the GSM8K dataset, Promptbreeder achieved 83.9% accuracy, exceeding hand-crafted Chain-of-Thought (66.5%) and optimization baselines like OPRO (80.2%). The system outperformed all comparative baselines across eight standard reasoning benchmarks and surpassed prior automated prompt generation methods on 21 of 24 instruction induction tasks. Ablation analyses confirmed that self-referential hyper-mutations and structured initializations were critical to driving performance gains.

These results demonstrate that language models can effectively self-improve in an entirely gradient-free, post-training setup using natural language as the operational substrate. By treating prompts as executable programs that guide model behavior, organizations can significantly elevate output accuracy, reduce reasoning errors, and automate prompt maintenance across complex domains without incurring the infrastructure costs or engineering constraints of full model fine-tuning.

Organizations deploying large language models should consider automated prompt evolution pipelines for complex or domain-specific tasks rather than relying exclusively on manual engineering. When adopting such approaches, practitioners should seed the search space with diverse heuristics and provide representative training samples to maximize exploration quality. Further engineering is warranted to explore dynamic, multi-step prompt topologies, conditional reasoning graphs, and self-evaluating execution frameworks.

arXiv: 2309.16797
Cover for Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution

Abstract

Popular prompt strategies like Chain-of-Thought Prompting can dramatically improve the reasoning abilities of Large Language Models (LLMs) in various domains. However, such hand-crafted prompt-strategies are often sub-optimal. In this paper, we present PROMPTBREEDER, a general-purpose self-referential self-improvement mechanism that evolves and adapts prompts for a given domain. Driven by an LLM, Promptbreeder mutates a population of task-prompts, evaluates them for fitness on a training set, and repeats this process over multiple generations to evolve task-prompts. Crucially, the mutation of these task-prompts is governed by mutation-prompts that the LLM generates and improves throughout evolution in a self-referential way. That is, Promptbreeder is not just improving task-prompts, but it is also improving the mutation-prompts that improve these task-prompts. Promptbreeder outperforms state-of-the-art prompt strategies such as Chain-of-Thought and Plan-and-Solve Prompting on commonly used arithmetic and commonsense reasoning benchmarks. Furthermore, Promptbreeder is able to evolve intricate task-prompts for the challenging problem of hate speech classification.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Promptbreeder
  • 3.1. Promptbreeder Initialization
  • 3.2. Mutation Operators
  • 3.2.1. DIRECT MUTATION
  • 3.2.2. ESTIMATION OF DISTRIBUTION MUTATION
  • 3.2.3. HYPERMUTATION: MUTATION OF MUTATION-PROMPTS
  • 3.2.4. LAMARCKIAN MUTATION
  • 3.2.5. PROMPT CROSSOVER AND CONTEXT SHUFFLING
  • 4. Experiments
  • 5. Results and Discussion
  • 6. Conclusion and Future Work
  • Impact Statement
  • References
  • α: An Example Evolutionary Run
  • A. Glossary
  • B. Ablations
  • C. Mutation Prompts
  • D. Thinking Styles
  • E. Task Prompts Generated on Initialization
  • F. Promptbreeder as Self-Referential Self-Improvement System
  • G. Using instrospection to generate thinking styles and initial mutation prompts
  • G.1. Level 1 introspection
  • G.2. Level 1 introspection elaborating on entry 1 above
  • G.3. Level 2 introspection elaborating on entry 1 above
  • G.4. Level 3 introspection elaborating on entry 1 above
  • H. Problem Descriptions
  • I. Lamarckian Mutation Example
  • J. Datasets
  • J.1. Control Task-Prompts
  • J.2. Output format strings
  • J.3. Arithmetic Reasoning
  • J.4. Commonsense Reasoning
  • J.5. Hate Speech Classification
  • J.6. Instruction Induction
  • K. Example Results
  • K.1. ETHOS Evolved Task Prompt
  • K.2. Prompt Evolution Maths results
  • K.3. Evolved Mutation Prompts
  • K.4. Mutation Operator Effectiveness
  • K.5. ADDSUB
  • K.6. AQUA
  • K.7. MULTIARITH
  • K.8. GSM8K
  • K.9. SINGLEEQ
  • K.10. SVAMP
  • L. APE Instruction Induction tasks
  • L.1. Best prompts and contexts
  • L.1.1. FIRST LETTER
  • L.1.2. SECOND LETTER
  • L.1.3. LIST LETTERS
  • L.1.4. STARTING WITH
  • L.1.5. PLURALIZATION
  • L.1.6. PASSIVIZATION
  • L.1.7. NEGATION
  • L.1.8. ANTONYMS
  • L.1.9. SYNONYMS
  • L.1.10. MEMBERSHIP
  • L.1.11. RHYMES
  • L.1.12. LARGER ANIMAL
  • L.1.13. CAUSE SELECTION
  • L.1.14. FORMALITY
  • L.1.15. SUM
  • L.1.16. DIFFERENCE
  • L.1.17. NUMBER TO WORD
  • L.1.18. TRANSLATION ENGLISH-GERMAN
  • L.1.19. TRANSLATION ENGLISH-SPANISH
  • L.1.20. TRANSLATION ENGLISH-FRENCH
  • L.1.21. SENTIMENT ANALYSIS
  • L.1.22. SENTENCE SIMILARITY
  • L.1.23. WORD IN CONTEXT
  • M. GPT 3.5 Results
  • N. The Effect of Using Poor Problem Descriptions
  • N.1. Neutral problem description
  • N.2. Misleading problem description
  • O. Experiments in describing the population to the LLM

Knowls

  1. Knowl 1 — Promptbreeder Self-Referential Prompt Evolution Framework

    model/method

    Promptbreeder is a self-referential prompt optimization system for large language models (LLMs) that uses natural language as the substrate for evolutionary self-improvement without requiring gradient computation or neural network parameter updates.

    The fundamental evolutionary entity is a unit of evolution, which consists of:

    1. A set of task-prompts (typically two sequentially applied prompts P1,P2P_1, P_2) designed to condition the LLM's context before answering a domain-specific question QQ.
    2. An associated mutation-prompt MM, which is an instruction directing the LLM on how to modify and improve a task-prompt.
    3. In few-shot evolution settings, a few-shot context CC containing a dynamic set of successful reasoning paths ("workings out") generated during previous successful task evaluations.

    Promptbreeder implements self-referential evolution across two coupled levels:

    • Task-Prompt Evolution: A task-prompt PP is mutated by feeding the concatenation of a mutation-prompt MM and PP into an LLM, generating an offspring task-prompt P′=LLM(M+P)P' = \text{LLM}(M + P), where ++ denotes string concatenation.
    • Mutation-Prompt Evolution (Hypermutation): A mutation-prompt MM is itself mutated by prompting the LLM with a hyper-mutation prompt HH concatenated with MM, generating an offspring mutation-prompt M′=LLM(H+M)M' = \text{LLM}(H + M).

    At inference time, Promptbreeder applies a two-stage prompt execution strategy: the first task-prompt P1P_1 prepended to question QQ elicits an intermediate reasoning continuation R=LLM(P1+Q)R = \text{LLM}(P_1 + Q). The second task-prompt P2P_2, continuation RR, and an output format string SfmtS_{\text{fmt}} are then passed to the LLM to extract the final answer A=LLM(R+P2+Sfmt)A = \text{LLM}(R + P_2 + S_{\text{fmt}}). Fitness is measured as accuracy across a randomly sampled batch of question-answer pairs from the training split.

  2. Knowl 2 — Promptbreeder Evolutionary Algorithm

    algorithm

    Promptbreeder optimizes task-prompts and mutation-prompts using a binary tournament genetic algorithm across discrete generations over a fixed population of evolutionary units.

    Input: Problem description DD, seed mutation-prompts M\mathcal{M}, seed thinking-styles T\mathcal{T}, hyper-mutation prompt HH, training dataset Dtrain\mathcal{D}_{\text{train}}, population size NN, batch size BB, number of generations GG
    Output: Fittest task-prompt pair (P1∗,P2∗)(P_1^*, P_2^*) and mutation-prompt M∗M^*
    Initialize population P←∅\mathcal{P} \leftarrow \emptyset
    for i=1i = 1 to NN do
        Sample M∼MM \sim \mathcal{M} and T1,T2∼TT_1, T_2 \sim \mathcal{T}
        P1←LLM(T1+M+"INSTRUCTION: "+D+" INSTRUCTION MUTANT: ")P_1 \leftarrow \text{LLM}(T_1 + M + \text{"INSTRUCTION: "} + D + \text{" INSTRUCTION MUTANT: "})
        P2←LLM(T2+M+"INSTRUCTION: "+D+" INSTRUCTION MUTANT: ")P_2 \leftarrow \text{LLM}(T_2 + M + \text{"INSTRUCTION: "} + D + \text{" INSTRUCTION MUTANT: "})
        C←∅C \leftarrow \emptyset
        ui←(P1,P2,M,C)u_i \leftarrow (P_1, P_2, M, C)
        Sample batch B∼Dtrain\mathcal{B} \sim \mathcal{D}_{\text{train}} of size BB
        f(ui)←EvaluateFitness(ui,B)f(u_i) \leftarrow \text{EvaluateFitness}(u_i, \mathcal{B})
        P←P∪{ui}\mathcal{P} \leftarrow \mathcal{P} \cup \{u_i\}
    end for
    for g=1g = 1 to GG do
        for step =1= 1 to NN do
            Sample ua,ub∼Pu_a, u_b \sim \mathcal{P} uniformly at random
            if f(ua)≥f(ub)f(u_a) \ge f(u_b) then
                uwinner←uau_{\text{winner}} \leftarrow u_a
                uloser←ubu_{\text{loser}} \leftarrow u_b
            else
                uwinner←ubu_{\text{winner}} \leftarrow u_b
                uloser←uau_{\text{loser}} \leftarrow u_a
            end if
            uchild←Copy(uwinner)u_{\text{child}} \leftarrow \text{Copy}(u_{\text{winner}})
            Sample mutation operator O∼Uniform(O)O \sim \text{Uniform}(\mathcal{O}) from the 9 available operators
            uchild←ApplyOperator(O,uchild,P)u_{\text{child}} \leftarrow \text{ApplyOperator}(O, u_{\text{child}}, \mathcal{P})
            if rand()<0.10\text{rand}() < 0.10 then
                Sample upartner∼Pu_{\text{partner}} \sim \mathcal{P} via fitness-proportionate selection
                uchild.P←upartner.Pu_{\text{child}}.P \leftarrow u_{\text{partner}}.P
            end if
            Sample batch B∼Dtrain\mathcal{B} \sim \mathcal{D}_{\text{train}} of size BB
            f(uchild)←EvaluateFitness(uchild,B)f(u_{\text{child}}) \leftarrow \text{EvaluateFitness}(u_{\text{child}}, \mathcal{B})
            Replace uloseru_{\text{loser}} with uchildu_{\text{child}} in P\mathcal{P}
        end for
    end for
    return u∗∈Pu^* \in \mathcal{P} with highest fitness on Dtrain\mathcal{D}_{\text{train}}

    In standard execution, the population size is N=50N = 50, the evaluation batch size is B=100B = 100 Q&A pairs sampled at random per evaluation step to avoid overfitting, and evolution proceeds for 20–3020\text{--}30 generations (1000–20001000\text{--}2000 total fitness evaluations).

  3. Knowl 3 — Promptbreeder Mutation Operators

    model/method

    Promptbreeder employs nine distinct mutation operators categorized into five classes. During each reproduction event, one operator is sampled uniformly at random:

    1. Direct Mutation:

      • Zero-order Prompt Generation: Concatenates problem description DD with a hint prompt "A list of 100 hints:" to generate a new task-prompt P′=LLM(D+"A list of 100 hints:")P' = \text{LLM}(D + \text{"A list of 100 hints:"}), restarting search directly from the problem description.
      • First-order Prompt Generation: Mutates parent task-prompt PP using the unit's mutation-prompt MM via P′=LLM(M+" INSTRUCTION: "+P+" INSTRUCTION MUTANT: ")P' = \text{LLM}(M + \text{" INSTRUCTION: "} + P + \text{" INSTRUCTION MUTANT: "}).
    2. Estimation of Distribution (EDA) Mutation:

      • EDA Mutation: Feeds a filtered, numbered list of current population task-prompts (excluding prompts with BERT embedding cosine similarity >0.95>0.95 to ensure diversity) in random order to the LLM to predict a continuation prompt.
      • EDA Rank and Index Mutation: Orders filtered population prompts in ascending fitness order but prefixes them with the inverted prompt: "INSTRUCTION: " + M + "\nA List of Responses in descending order of score. " + (K+1) + " is the best response. It resembles " + K + " more than it does (1)". This intentional contradiction prevents the LLM from copying the final prompt and induces high-fitness, diverse extrapolations.
      • Lineage-Based Mutation: Presents the chronological list of historical elite ancestors in the individual's lineage under the header "GENOTYPES FOUND IN ASCENDING ORDER OF QUALITY" to generate a novel prompt continuation.
    3. Hypermutation (Mutation of Mutation-Prompts):

      • Zero-order Hypermutation: Concatenates problem description DD with a sampled thinking style TT to synthesize a new mutation-prompt M′M', which is immediately applied to mutate task-prompt PP.
      • First-order Hypermutation: Applies hypermutation prompt H=H = "Please summarize and improve the following instruction:" to mutate MM via M′=LLM(H+M)M' = \text{LLM}(H + M), and then applies M′M' to update PP.
    4. Lamarckian Mutation:

      • Working Out to Task-Prompt: Reverse-engineers a new task-prompt from successful reasoning phenotypes by prompting the LLM with correct reasoning steps generated during inference via: "I gave a friend an instruction and some advice. Here are the correct examples of his workings out: " + W_{\text{correct}} + " The instruction was:".
    5. Prompt Crossover and Context Shuffling:

      • Prompt Crossover: With a 10% probability, replaces the task-prompt with a task-prompt from another individual chosen via fitness-proportionate selection.
      • Context Shuffling: Updates few-shot reasoning buffers with newly discovered correct solutions and resamples contexts at a 10% base rate.
  4. Knowl 4 — Combinatorial Initialization and Introspection of Prompts

    model/method

    To seed the initial population with diverse cognitive heuristics, Promptbreeder uses a combinatorial initialization method combining three natural language components:

    1. Problem Description (DD): A concise specification of the target domain (e.g., "Solve the math word problem, giving your answer as an arabic numeral.").
    2. Mutation Prompt (M∈MM \in \mathcal{M}): A directive instructing the model on how to vary a prompt (e.g., "Modify this instruction in a way that no self-respecting LLM would!" or "Make a variant of the prompt.").
    3. Thinking Style (T∈TT \in \mathcal{T}): A description of a general cognitive reasoning heuristic (e.g., "Critical Thinking: This style involves analyzing the problem from different perspectives..." or "Let's think step by step").

    An initial task-prompt is synthesized by prompting the LLM: P0=LLM(T+M+" INSTRUCTION: "+D+" INSTRUCTION MUTANT: ")P_0 = \text{LLM}(T + M + \text{" INSTRUCTION: "} + D + \text{" INSTRUCTION MUTANT: "}) For two-stage prompting setups, two task-prompts P0,1P_{0,1} and P0,2P_{0,2} are independently sampled using distinct thinking styles T1,T2∼TT_1, T_2 \sim \mathcal{T} alongside an initial mutation prompt M∼MM \sim \mathcal{M}.

    When domain-specific thinking styles and mutation prompts are unavailable, they can be generated automatically from the problem description alone through hierarchical introspection. The LLM is prompted with "List of 10 Diverse ideas helpful in solving tasks like this one: INSTRUCTION : " + D. The generated items are then recursively expanded across up to three hierarchical levels to yield complete libraries of thinking styles and mutation prompts.

  5. Knowl 5 — Benchmark Performance on Arithmetic and Commonsense Reasoning

    data/table

    Promptbreeder (PB) evaluated with PaLM 2-L outperforms hand-crafted prompt strategies (Zero-shot Chain-of-Thought [CoT], Plan-and-Solve [PS], Plan-and-Solve+ [PS+]), automated prompt search techniques (Automatic Prompt Engineer [APE], Optimization by PROmpting [OPRO]), and initialization baselines across eight arithmetic and commonsense reasoning benchmarks.

    Method (LLM) MultiArith* SingleEq* AddSub* SVAMP* SQA CSQA AQuA-RAT GSM8K
    Zero-shot
    CoT (PaLM 2-L) 99.3 92.0 74.2 86.7 37.3 71.9 37.4 66.5
    PS (PaLM 2-L) 97.7 90.6 72.4 83.8 50.0 77.9 40.2 59.0
    PS+ (PaLM 2-L) 92.5 94.7 74.4 86.3 50.1 73.3 39.4 60.5
    APE (PaLM 2-L) 95.8 82.2 72.2 73.0 38.4 67.3 45.7 77.9
    OPRO (PaLM 2-L) – – – – – – – 80.2
    Combinatorial Init (PaLM 2-L) 99.4 96.1 85.8 87.0 70.1 81.9 57.9 65.5
    PD baseline (PaLM 2-L) 84.0 94.7 87.8 86.0 15.9 85.3 59.4 60.1
    PB (ours) (PaLM 2-L) 99.7 96.4 87.8 90.2 71.8 85.4 62.2 83.9
    Few-shot
    Manual-CoT (PaLM 2-L) 65.7 48.0 74.2 47.7 79.1 87.4 59.4 57.0
    PB (ours) (PaLM 2-L) 100.0 98.9 89.3 93.7 80.2 85.9 64.6 83.5

    Experimental details: Datasets marked with an asterisk split the dataset into equal training and test portions. SQA corresponds to StrategyQA, CSQA to CommonsenseQA, and PD baseline uses the unmutated problem description for both task prompts.

    Key takeaways from the data:

    • In the zero-shot setting on GSM8K, Promptbreeder achieves 83.9%83.9\%, outperforming CoT (66.5%66.5\%), APE (77.9%77.9\%), and OPRO (80.2%80.2\%). The best evolved zero-shot prompt on GSM8K was the single word "SOLUTION".
    • Promptbreeder beats the Combinatorial Initialization baseline across all 8 datasets, demonstrating that ongoing evolutionary search produces substantial gains beyond good initialization (e.g., +18.4 pp+18.4\text{ pp} on GSM8K and +4.3 pp+4.3\text{ pp} on AQuA-RAT).
    • Few-shot Promptbreeder achieves 100.0%100.0\% accuracy on MultiArith, 98.9%98.9\% on SingleEq, 93.7%93.7\% on SVAMP, and 80.2%80.2\% on StrategyQA by co-evolving task-prompts and successful reasoning contexts.
  6. Knowl 6 — Ablation Analysis and Operator Effectiveness in Promptbreeder

    empirical result

    Ablation experiments isolating the self-referential and evolutionary components in Promptbreeder show that removing any operator impairs search performance across nearly all benchmark domains.

    Ablation effects on population mean fitness over 200 evaluation steps (measured on populations of 10 across 8 reasoning datasets):

    • Removal of Thinking-Style Guided Initialization (SR task-prompt): Causes the largest drop in fitness across all datasets (e.g., −13.68-13.68 on CommonsenseQA/StrategyQA, −13.01-13.01 on GSM8K, −11.17-11.17 on StrategyQA, −9.92-9.92 on MultiArith, −7.34-7.34 on SVAMP).
    • Removal of Hypermutation (Hyper): Replacing hypermutation with standard prompt mutation degrades performance across most tasks (e.g., −16.43-16.43 on StrategyQA, −4.36-4.36 on MultiArith, −4.14-4.14 on GSM8K).
    • Removal of Lamarckian Mutation (Lamarck): Removing working-out-to-prompt induction consistently reduces fitness (e.g., −8.12-8.12 on MultiArith, −7.10-7.10 on StrategyQA, −4.68-4.68 on GSM8K).
    • Removal of Mutation-Prompt Pool (SR mut-prompts): Using a static default mutation prompt ("Please summarize and improve the following instruction:") hurts performance on most benchmarks (e.g., −12.05-12.05 on StrategyQA, −2.50-2.50 on SVAMP), though it produced a +4.81+4.81 gain on GSM8K.

    Empirical success rates of individual mutation operators on GSM8K (percentage of applications producing an offspring with fitness higher than its parent):

    1. Zero-order Hyper-Mutation: 42%42\%
    2. Lineage-Based Mutation: 26%26\%
    3. First-order Hyper-Mutation: 23%23\%
    4. EDA Rank and Index Mutation: 12.7%12.7\%
    5. Direct Mutation: 12.0%12.0\%
    6. EDA Mutation: 10.7%10.7\%
    7. Lamarckian Mutation: 6.3%6.3\%

    The three highest-yielding operators all rely on self-referential hypermutation or historical evolutionary lineage.

  7. Knowl 7 — Performance on Instruction Induction Benchmarks

    empirical result

    When evaluated across the 24 Instruction Induction benchmark tasks, Few-shot Promptbreeder using PaLM 2-L (without instruction tuning) matches or outperforms Few-shot Automatic Prompt Engineer (APE, evaluated with text-davinci-002) on 21 out of 24 tasks.

    Notable task-level results:

    • Substantial Improvements: Synonyms (43%43\% vs. 14%14\% for Few-shot APE, +29 pp+29\text{ pp}), Membership (100%100\% vs. 79%79\%, +21 pp+21\text{ pp}), Rhymes (100%100\% vs. 61%61\%, +39 pp+39\text{ pp}), Second Letter (95%95\% vs. 69%69\%, +26 pp+26\text{ pp}), Sentence Similarity (56%56\% vs. 43%43\%, +13 pp+13\text{ pp}), and Word in Context (65%65\% vs. 63%63\%).
    • Perfect Accuracy (100%100\%): Achieved on First Letter, Pluralization, Passivization, Sum, Difference, Number to Word, Cause Selection, Rhymes, and Membership.
    • Failure Modes: Underperforms APE on Common Concept (0%0\% vs. 32%32\%) and Formality (7%7\% vs. 70%70\%), where few-shot context examples drifted away from the semantic objective.

    In tasks initialized without a problem description, Promptbreeder successfully initializes from raw input-output pairs and recovers accurate task instructions via evolutionary mutation and Lamarckian operators.

  8. Knowl 8 — Domain Adaptation to Hate Speech Classification on ETHOS

    empirical result

    Promptbreeder adapts to complex, subjective classification domains by evolving intricate multi-stage prompt rubrics.

    On the ETHOS multi-label hate speech detection dataset, the standard zero-shot baseline prompt "Determine whether a text contains hate speech" achieves 80%80\% accuracy. Promptbreeder evolves a two-stage prompt strategy that increases accuracy to 89%89\%:

    • Stage 1 Prompt (P1P_1): Specifies formal taxonomical definitions of hate speech, identifying derogatory language, negative generalizations, incitement to violence, and hostile/discriminatory statements against protected demographic groups.
    • Stage 2 Prompt (P2P_2): Instructs the model to evaluate the candidate text against a structured three-part decision rubric:
      1. Target Group Identification: Identifying whether the targeted entity involves protected characteristics (race, religion, gender, disability, sexual orientation).
      2. Harmful Speech Identification: Isolating abusive, threatening, or derogatory elements.
      3. Contextual Evaluation: Analyzing speaker intent, audience, and setting, explicitly instructing the model to distinguish genuine hate speech from humor or satire.
  9. Knowl 9 — Cross-Model Transferability and Robustness to Suboptimal Task Descriptions

    empirical result

    Promptbreeder generalizes across underlying language model architectures and demonstrates self-correcting robustness when initialized with uninformative or misleading problem descriptions:

    • Model Generalization on GSM8K:

      • On GPT-3.5-Turbo-0613, Promptbreeder evolves prompts achieving 65.5%65.5\% zero-shot test accuracy, outperforming Zero-shot CoT (52.5%52.5\%) and Plan-and-Solve+ (44.7%44.7\%).
      • On GPT-3.5-Turbo-1106, Promptbreeder achieves 63.9%63.9\% test accuracy (vs. 53.0%53.0\% for Zero-shot CoT).
      • Both zero-shot evolved results exceed OpenAI's published few-shot benchmark accuracy of 57.1%57.1\% for GPT-3.5.
    • Robustness to Flawed Problem Descriptions on GSM8K:

      • Neutral Description ("Write some text"): The fittest prompt at initialization achieves 57.5%57.5\% test accuracy. After 2,5002,500 evaluations, Promptbreeder evolves prompts that improve accuracy to 66.6%66.6\%.
      • Misleading Description ("Write a poem"): The initial population starts at 33.8%33.8\% test accuracy. Over 2,5002,500 evaluations, Promptbreeder self-corrects by evolving math-specific working directives (e.g., "INSTRUCTION MUTANT: Write working out in the answer sheet as well.") to reach 57.7%57.7\% test accuracy.
      • Although evolutionary search recovers from poor seed descriptions, accurate initial problem descriptions yield significantly higher performance (83.9%83.9\%).
  10. Knowl 10 — Limitations and Computational Resource Requirements of Promptbreeder

    limitation

    Promptbreeder possesses several specific structural and computational limitations:

    1. Fixed Prompt Topology: The overall execution structure is fixed as a static two-stage sequential pipeline (P1P_1 generates intermediate reasoning, P2P_2 extracts the final answer). The evolutionary process modifies the natural language text within prompt slots, but cannot dynamically alter the execution graph or create conditional branching architectures.
    2. Dependence on External Verification: While Promptbreeder self-referentially evolves task-prompts and mutation-prompts, the evaluation mechanism is not self-referential; it strictly depends on an externally provided ground-truth dataset and accuracy fitness function.
    3. Context Drift in Few-Shot Evolution: In few-shot evolution, the dynamic reasoning contexts frequently dominate the LLM's output distribution, allowing the accompanying task-prompts to degenerate into uninterpretable text strings without suffering immediate training fitness loss, which risks poorer out-of-distribution generalization.
    4. Token and Compute Cost: A standard run (N=50N = 50, batch size B=100B = 100, 20 generations) consumes approximately 301 million301\text{ million} input tokens and generates 6 million6\text{ million} output tokens, taking approximately 12 hours across 8–16 parallel inference models (amounting to $200–$300 USD on commercial LLM APIs).

Coverage note — None of the substantial contributed material was omitted. All primary algorithmic mechanisms, mutation operator definitions, benchmark tables, ablation analyses, and qualitative findings are fully covered.

References

  1. 1.Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., Chu, E., Clark, J. H., Shafey, L. E., Huang, Y., Meier-Hellstern, K., Mishra, G., Moreira, E., Omernick, M., Robinson, K., Ruder, S., Tay, Y., Xiao, K., Xu, Y., Zhang, Y., Abrego, G. H., Ahn, J., Austin, J., Barham, P., Botha, J., Bradbury, J., Brahma, S., Brooks, K., Catasta, M., Cheng, Y., Cherry, C., Choquette-Choo, C. A., Chowdhery, A., Crepy, C., Dave, S., Dehghani, M., Dev, S., Devlin, J., Dıaz, M., Du, N., Dyer, E., Feinberg, V., Feng, F., Fienber, V., Freitag, M., Garcia, X., Gehrmann, S., Gonzalez, L., Gur-Ari, G., Hand, S., Hashemi, H., Hou, L., Howland, J., Hu, A., Hui, J., Hurwitz, J., Isard, M., Ittycheriah, A., Jagielski, M., Jia, W., Kenealy, K., Krikun, M., Kudugunta, S., Lan, C., Lee, K., Lee, B., Li, E., Li, M., Li, W., Li, Y., Li, J., Lim, H., Lin, H., Liu, Z., Liu, F., Maggioni, M., Mahendru, A., Maynez, J., Misra, V., Moussalem, M., Nado, Z., Nham, J., Ni, E., Nystrom, A., Parrish, A., Pellat, M., Polacek, M., Polozov, A., Pope, R., Qiao, S., Reif, E., Richter, B., Riley, P., Ros, A. C., Roy, A., Saeta, B., Samuel, R., Shelby, R., Slone, A., Smilkov, D., So, D. R., Sohn, D., Tokumine, S., Valter, D., Vasudevan, V., Vodrahalli, K., Wang, X., Wang, P., Wang, Z., Wang, T., Wieting, J., Wu, Y., Xu, K., Xu, Y., Xue, L., Yin, P., Yu, J., Zhang, Q., Zheng, S., Zheng, C., Zhou, W., Zhou, D., Petrov, S., and Wu, Y. PaLM 2 Technical Report, September 2023.
  2. 2.Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Gianinazzi, L., Gajda, J., Lehmann, T., Podstawski, M., Niewiadomski, H., Nyczyk, P., and Hoefler, T. Graph of thoughts: Solving elaborate problems with large language models. CoRR, abs/2308.09687, 2023. doi: 10.48550/arXiv.2308.09687. URL https://doi.org/10.48550/arXiv.2308.09687.
  3. 3.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  4. 4.Chen, A., Dohan, D. M., and So, D. R. Evoprompting: Language models for code-level neural architecture search. CoRR, abs/2302.14838, 2023a. doi: 10.48550/arXiv.2302.14838. URL https://doi.org/10.48550/arXiv.2302.14838.
  5. 5.Chen, L., Chen, J., Goldstein, T., Huang, H., and Zhou, T. Instructzero: Efficient instruction optimization for black-box large language models, 2023b.
  6. 6.Chen, W., Ma, X., Wang, X., and Cohen, W. W. Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, November 2022.
  7. 7.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. CoRR, abs/2110.14168, 2021. URL https://arxiv.org/abs/2110.14168.
  8. 8.Dawkins, R. 13 - The evolution of evolvability. In Kumar, S. and Bentley, P. J. (eds.), On Growth, Form and Computers, pp. 239–255. Academic Press, London, January 2003. ISBN 978-0-12-428765-5. doi: 10.1016/B978-012428765-5/50046-3.
  9. 9.Devlin, J., Chang, M., Lee, K., and Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C., and Solorio, T. (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pp. 4171–4186. Association for Computational Linguistics, 2019. doi: 10.18653/v1/n19-1423. URL https://doi.org/10.18653/v1/n19-1423.
  10. 10.Gajewski, A., Clune, J., Stanley, K. O., and Lehman, J. Evolvability ES: scalable and direct optimization of evolvability. In Auger, A. and Stutzle, T. (eds.), Proceedings of the Genetic and Evolutionary Computation Conference, GECCO 2019, Prague, Czech Republic, July 13-17, 2019, pp. 107–115. ACM, 2019. doi: 10.1145/3321707.3321876. URL https://doi.org/10.1145/3321707.3321876.
  11. 11.Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J. Did aristotle use a laptop? A question answering benchmark with implicit reasoning strategies. Trans. Assoc. Comput. Linguistics, 9:346–361, 2021. doi: 10.1162/tacl_a_00370. URL https://doi.org/10.1162/tacl_a_00370.
  12. 12.Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., and Yang, Y. Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers, September 2023.
  13. 13.Harvey, I. The microbial genetic algorithm. In Advances in Artificial Life. Darwin Meets von Neumann: 10th European Conference, ECAL 2009, Budapest, Hungary, September 13-16, 2009, Revised Selected Papers, Part II 10, pp. 126–133. Springer, 2011.
  14. 14.Hauschild, M. and Pelikan, M. An introduction and survey of estimation of distribution algorithms. Swarm and evolutionary computation, 1(3):111–128, 2011.
  15. 15.Honovich, O., Shaham, U., Bowman, S. R., and Levy, O. Instruction induction: From few examples to natural language task descriptions. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 1935–1952. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.acl-long.108. URL https://doi.org/10.18653/v1/2023.acl-long.108.
  16. 16.Hosseini, M. J., Hajishirzi, H., Etzioni, O., and Kushman, N. Learning to solve arithmetic word problems with verb categorization. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 523–533, Doha, Qatar, October 2014. Association for Computational Linguistics. doi: 10.3115/v1/D14-1058. URL https://aclanthology.org/D14-1058.
  17. 17.Hsieh, C., Li, C., Yeh, C., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C., and Pfister, T. Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 8003–8017. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.findings-acl.507. URL https://doi.org/10.18653/v1/2023.findings-acl.507.
  18. 18.Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J. Large language models can self-improve. CoRR, abs/2210.11610, 2022. doi: 10.48550/arXiv.2210.11610. URL https://doi.org/10.48550/arXiv.2210.11610.
  19. 19.Irie, K., Schlag, I., Csordas, R., and Schmidhuber, J. A modern self-referential weight matrix that learns to modify itself. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 9660–9677. PMLR, 2022. URL https://proceedings.mlr.press/v162/irie22b.html.
  20. 20.Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., Fernando, C., and Kavukcuoglu, K. Population based training of neural networks. CoRR, abs/1711.09846, 2017a. URL http://arxiv.org/abs/1711.09846.
  21. 21.Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. Reinforcement learning with unsupervised auxiliary tasks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017b. URL https://openreview.net/forum?id=SJ6yPD5xg.
  22. 22.Jiang, M., Dennis, M., Parker-Holder, J., Foerster, J. N., Grefenstette, E., and Rocktaschel, T. Replay-guided adversarial environment design. In Ranzato, M., Beygelzimer, A., Dauphin, Y. N., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pp. 1884–1897, 2021a.
  23. 23.Jiang, M., Grefenstette, E., and Rocktaschel, T. Prioritized level replay. In Meila, M. and Zhang, T. (eds.), Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pp. 4940–4950. PMLR, 2021b.
  24. 24.Jiang, M., Rocktaschel, T., and Grefenstette, E. General intelligence requires rethinking exploration. CoRR, abs/2211.07819, 2022. doi: 10.48550/arXiv.2211.07819. URL https://doi.org/10.48550/arXiv.2211.07819.
  25. 25.Kirsch, L. and Schmidhuber, J. Eliminating meta optimization through self-referential meta learning. CoRR, abs/2212.14392, 2022. doi: 10.48550/arXiv.2212.14392. URL https://doi.org/10.48550/arXiv.2212.14392.
  26. 26.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. In NeurIPS, 2022.
  27. 27.Koncel-Kedziorski, R., Hajishirzi, H., Sabharwal, A., Etzioni, O., and Ang, S. D. Parsing algebraic word problems into equations. Transactions of the Association for Computational Linguistics, 3:585–597, 2015. doi: 10.1162/tacl a 00160. URL https://aclanthology.org/Q15-1042.
  28. 28.Lehman, J. and Stanley, K. O. Evolving a diversity of virtual creatures through novelty search and local competition. In Krasnogor, N. and Lanzi, P. L. (eds.), 13th Annual Genetic and Evolutionary Computation Conference, GECCO 2011, Proceedings, Dublin, Ireland, July 12-16, 2011, pp. 211–218. ACM, 2011a. doi: 10.1145/2001576.2001606. URL https://doi.org/10.1145/2001576.2001606.
  29. 29.Lehman, J. and Stanley, K. O. Abandoning Objectives: Evolution Through the Search for Novelty Alone. Evolutionary Computation, 19(2):189–223, June 2011b. ISSN 1063-6560. doi: 10.1162/EVCO a 00025.
  30. 30.Lehman, J., Gordon, J., Jain, S., Ndousse, K., Yeh, C., and Stanley, K. O. Evolution through large models. CoRR, abs/2206.08896, 2022. doi: 10.48550/arXiv.2206.08896. URL https://doi.org/10.48550/arXiv.2206.08896.
  31. 31.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. In Moens, M., Huang, X., Specia, L., and Yih, S. W. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pp. 3045–3059. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.emnlp-main.243. URL https://doi.org/10.18653/v1/2021.emnlp-main.243.
  32. 32.Lin, X., Wu, Z., Dai, Z., Hu, W., Shu, Y., Ng, S.-K., Jaillet, P., and Low, B. K. H. Use your instinct: Instruction optimization using neural bandits coupled with transformers, 2023.
  33. 33.Ling, W., Yogatama, D., Dyer, C., and Blunsom, P. Program induction by rationale generation: Learning to solve and explain algebraic word problems. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 158–167, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1015. URL https://aclanthology.org/P17-1015.
  34. 34.Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. Lost in the middle: How language models use long contexts. CoRR, abs/2307.03172, 2023. doi: 10.48550/arXiv.2307.03172. URL https://doi.org/10.48550/arXiv.2307.03172.
  35. 35.Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J. GPT understands, too. CoRR, abs/2103.10385, 2021. URL https://arxiv.org/abs/2103.10385.
  36. 36.Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Muresan, S., Nakov, P., and Villavicencio, A. (eds.), Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pp. 8086–8098. Association for Computational Linguistics, 2022. doi: 10.18653/v1/2022.acl-long.556. URL https://doi.org/10.18653/v1/2022.acl-long.556.
  37. 37.Madaan, A. and Yazdanbakhsh, A. Text and patterns: For effective chain of thought, it takes two to tango. CoRR, abs/2209.07686, 2022. doi: 10.48550/arXiv.2209.07686. URL https://doi.org/10.48550/arXiv.2209.07686.
  38. 38.Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Welleck, S., Majumder, B. P., Gupta, S., Yazdanbakhsh, A., and Clark, P. Self-refine: Iterative refinement with self-feedback. CoRR, abs/2303.17651, 2023. doi: 10.48550/arXiv.2303.17651. URL https://doi.org/10.48550/arXiv.2303.17651.
  39. 39.Meyerson, E., Nelson, M. J., Bradley, H., Moradi, A., Hoover, A. K., and Lehman, J. Language model crossover: Variation through few-shot prompting. CoRR, abs/2302.12170, 2023. doi: 10.48550/arXiv.2302.12170. URL https://doi.org/10.48550/arXiv.2302.12170.
  40. 40.Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M. G., Rao, K., Sadigh, D., and Zeng, A. Large language models as general pattern machines. CoRR, abs/2307.04721, 2023. doi: 10.48550/arXiv.2307.04721. URL https://doi.org/10.48550/arXiv.2307.04721.
  41. 41.Mollas, I., Chrysopoulou, Z., Karlos, S., and Tsoumakas, G. ETHOS: a multi-label hate speech detection dataset. Complex and Intelligent Systems, 8(6):4663–4678, jan 2022. doi: 10.1007/s40747-021-00608-2. URL https://doi.org/10.1007%2Fs40747-021-00608-2.
  42. 42.Moradi, M. and Samwald, M. Evaluating the robustness of neural language models to input perturbations. In Moens, M., Huang, X., Specia, L., and Yih, S. W. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pp. 1558–1570. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.emnlp-main.117. URL https://doi.org/10.18653/v1/2021.emnlp-main.117.
  43. 43.Mouret, J. and Clune, J. Illuminating search spaces by mapping elites. CoRR, abs/1504.04909, 2015a. URL http://arxiv.org/abs/1504.04909.
  44. 44.Mouret, J.-B. and Clune, J. Illuminating search spaces by mapping elites. arXiv preprint arXiv:1504.04909, 2015b.
  45. 45.Nasir, M. U., Earle, S., Togelius, J., James, S., and Cleghorn, C. Llmatic: Neural architecture search via large language models and quality-diversity optimization. arXiv preprint arXiv:2306.01102, 2023.
  46. 46.Nye, M. I., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A. Show your work: Scratchpads for intermediate computation with language models. CoRR, abs/2112.00114, 2021. URL https://arxiv.org/abs/2112.00114.
  47. 47.Ollinger, M. and Knoblich, G. Psychological research on insight problem solving. In Recasting reality: Wolfgang Pauli’s philosophical ideas and contemporary science, pp. 275–300. Springer, 2009.
  48. 48.OpenAI. GPT-4 Technical Report, March 2023.
  49. 49.Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. CoRR, abs/2304.03442, 2023. doi: 10.48550/arXiv.2304.03442. URL https://doi.org/10.48550/arXiv.2304.03442.
  50. 50.Parker-Holder, J., Jiang, M., Dennis, M., Samvelyan, M., Foerster, J. N., Grefenstette, E., and Rocktaschel, T. Evolving curricula with regret-based environment design. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, volume 162 of Proceedings of Machine Learning Research, pp. 17473–17498. PMLR, 2022. URL https://proceedings.mlr.press/v162/parker-holder22a.html.
  51. 51.Patel, A., Bhattamishra, S., and Goyal, N. Are NLP models really able to solve simple math word problems? In Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-Tur, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (eds.), Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pp. 2080–2094. Association for Computational Linguistics, 2021. doi: 10.18653/v1/2021.naacl-main.168. URL https://doi.org/10.18653/v1/2021.naacl-main.168.
  52. 52.Payne, J. L. and Wagner, A. The causes of evolvability and their evolution. Nature Reviews Genetics, 20(1): 24–38, January 2019. ISSN 1471-0064. doi: 10.1038/s41576-018-0069-z.
  53. 53.Pigliucci, M. Is evolvability evolvable? Nature Reviews Genetics, 9(1):75–82, January 2008. ISSN 1471-0064. doi: 10.1038/nrg2278.
  54. 54.Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M. Automatic prompt optimization with" gradient descent" and beam search. arXiv preprint arXiv:2305.03495, 2023.
  55. 55.Qin, G. and Eisner, J. Learning How to Ask: Querying LMs with Mixtures of Soft Prompts, April 2021.
  56. 56.Roy, S. and Roth, D. Solving general arithmetic word problems. arXiv preprint arXiv:1608.01413, 2016.
  57. 57.Schick, T., Dwivedi-Yu, J., Dessı, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. Toolformer: Language Models Can Teach Themselves to Use Tools, February 2023.
  58. 58.Schmidhuber, J. Making the world differentiable: On using fully recurrent self-supervised neural networks for dynamic reinforcement learning and planning in nonstationary environments. 1990.
  59. 59.Schmidhuber, J. Learning to Control Fast-Weight Memories: An Alternative to Dynamic Recurrent Networks. Neural Computation, 4(1):131–139, January 1992. ISSN 0899-7667. doi: 10.1162/neco.1992.4.1.131.
  60. 60.Schmidhuber, J. A ‘Self-Referential’ Weight Matrix. In Gielen, S. and Kappen, B. (eds.), ICANN ’93, pp. 446–450, London, 1993. Springer. ISBN 978-1-4471-2063-6. doi: 10.1007/978-1-4471-2063-6 107.
  61. 61.Schmidhuber, J. Godel machines: self-referential universal problem solvers making provably optimal self-improvements. arXiv preprint cs/0309048, 2003.
  62. 62.Secretan, J., Beato, N., D Ambrosio, D. B., Rodriguez, A., Campbell, A., and Stanley, K. O. Picbreeder: Evolving pictures collaboratively online. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’08, pp. 1759–1768, New York, NY, USA, April 2008. Association for Computing Machinery. ISBN 978-1-60558-011-1. doi: 10.1145/1357054.1357328.
  63. 63.Shir, O. M. and Back, T. Niching in evolution strategies. In Proceedings of the 7th annual conference on Genetic and evolutionary computation, pp. 915–916, 2005.
  64. 64.Shum, K., Diao, S., and Zhang, T. Automatic prompt augmentation and selection with chain-of-thought from labeled data. CoRR, abs/2302.12822, 2023. doi: 10.48550/arXiv.2302.12822. URL https://doi.org/10.48550/arXiv.2302.12822.
  65. 65.Talmor, A., Herzig, J., Lourie, N., and Berant, J. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4149–4158, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1421. URL https://aclanthology.org/N19-1421.
  66. 66.Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. Voyager: An open-ended embodied agent with large language models. CoRR, abs/2305.16291, 2023a. doi: 10.48550/arXiv.2305.16291. URL https://doi.org/10.48550/arXiv.2305.16291.
  67. 67.Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K., and Lim, E. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 2609–2634. Association for Computational Linguistics, 2023b. doi: 10.18653/v1/2023.acl-long.147. URL https://doi.org/10.18653/v1/2023.acl-long.147.
  68. 68.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  69. 69.Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language models with self-generated instructions. In Rogers, A., Boyd-Graber, J. L., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pp. 13484–13508. Association for Computational Linguistics, 2023c. doi: 10.18653/v1/2023.acl-long.754. URL https://doi.org/10.18653/v1/2023.acl-long.754.
  70. 70.Wang, Z., Cai, S., Liu, A., Ma, X., and Liang, Y. Describe, explain, plan and select: Interactive planning with large language models enables open-world multi-task agents. CoRR, abs/2302.01560, 2023d. doi: 10.48550/arXiv.2302.01560. URL https://doi.org/10.48550/arXiv.2302.01560.
  71. 71.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS, 2022.
  72. 72.Wu, Y., Prabhumoye, S., Min, S. Y., Bisk, Y., Salakhutdinov, R., Azaria, A., Mitchell, T. M., and Li, Y. SPRING: GPT-4 out-performs RL algorithms by studying papers and reasoning. CoRR, abs/2305.15486, 2023. doi: 10.48550/arXiv.2305.15486. URL https://doi.org/10.48550/arXiv.2305.15486.
  73. 73.Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X. Large language models as optimizers. CoRR, abs/2309.03409, 2023a. doi: 10.48550/arXiv.2309.03409. URL https://doi.org/10.48550/arXiv.2309.03409.
  74. 74.Yang, Z., Li, L., Wang, J., Lin, K., Azarnasab, E., Ahmed, F., Liu, Z., Liu, C., Zeng, M., and Wang, L. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv preprint arXiv:2303.11381, 2023b.
  75. 75.Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022.
  76. 76.Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K. Tree of Thoughts: Deliberate Problem Solving with Large Language Models, May 2023.
  77. 77.Zelikman, E., Wu, Y., Mu, J., and Goodman, N. D. Star: Bootstrapping reasoning with reasoning. In NeurIPS, 2022.
  78. 78.Zhang, J., Lehman, J., Stanley, K. O., and Clune, J. OMNI: open-endedness via models of human notions of interestingness. CoRR, abs/2306.01711, 2023a. doi: 10.48550/arXiv.2306.01711. URL https://doi.org/10.48550/arXiv.2306.01711.
  79. 79.Zhang, Z., Zhang, A., Li, M., and Smola, A. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023b. URL https://openreview.net/pdf?id=5NTt8GFjUHkr.
  80. 80.Zhou, D., Scharli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., et al. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022.
  81. 81.Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. Large language models are human-level prompt engineers. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. URL https://openreview.net/pdf?id=92gvk82DE-.

Citation

MLA
Fernando, C., et al. “Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution”. arXiv, 2023, http://arxiv.org/abs/2309.16797v1.
APA
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., & Rocktäschel, T. (2023). Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv. http://arxiv.org/abs/2309.16797v1
Chicago
Fernando, C., D. Banarse, H. Michalewski, S. Osindero, and T. Rocktäschel. 2023. “Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution”. arXiv. http://arxiv.org/abs/2309.16797v1.
Harvard
Fernando, C. et al. (2023) “Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.16797v1.
Vancouver
1. Fernando C, Banarse D, Michalewski H, Osindero S, Rocktäschel T (2023) Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution. arXiv

BibTeX

@article{fernando2023promptbreeder,
  title = {Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution},
  author = {Fernando, Chrisantha and Banarse, Dylan and Michalewski, Henryk and Osindero, Simon and Rocktäschel, Tim},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.16797v1},
  eprint = {2309.16797}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/