GPS: Genetic Prompt Search for Efficient Few-Shot Learning
Hanwei XuYujun ChenYulun DuNan ShaoYanggang WangHaiyu LiZhilin Yang
Proposes a gradient-free genetic algorithm that automatically discovers high-performing, fluent natural language prompts using only a tiny validation set, outperforming both manual prompt engineering and parameter-efficient tuning methods.
Deploying large language models across diverse business applications is often hindered by the high cost of manual data labeling and model fine-tuning. While prompting allows models to perform tasks with minimal data, manually written prompts are typically suboptimal and produce inconsistent results. The article introduces Genetic Prompt Search (GPS), an automated method that uses evolutionary algorithms to discover high-performing text prompts without updating the underlying model parameters.
The main objective of the article is to demonstrate that GPS can automatically optimize discrete prompts using only a tiny validation dataset, matching or exceeding the performance of parameter-tuning methods while reducing computational overhead. To evaluate this, the authors tested GPS across ten diverse natural language processing benchmark tasks using only 32 labeled examples per task. The approach begins with human-written seed prompts, mutates them using techniques like back-translation and generative sentence continuation, and iteratively selects top-performing candidates based on validation accuracy.
The findings show that GPS significantly improves task accuracy compared to existing prompting and parameter-efficient tuning methods. First, GPS achieved an average accuracy of 60.12%, outperforming manual prompts by 2.6 percentage points and beating parameter-efficient techniques like Prompt Tuning (58.56%) and Black-Box Tuning (57.82%). Second, GPS outperformed rule-based discrete prompt search methods such as GRIPS by 1.4 percentage points while generating semantically fluent, interpretable text. Third, sentence continuation powered by larger generative models proved to be the most effective mutation strategy. Finally, GPS required roughly one-tenth the computational operations of traditional tuning methods during search and training, all while keeping model weights frozen.
These results demonstrate that organizations can achieve superior few-shot task accuracy without the high storage, memory, and infrastructure costs required to maintain customized model checkpoints for each use case. Because GPS searches only for discrete text strings, a single frozen foundation model can serve multiple business applications simultaneously with zero added inference latency or task-specific parameter storage.
For enterprise systems operating under limited labeled data, adopting automated discrete prompt search offers a practical and cost-effective alternative to model fine-tuning. When implementing this approach, teams should prioritize generative sentence continuation over simple word-replacement rules. However, decision-makers should note that the evaluation was limited to 10 English benchmark tasks under a strict 32-sample regime. Where thousands of labeled samples are available, traditional fine-tuning may still provide higher peak accuracy, and further evaluation is warranted on specialized enterprise domains.
- Paper: Multitask Prompted Training Enables Zero-Shot Task Generalization, Victor Sanh et al. (2021). Read this account of T0’s multitask prompt training first, since GPS evaluates its search method within the T0 framework.
- Paper: How Can We Know What Language Models Know?, Zhengbao Jiang et al. (2019). Its automatic prompt generation via round-trip translation provides a direct methodological precursor to GPS’s use of back-translation to propose prompt variants.
- Paper: Making Pre-trained Language Models Better Few-shot Learners, Tianyu Gao et al. (2021). Its few-shot prompt-generation and selection methods establish the automated search setting that GPS develops into gradient-free evolutionary search.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). Understanding learned soft-prompt tuning first clarifies the parameter-efficient baseline that GPS contrasts with its discrete, human-readable prompts.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). Its P-Tuning results explain the continuous-prompt approach that GPS compares against while avoiding gradient-based updates.
- Paper: Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution, Chrisantha Fernando et al. (2024). Promptbreeder carries prompt evolution further by evolving the mutation instructions as well as the task prompts, extending GPS’s evolutionary search idea.
- Paper: Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs, Krista Opsahl-Ong et al. (2024). MIPRO extends prompt optimization beyond a single instruction by jointly selecting instructions and demonstrations across multi-stage language-model programs.
- Paper: InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models, Lichang Chen et al. (2024). InstructZero extends black-box prompt optimization by searching a compact continuous space that generates readable instructions for inaccessible models.
- Paper: Reprompting: Automated Chain-of-Thought Prompt Inference Through Gibbs Sampling, Weijia Xu et al. (2024). Reprompting carries automated iterative prompt search into chain-of-thought reasoning, discovering reasoning recipes from examples rather than optimizing general task instructions.
