Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
Krista Opsahl-OngMichael J. RyanJosh PurtellDavid BromanChristopher PottsMatei ZahariaOmar Khattab
Introduces MIPRO, a prompt optimization algorithm that jointly searches for effective instructions and few-shot demonstrations in multi-stage language model programs, boosting downstream task accuracy without requiring module-level labels or gradients.
Complex tasks in modern natural language processing increasingly rely on multi-stage language model pipelines, which connect multiple model calls to perform step-by-step reasoning and retrieval. However, developing these systems currently demands laborious manual prompt engineering to ensure every step works effectively together. Existing automated prompt optimization methods are generally limited to single-step models and fail in multi-stage settings where intermediate labels, gradients, and internal model probabilities are unavailable.
The article investigates how to automatically and efficiently optimize both free-form instructions and few-shot input-output demonstrations across multi-stage language model programs to maximize end-to-end task performance.
To address this challenge, the authors formalize language model program optimization around two primary bottlenecks: proposing high-quality candidate prompts and assigning credit to determine which pipeline stages drive overall performance gains. They evaluate several optimization algorithms across a diverse benchmark of seven single-stage and multi-stage tasks, including multi-hop question answering, fact verification, logical deduction, and structured classification. The primary experimental setup uses Llama-3-8B as the task model and GPT-3.5 or GPT-4 as the candidate-generation model, evaluating techniques such as grounded prompt generation and Bayesian surrogate modeling.
The findings establish that optimizing both instructions and few-shot demonstrations together via the newly introduced optimizer, MIPRO (Multi-prompt Instruction PRoposal Optimizer), yields the best overall performance, outperforming baseline optimizers on five of seven tasks by margins of up to 13% in accuracy. Second, generating and selecting high-performing few-shot demonstrations proves to be the single most impactful factor for most standard reasoning pipelines. Third, optimizing free-form instructions is essential for complex tasks with nuanced, conditional formatting rules where demonstrations alone cannot convey all requirements. Finally, grounding prompt proposals in dataset and program summaries generally enhances prompt quality, though optimal proposal strategies vary by specific task requirements.
These results demonstrate that automated prompt optimization can substantially lower the cost and development time of deploying multi-stage language model systems, replacing fragile manual prompt tuning with systematic, data-driven optimization. The findings also caution practitioners that intermediate prompt updates can occasionally overfit to task demonstrations, emphasizing the need for robust validation against end-to-end task metrics.
Organizations developing language model workflows should adopt joint instruction and demonstration optimizers like MIPRO when building multi-stage pipelines. For standard tasks with straightforward instructions, teams can achieve substantial performance gains simply by optimizing bootstrapped few-shot examples. When tasks involve strict conditional formatting or custom business logic, teams must optimize free-form instructions from an explicit initial prompt describing those rules. Further empirical work should evaluate these optimization strategies across wider ranges of budget constraints and alternative underlying language models.
The conclusions are supported by statistically significant improvements on standard benchmarks; however, current optimizers still struggle to infer complex task rules entirely from scratch without an initial human-written prompt, and performance across extreme low-budget or high-budget regimes remains an open area for further investigation.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Its Automatic Prompt Engineer framework establishes the black-box generation-and-scoring approach to instruction optimization that MIPRO extends to multi-module programs.
No sufficiently relevant recommendations were found.
