Large Language Models as Optimizers
Chengrun YangXuezhi WangYifeng LuHanxiao LiuQuoc V. LeDenny ZhouXinyun Chen
Introduces Optimization by PROmpting (OPRO), a technique that uses large language models as gradient-free optimizers to iteratively generate solutions, discovering prompts that outperform human-engineered instructions by up to 50% on Big-Bench Hard.
Organizations increasingly rely on large language models for complex analytical and reasoning tasks, but their output quality depends heavily on how input prompts are engineered. Manual prompt engineering is labor-intensive, ad hoc, and difficult to scale across diverse tasks, while formal mathematical optimization methods struggle with natural language because text is discrete and lacks mathematical gradients.
The article demonstrates and evaluates Optimization by PROmpting (OPRO), a simple, automated framework that uses large language models as derivative-free optimizers. The objective is to evaluate whether language models can iteratively generate and refine candidate solutions—particularly effective text instructions—based entirely on natural language problem descriptions and previous performance trajectories.
The evaluated approach uses an iterative feedback loop. In each step, an optimizer language model receives a "meta-prompt" containing a description of the task, a few exemplars, and a history of previously generated solutions paired with their evaluation scores. The model proposes new candidate solutions, which an evaluation model scores against a small training set before feeding the results back into the trajectory for subsequent iterations. The authors validated this method across simple continuous and discrete mathematical benchmarks (linear regression and traveling salesman problems) and applied it comprehensively to prompt optimization across multiple leading language models (including PaLM 2-L, text-bison, GPT-3.5-Turbo, and GPT-4) using standard reasoning benchmarks such as GSM8K and Big-Bench Hard.
The key findings demonstrate significant performance gains and broad versatility. First, prompts discovered through automated optimization substantially outperform standard human-designed baselines, achieving up to an 8% accuracy increase on math reasoning (GSM8K) and up to 50% gains across diverse Big-Bench Hard tasks compared to popular baselines like "Let's think step by step." Second, the optimization process exhibits strong sample efficiency, reaching high test performance using only small training subsets (such as 3.5% of GSM8K training examples). Third, instructions optimized on one reasoning dataset transfer effectively to other benchmarks within the same domain, generating consistent performance improvements. Finally, in mathematical case studies, the models successfully identified descent directions on small-scale problems, matching or exceeding hand-designed heuristics on small traveling salesman instances.
These results show that automated prompting can systematically maximize model accuracy and operational reliability without requiring costly human trial-and-error, manual prompt tuning, or task-specific model retraining. Crucially, the method functions via standard application programming interfaces (APIs) without requiring access to internal model parameters or gradients. However, the analysis shows that subtle stylistic variations in prompts cause dramatic swings in accuracy, confirming that intuition-based prompt writing is suboptimal compared to data-driven automated search.
Organizations deploying language models should adopt iterative, automated prompt optimization pipelines rather than relying on manual prompt authoring, especially for critical operational workflows. Practitioners should configure optimization meta-prompts to include task exemplars, maintain historical score rankings, and sample multiple candidates per step to maintain search stability and balance exploration with exploitation. Where feasible, teams should employ early stopping or holdout validation sets to monitor and prevent overfitting.
While highly effective for prompt optimization, the approach has notable limitations. It is not intended to replace specialized mathematical solvers, as performance degrades significantly on large-scale combinatorial problems and complex, irregular optimization landscapes due to context window limits and calculation errors. Additionally, while the optimizer effectively exploits trajectory patterns, it does not yet extract detailed insights directly from individual error cases. Stakeholders can have high confidence in the method's ability to boost prompt performance, but they should exercise caution before applying it to complex non-textual optimization problems.
- Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Its Automatic Prompt Engineer frames instruction generation as black-box search with model-scored candidates, providing a direct antecedent to OPRO’s iterative prompt optimization.
- Paper: Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs, Krista Opsahl-Ong et al. (2024). MIPRO extends automated instruction search to jointly optimize demonstrations across multi-stage language-model programs, carrying prompt optimization beyond OPRO’s single-step setting.
- Paper: Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution, Chrisantha Fernando et al. (2024). Promptbreeder advances automated prompt search with evolutionary populations that improve both task prompts and the instructions used to mutate them.
