Large Language Models as Optimizers

Chengrun YangXuezhi WangYifeng LuHanxiao LiuQuoc V. LeDenny ZhouXinyun Chen

article2023ICLR1,110 citations

Introduces Optimization by PROmpting (OPRO), a technique that uses large language models as gradient-free optimizers to iteratively generate solutions, discovering prompts that outperform human-engineered instructions by up to 50% on Big-Bench Hard.

Listen

Organizations increasingly rely on large language models for complex analytical and reasoning tasks, but their output quality depends heavily on how input prompts are engineered. Manual prompt engineering is labor-intensive, ad hoc, and difficult to scale across diverse tasks, while formal mathematical optimization methods struggle with natural language because text is discrete and lacks mathematical gradients.

The article demonstrates and evaluates Optimization by PROmpting (OPRO), a simple, automated framework that uses large language models as derivative-free optimizers. The objective is to evaluate whether language models can iteratively generate and refine candidate solutions—particularly effective text instructions—based entirely on natural language problem descriptions and previous performance trajectories.

The evaluated approach uses an iterative feedback loop. In each step, an optimizer language model receives a "meta-prompt" containing a description of the task, a few exemplars, and a history of previously generated solutions paired with their evaluation scores. The model proposes new candidate solutions, which an evaluation model scores against a small training set before feeding the results back into the trajectory for subsequent iterations. The authors validated this method across simple continuous and discrete mathematical benchmarks (linear regression and traveling salesman problems) and applied it comprehensively to prompt optimization across multiple leading language models (including PaLM 2-L, text-bison, GPT-3.5-Turbo, and GPT-4) using standard reasoning benchmarks such as GSM8K and Big-Bench Hard.

The key findings demonstrate significant performance gains and broad versatility. First, prompts discovered through automated optimization substantially outperform standard human-designed baselines, achieving up to an 8% accuracy increase on math reasoning (GSM8K) and up to 50% gains across diverse Big-Bench Hard tasks compared to popular baselines like "Let's think step by step." Second, the optimization process exhibits strong sample efficiency, reaching high test performance using only small training subsets (such as 3.5% of GSM8K training examples). Third, instructions optimized on one reasoning dataset transfer effectively to other benchmarks within the same domain, generating consistent performance improvements. Finally, in mathematical case studies, the models successfully identified descent directions on small-scale problems, matching or exceeding hand-designed heuristics on small traveling salesman instances.

These results show that automated prompting can systematically maximize model accuracy and operational reliability without requiring costly human trial-and-error, manual prompt tuning, or task-specific model retraining. Crucially, the method functions via standard application programming interfaces (APIs) without requiring access to internal model parameters or gradients. However, the analysis shows that subtle stylistic variations in prompts cause dramatic swings in accuracy, confirming that intuition-based prompt writing is suboptimal compared to data-driven automated search.

Organizations deploying language models should adopt iterative, automated prompt optimization pipelines rather than relying on manual prompt authoring, especially for critical operational workflows. Practitioners should configure optimization meta-prompts to include task exemplars, maintain historical score rankings, and sample multiple candidates per step to maintain search stability and balance exploration with exploitation. Where feasible, teams should employ early stopping or holdout validation sets to monitor and prevent overfitting.

While highly effective for prompt optimization, the approach has notable limitations. It is not intended to replace specialized mathematical solvers, as performance degrades significantly on large-scale combinatorial problems and complex, irregular optimization landscapes due to context window limits and calculation errors. Additionally, while the optimizer effectively exploits trajectory patterns, it does not yet extract detailed insights directly from individual error cases. Stakeholders can have high confidence in the method's ability to boost prompt performance, but they should exercise caution before applying it to complex non-textual optimization problems.

  • Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Its Automatic Prompt Engineer frames instruction generation as black-box search with model-scored candidates, providing a direct antecedent to OPRO’s iterative prompt optimization.
Cover for Large Language Models as Optimizers

Abstract

Optimization is ubiquitous. While derivative-based algorithms have been powerful tools for various problems, the absence of gradient imposes challenges on many real-world applications. In this work, we propose Optimization by PROmpting (OPRO), a simple and effective approach to leverage large language models (LLMs) as optimizers, where the optimization task is described in natural language. In each optimization step, the LLM generates new solutions from the prompt that contains previously generated solutions with their values, then the new solutions are evaluated and added to the prompt for the next optimization step. We first showcase OPRO on linear regression and traveling salesman problems, then move on to our main application in prompt optimization, where the goal is to find instructions that maximize the task accuracy. With a variety of LLMs, we demonstrate that the best prompts optimized by OPRO outperform human-designed prompts by up to 8% on GSM8K, and by up to 50% on Big-Bench Hard tasks. Code at this https URL.

Table of Contents

  • 1 Introduction
  • 2 OPRO: LLM as the Optimizer
  • 2.1 Desirables of Optimization by LLMs
  • 2.2 Meta-prompt Design
  • 2.3 Solution Generation
  • 3 Motivating Example: Mathematical Optimization
  • 3.1 Linear Regression
  • 3.2 Traveling Salesman Problem (TSP)
  • 4 Application: Prompt Optimization
  • 4.1 Problem Setup
  • 4.2 Meta-Prompt Design
  • 5 Prompt Optimization Experiments
  • 5.1 Evaluation Setup
  • 5.2 Main Results
  • 5.2.1 GSM8K
  • 5.2.2 BBH
  • 5.2.3 Semantically similar instructions may achieve drastically different accuracies
  • 5.2.4 Transferability of found instructions
  • 5.3 Ablation Studies
  • 5.4 Overfitting Analysis in Prompt Optimization
  • 5.5 Comparison with EvoPrompt
  • 6 Related Work
  • 7 Conclusion
  • References
  • A Some Failure Cases
  • B Prompting Formats for Scorer LLM
  • C Meta-Prompts
  • C.1 Meta-Prompt for Math Optimization
  • C.2 Meta-Prompt for Prompt Optimization
  • D Prompt Optimization Curves on the Remaining BBH Tasks
  • E Prompt Optimization on BBH Tasks – Tabulated Accuracies and Found Instructions
  • E.1 PaLM 2-L-IT as optimizer, optimization starting from the empty string
  • E.2 gpt-3.5-turbo as optimizer, optimization starting from the empty string
  • E.3 PaLM 2-L as scorer, gpt-3.5-turbo as optimizer, optimization starting from “Let’s solve the problem.”

Knowls

  1. Knowl 1 — OPRO iteratively optimizes solutions using an LLM and evaluated solution history

    algorithm

    Optimization by PROmpting (OPRO) uses an LLM as a black-box optimizer. The inputs are a natural-language description of the optimization task and constraints, an evaluator that assigns each candidate a score, an initial set of candidate solutions, and a maximum number of steps. At each step, the optimizer LLM receives a meta-prompt containing the task description and previously evaluated solution–score pairs, then generates candidate solutions. The evaluator scores the new candidates, and the resulting pairs are added to the optimization history for the next step. The prompt can retain only the highest-scoring history entries if context length is limited. The process stops when the optimizer cannot propose a better-scoring solution or the step limit is reached; the best-scoring solutions can then be returned. OPRO requires neither gradients nor a task-specific programmed update rule, and the optimization task can be specified in natural language.

  2. Knowl 2 — OPRO meta-prompts combine an optimization trajectory with task information

    model/method

    An OPRO meta-prompt has two essential components: a natural-language description of the objective and constraints, and an optimization trajectory consisting of earlier candidate solutions paired with their scores. In prompt optimization, it additionally includes examples from the task so the optimizer LLM can infer the input/output format and the desired instruction style. The trajectory is ordered from lower to higher score, allowing high-scoring solutions to appear nearer the end of the prompt; only the best-scoring entries need be retained to respect the context limit. Meta-instructions explain the goal, how to interpret the trajectory and examples, and any desired output format or informal constraints, such as requesting concise solutions. Generating multiple candidates per step is used to reduce instability from any one poor candidate and to explore several possibilities at once; the sampling temperature controls diversity, with lower values favoring smaller changes and higher values encouraging broader exploration.

  3. Knowl 3 — Prompt optimization treats an instruction as the candidate and task accuracy as its score

    model/method

    For prompt optimization, the solution is a natural-language instruction and its objective value is the task accuracy obtained when a scorer LLM uses that instruction. A training split supplies the accuracy used to guide the search; the optimized instruction is assessed on held-out test examples after optimization. The optimizer LLM proposes instructions, while the scorer LLM evaluates them, and the two models may be the same or different. The instruction is inserted into the scorer prompt in one of three positions: Q_begin places it before the question; Q_end places it after the question; A_begin places it at the beginning of the scorer’s answer. A_begin is used with pretrained models prompted in a question–answer sequence, while Q_begin and Q_end are used with instruction-tuned scorers. In the experiments, GSM8K optimization used a fixed random subset of 3.5% of its training data, and each Big-Bench Hard (BBH) task used 20% of its examples for optimization, with the remaining examples reserved for testing.

  4. Knowl 4 — OPRO improves GSM8K test accuracy across several optimizer LLMs

    empirical result

    On GSM8K, the scorer was pretrained PaLM 2-L and the optimized instruction was inserted at A_begin. The best instruction found using instruction-tuned PaLM 2-L as optimizer achieved 80.2% test accuracy; the instruction was “Take a deep breath and work on this problem step-by-step.” Other optimizer LLMs also found high-scoring instructions: PaLM 2-L achieved 79.9% with “Break this down.”, gpt-3.5-turbo achieved 78.5%, and gpt-4 achieved 74.5%. For comparison, the human-designed “Let’s think step by step.” instruction scored 71.8%, “Let’s work this out in a step by step way to be sure we have the right answer.” scored 58.8%, and the empty instruction scored 34.0%. Thus, the best OPRO instruction exceeded the step-by-step baseline by 8.4 percentage points in this evaluation. The optimizer models produced different instruction styles, and high performance was not restricted to instructions with the phrase “step-by-step.”

  5. Knowl 5 — OPRO yields large gains on many BBH tasks, with performance varying by task

    empirical result

    On BBH, prompt optimization began from an empty instruction and used 20% of each task’s examples for optimization; instructions were evaluated on the remaining examples. With PaLM 2-L as scorer, OPRO instructions exceeded the “Let’s think step by step.” baseline by more than 5 percentage points on 19 of 23 tasks; with text-bison as scorer, this occurred on 15 of 23 tasks. Relative to the empty instruction, improvements greater than 5 points occurred on 20 of 23 tasks with PaLM 2-L and 15 of 23 with text-bison. Examples of strong overall accuracies include 90.8% on movie_recommendation with PaLM 2-L and 91.6% on the same task with text-bison. Gains were not universal: for example, on tracking_shuffled_objects_seven_objects the optimized PaLM 2-L instruction achieved 19.6% overall accuracy, compared with 60.8% for “Let’s think step by step.” These results show both broad improvements and substantial task dependence.

  6. Knowl 6 — LLMs can make black-box progress on a small synthetic linear-regression problem

    empirical result

    The linear-regression case study optimized a slope ww and intercept bb for synthetic data generated as y=wtruex+btrue+ϵy=w_{\mathrm{true}}x+b_{\mathrm{true}}+\epsilon, with xx ranging from 1 to 50 and standard Gaussian noise ϵ\epsilon. Each instance contained 50 data points. Optimization began from five randomly sampled (w,b)(w,b) pairs in [10,20]×[10,20][10,20]\times[10,20]. At each step, an instruction-tuned optimizer LLM saw the best 20 historical pairs and their objective values, and generated up to eight new pairs at temperature 1.0; the analytic objective was not disclosed in the meta-prompt. Across five runs per setting, the tested LLMs reached the global optima while evaluating fewer unique pairs than exhaustive search. For a ground truth of (wtrue,btrue)=(15,14)(w_{\mathrm{true}},b_{\mathrm{true}})=(15,14), gpt-4 required a mean of 4.0±1.54.0\pm1.5 steps and explored 17.2±5.117.2\pm5.1 unique pairs. The search became harder as the ground truth moved farther from the starting region: for (36,−1)(36,-1), text-bison used 35.8±6.435.8\pm6.4 steps and explored 174.0±28.2174.0\pm28.2 unique pairs, while gpt-4 used 50.4±18.850.4\pm18.8 steps and explored 116.4±32.7116.4\pm32.7 unique pairs. These results are evidence of black-box optimization on this small problem, not a claim of superiority to specialized continuous optimizers.

  7. Knowl 7 — TSP performance is competitive on small instances but degrades as instance size grows

    data/table

    The TSP experiments sampled node coordinates independently from [−100,100][-100,100] and tested five problem instances at each size. OPRO began with five random tours and generated at most eight new tours per step; Gurobi supplied oracle optima. The table reports optimality gap in percent, and, for successful OPRO runs, mean optimization steps and the number of instances solved optimally. Nearest Neighbor (NN) and Farthest Insertion (FI) are heuristic baselines. OPRO with gpt-4 found every n=10n=10 optimum and had a 1.4% mean gap at n=20n=20, but none of the LLM optimizers found an optimum for any n=50n=50 instance. At n=50n=50, gpt-4’s 11.0% gap was comparable to FI’s 9.8%, while text-bison and gpt-3.5-turbo had much larger gaps.

    Could not parse LaTeX table

    For n=10n=10, gpt-4 reached the optimum in a mean of 9.6 steps, compared with 40.4 for text-bison and 46.8 for gpt-3.5-turbo. As node count increased, OPRO gaps generally rose, and FI outperformed all three LLM optimizers in gap at n=20n=20. The step statistics are reported for successful runs only; N/A indicates no successful run.

  8. Knowl 8 — Meta-prompt ordering, scores, examples, and sampling settings affect optimization

    empirical result

    Ablations used text-bison as scorer and PaLM 2-L as optimizer on GSM8K and BBH sports_understanding, with three optimization repetitions. The default meta-prompt ordered past instructions from lowest to highest score; this converged faster and reached better final accuracy than descending or randomized order. Showing scores also helped: rounding accuracy to integers (100 possible buckets) performed better than using 20 buckets or omitting scores. Task exemplars were important: three exemplars outperformed using none, while increasing the number to ten did not reliably help. Among batches of 1, 2, 4, 8, or 16 generated instructions per step, eight performed best overall under the tested fixed evaluation budget. Among optimizer temperatures 0.0,0.5,1.0,1.5,0.0, 0.5, 1.0, 1.5, and 2.02.0, temperature 1.0 performed best: lower temperatures often repeated the same instruction, while higher temperatures more often failed to exploit the prior trajectory. These findings support the experimental default of retaining the best 20 instructions, showing three exemplars, generating eight candidates per step, using temperature 1.0, and ordering the history from lower to higher scores.

  9. Knowl 9 — Iterative use of the trajectory outperforms one-shot instruction generation

    empirical result

    To test whether iterative optimization adds value, the authors compared OPRO with a baseline that generated 50 instructions in one call, without feeding evaluated instructions and scores back into later generations. With PaLM 2-L-IT as optimizer on GSM8K, the best one-shot instruction was still “Let’s solve the problem,” with 64.4% training and 60.8% test accuracy. OPRO found “Let’s do the math!” after five steps, with 78.2% training and 76.3% test accuracy. On BBH sports_understanding, the best one-shot instruction reached 84.0% training and 80.0% test accuracy, while OPRO reached 88.0% training and 84.5% test accuracy after four steps. In a separate comparison using gpt-3.5-turbo as optimizer, OPRO improved steadily on GSM8K while EvoPrompt’s genetic-algorithm and differential-evolution variants degraded performance from simple initial instructions; on sports_understanding with task-specific initial instructions, EvoPrompt’s differential-evolution variant improved but was less stable than OPRO. The results support using a scored optimization trajectory rather than generating all candidates at once.

  10. Knowl 10 — OPRO has context, landscape, and overfitting limitations

    limitation

    The mathematical-optimization case studies expose limits of OPRO: a finite LLM context makes large problem descriptions difficult to include, and bumpy objective landscapes can cause the optimizer to get stuck. In the reported Rosenbrock example, optimization often moved toward (0,0)(0,0) rather than the global optimum (20,400)(20,400), where the narrow valley made subsequent progress difficult. In prompt optimization, training accuracy can overstate held-out accuracy: the found instructions’ training scores were often 5–20 percentage points higher than test scores. In two validation experiments, however, validation accuracy tended to rise and fall alongside training accuracy, so training-score rankings remained informative in those settings. The authors note that larger training sets and early stopping may reduce overfitting. Prompt optimization also depends on having at least tens of training examples in their implementation, and the optimizer did not reliably use error cases to infer why an instruction failed; providing error cases instead of random training examples produced similar results.

Coverage note — The cross-dataset transfer results on MultiArith and AQuA are omitted as a separate knowl because the ten higher-priority knowls cover the core method, its main evaluations, ablations, and limitations.

References

  1. 1.Shun-ichi Amari. Backpropagation and stochastic gradient descent method. Neurocomputing, 5(4-5): 185–196, 1993.
  2. 2.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023.
  3. 3.David Applegate, Ribert Bixby, Vasek Chvatal, and William Cook. Concorde tsp solver, 2006.
  4. 4.Thomas Bäck and Hans-Paul Schwefel. An overview of evolutionary algorithms for parameter optimization. Evolutionary computation, 1(1):1–23, 1993.
  5. 5.Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022.
  6. 6.Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou. Large language models as tool makers. arXiv preprint arXiv:2305.17126, 2023.
  7. 7.Angelica Chen, David M Dohan, and David R So. Evoprompting: Language models for code-level neural architecture search. arXiv preprint arXiv:2302.14838, 2023a.
  8. 8.Angelica Chen, Jérémy Scheurer, Tomasz Korbak, Jon Ander Campos, Jun Shern Chan, Samuel R Bowman, Kyunghyun Cho, and Ethan Perez. Improving code generation by training with natural language feedback. arXiv preprint arXiv:2303.16749, 2023b.
  9. 9.Jiuhai Chen, Lichang Chen, Heng Huang, and Tianyi Zhou. When do you need chain-of-thought prompting for chatgpt? arXiv preprint arXiv:2304.03262, 2023c.
  10. 10.Lichang Chen, Jiuhai Chen, Tom Goldstein, Heng Huang, and Tianyi Zhou. Instructzero: Efficient instruction optimization for black-box large language models. arXiv preprint arXiv:2306.03082, 2023d.
  11. 11.Xinyun Chen and Yuandong Tian. Learning to perform local rewriting for combinatorial optimization. Advances in Neural Information Processing Systems, 32, 2019.
  12. 12.Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128, 2023e.
  13. 13.Yutian Chen, Xingyou Song, Chansoo Lee, Zi Wang, Richard Zhang, David Dohan, Kazuya Kawakami, Greg Kochanski, Arnaud Doucet, Marc’aurelio Ranzato, et al. Towards learning universal hyperparameter optimizers with transformers. Advances in Neural Information Processing Systems, 35:32053–32068, 2022.
  14. 14.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  15. 15.Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric P Xing, and Zhiting Hu. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548, 2022.
  16. 16.Michel Deudon, Pierre Cournut, Alexandre Lacoste, Yossiri Adulyasak, and Louis-Martin Rousseau. Learning heuristics for the tsp by policy gradient. In International Conference on the Integration of Constraint Programming, Artificial Intelligence, and Operations Research, pp. 170–181. Springer, 2018.
  17. 17.Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797, 2023.
  18. 18.Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas Liao, Kamile Lukoši ˙ ut¯ e, Anna Chen, ˙ Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al. The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459, 2023.
  19. 19.Tianyu Gao, Adam Fisch, and Danqi Chen. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723, 2020.
  20. 20.Bruce Golden, Lawrence Bodin, T Doyle, and W Stewart Jr. Approximate traveling salesman algorithms. Operations research, 28(3-part-ii):694–711, 1980.
  21. 21.Qingyan Guo, Rui Wang, Junliang Guo, Bei Li, Kaitao Song, Xu Tan, Guoqing Liu, Jiang Bian, and Yujiu Yang. Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. arXiv preprint arXiv:2309.08532, 2023.
  22. 22.Gregory Gutin and Abraham P Punnen. The traveling salesman problem and its variations, volume 12. Springer Science & Business Media, 2006.
  23. 23.Keld Helsgaun. An extension of the lin-kernighan-helsgaun tsp solver for constrained traveling salesman and vehicle routing problems. Roskilde: Roskilde University, 12, 2017.
  24. 24.Michael Jünger, Gerhard Reinelt, and Giovanni Rinaldi. The traveling salesman problem. Handbooks in operations research and management science, 7:225–330, 1995.
  25. 25.Geunwoo Kim, Pierre Baldi, and Stephen McAleer. Language models can solve computer tasks. arXiv preprint arXiv:2303.17491, 2023.
  26. 26.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015.
  27. 27.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  28. 28.Wouter Kool, Herke van Hoof, and Max Welling. Attention, learn to solve routing problems! In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=ByxBFsRqYm.
  29. 29.Joel Lehman, Jonathan Gordon, Shawn Jain, Kamal Ndousse, Cathy Yeh, and Kenneth O Stanley. Evolution through large models. arXiv preprint arXiv:2206.08896, 2022.
  30. 30.Brian Lester, Rami Al-Rfou, and Noah Constant. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691, 2021.
  31. 31.Xiang Lisa Li and Percy Liang. Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021.
  32. 32.Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. Program induction by rationale generation: Learning to solve and explain algebraic word problems. arXiv preprint arXiv:1705.04146, 2017.
  33. 33.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. Gpt understands, too. arXiv preprint arXiv:2103.10385, 2021.
  34. 34.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. arXiv preprint arXiv:2104.08786, 2021.
  35. 35.Xiao Ma, Swaroop Mishra, Ahmad Beirami, Alex Beutel, and Jilin Chen. Let’s do a thought experiment: Using counterfactuals to improve moral reasoning. arXiv preprint arXiv:2306.14308, 2023.
  36. 36.Aman Madaan and Amir Yazdanbakhsh. Text and patterns: For effective chain of thought, it takes two to tango. arXiv preprint arXiv:2209.07686, 2022.
  37. 37.Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. Self-refine: Iterative refinement with self-feedback. arXiv preprint arXiv:2303.17651, 2023.
  38. 38.Elliot Meyerson, Mark J Nelson, Herbie Bradley, Arash Moradi, Amy K Hoover, and Joel Lehman. Language model crossover: Variation through few-shot prompting. arXiv preprint arXiv:2302.12170, 2023.
  39. 39.Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montserrat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng. Large language models as general pattern machines. arXiv preprint arXiv:2307.04721, 2023.
  40. 40.Varun Nair, Elliot Schumacher, Geoffrey Tso, and Anitha Kannan. Dera: Enhancing large language model completions with dialog-enabled resolving agents. arXiv preprint arXiv:2303.17071, 2023.
  41. 41.MohammadReza Nazari, Afshin Oroojlooy, Lawrence Snyder, and Martin Takac. Reinforcement learning for solving the vehicle routing problem. In Advances in Neural Information Processing Systems, pp. 9861–9871, 2018.
  42. 42.Theo X Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, and Armando Solar-Lezama. Demystifying gpt self-repair for code generation. arXiv preprint arXiv:2306.09896, 2023.
  43. 43.Gurobi Optimization et al. Gurobi optimizer reference manual, 2020.
  44. 44.Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. Grips: Gradient-free, edit-based instruction search for prompting large language models. arXiv preprint arXiv:2203.07281, 2022.
  45. 45.Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng. Automatic prompt optimization with" gradient descent" and beam search. arXiv preprint arXiv:2305.03495, 2023.
  46. 46.Ning Qian. On the momentum term in gradient descent learning algorithms. Neural networks, 12(1): 145–151, 1999.
  47. 47.Guanghui Qin and Jason Eisner. Learning how to ask: Querying lms with mixtures of soft prompts. arXiv preprint arXiv:2104.06599, 2021.
  48. 48.Colin R Reeves. Modern heuristic techniques for combinatorial problems. John Wiley & Sons, Inc., 1993.
  49. 49.Laria Reynolds and Kyle McDonell. Prompt programming for large language models: Beyond the few-shot paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, pp. 1–7, 2021.
  50. 50.Luis Miguel Rios and Nikolaos V Sahinidis. Derivative-free optimization: a review of algorithms and comparison of software implementations. Journal of Global Optimization, 56:1247–1293, 2013.
  51. 51.Daniel J Rosenkrantz, Richard E Stearns, and Philip M Lewis, II. An analysis of several heuristics for the traveling salesman problem. SIAM journal on computing, 6(3):563–581, 1977.
  52. 52.Subhro Roy and Dan Roth. Solving general arithmetic word problems. arXiv preprint arXiv:1608.01413, 2016.
  53. 53.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761, 2023.
  54. 54.Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  55. 55.Noah Shinn, Beck Labash, and Ashwin Gopinath. Reflexion: an autonomous agent with dynamic memory and self-reflection. arXiv preprint arXiv:2303.11366, 2023.
  56. 56.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  57. 57.Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261, 2022.
  58. 58.Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291, 2023.
  59. 59.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdh￾ery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  60. 60.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  61. 61.Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, et al. Larger language models do in-context learning differently. arXiv preprint arXiv:2303.03846, 2023.
  62. 62.Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. arXiv preprint arXiv:2302.03668, 2023.
  63. 63.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244, 2023.
  64. 64.Hanwei Xu, Yujun Chen, Yulun Du, Nan Shao, Yanggang Wang, Haiyu Li, and Zhilin Yang. Gps: Genetic prompt search for efficient few-shot learning. arXiv preprint arXiv:2210.17041, 2022.
  65. 65.Weizhe Yuan, Kyunghyun Cho, and Jason Weston. System-level natural language feedback. arXiv preprint arXiv:2306.13588, 2023.
  66. 66.Tianjun Zhang, Xuezhi Wang, Denny Zhou, Dale Schuurmans, and Joseph E Gonzalez. Tempera: Test-time prompt editing via reinforcement learning. In The Eleventh International Conference on Learning Representations, 2023.
  67. 67.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pp. 12697–12706. PMLR, 2021.
  68. 68.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022a.
  69. 69.Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. Large language models are human-level prompt engineers. arXiv preprint arXiv:2211.01910, 2022b.

Citation

MLA
Yang, C., et al. “Large Language Models as Optimizers”. arXiv, 2023, http://arxiv.org/abs/2309.03409v3.
APA
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., & Chen, X. (2023). Large Language Models as Optimizers. arXiv. http://arxiv.org/abs/2309.03409v3
Chicago
Yang, C., X. Wang, Y. Lu, et al. 2023. “Large Language Models as Optimizers”. arXiv. http://arxiv.org/abs/2309.03409v3.
Harvard
Yang, C. et al. (2023) “Large Language Models as Optimizers”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2309.03409v3.
Vancouver
1. Yang C, Wang X, Lu Y, Liu H, Le QV, Zhou D, Chen X (2023) Large Language Models as Optimizers. arXiv

BibTeX

@article{yang2023large,
  title = {Large Language Models as Optimizers},
  author = {Yang, Chengrun and Wang, Xuezhi and Lu, Yifeng and Liu, Hanxiao and Le, Quoc V. and Zhou, Denny and Chen, Xinyun},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2309.03409v3},
  eprint = {2309.03409}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/publicdomain/zero/1.0/