Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs

Krista Opsahl-OngMichael J. RyanJosh PurtellDavid BromanChristopher PottsMatei ZahariaOmar Khattab

article2024EMNLP160 citations

Introduces MIPRO, a prompt optimization algorithm that jointly searches for effective instructions and few-shot demonstrations in multi-stage language model programs, boosting downstream task accuracy without requiring module-level labels or gradients.

Listen

Complex tasks in modern natural language processing increasingly rely on multi-stage language model pipelines, which connect multiple model calls to perform step-by-step reasoning and retrieval. However, developing these systems currently demands laborious manual prompt engineering to ensure every step works effectively together. Existing automated prompt optimization methods are generally limited to single-step models and fail in multi-stage settings where intermediate labels, gradients, and internal model probabilities are unavailable.

The article investigates how to automatically and efficiently optimize both free-form instructions and few-shot input-output demonstrations across multi-stage language model programs to maximize end-to-end task performance.

To address this challenge, the authors formalize language model program optimization around two primary bottlenecks: proposing high-quality candidate prompts and assigning credit to determine which pipeline stages drive overall performance gains. They evaluate several optimization algorithms across a diverse benchmark of seven single-stage and multi-stage tasks, including multi-hop question answering, fact verification, logical deduction, and structured classification. The primary experimental setup uses Llama-3-8B as the task model and GPT-3.5 or GPT-4 as the candidate-generation model, evaluating techniques such as grounded prompt generation and Bayesian surrogate modeling.

The findings establish that optimizing both instructions and few-shot demonstrations together via the newly introduced optimizer, MIPRO (Multi-prompt Instruction PRoposal Optimizer), yields the best overall performance, outperforming baseline optimizers on five of seven tasks by margins of up to 13% in accuracy. Second, generating and selecting high-performing few-shot demonstrations proves to be the single most impactful factor for most standard reasoning pipelines. Third, optimizing free-form instructions is essential for complex tasks with nuanced, conditional formatting rules where demonstrations alone cannot convey all requirements. Finally, grounding prompt proposals in dataset and program summaries generally enhances prompt quality, though optimal proposal strategies vary by specific task requirements.

These results demonstrate that automated prompt optimization can substantially lower the cost and development time of deploying multi-stage language model systems, replacing fragile manual prompt tuning with systematic, data-driven optimization. The findings also caution practitioners that intermediate prompt updates can occasionally overfit to task demonstrations, emphasizing the need for robust validation against end-to-end task metrics.

Organizations developing language model workflows should adopt joint instruction and demonstration optimizers like MIPRO when building multi-stage pipelines. For standard tasks with straightforward instructions, teams can achieve substantial performance gains simply by optimizing bootstrapped few-shot examples. When tasks involve strict conditional formatting or custom business logic, teams must optimize free-form instructions from an explicit initial prompt describing those rules. Further empirical work should evaluate these optimization strategies across wider ranges of budget constraints and alternative underlying language models.

The conclusions are supported by statistically significant improvements on standard benchmarks; however, current optimizers still struggle to infer complex task rules entirely from scratch without an initial human-written prompt, and performance across extreme low-budget or high-budget regimes remains an open area for further investigation.

arXiv: 2406.11695stanfordnlp/dspy

No sufficiently relevant recommendations were found.

Cover for Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs

Abstract

Language Model Programs, i.e. sophisticated pipelines of modular language model (LM) calls, are increasingly advancing NLP tasks, but they require crafting prompts that are jointly effective for all modules. We study prompt optimization for LM programs, i.e. how to update these prompts to maximize a downstream metric without access to module-level labels or gradients. To make this tractable, we factorize our problem into optimizing the free-form instructions and few-shot demonstrations of every module and introduce several strategies to craft task-grounded instructions and navigate credit assignment across modules. Our strategies include (i) program- and data-aware techniques for proposing effective instructions, (ii) a stochastic mini-batch evaluation function for learning a surrogate model of our objective, and (iii) a meta-optimization procedure in which we refine how LMs construct proposals over time. Using these insights we develop MIPRO, a novel algorithm for optimizing LM programs. MIPRO outperforms baseline optimizers on five of seven diverse multi-stage LM programs using a best-in-class open-source model (Llama-3-8B), by as high as 13% accuracy. We have released our new optimizers and benchmark in DSPy at this http URL

Table of Contents

  • 1 Introduction
  • 2 Problem Statement
  • 3 Designing LM Program Optimizers
  • 3.1 The Proposal Problem
  • 3.2 Credit Assignment
  • 4 Optimizers
  • 4.1 Bootstrap Random Search
  • 4.2 Module-Level OPRO
  • 4.3 MIPRO
  • 4.4 Other MIPRO variants
  • 4.5 Other OPRO variants
  • 5 Experimental Setup
  • 5.1 Benchmark
  • 5.2 Methods & Models
  • 6 Results & Discussion
  • 7 Related Work
  • 8 Conclusion
  • References
  • A Detailed Task Descriptions
  • B Experiment Setup Details
  • B.1 Data Splits
  • B.2 Optimizer Budget
  • B.3 Optimizer Hyperparameters
  • B.4 Language Model Hyperparameters
  • C Grounding Details
  • C.1 Instruction Proposal Program
  • C.2 Tips
  • C.3 Dataset Summary Generation Process
  • C.4 Example Dataset Summaries
  • C.5 Program Summarization Process
  • D Learned Feature Importances
  • D.1 ScoNe
  • D.2 HotpotQA
  • D.3 HoVeR
  • E DSPy LM Programs
  • E.1 ScoNe
  • E.2 HotpotQA
  • E.3 HoVeR
  • E.4 HotPotQA Conditional With Handwritten Seed Instructions
  • E.5 Iris
  • E.6 Iris-Typo
  • E.7 Heart Disease
  • F Algorithms
  • F.1 Bootstrap Demonstrations
  • F.2 Single-module OPRO
  • F.3 Module-Level History Based
  • F.4 Program-level History Based
  • F.5 Surrogate Model (MIPRO)
  • G Optimization Results
  • H Prompt Progressions

Knowls

  1. Knowl 1 — Prompt optimization for an LM program is a black-box assignment problem

    definition

    A language model program (LM program) Φ\Phi contains mm modules, each with a prompt template and one or more string-valued variables, such as an instruction or few-shot demonstrations. Let VV be the variables selected for optimization, let SS be the space of strings assignable to those variables, and let V↦SV\mapsto S denote a complete assignment. Given training examples D={(xj,xj′)}j=1∣D∣D=\{(x_j,x'_j)\}_{j=1}^{|D|}, where xjx_j is a program input and xj′x'_j is optional target or metadata, and a task-level metric μ\mu, the objective is

    Φ∗=arg⁡max⁡V↦S1∣D∣∑(x,x′)∈Dμ ⁣(ΦV↦S(x),x′).\Phi^* = \arg\max_{V\mapsto S}\frac{1}{|D|}\sum_{(x,x')\in D}\mu\!\left(\Phi_{V\mapsto S}(x),x'\right).

    The metric evaluates the complete program, not its individual modules. The formulation assumes no access to module-level labels, LM weights, gradients, or log probabilities; variables not selected for optimization remain fixed. The practical setting also has limited training data and a limited budget for executing the program.

  2. Knowl 2 — MIPRO jointly searches instruction and demonstration assignments using a Bayesian surrogate

    algorithm

    MIPRO (Multi-prompt Instruction Proposal Optimizer) optimizes a multi-module LM program using candidate instructions and bootstrapped few-shot demonstration sets, with task-level program scores as its feedback. For each module, it first constructs candidate instructions using a grounded proposal process and candidate demonstration sets by running training examples through the program and retaining successful traces. It represents each module’s instruction choice and demonstration-set choice as categorical variables, initialized with uniform priors.

    At each optimization trial, a multivariate Tree-structured Parzen Estimator (TPE) samples a joint assignment of candidates across the program’s variables. The resulting program is evaluated on a randomly sampled minibatch of BB training examples, and the observed score updates the TPE’s estimates of candidate quality and dependencies. Every SS trials, candidate assignments with the highest mean score across their minibatch evaluations are evaluated on the full training set. MIPRO returns the highest-scoring assignment among those fully evaluated. The procedure uses candidate-pool size, minibatch size, and full-evaluation interval as hyperparameters; their values vary by task in the experiments.

  3. Knowl 3 — Grounded proposal supplies task- and program-specific context for instruction generation

    model/method

    The instruction-proposal method uses a separate proposer language model to generate an instruction for a specified module. Its context can include a summary of the training data, a summary of the LM program and its control flow, examples of successful module inputs and outputs, previously evaluated instructions with their task-level scores, a basic instruction, and an optional prompt-engineering tip. The proposer is asked to produce a new instruction for the module rather than to assign credit to the modules based on the program score.

    Dataset summaries are constructed by asking the proposer to add observations while processing training examples in batches; the observation stage ends after five consecutive outputs indicating that there is nothing more to add. The accumulated observations are then condensed into a short summary. Program summaries are generated by giving the program code to a proposer and asking it to describe the task and how the program appears to solve it. The paper treats the benefit of supplying each context component as task-dependent, rather than assuming that more context always helps.

  4. Knowl 4 — Successful program traces provide candidate demonstrations for every module

    algorithm

    Demonstration bootstrapping turns executions of a complete LM program into candidate module-level few-shot examples. Given a training pair (x,x′)(x,x'), run the current program Φ(x)\Phi(x) and retain the execution trace when the final output meets a success criterion, such as μ(Φ(x),x′)≥λ\mu(\Phi(x),x')\geq\lambda, where λ\lambda is a chosen acceptance threshold. For each module in a retained trace, record its input and output as a candidate demonstration. Repeat this process until enough candidate demonstrations have been collected, then form demonstration sets of KK examples for each module.

    Bootstrap Random Search samples joint combinations of these per-module sets, evaluates each parameterized program on the training set or a designated validation split, and returns the highest-scoring assignment. Bayesian Bootstrap instead uses a Bayesian surrogate to select among demonstration candidates. Both methods use only final program-level success to filter traces; they do not require intermediate module labels.

  5. Knowl 5 — Credit assignment strategies trade off isolation, joint modeling, and proposal flexibility

    model/method

    The paper distinguishes three ways to allocate search effort when only a complete LM program receives a score. Greedy credit assignment changes and evaluates one module’s parameters at a time while holding the others fixed. This can reduce attribution ambiguity, but requires sequential evaluations and may miss changes whose benefits depend on improving another module; preliminary experiments found no effectiveness advantage sufficient to offset its cost.

    Surrogate-based credit assignment fits a Bayesian model to previously evaluated combinations of module parameters. The authors use Optuna’s multivariate Tree-structured Parzen Estimator, which can model joint choices and guide search toward promising combinations. Its candidate set is fixed during search, so evaluations do not directly improve how new instructions are proposed.

    History-based credit assignment gives a proposer LM records of previous instructions and their whole-program scores and asks it to propose better instructions. Module-Level OPRO applies the program score to each module’s instruction history, implicitly assuming that each instruction’s contribution is sufficiently reflected in the same score. Program-Level OPRO instead gives one proposer the full multi-module instruction history and relies on it to assign credit across stages. These approaches differ in their credit-assignment assumptions; the paper uses Module-Level OPRO in its experiments and reports no apparent added benefit from the more complex Program-Level variant.

  6. Knowl 6 — MIPRO++ learns how to construct instruction proposals

    model/method

    MIPRO++ applies Bayesian optimization to the settings of the instruction-proposal process rather than directly to the LM program’s instruction and demonstration choices. In the evaluated 0-Shot MIPRO++ variant, the searched settings include whether to provide a dataset summary and a program summary, proposer temperature, which prompt-engineering tip to use, and which bootstrapped demonstrations to show the proposer. A Bayesian surrogate with minibatch evaluation selects these settings, and the resulting proposed instruction is evaluated in the task program. The best fully evaluated program is returned. The full formulation also allows proposal settings for demonstration bootstrapping, but the experiments focus on optimizing instruction proposal.

  7. Knowl 7 — The benchmark tests seven task-program pairs with varied structure and feedback metrics

    experimental setup

    The benchmark comprises seven tasks: HotPotQA multi-hop question answering (two modules, three LM calls, exact match); HotPotQA Conditional, which adds answer-type-dependent formatting rules (two modules, three calls, custom metric); Iris flower classification (one module, one call, accuracy); Iris-Typo, the same classification setup with “versicolour” misspelled in the prompt (one module, one call, accuracy); Heart Disease classification using an ensemble of clinical opinions (two modules, four calls, accuracy); ScoNe entailment with negation reasoning (one module, one call, exact match); and HoVer multi-hop claim-evidence retrieval (four modules, four calls, Recall@21).

    Training/development/test sizes are respectively 500/500/2,000 for HotPotQA, 500/200/200 for HotPotQA Conditional, 75/no development set/75 for Iris, 120/no development set/183 for Heart Disease, 500/500/1,200 for ScoNe, and 500/500/1,520 for HoVer. The HoVer evaluation uses examples with three supporting facts and measures whether the required documents occur among the top-10 retrieved documents across three hops. Most experiments use Llama 3 8B as the task model and GPT-3.5 as the instruction proposer; GPT-4 is used as the proposer for ScoNe and HoVer, and as the demonstration teacher for those tasks. The study runs five trials per method and task, with optimization budgets of 50 full evaluations for HotPotQA and ScoNe, 30 for Iris, Heart Disease, and HotPotQA Conditional, and 20 for HoVer. Minibatching allows methods using it to conduct more minibatch evaluations within those budgets.

  8. Knowl 8 — Joint optimization has the highest test score on five of seven benchmark tasks

    empirical result

    The following are test-set scores on the paper’s 0–100 scale, averaged over five runs. The task order is ScoNe, HotPotQA, HoVer, HotPotQA Conditional, Iris, Iris-Typo, and Heart Disease; a dash indicates that the method was not evaluated for that task.

    • Unoptimized baseline: 69.1, 36.1, 25.3, 6, 40.9, 32, 26.8.
    • Module-Level OPRO without grounding: 76.1, 36.0, 25.7, —, —, —, —.
    • Module-Level OPRO: 73.5, 39.0, 32.5, —, —, —, —.
    • 0-Shot MIPRO: 71.5, 36.8, 33.1, 14.6, 36.4, 56.7, 25.8.
    • 0-Shot MIPRO++: 75.7, 39.3, 32.6, —, —, —, —.
    • Bootstrap Random Search: 75.4, 45.8, 37.2, 10.4, 94.1, 58.7, 79.2.
    • Bayesian Bootstrap: 77.4, 46.2, 37.6, —, —, —, —.
    • MIPRO: 79.4, 46.4, 39.0, 23.3, 88.6, 68.7, 74.2.

    MIPRO attains the highest reported test score on ScoNe, HotPotQA, HoVer, HotPotQA Conditional, and Iris-Typo. Bootstrap Random Search has the highest score on Iris and Heart Disease. The paper’s significance assessment uses Wilcoxon signed-rank tests on per-example test results; it marks multiple methods as comparable when significance over the next-highest average is not established.

  9. Knowl 9 — Demonstrations usually outperform instruction-only tuning, while conditional rules favor instructions

    empirical result

    Across the benchmark experiments, optimizing bootstrapped few-shot demonstrations alone generally performs better than optimizing instructions alone; the reported exception is HotPotQA Conditional. Joint optimization of instructions and demonstrations with MIPRO generally gives the strongest results, though it does not lead on Iris or Heart Disease and is not uniformly best across all tasks. The results support the paper’s qualified interpretation that demonstrations can encode useful reasoning behavior, while instructions can be especially important when a task has conditional rules that are difficult for the model to infer from a small set of examples.

    For HotPotQA Conditional, which requires different output formats for people, places, dates, and other answer types, instruction-only optimization exceeds demonstration-only optimization. For Iris-Typo, optimization improves performance over the misspelled seed prompt, consistent with the reported ability of instruction optimization to help correct that prompt. Grounding effects also vary: the Module-Level OPRO ablation indicates that grounding helps on HotPotQA and HoVer but hurts on ScoNe; 0-Shot MIPRO++ recovers ScoNe performance by learning proposal settings. These outcomes are empirical observations on the benchmark, not guarantees for other tasks.

  10. Knowl 10 — The study leaves model, budget, task-complexity, and seed-instruction generalization unresolved

    limitation

    The experiments use a fixed task model and a fixed proposer-model setup, and evaluate optimizers within a limited set of budgets rather than systematically studying extremely low- or high-budget regimes. The relative performance of methods may change when models or budgets change. The optimizers also have limited ability to infer all rules of a complex task without an informative handwritten seed instruction: dataset or program summaries and demonstrations may not expose every required rule. Finally, the seven-task benchmark does not establish performance on increasingly complex programs or tasks, so broader evaluation is needed.

Coverage note — The paper’s prompt-progression examples and per-task learned-feature-importance plots are omitted because they illustrate particular runs or proposal settings rather than adding a general result beyond the benchmark findings and MIPRO++ method.

References

  1. 1.AI@Meta. 2024. Llama 3 model card.
  2. 2.Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pages 2623–2631.
  3. 3.James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl. 2011. Algorithms for hyper-parameter optimization. Advances in neural information processing systems, 24.
  4. 4.Luca Beurer-Kellner, Marc Fischer, and Martin Vechev. 2023. Prompting is programming: A query language for large language models. Proceedings of the ACM on Programming Languages, 7(PLDI):1946–1969.
  5. 5.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. arXiv preprint arXiv:2211.12588.
  6. 6.Antonia Creswell and Murray Shanahan. 2022. Faithful reasoning using large language models. Preprint, arXiv:2208.14271.
  7. 7.Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, Yihan Wang, Han Guo, Tianmin Shu, Meng Song, Eric Xing, and Zhiting Hu. 2022. RLPrompt: Optimizing discrete text prompts with reinforcement learning. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3369–3391, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  8. 8.Robert C. Detrano, András Jánosi, Walter Steinbrunn, Matthias Emil Pfisterer, Johann-Jakob Schmid, Sarbjit Sandhu, Kern Guppy, Stella Lee, and Victor Froelicher. 1989. International application of a new probability algorithm for the diagnosis of coronary artery disease. The American journal of cardiology, 64 5:304–10.
  9. 9.David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A Saurous, Jascha Sohl-Dickstein, et al. 2022. Language model cascades. arXiv preprint arXiv:2207.10342.
  10. 10.Stefan Falkner, Aaron Klein, and Frank Hutter. 2018. BOHB: Robust and efficient hyperparameter optimization at scale. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 1437–1446. PMLR.
  11. 11.Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. 2023. Promptbreeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797.
  12. 12.Rory A. Fisher. 1936. The use of multiple measurements in taxonomic problems. Annals of Human Genetics, 7:179–188.
  13. 13.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  14. 14.Qingyan Guo, Rui Wang, Junliang Guo, Bei Li, Kaitao Song, Xu Tan, Guoqing Liu, Jiang Bian, and Yujiu Yang. 2024. Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. Preprint, arXiv:2309.08532.
  15. 15.Yaru Hao, Zewen Chi, Li Dong, and Furu Wei. 2022. Optimizing prompts for text-to-image generation. arXiv preprint arXiv:2212.09611.
  16. 16.Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Singh, and Mohit Bansal. 2020. HoVer: A dataset for many-hop fact extraction and claim verification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 3441–3460, Online. Association for Computational Linguistics.
  17. 17.Omar Khattab, Christopher Potts, and Matei Zaharia. 2021. Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval. In Thirty-Fifth Conference on Neural Information Processing Systems.
  18. 18.Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp. arXiv preprint arXiv:2212.14024.
  19. 19.Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan A, Saiful Haq, Ashutosh Sharma, Thomas T. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts. 2024. DSPy: Compiling declarative language model calls into state-of-the-art pipelines. In The Twelfth International Conference on Learning Representations.
  20. 20.Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4582–4597, Online. Association for Computational Linguistics.
  21. 21.Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023. Lost in the middle: How language models use long contexts. arXiv preprint arXiv:2307.03172.
  22. 22.Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr. 2024. Autonomous evaluation and refinement of digital agents. Preprint, arXiv:2404.06474.
  23. 23.Mohammadreza Pourreza and Davood Rafiei. 2023. Din-sql: Decomposed in-context learning of text-to-sql with self-correction. arXiv preprint arXiv:2304.11015.
  24. 24.Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng. 2023. Automatic prompt optimization with" gradient descent" and beam search. arXiv preprint arXiv:2305.03495.
  25. 25.Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023. Toolllm: Facilitating large language models to master 16000+ real-world apis. Preprint, arXiv:2307.16789.
  26. 26.Tal Ridnik, Dedy Kredo, and Itamar Friedman. 2024. Code generation with alphacodium: From prompt engineering to flow engineering. arXiv preprint arXiv:2401.08500.
  27. 27.Imanol Schlag, Sainbayar Sukhbaatar, Asli Celikyilmaz, Wen-tau Yih, Jason Weston, Jürgen Schmidhuber, and Xian Li. 2023. Large language model programs. arXiv preprint arXiv:2305.05364.
  28. 28.Jingyuan S. She, Christopher Potts, Samuel R. Bowman, and Atticus Geiger. 2023. ScoNe: Benchmarking negation reasoning in language models with fine-tuning and in-context learning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1803–1821, Toronto, Canada. Association for Computational Linguistics.
  29. 29.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  30. 30.Jasper Snoek, Hugo Larochelle, and Ryan P. Adams. 2012. Practical bayesian optimization of machine learning algorithms. Preprint, arXiv:1206.2944.
  31. 31.Alessandro Sordoni, Xingdi Yuan, Marc-Alexandre Côté, Matheus Pereira, Adam Trischler, Ziang Xiao, Arian Hosseini, Friederike Niedtner, and Nicolas Le Roux. 2023. Joint prompt optimization of stacked llms using variational inference. Preprint, arXiv:2306.12509.
  32. 32.Dilara Soylu, Christopher Potts, and Omar Khattab. 2024. Fine-tuning and prompt optimization: Two great steps that work better together. Preprint, arXiv:2407.10930.
  33. 33.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903.
  34. 34.Yuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2023. Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery. arXiv preprint arXiv:2302.03668.
  35. 35.Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems, pages 1–22.
  36. 36.Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen. 2023. Large language models as optimizers. arXiv preprint arXiv:2309.03409.
  37. 37.Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V. Le, Denny Zhou, and Xinyun Chen. 2024. Large language models as optimizers. Preprint, arXiv:2309.03409.
  38. 38.Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  39. 39.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. Preprint, arXiv:2210.03629.
  40. 40.Tianjun Zhang, Xuezhi Wang, Denny Zhou, Dale Schuurmans, and Joseph E Gonzalez. 2022. Tempera: Test-time prompt editing via reinforcement learning. In The Eleventh International Conference on Learning Representations.
  41. 41.Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. Sglang: Efficient execution of structured language model programs. Preprint, arXiv:2312.07104.
  42. 42.Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2022. Large language models are human-level prompt engineers. arXiv preprint arXiv:2211.01910.
  43. 43.Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba. 2023. Large language models are human-level prompt engineers. Preprint, arXiv:2211.01910.

Citation

MLA
Opsahl-Ong, K., et al. “Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 9340–66, https://doi.org/10.18653/v1/2024.emnlp-main.525.
APA
Opsahl-Ong, K., Ryan, M. J., Purtell, J., Broman, D., Potts, C., Zaharia, M., & Khattab, O. (2024). Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9340–9366. https://doi.org/10.18653/v1/2024.emnlp-main.525
Chicago
Opsahl-Ong, K., M. J. Ryan, J. Purtell, et al. 2024. “Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs”. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 9340–66. https://doi.org/10.18653/v1/2024.emnlp-main.525.
Harvard
Opsahl-Ong, K. et al. (2024) “Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs”, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 9340–9366. Available at: https://doi.org/10.18653/v1/2024.emnlp-main.525.
Vancouver
1. Opsahl-Ong K, Ryan MJ, Purtell J, Broman D, Potts C, Zaharia M, Khattab O (2024) Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 9340–9366

BibTeX

@inproceedings{opsahl-ong-etal-2024-optimizing,
    title = "Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs",
    author = "Opsahl-Ong, Krista  and
      Ryan, Michael J  and
      Purtell, Josh  and
      Broman, David  and
      Potts, Christopher  and
      Zaharia, Matei  and
      Khattab, Omar",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.525/",
    doi = "10.18653/v1/2024.emnlp-main.525",
    pages = "9340--9366"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/