InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models

Lichang ChenJiuhai ChenTom GoldsteinHeng HuangTianyi Zhou

article2024ICML83 citations

Proposes a framework that automatically optimizes instructions for black-box language models by applying Bayesian optimization to continuous soft prompts on an open-source model, significantly outperforming existing prompt generation methods on zero-shot benchmarks.

Listen

Commercial artificial intelligence deployments increasingly rely on black-box large language models accessed solely through application programming interfaces. While these models are highly versatile, their task performance depends heavily on the specific phrasing of textual prompts. Traditional prompt engineering relies on costly, slow trial-and-error by human experts. Furthermore, standard automated optimization techniques cannot be applied directly because proprietary model internals are inaccessible and natural language instructions represent an intractable, high-dimensional search space.

The article demonstrates a new framework, named InstructZero, designed to automate prompt optimization for black-box language models without needing internal model access or human intervention. The approach aims to generate high-performing, human-readable instructions by coupling continuous mathematical optimization with accessible open-source language models.

To accomplish this, the method delegates instruction creation to an open-source model such as Vicuna. Instead of searching through countless combinations of discrete words, the framework optimizes a low-dimensional numerical prompt vector. When paired with a small number of task examples, this prompt guides the open-source model to produce a complete natural language instruction. That instruction is then sent to the target black-box model, such as ChatGPT, for zero-shot evaluation on validation data. A Bayesian optimization algorithm uses the resulting performance score and a specialized kernel to iteratively propose better prompt vectors over successive cycles.

The evaluation across 32 natural language understanding tasks yielded several clear results. First, the proposed method achieved top performance on all 32 evaluated tasks, consistently outperforming established automated prompt baselines such as Automatic Prompt Engineer and uniform random exploration. Second, the accuracy gains were substantial on challenging benchmarks; the framework achieved improvements between 20 and 100 percentage points over baselines on a majority of the tested tasks. Third, ablation analyses showed that a 13-billion-parameter open-source model using InstructZero could generate prompts that outperformed instructions written by humans or generated by much larger systems. Finally, across iterative rounds, the quality and accuracy of the generated instructions steadily improved toward optimal clarity.

These findings indicate that organizations can significantly enhance proprietary model performance and consistency while eliminating manual prompt engineering overhead. By utilizing smaller, cost-effective open-source models to orchestrate inputs for commercial endpoints, teams can reduce trial-and-error costs, accelerate production timelines, and systematically discover high-performing phrasing strategies without fine-tuning underlying models.

Organizations seeking to optimize automated workflows should consider adopting automated prompt frameworks for downstream tasks, utilizing small validation sets of task examples to drive optimization cycles. However, decision-makers should recognize several operational boundaries. The framework relies on validation accuracy as a guiding signal, which depends on well-defined evaluation metrics and access to representative sample data. Additionally, generating and evaluating candidate prompts requires upfront computational runs and API usage. Overall, the evidence provides strong confidence that combining latent prompt optimization with open-source models is a robust, superior alternative to manual prompt engineering for black-box systems.

  • Paper: Large Language Models Are Human-Level Prompt Engineers, Yongchao Zhou et al. (2022). Its Automatic Prompt Engineer frames instruction generation and selection as black-box optimization, providing the direct baseline InstructZero improves upon.
  • Paper: RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning, Mingkai Deng et al. (2022). RLPrompt establishes reward-guided optimization of discrete prompts, clarifying the search-space challenge that InstructZero addresses with low-dimensional continuous vectors.
  • Paper: GPT Understands, Too, Xiao Liu et al. (2021). P-Tuning introduces continuous prompt representations, giving useful grounding for InstructZero’s optimization of numerical prompt vectors.
Cover for InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models

Abstract

Large language models (LLMs) are instruction followers but the performance varies under different instructions. It is challenging to create the best instruction, especially for black-box LLMs on which backpropagation is forbidden. Instead of directly optimizing the discrete instruction, we optimize a low-dimensional soft prompt applied to an open-source LLM to generate the instruction for the black-box LLM. In each optimization step of the proposed method INSTRUCTZERO, a soft prompt is converted into an instruction by the open-source LLM, which is then submitted to the black-box LLM for zero-shot evaluation, whose result is sent to Bayesian optimization to produce new soft prompts improving the zero-shot performance. We evaluate INSTRUCTZERO on different combinations of open-source LLMs and APIs including Vicuna and ChatGPT. INSTRUCTZERO outperforms SOTA auto-instruction methods across a variety of downstream tasks. Our code is available: https://github.com/Lichang-Chen/InstructZero.

Table of Contents

  • 1. Introduction
  • 2. Instruction Optimization
  • 2.1. Problem Formulation
  • 2.2. From Structured Combinatorial Search to Low-dimensional Continuous Optimization
  • 3. Bayesian optimization with InstructionCoupled Kernel
  • 3.1. Bayesian Optimization of Soft Prompt
  • 3.2. Instruction-Coupled Kernel
  • 4. Experiments
  • 4.1. Tasks, Datasets, Baselines, and Implementation
  • 4.2. Main Results
  • 4.3. Ablation Study
  • 4.4. Case Study
  • 5. Related Work
  • 6. Discussion, Conclusions, and Limitations
  • Impact Statement
  • Acknowledgement
  • References
  • A. Supplementary Material
  • B. Frequently Asked Questions
  • B.1. Why is the performance of APE quite poor on ChatGPT?
  • B.2. Code Availability
  • B.3. Choices of Kernel in Bayesian Optimization
  • B.4. Optimization process on more Tasks
  • C. Evaluation Metrics
  • D. Different combinations of API LLM + Open-source LLM
  • E. Comparison of InstructZero instructions and human instructions

Knowls

  1. Knowl 1 — Instruction optimization via a frozen open-source instruction generator

    model/method

    INSTRUCTZERO optimizes instructions for a task handled by a black-box language model without gradients or access to its parameters. Let ff be the black-box model, vv a natural-language instruction supplied with query XX, and hh a task score comparing the model output with the correct answer YY. For task data (X,Y)(X,Y) drawn from distribution DtD_t, the target is to find an instruction maximizing expected performance:

    max⁡v∈V  E(X,Y)∼Dt[h(f([v;X]),Y)].\max_{v\in\mathcal V}\;\mathbb E_{(X,Y)\sim D_t}\left[h\left(f([v;X]),Y\right)\right].

    Here [v;X][v;X] denotes the instruction and query presented together, and V\mathcal V is the space of readable, task-relevant instructions. Rather than search this discrete space directly, INSTRUCTZERO uses a frozen open-source model gg to turn a continuous soft prompt and κ\kappa task exemplars (xi,yi)(x_i,y_i) into an instruction. It evaluates the generated instruction with the frozen black-box model on task examples and uses those scores to guide soft-prompt search. Neither model is fine-tuned.

  2. Knowl 2 — Low-dimensional soft prompts parameterize instruction generation

    model/method

    INSTRUCTZERO represents the search variable as p∈Rdp\in\mathbb R^d and maps it to the soft-prompt embedding input expected by the open-source generator, whose embedding dimension is d′d', with d≪d′d\ll d'. A random linear projection AA produces the input prompt ApAp. Given κ\kappa exemplars (xi,yi)(x_i,y_i) from the target task, the generated instruction and its objective are

    v(p)=g([Ap;(xi,yi)i=1κ]),H(p)=E(X,Y)∼Dt[h(f([v(p);X]),Y)].v(p)=g\left([Ap;(x_i,y_i)_{i=1}^{\kappa}]\right),\qquad H(p)=\mathbb E_{(X,Y)\sim D_t}\left[h\left(f([v(p);X]),Y\right)\right].

    Here gg is the open-source instruction generator, ff is the black-box model being prompted, DtD_t is the target-task distribution, and hh is the task score. The entries of the random projection are sampled from a normal or uniform distribution; the experiments use a uniform distribution. The authors motivate dimension reduction by the distance-preserving behavior of random projections and by the generator's ability to produce diverse task-relevant instructions from the prompt together with exemplars.

  3. Knowl 3 — Instruction-coupled kernel transfers instruction similarity into soft-prompt search

    equation

    For mm previously evaluated soft prompts p1,…,pm∈Rdp_1,\ldots,p_m\in\mathbb R^d and their generated instructions v1,…,vmv_1,\ldots,v_m, INSTRUCTZERO combines a soft-prompt kernel with a kernel measuring similarity between the instructions' black-box predictions. Let l(pi,pj)l(p_i,p_j) be a kernel on soft prompts, let L∈Rm×mL\in\mathbb R^{m\times m} have entries Lab=l(pa,pb)L_{ab}=l(p_a,p_b), and let S∈Rm×mS\in\mathbb R^{m\times m} have entries

    Sab=s(va,vb)=EX∼Dt[sim⁡(f([va;X]),f([vb;X]))],S_{ab}=s(v_a,v_b)=\mathbb E_{X\sim D_t}\left[\operatorname{sim}\left(f([v_a;X]),f([v_b;X])\right)\right],

    where sim⁡\operatorname{sim} measures similarity of predictions, with exact match, F1, and BLEU given as examples. For any soft prompts pi,pjp_i,p_j, define vectors li=[l(pi,p1),…,l(pi,pm)]⊤l_i=[l(p_i,p_1),\ldots,l(p_i,p_m)]^\top and lj=[l(pj,p1),…,l(pj,pm)]⊤l_j=[l(p_j,p_1),\ldots,l(p_j,p_m)]^\top. The instruction-coupled kernel is

    k(pi,pj)=li⊤L−1SL−1lj.k(p_i,p_j)=l_i^\top L^{-1}SL^{-1}l_j.

    The formula uses the inverse of LL. On the observed prompts, the resulting kernel matrix recovers SS, so it encodes the observed instruction-prediction similarities; for new prompts it provides a smooth extrapolation based on the soft-prompt kernel. This couples Bayesian optimization in the continuous prompt space to similarity in the instruction space.

  4. Knowl 4 — Gaussian-process Bayesian optimization selects prompts by expected improvement

    equation

    INSTRUCTZERO models the black-box objective over soft prompts with a zero-mean Gaussian-process prior and a covariance kernel kk. After evaluating mm prompts p1,…,pm∈Rdp_1,\ldots,p_m\in\mathbb R^d, let yi=H(pi)y_i=H(p_i) be their scores, y1:m=[y1,…,ym]⊤y_{1:m}=[y_1,\ldots,y_m]^\top, and let K∈Rm×mK\in\mathbb R^{m\times m} have entries Kij=k(pi,pj)K_{ij}=k(p_i,p_j). For a candidate pp, define kp=[k(p,p1),…,k(p,pm)]⊤k_p=[k(p,p_1),\ldots,k(p,p_m)]^\top, let η\eta denote the observation-noise level, and let II be the m×mm\times m identity matrix. The posterior mean and variance are

    μ(p)=kp⊤(K+η2I)−1y1:m,σ2(p)=k(p,p)−kp⊤(K+η2I)−1kp.\mu(p)=k_p^\top(K+\eta^2 I)^{-1}y_{1:m},\qquad \sigma^2(p)=k(p,p)-k_p^\top(K+\eta^2 I)^{-1}k_p.

    The expected-improvement acquisition value is the posterior expectation of the improvement over the highest observed score, and the next prompt is chosen to maximize it:

    u(p)=EZ∼N(μ(p),σ2(p))[max⁡(0,Z−max⁡1≤i≤myi)],pm+1∈arg⁡max⁡p∈Rdu(p).u(p)=\mathbb E_{Z\sim\mathcal N(\mu(p),\sigma^2(p))}\left[\max\left(0,Z-\max_{1\le i\le m}y_i\right)\right], \qquad p_{m+1}\in\arg\max_{p\in\mathbb R^d}u(p).

    Here ZZ is the Gaussian-process posterior random variable for the objective at pp. The posterior mean and uncertainty jointly govern exploitation of promising prompts and exploration of uncertain ones.

  5. Knowl 5 — INSTRUCTZERO search procedure

    algorithm

    The procedure searches continuous prompts while evaluating only the natural-language instructions produced by the open-source generator. Its inputs are target-task exemplars, a validation set, frozen open-source and black-box models, a random projection, and a maximum search budget. Its output is the instruction associated with the highest validation score observed.

    Input: Exemplars (xi,yi)i=1κ(x_i,y_i)_{i=1}^{\kappa}, validation set DtD_t, frozen generator gg, frozen black-box model ff, random projection AA, maximum evaluations TT
    Initialize a prompt p1p_1 in the search domain; initialize the observed prompt, instruction, and score sets as empty
    For each evaluation, up to TT or until convergence:
        Map the current low-dimensional prompt to the generator input using ApA p
        Generate an instruction v=g([Ap;(xi,yi)i=1κ])v=g([A p;(x_i,y_i)_{i=1}^{\kappa}])
        Evaluate f([v;X])f([v;X]) on validation examples and compute the task score H(p)H(p)
        Store the prompt, generated instruction, and score
        Update the instruction-coupled kernel from the stored prompts and instructions
        Update the Gaussian-process posterior and expected-improvement acquisition function
        Select the next prompt by maximizing expected improvement
    Return the stored instruction with the largest validation score

    In the reported implementation, the search is batched: it selects 25 prompts with the largest acquisition values per iteration, using CMA-ES to find those candidates. The chosen instruction is the best-scoring one observed, not necessarily the last instruction evaluated.

  6. Knowl 6 — Experimental protocol and implementation

    experimental setup

    The primary experiments pair Vicuna-13B as the open-source instruction generator with ChatGPT as the black-box model. The benchmark contains 32 instruction-following tasks: 24 tasks used in prior auto-instruction work and 8 additional tasks. For each task, the authors draw 5 training examples as generator exemplars and 20 training examples as the validation set; final performance is measured on a held-out test set. The optimization score is computed from validation outputs, while task evaluation uses the applicable metric: exact match, exact set, containment, or F1.

    The low-dimensional prompt has d=10d=10. The number of soft-prompt tokens is selected from {3,5,10}\{3,5,10\} using validation performance, and projection entries are sampled uniformly from [−1,1][-1,1]. The implementation explores 25 candidate prompts per iteration and uses CMA-ES to search for candidates with high acquisition values. INSTRUCTZERO is compared with APE, whose instructions are generated by ChatGPT, and Uniform, which uses the same models as INSTRUCTZERO and draws the same total number of soft prompts uniformly without iterative Bayesian optimization. The reported experiments were run on an NVIDIA RTX A6000 GPU.

  7. Knowl 7 — INSTRUCTZERO leads on all 32 benchmark tasks

    empirical result

    On the 32-task benchmark, ChatGPT prompted with INSTRUCTZERO-generated instructions achieved the highest zero-shot test accuracy among INSTRUCTZERO, APE, and Uniform on every task (32/32), as reported in the benchmark comparison. On easy tasks where a baseline already reached 100% accuracy, INSTRUCTZERO was comparable at the ceiling; its largest relative advantages appeared on more challenging tasks.

    For the Taxonomy Animal example shown in the paper, the task is to select animals from an input list. The reported zero-shot accuracies were 0.04 for APE, 0.72 for Uniform, and 0.92 for INSTRUCTZERO. The associated INSTRUCTZERO instruction, asking for a list of the animals from the input list, was more task-precise than the APE instruction, which mistakenly described sorting and selecting list positions.

  8. Knowl 8 — Optimizing the soft prompt improves instruction-generation accuracy

    data/table

    This ablation compares INSTRUCTZERO with two instruction-generation inputs while keeping the broader evaluation setting: a human-written manual meta-prompt in place of the optimized soft prompt, and no prompt beyond the task exemplars. The reported execution accuracies show that the learned soft prompt substantially improves results on these seven tasks, although the no-prompt variant is better than the manual prompt on some tasks.

    Task Manual prompt Exemplars only INSTRUCTZERO
    Cause and effect 0.36 0.56 0.91
    Negation 0.27 0.01 0.80
    Translation EN–FR 0.02 0.47 0.89
    Sum 0.00 0.00 1.00
    Formality 0.59 0.31 0.63
    Letters list 0.00 0.15 1.00
    Larger animal 0.49 0.81 0.91

    The manual condition replaces the optimized prompt with the human-written meta-prompt used by APE; the exemplars-only condition removes the prompt entirely. INSTRUCTZERO has the highest score in all seven listed rows, including 1.00 versus 0.00 for the manual prompt on Letters List.

  9. Knowl 9 — Instruction-coupled kernel generally outperforms a soft-prompt-only kernel

    data/table

    The kernel ablation compares the proposed instruction-coupled kernel with a standard kernel that uses only structure in the soft-prompt space. Scores are the reported task performance; higher is better. The instruction-coupled version is better on 9 of the 10 listed tasks, while the standard kernel performs better on Unscrambling.

    Task Instruction-coupled kernel Standard kernel
    Sentiment 0.93 0.83
    Negation 0.80 0.39
    Larger animal 0.91 0.81
    Second Letter 0.62 0.33
    Formality 0.63 0.44
    Debugging 0.50 0.25
    Unscrambling 0.58 0.67
    Odd one out 0.92 0.90
    Ascii 0.33 0.13
    CS algorithm 0.38 0.26

    The consistent gains on most tasks support the value of incorporating similarities between instructions' predictions, but the Unscrambling result shows that this kernel is not uniformly superior on every task.

  10. Knowl 10 — Additional model pairings and human-instruction comparisons

    empirical result

    Supplementary evaluations test Vicuna-13B and WizardLM-13B as open-source generators with ChatGPT or GPT-4 as the API model, and compare selected results with human-written instructions. The scores below are reported accuracies. Across the Second Letter and Cause Selection tasks, the best automated result reaches 1.00; performance varies by model pairing and task, so these results demonstrate transfer across combinations rather than uniform gains for every pairing.

    Task Generator and API model Best-instruction accuracy Human-instruction accuracy
    Second Letter Vicuna-13B + ChatGPT 0.62 0.88
    Second Letter Vicuna-13B + GPT-4 0.99 0.96
    Second Letter WizardLM-13B + ChatGPT 0.99 0.88
    Second Letter WizardLM-13B + GPT-4 1.00 0.96
    Cause Selection Vicuna-13B + ChatGPT 0.86 0.52
    Cause Selection Vicuna-13B + GPT-4 1.00 0.72
    Cause Selection WizardLM-13B + ChatGPT 0.58 0.52
    Cause Selection WizardLM-13B + GPT-4 0.76 0.72

    The human comparison uses the corresponding human-instruction scores reported for ChatGPT or GPT-4. In a separate four-task comparison with the primary INSTRUCTZERO setting, the human versus INSTRUCTZERO scores were: Active to passive, 0.69 versus 1.00; Cause Selection, 0.52 versus 0.86; Taxonomy, 0.00 versus 0.82; and English-to-German translation, 0.74 versus 0.84. Thus, the reported INSTRUCTZERO instructions outperform the listed human instructions in those four cases, but the cross-pairing results also show that performance depends on the generator/API combination.

Coverage note — The optimization-trajectory visualizations and individual instruction catalog are omitted because they are illustrative rather than additional method or benchmark evidence; the supplementary GSM8K, AQUA, and SVAMP comparison with chain-of-thought prompts is also omitted as a secondary evaluation outside the 32-task instruction benchmark.

References

  1. 1.Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  2. 2.Chen, J., Chen, L., Huang, H., and Zhou, T. When do you need chain-of-thought prompting for chatgpt? arXiv preprint arXiv:2304.03262, 2023a.
  3. 3.Chen, J., Chen, L., and Zhou, T. It takes one to tango but more make trouble? in-context training with different number of demonstrations. arXiv preprint arXiv:2303.08119, 2023b.
  4. 4.Chen, L., Huang, H., and Cheng, M. Ptp: Boosting stability and performance of prompt tuning with perturbation-based regularizer. arXiv preprint arXiv:2305.02423, 2023c.
  5. 5.Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., and Hsieh, C.-J. Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. In Proceedings of the 10th ACM workshop on artificial intelligence and security, pp. 15–26, 2017.
  6. 6.Chen, X., Lin, M., Schärli, N., and Zhou, D. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128, 2023d.
  7. 7.Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  8. 8.Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022.
  9. 9.Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  10. 10.Deng, M., Wang, J., Hsieh, C.-P., Wang, Y., Guo, H., Shu, T., Song, M., Xing, E. P., and Hu, Z. Rlprompt: Optimizing discrete text prompts with reinforcement learning. arXiv preprint arXiv:2205.12548, 2022.
  11. 11.Deshwal, A. and Doppa, J. Combining latent space and structured kernels for bayesian optimization over combinatorial spaces. Advances in Neural Information Processing Systems, 34:8185–8200, 2021a.
  12. 12.Deshwal, A. and Doppa, J. Combining latent space and structured kernels for bayesian optimization over combinatorial spaces. Advances in Neural Information Processing Systems, 34:8185–8200, 2021b.
  13. 13.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Frazier, P. I. A tutorial on bayesian optimization. arXiv preprint arXiv:1807.02811, 2018.
  15. 15.Garcia, N., Ye, C., Liu, Z., Hu, Q., Otani, M., Chu, C., Nakashima, Y., and Mitamura, T. A dataset and baselines for visual question answering on art. In Proceedings of the European Conference in Computer Vision Workshops, 2020.
  16. 16.Gómez-Bombarelli, R., Wei, J. N., Duvenaud, D., Hernández-Lobato, J. M., Sánchez-Lengeling, B., Sheberla, D., Aguilera-Iparraguirre, J., Hirzel, T. D., Adams, R. P., and Aspuru-Guzik, A. Automatic chemical design using a data-driven continuous representation of molecules. ACS central science, 4(2):268–276, 2018.
  17. 17.Google. Palm-2-llm. https://blog.google/technology/ai/google-palm-2-ai-large-language-model/, 2023.
  18. 18.Hansen, N. The CMA evolution strategy: A tutorial. CoRR, abs/1604.00772, 2016. URL http://arxiv.org/abs/1604.00772.
  19. 19.Honovich, O., Shaham, U., Bowman, S. R., and Levy, O. Instruction induction: From few examples to natural language task descriptions. arXiv preprint arXiv:2205.10782, 2022.
  20. 20.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  21. 21.Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J. Large language models can self-improve. arXiv preprint arXiv:2210.11610, 2022.
  22. 22.Iyer, S., Lin, X. V., Pasunuru, R., Mihaylov, T., Simig, D., Yu, P., Shuster, K., Wang, T., Liu, Q., Koura, P. S., et al. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017, 2022.
  23. 23.Jin, W., Barzilay, R., and Jaakkola, T. Junction tree variational autoencoder for molecular graph generation. In International conference on machine learning, pp. 2323–2332. PMLR, 2018.
  24. 24.Kajino, H. Molecular hypergraph grammar with its application to molecular optimization. In International Conference on Machine Learning, pp. 3183–3191. PMLR, 2019.
  25. 25.Kleinberg, J. M. Two algorithms for nearest-neighbor search in high dimensions. In Proceedings of the twenty-ninth annual ACM symposium on Theory of computing, pp. 599–608, 1997.
  26. 26.Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  27. 27.Lester, B., Al-Rfou, R., and Constant, N. The power of scale for parameter-efficient prompt tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 3045–3059, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.243. URL https://aclanthology.org/2021.emnlp-main.243.
  28. 28.Li, X. L. and Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. In ACL 2021, pp. 4582–4597. Association for Computational Linguistics, 2021. URL https://doi.org/10.18653/v1/2021.acl-long.353.
  29. 29.Lin, X., Wu, Z., Dai, Z., Hu, W., Shu, Y., Ng, S.-K., Jaillet, P., and Low, B. K. H. Use your instinct: Instruction optimization using neural bandits coupled with transformers, 2023.
  30. 30.Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Computing Surveys, 55(9):1–35, 2023.
  31. 31.Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J. GPT understands, too. CoRR, abs/2103.10385, 2021. URL https://arxiv.org/abs/2103.10385.
  32. 32.Lu, X., Gonzalez, J., Dai, Z., and Lawrence, N. D. Structured variationally auto-encoded optimization. In International conference on machine learning, pp. 3267–3275. PMLR, 2018.
  33. 33.OpenAI. Sharegpt. https://sharegpt.com, 2023.
  34. 34.OpenAI. Chatgpt. https://openai.com/blog/chatgpt, 2023a.
  35. 35.OpenAI. Gpt-4 technical report. arXiv, 2023b.
  36. 36.Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  37. 37.Patel, A., Bhattamishra, S., and Goyal, N. Are NLP models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2080–2094, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.168. URL https://aclanthology.org/2021.naacl-main.168.
  38. 38.Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207, 2021.
  39. 39.Schrijver, A. et al. Combinatorial optimization: polyhedra and efficiency, volume 24. Springer, 2003.
  40. 40.Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. arXiv preprint arXiv:2010.15980, 2020.
  41. 41.Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023.
  42. 42.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  43. 43.Wang, Y., Du, S., Balakrishnan, S., and Singh, A. Stochastic zeroth-order optimization in high dimensions. In International conference on artificial intelligence and statistics, pp. 1356–1365. PMLR, 2018.
  44. 44.Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., et al. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv preprint arXiv:2204.07705, 2022.
  45. 45.Wang, Z., Hutter, F., Zoghi, M., Matheson, D., and De Feitas, N. Bayesian optimization in a billion dimensions via random embeddings. Journal of Artificial Intelligence Research, 55:361–387, 2016.
  46. 46.Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  47. 47.Wolsey, L. A. and Nemhauser, G. L. Integer and combinatorial optimization, volume 55. John Wiley & Sons, 1999.
  48. 48.Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244, 2023.
  49. 49.Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. Large language models are human-level prompt engineers. Arxiv, 2022.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/