GPS: Genetic Prompt Search for Efficient Few-Shot Learning

Hanwei XuYujun ChenYulun DuNan ShaoYanggang WangHaiyu LiZhilin Yang

article2022EMNLP51 citations

Proposes a gradient-free genetic algorithm that automatically discovers high-performing, fluent natural language prompts using only a tiny validation set, outperforming both manual prompt engineering and parameter-efficient tuning methods.

Listen

Deploying large language models across diverse business applications is often hindered by the high cost of manual data labeling and model fine-tuning. While prompting allows models to perform tasks with minimal data, manually written prompts are typically suboptimal and produce inconsistent results. The article introduces Genetic Prompt Search (GPS), an automated method that uses evolutionary algorithms to discover high-performing text prompts without updating the underlying model parameters.

The main objective of the article is to demonstrate that GPS can automatically optimize discrete prompts using only a tiny validation dataset, matching or exceeding the performance of parameter-tuning methods while reducing computational overhead. To evaluate this, the authors tested GPS across ten diverse natural language processing benchmark tasks using only 32 labeled examples per task. The approach begins with human-written seed prompts, mutates them using techniques like back-translation and generative sentence continuation, and iteratively selects top-performing candidates based on validation accuracy.

The findings show that GPS significantly improves task accuracy compared to existing prompting and parameter-efficient tuning methods. First, GPS achieved an average accuracy of 60.12%, outperforming manual prompts by 2.6 percentage points and beating parameter-efficient techniques like Prompt Tuning (58.56%) and Black-Box Tuning (57.82%). Second, GPS outperformed rule-based discrete prompt search methods such as GRIPS by 1.4 percentage points while generating semantically fluent, interpretable text. Third, sentence continuation powered by larger generative models proved to be the most effective mutation strategy. Finally, GPS required roughly one-tenth the computational operations of traditional tuning methods during search and training, all while keeping model weights frozen.

These results demonstrate that organizations can achieve superior few-shot task accuracy without the high storage, memory, and infrastructure costs required to maintain customized model checkpoints for each use case. Because GPS searches only for discrete text strings, a single frozen foundation model can serve multiple business applications simultaneously with zero added inference latency or task-specific parameter storage.

For enterprise systems operating under limited labeled data, adopting automated discrete prompt search offers a practical and cost-effective alternative to model fine-tuning. When implementing this approach, teams should prioritize generative sentence continuation over simple word-replacement rules. However, decision-makers should note that the evaluation was limited to 10 English benchmark tasks under a strict 32-sample regime. Where thousands of labeled samples are available, traditional fine-tuning may still provide higher peak accuracy, and further evaluation is warranted on specialized enterprise domains.

Cover for GPS: Genetic Prompt Search for Efficient Few-Shot Learning

Abstract

Prompt-based techniques have demonstrated great potential for improving the few-shot generalization of pretrained language models. However, their performance heavily relies on the manual design of prompts and thus requires a lot of human efforts. In this paper, we introduce Genetic Prompt Search (GPS) to improve few-shot learning with prompts, which utilizes a genetic algorithm to automatically search for high-performing prompts. GPS is gradient-free and requires no update of model parameters but only a small validation set. Experiments on diverse datasets proved the effectiveness of GPS, which outperforms manual prompts by a large margin of 2.6 points. Our method is also better than other parameter-efficient tuning methods such as prompt tuning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Few-Shot Tuning
  • 2.2 Prompting Enhancement
  • 3 Genetic Prompt Search
  • 3.1 Genetic Prompt Search Algorithm
  • 3.2 Prompt Generation Strategies
  • 4 Experiments
  • 4.1 Experimental Setups
  • 4.2 Datasets
  • 4.3 Baselines
  • 4.4 Implementation Details
  • 4.5 Main Results
  • 4.6 Ablation Study
  • 4.6.1 Prompt Generation Strategies
  • 4.6.2 The Size of Validation Set
  • 4.6.3 The Number of Prompt Search Iterations
  • 4.7 Case Study
  • 4.8 Overall Comparison of Different Few-Shot Methods
  • 5 Conclusions
  • 6 Limitations
  • References

Knowls

  1. Knowl 1 — GPS searches discrete prompts without tuning the language model

    model/method

    Genetic Prompt Search (GPS) improves few-shot task performance by searching over hard prompts: discrete text templates that are combined with task inputs. It keeps the pretrained language model frozen and does not update model parameters or optimize gradients. Instead, GPS uses a small labeled development set to score candidate prompts, retaining and generating prompts according to their scores. Its search therefore changes the prompt text, not the language model or a learned soft prompt.

  2. Knowl 2 — GPS iteratively selects and reproduces prompt candidates

    algorithm

    GPS takes handcrafted seed prompts, a labeled development set, a prompt-scoring function, a prompt-generation strategy, a selection size KK, and a search-iteration count TT. It returns a final group of high-scoring prompts. In the paper’s main experiments, KK was set to the number of initial prompts (5), the search ran for 6 iterations, top-pp sampling used p=0.9p=0.9, and duplicate prompts or prompts missing a required input placeholder were filtered out. A prompt-pool size of 30 was used in the computational-cost comparison.

    Input: Handcrafted prompt set G0, development set Ddev, score function fGPS, generation strategy gGPS, selection size K, iteration count T
    Output: Final optimized prompt set GT+1
    For each generation t from 0 through T:
        Store the current prompt set Gt
        Score every prompt in Gt using fGPS on Ddev
        Select the K highest-scoring prompts as the reproductive group Gt*
        Generate the next prompt set Gt+1 from Gt* using gGPS
        Remove generated prompts that duplicate existing prompts or lack a valid task-input placeholder
    Collect the top-K prompts from each stored generation
    Rescore the collected prompts and select the K highest-scoring prompts as GT+1
    Return GT+1

    The prompt-scoring function evaluates candidates on the development set; the generation strategy determines how new text prompts are produced. The paper does not specify a general computational-complexity bound for this procedure.

  3. Knowl 3 — GPS generates candidates by back translation, cloze completion, or sentence continuation

    model/method

    GPS investigates three ways to generate prompt candidates. Back translation (BT) translates an English prompt to Chinese, Japanese, Korean, French, Spanish, Italian, Russian, German, Arabic, Greek, and Cantonese, then translates it back to English. Cloze generation begins with a manual prompt, replaces randomly selected tokens with placeholders, and asks T5 to fill the blanks; the paper reports that directly applying an earlier automatic-template procedure did not work well in this no-parameter-update setting. Sentence continuation (SC) prompts a pretrained generator with “Write two sentences that mean the same thing. Sentence 1: [manual prompt], Sentence 2:” and uses its continuation as a candidate. The SC generators tested were GPT2-XL (1.5B parameters) and T5LM-XXL (11B parameters).

    Cloze candidates are scored by average logits on the development examples. BT and SC candidates are scored by development-set accuracy, since the paper does not use average logits for those strategies. These are prompt-selection scores, not test-set metrics.

  4. Knowl 4 — Evaluation uses ten held-out T0 tasks and a 32-example few-shot budget

    experimental setup

    GPS was evaluated on ten English tasks held out from T0’s prompted training tasks: ANLI R1, ANLI R2, ANLI R3, CB, RTE, WSC, Winogrande, COPA, HellaSwag, and WiC. The tasks cover natural-language inference, coreference resolution, sentence completion, and word-sense disambiguation. The few-shot setting used 32 labeled examples per task, divided across classes (for example, 8 examples per class for a four-class task). Experiments were repeated with three data splits, and reported few-shot results average performance across prompts and splits.

    For methods that do not tune parameters, the 32 examples formed the development set used to search for prompts. For parameter-tuning methods, the 32 examples were divided into 16 training and 16 validation examples. Seed prompts were taken from T0, and the same seed-prompt set was used for the compared baselines where applicable.

  5. Knowl 5 — GPS is the strongest parameter-frozen method in the main benchmark

    data/table

    The table reports accuracy (percent) on the ten held-out tasks. Few-shot results are means across the three data splits, with standard deviations across splits; the two zero-shot T0 columns have no reported standard deviation. GPS averages 60.12, improving on reproduced zero-shot T0 by 2.60 points and on GRIPS by 1.46 points. It is the strongest method among those that freeze all parameters, but full Model Tuning has the highest overall average at 61.73.

    TaskT0 zero-shot (original)T0 zero-shot (reproduced)BBTPTMTICLGRIPSGPS
    ANLI R143.5643.1642.97 ± 0.3243.21 ± 0.1546.73 ± 3.2537.87 ± 0.8044.41 ± 0.0144.06 ± 2.78
    ANLI R238.6838.6838.89 ± 0.0537.36 ± 1.6939.12 ± 3.1734.51 ± 1.5339.57 ± 0.2438.10 ± 1.68
    ANLI R341.2641.8741.32 ± 0.0540.85 ± 0.3642.20 ± 2.1135.98 ± 2.8042.96 ± 0.4241.51 ± 2.33
    CB70.1270.1272.06 ± 0.1471.31 ± 2.3983.97 ± 5.4059.21 ± 5.6576.55 ± 0.4180.12 ± 1.61
    RTE80.8380.9781.73 ± 0.4982.47 ± 0.8679.51 ± 2.0864.86 ± 10.2981.71 ± 0.1284.22 ± 1.02
    WSC61.4561.0660.74 ± 0.3163.30 ± 1.8264.65 ± 1.7961.03 ± 3.8661.47 ± 1.4463.62 ± 1.68
    Winogrande59.9459.7059.46 ± 0.3258.63 ± 0.7059.76 ± 1.4753.22 ± 1.5858.11 ± 0.2659.59 ± 2.06
    COPA90.0290.0290.51 ± 0.7492.33 ± 0.3992.54 ± 0.9882.82 ± 1.3991.75 ± 0.4293.50 ± 0.14
    HellaSwag33.5533.5233.47 ± 0.2037.28 ± 0.2949.75 ± 4.9827.35 ± 2.2833.09 ± 0.2038.85 ± 5.54
    WiC56.6856.1357.09 ± 0.4058.86 ± 1.2259.04 ± 1.0950.35 ± 0.8957.03 ± 0.6757.65 ± 1.18
    Average57.6057.5257.82 ± 0.0358.56 ± 0.2061.73 ± 0.0951.28 ± 1.6658.66 ± 0.3560.12 ± 1.40

    BBT is Black-Box Tuning, PT is Prompt Tuning, MT is Model Tuning, and ICL is In-Context Learning. The GPS average exceeds the parameter-frozen alternatives GRIPS (58.66), ICL (51.28), and reproduced zero-shot T0 (57.52), as well as the parameter-efficient tuning alternatives BBT (57.82) and PT (58.56); it does not exceed MT (61.73).

  6. Knowl 6 — Sentence continuation with T5LM yields the best prompt-generation results

    empirical result

    The generation-strategy comparison measures accuracy on the same ten tasks, using reproduced zero-shot T0 as a reference. Sentence continuation with the larger T5LM generator has the highest average (61.72), compared with back translation (60.65), sentence continuation with GPT2 (60.01), and cloze generation (57.65). Cloze is approximately at the zero-shot baseline overall, while T5LM sentence continuation is markedly better than GPT2 sentence continuation.

    TaskT0 zero-shot (reproduced)BTClozeSC (GPT2)SC (T5LM)
    ANLI R143.1644.9542.1044.6446.47
    ANLI R238.6840.1338.9539.1439.91
    ANLI R341.8742.8342.0042.7543.01
    CB70.1279.3873.2179.7180.00
    RTE80.9782.6082.6482.5383.86
    WSC61.0665.3864.3863.6565.48
    Winogrande59.7061.0953.5059.7661.96
    COPA90.0293.1289.7793.3193.43
    HellaSwag33.5236.3433.8336.7244.29
    WiC56.1360.7256.1557.9158.82
    Average57.5260.6557.6560.0161.72

    Qualitative examples show that successful candidates can remain fluent while changing task-relevant wording. For WSC, the T5LM-generated prompt removes the phrase “In the passage above” and asks whether the pronoun refers to “the person of” the candidate antecedent; its reported metric rises from 60.58 for the original prompt to 70.19. For HellaSwag, a T5LM continuation changes the question to ask what is “the most likely thing to happen next,” and the reported metric rises from 34.00 to 47.63.

  7. Knowl 7 — More validation examples generally improve prompt-search gains

    empirical result

    The paper varies the GPS development-set size from 8 to 128 examples. In general, prompt-search gains increase as more validation examples provide feedback for scoring candidates; WSC and HellaSwag follow this trend. GPS is reported to remain substantially better than manual prompts with only 8 validation examples. The experiments use 32 examples as the default to reflect a limited-label few-shot setting, despite the generally beneficial trend with larger validation sets.

  8. Knowl 8 — More search iterations improve aggregate performance, but some tasks peak early

    empirical result

    GPS was evaluated with up to 9 prompt-search iterations. Aggregate performance across datasets improves with additional iterations, whereas WSC and Winogrande reach their best results at an early iteration. The paper sets the default to 6 iterations as a trade-off between performance and search cost; the zero-iteration reference is reproduced zero-shot T0.

  9. Knowl 9 — GPS combines low deployment overhead with the second-best reported average

    data/table

    This comparison considers serving efficiency, tunable parameters, accuracy, and computational cost. GPS requires no tunable parameters and has the lowest nonzero reported search/training cost among the listed methods, while its 60.12 accuracy is second only to full Model Tuning (61.73). The reported cost is normalized to GPS at 1.0×; In-Context Learning is assigned 0× for training or prompt-search computation, although its long inference sequences are described as costly. A check mark indicates serving efficiency in the paper’s comparison.

    MethodServing efficiencyTunable parametersPerformance (accuracy)Computation cost
    Model TuningNo100%61.7311.1×
    Prompt TuningYes~0.01%58.5611.1×
    Black-Box TuningYes~0.001%57.829.3×
    In-Context LearningNo*0%51.280×
    GPSYes0%60.121.0×

    The absolute estimated costs were 48,000 equivalent forward passes for Model Tuning and Prompt Tuning, 40,000 for Black-Box Tuning, and 4,320 for GPS. The GPS estimate assumes 6 search iterations, a prompt pool of 30, and two equivalent forward passes to generate one prompt. The paper’s computation-cost measure combines training and prompt-search costs; it does not include the inference-sequence overhead of In-Context Learning.

  10. Knowl 10 — Evidence is limited to a small benchmark and does not explain prompt effectiveness

    limitation

    The paper identifies three limits on the scope of its findings: evaluation covers only the ten T0 test tasks, so performance on other datasets is unknown; comparisons use only 32 labeled examples per task, and methods that tune parameters could perform better and more stably with hundreds or thousands of examples; and the study does not explain why the automatically selected prompts work better. The reported results therefore do not establish how GPS compares with tuning methods at larger data budgets or what prompt properties cause the observed gains.

Coverage note — No substantial contributed material was omitted; the prompt examples are incorporated into the generation-strategy findings, and the paper’s stated limitations are included.

References

  1. 1.Elad Ben Zaken, Yoav Goldberg, and Shauli Ravfogel. 2022. BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1–9, Dublin, Ireland. Association for Computational Linguistics.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  3. 3.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  4. 4.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021a. A framework for few-shot language model evaluation.
  5. 5.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021b. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  6. 6.Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021. Ptr: Prompt tuning with rules for text classification.
  7. 7.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 2790–2799. PMLR.
  8. 8.Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models.
  9. 9.Teven Le Scao and Alexander Rush. 2021. How many data points is a prompt worth? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2627–2636, Online. Association for Computational Linguistics.
  10. 10.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning.
  11. 11.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021a. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  12. 12.Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021b. Gpt understands, too.
  13. 13.Swaroop Mishra, Daniel Khashabi, Chitta Baral, Yejin Choi, and Hannaneh Hajishirzi. 2021. Reframing instructional prompts to gptk’s language.
  14. 14.T. M. Mitchell. 1980. The need for biases in learning generalizations. Technical report, Computer Science Department, Rutgers University, New Brunswick, MA.
  15. 15.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback.
  16. 16.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 11054–11070.
  17. 17.Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal. 2022. Grips: Gradient-free, edit-based instruction search for prompting large language models. arXiv preprint arXiv:2203.07281.
  18. 18.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018. Language models are unsupervised multitask learners.
  19. 19.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.
  20. 20.Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Stella Biderman, Leo Gao, Tali Bers, Thomas Wolf, and Alexander M. Rush. 2021. Multitask prompted training enables zero-shot task generalization.
  21. 21.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online. Association for Computational Linguistics.
  22. 22.Timo Schick and Hinrich Schütze. 2021. Generating datasets with pretrained language models. Computing Research Repository, arXiv:2104.07540.
  23. 23.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  24. 24.Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu. 2022a. Black-box tuning for language-model-as-a-service.
  25. 25.Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu. 2022b. Black-box tuning for language-model-as-a-service. arXiv preprint arXiv:2201.03514.
  26. 26.Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. 2022. What language model architecture and pretraining objective work best for zero-shot generalization? arXiv preprint arXiv:2204.05832.
  27. 27.Albert Webson and Ellie Pavlick. 2021. Do prompt-based models really understand the meaning of their prompts? arXiv preprint arXiv:2109.01247.
  28. 28.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2021. Finetuned language models are zero-shot learners.
  29. 29.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan R Salakhutdinov, and Quoc V Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems, 32.

Citation

MLA
Xu, H., et al. “GPS: Genetic Prompt Search for Efficient Few-Shot Learning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 8162–71, https://doi.org/10.18653/v1/2022.emnlp-main.559.
APA
Xu, H., Chen, Y., Du, Y., Shao, N., Yanggang, W., Li, H., & Yang, Z. (2022). GPS: Genetic Prompt Search for Efficient Few-Shot Learning. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8162–8171. https://doi.org/10.18653/v1/2022.emnlp-main.559
Chicago
Xu, H., Y. Chen, Y. Du, et al. 2022. “GPS: Genetic Prompt Search for Efficient Few-Shot Learning”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 8162–71. https://doi.org/10.18653/v1/2022.emnlp-main.559.
Harvard
Xu, H. et al. (2022) “GPS: Genetic Prompt Search for Efficient Few-Shot Learning”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 8162–8171. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.559.
Vancouver
1. Xu H, Chen Y, Du Y, Shao N, Yanggang W, Li H, Yang Z (2022) GPS: Genetic Prompt Search for Efficient Few-Shot Learning. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 8162–8171

BibTeX

@inproceedings{xu-etal-2022-gps,
    title = "{GPS}: Genetic Prompt Search for Efficient Few-Shot Learning",
    author = "Xu, Hanwei  and
      Chen, Yujun  and
      Du, Yulun  and
      Shao, Nan  and
      Yanggang, Wang  and
      Li, Haiyu  and
      Yang, Zhilin",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.559/",
    doi = "10.18653/v1/2022.emnlp-main.559",
    pages = "8162--8171"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/