Skill-Based Few-Shot Selection for In-Context Learning

Shengnan AnBo ZhouZeqi LinQiang FuBei ChenNanning ZhengWeizhu ChenJian-Guang Lou

article2023EMNLP51 citations

Proposes SKILL-KNN, a training-free few-shot selection method that prompts large language models to convert inputs into skill-based descriptions before retrieval, preventing embedding models from being misled by irrelevant surface features and improving semantic parsing performance.

Listen

Adapting large language models to complex tasks relies heavily on providing a few relevant demonstration examples in the prompt. Standard selection approaches choose examples by matching the surface-level text of raw user queries. However, this raw matching frequently fails in structured domains like database querying because it focuses on superficial keywords and entities rather than the underlying operations and logic needed to solve the problem. While fine-tuning specialized retrieval models offers an alternative, it requires substantial computational training overhead and lacks flexibility when example pools change frequently.

The article demonstrates and evaluates SKILL-KNN, a training-free few-shot selection framework. The objective of SKILL-KNN is to retrieve demonstration examples based on the intrinsic task-specific operations or skills required by a query, rather than relying on surface phrasing or requiring dedicated model fine-tuning.

The framework implements a two-stage "rewrite-then-retrieve" strategy. In the rewriting phase, a lightweight, frozen language model receives the user query alongside a small set of manually annotated examples (such as 12 to 16 demonstrations) to convert the raw input into a natural language description of required operations. In the retrieval phase, standard off-the-shelf embedding models match these generated skill descriptions against a bank of candidate examples. To eliminate sensitivity to prompt ordering, the authors developed two aggregation variants: a consistency-based variant that averages multiple candidate embeddings, and a distinctiveness-based variant that selects the most informative candidate representation. The authors evaluated the approach across five complex semantic parsing benchmarks (including Spider, Dr. Spider, KaggleDBQA, BIRD, and COGS), one math reasoning benchmark (GSM8K), and six major language model backbones.

The experimental findings show that SKILL-KNN consistently outperforms standard raw-input selection methods across all tested language models and benchmarks, securing the top performance among non-oracle methods. For example, on the Spider database benchmark with text-chat-davinci-002, SKILL-KNN improved execution accuracy to 78.3%, outperforming the best raw retrieval baseline (74.6%) and matching oracle methods that access ground-truth outputs (78.6%). Furthermore, the method matched or exceeded the accuracy of specialized fine-tuning selection models while remaining completely training-free. SKILL-KNN also showed superior robustness against perturbed queries and database structures on the Dr. Spider benchmark. Finally, ablation analyses showed that the rewriting system generalized effectively even when the initial demonstration pool was reduced to only four examples or constrained to limited database types.

These findings indicate that optimizing the descriptive input fed into standard embedding tools is significantly more efficient than fine-tuning underlying retrieval architectures. Organizations can dramatically improve generative accuracy and robustness in structured querying and reasoning tasks while eliminating the cost, timeline delays, and maintenance risks associated with retraining custom retrieval models on dynamic databases.

For practical deployment, organizations implementing few-shot retrieval pipelines should adopt prompt-based skill rewriting rather than investing in custom embedding fine-tuning. Teams should select the distinctiveness variant when downstream models prefer simpler, highly targeted examples, or the consistency variant when models handle higher structural complexity well. Future engineering efforts should explore hybrid aggregation mechanisms that combine both consistency and distinctiveness within a single selection step.

The primary operational limitation is the computational cost of performing multiple generation calls per query, which required substantial graphics processing unit time across large evaluation sets. Additionally, the approach provides the greatest benefit in structured, skill-intensive tasks, meaning performance advantages may diminish in simpler classification settings where surface text similarity is sufficient.

arXiv: 2305.14210

No sufficiently relevant recommendations were found.

Cover for Skill-Based Few-Shot Selection for In-Context Learning

Abstract

In-context learning is the paradigm that adapts large language models to downstream tasks by providing a few examples. Few-shot selection—selecting appropriate examples for each test instance separately—is important for in-context learning. In this paper, we propose Skill-KNN, a skill-based few-shot selection method for in-context learning. The key advantages of Skill-KNN include: (1) it addresses the problem that existing methods based on pre-trained embeddings can be easily biased by surface natural language features that are not important for the target task; (2) it does not require training or fine-tuning of any models, making it suitable for frequently expanding or changing example banks. The key insight is to optimize the inputs fed into the embedding model, rather than tuning the model itself. Technically, Skill-KNN generates the skill-based descriptions for each test case and candidate example by utilizing a pre-processing few-shot prompting, thus eliminating unimportant surface features. Experimental results across five cross-domain semantic parsing datasets and six backbone models show that Skill-KNN significantly outperforms existing methods.

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 In-Context Learning with Few-Shot Selection
  • 2.2 Embedding-Based Few-Shot Selection
  • 3 SKILL-KNN
  • 3.1 Generating Skill-Based Descriptions
  • 3.2 Variants
  • 4 Experimental Setup
  • 4.1 Tasks
  • 4.2 Selection Methods
  • 4.3 Backbones and Hyper-parameters
  • 5 Main Results
  • 6 Analysis
  • 7 Related Work
  • 8 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgments
  • References
  • A More Experimental Results
  • A.1 Recall@ N Performance of SKILL-KNN
  • A.2 Comparison with More Baseline Methods
  • B More Analysis
  • B.1 Motivation Behind Two Variants of SKILL-KNN
  • B.2 The Choice of Hyper-Parameter m
  • B.3 Measuring Diversity of SKILL-KNN and Oracle methods
  • B.4 Why do skill-based descriptions perform better?
  • B.5 What factors cause the different performances of two variants?
  • C Case Study
  • D Detailed Settings of Experiments
  • D.1 Select Examples for Annotation
  • D.2 Inference Hyper-Parameters
  • D.3 Input-Output Formats
  • D.4 Evaluation on BIRD
  • D.5 Evaluation on COGS
  • D.6 Target Sketch Matching for SQL
  • D.7 T-SNE Visualization
  • E Annotated Demonstrations

Knowls

  1. Knowl 1 — SKILL-KNN rewrites inputs into task-skill descriptions before retrieval

    model/method

    SKILL-KNN is a training-free, rewrite-then-retrieve method for choosing in-context examples. Given a test input and an example bank of input-output pairs, a frozen large language model (LLM) uses a small set of human-annotated input-to-skill demonstrations to rewrite the test input and each bank input as natural-language descriptions of the operations or structures needed to solve them. An off-the-shelf embedding model encodes those skill descriptions; SKILL-KNN retrieves the kk bank examples with the highest cosine similarity to the test description. It then places the original input-output examples—not the skill descriptions—in the prompt for solving the test case, with more similar examples nearer the test input. The method therefore changes the text used for retrieval rather than training or fine-tuning the embedding model.

  2. Knowl 2 — Consistency and distinctiveness aggregate multiple generated skill descriptions differently

    model/method

    Because generated descriptions can depend on the order of the annotated demonstrations, SKILL-KNN has two variants. For an input xix_i, let Si={si1,…,sim}S_i=\{s_i^1,\ldots,s_i^m\} be the mm descriptions produced by reordering those demonstrations, and let v(s)v(s) be the embedding vector of description ss. For the consistency variant, define the mean embedding ei=1m∑j=1mv(sij)e_i=\frac{1}{m}\sum_{j=1}^{m}v(s_i^j). The similarity between a test input xtx_t and a bank input xix_i is the cosine similarity of their mean embeddings, cos⁡(et,ei)\operatorname{cos}(e_t,e_i). For the distinctiveness variant, the similarity is the largest pairwise cosine similarity across the two candidate sets: max⁡1≤j,k≤mcos⁡(v(stj),v(sik))\max_{1\leq j,k\leq m}\operatorname{cos}(v(s_t^j),v(s_i^k)). Both variants rank examples by their resulting similarity scores. The first aggregates the candidates around their central embedding; the second lets the closest, most distinctive pair determine the match.

  3. Knowl 3 — Evaluation covers cross-domain semantic parsing and six frozen LLM backbones

    experimental setup

    The evaluation uses Spider, Dr. Spider, KaggleDBQA, BIRD, and COGS. Spider's training examples serve as the example bank for Spider development, Dr. Spider, and KaggleDBQA; BIRD and COGS use their own training examples, with 2,000 COGS training examples sampled for the bank. The text-to-SQL evaluations use execution-with-values accuracy, and COGS uses exact-match accuracy on primitive-substitution and primitive-structural-alternation subtasks. The six frozen task-solving backbones are text-chat-davinci-002, code-davinci-002, text-davinci-003, code-cushman-002, gpt-35-turbo, and gpt-4. Skill descriptions are generated with gpt-3.5-turbo. Experiments select k=4k=4 examples; decoding uses temperature 0 and a maximum length of 200 tokens. The default annotation set has 16 demonstrations, except that BIRD uses 12 demonstrations with evidence.

  4. Knowl 4 — SKILL-KNN improves over raw-input retrieval across the evaluated tasks

    empirical result

    Across the evaluated backbones and tasks, the authors report that SKILL-KNN is the strongest non-oracle selection approach overall, including cases where SKILL-KNN with SBERT exceeds retrieval using OpenAI embedding models. On Spider with text-chat-davinci-002, the distinctiveness variant with SBERT reaches 78.3% execution accuracy, compared with 74.6% for the strongest raw-input baseline, MMR with OpenAI Ada; the oracle Target-KNN reaches 78.6%. On Spider with code-davinci-002, SKILL-KNN with SBERT reaches 77.4%, versus 74.8% for the strongest raw-input baseline, MMR with OpenAI Ada. The results span five semantic-parsing datasets, including perturbation tests and compositional generalization subtasks, rather than relying only on in-domain retrieval.

  5. Knowl 5 — SKILL-KNN is competitive with trained selectors and effective with newer backbones

    empirical result

    On Spider development, SKILL-KNN is competitive with methods that fine-tune a retrieval model, although it does not outperform every such method in every setting. With text-chat-davinci-002, SKILL-KNN with SBERT scores 76.8% execution accuracy, compared with 74.4% for EPR, 75.0% for CEIL, and 76.3% for TST. With text-davinci-003, the best reported SKILL-KNN variant scores 76.6%, compared with 69.7% for EPR; with code-cushman-002, EPR scores 74.6% and SKILL-KNN's consistency variant scores 74.7%. On the same Spider evaluation, gpt-4 with the consistency variant reaches 82.7%, compared with 76.7% for SBERT KNN and 76.1% for random selection; gpt-35-turbo with the distinctiveness variant reaches 76.8%, compared with 73.7% for SBERT KNN and 74.3% for random selection. In the paper's additional GSM8K math-reasoning evaluation using text-chat-davinci-002, accuracy is 71.0% for SKILL-KNN, 69.9% for SBERT KNN, and 69.1% for random selection.

  6. Knowl 6 — Skill-based retrieval is more robust to diagnostic perturbations

    empirical result

    On Dr. Spider, the authors interpret SKILL-KNN's results as evidence of greater robustness to database, natural-language-question, and SQL perturbations than raw-input retrieval. With text-chat-davinci-002, SBERT KNN scores below random selection on two of the three perturbation categories, whereas each of the three SKILL-KNN versions scores above random selection on all three. The accompanying embedding-space visualization compares Spider development inputs with bank examples: raw-input embeddings appear unevenly matched, while skill-description embeddings place the two sets in a more similar distribution and appear more clustered. The authors suggest that the clearer neighborhoods in the skill-description space may help retrieval remain robust to perturbations; the visualization is qualitative evidence, not a causal test.

  7. Knowl 7 — SKILL-KNN retains an advantage with fewer or constrained annotation examples

    empirical result

    In Spider experiments with the base SKILL-KNN method and SBERT, increasing the number of annotated demonstrations from 4 toward the default of 16 yields only marginal improvements, and the method remains better than raw-input KNN with just four demonstrations. The authors also constrain which examples can be annotated: one condition removes the procedure for ensuring coverage of SQL operations, and another draws annotations from only two databases. Both constraints cause only a minor performance decline, while SKILL-KNN retains a substantial advantage over raw-input selection. A manual check of 100 generated skills finds 86 exactly correct descriptions, 12 that are mostly correct but need partial edits, and 2 that are entirely wrong; the authors also observe descriptions of skills absent from the annotated examples.

  8. Knowl 8 — The two variants trade off resilience to diffuse noise and outliers

    empirical result

    An analysis of the variants treats prompt-order effects as noise in generated skill embeddings. In a synthetic selection test, consistency achieves 95.4% selection accuracy under zero-mean white noise, versus 90.4% for distinctiveness; under spike noise, distinctiveness achieves 87.9%, versus 81.9% for consistency. In a human review of 23 cases where the variants differed, examples won by consistency had an average of 0.4 spike-noise occurrences per candidate set, while examples won by distinctiveness had 1.6. The authors also find that the consistency variant selects more complex text-to-SQL examples than the distinctiveness variant, measured by average database-table count and SQL-query length. They report different backbone preferences: text-chat-davinci-002 and code-davinci-002 tend to favor distinctiveness, while text-davinci-003 and code-cushman-002 tend to favor consistency.

  9. Knowl 9 — Retrieved-example diversity can accompany performance competitive with oracle selection

    empirical result

    SKILL-KNN sometimes exceeds an oracle selector that can use target outputs: the reported cases include the Dr. Spider database-perturbation subtask and two COGS subtasks. The authors propose, as a possible explanation rather than a demonstrated cause, that the selected examples are more diverse. For four retrieved examples, the mean number of distinct databases is 2.83 for SKILL-KNN with distinctiveness, versus 2.18 for Target-KNN, on Spider; 2.84 versus 2.16 on Dr. Spider; and 2.88 versus 2.21 on KaggleDBQA. The comparison supports the hypothesis that example diversity may aid cross-domain generalization even when a selector's similarity to the test case is not maximal.

  10. Knowl 10 — The method's demonstrated scope and computational cost are limited

    limitation

    The evaluation focuses mainly on cross-domain semantic parsing, with one additional math-reasoning test; the authors expect the advantage of skill-based selection may diminish for tasks where surface-feature similarity is sufficient. The two proposed variants handle consistency and distinctiveness separately, and the authors leave a method combining both as future work. The experiments require substantial computation: they were run on a station with eight NVIDIA A100 GPUs, and generating results for 10,000 examples takes about two times eight GPU-hours; reproducing the main results is estimated to require 400–500 times eight GPU-hours.

Coverage note — No substantial contributed finding was omitted; the detailed per-example annotation inventories and auxiliary prompt-format examples were left out because they support implementation rather than add independent results.

References

  1. 1.Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen, Nanning Zheng, Jian-Guang Lou, and Dongmei Zhang. 2023. How do in-context examples affect compositional generalization?
  2. 2.Patrick Bareiß, Beatriz Souza, Marcelo d’Amorim, and Michael Pradel. 2022. Code generation tools (almost) for free? a study of few-shot, pre-trained language models on code. arXiv preprint arXiv:2206.01335.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared J Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  4. 4.Shuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang, Jiarong Jiang, Joseph Lilien, Steve Ash, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, and Bing Xiang. 2023. Dr.spider: A diagnostic evaluation benchmark towards text-to-SQL robustness. In The Eleventh International Conference on Learning Representations.
  5. 5.Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, and Weizhu Chen. 2023a. Codet: Code generation with generated tests. In The Eleventh International Conference on Learning Representations.
  6. 6.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  7. 7.Yanda Chen, Chen Zhao, Zhou Yu, Kathleen McKeown, and He He. 2023b. On the relation between sensitivity and accuracy in in-context learning.
  8. 8.Zhoujun Cheng, Tianbao Xie, Peng Shi, Chengzu Li, Rahul Nadkarni, Yushi Hu, Caiming Xiong, Dragomir Radev, Mari Ostendorf, Luke Zettlemoyer, Noah A. Smith, and Tao Yu. 2023. Binding language models in symbolic languages. In The Eleventh International Conference on Learning Representations.
  9. 9.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  10. 10.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. ArXiv, abs/2110.14168.
  11. 11.Antonia Creswell, Murray Shanahan, and Irina Higgins. 2023. Selection-inference: Exploiting large language models for interpretable logical reasoning. In The Eleventh International Conference on Learning Representations.
  12. 12.Li Dong and Mirella Lapata. 2016. Language to logical form with neural attention. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 33–43.
  13. 13.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830.
  14. 14.Jiaqi Guo, Zecheng Zhan, Yan Gao, Yan Xiao, Jian-Guang Lou, Ting Liu, and Dongmei Zhang. 2019. Towards complex text-to-sql in cross-domain database with intermediate representation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4524–4535.
  15. 15.Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. In International Conference on Learning Representations.
  16. 16.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  17. 17.Yushi Hu, Chia-Hsuan Lee, Tianbao Xie, Tao Yu, Noah A Smith, and Mari Ostendorf. 2022. In-context learning for few-shot dialogue state tracking. arXiv preprint arXiv:2203.08568.
  18. 18.Rohit J Kate, Yuk Wah Wong, Raymond J Mooney, et al. 2005. Learning to transform natural to formal languages. In AAAI, volume 5, pages 1062–1068.
  19. 19.Najoung Kim and Tal Linzen. 2020. Cogs: A compositional generalization challenge based on semantic interpretation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 9087–9105.
  20. 20.Chia-Hsuan Lee, Oleksandr Polozov, and Matthew Richardson. 2021. Kaggledbqa: Realistic evaluation of text-to-sql parsers. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2261–2273.
  21. 21.Itay Levy, Ben Bogin, and Jonathan Berant. 2022. Diverse demonstrations improve in-context compositional generalization. arXiv preprint arXiv:2212.06800.
  22. 22.Haoyang Li, Jing Zhang, Cuiping Li, and Hong Chen. 2023a. Resdsql: Decoupling schema linking and skeleton parsing for text-to-sql.
  23. 23.Jia Li, Yunfei Zhao, Yongmin Li, Ge Li, and Zhi Jin. 2023b. Towards enhancing in-context learning for code generation. arXiv preprint arXiv:2303.17780.
  24. 24.Jinyang Li, Binyuan Hui, Ge Qu, Binhua Li, Jiaxi Yang, Bowen Li, Bailin Wang, Bowen Qin, Rongyu Cao, Ruiying Geng, Nan Huo, Xuanhe Zhou, Chenhao Ma, Guoliang Li, Kevin C. C. Chang, Fei Huang, Reynold Cheng, and Yongbin Li. 2023c. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls.
  25. 25.Xiaonan Li and Xipeng Qiu. 2023. Finding supporting examples for in-context learning. arXiv preprint arXiv:2302.13539.
  26. 26.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022a. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336.
  27. 27.Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, et al. 2022b. Competition-level code generation with alphacode. Science, 378(6624):1092–1097.
  28. 28.Xi Victoria Lin, Richard Socher, and Caiming Xiong. 2020. Bridging textual and tabular data for cross-domain text-to-sql semantic parsing. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4870–4888.
  29. 29.Aiwei Liu, Xuming Hu, Lijie Wen, and Philip S. Yu. 2023. A comprehensive evaluation of chatgpt’s zero-shot text-to-sql capability.
  30. 30.Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pages 100–114.
  31. 31.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-play compositional reasoning with large language models. arXiv preprint arXiv:2304.09842.
  32. 32.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8086–8098.
  33. 33.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2022. Webgpt: Browser-assisted question-answering with human feedback.
  34. 34.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, Johannes Heidecke, Pranav Shyam, Boris Power, Tyna Eloundou Nekoul, Girish Sastry, Gretchen Krueger, David Schnurr, Felipe Petroski Such, Kenny Hsu, Madeleine Thompson, Tabarak Khan, Toki Sherbakov, Joanne Jang, Peter Welinder, and Lilian Weng. 2022. Text and code embeddings by contrastive pre-training.
  35. 35.Tai Nguyen and Eric Wong. 2023. In-context example selection with influences. arXiv preprint arXiv:2302.11042.
  36. 36.Roma Patel and Ellie Pavlick. 2021. Mapping language models to grounded conceptual spaces. In International Conference on Learning Representations.
  37. 37.Gabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. 2022. Synchromesh: Reliable code generation from pre-trained language models. In International Conference on Learning Representations.
  38. 38.Mohammadreza Pourreza and Davood Rafiei. 2023. Din-sql: Decomposed in-context learning of text-to-sql with self-correction.
  39. 39.Jiexing Qi, Jingyao Tang, Ziwei He, Xiangpeng Wan, Chenghu Zhou, Xinbing Wang, Quanshi Zhang, and Zhouhan Lin. 2022. Rasat: Integrating relational structures into pretrained seq2seq model for text-to-sql. arXiv preprint arXiv:2205.06983.
  40. 40.Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023. Is chatgpt a general-purpose natural language processing task solver? arXiv preprint arXiv:2302.06476.
  41. 41.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  42. 42.Nitarshan Rajkumar, Raymond Li, and Dzmitry Bahdanau. 2022. Evaluating the text-to-sql capabilities of large language models. arXiv preprint arXiv:2204.00498.
  43. 43.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992.
  44. 44.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2655–2671.
  45. 45.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. arXiv preprint arXiv:2302.04761.
  46. 46.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021. Picard: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9895–9901.
  47. 47.Ozan Sener and Silvio Savarese. 2018. Active learning for convolutional neural networks: A core-set approach.
  48. 48.Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023. Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface. arXiv preprint arXiv:2303.17580.
  49. 49.Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2023. Language models are multilingual chain-of-thought reasoners. In The Eleventh International Conference on Learning Representations.
  50. 50.Richard Shin, Christopher Lin, Sam Thomson, Charles Chen Jr, Subhro Roy, Emmanouil Antonios Platanios, Adam Pauls, Dan Klein, Jason Eisner, and Benjamin Van Durme. 2021. Constrained language models yield few-shot semantic parsers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7699–7715.
  51. 51.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  52. 52.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. 2022. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615.
  53. 53.Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-sne. Journal of machine learning research, 9(11).
  54. 54.Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. 2020. Rat-sql: Relation-aware schema encoding and linking for text-to-sql parsers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7567–7578.
  55. 55.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  56. 56.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171.
  57. 57.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022a. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  58. 58.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in language models. In Advances in Neural Information Processing Systems.
  59. 59.Zhiyong Wu, Yaoxiang Wang, Jiacheng Ye, and Lingpeng Kong. 2023. Self-adaptive in-context learning: An information compression perspective for in-context example selection and ordering.
  60. 60.Xiaojun Xu, Chang Liu, and Dawn Song. 2017. Sqlnet: Generating structured queries from natural language without reinforcement learning.
  61. 61.Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. 2023. Compositional exemplars for in-context learning. arXiv preprint arXiv:2302.05698.
  62. 62.Xi Ye and Greg Durrett. 2023. Explanation selection using unlabeled data for in-context learning. arXiv preprint arXiv:2302.04813.
  63. 63.Xi Ye, Srinivasan Iyer, Asli Celikyilmaz, Ves Stoyanov, Greg Durrett, and Ramakanth Pasunuru. 2022. Complementary explanations for effective in-context learning. arXiv preprint arXiv:2211.13892.
  64. 64.Tao Yu, Zifan Li, Zilin Zhang, Rui Zhang, and Dragomir Radev. 2018a. Typesql: Knowledge-based type-aware neural text-to-sql generation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 588–594.
  65. 65.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al. 2018b. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921.
  66. 66.John M Zelle and Raymond J Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the national conference on artificial intelligence, pages 1050–1055.
  67. 67.Luke S Zettlemoyer and Michael Collins. 2012. Learning to map sentences to logical form: Structured classification with probabilistic categorial grammars. arXiv preprint arXiv:1207.1420.
  68. 68.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022a. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  69. 69.Yiming Zhang, Shi Feng, and Chenhao Tan. 2022b. Active example selection for in-context learning. arXiv preprint arXiv:2211.04486.
  70. 70.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.
  71. 71.Victor Zhong, Mike Lewis, Sida I Wang, and Luke Zettlemoyer. 2020. Grounded adaptation for zero-shot executable semantic parsing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6869–6882.

Citation

MLA
An, S., et al. “Skill-Based Few-Shot Selection for In-Context Learning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 13472–92, https://doi.org/10.18653/v1/2023.emnlp-main.831.
APA
An, S., Zhou, B., Lin, Z., Fu, Q., Chen, B., Zheng, N., Chen, W., & Lou, J.-G. (2023). Skill-Based Few-Shot Selection for In-Context Learning. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13472–13492. https://doi.org/10.18653/v1/2023.emnlp-main.831
Chicago
An, S., B. Zhou, Z. Lin, et al. 2023. “Skill-Based Few-Shot Selection for In-Context Learning”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 13472–92. https://doi.org/10.18653/v1/2023.emnlp-main.831.
Harvard
An, S. et al. (2023) “Skill-Based Few-Shot Selection for In-Context Learning”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 13472–13492. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.831.
Vancouver
1. An S, Zhou B, Lin Z, Fu Q, Chen B, Zheng N, Chen W, Lou J-G (2023) Skill-Based Few-Shot Selection for In-Context Learning. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 13472–13492

BibTeX

@inproceedings{an-etal-2023-skill,
    title = "Skill-Based Few-Shot Selection for In-Context Learning",
    author = "An, Shengnan  and
      Zhou, Bo  and
      Lin, Zeqi  and
      Fu, Qiang  and
      Chen, Bei  and
      Zheng, Nanning  and
      Chen, Weizhu  and
      Lou, Jian-Guang",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.831/",
    doi = "10.18653/v1/2023.emnlp-main.831",
    pages = "13472--13492"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/