Skill-Based Few-Shot Selection for In-Context Learning
Shengnan AnBo ZhouZeqi LinQiang FuBei ChenNanning ZhengWeizhu ChenJian-Guang Lou
Proposes SKILL-KNN, a training-free few-shot selection method that prompts large language models to convert inputs into skill-based descriptions before retrieval, preventing embedding models from being misled by irrelevant surface features and improving semantic parsing performance.
Adapting large language models to complex tasks relies heavily on providing a few relevant demonstration examples in the prompt. Standard selection approaches choose examples by matching the surface-level text of raw user queries. However, this raw matching frequently fails in structured domains like database querying because it focuses on superficial keywords and entities rather than the underlying operations and logic needed to solve the problem. While fine-tuning specialized retrieval models offers an alternative, it requires substantial computational training overhead and lacks flexibility when example pools change frequently.
The article demonstrates and evaluates SKILL-KNN, a training-free few-shot selection framework. The objective of SKILL-KNN is to retrieve demonstration examples based on the intrinsic task-specific operations or skills required by a query, rather than relying on surface phrasing or requiring dedicated model fine-tuning.
The framework implements a two-stage "rewrite-then-retrieve" strategy. In the rewriting phase, a lightweight, frozen language model receives the user query alongside a small set of manually annotated examples (such as 12 to 16 demonstrations) to convert the raw input into a natural language description of required operations. In the retrieval phase, standard off-the-shelf embedding models match these generated skill descriptions against a bank of candidate examples. To eliminate sensitivity to prompt ordering, the authors developed two aggregation variants: a consistency-based variant that averages multiple candidate embeddings, and a distinctiveness-based variant that selects the most informative candidate representation. The authors evaluated the approach across five complex semantic parsing benchmarks (including Spider, Dr. Spider, KaggleDBQA, BIRD, and COGS), one math reasoning benchmark (GSM8K), and six major language model backbones.
The experimental findings show that SKILL-KNN consistently outperforms standard raw-input selection methods across all tested language models and benchmarks, securing the top performance among non-oracle methods. For example, on the Spider database benchmark with text-chat-davinci-002, SKILL-KNN improved execution accuracy to 78.3%, outperforming the best raw retrieval baseline (74.6%) and matching oracle methods that access ground-truth outputs (78.6%). Furthermore, the method matched or exceeded the accuracy of specialized fine-tuning selection models while remaining completely training-free. SKILL-KNN also showed superior robustness against perturbed queries and database structures on the Dr. Spider benchmark. Finally, ablation analyses showed that the rewriting system generalized effectively even when the initial demonstration pool was reduced to only four examples or constrained to limited database types.
These findings indicate that optimizing the descriptive input fed into standard embedding tools is significantly more efficient than fine-tuning underlying retrieval architectures. Organizations can dramatically improve generative accuracy and robustness in structured querying and reasoning tasks while eliminating the cost, timeline delays, and maintenance risks associated with retraining custom retrieval models on dynamic databases.
For practical deployment, organizations implementing few-shot retrieval pipelines should adopt prompt-based skill rewriting rather than investing in custom embedding fine-tuning. Teams should select the distinctiveness variant when downstream models prefer simpler, highly targeted examples, or the consistency variant when models handle higher structural complexity well. Future engineering efforts should explore hybrid aggregation mechanisms that combine both consistency and distinctiveness within a single selection step.
The primary operational limitation is the computational cost of performing multiple generation calls per query, which required substantial graphics processing unit time across large evaluation sets. Additionally, the approach provides the greatest benefit in structured, skill-intensive tasks, meaning performance advantages may diminish in simpler classification settings where surface text similarity is sufficient.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). KATE establishes query-conditioned nearest-neighbor retrieval of in-context examples, the direct retrieval baseline that SKILL-KNN refines by matching task operations rather than surface wording.
- Paper: Compositional Exemplars for In-context Learning, Jiacheng Ye et al. (2023). CEIL introduces learned, diversity-aware selection of demonstration sets, providing essential context for how SKILL-KNN’s skill-based retrieval differs from earlier example-selection strategies.
- Paper: Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering, Zhiyong Wu et al. (2023). This work frames per-query example selection and ordering as central ICL design problems, clarifying the selection challenge SKILL-KNN addresses through skill descriptions.
No sufficiently relevant recommendations were found.
