Compositional Exemplars for In-context Learning
Jiacheng YeZhiyong WuJiangtao FengTao YuLingpeng Kong
Proposes a determinantal point process framework that optimizes in-context exemplar selection by modeling both input relevance and inter-example diversity through contrastive learning, achieving state-of-the-art performance across 12 diverse NLP benchmarks.
Large language models can perform new tasks without updating their underlying parameters by learning directly from demonstration examples provided in their input context, a process known as in-context learning. However, the operational accuracy of this approach is highly volatile and depends heavily on which examples are selected. Traditional selection techniques evaluate examples individually based on simple heuristics or basic similarity, often introducing redundant demonstrations and failing to stay within strict prompt size limits.
The article introduces and evaluates CEIL (Compositional Exemplars for In-context Learning), a framework designed to select complete, high-performing subsets of demonstration examples. The main objective is to demonstrate that optimizing the joint selection of demonstration sets—balancing relevance to the query with diversity among the chosen examples—substantially improves task accuracy across various language models without altering model parameters.
To achieve this, the authors model set selection using conditional Determinantal Point Processes, a mathematical approach that captures interactions among examples to reward diversity while penalizing redundancy. The retrieval model employs contrastive learning aligned with feedback from the target language model, using a pair-wise ranking loss to distinguish effective subsets from less useful ones. During runtime, the optimal subset is chosen using an efficient greedy search algorithm over a narrowed candidate pool. The approach was evaluated across 12 benchmark datasets covering seven diverse natural language processing tasks, including sentiment analysis, question answering, and code generation, using language models ranging from 1.5 billion to 175 billion parameters.
Across the 12 evaluation benchmarks, CEIL achieved an average performance score of 56.76%, outperforming the previous state-of-the-art retriever by an absolute gain of 3.39 percentage points and surpassing learning-free baselines by more than 10 percentage points. The framework showed substantial improvements on complex natural language inference tasks, delivering gains of over 20 percentage points compared to learning-free methods. On compositional semantic parsing benchmarks, CEIL consistently outperformed competing retrievers by capturing complementary elements required for multi-step queries. Additionally, CEIL demonstrated extreme efficiency: when restricted to just four examples, it routinely surpassed the baseline models configured with 32 examples, significantly reducing input length.
These findings indicate that treating demonstration selection as an integrated subset problem yields far more effective prompts than ranking individual examples independently. Because CEIL selects more compact yet informative demonstration sets, organizations can significantly reduce computational overhead and latency, as processing shorter inputs requires less compute. Furthermore, the retriever trained on one model transfers successfully to others—including massive models like Codex—allowing organizations to boost performance on closed-source or expensive platforms without paying to retrain model-specific retrievers.
Organizations utilizing in-context learning should adopt set-level, diversity-aware retrieval strategies in place of individual similarity matching, particularly for multi-step reasoning, semantic parsing, and code generation. Teams should also leverage smaller demonstration sizes during deployment to lower operational inference costs. Looking ahead, technical teams should explore multi-task training to develop a single, generalized retriever capable of handling new domains without task-specific training data.
The primary limitation of this approach is the upfront computational cost required to train the retriever and generate candidate subset scores using language models. While the framework demonstrates solid transferability across several tasks and model architectures, cross-task generalization remains inconsistent between single-input and double-input formats, warranting localized validation before broad production deployment.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). KATE establishes query-conditioned nearest-neighbor retrieval of demonstrations, a key selection baseline that helps situate CEIL’s more structured subset-selection approach.
- Paper: MetaICL: Learning to Learn In Context, Sewon Min et al. (2022). MetaICL lays out how language models use demonstrations to infer unseen tasks, providing the ICL setting in which CEIL optimizes example selection.
- Paper: Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?, Sewon Min et al. (2022). This study tests which demonstration properties drive ICL performance, clarifying why selecting examples for their composition—not merely their labels—matters.
- Paper: Active Prompting with Chain-of-Thought for Large Language Models, Shizhe Diao et al. (2024). Active-Prompt carries example selection into reasoning by choosing informative questions for annotation, extending the selection problem to uncertainty-driven Chain-of-Thought demonstrations.
