Compositional Exemplars for In-context Learning

Jiacheng YeZhiyong WuJiangtao FengTao YuLingpeng Kong

article2023ICML216 citations

Proposes a determinantal point process framework that optimizes in-context exemplar selection by modeling both input relevance and inter-example diversity through contrastive learning, achieving state-of-the-art performance across 12 diverse NLP benchmarks.

Listen

Large language models can perform new tasks without updating their underlying parameters by learning directly from demonstration examples provided in their input context, a process known as in-context learning. However, the operational accuracy of this approach is highly volatile and depends heavily on which examples are selected. Traditional selection techniques evaluate examples individually based on simple heuristics or basic similarity, often introducing redundant demonstrations and failing to stay within strict prompt size limits.

The article introduces and evaluates CEIL (Compositional Exemplars for In-context Learning), a framework designed to select complete, high-performing subsets of demonstration examples. The main objective is to demonstrate that optimizing the joint selection of demonstration sets—balancing relevance to the query with diversity among the chosen examples—substantially improves task accuracy across various language models without altering model parameters.

To achieve this, the authors model set selection using conditional Determinantal Point Processes, a mathematical approach that captures interactions among examples to reward diversity while penalizing redundancy. The retrieval model employs contrastive learning aligned with feedback from the target language model, using a pair-wise ranking loss to distinguish effective subsets from less useful ones. During runtime, the optimal subset is chosen using an efficient greedy search algorithm over a narrowed candidate pool. The approach was evaluated across 12 benchmark datasets covering seven diverse natural language processing tasks, including sentiment analysis, question answering, and code generation, using language models ranging from 1.5 billion to 175 billion parameters.

Across the 12 evaluation benchmarks, CEIL achieved an average performance score of 56.76%, outperforming the previous state-of-the-art retriever by an absolute gain of 3.39 percentage points and surpassing learning-free baselines by more than 10 percentage points. The framework showed substantial improvements on complex natural language inference tasks, delivering gains of over 20 percentage points compared to learning-free methods. On compositional semantic parsing benchmarks, CEIL consistently outperformed competing retrievers by capturing complementary elements required for multi-step queries. Additionally, CEIL demonstrated extreme efficiency: when restricted to just four examples, it routinely surpassed the baseline models configured with 32 examples, significantly reducing input length.

These findings indicate that treating demonstration selection as an integrated subset problem yields far more effective prompts than ranking individual examples independently. Because CEIL selects more compact yet informative demonstration sets, organizations can significantly reduce computational overhead and latency, as processing shorter inputs requires less compute. Furthermore, the retriever trained on one model transfers successfully to others—including massive models like Codex—allowing organizations to boost performance on closed-source or expensive platforms without paying to retrain model-specific retrievers.

Organizations utilizing in-context learning should adopt set-level, diversity-aware retrieval strategies in place of individual similarity matching, particularly for multi-step reasoning, semantic parsing, and code generation. Teams should also leverage smaller demonstration sizes during deployment to lower operational inference costs. Looking ahead, technical teams should explore multi-task training to develop a single, generalized retriever capable of handling new domains without task-specific training data.

The primary limitation of this approach is the upfront computational cost required to train the retriever and generate candidate subset scores using language models. While the framework demonstrates solid transferability across several tasks and model architectures, cross-task generalization remains inconsistent between single-input and double-input formats, warranting localized validation before broad production deployment.

Cover for Compositional Exemplars for In-context Learning

Abstract

Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task via a prompt consisting of input-output examples as the demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we formulate in-context example selection as a subset selection problem. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through a carefully-designed contrastive learning objective to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, paraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation, and semantic parsing. Extensive experiments demonstrate not only the state-of-the-art performance but also the transferability and compositionality of CEIL, shedding new light on in-context learning. Our code is released at https://github.com/HKUNLP/icl-ceil.

Table of Contents

  • 1. Introduction
  • 2. Preliminary
  • 2.1. In-context Learning
  • 2.2. Determinantal Point Processes
  • 3. Model
  • 3.1. Modeling
  • 3.2. Training
  • 3.3. Inference
  • 4. Experiments
  • 4.1. Datasets and Evaluation
  • 4.2. Baselines
  • 4.3. Implementation Details
  • 4.4. Main Results
  • 4.5. Compositionality
  • 4.6. Transferability
  • 4.7. Analysis
  • 5. Related Work
  • 5.1. In-context Learning
  • 5.2. Determinantal Point Processes
  • 6. Conclusion
  • Acknowledgement
  • References
  • A. Experimental Setup
  • A.1. Datasets
  • A.2. Experimental Setup for Compositionality
  • B. Additional Experiments
  • B.1. Varying n at Inference Time
  • B.2. Number of In-context Examples
  • C. Limitation

Knowls

  1. Knowl 1 — CEIL models exemplar subsets with a relevance-conditioned DPP

    model/method

    CEIL (Compositional Exemplars for In-context Learning) scores a set of candidate demonstrations for a test input rather than scoring each demonstration independently. Let xx be a test input, let ZZ be the candidate exemplar index set, and let aia_i be the learned vector representation of candidate exemplar i∈Zi\in Z. Let LL be a positive-semidefinite similarity matrix with entries Lij=k(ai,aj)L_{ij}=k(a_i,a_j), and let qiq_i be the learned relevance score between aia_i and xx. CEIL forms a query-conditioned kernel using Dii=exp⁡(qi/(2λ))D_{ii}=\exp(q_i/(2\lambda)), where DD is diagonal and λ>0\lambda>0 controls the relevance–diversity trade-off:

    L′=DLDL'=DLD.

    The conditional DPP assigns a subset S⊆ZS\subseteq Z probability P(S∣x)=det⁡(LS′)/det⁡(L′+I)P(S\mid x)=\det(L'_S)/\det(L'+I), where LS′L'_S is the submatrix indexed by SS and II is the identity matrix. Its unnormalized log score is log⁡det⁡(LS′)=λ−1∑i∈Sqi+log⁡det⁡(LS)\log\det(L'_S)=\lambda^{-1}\sum_{i\in S}q_i+\log\det(L_S). Thus relevance contributes through the individual qiq_i scores, while the determinant rewards subsets whose exemplars are not redundant. CEIL uses learnable text encoders for inputs and exemplars and applies dot-product similarity in the embedding space; the language model used to score answers remains frozen.

  2. Knowl 2 — CEIL trains set selection using frozen-LM preferences

    model/method

    For each training instance (xi,yi)(x_i,y_i), CEIL constructs MM candidate subsets SijS_{ij} of demonstrations drawn from the training data, excluding the instance itself. Each subset receives a quality score sij=PLM(yi∣Sij,xi)s_{ij}=P_{\mathrm{LM}}(y_i\mid S_{ij},x_i) from the frozen inference language model; higher probability of the training answer means the subset is more helpful. Candidate subsets are sampled without replacement and contain no repeated exemplar, avoiding zero determinants.

    For pairs ordered so that the LM prefers S+S^+ to S−S^-, the retriever is trained with the pairwise margin loss

    Li=∑(S+,S−)∈Cimax⁡(0,log⁡Pθ(S−∣xi)−log⁡Pθ(S+∣xi)ci+ξ)\mathcal{L}_i=\sum_{(S^+,S^-)\in C_i}\max\left(0,\frac{\log P_\theta(S^-\mid x_i)-\log P_\theta(S^+\mid x_i)}{c_i}+\xi\right),

    where CiC_i is the set of sampled subsets for instance ii, PθP_\theta is the conditional DPP probability, and ci=max⁡S∈Cilog⁡Pθ(S∣xi)−min⁡S∈Cilog⁡Pθ(S∣xi)c_i=\max_{S\in C_i}\log P_\theta(S\mid x_i)-\min_{S\in C_i}\log P_\theta(S\mid x_i) rescales the DPP-score differences. The margin is ξ=γ(rank⁡(S−)−rank⁡(S+))\xi=\gamma(\operatorname{rank}(S^-)-\operatorname{rank}(S^+)), with ranks ordered from best LM score to worst and γ=1/∣Ci∣\gamma=1/|C_i|. Because both subsets have the same input, the DPP normalization term cancels in their log-probability difference. This lets the loss use graded LM preferences between subsets rather than treating every negative subset identically.

  3. Knowl 3 — Greedy DPP-MAP selects a fixed-size demonstration set

    algorithm

    Given a test input xx, a candidate exemplar pool of size nn, and a desired set size KK, CEIL first narrows the full training set to the nn nearest candidates using a dense retriever. It then initializes the selected set SS to empty and repeats until ∣S∣=K|S|=K: add the remaining candidate jj that maximizes the increase in log⁡det⁡(LS∪{j}′)\log\det(L'_{S\cup\{j\}}) relative to log⁡det⁡(LS′)\log\det(L'_S). Each exemplar can be selected only once. The resulting set is the greedy solution to DPP maximum-a-posteriori selection; global MAP optimization is NP-hard, so the greedy procedure is not a guarantee of the global optimum. An incremental Cholesky factorization updates the determinant scores, giving stated total complexity O(nK2)O(nK^2) rather than recomputing each determinant from scratch. The reported typical setting is n=100n=100, K=16K=16; the selected exemplars are arranged in ascending similarity to the input when placed in the prompt.

  4. Knowl 4 — Experimental design and training configuration

    experimental setup

    CEIL was evaluated on 12 datasets spanning sentiment analysis (SST-5), paraphrase detection (MRPC), natural-language inference (QNLI and MNLI), commonsense reasoning (CommonsenseQA/CMSQA and HellaSwag), open-domain question answering (WebQs), code generation (GeoQuery and NL2Bash), and semantic parsing (Break, MTOP, and SMCalFlow). Classification was scored by accuracy; generation was scored by exact match on WebQs, GeoQuery, NL2Bash, MTOP, and SMCalFlow, and by LF-EM on Break. Results were reported on validation sets.

    The primary inference model was GPT-Neo with 2.7 billion parameters; GPT2-XL (1.5 billion) and Codex (175 billion) were also used in transfer experiments. Prompts used up to 50 demonstrations, truncated to each model's context limit. Retriever encoders were BERT-base models initialized from EPR, with labels excluded from the text encoded for each example. Retriever training used at most 44,000 instances, 50 candidate subsets of 16 demonstrations per instance, Adam with batch size 128 and learning rate 10−510^{-5}, and 30 epochs on two NVIDIA A100 GPUs. The trade-off parameter λ\lambda was selected from {0.01,0.05,0.1}\{0.01,0.05,0.1\}. For candidate-subset construction, the reported standard setting used a top-100 candidate pool followed by random subset sampling.

  5. Knowl 5 — CEIL improves average in-context performance across 12 datasets

    empirical result

    With GPT-Neo as the inference model, CEIL was compared with learning-free retrievers and EPR, a learned retriever that selects individual demonstrations. Scores below are percentages; classification datasets use accuracy, generation datasets use exact match except Break, which uses LF-EM. CEIL exceeds EPR on every dataset and has the highest reported mean score. The mean is an unweighted average across the listed task scores.

    MethodSST-5MRPCQNLIMNLICMSQAHellaSwagWebQsGeoQueryNL2BashBreakMTOPSMCalFlowAvg.
    EPR42.8275.9880.7666.0636.7742.6119.5968.5756.8231.9064.2054.3053.37
    CEIL47.0580.1585.4171.7437.1843.2020.9273.2159.9134.1867.4360.7356.76
    CEIL minus EPR+4.23+4.17+4.65+5.68+0.41+0.59+1.33+4.64+3.09+2.28+3.23+6.43+3.39

    The best learning-free mean was 45.84, from TOPK-BM25. CEIL's per-dataset advantage over EPR was largest on natural-language inference and SMCalFlow; the improvement was smaller on commonsense reasoning, where retrieval methods generally performed near the random baseline. CEIL adds no parameters relative to the EPR retriever, although its set scoring and inference entail additional computation.

  6. Knowl 6 — CEIL retrieves useful demonstrations for compositional semantic parsing

    empirical result

    The compositionality evaluation used GeoQuery's Standard, Template, TMCD, and Length splits and SMCalFlow-CS's single-domain test split (0-S) and cross-domain splits (0-C, 8-C, 16-C, 32-C). The kk-C splits include kk compositional demonstrations in the prompt; 0-C has none. Scores are percentages. GPT-Neo results use retrievers trained on the corresponding GeoQuery and SMCalFlow data; the Codex experiment reused the GPT-Neo retriever trained on GeoQuery and SMCalFlow. Dashes indicate evaluations not reported.

    InferencerRetrieverGeoQuery StandardTemplateTMCDLengthSMCalFlow-CS 0-S0-C8-C16-C32-C
    GPT-NeoEPR68.5738.9544.0932.2757.780.000.00——
    GPT-NeoCEIL73.2140.7744.0932.7360.270.000.28——
    CodexEPR91.7087.9362.7373.4180.830.5635.5638.6148.06
    CodexCEIL93.2189.9863.6474.0981.391.6742.7848.0655.28

    Relative to EPR, CEIL improves GeoQuery Standard and Template performance by 4.64 and 1.82 points with GPT-Neo, and by 1.51 and 2.50 points with Codex. On Codex's cross-domain SMCalFlow-CS splits, the improvements are 1.11 points at 0-C, 7.22 at 8-C, 9.45 at 16-C, and 7.22 at 32-C. These results support the claim that jointly selecting demonstrations can help with compositional and longer programs. They do not show generalization to compositional tasks unseen by the retriever: the retriever was trained on the standard task data, and the kk-C settings supply compositional examples in context.

  7. Knowl 7 — Retriever transfer depends on the source and target task

    empirical result

    The paper tested whether a retriever trained on one dataset could be used on another without retraining. Each entry below is the absolute performance improvement, in percentage points, over TOPK-BERT on the target dataset. Rows are retriever-training datasets and columns are evaluation datasets; positive values indicate improvement.

    Trained on → evaluated onSST-5QNLIMNLIGeoQueryMTOPSMCalFlow
    SST-59.82.04.7-5.7-1.0-2.2
    QNLI-4.820.85.3-10.4-21.6-20.4
    MNLI-9.77.229.6-20.0-47.3-32.1
    GeoQuery-2.81.24.56.41.00.3
    MTOP-1.30.53.91.814.62.6
    SMCalFlow-2.90.94.4-1.81.215.2

    Retrievers trained on natural-language inference often transfer between QNLI and MNLI, but the NLI-trained retrievers usually do not help on other task types. The authors suggest that NLI's two-input format may contribute to this asymmetry. Among the listed single-input tasks, transfer is more evident between related code-generation and semantic-parsing tasks. A separate cross-inference-model test found that a GPT-Neo-trained retriever could be used with GPT2-XL and Codex without retraining; it was comparable to a GPT2-XL-trained retriever on several tested GPT2-XL tasks, and still improved over TOPK-BERT with Codex.

  8. Knowl 8 — Retriever initialization and contrastive objective affect task performance

    empirical result

    Ablations compared BERT initialization with EPR initialization and the pairwise margin loss with InfoNCE. Scores are percentages. EPR initialization substantially helps relative to training directly from BERT, while the pairwise loss is especially effective on generation tasks such as GeoQuery and MTOP. The selected CEIL configuration uses EPR initialization and the pairwise loss.

    ConfigurationSST-5MRPCQNLIGeoQueryMTOP
    TOPK-BERT37.2469.3664.6566.7952.13
    EPR42.8275.9880.7668.5764.20
    BERT initialization + InfoNCE31.3469.1263.9268.5747.43
    BERT initialization + pairwise loss35.5567.8965.0067.5041.30
    EPR initialization + InfoNCE49.1480.6485.5469.2961.92
    EPR initialization + pairwise loss (CEIL)47.0580.1585.4173.2167.43

    On GeoQuery and MTOP, the EPR-initialized pairwise model scores 3.92 and 5.51 points above its InfoNCE counterpart, respectively; the InfoNCE variant scores slightly higher on some classification tasks. Candidate construction also matters: top-100 retrieval followed by random subset sampling outperformed fixed-size k-DPP sampling at similar Top-K-candidate MRR. For GeoQuery and MTOP, increasing the number of sampled subsets from 10 to 50 improved scores from 67.86 to 73.21 and from 62.37 to 67.43, respectively. One-stage random sampling performed worse than the two-stage top-100-plus-random strategy on SST-5 and MTOP.

  9. Knowl 9 — CEIL can use fewer demonstrations, with a retrieval-size efficiency trade-off

    empirical result

    In experiments varying the number of demonstrations, CEIL generally retained or improved classification performance as the number of demonstrations decreased, whereas generation performance tended to improve with more demonstrations. The authors report that CEIL mostly exceeded EPR with 32 demonstrations using only 4, and mostly exceeded TOPK-BERT with 32 using only 1. This suggests that a carefully selected compact prompt can reduce the attention computation required by the inference model.

    A separate sweep varied nn, the size of the candidate pool used for greedy DPP-MAP, while retaining the same evaluation setup. Increasing nn improved the mean score only modestly at larger pool sizes, while adding retrieval latency on the SST-5 validation set.

    CEIL candidate pool nnSST-5 latencyMean score
    5030 s67.96
    10036 s68.86
    20055 s68.88
    40087 s69.28
    800118 s69.32

    The mean rises by 1.36 points from n=50n=50 to n=800n=800, while latency rises from 30 s to 118 s. The authors also observed that candidates beyond roughly the top 200 were not typically selected, consistent with diminishing performance gains at larger pools.

  10. Knowl 10 — CEIL requires task-specific training data and has slower subset scoring

    limitation

    CEIL is a learned retriever and therefore requires training data for each task in its standard application; constructing training preferences is slower than for EPR because the frozen inference language model must score a full demonstration subset rather than a single demonstration. Although the experiments show transfer across some datasets and inference models, the paper characterizes this transfer as preliminary and does not establish a retriever that works reliably for all tasks. The authors identify multitask training of a unified retriever as a possible way to reduce the need for retraining on each new task.

Coverage note — Detailed task-specific prompt examples and full plotted curves for every demonstration count are omitted as implementation or secondary analysis; the method, principal benchmark results, compositionality and transfer evaluations, training ablations, efficiency findings, and stated limitations are covered.

References

  1. 1.Agrawal, S., Zhou, C., Lewis, M., Zettlemoyer, L., and Ghazvininejad, M. In-context examples selection for machine translation. arXiv preprint arXiv:2212.02437, 2022.
  2. 2.An, C., Feng, J., Lv, K., Kong, L., Qiu, X., and Huang, X. Cont: Contrastive neural text generation. NeurIPS, 2022.
  3. 3.Andreas, J., Bufe, J., Burkett, D., Chen Jr, C., Clausman, J., Crawford, J., Crim, K., DeLoach, J., Dorner, L., Eisner, J., et al. Task-oriented dialogue as dataflow synthesis. Transactions of the Association for Computational Linguistics, 8:556–571, 2020.
  4. 4.Azadi, S., Feng, J., and Darrell, T. Learning detection with diverse proposals. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7149–7157, 2017.
  5. 5.Berant, J., Chou, A., Frostig, R., and Liang, P. Semantic parsing on Freebase from question-answer pairs. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pp. 1533–1544, Seattle, Washington, USA, October 2013. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/D13-1160.
  6. 6.Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021. URL https://doi.org/10.5281/zenodo.5297715.
  7. 7.Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. Language models are few-shot learners. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
  8. 8.Chen, L., Zhang, G., and Zhou, E. Fast greedy map inference for determinantal point process to improve recommendation diversity. Advances in Neural Information Processing Systems, 31, 2018.
  9. 9.Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021a.
  10. 10.Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021b.
  11. 11.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4171–4186, Minneapolis, Minnesota, 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/N19-1423.
  12. 12.Dolan, W. B., Quirk, C., and Brockett, C. Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources. In COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics, pp. 350–356, 2004.
  13. 13.FAIR, Bakhtin, A., Brown, N., Dinan, E., Farina, G., Flaherty, C., Fried, D., Goff, A., Gray, J., Hu, H., et al. Human-level play in the game of diplomacy by combining language models with strategic reasoning. Science (New York, NY), 378(6624):1067–1074, 2022.
  14. 14.Finegan-Dollak, C., Kummerfeld, J. K., Zhang, L., Ramanathan, K., Sadasivam, S., Zhang, R., and Radev, D. Improving text-to-SQL evaluation methodology. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 351–360, Melbourne, Australia, July 2018. Association for Computational Linguistics. doi: 10.18653/v1/P18-1033. URL https://aclanthology.org/P18-1033.
  15. 15.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. The pile: An 800gb dataset of diverse text for language modeling. ArXiv preprint, abs/2101.00027, 2021a. URL https://arxiv.org/abs/2101.00027.
  16. 16.Gao, T., Yao, X., and Chen, D. Simcse: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 6894–6910, 2021b.
  17. 17.Gillenwater, J., Kulesza, A., and Taskar, B. Near-optimal map inference for determinantal point processes. Advances in Neural Information Processing Systems, 25, 2012.
  18. 18.Gong, B., Chao, W.-L., Grauman, K., and Sha, F. Diverse sequential subset selection for supervised video summarization. Advances in neural information processing systems, 27, 2014.
  19. 19.Han, I., Kambadur, P., Park, K., and Shin, J. Faster greedy map inference for determinantal point processes. In International Conference on Machine Learning, pp. 1384–1393. PMLR, 2017.
  20. 20.Hasson, M. and Berant, J. Question decomposition with dependency graphs. In 3rd Conference on Automated Knowledge Base Construction, 2021.
  21. 21.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9729–9738, 2020.
  22. 22.Heilbron, F. C., Escorcia, V., Ghanem, B., and Niebles, J. C. Activitynet: A large-scale video benchmark for human activity understanding. In 2015 IEEE conference on computer vision and pattern recognition (CVPR), pp. 961–970. IEEE, 2015.
  23. 23.Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y. The curious case of neural text degeneration. In International Conference on Learning Representations, 2019.
  24. 24.Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bojanowski, P., Joulin, A., and Grave, E. Towards unsupervised dense information retrieval with contrastive learning. arXiv preprint arXiv:2112.09118, 2021.
  25. 25.Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 6769–6781, 2020.
  26. 26.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Proceedings of ICLR, 2015.
  27. 27.Ko, C.-W., Lee, J., and Queyranne, M. An exact algorithm for maximum entropy sampling. Operations Research, 43(4):684–691, 1995.
  28. 28.Kulesza, A. and Taskar, B. k-dpps: Fixed-size determinantal point processes. In ICML, 2011.
  29. 29.Kulesza, A., Taskar, B., et al. Determinantal point processes for machine learning. Foundations and Trends® in Machine Learning, 5(2–3):123–286, 2012.
  30. 30.Kulis, B. et al. Metric learning: A survey. Foundations and Trends® in Machine Learning, 5(4):287–364, 2013.
  31. 31.Levy, I., Bogin, B., and Berant, J. Diverse demonstrations improve in-context compositional generalization. arXiv preprint arXiv:2212.06800, 2022.
  32. 32.Li, H., Arora, A., Chen, S., Gupta, A., Gupta, S., and Mehdad, Y. Mtop: A comprehensive multilingual task-oriented semantic parsing benchmark. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pp. 2950–2962, 2021.
  33. 33.Li, Y., Lin, Z., Zhang, S., Fu, Q., Chen, B., Lou, J.-G., and Chen, W. On the advance of making language models better reasoners. arXiv preprint arXiv:2206.02336, 2022.
  34. 34.Lin, X. V., Wang, C., Zettlemoyer, L., and Ernst, M. D. Nl2bash: A corpus and semantic parser for natural language interface to the linux operating system. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation LREC 2018, Miyazaki (Japan), 7-12 May, 2018., 2018.
  35. 35.Liu, J., Shen, D., Zhang, Y., Dolan, W. B., Carin, L., and Chen, W. What makes good in-context examples for gpt-3? In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, pp. 100–114, 2022.
  36. 36.Liu, T.-Y. et al. Learning to rank for information retrieval. Foundations and Trends® in Information Retrieval, 3(3):225–331, 2009.
  37. 37.Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 8086–8098, 2022.
  38. 38.Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021.
  39. 39.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  40. 40.OpenAI, T. Chatgpt: Optimizing language models for dialogue. OpenAI, 2022.
  41. 41.Qiu, L., Shaw, P., Pasupat, P., Nowak, P., Linzen, T., Sha, F., and Toutanova, K. Improving compositional generalization with latent structure and data augmentation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4341–4362, Seattle, United States, July 2022a. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.323. URL https://aclanthology.org/2022.naacl-main.323.
  42. 42.Qiu, L., Shaw, P., Pasupat, P., Shi, T., Herzig, J., Pitler, E., Sha, F., and Toutanova, K. Evaluating the impact of model scale for compositional generalization in semantic parsing. arXiv preprint arXiv:2205.12253, 2022b.
  43. 43.Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I. Language models are unsupervised multitask learners. 2019.
  44. 44.Robertson, S. and Zaragoza, H. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends in Information Retrieval, 3:333–389, 01 2009. doi: 10.1561/1500000019.
  45. 45.Rohrbach, A., Torabi, A., Rohrbach, M., Tandon, N., Pal, C., Larochelle, H., Courville, A., and Schiele, B. Movie description. International Journal of Computer Vision, 123:94–120, 2017.
  46. 46.Rubin, O., Herzig, J., and Berant, J. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2655–2671, Seattle, United States, July 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.naacl-main.191. URL https://aclanthology.org/2022.naacl-main.191.
  47. 47.Scholkopf, B., Smola, A. J., Bach, F., et al. Learning with kernels: support vector machines, regularization, optimization, and beyond. MIT press, 2002.
  48. 48.Shaw, P., Chang, M.-W., Pasupat, P., and Toutanova, K. Compositional generalization and natural language variation: Can a semantic parsing approach handle both? In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pp. 922–938, Online, August 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.75. URL https://aclanthology.org/2021.acl-long.75.
  49. 49.Shi, F., Fried, D., Ghazvininejad, M., Zettlemoyer, L., and Wang, S. I. Natural language to code translation with execution. arXiv preprint arXiv:2204.11454, 2022.
  50. 50.Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pp. 1631–1642, 2013.
  51. 51.Su, H., Kasai, J., Wu, C. H., Shi, W., Wang, T., Xin, J., Zhang, R., Ostendorf, M., Zettlemoyer, L., Smith, N. A., et al. Selective annotation makes language models better few-shot learners. arXiv preprint arXiv:2209.01975, 2022.
  52. 52.Talmor, A., Herzig, J., Lourie, N., and Berant, J. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pp. 4149–4158, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1421. URL https://aclanthology.org/N19-1421.
  53. 53.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  54. 54.Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. Glue: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pp. 353–355, 2018.
  55. 55.Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  56. 56.Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022.
  57. 57.Williams, A., Nangia, N., and Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pp. 1112–1122, 2018.
  58. 58.Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38–45, Online, October 2020. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2020.emnlp-demos.6.
  59. 59.Wolfson, T., Geva, M., Gupta, A., Gardner, M., Goldberg, Y., Deutch, D., and Berant, J. Break it down: A question understanding benchmark. Transactions of the Association for Computational Linguistics, 8:183–198, 2020.
  60. 60.Wu, Z., Wang, Y., Ye, J., and Kong, L. Self-adaptive in-context learning. arXiv preprint arXiv:2212.10375, 2022.
  61. 61.Xie, P., Salakhutdinov, R., Mou, L., and Xing, E. P. Deep determinantal point process for large-scale multi-label classification. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Oct 2017.
  62. 62.Ye, J., Gao, J., Wu, Z., Feng, J., Yu, T., and Kong, L. ProGen: Progressive zero-shot dataset generation via in-context feedback. In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 3671–3683, Abu Dhabi, United Arab Emirates, December 2022a. Association for Computational Linguistics. URL https://aclanthology.org/2022.findings-emnlp.269.
  63. 63.Ye, J., Li, C., Kong, L., and Yu, T. Generating data for symbolic language with large language models. 2023.
  64. 64.Ye, X., Iyer, S., Celikyilmaz, A., Stoyanov, V., Durrett, G., and Pasunuru, R. Complementary explanations for effective in-context learning. arXiv preprint arXiv:2211.13892, 2022b.
  65. 65.Yin, P., Fang, H., Neubig, G., Pauls, A., Platanios, E. A., Su, Y., Thomson, S., and Andreas, J. Compositional generalization for neural semantic parsing via span-level supervised attention. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 2810–2823, 2021.
  66. 66.Zelle, J. M. and Mooney, R. J. Learning to parse database queries using inductive logic programming. In AAAI/IAAI, pp. 1050–1055, Portland, OR, August 1996. AAAI Press/MIT Press. URL http://www.cs.utexas.edu/users/ai-lab?zelle:aaai96.
  67. 67.Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  68. 68.Zhong, M., Liu, P., Chen, Y., Wang, D., Qiu, X., and Huang, X. Extractive summarization as text matching. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 6197–6208, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.552. URL https://aclanthology.org/2020.acl-main.552.

Citation

MLA
Ye, J., et al. “Compositional Exemplars for In-context Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 39818–33, https://proceedings.mlr.press/v202/ye23c.html.
APA
Ye, J., Wu, Z., Feng, J., Yu, T., & Kong, L. (2023). Compositional Exemplars for In-context Learning. International Conference on Machine Learning, 202, 39818–39833. https://proceedings.mlr.press/v202/ye23c.html
Chicago
Ye, J., Z. Wu, J. Feng, T. Yu, and L. Kong. 2023. “Compositional Exemplars for In-context Learning”. International Conference on Machine Learning 202: 39818–33. https://proceedings.mlr.press/v202/ye23c.html.
Harvard
Ye, J. et al. (2023) “Compositional Exemplars for In-context Learning”, International Conference on Machine Learning. PMLR, pp. 39818–39833. Available at: https://proceedings.mlr.press/v202/ye23c.html.
Vancouver
1. Ye J, Wu Z, Feng J, Yu T, Kong L (2023) Compositional Exemplars for In-context Learning. In: International Conference on Machine Learning. PMLR, pp 39818–39833

BibTeX

@InProceedings{pmlr-v202-ye23c,
  title = 	 {Compositional Exemplars for In-context Learning},
  author =       {Ye, Jiacheng and Wu, Zhiyong and Feng, Jiangtao and Yu, Tao and Kong, Lingpeng},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {39818--39833},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/ye23c/ye23c.pdf},
  url = 	 {https://proceedings.mlr.press/v202/ye23c.html},
  abstract = 	 {Large pretrained language models (LMs) have shown impressive In-Context Learning (ICL) ability, where the model learns to do an unseen task simply by conditioning on a prompt consisting of input-output examples as demonstration, without any parameter updates. The performance of ICL is highly dominated by the quality of the selected in-context examples. However, previous selection methods are mostly based on simple heuristics, leading to sub-optimal performance. In this work, we systematically formulate in-context example selection as a subset selection problem, and optimize it in an end-to-end fashion. We propose CEIL (Compositional Exemplars for In-context Learning), which is instantiated by Determinantal Point Processes (DPPs) to model the interaction between the given input and in-context examples, and optimized through carefully-designed contrastive learning to obtain preference from LMs. We validate CEIL on 12 classification and generation datasets from 7 distinct NLP tasks, including sentiment analysis, phraphrase detection, natural language inference, commonsense reasoning, open-domain question answering, code generation and semantic parsing. Extensive experiments demonstrate the effectiveness, transferability, compositionality of CEIL, shedding new lights on in-context leaning. Our code is released at https://github.com/HKUNLP/icl-ceil.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/