Unified Demonstration Retriever for In-Context Learning

Xiaonan LiKai LvHang YanTianyang LinWei ZhuYuan NiGuotong XieXiaoling WangXipeng Qiu

article2023ACL218 citations

Proposes a multi-task list-wise ranking framework that trains a single unified retriever using language model feedback to select effective in-context learning demonstrations across diverse unseen tasks and model scales.

Listen

Large language models increasingly rely on in-context learning to perform tasks by observing a few input-output demonstration examples without retraining model parameters. However, overall system performance depends heavily on the quality and relevance of the retrieved demonstrations. Existing approaches typically rely either on basic semantic similarity search or on specialized, task-specific retrievers that require separate engineering, training, and maintenance for each individual use case, driving up infrastructure and deployment costs.

The article evaluates and demonstrates a single, multi-task framework called the Unified Demonstration Retriever to retrieve high-quality demonstrations across diverse natural language processing applications. The goal is to provide a unified, parameter-efficient retrieval model that scales across multiple tasks without needing separate retrievers for each specific domain.

To achieve this, the authors developed a list-wise ranking training framework that uses feedback from a language model to assess candidate demonstrations across tasks. Rather than relying on rigid manual rules, the retriever uses an iterative mining strategy where it searches the full training dataset to identify strong positive examples and informative hard negative examples, training a two-tower neural retriever with task instructions. The system was trained and evaluated across more than 30 tasks spanning 13 task families, including classification, question answering, text summarization, semantic parsing, and code summarization, using language models ranging from 1.3 billion to 175 billion parameters.

The findings show that the unified retriever consistently outperformed standard baselines across benchmarks, achieving an average improvement of roughly 10 points over general text embeddings on classification tasks and approximately 7 points on text generation tasks. Furthermore, the retriever transferred effectively to completely unseen datasets, outperforming baseline semantic and keyword search retrievers by around 10 points on average. The method maintained robust performance when evaluated across various language model sizes—scaling smoothly from small models up to 175-billion-parameter systems—and demonstrated that demonstration quality is far more critical than quantity, with two high-quality examples often surpassing eight lower-quality demonstrations. Unlike random demonstration selection, the unified retriever's selections proved largely insensitive to prompt ordering, showing variations of less than 1 point across sequence arrangements.

These results indicate that organizations deploying large language models can replace costly, fragmented retrieval components with a single, lightweight retrieval model. This consolidation significantly cuts parameter storage overhead, simplifies deployment pipelines, and boosts inference accuracy and consistency across varied tasks. The stability across different prompt templates and prompt ordering also mitigates operational risks associated with brittle prompt engineering.

Organizations seeking to improve generative artificial intelligence performance should consider adopting a unified, model-feedback-driven retrieval approach instead of investing in separate task-specific retrievers or relying on basic similarity search. Practitioners can implement the unified retriever across diverse workflows to optimize demonstration quality while maintaining a streamlined footprint.

Decision-makers should note certain limitations: the retriever was trained using base BERT architectures, was evaluated by scoring individual demonstrations independently rather than modeling multi-demonstration interactions jointly, and operates largely as a black-box model whose underlying selection mechanics warrant deeper interpretability. Nonetheless, because results were validated across dozens of tasks and multiple independent model architectures, confidence in the retriever's generalizability and practical effectiveness remains high.

arXiv: 2305.04320
Cover for Unified Demonstration Retriever for In-Context Learning

Abstract

In-context learning is a new learning paradigm where a language model conditions on a few input-output pairs (demonstrations) and a test input, and directly outputs the prediction. It has been shown highly dependent on the provided demonstrations and thus promotes the research of demonstration retrieval: given a test input, relevant examples are retrieved from the training set to serve as informative demonstrations for in-context learning. While previous works focus on training task-specific retrievers for several tasks separately, these methods are often hard to transfer and scale on various tasks, and separately trained retrievers incur a lot of parameter storage and deployment cost. In this paper, we propose Unified Demonstration Retriever (UDR), a single model to retrieve demonstrations for a wide range of tasks. To train UDR, we cast various tasks' training signals into a unified list-wise ranking formulation by language model's feedback. Then we propose a multi-task list-wise ranking training framework, with an iterative mining strategy to find high-quality candidates, which can help UDR fully incorporate various tasks' signals. Experiments on 30+ tasks across 13 task families and multiple data domains show that UDR significantly outperforms baselines. Further analyses show the effectiveness of each proposed component and UDR's strong ability in various scenarios including different LMs (1.3B ~ 175B), unseen datasets, varying demonstration quantities, etc.

Table of Contents

  • 1 Introduction
  • 2 Unified Demonstration Retriever
  • 2.1 Bi-encoder with Task Instruction
  • 2.2 Learning from LM Feedback
  • 2.2.1 Ranking Candidates by LM
  • 2.2.2 Loss Function
  • 2.2.3 Iterative Candidate Mining
  • 2.3 Inference
  • 3 Experiment
  • 3.1 Experimental Settings
  • 3.2 Main Results
  • 3.3 Analysis
  • 3.3.1 Ablation Study
  • 3.3.2 Transferability across Different LMs
  • 3.3.3 Performance on Unseen Datasets
  • 3.3.4 The Order of Demonstrations
  • 3.3.5 The Impact of Demonstration Quantity
  • 4 Related Work
  • 5 Conclusion
  • Limitations
  • 6 Acknowledgements
  • References
  • A Task Overview
  • B Implementation Details and Hyper-Parameters
  • C Performance across varying inference templates
  • D Potential Risk

Knowls

  1. Knowl 1 — Bi-Encoder Architecture with Task Instructions for Demonstration Retrieval

    model/method

    The Unified Demonstration Retriever (UDR) uses a dual-encoder architecture based on Dense Passage Retrieval (DPR) to retrieve in-context demonstrations for multiple NLP tasks. To allow a single retriever to distinguish between different tasks while maintaining parameter efficiency, task-specific textual instructions are prepended to both query and candidate examples.

    Given an input query xx, a candidate demonstration z={x′,y′}z = \{x', y'\} consisting of an input-output pair, and a textual task instruction IiI_i corresponding to task TiT_i (e.g., "Summarize the text" or "Sentiment of the sentence:"), the query encoder EqE_q and demonstration encoder EdE_d encode the instruction-augmented inputs respectively. The retrieval score is calculated as the inner product of their pooled embeddings:

    sim(x,z)=Eq(Ii⊕x)⊤Ed(Ii⊕z)\text{sim}(x, z) = E_q(I_i \oplus x)^\top E_d(I_i \oplus z)

    where ⊕\oplus denotes string concatenation, and EqE_q and EdE_d are multi-layer Transformer encoders initialized with separate pre-trained BERT-base-uncased models (totaling approximately 220 million parameters) using [CLS] token pooling.

  2. Knowl 2 — Demonstration Candidate Scoring via Language Model Feedback

    model/method

    To train a demonstration retriever without task-specific heuristic labels, UDR ranks candidate demonstrations based on the feedback of an inference language model (LM) GG. For a given query xx, ground-truth target yy, and a candidate demonstration list Z={zi}i=1lZ = \{z_i\}_{i=1}^l, LM feedback scores each candidate by measuring how effectively conditioning on zjz_j enables GG to generate yy.

    For text generation tasks, the score s(zj)s(z_j) is the conditional likelihood of the ground-truth target given candidate zjz_j and query xx:

    sgen(zj)=pG(y∣zj,x)s_{\text{gen}}(z_j) = p_G(y \mid z_j, x)

    For classification and multiple-choice tasks with label space or choice set Y\mathcal{Y}, the score is normalized over all possible verbalizers/choices:

    scls(zj)=pG(y∣zj,x)∑y′∈YpG(y′∣zj,x)s_{\text{cls}}(z_j) = \frac{p_G(y \mid z_j, x)}{\sum_{y' \in \mathcal{Y}} p_G(y' \mid z_j, x)}

    The candidates in ZZ are then ranked according to their scores, where candidate zjz_j is assigned rank r(zj)=rank(s(zj)∣{s(zi)}i=1l)r(z_j) = \text{rank}(s(z_j) \mid \{s(z_i)\}_{i=1}^l) such that a higher LM score corresponds to a higher ranking position (with rank 1 being the best).

  3. Knowl 3 — Unified Multi-Task List-wise Ranking and In-Batch Negative Objective

    equation

    UDR optimizes demonstration retrieval across diverse tasks through a combined list-wise ranking loss and an in-batch negative loss.

    Inspired by LambdaRank, the list-wise ranking loss penalizes demonstration pairs whose predicted similarity order disagrees with their LM feedback ranking:

    Lrank=∑zi,zj∈Zw⋅log⁡(1+esim(x,zj)−sim(x,zi))\mathcal{L}_{\text{rank}} = \sum_{z_i, z_j \in Z} w \cdot \log\left(1 + e^{\text{sim}(x, z_j) - \text{sim}(x, z_i)}\right)

    where the pair weight w=max⁡(0,1r(zi)−1r(zj))w = \max\left(0, \frac{1}{r(z_i)} - \frac{1}{r(z_j)}\right) scales the gradient proportionally to the difference in reciprocal ranks when r(zi)<r(zj)r(z_i) < r(z_j).

    To exploit computational efficiency across batch samples, an in-batch negative contrastive loss is incorporated:

    Lib=−log⁡esim(x,z∗)∑z∈Zesim(x,z)\mathcal{L}_{\text{ib}} = -\log \frac{e^{\text{sim}(x, z^*)}}{\sum_{z \in \mathcal{Z}} e^{\text{sim}(x, z)}}

    where z∗z^* is the rank-1 candidate demonstration for query xx, and Z\mathcal{Z} denotes the set of all candidate demonstrations in the batch.

    The combined training loss is:

    L=λLrank+(1−λ)Lib\mathcal{L} = \lambda \mathcal{L}_{\text{rank}} + (1 - \lambda) \mathcal{L}_{\text{ib}}

    where λ=0.8\lambda = 0.8. To prevent bias toward high-resource tasks during multi-task training, each task TiT_i is sampled according to probability p(Ti)=qiα∑j=1Tqjαp(T_i) = \frac{q_i^\alpha}{\sum_{j=1}^T q_j^\alpha}, with qi=∣DTi∣∑j=1T∣DTj∣q_i = \frac{|\mathcal{D}_{T_i}|}{\sum_{j=1}^T |\mathcal{D}_{T_j}|} and smoothing parameter α=0.5\alpha = 0.5.

  4. Knowl 4 — Multi-Task List-Wise Ranking Training with Iterative Candidate Mining

    algorithm

    UDR employs an iterative mining strategy to discover both informative positive demonstrations and hard negatives from the entire training corpus rather than relying solely on initial surface-similarity heuristics.

    Input: Dual encoders EqE_q and EdE_d, language model GG, training sets for TT tasks {DTi}i=1T\{\mathcal{D}_{T_i}\}_{i=1}^T, number of iterations Niter=3N_{\text{iter}} = 3, candidate pool size K=50K = 50, sample list size l=8l = 8, task sampling distribution p(T)p(T).
    Output: Trained bi-encoder parameters EqE_q and EdE_d.
    Initialize EqE_q and EdE_d with BERT-base-uncased.
    For each task TiT_i, initialize candidate set Z∗Z^* of size KK for each example using BM25 (on inputs for classification/multi-choice; on targets for generation).
    Score all initialized candidates using LM GG.
    Train initial bi-encoder on scored candidates for 30 epochs.
    for iteration = 1 to NiterN_{\text{iter}} do
        for each training step do
            Sample task Ti∼p(T)T_i \sim p(T).
            Sample a batch of queries from DTi\mathcal{D}_{T_i}.
            For each query xx, randomly sample ll candidates z1…lz_{1 \dots l} from its current candidate pool Z∗Z^*, rank them by LM score s(z)s(z), and compute L=λLrank+(1−λ)Lib\mathcal{L} = \lambda \mathcal{L}_{\text{rank}} + (1 - \lambda) \mathcal{L}_{\text{ib}}.
            Update parameters of EqE_q and EdE_d via AdamW gradient descent.
        end for
        For each training example x∈DTix \in \mathcal{D}_{T_i}, mine updated candidate set Z∗=top-Kz∈DTisim(x,z)Z^* = \text{top-}K_{z \in \mathcal{D}_{T_i}} \text{sim}(x, z) using current EqE_q and EdE_d.
        Score new candidates in Z∗Z^* using LM GG.
        Train bi-encoder for 10 epochs on updated candidates.
    end for
    return Eq,EdE_q, E_d

    At test time, the full training set of target task TiT_i is encoded with Ed(Ii⊕⋅)E_d(I_i \oplus \cdot). For a test query xtestx_{\text{test}}, FAISS searches for the top-LL demonstrations ascendingly sorted by similarity sim(xtest,⋅)\text{sim}(x_{\text{test}}, \cdot), which are concatenated with xtestx_{\text{test}} for LM inference.

  5. Knowl 5 — Cross-Task Demonstration Retrieval Benchmark Results

    data/table

    UDR was evaluated across ~40 NLP tasks spanning 13 task families using GPT-Neo-2.7B as both the scoring and inference model. Baselines include Random demonstration selection, BM25, Sentence-BERT (SBERT, paraphrase-mpnet-base-v2), Instructor-base, DR-Target (BM25 target similarity + DPR), and EPR (single-task DPR with LM binary scoring).

    Retrieval Method SST-2 Yelp AGNews TREC MNLI COLA Cls Overall Gen Overall
    Random 57.7 25.3 74.2 42.6 34.1 52.8 47.3 15.8
    BM25 74.1 36.9 88.4 89.4 35.3 53.3 57.7 24.2
    SBERT 84.3 36.0 88.3 89.4 38.0 56.3 61.6 23.7
    Instructor 83.7 46.6 89.6 91.2 41.3 56.0 63.2 24.2
    DR-Target – – – – – – – 23.8
    EPR 87.9 49.6 89.9 95.2 54.3 70.3 68.8 27.7
    UDR 92.4 61.7 91.5 96.6 73.8 78.9 73.2 30.9

    On classification and multi-choice benchmarks, UDR achieves an overall average score of 73.2% accuracy, outperforming EPR (68.8%), Instructor (63.2%), SBERT (61.6%), and BM25 (57.7%). On generation benchmarks (including semantic parsing, summarization, commonsense generation, story generation, code summarization, text simplification, and data-to-text), UDR achieves an overall score of 30.9 points, outperforming EPR (27.7), Instructor (24.2), and BM25 (24.2).

  6. Knowl 6 — Ablation of UDR Training Components

    empirical result

    An ablation study evaluated the contribution of individual components of UDR on three representative datasets: SMCalFlow (Exact Match), SST-2 (Accuracy), and Java code summarization (BLEU-1).

    Configuration SMCalFlow SST-2 Java Average
    Full UDR 60.8 91.3 23.2 58.4
    - w/o Task Prompt 60.1 90.8 21.9 57.6
    - w/o MultiTask 60.9 91.0 22.9 58.3
    - w/o Rank Loss 56.7 89.2 21.1 55.7
    - w/o Self-Guided Iterative Mining 59.5 90.2 19.7 56.5

    Removing the list-wise ranking loss (replacing it with EPR's binary cross-entropy objective) caused the largest average drop of 2.7 points (58.4 to 55.7). Removing self-guided iterative candidate mining dropped average performance by 1.9 points (58.4 to 56.5). Removing task prompts led to a 0.8 point degradation. Training on single tasks separately (- w/o MultiTask) performed comparably on average (58.3 vs 58.4), showing that a single unified multi-task UDR model matches or slightly exceeds separately trained task-specific models while reducing deployment and parameter overhead.

  7. Knowl 7 — Transferability of UDR Demonstrations Across Language Model Architectures and Scales

    empirical result

    Demonstrations retrieved by UDR (trained exclusively with feedback from GPT-Neo-2.7B) generalize effectively across different inference language models ranging from 1.3B to 175B parameters, evaluated on 1,000 sampled test instances from SMCalFlow and E2E.

    SMCalFlow (EM) E2E (BLEU-4)
    Inference LM BM25 EPR UDR BM25 EPR UDR
    GPT-Neo-1.3B 44.8 52.9 59.5 29.9 29.7 31.9
    GPT-Neo-2.7B 46.5 53.7 62.2 29.2 29.1 32.6
    GPT-J (6B) 49.0 55.9 64.0 33.3 33.7 35.0
    Code-Davinci-002 (175B) 50.9 55.2 62.9 23.5 24.4 26.4
    Text-Davinci-003 (175B) 55.0 58.9 64.7 31.3 31.5 34.3

    Across all model families and parameter scales, UDR-retrieved demonstrations consistently outperform both BM25 and task-specific EPR retrievers. Furthermore, downstream performance scales positively with inference LM capacity when using UDR demonstrations.

  8. Knowl 8 — Zero-Shot Demonstration Retrieval on Unseen Datasets and Domains

    empirical result

    UDR generalizes to domains, dataset formats, and programming languages not seen during retriever training without fine-tuning. Experiments evaluated Twitter sentiment classification, Question-Answering NLI (QNLI), and code summarization for Ruby and JavaScript against universal dense (SBERT) and sparse (BM25) retrievers using GPT-Neo-2.7B as the inference LM.

    Retriever Twitter (Acc) QNLI (Acc) Ruby (BLEU-1) JavaScript (BLEU-1)
    BM25 50.0 54.1 9.2 12.7
    SBERT 51.6 53.7 8.7 15.9
    UDR 56.8 74.4 19.6 21.6

    UDR outperforms BM25 and SBERT by an average of approximately 10 points across these unseen datasets, demonstrating that the multi-task list-wise ranking objective teaches generalizable principles of demonstration quality.

  9. Knowl 9 — Robustness of In-Context Learning to Demonstration Permutation Order

    empirical result

    While random demonstration selection in few-shot in-context learning is known to cause severe performance fluctuations (varying up to tens of accuracy points depending on ordering), demonstrations retrieved by UDR are remarkably insensitive to permutation order.

    Ordering Strategy SST-2 TREC Reddit CommonGen
    Random-Order (Best of 10 seeds) 92.5 96.6 16.8 27.5
    Random-Order (Worst of 10 seeds) 92.0 96.2 16.2 26.6
    Descending-Order (Highest sim first) 92.2 96.6 16.2 27.0
    Ascending-Order (Lowest sim first) 92.4 96.6 16.3 27.3

    Across classification (SST-2, TREC) and generation tasks (Reddit, CommonGen), the performance gap across different orderings of UDR demonstrations remains within 1.0 point, showing that high-quality, query-relevant demonstrations stabilize in-context learning against prompt order sensitivity.

  10. Knowl 10 — Demonstration Quality versus Quantity in In-Context Learning

    empirical result

    Comparing in-context learning performance across varying numbers of retrieved demonstrations (K∈{1,2,3,4,5,6,7,8}K \in \{1, 2, 3, 4, 5, 6, 7, 8\}) revealed two key behavioral characteristics:

    1. Task Type Divergence: The quantity of demonstrations significantly impacts generation tasks (WikiAuto SARI and Java code summarization BLEU-1 continuously improve as demonstration count increases from 1 to 8), whereas classification tasks (Yelp and RTE accuracy) plateau rapidly after 2 to 3 demonstrations.
    2. Quality Dominates Quantity: UDR with only 2 demonstrations achieves equal or higher downstream performance than BM25 and EPR supplied with 8 demonstrations on Yelp, RTE, WikiAuto, and Java. This demonstrates that relevance and precision in demonstration selection outweigh raw demonstration volume.
  11. Knowl 11 — Limitations of Unified Demonstration Retriever

    limitation

    The UDR approach has three identified limitations:

    1. Encoder Initialization: UDR was evaluated only using dual BERT-base-uncased backbones (220M parameters) due to computational constraints; stronger pre-trained encoders such as RoBERTa or DeBERTa were not explored.
    2. Lack of Interpretability: UDR operates as a dense black-box neural retriever, offering limited transparency into the underlying linguistic or semantic principles that govern why specific examples serve as effective in-context prompts.
    3. Training vs. Inference Context Mismatch: During training, candidate demonstrations are scored individually and independently by the language model GG (pG(y∣zj,x)p_G(y \mid z_j, x)). However, during inference, the LM conditions on an ordered sequence of multiple demonstrations ([z1;z2;… ;zL;xtest][z_1; z_2; \dots; z_L; x_{\text{test}}]), leaving inter-demonstration cross-attention and combinatorial dependencies unmodeled during retriever training.

Coverage note — No substantial contributed material was omitted; the extraction covers UDR's dual-encoder instruction architecture, LM candidate scoring formulation, list-wise ranking objective, iterative mining algorithm, main cross-task results, ablation studies, cross-LM transferability, unseen domain zero-shot evaluation, ordering/quantity analyses, and limitations.

References

  1. 1.Sweta Agrawal, Chunting Zhou, Mike Lewis, Luke Zettlemoyer, and Marjan Ghazvininejad. 2022. In-context examples selection for machine translation. CoRR, abs/2212.02437.
  2. 2.Reinald Kim Amplayo, Arthur Brazinskas, Yoshi Suhara, Xiaolan Wang, and Bing Liu. 2022. Beyond opinion mining: Summarizing opinions of customer reviews. In SIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, July 11 - 15, 2022, pages 3447–3450. ACM.
  3. 3.Jacob Andreas, John Bufe, David Burkett, Charles Chen, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, Hao Fang, Alan Guo, David Hall, Kristin Hayes, Kellie Hill, Diana Ho, Wendy Iwaszuk, Smriti Jha, Dan Klein, Jayant Krishnamurthy, Theo Lanman, Percy Liang, Christo-pher H. Lin, Ilya Lintsbakh, Andy McGovern, Alek-sandr Nisnevich, Adam Pauls, Dmitrij Petters, Brent Read, Dan Roth, Subhro Roy, Jesse Rusak, Beth Short, Div Slomin, Ben Snyder, Stephon Striplin, Yu Su, Zachary Tellman, Sam Thomson, Andrei Vorobev, Izabela Witoszko, Jason Andrew Wolfe, Abby Wray, Yuchen Zhang, and Alexander Zotov. 2020. Task-oriented dialogue as dataflow synthesis. Trans. Assoc. Comput. Linguistics, 8:556–571.
  4. 4.Roy Bar-Haim, Ido Dagan, and Idan Szpektor. 2014. Benchmarking applied semantic inference: The PAS-CAL recognising textual entailment challenges. In Language, Culture, Computation. Computing - The-ory and Technology - Essays Dedicated to Yaacov Choueka on the Occasion of His 75th Birthday, Part I, volume 8001 of Lecture Notes in Computer Science, pages 409–424. Springer.
  5. 5.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big? In FAccT ’21: 2021 ACM Conference on Fairness, Accountability, and Transparency, Virtual Event / Toronto, Canada, March 3-10, 2021, pages 610–623. ACM.
  6. 6.Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  7. 7.Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. A large an-notated corpus for learning natural language infer-ence. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, EMNLP 2015, Lisbon, Portugal, September 17-21, 2015, pages 632–642. The Association for Computa-tional Linguistics.
  8. 8.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Ad-vances in Neural Information Processing Systems 33: Annual Conference on Neural Information Process-ing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  9. 9.Christopher J. C. Burges. 2010. From RankNet to Lamb-daRank to LambdaMART: An overview. Technical report, Microsoft Research.
  10. 10.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sas-try, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cum-mings, Matthias Plappert, Fotios Chantzis, Eliza-beth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021. Evaluat-ing large language models trained on code. CoRR, abs/2107.03374.
  11. 11.Yanda Chen, Chen Zhao, Zhou Yu, Kathleen R. McKe-own, and He He. 2022. On the relation between sen-sitivity and accuracy in in-context learning. CoRR, abs/2209.07661.
  12. 12.Qinyuan Cheng, Xiaogui Yang, Tianxiang Sun, Linyang Li, and Xipeng Qiu. 2023. Improving contrastive learning of sentence embeddings from AI feedback. CoRR, abs/2305.01918.
  13. 13.Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018. A discourse-aware attention model for abstractive summarization of long documents. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computa-tional Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 615–621, New Or-leans, Louisiana. Association for Computational Lin-guistics.
  14. 14.Alexis Conneau and Guillaume Lample. 2019. Cross-lingual language model pretraining. In Advances in Neural Information Processing Systems 32: An-nual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 7057–7067.
  15. 15.Rajarshi Das, Manzil Zaheer, Dung Thai, Ameya God-bole, Ethan Perez, Jay Yoon Lee, Lizhen Tan, Lazaros Polymenakos, and Andrew McCallum. 2021. Case-based reasoning for natural language queries over knowledge bases. In Proceedings of the 2021 Confer-ence on Empirical Methods in Natural Language Pro-cessing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 9594–9611. Association for Computational Linguis-tics.
  16. 16.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language under-standing. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech-nologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  17. 17.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui. 2022. A survey for in-context learning.
  18. 18.Ondˇrej Dušek, David M Howcroft, and Verena Rieser. 2019. Semantic Noise Matters for Neural Natural Language Generation. In Proceedings of the 12th International Conference on Natural Language Gen-eration (INLG 2019), pages 421–426, Tokyo, Japan.
  19. 19.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenRe-view.net.
  20. 20.Karl Moritz Hermann, Tomás Kociský, Edward Grefen-stette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015. Teaching machines to read and comprehend. In Advances in Neural Information Processing Systems 28: Annual Conference on Neu-ral Information Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 1693–1701.
  21. 21.Yushi Hu, Chia-Hsuan Lee, Tianbao Xie, Tao Yu, Noah A. Smith, and Mari Ostendorf. 2022. In-context learning for few-shot dialogue state tracking. CoRR, abs/2203.08568.
  22. 22.Lifu Huang, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. Cosmos QA: Machine reading comprehension with contextual commonsense rea-soning. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natu-ral Language Processing (EMNLP-IJCNLP), pages 2391–2401, Hong Kong, China. Association for Com-putational Linguistics.
  23. 23.Chao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong, and Wei Xu. 2020. Neural CRF model for sentence alignment in text simplification. In Proceed-ings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7943–7960, On-line. Association for Computational Linguistics.
  24. 24.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2021. Billion-scale similarity search with gpus. IEEE Trans. Big Data, 7(3):535–547.
  25. 25.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Nat-ural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6769–6781. Associa-tion for Computational Linguistics.
  26. 26.Byeongchang Kim, Hyunwoo Kim, and Gunhee Kim. 2019. Abstractive summarization of Reddit posts with multi-level memory networks. In Proceedings of the 2019 Conference of the North American Chap-ter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2519–2531, Minneapolis, Min-nesota. Association for Computational Linguistics.
  27. 27.Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N. Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick van Kleef, Sören Auer, and Christian Bizer. 2015. Dbpedia - A large-scale, multilingual knowledge base extracted from wikipedia. Semantic Web, 6(2):167–195.
  28. 28.Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021. MTOP: A comprehensive multilingual task-oriented semantic parsing benchmark. In Proceedings of the 16th Conference of the European Chapter of the Asso-ciation for Computational Linguistics: Main Volume, pages 2950–2962, Online. Association for Computa-tional Linguistics.
  29. 29.Xiaonan Li and Xipeng Qiu. 2023a. Finding sup-porting examples for in-context learning. CoRR, abs/2302.13539.
  30. 30.Xiaonan Li and Xipeng Qiu. 2023b. Mot: Pre-thinking and recalling enable chatgpt to self-improve with memory-of-thoughts.
  31. 31.Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020. CommonGen: A constrained text gen-eration challenge for generative commonsense rea-soning. In Findings of the Association for Computa-tional Linguistics: EMNLP 2020, pages 1823–1840, Online. Association for Computational Linguistics.
  32. 32.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2022. What makes good in-context examples for gpt-3? In Pro-ceedings of Deep Learning Inside Out: The 3rd Work-shop on Knowledge Extraction and Integration for Deep Learning Architectures, DeeLIO@ACL 2022, Dublin, Ireland and Online, May 27, 2022, pages 100–114. Association for Computational Linguistics.
  33. 33.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Man-dar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  34. 34.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenRe-view.net.
  35. 35.Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Li-dong Zhou, Linjun Shou, Long Zhou, Michele Tu-fano, Ming Gong, Ming Zhou, Nan Duan, Neel Sun-daresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. Codexglue: A machine learning bench-mark dataset for code understanding and generation. In Proceedings of the Neural Information Process-ing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual.
  36. 36.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Compu-tational Linguistics (Volume 1: Long Papers), pages 8086–8098, Dublin, Ireland. Association for Compu-tational Linguistics.
  37. 37.Julian J. McAuley and Jure Leskovec. 2013. Hidden factors and hidden topics: understanding rating di-mensions with review text. In Seventh ACM Confer-ence on Recommender Systems, RecSys ’13, Hong Kong, China, October 12-16, 2013, pages 165–172. ACM.
  38. 38.Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Noisy channel language model prompting for few-shot text classification. In Proceedings of the 60th Annual Meeting of the As-sociation for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 5316–5330. Association for Computa-tional Linguistics.
  39. 39.Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen. 2016. A corpus and cloze evaluation for deeper understanding of commonsense stories. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 839–849, San Diego, California. Association for Computational Linguis-tics.
  40. 40.Ibrahim Naji. 2012. TSATC: Twitter Sentiment Analy-sis Training Corpus. In thinknook.
  41. 41.Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xi-angru Tang, Aadit Vyas, Neha Verma, Pranav Kr-ishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani. 2021. DART: Open-domain structured data record to text generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computa-tional Linguistics: Human Language Technologies, pages 432–447, Online. Association for Computa-tional Linguistics.
  42. 42.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Car-roll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. CoRR, abs/2203.02155.
  43. 43.Bo Pang and Lillian Lee. 2004. A sentimental educa-tion: Sentiment analysis using subjectivity summa-rization based on minimum cuts. In Proceedings of the 42nd Annual Meeting of the Association for Com-putational Linguistics, 21-26 July, 2004, Barcelona, Spain, pages 271–278. ACL.
  44. 44.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In ACL 2005, 43rd Annual Meeting of the Association for Computational Lin-guistics, Proceedings of the Conference, 25-30 June 2005, University of Michigan, USA, pages 115–124. The Association for Computer Linguistics.
  45. 45.Gabriel Poesia, Alex Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gul-wani. 2022. Synchromesh: Reliable code generation from pre-trained language models. In The Tenth In-ternational Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. Open-Review.net.
  46. 46.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2018. Language models are unsupervised multitask learners.
  47. 47.Nils Reimers and Iryna Gurevych. 2019a. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natu-ral Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Com-putational Linguistics.
  48. 48.Nils Reimers and Iryna Gurevych. 2019b. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Nat-ural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3980–3990. Association for Computational Linguistics.
  49. 49.Stephen E. Robertson and Hugo Zaragoza. 2009. The probabilistic relevance framework: BM25 and be-yond. Found. Trends Inf. Retr., 3(4):333–389.
  50. 50.Melissa Roemmele, Cosmin Adrian Bejan, and An-drew S Gordon. 2011. Choice of plausible alter-natives: An evaluation of commonsense causal rea-soning. In 2011 AAAI Spring Symposium Series.
  51. 51.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning to retrieve prompts for in-context learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech-nologies, NAACL 2022, Seattle, WA, United States, July 10-15, 2022, pages 2655–2671. Association for Computational Linguistics.
  52. 52.Peng Shi, Rui Zhang, He Bai, and Jimmy Lin. 2022. XRICL: cross-lingual retrieval-augmented in-context learning for cross-lingual text-to-sql semantic pars-ing. CoRR, abs/2210.13693.
  53. 53.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013. Recursive deep mod-els for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP 2013, 18-21 October 2013, Grand Hyatt Seattle, Seattle, Washington, USA, A meeting of SIG-DAT, a Special Interest Group of the ACL, pages 1631–1642. ACL.
  54. 54.Hongjin Su, Weijia Shi, Jungo Kasai, Yushi Hu, Yizhong Wang, Mari Ostendorf, Wen tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. One embedder, any task: Instruction-finetuned text embeddings.
  55. 55.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Pro-cessing Systems, volume 30. Curran Associates, Inc.
  56. 56.Ellen M. Voorhees and Dawn M. Tice. 2000. Building a question answering test collection. In SIGIR 2000: Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, July 24-28, 2000, Athens, Greece, pages 200–207. ACM.
  57. 57.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a. GLUE: A multi-task benchmark and analysis plat-form for natural language understanding. In 7th In-ternational Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  58. 58.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Lan-guage Model. https://github.com/kingoflolz/mesh-transformer-jax.
  59. 59.Cunxiang Wang, Shuailong Liang, Yue Zhang, Xiao-nan Li, and Tian Gao. 2019b. Does it make sense? and why? A pilot study for sense making and ex-planation. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Vol-ume 1: Long Papers, pages 4020–4026. Association for Computational Linguistics.
  60. 60.Adina Williams, Nikita Nangia, and Samuel R. Bow-man. 2018. A broad-coverage challenge corpus for sentence understanding through inference. In Pro-ceedings of the 2018 Conference of the North Amer-ican Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics.
  61. 61.Tomer Wolfson, Mor Geva, Ankit Gupta, Yoav Gold-berg, Matt Gardner, Daniel Deutch, and Jonathan Berant. 2020. Break it down: A question understand-ing benchmark. Trans. Assoc. Comput. Linguistics, 8:183–198.
  62. 62.Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N. Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate nearest neigh-bor negative contrastive learning for dense text re-trieval. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  63. 63.Benfeng Xu, Quan Wang, Zhendong Mao, Yajuan Lyu, Qiaoqiao She, and Yongdong Zhang. 2023. knn prompting: Beyond-context learning with calibration-free nearest neighbor inference.
  64. 64.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text clas-sification. In Advances in Neural Information Pro-cessing Systems 28: Annual Conference on Neural In-formation Processing Systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pages 649–657.
  65. 65.Yiming Zhang, Shi Feng, and Chenhao Tan. 2022. Ac-tive example selection for in-context learning. CoRR, abs/2211.04486.

Citation

MLA
Li, X., et al. “Unified Demonstration Retriever for In-Context Learning”. arXiv, 2023, http://arxiv.org/abs/2305.04320v2.
APA
Li, X., Lv, K., Yan, H., Lin, T., Zhu, W., Ni, Y., Xie, G., Wang, X., & Qiu, X. (2023). Unified Demonstration Retriever for In-Context Learning. arXiv. http://arxiv.org/abs/2305.04320v2
Chicago
Li, X., K. Lv, H. Yan, et al. 2023. “Unified Demonstration Retriever for In-Context Learning”. arXiv. http://arxiv.org/abs/2305.04320v2.
Harvard
Li, X. et al. (2023) “Unified Demonstration Retriever for In-Context Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.04320v2.
Vancouver
1. Li X, Lv K, Yan H, Lin T, Zhu W, Ni Y, Xie G, Wang X, Qiu X (2023) Unified Demonstration Retriever for In-Context Learning. arXiv

BibTeX

@article{li2023unified,
  title = {Unified Demonstration Retriever for In-Context Learning},
  author = {Li, Xiaonan and Lv, Kai and Yan, Hang and Lin, Tianyang and Zhu, Wei and Ni, Yuan and Xie, Guotong and Wang, Xiaoling and Qiu, Xipeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.04320v2},
  eprint = {2305.04320}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/