Data Curation Alone Can Stabilize In-context Learning

Ting-Yun ChangRobin Jia

article2023ACL69 citations

Demonstrates that selecting high-value training subsets through individual example scoring significantly reduces performance variance and increases accuracy in in-context learning without requiring dynamic prompt retrieval or model calibration.

Listen

In-context learning allows large language models to tackle new classification tasks simply by receiving a few reference examples inside the prompt, without modifying model parameters. While this approach avoids costly retraining, standard in-context learning is notoriously fragile. Minor adjustments to the selected prompt examples, their order, or prompt formatting frequently cause substantial fluctuations in accuracy. Organizations deploying language models therefore face significant performance variability and operational risk.

The article demonstrates that carefully curating a small subset of high-quality training examples can resolve this instability without requiring complex inference-time retrieval systems, prompt calibration, or model fine-tuning. The primary objective is to evaluate whether automated data valuation techniques can extract small pools of consistently effective examples from larger training datasets to stabilize and improve few-shot performance.

To identify stable subsets, the authors evaluated two scoring methods on large language models including GPT-J (6 billion parameters) and OPT (13 billion parameters) across five classification benchmarks. The first method, conditional accuracy, scores each training example by its average validation accuracy when combined with other random examples. The second method, data modeling, fits linear regression models to predict language model output margins based on the presence and position of specific training examples. The highest-scoring examples per class are selected to form a curated subset of 20 examples. Prompts sampled from these curated pools were tested against baselines using uncurated datasets, calibration, one-shot scoring, and top-performing prompt unions.

The results show that data curation alone substantially boosts reliability and accuracy. Across all tasks and models, prompts randomly drawn from subsets curated via conditional accuracy and data modeling improved average accuracy by 7.7% and 6.3%, respectively, over sampling from the full, uncurated training set. These curated subsets raised worst-case performance floors and reduced variance across prompt combinations. When evaluated on out-of-distribution target datasets, curated prompts maintained superior performance, demonstrating that the selected examples teach task-level definitions rather than overfitting to specific distributions. Even in an unlabeled setup—where inputs were paired with all candidate labels—the curated selection method improved performance by 5.7% over the fully gold-labeled baseline. Surprisingly, analyses showed that stable examples are not defined by high text diversity or low perplexity; instead, effective examples cluster tightly in representation space.

These findings indicate that in-context learning sensitivity stems largely from low-quality data rather than an inherent flaw in few-shot prompting. For practitioners, establishing a fixed pool of curated examples removes the operational complexity, latency, and infrastructure costs associated with dynamic example-retrieval pipelines. Furthermore, the ability to curate high-performing prompts without gold labels lowers data annotation costs. Curated prompts also proved more compact, with 4 curated examples outperforming 16 to 24 uncurated examples in several tasks, conserving prompt context limits.

Organizations leveraging in-context learning should prioritize offline data curation to build vetted example pools for deployment prompts. If resources permit, teams should evaluate conditional accuracy scoring on a representative validation set before deployment. However, decision-makers should account for significant upfront computational costs: profiling example effectiveness required running tens of thousands of validation prompts, consuming hundreds of GPU hours in larger configurations. Additional exploration is recommended to test more efficient search methods during data collection and to determine if these findings transfer to generative tasks and models exceeding 100 billion parameters.

Chang et al (2023).pdf
Cover for Data Curation Alone Can Stabilize In-context Learning

Abstract

In-context learning (ICL) enables large language models (LLMs) to perform new tasks by prompting them with a sequence of training examples. However, it is known that ICL is very sensitive to the choice of training examples: randomly sampling examples from a training set leads to high variance in performance. In this paper, we show that carefully curating a subset of training data greatly stabilizes ICL performance without any other changes to the ICL algorithm (e.g., prompt retrieval or calibration). We introduce two methods to choose training subsets—both score training examples individually, then select the highest-scoring ones. CONDACC scores a training example by its average dev-set ICL accuracy when combined with random training examples, while DATAMODELS learns linear regressors that estimate how the presence of each training example influences LLM outputs. Across five tasks and two LLMs, sampling from stable subsets selected by CONDACC and DATAMODELS improves average accuracy over sampling from the entire training set by 7.7% and 6.3%, respectively. Surprisingly, the stable subset examples are not especially diverse in content or low in perplexity, in contrast with other work suggesting that diversity and perplexity are important when prompting LLMs.

Table of Contents

  • 1 Introduction
  • 2 Problem Setups
  • 3 Methods
  • 3.1 CondAcc
  • 3.2 Datamodels
  • 3.3 Select Training Examples
  • 4 Experiment
  • 4.1 Setups
  • 4.2 Evaluation and Baselines
  • 5 Results
  • 5.1 Main Results
  • 5.2 Single-Label Prompts
  • 5.3 Out-of-Distribution Tasks
  • 6 Analysis
  • 6.1 Sequence Length and Perplexity
  • 6.2 Diversity
  • 6.3 Do LLMs find the same stable examples?
  • 7 Discussion and Related Work
  • 7.1 Prompt Retrieval
  • 7.2 The influence of in-context labels
  • 7.3 Data Valuation
  • 8 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Connection with Data Shapley
  • A.2 Relation to Prompt Tuning
  • A.3 Evaluating Datamodels
  • A.4 Training Details of Datamodels
  • A.5 Implementation Details
  • A.6 More Experiments
  • A.7 Why not Instruction-Finetuned LLMs?
  • A.8 Larger Number of Shots

Knowls

  1. Knowl 1 — Stable in-context learning through task-level training subsets

    definition

    The paper defines a stable training subset EE as a subset of a larger training set DtrD_{tr}, with ∣E∣>K|E|>K, from which randomly sampled, ordered prompts of KK examples tend to produce both higher average and higher worst-case test accuracy than prompts sampled from the full training set. The proposed approach keeps the task template and the language model’s inference procedure fixed: it estimates each training example’s value using development-set in-context learning (ICL), then retains the highest-scoring examples within each class. At test time, prompts are sampled from this curated subset rather than retrieved separately for each test input.

  2. Knowl 2 — CONDACC scores examples by conditional development accuracy

    model/method

    For a training example ei=(xi,yi)e_i=(x_i,y_i), CONDACC estimates how accurately prompts perform when they contain that example. Let DICLD_{ICL} be a collection of prompts, each formed from KK randomly sampled training examples; let Acc(Z)Acc(Z) be the ICL accuracy on the development set for prompt ZZ. The score is

    sca(i)=EZ∼DICL[Acc(Z)∣ei∈Z].s_{ca}(i)=\mathbb{E}_{Z\sim D_{ICL}}[Acc(Z)\mid e_i\in Z].

    The collection is constructed so that each example appears in multiple prompts, in different orders and with different accompanying examples. The score therefore reflects average performance conditional on an example’s presence, without distinguishing its position in the prompt. The method selects the highest-scoring examples per class. The paper also establishes that this score ranks examples in the same order as a Data Shapley value defined for KK-example ICL prompts.

  3. Knowl 3 — DATAMODELS estimates example effects at each prompt position

    model/method

    DATAMODELS fits a linear predictor of an LLM’s output margin for each labeled development example. For a development input xˉ\bar{x} with correct label yˉ\bar{y} and a prompt Z=[z1,…,zK]Z=[z_1,\ldots,z_K], define the margin f(xˉ;Z)=o(yˉ∣Z,xˉ)−max⁡y∈C, y≠yˉo(y∣Z,xˉ)f(\bar{x};Z)=o(\bar{y}\mid Z,\bar{x})-\max_{y\in C,\,y\ne\bar{y}}o(y\mid Z,\bar{x}), where CC is the label set and o(y∣Z,xˉ)o(y\mid Z,\bar{x}) is the LLM logit for label yy. For each development example, the datamodel predicts this margin as

    gw,b(xˉ;Z)=∑j=1Kw(id⁡(zj),j)+b,g_{w,b}(\bar{x};Z)=\sum_{j=1}^{K}w(\operatorname{id}(z_j),j)+b,

    where id⁡(zj)\operatorname{id}(z_j) is the identity of the training example at prompt position jj, w(i,j)w(i,j) is its learned real-valued weight at that position, and bb is a real-valued bias. The parameters are fitted by mean-squared error over observed prompt–margin pairs. The example score sums positive weights across development examples and prompt positions: sdm(i)=∑(xˉ,yˉ)∈Ddev∑j=1K1{wxˉ(i,j)>0}s_{dm}(i)=\sum_{(\bar{x},\bar{y})\in D_{dev}}\sum_{j=1}^{K}\mathbf{1}\{w_{\bar{x}}(i,j)>0\}. Thus, an example scores highly when its presence is associated with a positive correct-class margin for many development examples, regardless of position. Training uses a shared first phase followed by separate fitting for each in-context label pattern; weights from these models are aggregated for scoring. On unseen prompts with unseen example combinations, the learned datamodels achieved Pearson correlations of 0.840–0.962 and mean absolute errors of 0.088–0.340 across the reported tasks and GPTJ-6B/OPT-13B.

  4. Knowl 4 — Curated subsets improve accuracy across five classification tasks

    empirical result

    The main evaluation used SST-2, BoolQ, Subj, Scicite, and AGNews, with class-balanced training and test sets of 1,000 examples each and 50 development examples per class. The selected subset contained 20 examples; prompts used 4 shots for binary tasks, 3 for Scicite, and 4 for AGNews. The authors evaluated 50 randomly sampled prompts per method on GPTJ-6B and OPT-13B, reporting mean accuracy over prompts and tasks. Average accuracies across the five tasks were:

    • GPTJ-6B: ALL 65.2%; TOPPROMPTS-5 69.1%; CONDACC 72.4%; DATAMODELS 72.4%.
    • OPT-13B: ALL 63.8%; TOPPROMPTS-5 66.1%; CONDACC 72.0%; DATAMODELS 69.1%.

    Across the two models, CONDACC improved average accuracy over ALL by 7.7 percentage points, and DATAMODELS by 6.3 points. The paper reports that the proposed methods also generally improved worst-case accuracy and reduced variability across prompts; they performed better than calibration in 8 of 10 model–task setups. TOPPROMPTS-5, which selects examples occurring in the five best development-set prompts, was the strongest baseline overall, but remained below the proposed methods on average.

  5. Knowl 5 — CONDACC works in the unlabeled training-data setup, even with some incorrect labels

    empirical result

    In the unlabeled setup, each training input was paired with every possible class label, and CONDACC selected among these input–label pairs using labeled development data. UN-CONDACC achieved average accuracy across the five tasks of 71.3% with GPTJ-6B and 69.2% with OPT-13B, compared with 65.2% and 63.8% for ALL using gold-labeled examples, and 61.8% and 58.3% for UN-ALL. The selected examples were not all correctly labeled: out of 20 selected examples, the numbers with gold labels for GPTJ-6B/OPT-13B were SST-2 20/19, BoolQ 11/10, Subj 20/20, Scicite 16/11, and AGNews 18/13. The majority-guess expectations were 10, 10, 10, 6.6, and 5, respectively. Correcting the mislabeled BoolQ examples reduced average and worst accuracy by 1.9 and 4.5 points for GPTJ-6B, and by 0.4 and 5.7 points for OPT-13B. The results show that selection can be effective without gold labels for every prompt example, while not establishing that incorrect labels are generally beneficial.

  6. Knowl 6 — Curated examples retain task signal in prompts containing only one label

    empirical result

    The authors tested binary classification prompts containing four examples all bearing the same label, evaluating SST-2 and BoolQ with 50 sampled prompts. ALL and TOPPROMPTS were generally near the 50% majority-guess baseline, consistent with the LLM being biased toward predicting the sole label shown in the prompt. Prompts sampled from CONDACC and DATAMODELS subsets exceeded majority guessing in every reported model–task–label condition. DATAMODELS achieved 72.8% on GPTJ-6B with all-negative SST-2 examples and 68.4% with all-positive examples; on BoolQ it achieved 61.7% with all-negative and 56.9% with all-positive examples. With OPT-13B, the best proposed-method accuracies for those four conditions were 66.3%, 84.5%, 69.2%, and 60.1%, respectively. These results support the authors’ interpretation that selected examples can convey task-level information even when the prompt does not exhibit the full label space.

  7. Knowl 7 — Curated source-task examples generalize to out-of-distribution test data

    empirical result

    The authors selected examples on SST-2 and evaluated on IMDb, and selected examples on BoolQ and evaluated on the BoolQ Contrast Set. Each result is the mean accuracy, standard deviation, and minimum accuracy over 50 sampled prompts. For GPTJ-6B on IMDb, ALL scored 86.5% / 5.7 / 63.6%, TOPPROMPTS 87.2% / 5.2 / 63.0%, CONDACC 90.5% / 1.8 / 84.8%, and DATAMODELS 91.6% / 1.5 / 84.0%; on BoolQ Contrast Set, the corresponding scores were 56.6% / 3.0 / 50.1%, 56.7% / 2.6 / 49.9%, 58.9% / 1.7 / 54.6%, and 57.6% / 1.9 / 54.0%. For OPT-13B on IMDb, the scores were 79.2% / 12.1 / 50.1%, 80.5% / 14.0 / 50.8%, 83.5% / 10.8 / 54.6%, and 84.1% / 9.3 / 58.9%; on BoolQ Contrast Set, they were 59.8% / 2.9 / 51.6%, 60.3% / 3.5 / 51.0%, 60.1% / 2.1 / 56.7%, and 60.6% / 3.3 / 54.3%. The curated methods outperformed the baselines on IMDb for both models and generally raised minimum accuracy, indicating that their benefits were not limited to the source test distribution.

  8. Knowl 8 — Selected examples are not characterized by high diversity, high perplexity, or unusual length

    empirical result

    The paper compared high- and low-scoring CONDACC examples and subsets against random subsets. Across tasks and models, example score showed little correlation with input length; the selected good examples did not occupy the extreme long-length tail, although exceptionally long examples showed a slight tendency to hurt ICL. Input perplexity, measured under the same LLM used for ICL, also showed no correlation with example score, and good examples were not unusually high-perplexity outliers. For subset diversity, the authors compared 5,000 random subsets with the selected subsets using DIV-I, a raw-text diversity measure, and DIV-F, a SentBERT-embedding diversity measure. Good subsets were no more diverse than random subsets and sometimes had low diversity. In datamodel-embedding visualizations, CONDACC and DATAMODELS selections formed tightly clustered groups and often selected similar examples. These analyses do not support using length, perplexity, or diversity alone as a proxy for the methods’ example scores.

  9. Knowl 9 — Stable examples transfer inconsistently between language models

    empirical result

    The authors compared CONDACC example scores and selected-subset overlap across models. For GPTJ-6B versus OPT-13B, score correlations / numbers of shared selected examples were SST-2 0.15 / 4, AGNews 0.46 / 1, BoolQ 0.41 / 2, Subj −0.03 / 2, and Scicite 0.02 / 2. For GPTJ-6B versus OPT-6.7B, the corresponding values were SST-2 0.27 / 3 and AGNews 0.08 / 1; for OPT-6.7B versus OPT-13B, they were SST-2 0.76 / 3 and AGNews 0.42 / 1. Thus, stability scores and selected examples were not generally transferable across different model families, though OPT-family models showed stronger agreement in some cases. As a specific exception, the four SST-2 examples shared by GPTJ-6B and OPT-13B produced high accuracy over all 24 orderings: 88.6% ± 3.7 for GPTJ-6B, 89.8% ± 3.0 for OPT-6.7B, and 87.3% ± 5.2 for OPT-13B.

  10. Knowl 10 — Computational cost and evaluation scope limit the demonstrated method

    limitation

    Both selection methods require collecting LLM ICL outcomes over many development-set prompts before example scores can be computed. The authors report that their prompt collection contained about 50,000 prompts per setup and that a preliminary study suggested at least 10,000 randomly sampled prompts were needed for good performance. The most expensive reported setup, OPT-13B on BoolQ, required more than 500 GPU-hours on one RTX A6000; inference memory use is also substantial. The experiments cover classification tasks, models no larger than 13 billion parameters, and fixed shot counts and prompt templates. The paper therefore does not establish whether the methods or findings extend to generation tasks, substantially larger models, or other prompting configurations.

Coverage note — The larger-shot MAXSHOT comparison and the full per-task appendix result matrices are omitted as secondary to the main curated-subset comparisons; the primary tasks, evaluation conditions, and headline results are retained.

References

  1. 1.Akari Asai, Mohammadreza Salehi, Matthew E. Peters, and Hannaneh Hajishirzi. 2022. Attentional mixtures of soft prompt tuning for parameter-efficient multitask knowledge sharing. ArXiv, abs/2205.11961.
  2. 2.Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Xiangru Tang, Mike Tian-Jian Jiang, and Alexander M. Rush. 2022. Promptsource: An integrated development environment and repository for natural language prompts.
  3. 3.Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman. 2021. GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.
  5. 5.Yanda Chen, Chen Zhao, Zhou Yu, Kathleen McKeown, and He He. 2022. On the relation between sensitivity and accuracy in in-context learning. arXiv preprint arXiv:2209.07661.
  6. 6.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  8. 8.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044.
  9. 9.Arman Cohan, Waleed Ammar, Madeleine van Zuylen, and Field Cady. 2019. Structural scaffolds for citation intent classification in scientific publications. In North American Chapter of the Association for Computational Linguistics.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  11. 11.Tianyu Gao, Adam Fisch, and Danqi Chen. 2020. Making pre-trained language models better few-shot learners. arXiv preprint arXiv:2012.15723.
  12. 12.Matt Gardner, Yoav Artzi, Victoria Basmova, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, et al. 2020. Evaluating models’ local decision boundaries via contrast sets. arXiv preprint arXiv:2004.02709.
  13. 13.Amirata Ghorbani and James Zou. 2019. Data shapley: Equitable valuation of data for machine learning. In International Conference on Machine Learning, pages 2242–2251. PMLR.
  14. 14.Hila Gonen, Srini Iyer, Terra Blevins, Noah A Smith, and Luke Zettlemoyer. 2022. Demystifying prompts in language models via perplexity estimation. arXiv preprint arXiv:2212.04037.
  15. 15.Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry. 2022. Datamodels: Predicting predictions from training data. arXiv preprint arXiv:2202.00622.
  16. 16.Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Dániel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, et al. 2022. Opt-iml: Scaling language model instruction meta learning through the lens of generalization. arXiv preprint arXiv:2212.12017.
  17. 17.Junyeob Kim, Hyuhng Joon Kim, Hyunsoo Cho, Hwiyeol Jo, Sang-Woo Lee, Sang goo Lee, Kang Min Yoo, and Taeuk Kim. 2022. Ground-truth labels matter: A deeper look into input-label demonstrations. ArXiv, abs/2205.12685.
  18. 18.Pang Wei Koh and Percy Liang. 2017. Understanding black-box predictions via influence functions. In International conference on machine learning, pages 1885–1894. PMLR.
  19. 19.Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The power of scale for parameter-efficient prompt tuning. arXiv preprint arXiv:2104.08691.
  20. 20.Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, and Weizhu Chen. 2021. What makes good in-context examples for gpt-3? arXiv preprint arXiv:2101.06804.
  21. 21.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2021. Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Annual Meeting of the Association for Computational Linguistics.
  22. 22.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, pages 142–150, Portland, Oregon, USA. Association for Computational Linguistics.
  23. 23.Leland McInnes, John Healy, and James Melville. 2018. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426.
  24. 24.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837.
  25. 25.Tai Nguyen and Eric Wong. 2023. In-context example selection with influences. arXiv preprint arXiv:2302.11042.
  26. 26.Bo Pang and Lillian Lee. 2004. A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts. In Annual Meeting of the Association for Computational Linguistics.
  27. 27.Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021. True few-shot learning with language models. Advances in Neural Information Processing Systems, 34:11054–11070.
  28. 28.Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. 2020. Estimating training data influence by tracing gradient descent. Advances in Neural Information Processing Systems, 33:19920–19930.
  29. 29.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446.
  30. 30.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. ArXiv, abs/1908.10084.
  31. 31.Stephen Robertson, Hugo Zaragoza, et al. 2009. The probabilistic relevance framework: Bm25 and beyond. Foundations and Trends® in Information Retrieval, 3(4):333–389.
  32. 32.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2021. Learning to retrieve prompts for in-context learning. arXiv preprint arXiv:2112.08633.
  33. 33.Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. 2021. Multitask prompted training enables zero-shot task generalization. arXiv preprint arXiv:2110.08207.
  34. 34.Timo Schick and Hinrich Schütze. 2020. Exploiting cloze questions for few shot text classification and natural language inference. arXiv preprint arXiv:2001.07676.
  35. 35.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. 2022. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990.
  36. 36.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.
  37. 37.Hongjin Su, Jungo Kasai, Chen Henry Wu, Weijia Shi, Tianlu Wang, Jiayi Xin, Rui Zhang, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, et al. 2022. Selective annotation makes language models better few-shot learners. arXiv preprint arXiv:2209.01975.
  38. 38.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. 2022. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239.
  39. 39.Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model. https://github.com/kingoflolz/mesh-transformer-jax.
  40. 40.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al. 2022. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682.
  41. 41.Xi Ye, Srini Iyer, Asli Celikyilmaz, Ves Stoyanov, Greg Durrett, and Ramakanth Pasunuru. 2022. Complementary explanations for effective in-context learning. ArXiv, abs/2211.13892.
  42. 42.Jinsung Yoon, Sercan Arik, and Tomas Pfister. 2020. Data valuation using reinforcement learning. In International Conference on Machine Learning, pages 10842–10851. PMLR.
  43. 43.Michelle Yuan, Hsuan-Tien Lin, and Jordan Boyd-Graber. 2020. Cold-start active learning through self-supervised language modeling. arXiv preprint arXiv:2010.09535.
  44. 44.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022a. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  45. 45.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28.
  46. 46.Yiming Zhang, Shi Feng, and Chenhao Tan. 2022b. Active example selection for in-context learning. ArXiv, abs/2211.04486.
  47. 47.Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In International Conference on Machine Learning, pages 12697–12706. PMLR.
  48. 48.Fedor Zhdanov. 2019. Diverse mini-batch active learning. arXiv preprint arXiv:1901.05954.

Citation

MLA
Chang, T.-Y., and R. Jia. “Data Curation Alone Can Stabilize In-context Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 8123–44, https://doi.org/10.18653/v1/2023.acl-long.452.
APA
Chang, T.-Y., & Jia, R. (2023). Data Curation Alone Can Stabilize In-context Learning. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8123–8144. https://doi.org/10.18653/v1/2023.acl-long.452
Chicago
Chang, T.-Y., and R. Jia. 2023. “Data Curation Alone Can Stabilize In-context Learning”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8123–44. https://doi.org/10.18653/v1/2023.acl-long.452.
Harvard
Chang, T.-Y. and Jia, R. (2023) “Data Curation Alone Can Stabilize In-context Learning”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8123–8144. Available at: https://doi.org/10.18653/v1/2023.acl-long.452.
Vancouver
1. Chang T-Y, Jia R (2023) Data Curation Alone Can Stabilize In-context Learning. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8123–8144

BibTeX

@inproceedings{chang-jia-2023-data,
    title = "Data Curation Alone Can Stabilize In-context Learning",
    author = "Chang, Ting-Yun  and
      Jia, Robin",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.452/",
    doi = "10.18653/v1/2023.acl-long.452",
    pages = "8123--8144"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/