Understanding In-Context Learning via Supportive Pretraining Data

Xiaochuang HanDaniel SimigTodor MihaylovYulia TsvetkovAsli CelikyilmazTianlu Wang

article2023ACL52 citations

Reveals that in-context learning in large language models is driven by specific, challenging pretraining instances rich in long-tail tokens and difficult long-range contexts rather than domain-relevant text, providing actionable criteria to guide future pretraining data selection.

Listen

Large language models frequently perform downstream tasks through in-context learning, an operational capability where models infer solutions from just a few demonstrated examples provided in the prompt without updating their underlying parameters. Understanding where this ability originates during pretraining is essential for building more reliable systems and optimizing the massive datasets used to train them.

The article investigates how pretraining data drives in-context learning capabilities. Specifically, it identifies which specific pretraining instances support few-shot performance on downstream tasks and analyzes the key statistical and informational properties that distinguish supportive data from standard pretraining text.

The researchers evaluated the open-source OPT-6.7B model across six diverse natural language classification tasks selected from a broader instruction benchmark. Using an iterative, gradient-based search method called ORCA-ICL, the team scanned a subset of 2.5 million pretraining documents (comprising roughly 5 billion tokens) to locate training instances that share gradient similarity with task demonstration data. The team then performed targeted, one-pass continued pretraining on these small identified subsets (under 2,000 instances) to verify their impact, benchmarking the resulting models against both the original baseline and control models trained on randomly sampled pretraining data.

The evaluation revealed three key findings regarding supportive data. First, continued pretraining on the small supportive subset significantly boosted few-shot accuracy by up to 18% compared to random training subsets, while having no positive effect on zero-shot performance without demonstrations. Second, domain analysis using distributional text similarity metrics showed that supportive pretraining instances share no greater domain relevance to downstream tasks than random text, indicating that in-context learning relies on domain-invariant reasoning mechanisms rather than specific topic knowledge. Third, structural analysis demonstrated that supportive documents feature a higher concentration of rare, long-tail vocabulary terms and present lower information gain from extended context. This indicates that supportive instances are intrinsically challenging texts that force the model to separate relevant cues from long-range confounding context.

These findings indicate that in-context learning functions as an abstract meta-capability shaped by training on complex, linguistically rich text rather than simple exposure to task-specific subject matter. For AI practitioners and engineering leaders, this shifts data strategy away from naive domain-matching toward curating conceptually challenging training material that emphasizes rare terms and difficult context-filtering scenarios. Such targeted data curation has the potential to substantially reduce overall pretraining volume and compute costs while improving prompt-based model performance.

Organizations developing or refining foundation models should pilot pretraining data curation pipelines that screen for lower token-frequency concentration and higher contextual complexity. However, because the current search process requires significant computation—taking approximately one week across 32 enterprise GPUs per source task—practitioners should prioritize exploring more efficient gradient approximations, generation-based synthetic data methods, or applying these filtering criteria during upstream data acquisition.

Confidence in these findings is strong for autoregressive language models performing classification tasks under few-shot prompting. Readers should nevertheless exercise caution regarding generalizability, as the study focused exclusively on classification benchmarks, analyzed data relative to a finalized model checkpoint rather than early training phases, and relied primarily on correlational data properties that require further causal validation before establishing definitive pretraining guidelines.

arXiv: 2306.15091
Cover for Understanding In-Context Learning via Supportive Pretraining Data

Abstract

In-context learning (ICL) improves language models’ performance on a variety of NLP tasks by simply demonstrating a handful of examples at inference time. It is not well understood why ICL ability emerges, as the model has never been specifically trained on such demonstrations. Unlike prior work that explores implicit mechanisms behind ICL, we study ICL via investigating the pretraining data. Specifically, we first adapt an iterative, gradient-based approach to find a small subset of pretraining data that supports ICL. We observe that a continued pretraining on this small subset significantly improves the model’s ICL ability, by up to 18%. We then compare the supportive subset contrastively with random subsets of pretraining data and discover: (1) The supportive pretraining data to ICL do not have a higher domain relevance to downstream tasks. (2) The supportive pretraining data have a higher mass of rarely occurring, long-tail tokens. (3) The supportive pretraining data are challenging examples where the information gain from long-range context is below average, indicating learning to incorporate difficult long-range context encourages ICL. Our work takes a first step towards understanding ICL via analyzing instance-level pretraining data. Our insights have a potential to enhance the ICL ability of language models by actively guiding the construction of pretraining data in the future.

Table of Contents

  • 1 Introduction
  • 2 Finding supportive pretraining data for in-context learning
  • 2.1 Methodology
  • 2.2 Setup
  • 2.3 Results
  • 3 Analyzing supportive pretraining data for in-context learning
  • 3.1 Domain relevance
  • 3.2 Token frequency distribution
  • 3.3 Information gain from long-range context
  • 3.4 Future work
  • 4 Related Work
  • 5 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Qualitative examples

Knowls

  1. Knowl 1 — ORCA-ICL selects pretraining instances by gradient alignment with in-context task loss

    algorithm

    ORCA-ICL adapts iterative gradient-based data selection to identify pretraining examples whose training gradients align with the gradient that would reduce a model’s in-context learning (ICL) loss. Let θ0\theta_0 be the original autoregressive language model, ww a pretraining instance, and (x,y)(x,y) a downstream task example, with demonstrations included in the input context when calculating the ICL loss. The pretraining loss LθPT(w)L^{\mathrm{PT}}_\theta(w) is the negative log-likelihood of the tokens in ww; the ICL loss LθICL(x,y)L^{\mathrm{ICL}}_\theta(x,y) is the negative log-likelihood of the correct answer yy given task input xx and its demonstrations. For each candidate ww, the method ranks the cosine similarity between its pretraining-loss gradient at θ0\theta_0 and the gradient of the summed ICL loss over the guidance examples. It selects the top-scoring instances, then repeats selection using an updated guidance model while keeping the pretraining-instance gradients anchored at θ0\theta_0.

    Input: Pretrained model θ0\theta_0, pretraining instances DPTD_{PT}, ICL guidance examples DtaskD_{task}, iterations MM, instances per iteration kk
    Output: Supportive pretraining set SS
    Initialize S0=∅S_0 = \varnothing
    For i=1,…,Mi = 1, \ldots, M:
        If i=1i = 1, set guidance model g=θ0g = \theta_0; otherwise set g=θi−1g = \theta_{i-1}
        For each w∈DPTw \in D_{PT}, score ww by the cosine similarity of ∇θLθ0PT(w)\nabla_\theta L^{PT}_{\theta_0}(w) and ∇θ∑(x,y)∈DtaskLgICL(x,y)\nabla_\theta \sum_{(x,y)\in D_{task}} L^{ICL}_{g}(x,y)
        Let SiS_i be the kk highest-scoring instances
        Set Siall=⋃j=1iSjS_i^{all} = \bigcup_{j=1}^{i} S_j
        Starting from θ0\theta_0, perform one-pass SGD on SiallS_i^{all} to obtain θi\theta_i
    Return S=⋃i=1MSiS = \bigcup_{i=1}^{M} S_i

    In the experiments, the model was OPT-6.7B, the candidate pool contained 2.5 million pretraining instances of 2,048 tokens each, and the run used at most M=5M=5 iterations with k=400k=400. The one-pass updates used Adam, batch size 16, and learning rate 2×10−52\times10^{-5}; the resulting perturbation comprised 125 updates. The selected set was used as a small, perturbative continued-pretraining set, not as a replacement for full pretraining.

  2. Knowl 2 — Selected pretraining data generally improves ICL relative to random continued pretraining

    data/table

    For each row, supportive pretraining instances were selected using the named source task and then used for perturbative continued pretraining; the resulting model was evaluated on all six tasks with demonstrations. Each cell reports the selected-data accuracy first, followed by the accuracy mean ±\pm standard deviation from five random-data selections of the same size. The original OPT-6.7B ICL accuracies are shown in the first row. The selected-data run used the best result among M∈{1,2,3,4,5}M\in\{1,2,3,4,5\} iterations. The paper treated a selected-data result as stronger than random selection when the difference exceeded one random-selection standard deviation. Same-task results improved under this criterion for four of six source tasks; for five of the six source tasks, selected data improved results on at least three evaluation tasks. Cross-task transfer was therefore common but not universal.

    Source / evaluation SST-2 AG News Story Cloze SMS Spam Sentiment 140 TweetQA
    Original OPT-6.7B ICL 75.47 74.12 66.09 45.07 67.23 62.36
    SST-2 83.15 (75.87 ±\pm 1.64) 74.91 (73.24 ±\pm 1.24) 67.76 (66.24 ±\pm 1.25) 52.48 (49.82 ±\pm 4.50) 69.03 (66.23 ±\pm 1.24) 62.20 (61.75 ±\pm 0.26)
    AG News 79.04 (74.99 ±\pm 0.77) 75.40 (73.77 ±\pm 0.41) 68.34 (66.38 ±\pm 0.69) 59.24 (46.55 ±\pm 4.24) 68.96 (66.23 ±\pm 1.24) 61.86 (62.02 ±\pm 0.55)
    Story Cloze 75.33 (72.50 ±\pm 2.53) 74.12 (73.77 ±\pm 0.41) 67.47 (65.25 ±\pm 1.52) 51.36 (47.15 ±\pm 4.90) 69.92 (66.23 ±\pm 1.24) 62.33 (62.02 ±\pm 0.55)
    SMS Spam 73.88 (75.87 ±\pm 1.64) 72.78 (73.77 ±\pm 0.41) 67.25 (65.25 ±\pm 1.52) 64.69 (46.55 ±\pm 4.24) 63.70 (66.33 ±\pm 1.34) 62.13 (61.75 ±\pm 0.26)
    Sentiment 140 77.56 (73.49 ±\pm 2.33) 72.78 (73.77 ±\pm 0.41) 66.78 (66.38 ±\pm 0.69) 51.64 (44.52 ±\pm 2.45) 66.66 (66.00 ±\pm 1.41) 62.93 (61.64 ±\pm 0.21)
    TweetQA 75.22 (72.50 ±\pm 2.53) 71.52 (73.01 ±\pm 1.42) 66.27 (64.91 ±\pm 2.01) 43.09 (44.52 ±\pm 2.45) 66.76 (66.33 ±\pm 1.34) 61.31 (61.33 ±\pm 0.80)

    Accuracy values are reported as in the experiment. The gains and regressions across rows show that supportive instances can transfer between tasks, but the benefit depends on the source and evaluation tasks.

  3. Knowl 3 — Task selection and ICL evaluation protocol

    experimental setup

    The experiments used the autoregressive OPT-6.7B model and classification tasks from Natural Instructions v2. The researchers first screened 48 classification tasks, retained 17 for which ICL performance was at least 10% better than zero-shot performance, and chose six representative tasks: SST-2 (movie sentiment), AG News (news topic), Story Cloze (story coherence), SMS Spam, Sentiment 140 (tweet sentiment), and TweetQA (answer verification).

    For each task, 500 examples with balanced class distributions guided supportive-data selection; evaluation used the full task dataset. ICL prompts used the task’s Natural Instructions v2 template and verbalizer but omitted task instructions. For each test instance, four demonstration examples were independently sampled per candidate class and then randomly shuffled. The pretraining candidate pool consisted of 2.5 million 2,048-token instances, about 5 billion tokens or 3% of OPT’s 180-billion-token pretraining data. The study therefore tests targeted, small-scale continued-pretraining perturbations rather than retraining the full model.

  4. Knowl 4 — ICL-supportive updates do not generally improve zero-shot task performance

    empirical result

    To test whether selected pretraining data helped ICL specifically rather than merely improving task-input performance, the authors evaluated the original OPT-6.7B and the model continued-pretrained on ICL-supportive data without demonstrations. The reported zero-shot scores were:

    Task Original ICL-supportive continued pretraining
    SST-2 46.82 46.83
    AG News 46.14 44.05
    Story Cloze 50.43 51.39
    SMS Spam 44.41 43.84
    Sentiment 140 55.84 54.90
    TweetQA 50.44 50.32

    Only small increases were reported for SST-2 and Story Cloze; the other four scores decreased. The authors conclude that the selected data generally support performance in the demonstration-based setting without producing a corresponding zero-shot improvement.

  5. Knowl 5 — Supportive instances are not more domain-relevant to downstream tasks

    empirical result

    The domain-relevance analysis compared MAUVE scores between each downstream task’s text and (a) its selected supportive pretraining instances and (b) randomly sampled pretraining instances. The reported quantity was the supportive-data MAUVE score minus the random-data MAUVE score; positive values would indicate greater domain relevance for supportive instances. MAUVE embeddings were produced with RoBERTa, and 95% confidence intervals were estimated using 32 random seeds.

    For five of the six tasks, the analysis found no significant domain-relevance difference between supportive and random pretraining data; for SST-2, supportive data had a lower MAUVE score. Thus, the selected examples were not more closely matched to the downstream task domain than general pretraining examples. The authors suggest that this pattern is consistent with support for a domain-independent learning mechanism, but the analysis is correlational and does not establish that mechanism.

  6. Knowl 6 — Supportive instances have flatter token-frequency distributions

    empirical result

    For each 2,048-token pretraining instance, the authors fit a Zipfian distribution by regressing log token frequency on log token rank; the Zipf coefficient is the magnitude of the negative slope. A lower coefficient indicates a flatter distribution with relatively more mass on rare tokens. They compared groups of 2,000 supportive and random instances, with 95% confidence intervals estimated from 32 random seeds.

    For all six downstream tasks, supportive instances had significantly lower Zipf coefficients than random instances. The result indicates a relatively greater concentration of rarely occurring, long-tail tokens in the supportive data. The authors also note that a flatter distribution corresponds to higher token entropy, which may make these instances harder for the language model to fit; that proposed connection is not itself a causal test.

  7. Knowl 7 — Long-context information gain measures the added predictive value of context

    definition

    The paper defines information gain for next-token prediction as the decrease in a token’s cross-entropy when a longer context is available rather than a shorter one:

    IG(l,s;wi)=CE(wi∣Cs(i))−CE(wi∣Cl(i)).\mathrm{IG}(l,s;w_i)=\mathrm{CE}(w_i\mid C_s(i))-\mathrm{CE}(w_i\mid C_l(i)).

    Here, wiw_i is a target token in a pretraining sequence; ss and ll are short and long context lengths measured in tokens, with l>sl>s; Cs(i)C_s(i) and Cl(i)C_l(i) are the contexts used to predict wiw_i; and CE(wi∣C)\mathrm{CE}(w_i\mid C) is the token’s cross-entropy under the language model conditioned on context CC. A positive value means the longer context lowers the prediction cross-entropy. Because computing fixed-length contexts for every token was costly, the experiments approximated the measure by dividing sequences into blocks and using context lengths with the corresponding average lengths. They set s=128s=128 and tested l∈{256,512,1024}l\in\{256,512,1024\}.

  8. Knowl 8 — Supportive instances gain less predictive value from longer context

    empirical result

    The authors compared average long-context information gain in supportive versus random pretraining instances, calculating cross-entropies with OPT-1.3B rather than the OPT-6.7B model used for ICL; OPT-350M experiments reportedly showed a similar trend. With short context length fixed at 128 tokens and long context lengths of 256, 512, and 1,024 tokens, the supportive instances had lower relative information gain than random instances across all six tasks, with the difference tending to become more negative as long context length increased.

    The authors interpret this as evidence that supportive instances are more challenging for incorporating long-range context during language modeling. They propose that such instances may require the model to distinguish useful context from confounding spans, a capability that could be relevant to learning from demonstrations. This explanation is a hypothesis: low information gain may also result when long-range context is inherently irrelevant, so the converse claim—that every low-gain example is usefully challenging—does not follow.

  9. Knowl 9 — Scope and computational limitations

    limitation

    The identified supportive instances are relative to the particular language-model checkpoint: examples not selected for the final checkpoint may have been supportive for earlier checkpoints during pretraining. The study’s analyses of domain relevance, token frequency, and context information gain are correlational, so they do not show that deliberately increasing the measured features would causally improve ICL. The experiments also cover classification tasks only; applying the method to generation tasks would require defining an ICL loss over sequence probabilities.

    The gradient-based search is computationally expensive. In the reported setup, each source task’s search took about a week on four nodes with eight Nvidia V100 GPUs per node; its cost was described as comparable to training on the candidate pretraining data with batch size one.

Coverage note — The appendix’s qualitative pretraining-text examples are omitted because they are illustrative anecdotes rather than additional systematic findings; the quantitative comparisons and stated limitations are included.

References

  1. 1.Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou. 2022. What learning algorithm is in-context learning? investigations with linear models. arXiv preprint arXiv:2211.15661.
  2. 2.Tiago A Almeida, José María G Hidalgo, and Akebo Yamakami. 2011. Contributions to the study of sms spam filtering: new collection and results. In Proceedings of the 11th ACM symposium on Document engineering, pages 259–262.
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  4. 4.Stephanie CY Chan, Adam Santoro, Andrew Kyle Lampinen, Jane X Wang, Aaditya K Singh, Pierre Harvey Richemond, James McClelland, and Felix Hill. 2022. Data distributional properties drive emergent in-context learning in transformers. In Advances in Neural Information Processing Systems.
  5. 5.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  6. 6.Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. 2022. Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers. arXiv preprint arXiv:2212.10559.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proc. NAACL-HLT.
  8. 8.Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey for in-context learning. arXiv preprint arXiv:2301.00234.
  9. 9.Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant. 2022. What can transformers learn in-context? a case study of simple function classes. arXiv preprint arXiv:2208.01066.
  10. 10.Alec Go, Richa Bhayani, and Lei Huang. 2009. Twitter sentiment classification using distant supervision. CS224N project report, Stanford, 1(12):2009.
  11. 11.Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2023. Pre-training to learn in context.
  12. 12.Kelvin Guu, Albert Webson, Elizabeth-Jane Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi. 2023. Simfluence: Modeling the influence of individual training examples by simulating training runs. ArXiv, abs/2303.08114.
  13. 13.Xiaochuang Han and Yulia Tsvetkov. 2022. Orca: Interpreting prompted language models via locating supporting data evidence in the ocean of pretraining data. arXiv preprint arXiv:2205.12600.
  14. 14.Tianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, and Yulia Tsvetkov. 2022. On the blind spots of model-based evaluation metrics for text generation. arXiv preprint arXiv:2212.10020.
  15. 15.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556.
  16. 16.Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019. The curious case of neural text degeneration. In International Conference on Learning Representations.
  17. 17.Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980.
  18. 18.Pang Wei Koh and Percy Liang. 2017. Understanding black-box predictions via influence functions. In Proc. ICML.
  19. 19.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  20. 20.Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2022. Rethinking the role of demonstrations: What makes in-context learning work? In EMNLP.
  21. 21.Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. 2017. Ls-dsem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pages 46–51.
  22. 22.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al. 2022. In-context learning and induction heads. arXiv preprint arXiv:2209.11895.
  23. 23.Myle Ott, Sam Shleifer, Min Xu, Priya Goyal, Quentin Duval, and Vittorio Caggiano. 2021. Fully sharded data parallel: faster ai training with fewer gpus. https://engineering.fb.com/2021/07/15/open-source/fsdp/.
  24. 24.Jane Pan, Tianyu Gao, Howard Chen, and Danqi Chen. 2023. What in-context learning"learns"in-context: Disentangling task recognition and task learning.
  25. 25.Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2019. Language models as knowledge bases? In Proc. EMNLP.
  26. 26.Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaïd Harchaoui. 2021. Mauve: Measuring the gap between neural text and human text using divergence frontiers. In Proc. NeurIPS.
  27. 27.Garima Pruthi, Frederick Liu, Mukund Sundararajan, and Satyen Kale. 2020. Estimating training data influence by tracking gradient descent. In Proc. NeurIPS.
  28. 28.Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. 2022. Impact of pretraining term frequencies on few-shot reasoning. ArXiv, abs/2202.07206.
  29. 29.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proc. EACL.
  30. 30.Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woomyoung Park, Jung-Woo Ha, et al. 2022. On the effect of pretraining corpora on in-context learning by a large-scale language model. arXiv preprint arXiv:2204.13509.
  31. 31.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053.
  32. 32.Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642.
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  34. 34.Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. 2022. Transformers learn in-context by gradient descent. arXiv preprint arXiv:2212.07677.
  35. 35.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, et al. 2022. Benchmarking generalization via in-context instructions on 1,600+ language tasks. arXiv preprint arXiv:2204.07705.
  36. 36.Jerry W. Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma. 2023. Larger language models do in-context learning differently. ArXiv, abs/2303.03846.
  37. 37.Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. 2022. An explanation of in-context learning as implicit bayesian inference. In International Conference on Learning Representations.
  38. 38.Wenhan Xiong, Jiawei Wu, Hong Wang, Vivek Kulkarni, Mo Yu, Shiyu Chang, Xiaoxiao Guo, and William Yang Wang. 2019. Tweetqa: A social media focused question answering dataset. arXiv preprint arXiv:1907.06292.
  39. 39.Hongxin Zhang, Yanzhe Zhang, Ruiyi Zhang, and Diyi Yang. 2022a. Robustness of demonstration-based learning under limited data scenario. arXiv preprint arXiv:2210.10693.
  40. 40.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022b. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068.
  41. 41.Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proc. NeurIPS.

Citation

MLA
Han, X., et al. “Understanding In-Context Learning via Supportive Pretraining Data”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 12660–73, https://doi.org/10.18653/v1/2023.acl-long.708.
APA
Han, X., Simig, D., Mihaylov, T., Tsvetkov, Y., Celikyilmaz, A., & Wang, T. (2023). Understanding In-Context Learning via Supportive Pretraining Data. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12660–12673. https://doi.org/10.18653/v1/2023.acl-long.708
Chicago
Han, X., D. Simig, T. Mihaylov, Y. Tsvetkov, A. Celikyilmaz, and T. Wang. 2023. “Understanding In-Context Learning via Supportive Pretraining Data”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12660–73. https://doi.org/10.18653/v1/2023.acl-long.708.
Harvard
Han, X. et al. (2023) “Understanding In-Context Learning via Supportive Pretraining Data”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 12660–12673. Available at: https://doi.org/10.18653/v1/2023.acl-long.708.
Vancouver
1. Han X, Simig D, Mihaylov T, Tsvetkov Y, Celikyilmaz A, Wang T (2023) Understanding In-Context Learning via Supportive Pretraining Data. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 12660–12673

BibTeX

@inproceedings{han-etal-2023-understanding,
    title = "Understanding In-Context Learning via Supportive Pretraining Data",
    author = "Han, Xiaochuang  and
      Simig, Daniel  and
      Mihaylov, Todor  and
      Tsvetkov, Yulia  and
      Celikyilmaz, Asli  and
      Wang, Tianlu",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.708/",
    doi = "10.18653/v1/2023.acl-long.708",
    pages = "12660--12673"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/