Understanding In-Context Learning via Supportive Pretraining Data
Xiaochuang HanDaniel SimigTodor MihaylovYulia TsvetkovAsli CelikyilmazTianlu Wang
Reveals that in-context learning in large language models is driven by specific, challenging pretraining instances rich in long-tail tokens and difficult long-range contexts rather than domain-relevant text, providing actionable criteria to guide future pretraining data selection.
Large language models frequently perform downstream tasks through in-context learning, an operational capability where models infer solutions from just a few demonstrated examples provided in the prompt without updating their underlying parameters. Understanding where this ability originates during pretraining is essential for building more reliable systems and optimizing the massive datasets used to train them.
The article investigates how pretraining data drives in-context learning capabilities. Specifically, it identifies which specific pretraining instances support few-shot performance on downstream tasks and analyzes the key statistical and informational properties that distinguish supportive data from standard pretraining text.
The researchers evaluated the open-source OPT-6.7B model across six diverse natural language classification tasks selected from a broader instruction benchmark. Using an iterative, gradient-based search method called ORCA-ICL, the team scanned a subset of 2.5 million pretraining documents (comprising roughly 5 billion tokens) to locate training instances that share gradient similarity with task demonstration data. The team then performed targeted, one-pass continued pretraining on these small identified subsets (under 2,000 instances) to verify their impact, benchmarking the resulting models against both the original baseline and control models trained on randomly sampled pretraining data.
The evaluation revealed three key findings regarding supportive data. First, continued pretraining on the small supportive subset significantly boosted few-shot accuracy by up to 18% compared to random training subsets, while having no positive effect on zero-shot performance without demonstrations. Second, domain analysis using distributional text similarity metrics showed that supportive pretraining instances share no greater domain relevance to downstream tasks than random text, indicating that in-context learning relies on domain-invariant reasoning mechanisms rather than specific topic knowledge. Third, structural analysis demonstrated that supportive documents feature a higher concentration of rare, long-tail vocabulary terms and present lower information gain from extended context. This indicates that supportive instances are intrinsically challenging texts that force the model to separate relevant cues from long-range confounding context.
These findings indicate that in-context learning functions as an abstract meta-capability shaped by training on complex, linguistically rich text rather than simple exposure to task-specific subject matter. For AI practitioners and engineering leaders, this shifts data strategy away from naive domain-matching toward curating conceptually challenging training material that emphasizes rare terms and difficult context-filtering scenarios. Such targeted data curation has the potential to substantially reduce overall pretraining volume and compute costs while improving prompt-based model performance.
Organizations developing or refining foundation models should pilot pretraining data curation pipelines that screen for lower token-frequency concentration and higher contextual complexity. However, because the current search process requires significant computation—taking approximately one week across 32 enterprise GPUs per source task—practitioners should prioritize exploring more efficient gradient approximations, generation-based synthetic data methods, or applying these filtering criteria during upstream data acquisition.
Confidence in these findings is strong for autoregressive language models performing classification tasks under few-shot prompting. Readers should nevertheless exercise caution regarding generalizability, as the study focused exclusively on classification benchmarks, analyzed data relative to a finalized model checkpoint rather than early training phases, and relied primarily on correlational data properties that require further causal validation before establishing definitive pretraining guidelines.
- Paper: On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model, Seongjin Shin et al. (2022). Read this earlier investigation of how pretraining corpora shape in-context learning first; it establishes the data-source question that this paper narrows to individual supportive training instances.
- Paper: Pre-Training to Learn in Context, Yuxian Gu et al. (2023). Its framework for preparing models to learn in context from unstructured text provides a useful prerequisite for understanding this paper’s search for ICL-supportive pretraining instances.
- Paper: LESS: Selecting Influential Data for Targeted Instruction Tuning, Mengzhou Xia et al. (2024). This later work carries gradient-based data selection into targeted instruction tuning, extending the source’s approach to identifying training examples that improve a specific downstream capability.
