Built independently by an author, for readers. Read the story and support ChapterPal

keyword

supportive pretraining data

Supportive pretraining data refers to a subset of training examples within a machine learning model pretraining corpus that directly facilitates and enhances specific emergent capabilities, such as in-context learning and downstream task adaptation. Rather than being selected purely based on domain relevance or topic similarity to target tasks, supportive pretraining data is typically identified by analyzing its functional contribution to model behavior through gradient-based attribution or data influence methods. These high-utility examples often possess distinct compositional traits, such as an enriched concentration of rare or long-tail tokens and challenging contextual structures that require the model to resolve long-range dependencies during unsupervised pretraining.

1 item

Understanding In-Context Learning via Supportive Pretraining Data

Understanding In-Context Learning via Supportive Pretraining Data

Xiaochuang Han, Daniel Simig, Todor Mihaylov, Yulia Tsvetkov, Asli Celikyilmaz, Tianlu Wang

OrganizationsMetaUniversity of Washington

Why you should read this

Reveals that in-context learning in large language models is driven by specific, challenging pretraining instances rich in long-tail tokens and difficult long-range contexts rather than domain-relevant text, providing actionable criteria to guide future pretraining data selection.

In-context learning (ICL) improves language models’ performance on a variety of NLP tasks by simply demonstrating a handful of examples at inference time. It is not well understood why ICL ability emerges, as the model has never been specifically trained on such demonstrations. Unlike prior work that explores implicit mechanisms behind ICL, we study ICL via investigating the pretraining data. Specifically, we first adapt an iterative, gradient-based approach to find a small subset of pretraining data that supports ICL. We observe that a continued pretraining on this small subset significantly improves the model’s ICL ability, by up to 18%. We then compare the supportive subset contrastively with random subsets of pretraining data and discover: (1) The supportive pretraining data to ICL do not have a higher domain relevance to downstream tasks. (2) The supportive pretraining data have a higher mass of rarely occurring, long-tail tokens. (3) The supportive pretraining data are challenging examples where the information gain from long-range context is below average, indicating learning to incorporate difficult long-range context encourages ICL. Our work takes a first step towards understanding ICL via analyzing instance-level pretraining data. Our insights have a potential to enhance the ICL ability of language models by actively guiding the construction of pretraining data in the future.

Added

2026-10-03