On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model
Seongjin ShinSang-Woo LeeHwijeen AhnSungdong KimHyoungSeok KimBoseop KimKyunghyun ChoGichang LeeWoo-Myoung ParkJung-Woo Ha
Reveals that large language model in-context learning capabilities depend heavily on pretraining data sources and combinations rather than corpus size alone, showing that low validation perplexity does not reliably predict few-shot performance.
Large language models can perform downstream tasks through in-context learning—solving tasks using task prompts and a few reference examples without updating internal model weights. However, training these models requires massive computational and financial resources, and engineering teams currently lack clear insights into which pretraining text sources drive this emergent capability.
The article systematically analyzes how the source domain, volume, and combinations of pretraining text affect the emergence of zero- and few-shot in-context learning in large language models. It also investigates whether standard validation loss measurements reliably predict how well a model will perform on target tasks.
To conduct this evaluation, the researchers trained multiple 1.3-billion parameter variants of the Korean-centric HyperCLOVA model across seven distinct text domains (blogs, community forums, news, comments, question-and-answer boards, a curated government corpus, and encyclopedias) as well as combinations of these datasets, alongside baseline tests on a 6.9-billion parameter model. The models were trained under controlled token budgets (up to 150 billion tokens) and evaluated across four standardized tasks: sentiment classification, reading comprehension, machine translation, and topic classification.
The investigation produced four central findings. First, the pretraining source domain heavily governs in-context learning capability regardless of size; for example, a model trained on 150 billion blog tokens achieved competitive few-shot performance comparable to the full multi-domain model, whereas models trained on large volumes of news or forum text failed to develop few-shot capabilities. Second, combining ineffective single-domain datasets can trigger the emergence of strong in-context learning; mixing question-and-answer and encyclopedia texts produced strong reading comprehension and translation abilities that neither corpus could achieve alone. Third, domain relevance between pretraining text and target tasks improves zero-shot performance but fails to guarantee few-shot capability. Fourth, validation perplexity—a standard measure of how well a model predicts words—is not a reliable cross-model predictor of few-shot task accuracy, as models with low perplexity frequently failed downstream few-shot evaluations.
These findings challenge the common industry assumption that simply gathering more data or minimizing general validation loss will ensure high-performing few-shot models. For organizations building and deploying language models, strategic curation and diverse combination of text domains are far more critical than raw volume, offering substantial opportunities to lower compute costs, energy usage, and training timelines.
Practitioners should prioritize domain diversity and strategic data mixing rather than relying solely on single-domain scaling or validation loss metrics during pretraining. When preparing models for specialized applications, teams should pilot combinations of complementary text types and evaluate downstream task performance directly.
These conclusions are primarily bounded by evaluations conducted on a Korean-language architecture and models up to 6.9 billion parameters. While confidence in these empirical results is high, stakeholders should exercise caution before generalizing specific domain behaviors to other languages or massive models exceeding tens of billions of parameters without preliminary validation.
- Paper: Language Models are Few-Shot Learners, T. B. Brown et al. (2020). Brown et al. establish GPT-3’s zero-, one-, and few-shot in-context learning paradigm, the foundation this study probes by varying pretraining corpora.
- Paper: Pre-Training to Learn in Context, Yuxian Gu et al. (2023). PICL turns the source’s finding that pretraining data shapes in-context learning into a training method, using retrieved passages from unstructured text to cultivate the capability.
