Scaling Data-Constrained Language Models
Niklas MuennighoffAlexander M. RushBoaz BarakTeven Le ScaoAleksandra PiktusNouamane TaziSampo PyysaloThomas WolfColin Raffel
Establishes empirical scaling laws for language models trained in data-constrained regimes, demonstrating that repeating data for up to four epochs incurs negligible performance loss and provides a compute-optimal allocation strategy when unique text data is exhausted.
The rapid scaling of large language models has historically depended on expanding both model parameter counts and training dataset sizes simultaneously. However, researchers project that high-quality public text data on the internet may soon be exhausted, creating a severe data bottleneck for training future systems. Standard scaling frameworks, such as the Chinchilla scaling laws, assume an unlimited supply of unique data and advocate for single-epoch training. This article addresses the critical operational question facing artificial intelligence development: how to allocate computational resources effectively when unique text data is constrained.
The main objective of the article is to evaluate how repeating training data over multiple passes (epochs) affects model performance and to determine the optimal allocation of computing budgets across model size and training duration under data scarcity.
To investigate this, the authors executed over 400 empirical pre-training runs utilizing up to 900 billion total tokens and models scaling up to 8.7 billion parameters across high-performance supercomputing infrastructure. They systematically varied the degree of data repetition, total computational budget, and model parameter sizes using standardized web text corpora, primarily C4 and OSCAR. In addition to testing data repetition, the study assessed complementary data augmentation strategies, such as incorporating source code and modifying data-filtering thresholds, evaluating outcomes through held-out test loss and 19 downstream natural language benchmarks.
The article establishes several key findings. First, training on repeated data for up to 4 epochs results in negligible performance degradation compared to training on entirely fresh data, showing only about a 0.5% difference in held-out loss and preserving downstream task accuracy. Second, while repetition yields diminishing returns that decay toward zero past roughly 16 epochs, multi-epoch training remains far more viable than previously assumed. Third, when operating under fixed data constraints, the optimal resource allocation strategy diverges from traditional guidelines: additional compute should be disproportionately directed toward running smaller models for more epochs rather than simply expanding parameter counts. Finally, augmenting natural language datasets with up to 50% Python source code effectively doubles available training tokens without degrading natural language task performance, while also boosting reasoning and state-tracking capabilities.
These findings have immediate implications for capital deployment, hardware planning, and model architecture strategies. Organizations facing data scarcity can achieve better model performance with substantially smaller parameter footprints by training over multiple epochs, which directly reduces downstream inference costs and deployment risks. This contradicts prior industry practices, such as the 120-billion-parameter Galactica model, which over-allocated compute to parameters rather than epochs. Furthermore, the analysis indicates that aggressive data deduplication filters can needlessly discard useful training tokens on clean corpora, whereas filtering is primarily beneficial only on noisy source data.
Based on these results, practitioners facing data limits should plan pre-training budgets around smaller architectures trained across 4 to 16 epochs, integrate source code to safely double effective token volume, and reserve strict filtering for heavily contaminated datasets. Looking forward, further validation is warranted to explore how specific regularizers impact excess parameter decay and to measure scaling behaviors across diverse modalities, multi-language mixes, and non-transformer architectures.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). This seminal paper introduces the compute-optimal Chinchilla scaling laws and single-epoch baseline assumptions that the source article directly re-evaluates under data constraints.
- Paper: Scaling Laws for Neural Language Models, Jared Kaplan et al. (2020). This foundational work establishes empirical power-law relationships between model size, dataset size, and compute that serve as the standard framework adapted throughout the source.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). This study establishes standardized benchmarking for language model data curation and filtering strategies, providing direct context for the source's findings on data filtering thresholds.
- Paper: Scaling Language Models: Methods, Analysis & Insights from Training Gopher, Jack W. Rae et al. (2021). This work demonstrates large-scale model pre-training dynamics and parameter allocations on web text corpora, providing empirical baselines referenced in compute-allocation analyses.
- Paper: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, Stella Biderman et al. (2023). This suite systematically analyzes the effects of deduplication and training steps on memorization and capability across scales, motivating the study of data repetition.
- Paper: DsDm: Model-Aware Dataset Selection with Datamodels, Logan Engstrom et al. (2024). This paper analyzes the trade-offs of pre-training data filtering heuristics versus model performance, informing the data curation analysis in the source.
- Paper: Bridging Compute- and Data-Optimal Pretraining, Tian Qin et al.. This work formalizes and extends multi-epoch and data-constrained scaling into a unified compute-data Pareto framework with explicit token-effectiveness functions.
- Paper: Skaling: Chinchilla's Exponents Meet Kaplan's Coupling, Mathurin Videau et al. (2026). This research builds upon the failure modes of standard scaling laws in data-imbalanced regimes by formulating coupled scaling equations that improve predictions under resource constraints.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). This work extends the source's insights on web data filtering by evaluating scaling dynamics and crossing points when training dense models on unfiltered versus filtered Common Crawl pools.
- Paper: Reasoning to Learn from Latent Thoughts, Yangjun Ruan et al. (2025). This paper proposes a complementary solution to the data bottleneck by generating synthetic latent thoughts to extract richer learning signals from fixed reasoning corpora.
