Scaling Data-Constrained Language Models

Niklas MuennighoffAlexander M. RushBoaz BarakTeven Le ScaoAleksandra PiktusNouamane TaziSampo PyysaloThomas WolfColin Raffel

article2025JMLR462 citationsOutstanding Paper Runner-up

Establishes empirical scaling laws for language models trained in data-constrained regimes, demonstrating that repeating data for up to four epochs incurs negligible performance loss and provides a compute-optimal allocation strategy when unique text data is exhausted.

Listen

The rapid scaling of large language models has historically depended on expanding both model parameter counts and training dataset sizes simultaneously. However, researchers project that high-quality public text data on the internet may soon be exhausted, creating a severe data bottleneck for training future systems. Standard scaling frameworks, such as the Chinchilla scaling laws, assume an unlimited supply of unique data and advocate for single-epoch training. This article addresses the critical operational question facing artificial intelligence development: how to allocate computational resources effectively when unique text data is constrained.

The main objective of the article is to evaluate how repeating training data over multiple passes (epochs) affects model performance and to determine the optimal allocation of computing budgets across model size and training duration under data scarcity.

To investigate this, the authors executed over 400 empirical pre-training runs utilizing up to 900 billion total tokens and models scaling up to 8.7 billion parameters across high-performance supercomputing infrastructure. They systematically varied the degree of data repetition, total computational budget, and model parameter sizes using standardized web text corpora, primarily C4 and OSCAR. In addition to testing data repetition, the study assessed complementary data augmentation strategies, such as incorporating source code and modifying data-filtering thresholds, evaluating outcomes through held-out test loss and 19 downstream natural language benchmarks.

The article establishes several key findings. First, training on repeated data for up to 4 epochs results in negligible performance degradation compared to training on entirely fresh data, showing only about a 0.5% difference in held-out loss and preserving downstream task accuracy. Second, while repetition yields diminishing returns that decay toward zero past roughly 16 epochs, multi-epoch training remains far more viable than previously assumed. Third, when operating under fixed data constraints, the optimal resource allocation strategy diverges from traditional guidelines: additional compute should be disproportionately directed toward running smaller models for more epochs rather than simply expanding parameter counts. Finally, augmenting natural language datasets with up to 50% Python source code effectively doubles available training tokens without degrading natural language task performance, while also boosting reasoning and state-tracking capabilities.

These findings have immediate implications for capital deployment, hardware planning, and model architecture strategies. Organizations facing data scarcity can achieve better model performance with substantially smaller parameter footprints by training over multiple epochs, which directly reduces downstream inference costs and deployment risks. This contradicts prior industry practices, such as the 120-billion-parameter Galactica model, which over-allocated compute to parameters rather than epochs. Furthermore, the analysis indicates that aggressive data deduplication filters can needlessly discard useful training tokens on clean corpora, whereas filtering is primarily beneficial only on noisy source data.

Based on these results, practitioners facing data limits should plan pre-training budgets around smaller architectures trained across 4 to 16 epochs, integrate source code to safely double effective token volume, and reserve strict filtering for heavily contaminated datasets. Looking forward, further validation is warranted to explore how specific regularizers impact excess parameter decay and to measure scaling behaviors across diverse modalities, multi-language mixes, and non-transformer architectures.

  • Paper: Bridging Compute- and Data-Optimal Pretraining, Tian Qin et al.. This work formalizes and extends multi-epoch and data-constrained scaling into a unified compute-data Pareto framework with explicit token-effectiveness functions.
  • Paper: Skaling: Chinchilla's Exponents Meet Kaplan's Coupling, Mathurin Videau et al. (2026). This research builds upon the failure modes of standard scaling laws in data-imbalanced regimes by formulating coupled scaling equations that improve predictions under resource constraints.
  • Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). This work extends the source's insights on web data filtering by evaluating scaling dynamics and crossing points when training dense models on unfiltered versus filtered Common Crawl pools.
  • Paper: Reasoning to Learn from Latent Thoughts, Yangjun Ruan et al. (2025). This paper proposes a complementary solution to the data bottleneck by generating synthetic latent thoughts to extract richer learning signals from fixed reasoning corpora.
Cover for Scaling Data-Constrained Language Models

Abstract

The current trend of scaling language models involves increasing both parameter count and training data set size. Extrapolating this trend suggests that training data set size may soon be limited by the amount of text data available on the internet. Motivated by this limit, we investigate scaling language models in data-constrained regimes. Specifically, we run a large set of experiments varying the extent of data repetition and compute budget, ranging up to 900 billion training tokens and 9 billion parameter models. We find that with constrained data for a fixed compute budget, training with up to 4 epochs of repeated data yields negligible changes to loss compared to having unique data. However, with more repetition, the value of adding compute eventually decays to zero. We propose and empirically validate a scaling law for compute optimality that accounts for the decreasing value of repeated tokens and excess parameters. Finally, we experiment with approaches mitigating data scarcity, including augmenting the training data set with code data or removing commonly used filters. Models and data sets from our 400 training runs are freely available at https://github.com/huggingface/datablation.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Method: Data-Constrained Scaling Laws
  • 3.1 Parametric Fit
  • 4. Experimental Setup
  • 5. Results: Resource Allocation for Data-Constrained Scaling
  • 5.1 Double Descent
  • 5.2 Repeating on Heavily Deduplicated Data
  • 5.3 Do Excess Parameters Hurt, Plateau or Help?
  • 6. Results: Resource Return for Data-Constrained Scaling
  • 6.1 Training Loss
  • 6.2 Scaling Curves on the OSCAR Corpus
  • 6.3 Validation Loss by Epoch
  • 7. Case Study: Galactica
  • 8. Results: Complementary Strategies for Obtaining Additional Data
  • 8.1 Loss Curves for Complementary Strategies
  • 9. Related Work
  • 9.1 Large Language Models
  • 9.2 Scaling Laws
  • 9.3 Alternative Data Strategies
  • 10. Limitations and Future Work
  • 10.1 Repeating Fractions of the Data
  • 10.2 Sensitivity to Hyperparameters
  • 10.3 Other Data sets
  • 10.4 Other Modalities or Architectures
  • 10.5 Other Strategies
  • 11. Conclusion
  • Acknowledgments
  • Appendix A. Contributions
  • Appendix B. Derivation of Data-Constrained Scaling Laws
  • B.1 Analytical properties of compute-optimal point
  • Appendix C. C4 Scaling Coefficients
  • Appendix D. Evaluation Details
  • D.1 Loss Evaluation
  • D.2 Downstream Evaluation
  • Appendix E. Downstream Repetition Results
  • Appendix F. Detailed Code Augmentation Results
  • Appendix G. Filtering Procedure
  • G.1 Perplexity filtering
  • G.2 Deduplication
  • G.3 ROOTS filter
  • Appendix H. Detailed Filtering Results
  • Appendix I. Hyperparameters and Setup
  • Appendix J. Prompts and Samples
  • Appendix K. Other Experiments
  • K.1 UL2
  • K.2 The Pile
  • Appendix L. Release of Artifacts
  • References

Knowls

  1. Knowl 1 — Data-Constrained Scaling Law Formulation

    equation

    The data-constrained scaling law generalizes the compute-optimal parametric loss formula to settings where training data is repeated across multiple epochs and models may have excess parameters. For a language model with NN parameters trained on DD total processed tokens drawn from a constrained dataset containing DCD_C unique tokens, cross-entropy test loss LL is modeled as:

    L(UN,UD,RN,RD)=A(UN+UNRN∗(1−e−RN/RN∗))α+B(UD+UDRD∗(1−e−RD/RD∗))β+EL(U_N, U_D, R_N, R_D) = \frac{A}{\left(U_N + U_N R_N^* \left(1 - e^{-R_N / R_N^*}\right)\right)^\alpha} + \frac{B}{\left(U_D + U_D R_D^* \left(1 - e^{-R_D / R_D^*}\right)\right)^\beta} + E

    where:

    • UD=min⁡(DC,D)U_D = \min(D_C, D) is the count of unique tokens utilized.
    • RD=DUD−1R_D = \frac{D}{U_D} - 1 is the number of data repetitions (RD=0R_D = 0 corresponds to single-epoch training, with total epochs equal to RD+1R_D + 1).
    • UN=min⁡(Nopt(UD),N)U_N = \min(N_{\text{opt}}(U_D), N) represents the number of unique parameters compute-optimal for UDU_D tokens, defined by UN=min⁡(((UD⋅G)β/α)⋅G, N)U_N = \min\left(\left((U_D \cdot G)^{\beta/\alpha}\right) \cdot G, \, N\right) with G=(αAβB)1α+βG = \left(\frac{\alpha A}{\beta B}\right)^{\frac{1}{\alpha + \beta}}.
    • RN=max⁡(NUN−1, 0)R_N = \max\left(\frac{N}{U_N} - 1, \, 0\right) is the number of excess parameter repetitions.
    • RD∗R_D^* and RN∗R_N^* are learned decay constants characterizing the effective retention scale of data repetitions and excess parameters, respectively.
    • A,B,E,α,βA, B, E, \alpha, \beta are scaling parameters where EE is the irreducible loss.

    When RD=0R_D = 0 and RN=0R_N = 0, the equation reduces to the standard non-repeated Chinchilla scaling law L(N,D)=ANα+BDβ+EL(N, D) = \frac{A}{N^\alpha} + \frac{B}{D^\beta} + E. As repetitions increase to infinity (RD→∞,RN→∞R_D \to \infty, R_N \to \infty), the effective data and parameters saturate at UD(1+RD∗)U_D(1 + R_D^*) and UN(1+RN∗)U_N(1 + R_N^*).

  2. Knowl 2 — Fitted Scaling Coefficients on C4 under Data Constraints

    empirical result

    Fitting the data-constrained scaling law on 182 GPT-2 architecture training runs on C4 subsets (ranging from 7M to 9B parameters and 1 to 500 epochs) via Huber loss minimization with L-BFGS yields the following parameterized loss function:

    L(UD,RN,RD)=521(UN+5.3⋅UN(1−e−RN/5.3))0.35+1488(UD+15.4⋅UD(1−e−RD/15.4))0.35+1.87L(U_D, R_N, R_D) = \frac{521}{\left(U_N + 5.3 \cdot U_N \left(1 - e^{-R_N / 5.3}\right)\right)^{0.35}} + \frac{1488}{\left(U_D + 15.4 \cdot U_D \left(1 - e^{-R_D / 15.4}\right)\right)^{0.35}} + 1.87

    with the base optimal parameter relation given by UN=0.051⋅UDU_N = 0.051 \cdot U_D.

    The fitted baseline coefficients are A=521A = 521, B=1488B = 1488, E=1.87E = 1.87, and α=β=0.353\alpha = \beta = 0.353. The learned repetition constants are:

    • RD∗≈15.39R_D^* \approx 15.39: The decay scale for token repetitions, indicating that repeated tokens retain value up to roughly 15 repetitions (16 epochs).
    • RN∗≈5.31R_N^* \approx 5.31: The decay scale for excess parameters.

    Because RD∗>RN∗R_D^* > R_N^*, excess parameters suffer diminishing returns more rapidly than repeated tokens.

  3. Knowl 3 — Optimal Compute Allocation in Data-Constrained Regimes

    empirical result

    Under standard Chinchilla scaling laws with unconstrained unique data, compute-optimal training requires scaling model parameters NN and tokens DD in equal proportion (Nopt∝C0.5,Dopt∝C0.5N_{\text{opt}} \propto C^{0.5}, D_{\text{opt}} \propto C^{0.5}). In contrast, when the budget of unique data DCD_C is constrained and additional compute is spent by repeating data (RD>0R_D > 0), optimal resource allocation deviates from Chinchilla scaling by allocating more compute to training smaller models for more epochs.

    Because excess parameters decay faster in marginal utility than repeated tokens (RN∗≈5.31<RD∗≈15.39R_N^* \approx 5.31 < R_D^* \approx 15.39), the compute-efficient frontier favors higher repetition counts over larger parameter counts:

    • For a total compute budget of 9.3×10219.3 \times 10^{21} FLOPs and a unique token constraint of 25 billion tokens, the data-constrained compute-optimal design trains a 6.3B parameter model for 9.7 epochs (242B total tokens), achieving a held-out test loss of 2.359 and an average downstream benchmark score of 25.9%.
    • The standard Chinchilla allocation (which assumes fresh data) prescribes an 8.7B parameter model trained for 7.1 epochs (178B total tokens), achieving a worse test loss of 2.376 and an average downstream score of 23.5%.
  4. Knowl 4 — Diminishing Returns and the 4-Epoch Threshold for Repeated Data

    empirical result

    Evaluating language models (2.8B, 4.2B, and 8.7B parameters) trained under fixed compute budgets (9.3×10209.3 \times 10^{20}, 2.1×10212.1 \times 10^{21}, and 9.3×10219.3 \times 10^{21} FLOPs) across data budgets spanning 1 to 44 epochs reveals the following behavior:

    1. Training on repeated data for up to 4 epochs (where unique tokens constitute 25% of total processed tokens) incurs a negligible held-out test loss penalty compared to training on unique data for 1 epoch. For an 8.7B parameter model trained for 4 epochs on 44B unique tokens, final validation loss is only 0.5% higher than a 1-epoch model on 178B unique tokens, with comparable downstream task performance (26.8% vs. 26.2%).
    2. Training for more than 4 epochs yields diminishing returns on additional compute. Marginal gains decay sharply as repetition approaches the decay constant RD∗≈15R_D^* \approx 15 repetitions (16 epochs), beyond which effective data saturates.
    3. Excessive repetition (such as 44 epochs) leads to severe degradation, significant loss inflation, or training instability.
  5. Knowl 5 — Comparison of Parametric Decay Formulations for Repeated Data Scaling

    data/table

    Multiple mathematical formulations were fitted on 182 training runs (7M to 9B parameters, 1 to 500 epochs on C4) to evaluate how accurately they capture test loss under data repetition and excess parameters. Model fit was evaluated using fitting loss and the coefficient of determination (R2R^2):

    Parametric Fit RD∗R_D^* RN∗R_N^* Loss (↓\downarrow) R2R^2 (↑\uparrow)
    No decay (Standard Chinchilla) - - - 0.1430
    Decay NN only - 713.0015 0.0241 0.1671
    Decay DD only 2.9157 - 0.0169 0.7395
    Exponential decay model (Eq. 15) 15.3878 5.3097 0.0158 0.7810
    Geometric series decay (Eq. 11) for NN and DD 0.0104 0.3676 0.0155 0.8062
    Explicit exponential decay (Eq. 19) for NN and DD 0.0105 0.3676 0.0155 0.8061

    The data shows that:

    1. Assuming standard Chinchilla scaling without decay fails to fit multi-epoch training (R2=0.1430R^2 = 0.1430).
    2. Decaying only data (R2=0.7395R^2 = 0.7395) or only parameters (R2=0.1671R^2 = 0.1671) is inferior to decaying both tokens and parameters simultaneously (R2=0.7810R^2 = 0.7810).
    3. The simplified exponential decay approximation (R2=0.7810R^2 = 0.7810) performs comparably to the exact geometric series formulation (R2=0.8062R^2 = 0.8062) while maintaining directly interpretable decay constants RD∗R_D^* and RN∗R_N^*.
  6. Knowl 6 — Code Augmentation for Data-Scarcity Mitigation in LLM Pre-training

    empirical result

    When natural language text data is scarce, supplementing pre-training data (C4) with Python source code (The Stack) expands the effective token volume without hurting natural language task accuracy:

    • For 4.2B parameter models trained on 84B total tokens and 2.8B parameter models trained on 55B total tokens, substituting up to 50% of the training tokens with Python code produces no deterioration in average zero-to-five-shot accuracy across 19 natural language benchmarks (averaging 23.8% with 50% code vs. 22.1% with 0% code for 4.2B models).
    • Adding code data provides substantial performance gains on structured reasoning and state-tracking benchmarks: on the 20-task bAbI benchmark, models trained exclusively on C4 achieve 0.0% accuracy, whereas mixing in 10% to 50% Python code increases bAbI accuracy to 12.5%–23.2% for 4.2B models.
    • Code augmentation and multi-epoch repetition combine multiplicatively: doubling available data with 50% code and training for 4 epochs yields an 8×8\times increase in effective training tokens with negligible performance degradation.
  7. Knowl 7 — Effects of Perplexity and Deduplication Filtering in Data-Constrained Regimes

    empirical result

    Under data constraints, aggressive filtering removes unique tokens and forces higher repetition to satisfy a fixed compute budget. Benchmarking filtering approaches on C4 (a curated dataset) and OSCAR (a noisier dataset) reveals differing outcomes:

    1. Perplexity Filtering: Retaining only the top 25% lowest-perplexity documents (using a KenLM 5-gram language model trained on Wikipedia) and repeating the remaining tokens for ≈2\approx 2 epochs improves downstream evaluation scores for 4.2B parameter models on both C4 (average accuracy improves from 22.2% to 25.3%) and OSCAR (improves from 21.4% to 23.3%).
    2. Deduplication Filtering: Removing documents with ≥100\ge 100-character exact overlaps reduces C4 from 84B to 21B unique tokens (requiring 4 epochs to reach an 84B token budget). On C4, deduplication reduces average downstream performance from 22.2% to 20.5%. On OSCAR, deduplication removes repetitive web noise and improves average performance from 21.4% to 22.8%.
    3. Practical Guideline: Deduplication and aggressive filtering should be applied to noisy web datasets, but avoided on clean datasets where preserving unique token count is more advantageous than repeating heavily filtered subsets.
  8. Knowl 8 — Epoch-Wise Double Descent in Multi-Epoch LLM Training

    empirical result

    When dense autoregressive language models (14M to 83M parameters) are trained on a fixed dataset of 100 million unique tokens over hundreds to thousands of epochs, held-out test loss exhibits epoch-wise double descent rather than monotonic improvement:

    • Validation loss initially decreases up to 59–75 epochs.
    • At approximately 200 epochs, validation loss increases sharply across all model sizes (e.g., relative test loss increases back to ≈0.90\approx 0.90 from an earlier low of ≈0.82\approx 0.82 for 14M models).
    • Training beyond 200 epochs triggers a second descent phase where loss resumes decreasing, reaching a lower final test loss at 910–9000 epochs.

    Because standard scaling formulas assume monotonic loss reduction with respect to repetitions, training runs located within the double descent peak behave as non-monotonic outliers and must be excluded from parametric scaling law regression.

  9. Knowl 9 — Suboptimal Parameter Allocation in Galactica Due to Single-Epoch Scaling Assumptions

    empirical result

    The Galactica 120B model was trained on 106 billion unique scientific and code tokens for 4.25 epochs (450 billion total tokens), consuming approximately 5.4×10225.4 \times 10^{22} FLOPs. Evaluating this design against data-constrained scaling laws reveals substantial parameter overallocation:

    • Under standard single-epoch Chinchilla scaling (or assuming repeated data matches fresh data), a 5.4×10225.4 \times 10^{22} FLOPs budget suggests allocating compute symmetrically between parameters and data, resulting in a large model trained for few epochs.
    • Because excess parameters decay in marginal value faster than repeated tokens (RN∗<RD∗R_N^* < R_D^*), data-constrained scaling laws indicate that the optimal configuration for 5.4×10225.4 \times 10^{22} FLOPs on 106B unique tokens requires roughly 3×3\times fewer parameters and 3×3\times more epochs: specifically, a ∼40B\sim 40\text{B} parameter model trained for 1.35 trillion tokens (12.75 epochs).
    • Additionally, early stopping at 4.25 epochs to avoid transient validation spikes at epoch boundaries was premature, as validation spikes at epoch transitions in multi-epoch training are temporary before loss resumes decreasing.
  10. Knowl 10 — Multi-Epoch Language Model Training and Evaluation Setup

    experimental setup

    The empirical scaling study evaluates over 400 dense decoder-only transformer language models using the GPT-2 architecture and byte-pair encoding tokenizer:

    • Scale: Parameter counts span 7 million to 8.7 billion parameters (ranging from 3 layers / hidden size 128 to 42 layers / hidden size 4096, context length 2048, vocab size 50,257) trained on up to 900 billion total tokens (1 to 1500 epochs). Training compute is estimated as FLOPs(N,D)≈6ND\text{FLOPs}(N, D) \approx 6ND, where NN is parameter count and DD is total tokens processed.
    • Optimization: Models are trained using Megatron-DeepSpeed across up to 256 AMD Instinct MI250X GPUs. Training uses cosine learning rate schedules decaying 10×10\times across each run's total scheduled steps without early stopping.
    • Data Subsets and Repetition: Models are trained on subsets of C4 and OSCAR. For controlled multi-epoch comparisons, runs with fewer unique tokens strictly train on a direct subset of the data used in runs with more unique tokens, with complete dataset passes reshuffled after each epoch.
    • Evaluation Protocol: Held-out loss is evaluated on a fixed 210 million token validation split from C4. Downstream performance is benchmarked on 19 NLP datasets (14 classification/QA datasets evaluated via accuracy, 4 generation tasks via ROUGE-2, and bAbI via exact match) evaluated across 0 to 5 in-context few-shot exemplars with scores normalized against random baselines.

Coverage note — Omitted specific prompt templates for the 19 downstream tasks (Appendix J), full architectural dimension tables for every intermediate model size (Table 15), and exploratory negative results on the UL2 pre-training objective (Appendix K.1), as they do not constitute load-bearing theoretical or empirical scaling contributions.

References

  1. 1.Armen Aghajanyan, Lili Yu, Alexis Conneau, Wei-Ning Hsu, Karen Hambardzumyan, Susan Zhang, Stephen Roller, Naman Goyal, Omer Levy, and Luke Zettlemoyer. Scaling laws for generative mixed-modal language models. arXiv preprint arXiv:2301.03728, 2023.
  2. 2.Ibrahim M Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai. Revisiting neural scaling laws in language and vision. Advances in Neural Information Processing Systems, 35:22300–22312, 2022.
  3. 3.Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, et al. Santacoder: don’t reach for the stars! arXiv preprint arXiv:2301.03988, 2023.
  4. 4.Miltiadis Allamanis, Earl T Barr, Christian Bird, and Charles Sutton. Suggesting accurate method and class names. In Proceedings of the 2015 10th joint meeting on foundations of software engineering, pages 38–49, 2015.
  5. 5.Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Xiangru Tang, Mike Tian-Jian Jiang, and Alexander M. Rush. Promptsource: An integrated development environment and repository for natural language prompts, 2022.
  6. 6.Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma. Explaining neural scaling laws. arXiv preprint arXiv:2102.06701, 2021.
  7. 7.Yamini Bansal, Behrooz Ghorbani, Ankush Garg, Biao Zhang, Colin Cherry, Behnam Neyshabur, and Orhan Firat. Data scaling laws in nmt: The effect of noise and architecture. In International Conference on Machine Learning, pages 1466–1482. PMLR, 2022.
  8. 8.Zhengda Bian, Hongxin Liu, Boxiang Wang, Haichen Huang, Yongbin Li, Chuanrui Wang, Fan Cui, and Yang You. Colossal-ai: A unified deep learning system for large-scale parallel training. arXiv preprint arXiv:2110.14883, 2021.
  9. 9.Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raf. Emergent and predictable memorization in large language models. arXiv preprint arXiv:2304.11158, 2023.
  10. 10.Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. Piqa: Reasoning about physical commonsense in natural language. In Thirty-Fourth AAAI Conference on Artificial Intelligence, 2020.
  11. 11.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. Gpt-neox-20b: An open-source autoregressive language model. arXiv preprint arXiv:2204.06745, 2022.
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  13. 13.Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. Quantifying memorization across neural language models. arXiv preprint arXiv:2202.07646, 2022.
  14. 14.Thiago Castro Ferreira, Claire Gardent, Nikolai Ilinykh, Chris van der Lee, Simon Mille, Diego Moussallem, and Anastasia Shimorina. The 2020 bilingual, bi-directional webnlg+ shared task overview and evaluation results (webnlg+ 2020). In Proceedings of the 3rd WebNLG Workshop on Natural Language Generation from the Semantic Web (WebNLG+ 2020), pages 55–76, Dublin, Ireland (Virtual), 2020. Association for Computational Linguistics.
  15. 15.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In International conference on machine learning, pages 1691–1703. PMLR, 2020.
  16. 16.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311, 2022.
  17. 17.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Zhao, Yanping Huang, Andrew Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416, 2022. URL https://arxiv.org/abs/2210.11416.
  18. 18.Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. In NAACL, 2019.
  19. 19.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1, 2018.
  20. 20.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116, 2019.
  21. 21.Ido Dagan, Oren Glickman, and Bernardo Magnini. The pascal recognising textual entailment challenge. In Machine Learning Challenges. Evaluating Predictive Uncertainty, Visual Object Classification, and Recognising Tectual Entailment: First PASCAL Machine Learning Challenges Workshop, MLCW 2005, Southampton, UK, April 11-13, 2005, Revised Selected Papers, pages 177–190. Springer, 2006.
  22. 22.Marie-Catherine De Marneffe, Mandy Simons, and Judith Tonhauser. The commitmentbank: Investigating projection in naturally occurring discourse. In proceedings of Sinn und Bedeutung, volume 23, pages 107–124, 2019.
  23. 23.Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. Universal transformers. arXiv preprint arXiv:1807.03819, 2018.
  24. 24.Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. Scaling vision transformers to 22 billion parameters. arXiv preprint arXiv:2302.05442, 2023.
  25. 25.Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, et al. Glam: Efficient scaling of language models with mixture-of-experts. In International Conference on Machine Learning, pages 5547–5569. PMLR, 2022.
  26. 26.Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge. Computer Speech & Language, 59:123–156, January 2020. doi: 10.1016/j.csl.2019.06.009.
  27. 27.William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res, 23:1–40, 2021.
  28. 28.Samir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan, Mitchell Wortsman, Rulin Shao, Jean Mercat, Alex Fang, Jeffrey Li, Sedrick Keh, Rui Xin, Marianna Nezhurina, Igor Vasiljevic, Jenia Jitsev, Luca Soldaini, Alexandros G. Dimakis, Gabriel Ilharco, Pang Wei Koh, Shuran Song, Thomas Kollar, Yair Carmon, Achal Dave, Reinhard Heckel, Niklas Muennighoff, and Ludwig Schmidt. Language models scale reliably with over-training and on downstream tasks, 2024. URL https://arxiv.org/abs/2403.08540.
  29. 29.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The Pile: An 800GB dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  30. 30.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A framework for few-shot language model evaluation, September 2021. URL https://doi.org/10.5281/zenodo.5371628.
  31. 31.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. arXiv preprint arXiv:2009.11462, 2020.
  32. 32.Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D Dhole, et al. The gem benchmark: Natural language generation, its evaluation and metrics. arXiv preprint arXiv:2102.01672, 2021.
  33. 33.Behrooz Ghorbani, Orhan Firat, Markus Freitag, Ankur Bapna, Maxim Krikun, Xavier Garcia, Ciprian Chelba, and Colin Cherry. Scaling laws for neural machine translation. arXiv preprint arXiv:2109.07740, 2021.
  34. 34.Himanshu Gupta, Saurabh Arjun Sawant, Swaroop Mishra, Mutsumi Nakamura, Arindam Mitra, Santosh Mashetty, and Chitta Baral. Instruction tuned models are quick learners. arXiv preprint arXiv:2306.05539, 2023.
  35. 35.Kenneth Heafield. KenLM: Faster and smaller language model queries. In Proceedings of the Sixth Workshop on Statistical Machine Translation, pages 187–197, Edinburgh, Scotland, July 2011. Association for Computational Linguistics. URL https://aclanthology.org/W11-2123.
  36. 36.Peter Henderson, Mark Krass, Lucia Zheng, Neel Guha, Christopher D Manning, Dan Jurafsky, and Daniel Ho. Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset. Advances in Neural Information Processing Systems, 35:29217–29234, 2022.
  37. 37.Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al. Scaling laws for autoregressive generative modeling. arXiv preprint arXiv:2010.14701, 2020.
  38. 38.Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish. Scaling laws for transfer. arXiv preprint arXiv:2102.01293, 2021.
  39. 39.Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487, 2022.
  40. 40.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  41. 41.Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M Dai, Matthew D Hoffman, Monica Dinculescu, and Douglas Eck. Music transformer. arXiv preprint arXiv:1809.04281, 2018.
  42. 42.Berivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas, Sergei Vassilvitskii, and Sanmi Koyejo. Scaling laws for downstream task performance of large language models, 2024. URL https://arxiv.org/abs/2402.04177.
  43. 43.Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer. Summarizing source code using a neural attention model. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2073–2083, 2016.
  44. 44.Nikhil Kandpal, Eric Wallace, and Colin Raffel. Deduplicating training data mitigates privacy risks in language models, 2022.
  45. 45.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  46. 46.Mikhail Khrushchev, Ruslan Vasilev, Alexey Petrov, and Nikolay Zinov. YaLM 100B, 6 2022. URL https://github.com/yandex/YaLM-100B.
  47. 47.Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, et al. The stack: 3 tb of permissively licensed source code. arXiv preprint arXiv:2211.15533, 2022.
  48. 48.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. arXiv preprint arXiv:2205.11916, 2022.
  49. 49.Aran Komatsuzaki. One epoch is all you need. arXiv preprint arXiv:1906.06669, 2019.
  50. 50.Taku Kudo and John Richardson. SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 66–71, Brussels, Belgium, November 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-2012. URL https://aclanthology.org/D18-2012.
  51. 51.Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. Wikilingua: A new benchmark dataset for cross-lingual abstractive summarization. arXiv preprint arXiv:2010.03093, 2020.
  52. 52.Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. The bigscience roots corpus: A 1.6 tb composite multilingual dataset. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022.
  53. 53.Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499, 2021.
  54. 54.Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. Starcoder: may the source be with you! arXiv preprint arXiv:2305.06161, 2023.
  55. 55.Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  56. 56.Opher Lieber, Or Sharir, Barak Lenz, and Yoav Shoham. Jurassic-1: Technical details and evaluation. White Paper. AI21 Labs, 1, 2021.
  57. 57.Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004.
  58. 58.Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. Few-shot learning with multilingual language models. arXiv preprint arXiv:2112.10668, 2021.
  59. 59.Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, et al. Choosing transfer languages for cross-lingual learning. arXiv preprint arXiv:1905.12688, 2019.
  60. 60.Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, et al. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688, 2023a.
  61. 61.Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, Xinyi Wu, Enrico Shippole, Kurt Bollacker, Tongshuang Wu, Luis Villa, Sandy Pentland, Deb Roy, and Sara Hooker. The data provenance initiative: A large scale audit of dataset licensing & attribution in ai, 2023b.
  62. 62.Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, and Daphne Ippolito. A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity, 2023c.
  63. 63.Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian McAuley, Han Hu, Torsten Scholak, Sebastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, Mostofa Patwary, Nima Tajbakhsh, Yacine Jernite, Carlos Muñoz Ferrandis, Lingming Zhang, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. Starcoder 2 and the stack v2: The next generation, 2024. URL https://arxiv.org/abs/2402.19173.
  64. 64.Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Le Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki Heinonen, Aija Vahtola, Samuel Antao, and Sampo Pyysalo. FinGPT: Large generative models for a small language. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
  65. 65.Ali Madani, Bryan McCann, Nikhil Naik, Nitish Shirish Keskar, Namrata Anand, Raphael R Eguchi, Po-Ssu Huang, and Richard Socher. Progen: Language modeling for protein generation. arXiv preprint arXiv:2004.03497, 2020.
  66. 66.Bryan McCann, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher. The natural language decathlon: Multitask learning as question answering. CoRR, abs/1806.08730, 2018. URL http://arxiv.org/abs/1806.08730.
  67. 67.Sewon Min, Mike Lewis, Luke Zettlemoyer, and Hannaneh Hajishirzi. Metaicl: Learning to learn in context. arXiv preprint arXiv:2110.15943, 2021.
  68. 68.Nasrin Mostafazadeh, Michael Roth, Annie Louis, Nathanael Chambers, and James Allen. Lsdsem 2017 shared task: The story cloze test. In Proceedings of the 2nd Workshop on Linking Models of Lexical, Sentential and Discourse-level Semantics, pages 46–51, 2017.
  69. 69.Niklas Muennighoff. Vilio: State-of-the-art visio-linguistic models applied to hateful memes. arXiv preprint arXiv:2012.07788, 2020.
  70. 70.Niklas Muennighoff. Sgpt: Gpt sentence embeddings for semantic search. arXiv preprint arXiv:2202.08904, 2022.
  71. 71.Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, et al. Crosslingual generalization through multitask finetuning. arXiv preprint arXiv:2211.01786, 2022.
  72. 72.Niklas Muennighoff, Qian Liu, Armel Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre. Octopack: Instruction tuning code large language models. arXiv preprint arXiv:2308.07124, 2023.
  73. 73.Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. Deep double descent: Where bigger models and more data hurt. Journal of Statistical Mechanics: Theory and Experiment, 2021(12):124003, 2021a.
  74. 74.Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi. The deep bootstrap framework: Good online learners are good offline generalizers. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021b. URL https://openreview.net/forum?id=guetrIHLFGI.
  75. 75.Shashi Narayan, Shay B. Cohen, and Mirella Lapata. Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 2018.
  76. 76.Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al. Efficient large-scale language model training on gpu clusters using megatron-lm. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, pages 1–15, 2021.
  77. 77.Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. Adversarial nli: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2020.
  78. 78.Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. Codegen: An open large language model for code with multi-turn program synthesis. arXiv preprint arXiv:2203.13474, 2022.
  79. 79.nostalgebraist. chinchilla’s wild implications. lesswrong, 2022.
  80. 80.Gabriel Orlanski, Kefan Xiao, Xavier Garcia, Jeffrey Hui, Joshua Howland, Jonathan Malmaud, Jacob Austin, Rishah Singh, and Michele Catasta. Measuring the impact of programming language distribution. arXiv preprint arXiv:2302.01973, 2023.
  81. 81.Pedro Javier Ortiz Su’arez, Benoit Sagot, and Laurent Romary. Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures. Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-7) 2019. Cardiff, 22nd July 2019, pages 9 – 16, Mannheim, 2019. Leibniz-Institut f"ur Deutsche Sprache. doi: 10.14618/ids-pub-9021. URL http://nbn-resolving.de/urn:nbn:de:bsz:mh39-90215.
  82. 82.Pedro Javier Ortiz Su’arez, Laurent Romary, and Benoit Sagot. A monolingual approach to contextualized word embeddings for mid-resource languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1703–1714, Online, July 2020. Association for Computational Linguistics. URL https://www.aclweb.org/anthology/2020.acl-main.156.
  83. 83.Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155, 2022.
  84. 84.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116, 2023.
  85. 85.Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. Adding instructions during pretraining: Effective way of controlling toxicity in language models. arXiv preprint arXiv:2302.07388, 2023.
  86. 86.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019.
  87. 87.Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446, 2021.
  88. 88.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(140):1–67, 2020.
  89. 89.Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 3505–3506, 2020.
  90. 90.Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In AAAI spring symposium: logical formalizations of commonsense reasoning, pages 90–95, 2011.
  91. 91.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9):99–106, 2021.
  92. 92.Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al. Multitask prompted training enables zero-shot task generalization. In The Tenth International Conference on Learning Representations, 2022.
  93. 93.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022a.
  94. 94.Teven Le Scao, Thomas Wang, Daniel Hesslow, Lucile Saulnier, Stas Bekman, M Saiful Bari, Stella Bideman, Hady Elsahar, Niklas Muennighoff, Jason Phang, et al. What language model to train if you have one million gpu hours? arXiv preprint arXiv:2210.15424, 2022b.
  95. 95.Seongjin Shin, Sang-Woo Lee, Hwijeen Ahn, Sungdong Kim, HyoungSeok Kim, Boseop Kim, Kyunghyun Cho, Gichang Lee, Woomyoung Park, Jung-Woo Ha, et al. On the effect of pretraining corpora on in-context learning by a large-scale language model. arXiv preprint arXiv:2204.13509, 2022.
  96. 96.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv:1909.08053, 2019.
  97. 97.Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al. Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model. arXiv preprint arXiv:2201.11990, 2022.
  98. 98.Saleh Soltan, Shankar Ananthakrishnan, Jack FitzGerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith Peris, Stephen Rawls, Andy Rosenbaum, Anna Rumshisky, et al. Alexatm 20b: Few-shot learning using a large-scale multilingual seq2seq model. arXiv preprint arXiv:2208.01448, 2022.
  99. 99.Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S Morcos. Beyond neural scaling laws: beating power law scaling via data pruning. arXiv preprint arXiv:2206.14486, 2022.
  100. 100.Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. arXiv preprint arXiv:2206.04615, 2022.
  101. 101.Hui Su, Xiao Zhou, Houjing Yu, Yuwen Chen, Zilin Zhu, Yang Yu, and Jie Zhou. Welm: A well-read pre-trained language model for chinese. arXiv preprint arXiv:2209.10372, 2022.
  102. 102.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. Scale efficiently: Insights from pre-training and fine-tuning transformers. arXiv preprint arXiv:2109.10686, 2021.
  103. 103.Yi Tay, Mostafa Dehghani, Samira Abnar, Hyung Won Chung, William Fedus, Jinfeng Rao, Sharan Narang, Vinh Q Tran, Dani Yogatama, and Donald Metzler. Scaling laws vs model architectures: How does inductive bias influence scaling? arXiv preprint arXiv:2207.10551, 2022a.
  104. 104.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131, 2022b.
  105. 105.Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q Tran, David R So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, et al. Transcending scaling laws with 0.1% extra compute. arXiv preprint arXiv:2210.11399, 2022c.
  106. 106.Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085, 2022.
  107. 107.Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al. Lamda: Language models for dialog applications. arXiv preprint arXiv:2201.08239, 2022.
  108. 108.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  109. 109.Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho. Will we run out of data? an analysis of the limits of scaling datasets in machine learning. arXiv preprint arXiv:2211.04325, 2022.
  110. 110.Antti Virtanen, Jenna Kanerva, Rami Ilo, Jouni Luoma, Juhani Luotolahti, Tapio Salakoski, Filip Ginter, and Sampo Pyysalo. Multilingual is not enough: Bert for finnish. arXiv preprint arXiv:1912.07076, 2019.
  111. 111.Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. CoRR, abs/1905.00537, 2019. URL http://arxiv.org/abs/1905.00537.
  112. 112.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language model with self generated instructions. arXiv preprint arXiv:2212.10560, 2022.
  113. 113.Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, et al. How far can camels go? exploring the state of instruction tuning on open resources. arXiv preprint arXiv:2306.04751, 2023.
  114. 114.Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021.
  115. 115.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903, 2022.
  116. 116.Johannes Welbl, Nelson F Liu, and Matt Gardner. Crowdsourcing multiple choice science questions. arXiv preprint arXiv:1707.06209, 2017.
  117. 117.Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. Ccnet: Extracting high quality monolingual datasets from web crawl data. arXiv preprint arXiv:1911.00359, 2019.
  118. 118.Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M Rush, Bart Van Merriënboer, Armand Joulin, and Tomas Mikolov. Towards ai-complete question answering: A set of prerequisite toy tasks. arXiv preprint arXiv:1502.05698, 2015.
  119. 119.Mengzhou Xia, Guoqing Zheng, Subhabrata Mukherjee, Milad Shokouhi, Graham Neubig, and Ahmed Hassan Awadallah. Metaxl: Meta representation transformation for low-resource cross-lingual learning. arXiv preprint arXiv:2104.07908, 2021.
  120. 120.Mengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin, Ramakanth Pasunuru, Danqi Chen, Luke Zettlemoyer, and Ves Stoyanov. Training trajectories of language models across scales. arXiv preprint arXiv:2212.09803, 2022.
  121. 121.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. Wizardlm: Empowering large language models to follow complex instructions. arXiv preprint arXiv:2304.12244, 2023.
  122. 122.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer. arXiv preprint arXiv:2010.11934, 2020.
  123. 123.Ge Yang, Edward Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao. Tuning large neural networks via zero-shot hyperparameter transfer. Advances in Neural Information Processing Systems, 34:17084–17097, 2021.
  124. 124.Zheng-Xin Yong, Hailey Schoelkopf, Niklas Muennighoff, Alham Fikri Aji, David Ifeoluwa Adelani, Khalid Almubarak, M Saiful Bari, Lintang Sutawika, Jungo Kasai, Ahmed Baruwa, et al. Bloom+ 1: Adding language support to bloom for zero-shot prompting. arXiv preprint arXiv:2212.09535, 2022.
  125. 125.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.
  126. 126.Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al. Glm-130b: An open bilingual pre-trained model. arXiv preprint arXiv:2210.02414, 2022.
  127. 127.Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al. Pangu-alpha: Large-scale autoregressive pretrained chinese language models with auto-parallel computation. arXiv preprint arXiv:2104.12369, 2021.
  128. 128.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. Opt: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068, 2022.
  129. 129.Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. A survey of large language models. arXiv preprint arXiv:2303.18223, 2023.
  130. 130.Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al. Lima: Less is more for alignment. arXiv preprint arXiv:2305.11206, 2023.
  131. 131.Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. Designing effective sparse expert models. arXiv preprint arXiv:2202.08906, 2022.

Citation

MLA
Muennighoff, N., et al. “Scaling Data-Constrained Language Models”. Journal of Machine Learning Research, vol. 26, no. 53, 2025, pp. 1–6, https://www.jmlr.org/papers/v26/24-1000.html.
APA
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., & Raffel, C. (2025). Scaling Data-Constrained Language Models. Journal of Machine Learning Research, 26(53), 1–66. https://www.jmlr.org/papers/v26/24-1000.html
Chicago
Muennighoff, N., A. M. Rush, B. Barak, et al. 2025. “Scaling Data-Constrained Language Models”. Journal of Machine Learning Research 26 (53): 1–66. https://www.jmlr.org/papers/v26/24-1000.html.
Harvard
Muennighoff, N. et al. (2025) “Scaling Data-Constrained Language Models”, Journal of Machine Learning Research, 26(53), pp. 1–66. Available at: https://www.jmlr.org/papers/v26/24-1000.html.
Vancouver
1. Muennighoff N, Rush AM, Barak B, Scao TL, Piktus A, Tazi N, Pyysalo S, Wolf T, Raffel C (2025) Scaling Data-Constrained Language Models. Journal of Machine Learning Research 26:1–66

BibTeX

@article{JMLR:v26:24-1000,
  author  = {Niklas Muennighoff and Alexander M. Rush and Boaz Barak and Teven Le Scao and Aleksandra Piktus and Nouamane Tazi and Sampo Pyysalo and Thomas Wolf and Colin Raffel},
  title   = {Scaling Data-Constrained Language Models},
  journal = {Journal of Machine Learning Research},
  year    = {2025},
  volume  = {26},
  number  = {53},
  pages   = {1--66},
  url     = {http://jmlr.org/papers/v26/24-1000.html}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/