DataComp-LM: In search of the next generation of training sets for language models
Jeffrey LiAlex FangGeorgios SmyrnisMaor IvgiMatt JordanSamir Yitzhak GadreHritik BansalEtash Kumar GuhaSedrick Scott KehKushal Arora
Introduces a standardized benchmark and 240-trillion-token corpus for controlled data curation experiments, yielding an open training dataset that enables a 7B-parameter language model to match leading open-weight models while requiring up to six times less compute.
The rapid growth of large language models has led to soaring computational costs, shifting research focus toward optimizing training data quality rather than merely increasing scale. However, rigorous progress has been severely impeded by a lack of controlled comparisons, as researchers frequently evaluate datasets using inconsistent model architectures, hyperparameters, and compute budgets. Furthermore, details and underlying data for leading models remain largely closed and inaccessible to the public.
The article introduces DataComp for Language Models (DCLM), a standardized, large-scale experimental testbed designed to systematically evaluate data curation strategies. The primary objective is to isolate and demonstrate how techniques like text extraction, deduplication, quality filtering, and data mixing directly impact downstream language model performance across multiple compute scales.
To establish this benchmark, the researchers built an unfiltered corpus of 240 trillion tokens extracted from Common Crawl (DCLM-POOL) and created a suite of 53 evaluation tasks. They tested diverse curation methods across 416 controlled baseline experiments using standardized decoder-only Transformer models ranging from 412 million to 7 billion parameters. They systematically varied individual processing stages, comparing different extraction tools, deduplication algorithms, and model-based filtering approaches under identical training conditions.
The findings establish that dataset curation is decisive for model capability and training efficiency. Model-based quality filtering emerged as the most critical factor, where a fastText classifier trained on instruction-style reference data and set to retain the top 10% of documents produced the best results, boosting core evaluation scores by 3.5 percentage points over standard reference data. High-quality HTML text extraction using resiliparse delivered over a 2.5-point boost in core accuracy compared to standard pre-extracted web text while running eight times faster than competing tools. Combining these optimal techniques yielded DCLM-BASELINE, a dataset enabling a 7-billion parameter model to achieve 64% 5-shot accuracy on MMLU when trained on 2.6 trillion tokens. This represents a 6.6 percentage point gain over previous open-data models with 40% less compute, achieving parity with leading closed-data models like Llama 3 8B while using over six times less compute. In addition, performance rankings remained highly consistent across compute sizes (with correlation up to r = 0.98), proving that small-scale experiments reliably predict large-scale results.
These results demonstrate that disciplined data curation significantly reduces the cost, time, and computational resources needed to achieve frontier-level language model performance. Counter to common practice, mixing in secondary human-curated sources like Wikipedia actually diminished the performance of a rigorously filtered web corpus, showing that aggressive filtering of raw web data can outperform conventional mixing recipes.
Organizations developing or fine-tuning foundation models should adopt model-based filtering and high-performance text extraction while utilizing small 400M-to-1B parameter models to rapidly prototype data mixtures before committing large training budgets. Looking ahead, practitioners should expand these curation frameworks to target domain gaps such as mathematics, coding, safety, and multilingual capabilities.
Confidence in these findings is high due to the standardized testbed and extensive multi-scale ablations. However, readers should note that computational limits prevented full hyperparameter sweeps and multi-seed replications at the largest scales, and evaluation focused primarily on general natural language understanding rather than specialized code or math performance.
- Paper: DataComp: In search of the next generation of multimodal datasets, Samir Yitzhak Gadre et al. (2023). Introduces the foundational DataComp standardized benchmarking paradigm for controlled, data-centric pretraining experiments that DataComp-LM directly adapts and expands from multimodal to language models.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Establishes compute-optimal pretraining scaling laws that guide the experimental design and standardized token allocations in DataComp-LM.
- Paper: Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling, Stella Biderman et al. (2023). Demonstrates the necessity of standardized public model suites and controlled pretraining data sequences for reproducible scientific analysis of language model scaling.
- Paper: LLaMA: Open and Efficient Foundation Language Models, Hugo Touvron et al. (2023). Pioneered the methodology of training performant open foundation models entirely on publicly available web-scale corpora like Common Crawl.
- Paper: Mistral 7B, Albert Q. Jiang et al. (2023). Provides a crucial 7B architectural and performance baseline against which DataComp-LM measures the efficiency of its curated data recipes.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). Directly critiques and stress-tests the model-filtering premises established in DataComp-LM by evaluating whether DCLM-Baseline remains optimal when compute and training steps scale freely.
- Paper: Bridging Compute- and Data-Optimal Pretraining, Tian Qin et al.. Extends the study of pretraining data efficiency into data-constrained regimes by developing scaling laws for repeating and paraphrasing finite curated corpora.
- Paper: Tulu 3: Pushing Frontiers in Open Language Model Post-Training, Nathan Lambert et al. (2024). Builds on open base-model pretraining recipes by presenting a fully transparent post-training alignment pipeline to push open model capabilities further.
