DsDm: Model-Aware Dataset Selection with Datamodels
Logan EngstromAxel FeldmannAleksander Madry
Proposes a model-aware data selection framework using datamodels to optimize training subsets for downstream task performance, yielding language models that match the accuracy of standard baselines with half the training compute.
Training modern large-scale machine learning models requires massive amounts of text data, but standard industry heuristics for filtering this data rely heavily on intuitive notions of quality, such as text cleanliness or similarity to curated references like Wikipedia. The article addresses the emerging problem that these heuristic and similarity-based filters often fail in practice, frequently yielding models that perform no better than—or even worse than—those trained on randomly selected web data.
The main objective of the article is to demonstrate that pretraining data selection should be framed as an optimization problem aimed directly at maximizing model performance on target tasks, and to evaluate a new, model-aware selection framework that approximates this optimal subset.
To achieve this, the authors introduce a method called Dataset Selection with Datamodels (DSDM). Rather than evaluating text similarity, DSDM uses an efficient linear influence estimator to predict how including specific candidate documents will affect the model's loss on representative target tasks. The approach was evaluated by training GPT-style language models (ranging from 125 million to 1.8 billion parameters) on selections drawn from a massive public web crawl of over 200 million document chunks. To keep data curation feasible, small 125-million parameter proxy models were used to generate dataset rankings, which were then applied directly to larger training runs.
The analysis produced three primary findings. First, DSDM consistently improved model performance across targeted evaluation tasks, matching or exceeding the performance of models trained on random data with ten times the compute budget. Second, when proxy tasks were combined to select data for broad, general-purpose capabilities across fifteen diverse benchmarks, DSDM delivered a twofold compute multiplier, allowing a 1.3-billion parameter model to match the accuracy of a 1.8-billion parameter model trained on twice as much compute. Third, standard similarity-based baselines failed to beat random data selection across training budgets, revealing that text which intuitively looks high quality often trades away beneficial data diversity.
These findings indicate that data curation can serve as a highly effective intervention point in the model training pipeline to steer downstream capabilities and improve training efficiency. By achieving the same accuracy with half the training compute, organizations can substantially reduce training timelines, energy consumption, and infrastructure costs, while also uncovering counterintuitive training data that heuristics mistakenly discard.
Organizations training foundation models should consider adopting model-aware data selection rather than relying purely on qualitative filtering rules. Next steps should include testing representative target-task mixtures aligned with intended deployment goals and establishing small-scale proxy setups to evaluate influence scores before scaling to multi-billion parameter runs.
The primary limitations involve the computational overhead of computing gradients across candidate pools, although this cost can be amortized across multiple training lifecycles. Furthermore, while the linear model approximation transfers well across model scales, careful selection of target tasks remains necessary to avoid unexpected performance trade-offs in unrelated task categories.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). It establishes the standardized benchmark and data curation baselines for language model pretraining that motivate the need for model-aware dataset selection.
- Paper: The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only, Guilherme Penedo et al. (2023). It introduces standard web-data filtering heuristics and datasets that DsDm directly analyzes and seeks to improve upon using influence estimation.
- Paper: DataComp: In search of the next generation of multimodal datasets, Samir Yitzhak Gadre et al. (2023). It provides foundational methodology on controlled, data-centric benchmarking for large-scale dataset filtering and evaluation across fixed compute budgets.
- Paper: Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks, Suchin Gururangan et al. (2020). It formalizes the relationship between targeted domain/task data selection and downstream language model evaluation performance.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). It defines compute-optimal scaling principles between parameter scale and training tokens that contextualize compute multiplier gains in pretraining curation.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). It investigates the asymptotic limits and compute scaling of web text filtering, providing a direct critique and continuation of pretraining selection strategies like DsDm.
- Paper: Data Shapley in One Training Run, Jiachen T. Wang et al. (2025). It advances data valuation and influence attribution during foundation model pretraining by computing data Shapley values in a single training run.
- Paper: Bridging Compute- and Data-Optimal Pretraining, Tian Qin et al.. It extends pretraining data-efficiency dynamics into compute- and data-bound regimes where token effectiveness diminishes across scales.
- Paper: Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs, Wei Zhou et al. (2026). It provides a broad survey of emerging LLM-driven and automated data preparation paradigms that generalize beyond classical heuristic filtering.
