DsDm: Model-Aware Dataset Selection with Datamodels

Logan EngstromAxel FeldmannAleksander Madry

article2024ICML103 citations

Proposes a model-aware data selection framework using datamodels to optimize training subsets for downstream task performance, yielding language models that match the accuracy of standard baselines with half the training compute.

Listen

Training modern large-scale machine learning models requires massive amounts of text data, but standard industry heuristics for filtering this data rely heavily on intuitive notions of quality, such as text cleanliness or similarity to curated references like Wikipedia. The article addresses the emerging problem that these heuristic and similarity-based filters often fail in practice, frequently yielding models that perform no better than—or even worse than—those trained on randomly selected web data.

The main objective of the article is to demonstrate that pretraining data selection should be framed as an optimization problem aimed directly at maximizing model performance on target tasks, and to evaluate a new, model-aware selection framework that approximates this optimal subset.

To achieve this, the authors introduce a method called Dataset Selection with Datamodels (DSDM). Rather than evaluating text similarity, DSDM uses an efficient linear influence estimator to predict how including specific candidate documents will affect the model's loss on representative target tasks. The approach was evaluated by training GPT-style language models (ranging from 125 million to 1.8 billion parameters) on selections drawn from a massive public web crawl of over 200 million document chunks. To keep data curation feasible, small 125-million parameter proxy models were used to generate dataset rankings, which were then applied directly to larger training runs.

The analysis produced three primary findings. First, DSDM consistently improved model performance across targeted evaluation tasks, matching or exceeding the performance of models trained on random data with ten times the compute budget. Second, when proxy tasks were combined to select data for broad, general-purpose capabilities across fifteen diverse benchmarks, DSDM delivered a twofold compute multiplier, allowing a 1.3-billion parameter model to match the accuracy of a 1.8-billion parameter model trained on twice as much compute. Third, standard similarity-based baselines failed to beat random data selection across training budgets, revealing that text which intuitively looks high quality often trades away beneficial data diversity.

These findings indicate that data curation can serve as a highly effective intervention point in the model training pipeline to steer downstream capabilities and improve training efficiency. By achieving the same accuracy with half the training compute, organizations can substantially reduce training timelines, energy consumption, and infrastructure costs, while also uncovering counterintuitive training data that heuristics mistakenly discard.

Organizations training foundation models should consider adopting model-aware data selection rather than relying purely on qualitative filtering rules. Next steps should include testing representative target-task mixtures aligned with intended deployment goals and establishing small-scale proxy setups to evaluate influence scores before scaling to multi-billion parameter runs.

The primary limitations involve the computational overhead of computing gradients across candidate pools, although this cost can be amortized across multiple training lifecycles. Furthermore, while the linear model approximation transfers well across model scales, careful selection of target tasks remains necessary to avoid unexpected performance trade-offs in unrelated task categories.

arXiv: 2401.12926
Cover for DsDm: Model-Aware Dataset Selection with Datamodels

Abstract

When selecting data for training large-scale models, standard practice is to filter for examples that match human notions of data quality. Such filtering yields qualitatively clean datapoints that intuitively should improve model behavior. However, in practice the opposite can often happen: we find that selecting according to similarity with “high quality” data sources may not increase (and can even hurt) performance compared to randomly selecting data. To develop better methods for selecting data, we start by framing dataset selection as an optimization problem that we can directly solve for: given target tasks, a learning algorithm, and candidate data, select the subset that maximizes model performance. This framework thus avoids handpicked notions of data quality, and instead models explicitly how the learning process uses train datapoints to predict on the target tasks. Our resulting method greatly improves language model (LM) performance on both pre-specified tasks and previously unseen tasks. Specifically, choosing target tasks representative of standard LM problems and evaluating on diverse held-out benchmarks, our selected datasets provide a 2x compute multiplier over baseline methods.

Citation

MLA
Engstrom, L., et al. “DsDm: Model-Aware Dataset Selection with Datamodels”. arXiv, 2024, http://arxiv.org/abs/2401.12926v1.
APA
Engstrom, L., Feldmann, A., & Madry, A. (2024). DsDm: Model-Aware Dataset Selection with Datamodels. arXiv. http://arxiv.org/abs/2401.12926v1
Chicago
Engstrom, L., A. Feldmann, and A. Madry. 2024. “DsDm: Model-Aware Dataset Selection with Datamodels”. arXiv. http://arxiv.org/abs/2401.12926v1.
Harvard
Engstrom, L., Feldmann, A. and Madry, A. (2024) “DsDm: Model-Aware Dataset Selection with Datamodels”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.12926v1.
Vancouver
1. Engstrom L, Feldmann A, Madry A (2024) DsDm: Model-Aware Dataset Selection with Datamodels. arXiv

BibTeX

@article{engstrom2024dsdm,
  title = {DsDm: Model-Aware Dataset Selection with Datamodels},
  author = {Engstrom, Logan and Feldmann, Axel and Madry, Aleksander},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.12926v1},
  eprint = {2401.12926}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/