Built independently by an author, for readers. Read the story and support ChapterPal

keyword

data selection

Data selection is the process in machine learning of choosing a subset of training examples from a larger pool of candidate data to improve a model training efficiency, reduce computational costs, and enhance downstream performance. Rather than using an entire uncurated or web-scale corpus, which frequently contains noisy, redundant, or uninformative samples, practitioners apply data selection to identify instances based on criteria such as quality, task relevance, informativeness, diversity, or potential to reduce model loss. This curation can occur either prior to training through dataset filtering and heuristic scoring, or dynamically during training via specialized sampling and curriculum strategies, enabling models to achieve higher accuracy with fewer training steps.

4 items

Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt

Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt

Sören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma, Andreas Kirsch, Winnie Xu, Benedikt Höltgen, Aidan N. Gomez, Adrien Morisot, Sebastian Farquhar, Yarin Gal

OrganizationsCohereUniversity of OxfordUniversity of Toronto

Why you should read this

Proposes Reducible Holdout Loss Selection (RHO-LOSS), an effective data selection technique that accelerates deep model training by prioritizing examples that maximize generalization gains while filtering out noisy, irrelevant, and already learned points.

Training on web-scale data can take months. But most computation and time is wasted on redundant and noisy points that are already learnt or not learnable. To accelerate training, we introduce Reducible Holdout Loss Selection (RHO-LOSS), a simple but principled technique which selects approximately those points for training that most reduce the model's generalization loss. As a result, RHO-LOSS mitigates the weaknesses of existing data selection methods: techniques from the optimization literature typically select 'hard' (e.g. high loss) points, but such points are often noisy (not learnable) or less task-relevant. Conversely, curriculum learning prioritizes 'easy' points, but such points need not be trained on once learned. In contrast, RHO-LOSS selects points that are learnable, worth learning, and not yet learnt. RHO-LOSS trains in far fewer steps than prior art, improves accuracy, and speeds up training on a wide range of datasets, hyperparameters, and architectures (MLPs, CNNs, and BERT). On the large web-scraped image dataset Clothing-1M, RHO-LOSS trains in 18x fewer steps and reaches 2% higher final accuracy than uniform data shuffling.

Added

2026-09-28

QuRating: Selecting High-Quality Data for Training Language Models

QuRating: Selecting High-Quality Data for Training Language Models

Alexander Wettig, Aatmik Gupta, Saumya Malik, Danqi Chen

OrganizationsPrinceton University

Why you should read this

Presents a scalable data selection method that trains compact rating models on LLM pairwise judgments across qualitative criteria like educational value, enabling language models to match the performance of baselines trained on 50% more data.

Selecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics. We introduce QuRating, a method for selecting pre-training data that can capture human intuitions about data quality. In this paper, we investigate four qualities—writing style, required expertise, facts & trivia, and educational value—and find that LLMs are able to discern these qualities, especially when making pairwise judgments of texts. We train a QuRater model to learn scalar ratings from pairwise judgments, and use it to annotate a 260B training corpus with quality ratings for each of the four criteria. In our experiments, we select 30B tokens according to the different quality ratings and train 1.3B-parameter language models on the selected data. We find that it is important to balance quality and diversity. When we sample using quality ratings as logits over documents, our models obtain lower perplexity and stronger in-context learning performance than baselines. Our best model is based on educational value and performs similarly to a model trained with uniform sampling for 50% more steps. Beyond data selection, we use the quality ratings to construct a training curriculum which improves performance without changing the training dataset. We extensively analyze the quality ratings and discuss their characteristics, biases, and wider implications.2

Added

2026-09-26

DataComp: In search of the next generation of multimodal datasets

DataComp: In search of the next generation of multimodal datasets

Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Eyal Orgad, Rahim Entezari, Giannis Daras, Sarah M. Pratt, Vivek Ramanujan, Yonatan Bitton, Kalyani Marathe, Stephen Mussmann, Richard Vencu, Mehdi Cherti, Ranjay Krishna, Pang Wei Koh, Olga Saukh, Alexander J. Ratner, Shuran Song, Hannaneh Hajishirzi, Ali Farhadi, Romain Beaumont, Sewoong Oh, Alex Dimakis, Jenia Jitsev, Yair Carmon, Vaishaal Shankar, Ludwig Schmidt

OrganizationsAllen Institute for AIAppleColumbia UniversityForschungszentrum JülichGoogleGraz University of TechnologyLAIONSnorkel AITel Aviv UniversityThe Hebrew University of JerusalemUniversity of Illinois Urbana-ChampaignUniversity of Texas at AustinUniversity of Washington

Why you should read this

Introduces DataComp, a standardized benchmark centered on 12.8 billion candidate image-text pairs that allows researchers to rigorously evaluate multimodal dataset filtering methods across multiple compute scales and produce CLIP models that surpass OpenAI's original zero-shot ImageNet accuracy.

Multimodal datasets are a critical component in recent breakthroughs such as Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the ML ecosystem, we introduce DataComp, a testbed for dataset experiments centered around a new candidate pool of 12.8 billion image-text pairs from Common Crawl. Participants in our benchmark design new filtering techniques or curate new data sources and then evaluate their new dataset by running our standardized CLIP training code and testing the resulting model on 38 downstream test sets. Our benchmark consists of multiple compute scales spanning four orders of magnitude, which enables the study of scaling trends and makes the benchmark accessible to researchers with varying resources. Our baseline experiments show that the DataComp workflow leads to better training sets. In particular, our best baseline, DataComp-1B, enables training a CLIP ViT-L/14 from scratch to 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI's CLIP ViT-L/14 by 3.7 percentage points while using the same training procedure and compute. We release DataComp and all accompanying code at this http URL.

Added

2026-09-26