Built independently by an author, for readers. Read the story and support ChapterPal

keyword

AboutMe dataset

The AboutMe dataset is a collection of web-scraped self-descriptions and personal biographical profiles used to study demographic representation and the social impacts of data curation in natural language processing. By associating online text with creator-level metadata such as geographic affiliations, social roles, and topical interests, the dataset enables researchers to trace web content to its authors. It primarily serves as an analytical resource for evaluating how automated quality filters, language identification algorithms, and selection heuristics systematically alter the representation of diverse populations and demographic groups during the creation of pretraining corpora for large language models.

1 item

QuRating: Selecting High-Quality Data for Training Language Models

QuRating: Selecting High-Quality Data for Training Language Models

Alexander Wettig, Aatmik Gupta, Saumya Malik, Danqi Chen

OrganizationsPrinceton University

Why you should read this

Presents a scalable data selection method that trains compact rating models on LLM pairwise judgments across qualitative criteria like educational value, enabling language models to match the performance of baselines trained on 50% more data.

Selecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics. We introduce QuRating, a method for selecting pre-training data that can capture human intuitions about data quality. In this paper, we investigate four qualities—writing style, required expertise, facts & trivia, and educational value—and find that LLMs are able to discern these qualities, especially when making pairwise judgments of texts. We train a QuRater model to learn scalar ratings from pairwise judgments, and use it to annotate a 260B training corpus with quality ratings for each of the four criteria. In our experiments, we select 30B tokens according to the different quality ratings and train 1.3B-parameter language models on the selected data. We find that it is important to balance quality and diversity. When we sample using quality ratings as logits over documents, our models obtain lower perplexity and stronger in-context learning performance than baselines. Our best model is based on educational value and performs similarly to a model trained with uniform sampling for 50% more steps. Beyond data selection, we use the quality ratings to construct a training curriculum which improves performance without changing the training dataset. We extensively analyze the quality ratings and discuss their characteristics, biases, and wider implications.2

Added

2026-09-26