Built independently by an author, for readers. Read the story and support ChapterPal

keyword

QuRater models

A QuRater model is a machine learning model designed to evaluate and assign continuous numerical quality scores to text documents based on specific qualitative criteria, such as educational value, writing style, factual density, or required expertise. Typically trained on pairwise comparative judgments generated by larger language models rather than absolute scoring heuristics, a QuRater model uses preference-modeling techniques to translate relative textual comparisons into calibrated scalar ratings. In natural language processing, these models are primarily employed to evaluate, filter, and sample massive datasets for training language models, helping balance data quality and topic diversity while facilitating the construction of effective learning curricula.

1 item

QuRating: Selecting High-Quality Data for Training Language Models

QuRating: Selecting High-Quality Data for Training Language Models

Alexander Wettig, Aatmik Gupta, Saumya Malik, Danqi Chen

OrganizationsPrinceton University

Why you should read this

Presents a scalable data selection method that trains compact rating models on LLM pairwise judgments across qualitative criteria like educational value, enabling language models to match the performance of baselines trained on 50% more data.

Selecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics. We introduce QuRating, a method for selecting pre-training data that can capture human intuitions about data quality. In this paper, we investigate four qualities—writing style, required expertise, facts & trivia, and educational value—and find that LLMs are able to discern these qualities, especially when making pairwise judgments of texts. We train a QuRater model to learn scalar ratings from pairwise judgments, and use it to annotate a 260B training corpus with quality ratings for each of the four criteria. In our experiments, we select 30B tokens according to the different quality ratings and train 1.3B-parameter language models on the selected data. We find that it is important to balance quality and diversity. When we sample using quality ratings as logits over documents, our models obtain lower perplexity and stronger in-context learning performance than baselines. Our best model is based on educational value and performs similarly to a model trained with uniform sampling for 50% more steps. Beyond data selection, we use the quality ratings to construct a training curriculum which improves performance without changing the training dataset. We extensively analyze the quality ratings and discuss their characteristics, biases, and wider implications.2

Added

2026-09-26