QuRating: Selecting High-Quality Data for Training Language Models
Alexander WettigAatmik GuptaSaumya MalikDanqi Chen
Presents a scalable data selection method that trains compact rating models on LLM pairwise judgments across qualitative criteria like educational value, enabling language models to match the performance of baselines trained on 50% more data.
Training highly capable large language models requires vast amounts of text, yet standard pre-training pipelines still depend on blunt heuristics or coarse domain filters to curate data. These conventional methods often fail to capture subtle qualitative dimensions of text and can inadvertently discard valuable knowledge. As high-quality web text becomes scarcer and computing costs escalate, developing scalable techniques that can identify genuinely useful training material has become a critical operational challenge.
The article introduces and evaluates QuRating, a framework designed to capture human intuitions about text quality to guide language model pre-training. Specifically, the article analyzes four criteria: educational value, writing style, factual density, and required expertise. To build this system, an automated judge evaluated 250,000 text pairs to establish relative quality preferences. These comparisons were used to train a compact 1.3-billion-parameter rating model that scored a 260-billion-token corpus. Using probabilistic temperature sampling to balance quality and diversity, the authors selected 30-billion-token subsets and trained new language models from scratch to assess downstream capabilities across ten diverse benchmark tasks.
The core findings demonstrate that educational value is the most powerful selection criterion, delivering an average gain of 1.8 percentage points in in-context learning across all ten benchmarks compared to uniform data selection. Notably, a model trained on educationally curated data achieved performance comparable to a standard model trained on 50 percent more data and compute. The analysis also revealed that probabilistic sampling consistently outperforms rigid top-tier filtering, which excessively strips away sample diversity and hurts general performance. Furthermore, optimizing for writing style achieved the lowest perplexity but delivered negligible gains on downstream tasks, proving that standard perplexity metrics do not reliably predict reasoning and knowledge capabilities. Finally, sequencing data from low to high required expertise created an effective curriculum that enhanced performance using the same underlying data.
These insights show that data selection grounded in educational and explanatory qualities can substantially reduce training timelines, energy usage, and compute costs while producing stronger models. Decision-makers should prioritize educational value over stylistic polish when designing pre-training datasets. Technical teams should adopt soft probabilistic sampling rather than hard quality thresholds to avoid catastrophic loss of data diversity, and they should leverage quality scores to structure progressive training curricula.
Leaders should nevertheless weigh several limitations before large-scale adoption. The experimental findings were established using 1.3-billion-parameter models, meaning validation at larger model scales is recommended before making major infrastructure commitments. Additionally, because the rating model inherits biases from automated judgments, the selection process can exhibit subtle geographic, social, and linguistic skews. Organizations deploying this approach should combine automated selection with deliberate bias evaluations and manual domain balancing.
- Paper: Textbooks Are All You Need, Suriya Gunasekar et al. (2023). Provides the foundational demonstration that filtering web corpora for high educational value dramatically enhances language model sample efficiency.
- Paper: Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection, Suchin Gururangan et al. (2022). Establishes how standard automated quality classifiers introduce demographic, stylistic, and ideological biases, motivating the need for more nuanced, multidimensional quality rating frameworks.
- Paper: Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Lianmin Zheng et al. (2023). Introduces and validates the methodology of using LLM pairwise comparisons as automated judges, which QuRating adapts to train its compact quality rater.
- Paper: The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data Only, Guilherme Penedo et al. (2023). Establishes the baseline heuristic and rule-based web data filtering pipelines that QuRating aims to replace with nuanced quality scoring.
- Paper: Pretraining Language Models with Human Preferences, Tomasz Korbak et al. (2023). Explores incorporating human preference modeling directly into the pre-training phase, establishing the conceptual groundwork for preference-guided pre-training data curation.
- Paper: Training Compute-Optimal Large Language Models, Jordan Hoffmann et al. (2022). Defines the standard compute-optimal pre-training scaling laws against which QuRating's data-efficiency gains are measured.
- Paper: DataComp-LM: In search of the next generation of training sets for language models, Jeffrey Li et al. (2024). Provides a large-scale, standardized experimental testbed to systematically benchmark model-based quality filtering and data curation techniques across compute scales.
- Paper: DsDm: Model-Aware Dataset Selection with Datamodels, Logan Engstrom et al. (2024). Extends pre-training data selection from attribute-based scoring models to influence-based datamodels optimized directly for downstream target tasks.
- Paper: A Bitter Lesson for Data Filtering, Christopher Mohri et al. (2026). Critically re-examines the long-term utility of pre-training data filters, showing that unfiltered web data can surpass filtered pools when compute and tokens scale freely.
- Paper: Scaling Data-Constrained Language Models, Niklas Muennighoff et al. (2025). Investigates how to allocate compute and repeat filtered, high-quality data over multiple epochs when facing severe web-data constraints.
- Paper: From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge, Dawei Li et al. (2025). Synthesizes advanced methodologies, biases, and applications of LLM-as-a-judge frameworks used to generate preference and quality annotations.
- Paper: Evaluating Large Language Models at Evaluating Instruction Following, Zhiyuan Zeng et al. (2024). Evaluates the vulnerability of LLM evaluators to superficial formatting and surface polish, addressing a key limitation noted in automated quality judgments.
