keyword
training curriculum
A training curriculum is a structured optimization strategy in machine learning where training examples, tasks, or problem difficulties are presented to a model in a predetermined, progressive sequence rather than through uniform random sampling. Inspired by human pedagogical learning, this approach systematically schedules data exposure over the course of training, often transitioning from simpler, cleaner, or higher-quality inputs to progressively more challenging, noisy, or complex instances. By guiding the learning trajectory step-by-step, a training curriculum helps the model navigate complex loss landscapes, accelerate convergence, stabilize optimization, and achieve stronger overall generalization performance.
2 items

QuRating: Selecting High-Quality Data for Training Language Models
Alexander Wettig, Aatmik Gupta, Saumya Malik, Danqi Chen
Why you should read this
Presents a scalable data selection method that trains compact rating models on LLM pairwise judgments across qualitative criteria like educational value, enabling language models to match the performance of baselines trained on 50% more data.
Selecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics. We introduce QuRating, a method for selecting pre-training data that can capture human intuitions about data quality. In this paper, we investigate four qualities—writing style, required expertise, facts & trivia, and educational value—and find that LLMs are able to discern these qualities, especially when making pairwise judgments of texts. We train a QuRater model to learn scalar ratings from pairwise judgments, and use it to annotate a 260B training corpus with quality ratings for each of the four criteria. In our experiments, we select 30B tokens according to the different quality ratings and train 1.3B-parameter language models on the selected data. We find that it is important to balance quality and diversity. When we sample using quality ratings as logits over documents, our models obtain lower perplexity and stronger in-context learning performance than baselines. Our best model is based on educational value and performs similarly to a model trained with uniform sampling for 50% more steps. Beyond data selection, we use the quality ratings to construct a training curriculum which improves performance without changing the training dataset. We extensively analyze the quality ratings and discuss their characteristics, biases, and wider implications.2
Added
2026-09-26

Diffusion as a Training Curriculum for Timestep-Free Iterative Reasoning
Mariia Drozdova, Aidan Sirbu, Pietro Miotti, Robert Obryk, Mayalen Etcheverry, Eyvind Niklasson, Blake Richards
Why you should read this
Demonstrates that diffusion acts as a training curriculum rather than an inference procedure, allowing recurrent networks with persistent hidden states to achieve near-perfect accuracy on hard reasoning tasks through arbitrary-depth iteration and constant noise injection without requiring search, verifiers, or denoising schedules.
Diffusion models and recursive reasoners are both iterative, but they carry information across iterations differently. We add a persistent hidden state to a diffusion denoiser and remove its timestep conditioning, leaving a single shared update that can be run to arbitrary depth. The result is an anytime solver: accuracy keeps improving with inference depth far beyond the rollout lengths and backpropagation window used in training, reaching 99.90% exact solve on Sudoku-Extreme. We also obtain 98.93% solve rate on Maze-Unique. Surprisingly, progressive denoising is unnecessary at inference: holding corruption at its maximum by replacing every non-clue variable with fresh Gaussian noise at each step retains near-perfect solving and converges to stable solutions. This simple noise-injection mechanism enables a single trajectory to efficiently explore the solution space and settle on the correct answer without parallel rollouts, candidate selection, or external verifiers required by prior reasoning models. Nonetheless, ordered annealed corruption remains critical during training, which suggests that diffusion's primary contribution to our anytime solver is not a sampling procedure at inference, but a denoising training curriculum.
Added
2026-09-10

