Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum
Nived RajaramanAudrey HuangMiro DudíkRobert E. SchapireDylan FosterAkshay Krishnamurthy
Proves that autocurriculum methods drastically reduce the cost of training chain-of-thought reasoning models by requiring exponentially fewer supervised demonstrations and decoupling reinforcement learning compute from reference model quality.
Training modern language models to perform complex chain-of-thought reasoning requires enormous data and computational budgets. Current workflows rely heavily on gathering expensive expert reasoning demonstrations for supervised fine-tuning or running compute-intensive reinforcement learning pipelines that generate vast numbers of trial reasoning traces. As models grow, these scaling costs threaten to become unsustainable without algorithmic improvements.
The article demonstrates that autocurriculum—a framework where a learning model adaptively chooses which problems to practice based on its own ongoing performance—provably reduces the data and compute required to train reasoning models. Specifically, the analysis evaluates how adaptive prompt selection performs across both supervised fine-tuning and reinforcement learning with verifiable outcome rewards.
The authors analyze this problem using theoretical frameworks for autoregressive learning, drawing algorithmic inspiration from classical machine learning techniques including boosting-by-filtering and learning from counterexamples. In the supervised setting, the algorithm uses a cheap outcome verifier to check answers and selectively requests full reasoning demonstrations only for prompts the current model fails to solve. In the reinforcement learning setting, the learner applies an autocurriculum over a pre-trained reference model using rejection sampling, iteratively sharpening its capabilities on errors without requiring manual prompt filtering or structural assumptions about prompt difficulty.
The analysis establishes several primary findings. First, for supervised fine-tuning, autocurriculum achieves an exponential reduction in the number of required teacher reasoning demonstrations, dropping from a quantity that scales inversely with the target error to one that is nearly independent of target accuracy. Second, for reinforcement learning fine-tuning, autocurriculum decouples the total computational cost from the initial coverage quality of the reference model. Instead of multiplying total compute by coverage difficulty across all training steps, coverage costs are reduced to a fixed startup burn-in; beyond this initial phase, the compute required to drive higher accuracy matches that of an optimal reference model. Third, these sample and compute efficiencies hold for both deterministic and stochastic model classes.
These findings indicate that the steep costs currently associated with post-training reasoning models are not fundamental theoretical barriers, but rather artifacts of non-adaptive data collection. By shifting from static, uniform training pipelines to adaptive curricula guided by automated answer verification, organizations can significantly cut labeling expenses, reduce training runtimes, and lower overall compute overhead in domains with verifiable outcomes such as mathematics and software development.
Teams developing reasoning models should implement adaptive prompt-selection loops in their post-training architectures rather than relying on uniform, batch data collection. When high-accuracy guarantees are required under stochastic models, practitioners should combine adaptive ensembling with consensus voting at inference time. Further development is recommended to extend these theoretical frameworks to online policy gradient methods and to test whether self-improving curricula can expand coverage onto tasks where base models currently show zero success.
Confidence in these mathematical findings is high within the stated boundary conditions. However, decision-makers should note that the theoretical results assume access to a perfect outcome verifier and a realizable optimal model. Further research is necessary before deploying these exact guarantees in subjective or open-ended reasoning tasks where automated verification is imperfect or noisy.
- Paper: Curriculum learning, Yoshua Bengio et al. (2009). Read this foundational account of curriculum learning first to understand the staged-training ideas that the source adapts into model-driven selection of difficult prompts.
- Paper: Self-Paced Learning for Latent Variable Models, M. P. Kumar et al. (2010). Its self-paced learning method makes the key adaptive premise concrete: select examples according to the learner’s current ability, a direct precursor to autocurriculum.
- Paper: Active Prompting with Chain-of-Thought for Large Language Models, Shizhe Diao et al. (2024). Active-Prompt shows how model uncertainty can guide which reasoning problems receive costly supervision, clarifying the source’s adaptive selection strategy for SFT.
- Paper: Active Instruction Tuning: Improving Cross-Task Generalization by Training on Prompt Sensitive Tasks, Po-Nien Kung et al. (2023). Its prompt-sensitivity criterion illustrates how a model’s own signs of difficulty can prioritize training, preparing readers for the source’s focus on prompts where the model struggles.
- Paper: LESS: Selecting Influential Data for Targeted Instruction Tuning, Mengzhou Xia et al. (2024). LESS provides a targeted-data-selection approach for instruction tuning, useful groundwork for understanding the source’s claim that adaptive selection can reduce SFT data needs.
No sufficiently relevant recommendations were found.
