Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

Nived RajaramanAudrey HuangMiro DudíkRobert E. SchapireDylan FosterAkshay Krishnamurthy

article2026Annual Conference Computational Learning Theory2 citations

Proves that autocurriculum methods drastically reduce the cost of training chain-of-thought reasoning models by requiring exponentially fewer supervised demonstrations and decoupling reinforcement learning compute from reference model quality.

Listen

Training modern language models to perform complex chain-of-thought reasoning requires enormous data and computational budgets. Current workflows rely heavily on gathering expensive expert reasoning demonstrations for supervised fine-tuning or running compute-intensive reinforcement learning pipelines that generate vast numbers of trial reasoning traces. As models grow, these scaling costs threaten to become unsustainable without algorithmic improvements.

The article demonstrates that autocurriculum—a framework where a learning model adaptively chooses which problems to practice based on its own ongoing performance—provably reduces the data and compute required to train reasoning models. Specifically, the analysis evaluates how adaptive prompt selection performs across both supervised fine-tuning and reinforcement learning with verifiable outcome rewards.

The authors analyze this problem using theoretical frameworks for autoregressive learning, drawing algorithmic inspiration from classical machine learning techniques including boosting-by-filtering and learning from counterexamples. In the supervised setting, the algorithm uses a cheap outcome verifier to check answers and selectively requests full reasoning demonstrations only for prompts the current model fails to solve. In the reinforcement learning setting, the learner applies an autocurriculum over a pre-trained reference model using rejection sampling, iteratively sharpening its capabilities on errors without requiring manual prompt filtering or structural assumptions about prompt difficulty.

The analysis establishes several primary findings. First, for supervised fine-tuning, autocurriculum achieves an exponential reduction in the number of required teacher reasoning demonstrations, dropping from a quantity that scales inversely with the target error to one that is nearly independent of target accuracy. Second, for reinforcement learning fine-tuning, autocurriculum decouples the total computational cost from the initial coverage quality of the reference model. Instead of multiplying total compute by coverage difficulty across all training steps, coverage costs are reduced to a fixed startup burn-in; beyond this initial phase, the compute required to drive higher accuracy matches that of an optimal reference model. Third, these sample and compute efficiencies hold for both deterministic and stochastic model classes.

These findings indicate that the steep costs currently associated with post-training reasoning models are not fundamental theoretical barriers, but rather artifacts of non-adaptive data collection. By shifting from static, uniform training pipelines to adaptive curricula guided by automated answer verification, organizations can significantly cut labeling expenses, reduce training runtimes, and lower overall compute overhead in domains with verifiable outcomes such as mathematics and software development.

Teams developing reasoning models should implement adaptive prompt-selection loops in their post-training architectures rather than relying on uniform, batch data collection. When high-accuracy guarantees are required under stochastic models, practitioners should combine adaptive ensembling with consensus voting at inference time. Further development is recommended to extend these theoretical frameworks to online policy gradient methods and to test whether self-improving curricula can expand coverage onto tasks where base models currently show zero success.

Confidence in these mathematical findings is high within the stated boundary conditions. However, decision-makers should note that the theoretical results assume access to a perfect outcome verifier and a realizable optimal model. Further research is necessary before deploying these exact guarantees in subjective or open-ended reasoning tasks where automated verification is imperfect or noisy.

arXiv: 2603.18325

No sufficiently relevant recommendations were found.

Cover for Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

Abstract

Chain-of-thought reasoning, where language models expend additional computation by producing thinking tokens prior to final responses, has driven significant advances in model capabilities. However, training these reasoning models is extremely costly in terms of both data and compute, as it involves collecting long traces of reasoning behavior from humans or synthetic generators and further post-training the model via reinforcement learning. Are these costs fundamental, or can they be reduced through better algorithmic design? We show that autocurriculum, where the model uses its own performance to decide which problems to focus training on, provably improves upon standard training recipes for both supervised fine-tuning (SFT) and reinforcement learning (RL). For SFT, we show that autocurriculum requires exponentially fewer reasoning demonstrations than non-adaptive fine-tuning, by focusing teacher supervision on prompts where the current model struggles. For RL fine-tuning, autocurriculum decouples the computational cost from the quality of the reference model, reducing the latter to a burn-in cost that is nearly independent of the target accuracy. These improvements arise purely from adaptive data selection, drawing on classical techniques from boosting and learning from counterexamples, and requiring no assumption on the distribution or difficulty of prompts.

Table of Contents

  • 1 Introduction
  • 1.1 Contributions
  • 1.2 Organization
  • 2 Preliminaries
  • 3 SFT: Fine-Tuning with Teacher Supervision
  • 3.1 Prior Work: SFT without Curriculum
  • 3.2 Reducing the Cost of Supervision: Autocurriculum
  • 3.3 Main Result: Exponential Improvement in CoT Supervision via Autocurriculum
  • 3.4 Extension to General Models
  • 4 RL: Improving a Reference Model with Verifier Guidance
  • 4.1 Baseline: Learning via Rejection Sampling
  • 4.2 Main Result: Autocurriculum Decouples Coverage from Accuracy
  • 5 Discussion
  • References
  • A Related Work
  • B Proofs for Main Results
  • B.1 Autocurriculum for SFT (Deterministic Π\Pi): Proof of
  • B.1.1 Proof of
  • B.2 Autocurriculum for SFT (General Π\Pi): Proof of
  • B.2.1 Proof of
  • B.3 Improving a Reference Model Satisfying Coverage: Proof of
  • B.3.1 Proof of
  • B.4 Autocurriculum for Fine-Tuning a Reference Model: Proof of
  • C Proofs for Supporting Lemmas
  • C.1 Proofs for Lemmas from
  • C.1.1 Proof of
  • C.1.2 Proof of
  • C.1.3 Proof of
  • C.2 Proofs for Lemmas from
  • C.2.1 Proof of
  • C.2.2 Proof of
  • C.2.3 Proof of
  • C.2.4 Proof of
  • C.2.5 Proof of
  • C.3 Proofs for Lemmas from
  • C.3.1 Proof of

Knowls

  1. Knowl 1 — Deterministic SFT autocurriculum reduces teacher demonstrations exponentially

    theoretical result

    Suppose prompts are drawn i.i.d. from a distribution ρ\rho, the model class Π\Pi consists of deterministic autoregressive models with Natarajan dimension dd, the verifier has a unique correct answer for each prompt, and Π\Pi contains a model with perfect accuracy. For target error ε∈(0,1)\varepsilon\in(0,1) and failure probability δ∈(0,1/2)\delta\in(0,1/2), AutoTune achieves outcome accuracy at least 1−ε1-\varepsilon with probability at least 1−δ1-\delta when the prompt dataset has size n≥(d/ε) polylog⁡(ε−1,δ−1,T,∣Σ∣)n\ge (d/\varepsilon)\,\operatorname{polylog}(\varepsilon^{-1},\delta^{-1},T,|\Sigma|). It uses at most nn verifier queries and at most d polylog⁡(ε−1,δ−1,T,∣Σ∣)d\,\operatorname{polylog}(\varepsilon^{-1},\delta^{-1},T,|\Sigma|) teacher chain-of-thought (CoT) demonstrations. Its output is an improper predictor formed by taking a plurality vote over the ensemble's final answers. By comparison, non-adaptive next-token fine-tuning requires O~(d/ε)\widetilde O(d/\varepsilon) CoT demonstrations in the worst case. Thus adaptive selection makes the demonstration count nearly independent of target accuracy, without assumptions on prompt difficulty or distribution beyond sampling from ρ\rho.

  2. Knowl 2 — RL autocurriculum separates reference coverage cost from target accuracy

    theoretical result

    Suppose Π\Pi is deterministic with Natarajan dimension dd, contains a perfectly accurate model, the verifier has unique correct answers, and a fixed reference model assigns probability at least 1/Cseq1/C_{\mathrm{seq}} to the optimal length-TT reasoning trace for every prompt. AutoTune.RL uses a prompt dataset of size n≥(d/ε) polylog⁡(Cseq,ε−1,δ−1,T)n\ge (d/\varepsilon)\,\operatorname{polylog}(C_{\mathrm{seq}},\varepsilon^{-1},\delta^{-1},T) to return an outcome-level predictor with accuracy at least 1−ε1-\varepsilon with probability at least 1−δ1-\delta. Up to polylogarithmic factors in these parameters, both the total number of length-TT rollouts generated from the reference and learner models and the number of verifier calls are at most O~(dCseq+d/ε)\widetilde O(dC_{\mathrm{seq}}+d/\varepsilon). The non-curriculum rejection-sampling approach costs O~(dCseq/ε)\widetilde O(dC_{\mathrm{seq}}/\varepsilon) rollouts. Autocurriculum therefore confines the coverage-dependent expense to a burn-in term, while the cost of further accuracy improvement is nearly independent of reference coverage.

  3. Knowl 3 — AutoTune routes teacher supervision toward current ensemble errors

    algorithm

    AutoTune takes a prompt dataset, a perfect outcome verifier, a deterministic model class, and an SFT learner that trains from teacher CoTs. Set the weak-learner error target to e∗=1/4e_*=1/4 and the number of phases to k=⌈Clog⁡(1/ε)/e∗2⌉k=\lceil C\log(1/\varepsilon)/e_*^2\rceil, for a sufficiently large absolute constant CC. Divide the prompts into kk parts. In phase jj, for each prompt xx in that phase's part, let rj(x)r_j(x) be the number of the jj previously trained models whose final answers pass the verifier. Define weights by the backward recursion βrk,k=1{r≤k/2}\beta_r^{k,k}=\mathbf{1}\{r\le k/2\} and βrj,k=e∗βrj+1,k+(1−e∗)βr+1j+1,k\beta_r^{j,k}=e_*\beta_r^{j+1,k}+(1-e_*)\beta_{r+1}^{j+1,k} for j<kj<k, then set αrj,k=βrj+1,k−βr+1j+1,k\alpha_r^{j,k}=\beta_r^{j+1,k}-\beta_{r+1}^{j+1,k} and wj(x)=αrj(x)j,kw_j(x)=\alpha_{r_j(x)}^{j,k}. Accept each prompt with probability wj(x)/Mjw_j(x)/M_j, where Mj=max⁡0≤r≤jαrj,kM_j=\max_{0\le r\le j}\alpha_r^{j,k}; this is rejection sampling that upweights prompts with ensemble ranks favorable to correcting the final vote. Collect accepted prompts until there are enough for the base learner's sample complexity at error e∗e_* and failure probability δ/k\delta/k, or until that phase's prompt part is exhausted. Request teacher CoTs only for the accepted prompts, train one new model, and add it to the ensemble. Return the plurality of the kk models' final answers. With the prompt-dataset size in the deterministic SFT guarantee, the procedure uses O(log⁡(1/ε))O(\log(1/\varepsilon)) phases and at most d polylog⁡(ε−1,δ−1,T,∣Σ∣)d\,\operatorname{polylog}(\varepsilon^{-1},\delta^{-1},T,|\Sigma|) teacher demonstrations.

  4. Knowl 4 — Stochastic SFT autocurriculum guarantees moderate per-prompt accuracy

    theoretical result

    For a stochastic model class Π\Pi with a realizable perfectly accurate model, AutoTune's stochastic variant requires a prompt dataset of size n≥(log⁡∣Π∣/ε) polylog⁡(ε−1,δ−1,T)n\ge (\log|\Pi|/\varepsilon)\,\operatorname{polylog}(\varepsilon^{-1},\delta^{-1},T). With probability at least 1−δ1-\delta, it uses at most log⁡∣Π∣ polylog⁡(ε−1,δ−1,T)\log|\Pi|\,\operatorname{polylog}(\varepsilon^{-1},\delta^{-1},T) teacher CoTs and returns an outcome-level model f^\hat f satisfying Pr⁡x∼ρ(Acc⁡x(f^)≥3/5)≥1−ε\Pr_{x\sim\rho}(\operatorname{Acc}_x(\hat f)\ge 3/5)\ge 1-\varepsilon, where Acc⁡x\operatorname{Acc}_x is the probability that a sampled final answer passes the verifier. The procedure estimates each ensemble model's per-prompt accuracy from a sufficiently large constant number of independent answer samples; its rank counts models whose estimated accuracy is at least 9/109/10. It uses a boosting-style reweighting with weak error parameter 1/101/10 and terminal rank threshold 4k/54k/5, while training each model to base error 1/4001/400, and returns the uniform mixture of the trained models. It requires O(log⁡(1/ε))O(\log(1/\varepsilon)) calls to a log-loss ERM oracle. If correct answers are unique, drawing N=Ω(log⁡(1/ε))N=\Omega(\log(1/\varepsilon)) independent answers from this mixture and returning their consensus vote sharpens the guarantee to accuracy at least 1−O(ε)1-O(\varepsilon).

  5. Knowl 5 — Rejection sampling provides a reference-model baseline for RLVR

    algorithm

    Assume a deterministic realizable model class, unique correct answers, and a reference model that assigns probability at least 1/Cseq1/C_{\mathrm{seq}} to the optimal full reasoning trace for every prompt. For each of nn sampled prompts xix_i, RLFineTune draws and deduplicates m=Cseqlog⁡(4nCseq/δ)m=C_{\mathrm{seq}}\log(4nC_{\mathrm{seq}}/\delta) length-TT traces from the reference. It retains traces yy satisfying both V(xi,yT)=1V(x_i,y_T)=1 and πref(y∣xi)≥1/Cseq\pi_{\mathrm{ref}}(y\mid x_i)\ge 1/C_{\mathrm{seq}}. It returns a model in Π\Pi minimizing the fraction of prompts on which its generated trace is outside the retained set. For a sufficiently large absolute constant C1C_1, a prompt sample size n≥C1[dlog⁡(T∣Σ∣Cseq)log⁡(1/ε)+log⁡(1/δ)]/εn\ge C_1[d\log(T|\Sigma|C_{\mathrm{seq}})\log(1/\varepsilon)+\log(1/\delta)]/\varepsilon suffices for accuracy at least 1−ε1-\varepsilon with probability at least 1−δ1-\delta. The procedure generates at most O(nCseqlog⁡(nCseq/δ))O(nC_{\mathrm{seq}}\log(nC_{\mathrm{seq}}/\delta)) reference traces. Its reference-generation cost scales linearly in coverage, even though the prompt sample complexity depends only logarithmically on CseqC_{\mathrm{seq}}.

  6. Knowl 6 — Sequence-level coverage measures whether the reference can generate the target reasoning trace

    definition

    For a prompt xx, let yx∗y_x^* be the length-TT chain of thought generated by the optimal deterministic model. A reference model πref\pi_{\mathrm{ref}} has sequence-level coverage CseqC_{\mathrm{seq}} if πref(yx∗∣x)≥Cseq−1\pi_{\mathrm{ref}}(y_x^*\mid x)\ge C_{\mathrm{seq}}^{-1} for every prompt under consideration. The parameter CseqC_{\mathrm{seq}} quantifies the cost of obtaining a correct trace by sampling the reference: larger values mean weaker coverage. The main RL autocurriculum guarantee assumes this condition on every prompt; the paper also analyzes an extension where it holds on all but a small prompt mass.

  7. Knowl 7 — Outcome-level accuracy does not imply learning the teacher's full reasoning trace

    limitation

    The O~(d)\widetilde O(d) CoT-query guarantee for deterministic SFT concerns only whether the model predicts the correct final answer, as assessed by the outcome verifier. It does not guarantee that the model reproduces the teacher's intermediate reasoning tokens. If success is instead defined as matching the entire teacher CoT, the paper states that the query complexity degrades to Ω~(d/ε)\widetilde\Omega(d/\varepsilon): the outcome verifier supplies no supervision about intermediate tokens, so the final-answer autocurriculum cannot provide the same savings for full-trace imitation.

  8. Knowl 8 — Reference-model learning extends to partial sequence-level coverage

    theoretical result

    In the RLFineTune setting, suppose the reference covers the optimal full trace with probability at least 1/Cseq1/C_{\mathrm{seq}} on a prompt mass of at least 1−η1-\eta, rather than on every prompt. The paper's partial-coverage guarantee gives, for absolute constants C1,C2>0C_1,C_2>0, accuracy at least 1−C2(ε+η)1-C_2(\varepsilon+\eta) with probability at least 1−δ1-\delta when n≥C1[dlog⁡(TCseq)log⁡(n)+log⁡(1/δ)]/εn\ge C_1[d\log(TC_{\mathrm{seq}})\log(n)+\log(1/\delta)]/\varepsilon. RLFineTune generates at most O(nCseqlog⁡(nCseq/δ))O(nC_{\mathrm{seq}}\log(nC_{\mathrm{seq}}/\delta)) reference traces. Thus the uncovered prompt mass contributes additively to the achievable error, while the covered portion is learned using the same filtering approach.

  9. Knowl 9 — The theory assumes perfect verification, fixed coverage, and batch training

    limitation

    The analysis assumes a perfect outcome verifier. Its main RL autocurriculum result also uses a fixed reference model whose optimal-trace coverage holds on every prompt; when coverage fails, the filtering method does not learn those prompts, and the paper does not establish that training expands coverage to harder problems. The algorithms are batch procedures that train and then reevaluate on a fixed prompt dataset, rather than online policy-gradient methods in which the model and prompt distribution evolve together. The authors identify noisy or learned verifiers, self-play with improving reference coverage, and online RL as directions not resolved by these results.

Coverage note — No substantial contributed result was omitted; proof details and background or related-work discussion are excluded, and the paper reports no experiments.

References

  1. 1.Marah Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Javaheripi, Neel Joshi, et al. Phi-4-reasoning technical report. arXiv:2504.21318, 2025.
  2. 2.Dana Angluin. Learning regular sets from queries and counterexamples. Information and Computation, 1987.
  3. 3.Dana Angluin and Tyler Dohrn. The power of random counterexamples. In International Conference on Algorithmic Learning Theory, 2017.
  4. 4.Maria Florina Balcan and Steve Hanneke. Robust interactive learning. In Shie Mannor, Nathan Srebro, and Robert C. Williamson, editors, Proceedings of the 25th Annual Conference on Learning Theory, volume 23 of Proceedings of Machine Learning Research, pages 20.1–20.34, Edinburgh, Scotland, 25–27 Jun 2012. PMLR. URL https://proceedings.mlr.press/v23/balcan12c.html.
  5. 5.Peter L. Bartlett, Olivier Bousquet, and Shahar Mendelson. Local rademacher complexities. The Annals of Statistics, 2005.
  6. 6.Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. Curriculum learning. In Proceedings of the 26th International Conference on Machine Learning (ICML), 2009.
  7. 7.Fan Chen, Audrey Huang, Noah Golowich, Sadhika Malladi, Adam Block, Jordan T Ash, Akshay Krishnamurthy, and Dylan J Foster. The coverage principle: How pre-training enables post-training. arXiv preprint arXiv:2510.15020, 2025.
  8. 8.Mark Chen. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374, 2021.
  9. 9.Ehsan Emamjomeh-Zadeh and David Kempe. A general framework for robust interactive learning. Advances in Neural Information Processing Systems, 30, 2017.
  10. 10.Yang Fan, Fei Tian, Tao Qin, Xiang-Yang Li, and Tie-Yan Liu. Learning to teach. In International Conference on Learning Representations (ICLR), 2018. URL https://openreview.net/forum?id=HJewuJWCZ.
  11. 11.Dylan J Foster, Adam Block, and Dipendra Misra. Is behavior cloning all you need? understanding horizon in imitation learning. Advances in Neural Information Processing Systems, 2024.
  12. 12.Dylan J Foster, Zakaria Mhammedi, and Dhruv Rohatgi. Is a good foundation necessary for efficient reinforcement learning? the computational role of the base model in exploration. arXiv preprint arXiv:2503.07453, 2025.
  13. 13.Yoav Freund. Boosting a weak learning algorithm by majority. Information and Computation, 1995.
  14. 14.Yoav Freund and Robert E Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of computer and system sciences, 55(1):119–139, 1997.
  15. 15.Yoav Freund, Robert Schapire, and Naoki Abe. A short introduction to boosting. Journal-Japanese Society For Artificial Intelligence, 14(771-780):1612, 1999.
  16. 16.Zhaolin Gao, Joongwon Kim, Wen Sun, Thorsten Joachims, Sid Wang, Richard Yuanzhe Pang, and Liang Tan. Prompt curriculum learning for efficient llm post-training. arXiv preprint arXiv:2510.01135, 2025.
  17. 17.George Giapitzakis, Kimon Fountoulakis, Eshaan Nichani, and Jason D Lee. On the statistical query complexity of learning semiautomata: a random walk approach. arXiv preprint arXiv:2510.04115, 2025.
  18. 18.Grzegorz Gluch and Ruediger Urbanke. Exponential separation between two learning models and adversarial robustness. In Advances in Neural Information Processing Systems, 2021.
  19. 19.Etash Guha, Ryan Marten, Sedrick Keh, Negin Raoof, Georgios Smyrnis, Hritik Bansal, Marianna Nezhurina, Jean Mercat, Trung Vu, Zayne Sprague, et al. Openthoughts: Data recipes for reasoning models. arXiv:2506.04178, 2025.
  20. 20.Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, et al. Reinforced self-training (rest) for language modeling. arXiv preprint arXiv:2308.08998, 2023.
  21. 21.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, et al. Deepseek-r1 incentivizes reasoning in llms through reinforcement learning. Nature, 2025.
  22. 22.Guy Hacohen and Daphna Weinshall. On the power of curriculum learning in training deep networks, 2019. URL https://arxiv.org/abs/1904.03626.
  23. 23.Steve Hanneke. A bound on the label complexity of agnostic active learning. In Proceedings of the 24th international conference on Machine learning, pages 353–360, 2007.
  24. 24.Steve Hanneke and Liu Yang. Minimax analysis of active learning. J. Mach. Learn. Res., 16(1):3487–3602, 2015.
  25. 25.David Haussler and Philip M Long. A generalization of sauer’s lemma. Journal of Combinatorial Theory, Series A, 71(2):219–240, 1995.
  26. 26.Jingcheng Hu, Yinmin Zhang, Qi Han, Daxin Jiang, Xiangyu Zhang, and Heung-Yeung Shum. Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model. arXiv preprint arXiv:2503.24290, 2025.
  27. 27.Audrey Huang, Adam Block, Qinghua Liu, Nan Jiang, Akshay Krishnamurthy, and Dylan J Foster. Is best-of-n the best of them? coverage, scaling, and optimality in inference-time alignment. International Conference on Machine Learning (ICML), 2025.
  28. 28.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
  29. 29.Nan Jiang and Tengyang Xie. Offline reinforcement learning in large state spaces: Algorithms and guarantees. Statistical Science, 40(4):570–596, 2025.
  30. 30.Nirmit Joshi, Gal Vardi, Adam Block, Surbhi Goel, Zhiyuan Li, Theodor Misiakiewicz, and Nathan Srebro. A theory of learning with autoregressive chain of thought. arXiv:2503.07932, 2025.
  31. 31.Devvrit Khatri, Lovish Madaan, Rishabh Tiwari, Rachit Bansal, Sai Surya Duvvuri, Manzil Zaheer, Inderjit S Dhillon, David Brandfonbrener, and Rishabh Agarwal. The art of scaling reinforcement learning compute for llms. arXiv preprint arXiv:2510.13786, 2025.
  32. 32.M. Pawan Kumar, Benjamin Packer, and Daphne Koller. Self-paced learning for latent variable models. In Advances in Neural Information Processing Systems (NeurIPS), 2010.
  33. 33.Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, et al. Tulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024.
  34. 34.Nayoung Lee, Ziyang Cai, Avi Schwarzschild, Kangwook Lee, and Dimitris Papailiopoulos. Self-improving transformers overcome easy-to-hard and length generalization challenges. arXiv preprint arXiv:2502.01612, 2025.
  35. 35.Yong Lin, Shange Tang, Bohan Lyu, Jiayun Wu, Hongzhou Lin, Kaiyu Yang, Jia Li, Mengzhou Xia, Danqi Chen, Sanjeev Arora, et al. Goedel-prover: A frontier model for open-source automated theorem proving. arXiv preprint arXiv:2502.07640, 2025.
  36. 36.Mingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu, Xin Dong, Yejin Choi, Jan Kautz, and Yi Dong. Prorl: Prolonged reinforcement learning expands reasoning boundaries in large language models. arXiv preprint arXiv:2505.24864, 2025.
  37. 37.Deyu Meng, Qian Zhao, Lu Jiang, et al. What objective does self-paced learning indeed optimize?, 2015. URL https://arxiv.org/abs/1511.06049.
  38. 38.Sumeet Ramesh Motwani, Alesia Ivanova, Ziyang Cai, Philip Torr, Riashat Islam, Shital Shah, Christian Schroeder de Witt, and Charles London. h1: Bootstrapping llms to reason over longer horizons via reinforcement learning. arXiv preprint arXiv:2510.07312, 2025.
  39. 39.Balas K Natarajan. On learning sets and functions. Machine Learning, 4(1):67–97, 1989.
  40. 40.Team Olmo, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, et al. Olmo 3. arXiv preprint arXiv:2512.13961, 2025.
  41. 41.Gabriel Poesia, David Broman, Nick Haber, and Noah Goodman. Learning formal mathematics from intrinsic motivation. Advances in Neural Information Processing Systems, 37:43032–43057, 2024.
  42. 42.Jatin Prakash and Anirudh Buvanesh. What can you do when you have zero rewards during rl? arXiv preprint arXiv:2510.03971, 2025.
  43. 43.Yuxiao Qu, Matthew YR Yang, Amrith Setlur, Lewis Tunstall, Edward Emanuel Beeching, Ruslan Salakhutdinov, and Aviral Kumar. Optimizing test-time compute via meta reinforcement fine-tuning. arXiv preprint arXiv:2503.07572, 2025.
  44. 44.Nived Rajaraman, Audrey Huang, Miroslav Dudik, Robert Schapire, Dylan Foster, and Akshay Krishnamurthy. Learning to reason with curriculum ii: Compositional generalization, 2026. URL https://arxiv.org/abs/2606.27721.
  45. 45.Paria Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao, and Stuart Russell. Bridging offline reinforcement learning and imitation learning: A tale of pessimism. Advances in Neural Information Processing Systems, 34:11702–11716, 2021.
  46. 46.Dhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li, Ankur Moitra, Andrej Risteski, and Dylan J Foster. Taming imperfect process verifiers: A sampling perspective on backtracking. arXiv preprint arXiv:2510.03149, 2025.
  47. 47.Robert E Schapire and Yoav Freund. Boosting: Foundations and algorithms. MIT Press, 2013.
  48. 48.Amrith Setlur, Matthew YR Yang, Charlie Snell, Jeremy Greer, Ian Wu, Virginia Smith, Max Simchowitz, and Aviral Kumar. e3: Learning to explore enables extrapolation of test-time compute for llms. arXiv preprint arXiv:2506.09026, 2025.
  49. 49.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv:2402.03300, 2024.
  50. 50.Yuda Song, Gokul Swamy, Aarti Singh, J Bagnell, and Wen Sun. The importance of online data: Understanding preference fine-tuning via coverage. Advances in Neural Information Processing Systems, 37:12243–12270, 2024.
  51. 51.Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instructions. In Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers), pages 13484–13508, 2023.
  52. 52.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  53. 53.Wei Xiong, Chenlu Ye, Baohao Liao, Hanze Dong, Xinxing Xu, Christof Monz, Jiang Bian, Nan Jiang, and Tong Zhang. Reinforce-ada: An adaptive sampling framework under non-linear rl objectives. arXiv preprint arXiv:2510.04996, 2025.
  54. 54.Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. Wizardlm: Empowering large pre-trained language models to follow complex instructions. In The Twelfth International Conference on Learning Representations, 2024.
  55. 55.Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, et al. Dapo: An open-source llm reinforcement learning system at scale. arXiv preprint arXiv:2503.14476, 2025.
  56. 56.Zhiyuan Zeng, Hamish Ivison, Yiping Wang, Lifan Yuan, Shuyue Stella Li, Zhuorui Ye, Siting Li, Jacqueline He, Runlong Zhou, Tong Chen, et al. Rlve: Scaling up reinforcement learning for language models with adaptive verifiable environments. arXiv preprint arXiv:2511.07317, 2025.
  57. 57.Yuzhen Zhou, Jiajun Li, Yusheng Su, Gowtham Ramesh, Zilin Zhu, Xiang Long, Chenyang Zhao, Jin Pan, Xiaodong Yu, Ze Wang, Kangrui Du, Jialian Wu, Ximeng Sun, Jiang Liu, Qiaolin Yu, Hao Chen, Zicheng Liu, and Emad Barsoum. April: Active partial rollouts in reinforcement learning to tame long-tail generation, 2025. URL https://arxiv.org/abs/2509.18521.
  58. 58.Banghua Zhu, Michael Jordan, and Jiantao Jiao. Principled reinforcement learning with human feedback from pairwise or k-wise comparisons. In International Conference on Machine Learning, pages 43037–43067. PMLR, 2023.

Citation

MLA
Rajaraman, N., et al. “Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum”. arXiv, 2026, http://arxiv.org/abs/2603.18325v2.
APA
Rajaraman, N., Huang, A., Dudik, M., Schapire, R., Foster, D. J., & Krishnamurthy, A. (2026). Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum. arXiv. http://arxiv.org/abs/2603.18325v2
Chicago
Rajaraman, N., A. Huang, M. Dudik, R. Schapire, D. J. Foster, and A. Krishnamurthy. 2026. “Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum”. arXiv. http://arxiv.org/abs/2603.18325v2.
Harvard
Rajaraman, N. et al. (2026) “Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2603.18325v2.
Vancouver
1. Rajaraman N, Huang A, Dudik M, Schapire R, Foster DJ, Krishnamurthy A (2026) Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum. arXiv

BibTeX

@article{rajaraman2026learning,
  title = {Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum},
  author = {Rajaraman, Nived and Huang, Audrey and Dudik, Miro and Schapire, Robert and Foster, Dylan J. and Krishnamurthy, Akshay},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2603.18325v2},
  eprint = {2603.18325}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/