Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning

Wenkai YangShuming MaYankai LinFuru Wei

article2025NeurIPS141 citations

Reveals that excessively scaling chain-of-thought length can impair mathematical reasoning and introduces a test-time compute strategy that learns domain-optimal thinking lengths to match the performance of leading reasoning models on competitive benchmarks.

Listen

Recent advances in artificial intelligence have focused on test-time scaling, a method that encourages large language models to spend more computational time "thinking" by generating extended chains of thought before answering. While this approach has improved performance on complex reasoning tasks, it introduces severe inefficiencies and potential performance risks. The article addresses whether excessively extending these reasoning chains causes adverse effects beyond mere computational waste, investigating how reasoning length directly influences model accuracy across tasks of varying difficulty.

The main objective of the article is to demonstrate that over-scaling reasoning lengths can actively impair accuracy, especially on simpler problems, and to propose a practical strategy that dynamically scales reasoning effort based on problem difficulty.

To evaluate this, the authors analyzed leading reasoning models across standardized mathematics and general knowledge benchmarks. They then created a controlled experimental setup using open-source models (including 8-billion and 32-billion parameter models). The researchers fine-tuned an initial base model on a small seed dataset (around 1,300 problems) with three distinct reasoning lengths—low, medium, and high. This produced an intermediate model capable of adjusting its reasoning effort on command. They then used this model to generate multiple candidate solutions across tens of thousands of problems and selected the shortest correct answer for each to train the final self-improved model.

The findings reveal that longer chains of thought do not universally improve accuracy and can actively degrade performance on easier tasks. Detailed analysis shows that longer reasoning paths accumulate more intermediate errors; although learning error correction is beneficial, training on excessive erroneous steps harms model reasoning. Furthermore, there is an optimal reasoning length for each task difficulty: lower effort succeeds best on straightforward problems, while higher effort is necessary only for complex tasks. Applying this insight via the proposed strategy—termed Thinking-Optimal Scaling (TOPS)—enabled a 32-billion parameter model to achieve 95.82% accuracy on basic math (GSM8K) and 46.00% on advanced competition math (AIME 2024) after preference optimization, matching or exceeding much larger or heavily distilled models while consuming far fewer tokens.

These results have direct operational and cost implications. For organizations deploying reasoning models, unchecked test-time compute inflates operational costs and latency without guaranteeing higher accuracy. Implementing adaptive reasoning depth significantly reduces inference compute expenses and mitigates the risk of hallucination or circular reasoning caused by overthinking simple prompts.

Decision-makers and engineering teams should avoid uniform prompts that force maximum reasoning depth across all queries. Instead, deployments should adopt adaptive pipelines that select the most concise valid reasoning path. When curating training data from long-thinking models, teams should apply loss masking or pruning to suppress erroneous reasoning steps rather than training unconditionally on full reasoning traces.

The study's primary limitations include a heavy focus on mathematical reasoning benchmarks, with preliminary rather than comprehensive validation across broader natural language domains. Additionally, the evaluations were conducted predominantly within supervised fine-tuning environments rather than pure reinforcement learning frameworks. Nevertheless, the empirical findings are supported by consistent results across multiple base models and benchmarks, providing strong confidence in the core conclusion that reasoning compute must be tailored to task difficulty.

  • Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). This compute-optimal framework establishes how inference-time compute can be allocated by task difficulty, the foundation for understanding why the source seeks domain-specific optimal reasoning lengths.
  • Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). Its budget-forcing method and positive results from extending reasoning provide the test-time scaling baseline that the source qualifies by finding that longer thinking can sometimes hurt.
  • Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Its teacher-generated reasoning distillation pipeline introduces the core training approach that the source adapts to teach models different reasoning efforts.
  • Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). STaR’s iterative self-training on correct model-generated rationales prepares readers for the source’s use of model self-improvement after reasoning-length training.
Cover for Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning

Abstract

Recent studies have shown that making a model spend more time thinking through longer Chain of Thoughts (CoTs) enables it to gain significant improvements in complex reasoning tasks. While current researches continue to explore the benefits of increasing test-time compute by extending the CoT lengths of Large Language Models (LLMs), we are concerned about a potential issue hidden behind the current pursuit of test-time scaling: Would excessively scaling the CoT length actually bring adverse effects to a model's reasoning performance? Our explorations on mathematical reasoning tasks reveal an unexpected finding that scaling with longer CoTs can indeed impair the reasoning performance of LLMs in certain domains. Moreover, we discover that there exists an optimal scaled length distribution that differs across different domains. Based on these insights, we propose a Thinking-Optimal Scaling strategy. Our method first uses a small set of seed data with varying response length distributions to teach the model to adopt different reasoning efforts for deep thinking. Then, the model selects its shortest correct response under different reasoning efforts on additional problems for self-improvement. Our self-improved models built upon Qwen2.5-32B-Instruct outperform other distillation-based 32B o1-like models across various math benchmarks, and achieve performance on par with the teacher model QwQ-32B-Preview that produces the seed data.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 3 The impact of scaling efforts on the effectiveness of test-time scaling
  • 3.1 Preliminary analysis on existing o1-like models
  • 3.2 Deeper explorations on the scaling process of CoT length
  • 3.3 Analysis on the adverse effects of excessive length scaling
  • 4 Thinking-optimal test-time scaling
  • 5 Experiments and analysis
  • 5.1 Experimental settings
  • 5.2 Main results
  • 5.3 Results of iterative self-improvement
  • 5.4 Results on LLaMA3.1-8B-Instruct
  • 6 Conclusion
  • 7 Limitations
  • References
  • A Impact statement
  • B Details on estimating the number of tokens in hidden CoT of o1-mini by Qwen2.5 tokenizer
  • C Prompts for generating reasoning responses under different reasoning efforts
  • D User prompt for gpt-4o to determine the number of (erroneous) reasoning rounds
  • E Experimental settings
  • E.1 Training settings for tag models
  • E.2 Training settings in format imitation
  • E.3 Training settings in self-improvement
  • E.4 Training settings in iterative self-improvement
  • E.5 Evaluation settings
  • F Breakdown results of tag models on MATH500
  • G Response length distribution of QwQ-32B-Preview under different reasoning effort-based system prompts
  • H Performance of QwQ-32B-Preview and QwQ-32B when directly prompted with different reasoning effort-based prompts
  • I Results of fine-tuning on samples after filtering out solutions with erroneous steps
  • J Results on general reasoning tasks
  • K Standard deviation results

Knowls

  1. Knowl 1 — Thinking-Optimal Scaling selects a model-generated shortest correct reasoning path

    model/method

    Thinking-Optimal Scaling (TOPS) is a self-improvement method for teaching a language model to spend different amounts of test-time reasoning effort according to the problem. It has three stages:

    1. Format imitation: Fine-tune the base model on a small seed set of System-2-style solutions supplied under different reasoning-effort conditions. This teaches the model to produce reasoning paths of varying lengths.
    2. Reasoning-effort-conditioned generation: Use the resulting tag model to solve additional problems under each effort condition.
    3. Self-improvement: For each problem, retain the shortest generated response whose final answer is correct, then supervised-fine-tune the original base model on the retained problem-response pairs.

    The selection criterion is intended to avoid both insufficient reasoning, which may yield an incorrect answer, and unnecessary continuation after a correct answer is available, which may increase overthinking and erroneous steps. TOPS therefore learns from the model's own correctness-checked outputs rather than assuming that one fixed response-length distribution is appropriate for every problem.

  2. Knowl 2 — Definition of a System-2 thinking-optimal response

    definition

    A System-2 thinking-optimal response is the shortest response a model can generate using deliberate reasoning that still gives the correct answer. In the paper's operational framing, a shorter response may reflect underthinking and produce a wrong answer, whereas continuing beyond the shortest correct response may add unnecessary reasoning or overthinking.

  3. Knowl 3 — Reasoning effort has task-dependent optima, and excessive effort can reduce accuracy

    empirical result

    The authors trained LLaMA3.1-8B and Qwen2.5-32B tag models on the same problems paired with short, medium, and long QwQ-32B-Preview solutions. To control for the original effort prompts not reliably determining output length, they reordered each problem's responses by actual length and retained cases where adjacent lengths differed by more than 300 tokens. Accuracy is reported as a percentage; “distinct answers” is the average number of different answers among five samples per prompt. The results show that low effort was best on GSM8K for both models, while medium effort was best on MATH500 for both; the best AIME2024 effort differed between the two models. In particular, the high-effort Qwen tag model was less accurate than its low-effort counterpart on GSM8K and less accurate than its medium-effort counterpart on MATH500 and AIME2024, despite using longer responses. The lowest answer diversity often coincided with the best-performing effort, though not in every model-benchmark pair.

    Model GSM8K Acc. Distinct MATH500 Acc. Distinct AIME2024 Acc. Distinct
    LLaMA3.1-8B-Tag-Low 87.26 1.37 59.00 2.34 7.33 4.27
    LLaMA3.1-8B-Tag-Medium 87.06 1.42 61.12 2.54 7.33 4.27
    LLaMA3.1-8B-Tag-High 86.89 1.47 59.36 2.59 10.00 4.00
    Qwen2.5-32B-Tag-Low 95.53 1.33 90.60 1.39 34.67 3.20
    Qwen2.5-32B-Tag-Medium 94.33 1.14 91.48 1.36 42.00 3.13
    Qwen2.5-32B-Tag-High 93.31 1.15 90.56 1.40 41.33 3.30

    These controlled results support a domain- and difficulty-dependent optimum, rather than a rule that more reasoning tokens always improve performance.

  4. Knowl 4 — TOPS improves the accuracy-efficiency trade-off on Qwen2.5-32B

    empirical result

    On GSM8K (1,319 grade-school problems), MATH500 (500 problems), and AIME2024 (30 problems), Qwen2.5-32B-TOPS was evaluated at decoding temperature 1.0 and its results were averaged over five random seeds. Accuracy is in percent and token counts are the reported average response-token counts. Compared with the random-correct-response self-training baseline, TOPS used fewer tokens and had higher accuracy on all three benchmarks. Compared with the distillation-based STILL-2-32B and Sky-T1-32B-Preview models, TOPS was more accurate on GSM8K and MATH500; on AIME2024, its 43.33% accuracy was below STILL-2-32B's 45.33%. For STILL-2-32B and Sky-T1-32B-Preview, the reported token count covers the thought portion, not the summary portion.

    Model GSM8K Acc. Tokens MATH500 Acc. Tokens AIME2024 Acc. Tokens
    Qwen2.5-32B-Instruct (Temp. 0.0) 95.91 295.01 84.20 576.89 16.67 1407.43
    Qwen2.5-32B-Instruct (Temp. 1.0) 95.30 296.98 82.84 555.65 14.67 855.62
    QwQ-32B-Preview 95.23 761.01 92.02 2416.23 45.33 7636.63
    STILL-2-32B 95.47 570.64 91.40 2005.28 45.33 6656.11
    Sky-T1-32B-Preview 94.82 695.66 89.48 2022.07 35.33 5351.29
    Qwen2.5-32B-Random 95.00 938.45 90.16 2670.19 39.33 7691.30
    Qwen2.5-32B-TOPS 95.82 412.24 91.48 1883.29 43.33 7260.26

    TOPS used fewer tokens than QwQ-32B-Preview on GSM8K and MATH500 while approaching its accuracy, and concentrated substantially more reasoning on the harder AIME2024 task than on GSM8K.

  5. Knowl 5 — TOPS training data and implementation

    experimental setup

    For the principal Qwen2.5-32B experiment, the tag model was trained on about 1.3K NuminaMath problems with three QwQ-32B-Preview responses per problem—about 3.9K responses total—under low, medium, and high reasoning-effort conditions. The corresponding Qwen2.5 tag set contained 1,312 problems; its mean response lengths were 1,588.23, 2,535.65, and 3,767.92 tokens for low, medium, and high effort. The authors generated one response per effort condition for each of an additional 50K NuminaMath problems, retained the shortest correct response when one was available, and added the low-effort seed responses to form a self-improvement set of about 26K samples.

    The tag model was fine-tuned with learning rate 1×10−51\times10^{-5}, batch size 32, and three epochs. The TOPS base model was fine-tuned with learning rate 1×10−51\times10^{-5}, batch size 96, and two epochs. Training used NVIDIA H100 GPUs. Evaluation of reasoning models used temperature 1.0, a maximum generation length of 16,384 tokens, and five random seeds. The setup instantiates the method using a modest set of effort-conditioned seed examples followed by correctness-based selection over a much larger problem set.

  6. Knowl 6 — Excessive reasoning is associated with more erroneous reasoning rounds

    empirical result

    To investigate why longer reasoning could harm learning, the authors examined responses for 100 randomly selected tag-model training problems. GPT-4o was used to count complete reasoning or verification rounds and identify rounds containing an erroneous step or wrong final answer. Although the evaluated responses all had correct final answers, the average number of rounds and the number and proportion of erroneous rounds increased as the prompted effort rose from low to high. This is consistent with the authors' explanation that some reflection and correction can be useful, but training on an excess of erroneous steps may be harmful.

    A separate filtering experiment removed high-effort Qwen2.5-32B-Tag training solutions identified by GPT-4.1 as containing erroneous steps. Filtering shortened generated responses and improved GSM8K accuracy, but reduced accuracy on the harder MATH500 and AIME2024 benchmarks. The result suggests that removing entire solutions can also remove useful learning about deeper reflection and correction. In a controlled LLaMA3.1-8B experiment with long chains containing initial mistakes and subsequent corrections, masking the loss on identified erroneous steps performed better on MATH500 than training on all steps. Together, these findings favor retaining correction behavior without directly learning from the erroneous steps themselves.

    Model GSM8K Acc. Tokens MATH500 Acc. Tokens AIME2024 Acc. Tokens
    Qwen2.5-32B-Tag-High 93.31 1820.17 90.56 3185.75 41.33 8753.87
    Qwen2.5-32B-Tag-High-Filtered 94.87 1478.10 90.00 2783.98 36.33 8049.29
  7. Knowl 7 — Iterative self-improvement with shortest-correct choices and preference optimization

    algorithm

    The iterative TOPS procedure starts from Qwen2.5-32B-TOPS and uses additional 4,500 MATH problems plus AIME problems from 1983–2023. It samples eight responses per problem and chooses the shortest correct response. One variant supervised-fine-tunes on these chosen responses. The preference-optimization variant also forms rejected examples from incorrect responses: it uses the longest incorrect response when available, and includes the shortest incorrect response if it is shorter than the chosen correct one, to discourage underthinking. The authors apply Direct Preference Optimization (DPO) to the resulting preference pairs.

    Further supervised fine-tuning mainly shortened responses and did not consistently raise accuracy. The DPO variant improved the overall efficiency-effectiveness balance and reached 46.00% on AIME2024, compared with 43.33% for the initial TOPS model and 45.33% for QwQ-32B-Preview. The reported metrics were:

    Model GSM8K Acc. Tokens MATH500 Acc. Tokens AIME2024 Acc. Tokens
    Qwen2.5-32B-TOPS 95.82 412.24 91.48 1883.29 43.33 7260.26
    Qwen2.5-32B-TOPS-Iter-SFT 95.45 366.14 90.76 1701.11 44.00 6611.89
    Qwen2.5-32B-TOPS-Iter-DPO 95.80 384.81 91.60 1731.72 46.00 6426.62

    The iterative SFT used one epoch, learning rate 1×10−61\times10^{-6}, and batch size 32. Iterative DPO used three epochs, learning rate 5×10−75\times10^{-7}, and batch size 32.

  8. Knowl 8 — TOPS also improves LLaMA3.1-8B over random response selection

    empirical result

    The authors applied the same TOPS self-improvement approach to LLaMA3.1-8B-Instruct. Relative to random selection among correct responses, TOPS-SFT achieved higher accuracy on GSM8K, MATH500, and AIME2024, while using fewer reported tokens on each benchmark. This supports transfer of the approach beyond the Qwen model family. Accuracy is in percent; token counts and benchmark conditions follow the paper's evaluation setup.

    Model GSM8K Acc. Tokens MATH500 Acc. Tokens AIME2024 Acc. Tokens
    LLaMA3.1-8B-Instruct (Temp. 0.0) 82.18 262.23 47.00 1801.76 6.67 5506.30
    LLaMA3.1-8B-Instruct (Temp. 1.0) 76.21 233.08 39.60 733.56 4.67 1691.88
    LLaMA3.1-8B-Random-SFT 87.94 1051.05 60.52 3627.23 4.67 8165.69
    LLaMA3.1-8B-TOPS-SFT 88.54 571.10 61.28 3254.01 8.00 7392.59
  9. Knowl 9 — General-reasoning experiments show effort-dependent performance and TOPS gains

    empirical result

    The authors extended the effort-conditioned generation and TOPS comparison to general reasoning. They used effort-varied responses generated for WebInstruct-verified examples to train Qwen2.5-7B-Tag-General, then evaluated on MMLU-Pro and GPQA-Diamond, sampling once per prompt. Higher effort increased token use but did not monotonically improve accuracy: MMLU-Pro accuracy was nearly unchanged across the three effort settings, while medium effort performed best on GPQA-Diamond. On held-out general-reasoning data, TOPS outperformed random correct-response selection on accuracy for both benchmarks and used fewer tokens than the effort-conditioned tag variants. The results are preliminary evidence beyond mathematics, not a comprehensive evaluation of all reasoning domains.

    Model MMLU-Pro Acc. Tokens GPQA-Diamond Acc. Tokens
    Qwen2.5-7B-Instruct (Temp. 0.0) 52.46 401.38 34.85 592.73
    Qwen2.5-7B-Instruct (Temp. 1.0) 51.49 379.84 33.84 537.41
    Qwen2.5-7B-Tag-General-Low 56.00 1674.60 31.82 2808.88
    Qwen2.5-7B-Tag-General-Medium 55.92 2341.27 36.87 3931.13
    Qwen2.5-7B-Tag-General-High 55.81 2632.05 32.83 4238.11
    Qwen2.5-7B-Random-General 55.86 1960.87 34.34 3385.70
    Qwen2.5-7B-TOPS-General 56.50 1788.74 38.38 3083.76
  10. Knowl 10 — Scope and limitations of the evidence

    limitation

    The paper's main analysis and experiments focus on mathematical reasoning, where final-answer correctness can be checked relatively reliably. Its general-reasoning experiments are preliminary, so the effects of long chains of thought across other domains remain insufficiently established. The training approach is primarily supervised fine-tuning; the paper does not test its central claims in a reinforcement-learning setting. The suggestion that reinforcement learning might over-reward longer correct answers containing more erroneous intermediate reasoning, and thereby favor unnecessary self-correction, is presented as a prospective concern rather than an experimentally demonstrated result.

Coverage note — The preliminary comparison of existing System-1 and o1-like models and the prompt-specific length-rank statistics were omitted because they motivate the controlled experiments but add less to reconstructing the paper's core method and findings.

References

  1. 1.Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787, 2024.
  2. 2.Lingjiao Chen, Jared Quincy Davis, Boris Hanin, Peter Bailis, Ion Stoica, Matei Zaharia, and James Zou. Are more LLM calls all you need? towards the scaling properties of compound AI systems. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=m5106RRLgx.
  3. 3.Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, et al. Do not think that much for 2+ 3=? on the overthinking of o1-like llms. arXiv preprint arXiv:2412.21187, 2024.
  4. 4.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  5. 5.Ganqu Cui, Lifan Yuan, Zefan Wang, Hanbin Wang, Wendi Li, Bingxiang He, Yuchen Fan, Tianyu Yu, Qixin Xu, Weize Chen, Jiarui Yuan, Huayu Chen, Kaiyan Zhang, Xingtai Lv, Shuo Wang, Yuan Yao, Hao Peng, Yu Cheng, Zhiyuan Liu, Maosong Sun, Bowen Zhou, and Ning Ding. Process reinforcement through implicit rewards, 2025. Notion Blog.
  6. 6.DeepSeek. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 1 2025. URL https://github.com/deepseek-ai/DeepSeek-R1/blob/main/DeepSeek_R1.pdf.
  7. 7.Kanishk Gandhi, Denise Lee, Gabriel Grand, Muxin Liu, Winson Cheng, Archit Sharma, and Noah D Goodman. Stream of search (sos): Learning to search in language. arXiv preprint arXiv:2404.03683, 2024.
  8. 8.Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, et al. Omni-math: A universal olympiad level mathematic benchmark for large language models. arXiv preprint arXiv:2410.07985, 2024.
  9. 9.Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, et al. Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai. arXiv preprint arXiv:2411.04872, 2024.
  10. 10.Google. Gemini 2.0 flash experimental, 2024. URL https://deepmind.google/technologies/gemini/flash/.
  11. 11.Google. Gemini 2.0 flash thinking mode, 2024. URL https://ai.google.dev/gemini-api/docs/thinking-mode.
  12. 12.Xinyu Guan, Li Lyna Zhang, Yifei Liu, Ning Shang, Youran Sun, Yi Zhu, Fan Yang, and Mao Yang. rstar-math: Small llms can master math reasoning with self-evolved deep thinking. arXiv preprint arXiv:2501.04519, 2025.
  13. 13.Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. Measuring mathematical problem solving with the MATH dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. URL https://openreview.net/forum?id=7Bywt2mQsCe.
  14. 14.Zhen Huang, Haoyang Zou, Xuefeng Li, Yixiu Liu, Yuxiang Zheng, Ethan Chern, Shijie Xia, Yiwei Qin, Weizhe Yuan, and Pengfei Liu. O1 replication journey–part 2: Surpassing o1-preview through simple distillation, big progress or bitter lesson? arXiv preprint arXiv:2411.16489, 2024.
  15. 15.Kimi Team. Kimi k1.5: Scaling reinforcement learning with llms, 1 2025. URL https://github.com/MoonshotAI/Kimi-k1.5/blob/main/Kimi_k1.5.pdf.
  16. 16.Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia. Step-dpo: Step-wise preference optimization for long-chain reasoning of llms. arXiv preprint arXiv:2406.18629, 2024.
  17. 17.Jia Li, Edward Beeching, Lewis Tunstall, Ben Lipkin, Roman Soletskyi, Shengyi Huang, Kashif Rasul, Longhui Yu, Albert Q Jiang, Ziju Shen, et al. Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions. Hugging Face repository, 13, 2024.
  18. 18.Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. Making language models better reasoners with step-aware verifier. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5315–5333, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.291. URL https://aclanthology.org/2023.acl-long.291/.
  19. 19.Zongzhao Li, Zongyang Ma, Mingze Li, Songyou Li, Yu Rong, Tingyang Xu, Ziqi Zhang, Deli Zhao, and Wenbing Huang. Star-r1: Spatial transformation reasoning by reinforcing multimodal llms. arXiv preprint arXiv:2505.15804, 2025.
  20. 20.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  21. 21.Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024.
  22. 22.Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao. O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning. arXiv preprint arXiv:2501.12570, 2025.
  23. 23.Xueguang Ma, Qian Liu, Dongfu Jiang, Ge Zhang, Zejun Ma, and Wenhu Chen. General-reasoner: Advancing llm reasoning across all domains. arXiv preprint arXiv:2505.14652, 2025.
  24. 24.MetaAI. Introducing llama 3.1: Our most capable models to date. https://ai.meta.com/blog/meta-llama-3-1/, 2024.
  25. 25.Yingqian Min, Zhipeng Chen, Jinhao Jiang, Jie Chen, Jia Deng, Yiwen Hu, Yiru Tang, Jiapeng Wang, Xiaoxue Cheng, Huatong Song, et al. Imitate, explore, and self-improve: A reproduction report on slow-thinking reasoning systems. arXiv preprint arXiv:2412.09413, 2024.
  26. 26.NovaSky Team. Sky-t1: Train your own o1 preview model within $450. https://novasky-ai.github.io/posts/sky-t1, 2025. Accessed: 2025-01-09.
  27. 27.OpenAI. Hello gpt-4o, 2024. URL https://openai.com/index/hello-gpt-4o.
  28. 28.OpenAI. Learning to reason with llms, 2024. URL https://openai.com/index/learning-to-reason-with-llms.
  29. 29.Qwen Team. Qwen2. 5: A party of foundation models. Qwen (Sept. 2024). url: https://qwenlm.github. io/blog/qwen2, 5, 2024.
  30. 30.Qwen Team. Qwq: Reflect deeply on the boundaries of the unknown, November 2024. URL https://qwenlm.github.io/blog/qwq-32b-preview/.
  31. 31.Qwen Team. Qwq-32b: Embracing the power of reinforcement learning, November 2025. URL https://qwenlm.github.io/blog/qwq-32b/.
  32. 32.Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 2024.
  33. 33.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022, 2023.
  34. 34.Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al. Code llama: Open foundation models for code. arXiv preprint arXiv:2308.12950, 2023.
  35. 35.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  36. 36.Skywork-o1. Skywork-o1 open series. https://huggingface.co/Skywork, November 2024. URL https://huggingface.co/Skywork.
  37. 37.Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314, 2024.
  38. 38.Lin Sun, Guangxiang Zhao, Xiaoqi Jian, Yuhan Wu, Weihong Lin, Yongfu Zhu, Linglin Zhang, Jinzhu Wu, Junfeng Ran, Sai-er Hu, et al. Tinyr1-32b-preview: Boosting accuracy with branch-merge distillation. arXiv preprint arXiv:2503.04872, 2025.
  39. 39.Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. Math-shepherd: Verify and reinforce llms step-by-step without human annotations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9426–9439, 2024.
  40. 40.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=1PL1NIMMrw.
  41. 41.Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024.
  42. 42.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  43. 43.Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang. An empirical analysis of compute-optimal inference for problem-solving with language models. 2024.
  44. 44.Violet Xiang, Charlie Snell, Kanishk Gandhi, Alon Albalak, Anikait Singh, Chase Blagden, Duy Phung, Rafael Rafailov, Nathan Lile, Dakota Mahan, et al. Towards system 2 reasoning in llms: Learning how to think with meta chain-of-though. arXiv preprint arXiv:2501.04682, 2025.
  45. 45.An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, et al. Qwen2. 5-math technical report: Toward mathematical expert model via self-improvement. arXiv preprint arXiv:2409.12122, 2024.
  46. 46.Wenkai Yang, Jingwen Chen, Yankai Lin, and Ji-Rong Wen. Deepcritic: Deliberate critique with large language models. arXiv preprint arXiv:2505.00662, 2025.
  47. 47.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R Narasimhan. Tree of thoughts: Deliberate problem solving with large language models. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum?id=5Xc1ecxO1h.
  48. 48.Longhui Yu, Weisen Jiang, Han Shi, Jincheng YU, Zhengying Liu, Yu Zhang, James Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. Metamath: Bootstrap your own mathematical questions for large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=N8N0hgNDRt.
  49. 49.Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. Star: Bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems, 35:15476–15488, 2022.
  50. 50.Di Zhang, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, and Wanli Ouyang. Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b. arXiv preprint arXiv:2406.07394, 2024.
  51. 51.Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought prompting in large language models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=5NTt8GFjUHkr.
  52. 52.Yu Zhao, Huifeng Yin, Bo Zeng, Hao Wang, Tianqi Shi, Chenyang Lyu, Longyue Wang, Weihua Luo, and Kaifu Zhang. Marco-o1: Towards open reasoning models for open-ended solutions. arXiv preprint arXiv:2411.14405, 2024.

Citation

MLA
Yang, W., et al. “Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning”. Advances in Neural Information Processing Systems 38, 2025, pp. 48704–30, https://doi.org/10.52202/085713-1452.
APA
Yang, W., Ma, S., Lin, Y., & Wei, F. (2025). Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning. Advances in Neural Information Processing Systems 38, 48704–48730. https://doi.org/10.52202/085713-1452
Chicago
Yang, W., S. Ma, Y. Lin, and F. Wei. 2025. “Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning”. Advances in Neural Information Processing Systems 38, 48704–30. https://doi.org/10.52202/085713-1452.
Harvard
Yang, W. et al. (2025) “Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning”, Advances in Neural Information Processing Systems 38. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp. 48704–48730. Available at: https://doi.org/10.52202/085713-1452.
Vancouver
1. Yang W, Ma S, Lin Y, Wei F (2025) Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning. In: Advances in Neural Information Processing Systems 38. Neural Information Processing Systems Foundation, Inc. (NeurIPS), pp 48704–48730

BibTeX

@inproceedings{Yang_2025, series={NeurIPS 2025}, title={Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning}, url={http://dx.doi.org/10.52202/085713-1452}, DOI={10.52202/085713-1452}, booktitle={Advances in Neural Information Processing Systems 38}, publisher={Neural Information Processing Systems Foundation, Inc. (NeurIPS)}, author={Yang, Wenkai and Ma, Shuming and Lin, Yankai and Wei, Furu}, year={2025}, pages={48704–48730}, collection={NeurIPS 2025} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/