Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
Wenkai YangShuming MaYankai LinFuru Wei
Reveals that excessively scaling chain-of-thought length can impair mathematical reasoning and introduces a test-time compute strategy that learns domain-optimal thinking lengths to match the performance of leading reasoning models on competitive benchmarks.
Recent advances in artificial intelligence have focused on test-time scaling, a method that encourages large language models to spend more computational time "thinking" by generating extended chains of thought before answering. While this approach has improved performance on complex reasoning tasks, it introduces severe inefficiencies and potential performance risks. The article addresses whether excessively extending these reasoning chains causes adverse effects beyond mere computational waste, investigating how reasoning length directly influences model accuracy across tasks of varying difficulty.
The main objective of the article is to demonstrate that over-scaling reasoning lengths can actively impair accuracy, especially on simpler problems, and to propose a practical strategy that dynamically scales reasoning effort based on problem difficulty.
To evaluate this, the authors analyzed leading reasoning models across standardized mathematics and general knowledge benchmarks. They then created a controlled experimental setup using open-source models (including 8-billion and 32-billion parameter models). The researchers fine-tuned an initial base model on a small seed dataset (around 1,300 problems) with three distinct reasoning lengths—low, medium, and high. This produced an intermediate model capable of adjusting its reasoning effort on command. They then used this model to generate multiple candidate solutions across tens of thousands of problems and selected the shortest correct answer for each to train the final self-improved model.
The findings reveal that longer chains of thought do not universally improve accuracy and can actively degrade performance on easier tasks. Detailed analysis shows that longer reasoning paths accumulate more intermediate errors; although learning error correction is beneficial, training on excessive erroneous steps harms model reasoning. Furthermore, there is an optimal reasoning length for each task difficulty: lower effort succeeds best on straightforward problems, while higher effort is necessary only for complex tasks. Applying this insight via the proposed strategy—termed Thinking-Optimal Scaling (TOPS)—enabled a 32-billion parameter model to achieve 95.82% accuracy on basic math (GSM8K) and 46.00% on advanced competition math (AIME 2024) after preference optimization, matching or exceeding much larger or heavily distilled models while consuming far fewer tokens.
These results have direct operational and cost implications. For organizations deploying reasoning models, unchecked test-time compute inflates operational costs and latency without guaranteeing higher accuracy. Implementing adaptive reasoning depth significantly reduces inference compute expenses and mitigates the risk of hallucination or circular reasoning caused by overthinking simple prompts.
Decision-makers and engineering teams should avoid uniform prompts that force maximum reasoning depth across all queries. Instead, deployments should adopt adaptive pipelines that select the most concise valid reasoning path. When curating training data from long-thinking models, teams should apply loss masking or pruning to suppress erroneous reasoning steps rather than training unconditionally on full reasoning traces.
The study's primary limitations include a heavy focus on mathematical reasoning benchmarks, with preliminary rather than comprehensive validation across broader natural language domains. Additionally, the evaluations were conducted predominantly within supervised fine-tuning environments rather than pure reinforcement learning frameworks. Nevertheless, the empirical findings are supported by consistent results across multiple base models and benchmarks, providing strong confidence in the core conclusion that reasoning compute must be tailored to task difficulty.
- Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). This compute-optimal framework establishes how inference-time compute can be allocated by task difficulty, the foundation for understanding why the source seeks domain-specific optimal reasoning lengths.
- Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). Its budget-forcing method and positive results from extending reasoning provide the test-time scaling baseline that the source qualifies by finding that longer thinking can sometimes hurt.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Its teacher-generated reasoning distillation pipeline introduces the core training approach that the source adapts to teach models different reasoning efforts.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). STaR’s iterative self-training on correct model-generated rationales prepares readers for the source’s use of model self-improvement after reasoning-length training.
- Paper: Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?, Zhiyuan Zeng 0004 et al. (2025). This later evaluation tests whether longer reasoning actually improves accuracy, extending the source’s warning about harmful test-time scaling into broader model and benchmark comparisons.
- Paper: Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models, Xingyu Chen et al. (2025). Its analysis of overthinking and training for concise solutions carries the source’s effort-allocation insight into a method for reducing redundant reasoning.
- Paper: CoT-Valve: Length-Compressible Chain-of-Thought Tuning, Xinyin Ma et al. (2025). CoT-Valve extends adaptive reasoning-length control by making chain-of-thought output elastically compressible while aiming to preserve accuracy.
- Paper: Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching, Simon A. Aytes et al. (2025). Sketch-of-Thought applies the broader efficiency goal to compact, task-adaptive reasoning formats that reduce token use without sacrificing performance.
