Fractured Chain-of-Thought Reasoning
Baohao LiaoHanze DongYuhui XuDoyen SahooChristof MonzJunnan LiCaiming Xiong
Introduces Fractured Sampling, an inference-time scaling method that strategically truncates reasoning traces to match full Chain-of-Thought accuracy while drastically cutting token costs across large language models.
Large language models have recently achieved substantial breakthroughs in complex problem-solving by utilizing extended internal reasoning chains before generating answers. However, generating long reasoning trajectories consumes massive token budgets, leading to severe latency, steep computational costs, and serving bottlenecks that hinder deployment in time-sensitive production environments.
The article aims to evaluate whether full reasoning chains are genuinely necessary for high accuracy and demonstrates how sampling across intermediate reasoning stages can optimize the trade-off between computational cost and model performance.
To investigate this, the authors evaluated multiple reasoning models—including DeepSeek-R1 variants, Qwen3, Skywork-OR1, DeepScaler, and GPT-OSS—across five standard mathematics and science reasoning benchmarks. The approach introduced Fractured Sampling, a unified inference framework that systematically explores three dimensions: the number of independent reasoning paths, the number of final candidate solutions per path, and the reasoning depth at which intermediate chains are truncated. The evaluation tracked pass rates relative to overall token budgets and examined practical selection methods such as majority voting, process reward models, and early-stopping rules.
The evaluation yielded several key findings. First, truncating reasoning chains before full completion matched or exceeded the accuracy of complete chains while consuming significantly fewer tokens. Second, sampling across intermediate reasoning depths yielded the steepest improvements per token compared to simply generating more independent paths or multiple final answers. Third, pairing intermediate sampling with practical selection strategies—such as retaining only later reasoning steps or applying linear depth-weighted aggregation—improved average accuracy on a 7-billion parameter model from 60.4% to 70.8%, surpassing a standard baseline model with 14 billion parameters (68.3%). Finally, an automated early-stopping mechanism reduced token usage by approximately 20% across evaluated models while preserving overall accuracy.
These results demonstrate that long reasoning processes suffer from redundancy and that model errors across different reasoning depths are largely uncorrelated. Organizations can exploit this structure to capture diverse solutions earlier without paying the full computational cost of lengthy chains. This directly lowers operational serving costs, decreases user-facing latency, and provides a way for smaller, resource-efficient models to outperform standard larger models.
Decision-makers should consider adopting fractured sampling strategies and training-free early-stopping mechanisms to reduce inference costs. Teams can deploy linear depth-weighted aggregation or retain the final segment of intermediate solutions to avoid noise from early reasoning stages. Furthermore, engineering teams should explore integrating these efficient intermediate sampling techniques into reinforcement learning training pipelines, where large sample volumes are required.
Confidence in these findings is high across the tested mathematical and scientific reasoning benchmarks. However, leaders should note that the primary tests relied on specific open reasoning model families and targeted reasoning-heavy benchmarks. Practical gains in production will depend on task structure and should be validated through pilot deployments in target operational workflows.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces the foundational Chain-of-Thought prompting paradigm that Fractured Chain-of-Thought builds upon and aims to optimize at inference time.
- Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). Provides the foundational framework for analyzing test-time compute allocation and trade-offs between sequential and parallel search strategies.
- Paper: Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models, Xingyu Chen et al. (2025). Demonstrates the phenomenon of overthinking and token redundancy in long CoT reasoning, motivating the need for truncated and compute-efficient reasoning.
- Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). Establishes simple test-time scaling and introduces budget forcing techniques to control thinking token generation at inference time.
- Paper: Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?, Zhiyuan Zeng 0004 et al. (2025). Critically examines the limits of sequential test-time scaling versus parallel sampling in reasoning models, directly contextualizing fractured sampling's multi-axis formulation.
- Paper: Let's Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs, Pranjal Aggarwal et al. (2023). Explores inference-time sampling efficiency and early stopping criteria in reasoning models, prefiguring fractured trajectory truncation.
- Paper: CoT-Valve: Length-Compressible Chain-of-Thought Tuning, Xinyin Ma et al. (2025). Studies compressible reasoning chains and demonstrates that intermediate reasoning steps can be shortened without sacrificing accuracy.
- Paper: TokenSkip: Controllable Chain-of-Thought Compression in LLMs, Heming Xia et al. (2025). Analyzes token redundancy in reasoning chains and provides evidence that models can arrive at correct answers even after dropping intermediate tokens.
- Paper: SPIRAL: Learning to Search and Aggregate, Jubayer Ibn Hamid et al. (2026). Extends inference-time compute scaling by training models with reinforcement learning to jointly optimize and aggregate sequential and parallel reasoning traces.
- Paper: Reasoning with Sampling: Your Base Model is Smarter Than You Think, Aayush Karan et al. (2026). Expands on inference-time sampling strategies by demonstrating how iterative resampling from base models achieves high-quality reasoning without retraining.
- Paper: Efficient Reasoning on the Edge, Yelysei Bondarenko et al. (2026). Applies efficient inference and budget-constrained reasoning techniques to resource-limited edge deployments using parallel generation and switching mechanisms.
- Paper: Prefix Sliding for efficient test-time scaling, Niklas Muennighoff et al. (2026). Investigates context eviction and sliding-window test-time scaling to further reduce memory and computational costs during extended reasoning.
- Paper: Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers, Ying Fan et al. (2026). Replaces explicit sequential token generation with parallel latent space refinement to overcome the latency bottlenecks of step-by-step reasoning.
