Fractured Chain-of-Thought Reasoning

Baohao LiaoHanze DongYuhui XuDoyen SahooChristof MonzJunnan LiCaiming Xiong

article2026arXiv7 citations

Introduces Fractured Sampling, an inference-time scaling method that strategically truncates reasoning traces to match full Chain-of-Thought accuracy while drastically cutting token costs across large language models.

Listen

Large language models have recently achieved substantial breakthroughs in complex problem-solving by utilizing extended internal reasoning chains before generating answers. However, generating long reasoning trajectories consumes massive token budgets, leading to severe latency, steep computational costs, and serving bottlenecks that hinder deployment in time-sensitive production environments.

The article aims to evaluate whether full reasoning chains are genuinely necessary for high accuracy and demonstrates how sampling across intermediate reasoning stages can optimize the trade-off between computational cost and model performance.

To investigate this, the authors evaluated multiple reasoning models—including DeepSeek-R1 variants, Qwen3, Skywork-OR1, DeepScaler, and GPT-OSS—across five standard mathematics and science reasoning benchmarks. The approach introduced Fractured Sampling, a unified inference framework that systematically explores three dimensions: the number of independent reasoning paths, the number of final candidate solutions per path, and the reasoning depth at which intermediate chains are truncated. The evaluation tracked pass rates relative to overall token budgets and examined practical selection methods such as majority voting, process reward models, and early-stopping rules.

The evaluation yielded several key findings. First, truncating reasoning chains before full completion matched or exceeded the accuracy of complete chains while consuming significantly fewer tokens. Second, sampling across intermediate reasoning depths yielded the steepest improvements per token compared to simply generating more independent paths or multiple final answers. Third, pairing intermediate sampling with practical selection strategies—such as retaining only later reasoning steps or applying linear depth-weighted aggregation—improved average accuracy on a 7-billion parameter model from 60.4% to 70.8%, surpassing a standard baseline model with 14 billion parameters (68.3%). Finally, an automated early-stopping mechanism reduced token usage by approximately 20% across evaluated models while preserving overall accuracy.

These results demonstrate that long reasoning processes suffer from redundancy and that model errors across different reasoning depths are largely uncorrelated. Organizations can exploit this structure to capture diverse solutions earlier without paying the full computational cost of lengthy chains. This directly lowers operational serving costs, decreases user-facing latency, and provides a way for smaller, resource-efficient models to outperform standard larger models.

Decision-makers should consider adopting fractured sampling strategies and training-free early-stopping mechanisms to reduce inference costs. Teams can deploy linear depth-weighted aggregation or retain the final segment of intermediate solutions to avoid noise from early reasoning stages. Furthermore, engineering teams should explore integrating these efficient intermediate sampling techniques into reinforcement learning training pipelines, where large sample volumes are required.

Confidence in these findings is high across the tested mathematical and scientific reasoning benchmarks. However, leaders should note that the primary tests relied on specific open reasoning model families and targeted reasoning-heavy benchmarks. Practical gains in production will depend on task structure and should be validated through pilot deployments in target operational workflows.

Cover for Fractured Chain-of-Thought Reasoning

Abstract

Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining. Similarly, Chain-of-Thought (CoT) prompting and its extension, Long CoT, improve accuracy by generating rich intermediate reasoning trajectories, but these approaches incur substantial token costs that impede their deployment in latency-sensitive settings. In this work, we first show that truncated CoT, which stops reasoning before completion and directly generates the final answer, often matches the full CoT sampling while using dramatically fewer tokens. Building on this insight, we introduce Fractured Sampling, a unified inference-time strategy that interpolates between full CoT and solution-only sampling along three orthogonal axes: (1) the number of reasoning trajectories, (2) the number of final solutions per trajectory, and (3) the depth at which reasoning traces are truncated. Through extensive experiments on five diverse reasoning benchmarks and several model scales, we demonstrate that Fractured Sampling consistently achieves superior accuracy-cost trade-offs, yielding steep log-linear scaling gains in Pass@k versus token budget. Our analysis reveals how to allocate computation across these dimensions to maximize performance, paving the way for more efficient and scalable LLM reasoning. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Preliminary
  • 3 Fractured sampling for long-CoT reasoning
  • 3.1 Analysis of fractured sampling
  • 3.2 Scaling laws along the trajectory dimension
  • 4 Empirical results
  • 4.1 Scaling law for each dimension
  • 4.2 Scaling law across dimensions
  • 4.3 Accuracy across dimensions
  • 4.4 Early stopping for efficient generation
  • 4.5 More results
  • 5 Related work
  • 6 Conclusion
  • References
  • A Proof of the diversity lower bound
  • B More results
  • B.1 Linear depth-weighted aggregation
  • B.2 Latency

Knowls

  1. Knowl 1 — Fractured Sampling Framework for Long-Chain-of-Thought Reasoning

    model/method

    Fractured Sampling is an inference-time decoding strategy for reasoning Large Language Models (LLMs) that operates along three orthogonal axes: trajectory diversity (nn), reasoning depth diversity (HH), and solution diversity (mm). Instead of generating only complete reasoning chains or direct solutions, the framework truncates intermediate reasoning traces at multiple depth segments and generates candidate final solutions directly from these partial traces.

    Let xx denote an input prompt and ε\varepsilon denote stochastic random seeds. A reasoning model generates an intermediate reasoning trace h=[h1ε,…,hHε]=fh(x,ε)h = [h_1^\varepsilon, \dots, h_H^\varepsilon] = f_h(x, \varepsilon) of HH steps, and conditioned on a partial reasoning trace up to step tt, denoted h1:tε=fht(x,ε)h_{1:t}^\varepsilon = f_h^t(x, \varepsilon), the final solution is generated as z=fo(x,fht(x,ε),ε′)z = f_o(x, f_h^t(x, \varepsilon), \varepsilon'), with answer parser g(z)g(z). The full candidate answer set under Fractured Sampling is defined as:

    Fn,m,H(x,ε1:n,ε1:m,1:H)={g∘fo(x,fht(x,εi),εj,t)  |  i=1,…,n; j=1,…,m; t=1,…,H}F_{n,m,H}(x, \varepsilon_{1:n}, \varepsilon_{1:m, 1:H}) = \left\{ g \circ f_o\left(x, f_h^t(x, \varepsilon_i), \varepsilon_{j,t}\right) \;\middle|\; i = 1, \dots, n;\, j = 1, \dots, m;\, t = 1, \dots, H \right\}

    where:

    • nn controls the number of independent full reasoning chains sampled (trajectory diversity);
    • HH controls the number of intermediate depth intervals along each chain at which generation is branched (reasoning depth diversity);
    • mm controls the number of independent final answers generated per truncated reasoning prefix (solution diversity).
  2. Knowl 2 — Three-Dimensional Fractured Sampling Procedure

    algorithm

    The 3D Fractured Sampling algorithm decomposes reasoning chains into intermediate depth segments and samples multiple candidate solutions per segment before final aggregation.

    Input: prompt xx, number of trajectories nn, depth segments HH, solutions per depth mm, answer selector S\mathcal{S}
    Output: final selected answer y^\hat{y}
    C←∅\mathcal{C} \leftarrow \emptyset
    for i=1i = 1 to nn do
        Ti←MODEL.GenerateFullCoT(x)T_i \leftarrow \text{MODEL.GenerateFullCoT}(x)
        Tokenize TiT_i into tokens {t1,…,tL}\{t_1, \dots, t_L\}
        s←max⁡(1,⌊L/H⌋)s \leftarrow \max(1, \lfloor L / H \rfloor)
        for t=1t = 1 to HH do
            pi,t←detokenize({t1,…,tt⋅s})p_{i,t} \leftarrow \text{detokenize}(\{t_1, \dots, t_{t \cdot s}\})
            for j=1j = 1 to mm do
                y~i,t,j←MODEL.GenerateSolution(pi,t,x)\tilde{y}_{i,t,j} \leftarrow \text{MODEL.GenerateSolution}(p_{i,t}, x)
                ai,t,j←ExtractAnswer(y~i,t,j)a_{i,t,j} \leftarrow \text{ExtractAnswer}(\tilde{y}_{i,t,j})
                C←C∪{ai,t,j}\mathcal{C} \leftarrow \mathcal{C} \cup \{a_{i,t,j}\}
            end for
        end for
    end for
    y^←S(C)\hat{y} \leftarrow \mathcal{S}(\mathcal{C})
    return y^\hat{y}

    The selector S\mathcal{S} can be majority voting or scoring via a process/outcome reward model (Best-of-NN). When H=1H=1 and m=1m=1, the procedure reduces to standard multi-trajectory sampling (Vanilla CoT).

  3. Knowl 3 — Diversity Lower Bound via Reasoning Depth Decorrelation

    theoretical result

    Let Fk∈{0,1}F_k \in \{0, 1\} be the binary indicator of failure for branch sample k∈{1,…,K}k \in \{1, \dots, K\}, where K=mHK = mH branch samples are drawn across HH depth levels with mm solutions per depth, and let qk=P(Fk=1)q_k = \mathbb{P}(F_k = 1). The overall fractured-sampling success probability psegp_{\text{seg}} satisfies:

    pseg=1−P(⋀k=1KFk=1)=1−E[∏k=1KFk]p_{\text{seg}} = 1 - \mathbb{P}\left(\bigwedge_{k=1}^K F_k = 1\right) = 1 - \mathbb{E}\left[ \prod_{k=1}^K F_k \right]

    Applying the inclusion-exclusion joint cumulant expansion yields:

    E[∏k=1KFk]=∏k=1Kqk+∑i<jCov(Fi,Fj)+∑i<j<kκijk+…\mathbb{E}\left[ \prod_{k=1}^K F_k \right] = \prod_{k=1}^K q_k + \sum_{i < j} \text{Cov}(F_i, F_j) + \sum_{i < j < k} \kappa_{ijk} + \dots

    where Cov(Fi,Fj)=E[FiFj]−qiqj\text{Cov}(F_i, F_j) = \mathbb{E}[F_i F_j] - q_i q_j is the pairwise covariance between branch failures.

    • Independent Regime: If all failure events FkF_k are mutually independent, the success probability matches the product-of-marginals baseline: pseg=1−∏t=1H(1−pt)mp_{\text{seg}} = 1 - \prod_{t=1}^H (1 - p_t)^m, where pt=1−qtp_t = 1 - q_t is the marginal success rate at depth tt.
    • Intermediate Regime with Error Decorrelation: When failure events at different reasoning depths exhibit negative covariance (Cov(Fi,Fj)<0\text{Cov}(F_i, F_j) < 0), distinct reasoning depths provide non-coinciding, diverse error modes. This reduces the joint all-fail probability below ∏k=1Kqk\prod_{k=1}^K q_k, boosting psegp_{\text{seg}} strictly above the independent baseline.
  4. Knowl 4 — Log-Linear Test-Time Scaling Laws and Dominance of Depth Branching

    empirical result

    The computational token cost B(n,m,H)B(n, m, H) consumed by Fractured Sampling is modeled as:

    B(n,m,H)=n⋅Cthinking+n⋅m⋅H⋅Csolution=n(Cthinking+mHCsolution)B(n, m, H) = n \cdot C_{\text{thinking}} + n \cdot m \cdot H \cdot C_{\text{solution}} = n \left( C_{\text{thinking}} + m H C_{\text{solution}} \right)

    where CthinkingC_{\text{thinking}} is the average token length of a full reasoning prefix and CsolutionC_{\text{solution}} is the token cost of generating a final candidate solution.

    Under single-axis scaling where budget is swept along one parameter while holding the others fixed (Bn=B(n,1,1)B_n = B(n, 1, 1), Bm=B(1,m,1)B_m = B(1, m, 1), BH=B(1,1,H)B_H = B(1, 1, H)), the success rate pass@k\text{pass}@k follows a log-linear scaling behavior:

    pass@k(B∗)≈C∗log⁡B∗+c∗,∗∈{n,m,H}\text{pass}@k(B_*) \approx C_* \log B_* + c_*, \quad * \in \{n, m, H\}

    Empirical evaluation across mathematical and scientific reasoning benchmarks shows that:

    CH≥max⁡{Cn,Cm}C_H \ge \max\{C_n, C_m\}

    Allocating computational tokens to reasoning depth diversity (HH) yields a steeper scaling slope per token than allocating tokens to trajectory diversity (nn) or solution diversity (mm). The scaling behavior with respect to wall-clock time exhibits the same log-linear profile.

  5. Knowl 5 — Linear Depth-Weighted Aggregation for Fractured Sampling

    model/method

    Because solutions generated from early, shallow reasoning segments contain higher noise than solutions generated from deeper segments, linear depth-weighted aggregation scales candidate scores proportionally to the reasoning depth t∈{1,…,H}t \in \{1, \dots, H\}.

    For a candidate answer aa extracted at depth step tt, with base selector score s(a,t)s(a, t) (such as a Process Reward Model score, or s(a,t)=1s(a, t) = 1 for majority voting), the normalized depth weight wtw_t is defined as:

    wt=t∑k=1Hk=2tH(H+1)w_t = \frac{t}{\sum_{k=1}^H k} = \frac{2t}{H(H + 1)}

    The total aggregate score Sagg(a)S_{\text{agg}}(a) for answer aa across all occurrences and depths is:

    Sagg(a)=∑all occurrences t of awt⋅s(a,t)S_{\text{agg}}(a) = \sum_{\text{all occurrences } t \text{ of } a} w_t \cdot s(a, t)

    The final prediction a^\hat{a} is selected via:

    a^=arg⁡max⁡aSagg(a)\hat{a} = \arg\max_a S_{\text{agg}}(a)

  6. Knowl 6 — Consistency-Based Early Stopping for Long-CoT Generation

    model/method

    Consistency-based early stopping terminates reasoning generation before reaching the maximum token limit by checking the agreement of intermediate solutions across depth intervals.

    1. Initialization: The model generates an initial prefix of 6,144 reasoning tokens. A candidate solution is generated from this prefix, and its final answer is extracted.
    2. Periodic Check: The model continues generating reasoning tokens in increments of 2,048 tokens. After each 2,048-token chunk, a candidate solution is branched and an intermediate prediction is parsed.
    3. Termination Condition: Generation halts immediately if the exact same answer prediction appears more than once across evaluated intervals. If no repeat occurs before reaching max_tokens, the final prediction from the complete trace is selected.
  7. Knowl 7 — Token Savings and Accuracy of Consistency-Based Early Stopping

    data/table

    Evaluating consistency-based early stopping against vanilla full generation across multiple reasoning benchmarks (MATH500, AIME25, AIMO2, GPQA) demonstrates an average token reduction of approximately 20% while preserving or improving task accuracy.

    Model Method Accuracy (%) ↑\uparrow Avg Tokens / Question (K) ↓\downarrow
    MATH500 AIME25 AIMO2 GPQA Avg.
    DS-R1-1.5B Vanilla 70.8 27.5 15.0 34.1 36.9 14.3
    Early Stop +1.2 -0.0 +10.6 -0.3 +2.9 -2.9
    DSR-1.5B Vanilla 76.5 41.7 20.0 19.2 39.4 8.0
    Early Stop -0.9 -0.0 -0.0 +0.4 -0.1 -1.5
    SW-OR1-7B Vanilla 89.0 45.0 47.5 48.6 57.5 10.9
    Early Stop -0.4 -0.0 -0.0 +1.1 +0.2 -2.2

    Relative changes are shown for Early Stop compared to Vanilla sampling. DS-R1-1.5B refers to DeepSeek-R1-Distill-Qwen-1.5B, DSR-1.5B refers to DeepScaleR-1.5B-Preview, and SW-OR1-7B refers to Skywork-OR1-7B.

  8. Knowl 8 — Best-of-N and Majority Voting Accuracy under Depth-Weighted and Denoised Aggregation

    data/table

    Comparison of Best-of-NN (BoN using Qwen2.5-Math-PRM-72B scoring) and Majority Voting (Maj) on DS-R1-Qwen-7B with n=16n=16 trajectories across MATH500 Level 5, AIME24, AIME25, AIMO2, and GPQA. Naive uniform sampling across all H=16H=16 depths introduces early-stage noise, whereas retaining only the last 4 depths (H=−4H=-4) or applying linear depth-weighted aggregation outperforms both the H=1,m=1H=1, m=1 baseline and a 2×2\times larger model (DS-R1-Qwen-14B with H=1,m=1H=1, m=1).

    Metric Method HH mm MATH500 L5 AIME24 AIME25 AIMO2 GPQA Avg.
    DS-R1-Qwen-7B
    BoN Original 1 1 90.3 63.3 53.3 40.0 55.1 60.4
    Original 16 1 90.3 70.0 53.3 40.0 53.5 61.4
    Original -4 1 93.3 73.3 60.0 60.0 53.5 68.0
    Linear depth-weighted 16 1 96.3 76.7 60.0 60.0 56.1 69.8
    Original -4 4 94.8 73.3 60.0 70.0 56.1 70.8
    DS-R1-Qwen-14B Baseline
    BoN Original 1 1 91.8 80.0 60.0 50.0 59.6 68.3
    DS-R1-Qwen-7B
    Maj Original 1 1 95.5 76.7 60.0 50.0 51.5 66.7
    Original 16 1 94.0 73.3 60.0 50.0 49.0 65.3
    Original -4 1 96.3 76.7 63.3 60.0 53.0 69.9
    Linear depth-weighted 16 1 95.9 76.7 66.7 60.0 52.2 70.3
    Original -4 4 96.3 80.0 66.7 60.0 53.5 71.3
    DS-R1-Qwen-14B Baseline
    Maj Original 1 1 94.0 83.3 53.3 50.0 59.6 68.0

    H=−4H=-4 denotes discarding reasoning depths t=1,…,11t=1,\dots,11 and retaining only the final four depths t=12,…,16t=12,\dots,16.

  9. Knowl 9 — Experimental Benchmarking Setup for Fractured Sampling

    experimental setup

    Fractured Sampling is evaluated across five challenging reasoning benchmarks:

    • MATH500 Level 5: Challenging subset of the MATH dataset;
    • AIME24 and AIME25: American Invitational Mathematics Examination (2024 and 2025 editions);
    • AIMO2: Reference problems from the AI Mathematical Olympiad Progress Prize 2;
    • GPQA Diamond: High-difficulty Google-proof graduate-level science question-answering benchmark.

    Models Tested: DeepSeek-R1 (DS-R1), DeepSeek-R1-Distill-Qwen (1.5B, 7B, 14B), DeepScaleR-1.5B-Preview (DSR-1.5B), Skywork-OR1-7B (SW-OR1-7B), Qwen3-1.7B, and GPT-OSS-20B.

    Generation and Hardware Configuration: Inference runs on NVIDIA A100-80GB GPUs using the vLLM framework with decoding parameters temperature =0.6= 0.6, top-p=0.95p = 0.95, and maximum generation length =32,768= 32,768 tokens. The default multidimensional sampling configuration uses n=16n = 16 trajectories, H=16H = 16 equal token-count depth segments, and m=4m = 4 candidate solutions per segment.

Coverage note — None was omitted; all key theoretical propositions, algorithms, empirical scaling laws, and performance tables were included.

References

  1. 1.Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, et al. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925, 2025.
  2. 2.Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. Gqa: Training generalized multi-query transformer models from multi-head checkpoints. arXiv preprint arXiv:2305.13245, 2023.
  3. 3.AI Anthropic. Claude 3.5 sonnet model card addendum. Claude-3.5 Model Card, 3, 2024.
  4. 4.Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling. arXiv preprint arXiv:2407.21787, 2024.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  6. 6.Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318, 2023a.
  7. 7.Ziyi Chen, Xiaocong Yang, Jiacheng Lin, Chenkai Sun, Kevin Chen-Chuan Chang, and Jie Huang. Cascade speculative drafting for even faster llm inference. arXiv preprint arXiv:2312.11462, 2023b.
  8. 8.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  9. 9.Hanze Dong, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, KaShun SHUM, and Tong Zhang. RAFT: Reward ranked finetuning for generative foundation model alignment. Transactions on Machine Learning Research, 2023. ISSN 2835-8856. URL https://openreview.net/forum?id=m7p5O7zblY.
  10. 10.Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, and Tong Zhang. Rlhf workflow: From reward modeling to online rlhf. arXiv preprint arXiv:2405.07863, 2024.
  11. 11.Gongfan Fang, Xinyin Ma, and Xinchao Wang. Thinkless: Llm learns when to think. Advances in neural information processing systems, 38:151268–151295, 2026.
  12. 12.Simon Frieder, Sam Bealing, Arsenii Nikolaiev, Geoff C. Smith, Kevin Buzzard, Timothy Gowers, Peter J. Liu, Po-Shen Loh, Lester Mackey, Leonardo de Moura, Dan Roberts, D. Sculley, Terence Tao, David Balduzzi, Simon Coyle, Alex Gerko, Ryan Holbrook, Addison Howard, and XTX Markets. Ai mathematical olympiad - progress prize 2, 2024. Kaggle.
  13. 13.Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. Model tells you what to discard: Adaptive kv cache compression for llms. arXiv preprint arXiv:2310.01801, 2023.
  14. 14.Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.
  15. 15.Jujie He, Jiacai Liu, Chris Yuhao Liu, Rui Yan, Chaojie Wang, Peng Cheng, Xiaoyu Zhang, Fuxiang Zhang, Jiacheng Xu, Wei Shen, Siyuan Li, Liang Zeng, Tianwen Wei, Cheng Cheng, Yang Liu, and Yahui Zhou. Skywork open reasoner series, 2025. Notion Blog.
  16. 16.Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409, 2017.
  17. 17.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  18. 18.Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024.
  19. 19.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. Openai o1 system card. arXiv preprint arXiv:2412.16720, 2024.
  20. 20.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  21. 21.Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. Large language models are zero-shot reasoners. Advances in neural information processing systems, 35:22199–22213, 2022.
  22. 22.Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, et al. Training language models to self-correct via reinforcement learning. arXiv preprint arXiv:2409.12917, 2024.
  23. 23.Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023.
  24. 24.Yaniv Leviathan, Matan Kalman, and Yossi Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pp. 19274–19286. PMLR, 2023.
  25. 25.Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. Snapkv: Llm knows what you are looking for before generation. arXiv preprint arXiv:2404.14469, 2024a.
  26. 26.Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. Eagle: Speculative sampling requires rethinking feature uncertainty. arXiv preprint arXiv:2401.15077, 2024b.
  27. 27.Baohao Liao, Yuhui Xu, Hanze Dong, Junnan Li, Christof Monz, Silvio Savarese, Doyen Sahoo, and Caiming Xiong. Reward-guided speculative decoding for efficient llm reasoning. arXiv preprint arXiv:2501.19324, 2025.
  28. 28.Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step. arXiv preprint arXiv:2305.20050, 2023.
  29. 29.Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu. Kivi: A tuning-free asymmetric 2bit quantization for kv cache. arXiv preprint arXiv:2402.02750, 2024.
  30. 30.Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Meiqi Guo, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, et al. Improve mathematical reasoning in language models by automated process supervision. arXiv preprint arXiv:2406.06592, 2024.
  31. 31.Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica. Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025. Notion Blog.
  32. 32.Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, and Matei Zaharia. Reasoning models can be effective without thinking. arXiv preprint arXiv:2504.09858, 2025.
  33. 33.MAA Committees. AIME Problems and Solutions. https://artofproblemsolving.com/wiki/index.php/AIME_Problems_and_Solutions, 2025.
  34. 34.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022.
  35. 35.David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level googleproof q&a benchmark. In First Conference on Language Modeling, 2024.
  36. 36.Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Y Wu, et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.
  37. 37.Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. Scaling llm test-time compute optimally can be more effective than scaling model parameters. arXiv preprint arXiv:2408.03314, 2024.
  38. 38.Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit. Blockwise parallel decoding for deep autoregressive models. Advances in Neural Information Processing Systems, 31, 2018.
  39. 39.Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. Learning to summarize with human feedback. Advances in neural information processing systems, 33:3008–3021, 2020.
  40. 40.Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen. Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding. arXiv preprint arXiv:2404.11912, 2024.
  41. 41.Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024.
  42. 42.Qwen Team. Qwen3, April 2025. URL https://qwenlm.github.io/blog/qwen3/.
  43. 43.Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, et al. Efficient large language models: A survey. arXiv preprint arXiv:2312.03863, 1, 2023.
  44. 44.Peiyi Wang, Lei Li, Zhihong Shao, RX Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. Math-shepherd: Verify and reinforce llms step-by-step without human annotations. arXiv preprint arXiv:2312.08935, 2023.
  45. 45.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171, 2022.
  46. 46.Yiming Wang, Pei Zhang, Siyuan Huang, Baosong Yang, Zhuosheng Zhang, Fei Huang, and Rui Wang. Sampling-efficient test-time scaling: Self-estimating the best-of-n sampling in early decoding. Advances in Neural Information Processing Systems, 38:162137–162174, 2026.
  47. 47.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022.
  48. 48.Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, and Zhifang Sui. Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding. arXiv preprint arXiv:2401.07851, 2024.
  49. 49.Heming Xia, Chak Tou Leong, Wenjie Wang, Yongqi Li, and Wenjie Li. Tokenskip: Controllable chain-of-thought compression in llms. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 3351–3363, 2025.
  50. 50.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. arXiv preprint arXiv:2309.17453, 2023.
  51. 51.Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, and Tong Zhang. Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint. arXiv preprint arXiv:2312.11456, 2023.
  52. 52.Wei Xiong, Jiarui Yao, Yuhui Xu, Bo Pang, Lei Wang, Doyen Sahoo, Junnan Li, Nan Jiang, Tong Zhang, Caiming Xiong, et al. A minimalist approach to llm reasoning: from rejection sampling to reinforce. arXiv preprint arXiv:2504.11343, 2025.
  53. 53.Yuhui Xu, Zhanming Jie, Hanze Dong, Lei Wang, Xudong Lu, Aojun Zhou, Amrita Saha, Caiming Xiong, and Doyen Sahoo. Think: Thinner key cache by query-driven pruning. arXiv preprint arXiv:2407.21018, 2024.
  54. 54.Dongjie Yang, XiaoDong Han, Yan Gao, Yao Hu, Shilin Zhang, and Hai Zhao. Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference. arXiv preprint arXiv:2405.12532, 2024.
  55. 55.Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR), 2023.
  56. 56.Hanning Zhang, Pengcheng Wang, Shizhe Diao, Yong Lin, Rui Pan, Hanze Dong, Dylan Zhang, Pavlo Molchanov, and Tong Zhang. Entropy-regularized process reward model. arXiv preprint arXiv:2412.11006, 2024a.
  57. 57.Jintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo, Shuofei Qiao, Lun Du, Da Zheng, Huajun Chen, and Ningyu Zhang. Lightthinker: Thinking step-by-step compression. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 13318–13339, 2025a.
  58. 58.Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra. Draft & verify: Lossless large language model acceleration via self-speculative decoding. arXiv preprint arXiv:2309.08168, 2023.
  59. 59.Yichi Zhang, Bofei Gao, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, Wen Xiao, et al. Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling. arXiv preprint arXiv:2406.02069, 2024b.
  60. 60.Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. The lessons of developing process reward models in mathematical reasoning. arXiv preprint arXiv:2501.07301, 2025b.
  61. 61.Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al. H2o: Heavy-hitter oracle for efficient generative inference of large language models. Advances in Neural Information Processing Systems, 36, 2024c.
  62. 62.Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc Le, et al. Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625, 2022.

Citation

MLA
Liao, B., et al. “Fractured Chain-of-Thought Reasoning”. arXiv, 2025, http://arxiv.org/abs/2505.12992v4.
APA
Liao, B., Dong, H., Xu, Y., Sahoo, D., Monz, C., Li, J., & Xiong, C. (2025). Fractured Chain-of-Thought Reasoning. arXiv. http://arxiv.org/abs/2505.12992v4
Chicago
Liao, B., H. Dong, Y. Xu, et al. 2025. “Fractured Chain-of-Thought Reasoning”. arXiv. http://arxiv.org/abs/2505.12992v4.
Harvard
Liao, B. et al. (2025) “Fractured Chain-of-Thought Reasoning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2505.12992v4.
Vancouver
1. Liao B, Dong H, Xu Y, Sahoo D, Monz C, Li J, Xiong C (2025) Fractured Chain-of-Thought Reasoning. arXiv

BibTeX

@article{liao2025fractured,
  title = {Fractured Chain-of-Thought Reasoning},
  author = {Liao, Baohao and Dong, Hanze and Xu, Yuhui and Sahoo, Doyen and Monz, Christof and Li, Junnan and Xiong, Caiming},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2505.12992v4},
  eprint = {2505.12992}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission