CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Xinyin MaGuangnian WanRunpeng YuGongfan FangXinchao Wang

article2025ACL159 citations

Proposes CoT-Valve, a parameter-space tuning method that enables a single reasoning model to dynamically control and compress chain-of-thought length, slashing token costs on benchmarks like GSM8K and AIME with minimal accuracy loss.

Listen

Modern artificial intelligence models rely heavily on multi-step reasoning chains to solve complex mathematical and logical tasks. While this reasoning process boosts accuracy, it frequently leads models to generate excessively lengthy intermediate steps even for trivial questions, causing high computational overhead, increased latency, and inflated operational costs. Existing methods to reduce reasoning length, such as prompt-based instructions, struggle to reliably control output size and fail to generate highly compact explanations.

The article introduces and evaluates CoT-Valve, a novel tuning and inference framework designed to elastically control the length of reasoning paths in a single model by manipulating a specific update direction in its parameter space.

The researchers implemented CoT-Valve by isolating length-controlling update directions using lightweight low-rank adaptation modules. They constructed a multi-length reasoning dataset called MixChain, pairing long and short valid explanations for identical questions. Using this data, they evaluated two strategies: a precise continuous tuning method and a progressive compression approach that iteratively trains models on incrementally shorter reasoning paths. The evaluations spanned multiple open model architectures, including standard base models, reasoning-tuned models, and distilled variants, benchmarked across elementary and competitive mathematical datasets.

The analysis produced several critical findings. First, CoT-Valve compressed the average reasoning output of a thirty-two-billion-parameter reasoning model on elementary math from 741 tokens down to 225 tokens while maintaining accuracy above 94.9%, substantially outperforming prompt-based controls. Second, on challenging competition mathematics, the method reduced reasoning length from 6,827 tokens to 4,629 tokens with only one additional incorrect answer. Third, on simpler tasks, shorter reasoning paths frequently outperformed long chains, particularly in smaller models where direct training on verbose reasoning degraded task accuracy. Finally, process reward evaluations showed that intermediate reasoning steps generated through this compression method exhibited higher correctness scores by eliminating noisy and redundant reasoning.

These findings indicate that organizations deploying advanced reasoning models can dynamically tune inference cost and speed without retraining separate architectures for simple and complex user queries. By selectively scaling down token usage on less demanding tasks, teams can achieve significant cost savings and latency reductions while maintaining strong baseline performance on standard non-reasoning benchmarks.

Decision-makers should consider pilot implementations of parameter-based length modulation to optimize serving costs across varying task difficulties. Future development should explore fine-grained token compression that selectively shortens trivial segments of a reasoning path while preserving long-chain deliberation for complex intermediate steps.

The study's primary limitation is that aggressive compression on highly intricate problems can lead to degraded solution accuracy, as complex tasks still depend on extended deliberation. Additionally, the approach relies on the base model's initial generation capability, meaning that poor underlying reasoning cannot be effectively compressed. Confidence in the reported efficiency gains remains high across standard reasoning benchmarks, though careful task-level evaluation is recommended before applying extreme compression.

arXiv: 2502.09601
Cover for CoT-Valve: Length-Compressible Chain-of-Thought Tuning

Abstract

Chain-of-Thought significantly enhances a model’s reasoning capability, but it also comes with a considerable increase in inference costs due to long chains. With the observation that the reasoning path can be easily compressed under easy tasks but struggles on hard tasks, we explore the feasibility of elastically controlling the length of reasoning paths with only one model, thereby reducing the inference overhead of reasoning models dynamically based on task difficulty. We introduce a new tuning and inference strategy named CoT-Valve, designed to allow models to generate reasoning chains of varying lengths. To achieve this, we propose to identify a direction in the parameter space that, when manipulated, can effectively control the length of generated CoT. Moreover, we show that this property is valuable for compressing the reasoning chain. We construct datasets with chains from long to short for the same questions and explore two enhanced strategies for CoT-Valve: (1) a precise length-compressible CoT tuning method, and (2) a progressive chain length compression approach. Our experiments show that CoT-Valve successfully enables controllability and compressibility of the chain and shows better performance than the prompt-based control. We applied this method to QwQ-32B-Preview, reducing reasoning chains on GSM8K from 741 to 225 tokens with a minor performance drop (95.07% to 94.92%) and on AIME from 6827 to 4629 tokens, with only one additional incorrect answer.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Length-Compressible CoT Tuning
  • 3.2 Construct the MixChain Dataset
  • 3.3 Improved Tuning for CoT-Valve
  • Progressive Chain Compression: CoT-Valve+P.
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Datasets
  • 4.3 From Long-CoT to Short-CoT.
  • 4.4 From Short-CoT to Long-CoT & Short-Long-Short CoT
  • 4.5 Observations
  • 4.6 Analysis
  • 5 Conclusion
  • 6 Limitations
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Length-Compressible Chain-of-Thought Tuning (CoT-Valve)

    model/method

    CoT-Valve controls the reasoning length of large language models by identifying and adjusting an update direction Δθ\Delta \theta in the parameter space, implemented as a Low-Rank Adaptation (LoRA) branch. Given a reasoning model parameterized by θ\theta, an input prompt/question qq, and intermediate reasoning thought tokens {ti}i=1m\{t_i\}_{i=1}^m leading to a final answer aa, the parameter update Δθ\Delta \theta is learned by maximizing the likelihood of producing the correct answer with a shorter reasoning sequence:

    max⁡ΔθE(q,a)∼D[p(a∣t1,…,tm,q; θ+Δθ)∏i=1mp(ti∣t<i,q; θ+Δθ)]\max_{\Delta \theta} \mathbb{E}_{(q,a) \sim \mathcal{D}} \left[ p(a \mid t_1, \dots, t_m, q;\, \theta + \Delta \theta) \prod_{i=1}^m p(t_i \mid t_{<i}, q;\, \theta + \Delta \theta) \right]

    where m<nm < n denotes the length of the shorter reasoning path relative to the original chain length nn. At inference time, the parameter update is modulated by a scaling factor α\alpha as θ+αΔθ\theta + \alpha \Delta \theta. Setting α∈(0,1)\alpha \in (0, 1) interpolates the parameter state to smoothly transition between longer and shorter reasoning trajectories. Setting α>1\alpha > 1 extrapolates the update direction, compelling the model to generate reasoning paths shorter than those observed during fine-tuning.

  2. Knowl 2 — MixChain Dataset Construction

    model/method

    MixChain is a dataset comprising questions paired with multiple correct reasoning paths whose sequence lengths progressively decrease from long to short. It is constructed without repeated stochastic sampling via two distinct initialization strategies:

    1. Cold-Start Construction (MixChain-C): When human-annotated or ground-truth step-by-step solutions are available (e.g., from GSM8K or PRM800k), a model is first fine-tuned on the short solutions to define an initial update direction Δθ\Delta \theta. Adjusting the scaling factor α\alpha during inference produces diverse reasoning paths of systematically varied lengths for each question.

    2. Zero-Shot Parameter Interpolation (MixChain-Z): When short step-by-step reasoning annotations are unavailable, an update direction is obtained from the weight difference Δθ=θ2−θ1\Delta \theta = \theta_2 - \theta_1 between a base language model θ1\theta_1 (e.g., LLaMA-3.1-8B or Qwen-32B-Instruct) and its reasoning-specialized counterpart θ2\theta_2 (e.g., DeepSeek-R1-Distill-Llama-8B or QwQ-32B-Preview). Interpolating across α∈[0,1]\alpha \in [0, 1] between θ1\theta_1 and θ2\theta_2 generates chains spanning multiple length tiers. Incorrect solutions are filtered out.

  3. Knowl 3 — CoT-Valve++ Continuous Length-Constrained Tuning

    model/method

    CoT-Valve++ is an enhanced training formulation that resolves train-test discrepancies by training the parameter direction Δθ′\Delta \theta' to adhere to variable length constraints across all intermediate scaling factors. Given the MixChain dataset D′\mathcal{D}', each training instance consists of a question qq, target answer aa, candidate reasoning sequence {ti}i=1m\{t_i\}_{i=1}^m, and a normalized length ratio β∈[0,1]\beta \in [0, 1] calculated as:

    β=1−m−mmin⁡mmax⁡−mmin⁡\beta = 1 - \frac{m - m_{\min}}{m_{\max} - m_{\min}}

    where mmin⁡m_{\min} and mmax⁡m_{\max} represent the lengths of the shortest and longest candidate reasoning chains available for question qq. The optimization objective trains Δθ′\Delta \theta' subject to sample-specific β\beta scaling:

    max⁡Δθ′E(q,a)∼D′[p(a∣t<m,q; θ+βΔθ′)∏i=1mp(ti∣t<i,q; θ+βΔθ′)]\max_{\Delta \theta'} \mathbb{E}_{(q,a) \sim \mathcal{D}'} \left[ p\left(a \mid t_{<m}, q;\, \theta + \beta \Delta \theta'\right) \prod_{i=1}^m p\left(t_i \mid t_{<i}, q;\, \theta + \beta \Delta \theta'\right) \right]

    This continuous constraint enables finer and more predictable control over output reasoning length during inference.

  4. Knowl 4 — Progressive Chain Compression (CoT-Valve+P)

    model/method

    Progressive Chain Compression (CoT-Valve+P) is a multi-stage fine-tuning strategy that shortens reasoning paths iteratively rather than training directly on the shortest sequence. Using the length-ordered trajectories in the MixChain dataset (from longest Solution 4 down to shortest Solution 0 / ground-truth), the model is trained sequentially across successive stages or epochs, where each stage trains on the next shorter solution tier. This stepwise compression preserves intermediate reasoning quality, avoiding the severe performance degradation that occurs when attempting single-step distillation directly into concise outputs.

  5. Knowl 5 — Accuracy per Computation Unit (ACU)

    definition

    Accuracy per Computation Unit (ACU) is a composite evaluation metric that quantifies the trade-off between reasoning accuracy, model capacity, and inference token overhead:

    ACU=Accuracy#Params×#Tokens\text{ACU} = \frac{\text{Accuracy}}{\#\text{Params} \times \#\text{Tokens}}

    where Accuracy\text{Accuracy} is the task accuracy, #Params\#\text{Params} is the total parameter count of the model, and #Tokens\#\text{Tokens} is the average number of generated output tokens. ACU values are typically reported multiplied by 10210^2 for readability.

  6. Knowl 6 — Chain-of-Thought Compression Performance on GSM8K and AIME24

    data/table

    Applying CoT-Valve variants to QwQ-32B-Preview reduces token generation substantially with minimal accuracy impact, outperforming prompt-based length constraints and preference optimization methods across math benchmarks.

    Benchmark Method / Configuration Accuracy #Tokens ACU (×102\times 10^2)
    GSM8K QwQ-32B-Preview (Baseline) 95.1% 741.1 0.40
    GSM8K Prompt Control (Han et al., 2024) 93.6% 355.5 0.82
    GSM8K Prompt Control (Ding et al., 2024) 95.5% 617.7 0.48
    GSM8K Overthink SimPO (Chen et al., 2024) 94.8% 326.2 0.91
    GSM8K O1-Pruner (Luo et al., 2025) 96.5% 534.0 0.56
    GSM8K CoT-Valve (Ground-Truth) 94.0% 352.8 0.83
    GSM8K CoT-Valve++ (MixChain-C) 94.4% 276.3 1.07
    GSM8K CoT-Valve+P (MixChain-Z, 5-stage) 94.9% 225.5 1.32
    AIME24 QwQ-32B-Preview (Baseline) 14/30 6827.3 0.021
    AIME24 Prompt Control (Han et al., 2024) 13/30 6102.5 0.022
    AIME24 Prompt Control (Ding et al., 2024) 13/30 5562.3 0.024
    AIME24 Overthink (Chen et al., 2024) 13/30 5154.5 0.026
    AIME24 CoT-Valve++ (MixChain-C) 13/30 5360.5 0.025
    AIME24 CoT-Valve+P (MixChain-Z) 13/30 4629.6 0.029

    On GSM8K, CoT-Valve+P compresses reasoning length by over 69% (741.1 to 225.5 tokens) with only a 0.15% drop in accuracy (95.07% to 94.92%), more than tripling ACU from 0.40 to 1.32. On AIME24 under greedy decoding, CoT-Valve+P reduces the average token count from 6827.3 to 4629.6 tokens while maintaining 13/30 correct answers.

  7. Knowl 7 — Reasoning Distillation into Small Language Models with CoT-Valve

    empirical result

    When distilling reasoning capabilities from QwQ-32B-Preview into smaller language models (LLaMA-3.2-1B-Instruct and LLaMA-3.1-8B), CoT-Valve enables the small models to generate shorter reasoning paths that outperform standard full-length distillation on GSM8K:

    • On LLaMA-3.2-1B-Instruct, standard SFT on raw QwQ distillations achieves 52.7% accuracy with 759.3 tokens. CoT-Valve on QwQ distillations increases accuracy to 55.5% at 267.0 tokens. Fine-tuning with CoT-Valve on the concise tier of MixChain-Z achieves 58.9% accuracy with 275.4 tokens (compared to 43.8% accuracy and 137.7 tokens for SFT directly on GSM8K ground-truth).
    • On LLaMA-3.1-8B, SFT on QwQ distillations achieves 76.3% accuracy at 644.8 tokens. CoT-Valve achieves 77.5% accuracy at 569.8 tokens, and CoT-Valve+P on MixChain-Z achieves 77.1% accuracy at 371.2 tokens.
  8. Knowl 8 — Non-Monotonic Relationship Between Chain Length and Small Model Learnability

    empirical result

    Fine-tuning LLaMA-3.2-1B-Instruct on single solution subsets of varying token lengths from the MixChain-Z dataset on GSM8K reveals that intermediate-length reasoning paths yield optimal learning performance:

    • Ground-Truth / Solution 0 (average length: 116.0 tokens): 43.8% accuracy, 139.4 output tokens.
    • Solution 1 (average length: 279.6 tokens): 57.0% accuracy, 288.4 output tokens.
    • Solution 2 (average length: 310.7 tokens): 55.1% accuracy, 330.0 output tokens.
    • Solution 3 (average length: 386.7 tokens): 56.5% accuracy, 414.6 output tokens.
    • Solution 4 (average length: 497.2 tokens): 52.5% accuracy, 558.3 output tokens.

    Neither the shortest human-annotated chains nor the longest synthetic chains optimize performance; moderately short reasoning chains provide the necessary structural guidance for smaller models while avoiding unhelpful redundancy.

  9. Knowl 9 — Step-Level Correctness Verification of Compressed Chains Using PRMs

    empirical result

    Step-level reward evaluation of the MixChain-Z dataset derived from GSM8K using two process reward models (Qwen2.5-Math-PRM-72B and Skywork-o1-Open-PRM-7B) shows that shorter, compressed reasoning paths exhibit higher average step correctness than uncompressed chains:

    • Ground-truth GSM8K chains (121.8 tokens): 88.02% average PRM reward (Qwen2.5-Math-PRM) and 77.93% (Skywork-o1-PRM).
    • QwQ-32B-Preview raw chains (737.3 tokens): 98.04% (Qwen2.5-Math-PRM) and 89.16% (Skywork-o1-PRM).
    • Solution 4 (497.2 tokens): 99.00% (+0.96) and 91.76% (+2.60).
    • Solution 3 (386.7 tokens): 99.53% (+1.49) and 92.66% (+3.50).
    • Solution 2 (310.7 tokens): 99.82% (+1.78) and 93.79% (+4.63).
    • Solution 1 (279.6 tokens): 99.85% (+1.81) and 93.47% (+4.31).

    Systematic chain compression removes redundant or noisy exploration steps, increasing the step correctness score across the entire reasoning trace.

  10. Knowl 10 — Limitations of CoT-Valve

    limitation

    CoT-Valve exhibits several key limitations:

    1. Accuracy Degradation Under Extreme Compression on Hard Tasks: While elementary reasoning tasks tolerate aggressive chain reduction, complex multi-step reasoning tasks (such as AIME24) suffer accuracy degradation when reasoning length is excessively compressed due to insufficient test-time compute.
    2. Dependency on Base Model Generation: The construction of MixChain and the derivation of the parameter vector rely on the existence of a high-quality reasoning model; if the base model cannot generate valid solutions, effective short-chain dataset construction fails.
    3. Lack of Theoretical Interpretability: The linear modulation of reasoning sequence length via directional arithmetic in LoRA parameter space is empirical, and its theoretical foundations remain insufficiently understood.

Coverage note — Omitted minor secondary evaluations on general non-reasoning NLP benchmarks (MMLU, PIQA, HellaSwag in Table 9), which served only as side sanity checks confirming that CoT-Valve does not cause catastrophic forgetting.

References

  1. 1.Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905.
  2. 2.Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, et al. 2024. Do not think that much for 2+ 3=? on the overthinking of o1-like llms. arXiv preprint arXiv:2412.21187.
  3. 3.Jeffrey Cheng and Benjamin Van Durme. 2024. Compressed chain of thought: Efficient reasoning through dense representations. Preprint, arXiv:2412.13171.
  4. 4.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021a. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  5. 5.Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021b. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  6. 6.DeepSeek-AI. 2025. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. Preprint, arXiv:2501.12948.
  7. 7.Yuntian Deng, Yejin Choi, and Stuart Shieber. 2024. From explicit cot to implicit cot: Learning to internalize cot step by step. Preprint, arXiv:2405.14838.
  8. 8.Yuntian Deng, Kiran Prasad, Roland Fernandez, Paul Smolensky, Vishrav Chaudhary, and Stuart Shieber. 2023. Implicit chain of thought reasoning via knowledge distillation. arXiv preprint arXiv:2311.01460.
  9. 9.Mengru Ding, Hanmeng Liu, Zhizhang Fu, Jian Song, Wenbo Xie, and Yue Zhang. 2024. Break the chain: Large language models can be shortcut reasoners. arXiv preprint arXiv:2406.06580.
  10. 10.Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
  11. 11.Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020. Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning, pages 3259–3269. PMLR.
  12. 12.Xinyu Guan, Li Lyna Zhang, Yifei Liu, Ning Shang, Youran Sun, Yi Zhu, Fan Yang, and Mao Yang. 2025. rstar-math: Small llms can master math reasoning with self-evolved deep thinking. Preprint, arXiv:2501.04519.
  13. 13.Tingxu Han, Chunrong Fang, Shiyu Zhao, Shiqing Ma, Zhenyu Chen, and Zhenting Wang. 2024. Token-budget-aware llm reasoning. arXiv preprint arXiv:2412.18547.
  14. 14.Shibo Hao, Sainbayar Sukhbaatar, DiJia Su, Xian Li, Zhiting Hu, Jason Weston, and Yuandong Tian. 2024. Training large language models to reason in a continuous latent space. Preprint, arXiv:2412.06769.
  15. 15.Jujie He, Tianwen Wei, Rui Yan, Jiacai Liu, Chaojie Wang, Yimeng Gan, Shiwen Tu, Chris Yuhao Liu, Liang Zeng, Xiaokun Wang, Boyang Wang, Yongcong Li, Fuxiang Zhang, Jiacheng Xu, Bo An, Yang Liu, and Yahui Zhou. 2024. Skywork-o1 open series. https://huggingface.co/Skywork.
  16. 16.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training compute-optimal large language models. Preprint, arXiv:2203.15556.
  17. 17.Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations.
  18. 18.Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023. Editing models with task arithmetic. In The Eleventh International Conference on Learning Representations.
  19. 19.Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024. Openai o1 system card. arXiv preprint arXiv:2412.16720.
  20. 20.Mingyu Jin, Qinkai Yu, Dong Shu, Haiyan Zhao, Wenyue Hua, Yanda Meng, Yongfeng Zhang, and Mengnan Du. 2024. The impact of reasoning step length on large language models. In Findings of the Association for Computational Linguistics ACL 2024, pages 1830–1842, Bangkok, Thailand and virtual meeting.
  21. 21.Brihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, and Xiang Ren. 2023. Are machine rationales (not) useful to humans? measuring and improving human utility of free-text rationales. arXiv preprint arXiv:2305.07095.
  22. 22.Yu Kang, Xianghui Sun, Liangyu Chen, and Wei Zou. 2024. C3ot: Generating shorter chain-of-thought without compromising effectiveness. Preprint, arXiv:2412.11664.
  23. 23.Levente Kocsis and Csaba Szepesvari. 2006. Bandit based monte-carlo planning. In European Conference on Machine Learning.
  24. 24.Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. 2024. Let’s verify step by step. In The Twelfth International Conference on Learning Representations.
  25. 25.Shih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov, Yu-Chiang Frank Wang, Kwang-Ting Cheng, and Min-Hung Chen. 2024a. Dora: Weight-decomposed low-rank adaptation. In ICML.
  26. 26.Tengxiao Liu, Qipeng Guo, Xiangkun Hu, Cheng Jiayang, Yue Zhang, Xipeng Qiu, and Zheng Zhang. 2024b. Can language models learn to skip steps? In The Thirty-eighth Annual Conference on Neural Information Processing Systems.
  27. 27.Haotian Luo, Li Shen, Haiying He, Yibo Wang, Shiwei Liu, Wei Li, Naiqiang Tan, Xiaochun Cao, and Dacheng Tao. 2025. O1-pruner: Length-harmonizing fine-tuning for o1-like reasoning pruning. arXiv preprint arXiv:2501.12570.
  28. 28.Liangchen Luo, Yinxiao Liu, Rosanne Liu, Samrat Phatale, Harsh Lara, Yunxuan Li, Lei Shu, Yun Zhu, Lei Meng, Jiao Sun, et al. 2024. Improve mathematical reasoning in language models by automated process supervision. arXiv preprint arXiv:2406.06592.
  29. 29.Nanye Ma, Shangyuan Tong, Haolin Jia, Hexiang Hu, Yu-Chuan Su, Mingda Zhang, Xuan Yang, Yandong Li, Tommi Jaakkola, Xuhui Jia, and Saining Xie. 2025. Inference-time scaling for diffusion models beyond scaling denoising steps. Preprint, arXiv:2501.09732.
  30. 30.Yu Meng, Mengzhou Xia, and Danqi Chen. 2024. Simpo: Simple preference optimization with a reference-free reward. In Advances in Neural Information Processing Systems (NeurIPS).
  31. 31.Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz. 2017. Pruning convolutional neural networks for resource efficient inference. In International Conference on Learning Representations.
  32. 32.Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Ouyang Long, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021. Webgpt: Browser-assisted question-answering with human feedback. ArXiv, abs/2112.09332.
  33. 33.Hao Shao, Shengju Qian, Han Xiao, Guanglu Song, Zhuofan Zong, Letian Wang, Yu Liu, and Hongsheng Li. 2024. Visual cot: Unleashing chain-of-thought reasoning in multi-modal language models. Preprint, arXiv:2403.16999.
  34. 34.Zayne Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett. 2024. To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning. arXiv preprint arXiv:2409.12183.
  35. 35.Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.
  36. 36.Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, Chenzhuang Du, Chonghua Liao, et al. 2025. Kimi k1. 5: Scaling reinforcement learning with llms. arXiv preprint arXiv:2501.12599.
  37. 37.Qwen Team. 2024a. Qwen2.5: A party of foundation models.
  38. 38.Qwen Team. 2024b. Qwq: Reflect deeply on the boundaries of the unknown.
  39. 39.Jean-Francois Ton, Muhammad Faaiz Taufiq, and Yang Liu. 2024. Understanding chain-of-thought in llms through information theory. Preprint, arXiv:2411.11984.
  40. 40.Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. 2024. Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9426–9439, Bangkok, Thailand. Association for Computational Linguistics.
  41. 41.Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
  42. 42.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.
  43. 43.Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik R Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. In Thirty-seventh Conference on Neural Information Processing Systems.
  44. 44.Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu. 2025. Limo: Less is more for reasoning. arXiv preprint arXiv:2502.03387.
  45. 45.Ping Yu, Jing Xu, Jason Weston, and Ilia Kulikov. 2024. Distilling system 2 into system 1. arXiv preprint arXiv:2407.06023.
  46. 46.Di Zhang, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, and Wanli Ouyang. 2024. Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b. Preprint, arXiv:2406.07394.
  47. 47.Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2025a. The lessons of developing process reward models in mathematical reasoning. Preprint, arXiv:2501.07301.
  48. 48.Zhenru Zhang, Chujie Zheng, Yangzhen Wu, Beichen Zhang, Runji Lin, Bowen Yu, Dayiheng Liu, Jingren Zhou, and Junyang Lin. 2025b. The lessons of developing process reward models in mathematical reasoning. arXiv preprint arXiv:2501.07301.

Citation

MLA
Ma, X., et al. “CoT-Valve: Length-Compressible Chain-of-Thought Tuning”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 6025–35, https://doi.org/10.18653/v1/2025.acl-long.300.
APA
Ma, X., Wan, G., Yu, R., Fang, G., & Wang, X. (2025). CoT-Valve: Length-Compressible Chain-of-Thought Tuning. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6025–6035. https://doi.org/10.18653/v1/2025.acl-long.300
Chicago
Ma, X., G. Wan, R. Yu, G. Fang, and X. Wang. 2025. “CoT-Valve: Length-Compressible Chain-of-Thought Tuning”. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6025–35. https://doi.org/10.18653/v1/2025.acl-long.300.
Harvard
Ma, X. et al. (2025) “CoT-Valve: Length-Compressible Chain-of-Thought Tuning”, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6025–6035. Available at: https://doi.org/10.18653/v1/2025.acl-long.300.
Vancouver
1. Ma X, Wan G, Yu R, Fang G, Wang X (2025) CoT-Valve: Length-Compressible Chain-of-Thought Tuning. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6025–6035

BibTeX

@inproceedings{ma-etal-2025-cot,
    title = "{C}o{T}-Valve: Length-Compressible Chain-of-Thought Tuning",
    author = "Ma, Xinyin  and
      Wan, Guangnian  and
      Yu, Runpeng  and
      Fang, Gongfan  and
      Wang, Xinchao",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.300/",
    doi = "10.18653/v1/2025.acl-long.300",
    pages = "6025--6035",
    ISBN = "979-8-89176-251-0"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/