CoT-Valve: Length-Compressible Chain-of-Thought Tuning
Xinyin MaGuangnian WanRunpeng YuGongfan FangXinchao Wang
Proposes CoT-Valve, a parameter-space tuning method that enables a single reasoning model to dynamically control and compress chain-of-thought length, slashing token costs on benchmarks like GSM8K and AIME with minimal accuracy loss.
Modern artificial intelligence models rely heavily on multi-step reasoning chains to solve complex mathematical and logical tasks. While this reasoning process boosts accuracy, it frequently leads models to generate excessively lengthy intermediate steps even for trivial questions, causing high computational overhead, increased latency, and inflated operational costs. Existing methods to reduce reasoning length, such as prompt-based instructions, struggle to reliably control output size and fail to generate highly compact explanations.
The article introduces and evaluates CoT-Valve, a novel tuning and inference framework designed to elastically control the length of reasoning paths in a single model by manipulating a specific update direction in its parameter space.
The researchers implemented CoT-Valve by isolating length-controlling update directions using lightweight low-rank adaptation modules. They constructed a multi-length reasoning dataset called MixChain, pairing long and short valid explanations for identical questions. Using this data, they evaluated two strategies: a precise continuous tuning method and a progressive compression approach that iteratively trains models on incrementally shorter reasoning paths. The evaluations spanned multiple open model architectures, including standard base models, reasoning-tuned models, and distilled variants, benchmarked across elementary and competitive mathematical datasets.
The analysis produced several critical findings. First, CoT-Valve compressed the average reasoning output of a thirty-two-billion-parameter reasoning model on elementary math from 741 tokens down to 225 tokens while maintaining accuracy above 94.9%, substantially outperforming prompt-based controls. Second, on challenging competition mathematics, the method reduced reasoning length from 6,827 tokens to 4,629 tokens with only one additional incorrect answer. Third, on simpler tasks, shorter reasoning paths frequently outperformed long chains, particularly in smaller models where direct training on verbose reasoning degraded task accuracy. Finally, process reward evaluations showed that intermediate reasoning steps generated through this compression method exhibited higher correctness scores by eliminating noisy and redundant reasoning.
These findings indicate that organizations deploying advanced reasoning models can dynamically tune inference cost and speed without retraining separate architectures for simple and complex user queries. By selectively scaling down token usage on less demanding tasks, teams can achieve significant cost savings and latency reductions while maintaining strong baseline performance on standard non-reasoning benchmarks.
Decision-makers should consider pilot implementations of parameter-based length modulation to optimize serving costs across varying task difficulties. Future development should explore fine-grained token compression that selectively shortens trivial segments of a reasoning path while preserving long-chain deliberation for complex intermediate steps.
The study's primary limitation is that aggressive compression on highly intricate problems can lead to degraded solution accuracy, as complex tasks still depend on extended deliberation. Additionally, the approach relies on the base model's initial generation capability, meaning that poor underlying reasoning cannot be effectively compressed. Confidence in the reported efficiency gains remains high across standard reasoning benchmarks, though careful task-level evaluation is recommended before applying extreme compression.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Introduces chain-of-thought prompting for multi-step reasoning in language models, establishing the fundamental paradigm that CoT-Valve compresses.
- Paper: Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters, Charlie Snell et al. (2024). Analyzes the trade-offs of test-time compute allocation across difficulty tiers, directly motivating length modulation for variable-complexity tasks.
- Paper: Teaching Small Language Models to Reason, Lucie Charlotte Magister et al. (2023). Demonstrates reasoning distillation into smaller language models, providing key context for how verbose chains impact smaller architectures.
- Paper: Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations, Peiyi Wang et al. (2024). Establishes automated step-level process reward modeling to evaluate intermediate reasoning steps, which CoT-Valve employs to benchmark chain correctness.
- Paper: Training Large Language Models to Reason in a Continuous Latent Space, Shibo Hao et al. (2024). Explores latent-space continuous reasoning to reduce chain-of-thought token overhead, offering an essential counterpart to parameter-space length control.
- Paper: Self-Consistency Improves Chain of Thought Reasoning in Language Models, Xuezhi Wang et al. (2023). Introduces majority-voting self-consistency over multiple reasoning paths, illustrating the trade-off between chain-of-thought sampling cost and task accuracy.
- Paper: Composable Text Controls in Latent Space with ODEs, Guangyi Liu et al. (2023). Presents methods for continuous, low-rank parameter-space text attribute control that inform continuous parameter manipulation techniques.
- Paper: s1: Simple test-time scaling, Niklas Muennighoff et al. (2025). Introduces budget forcing to explicitly regulate test-time reasoning token lengths, continuing the exploration of controllable inference compute.
- Paper: Efficient Reasoning on the Edge, Yelysei Bondarenko et al. (2026). Applies LoRA-based fine-tuning and length-penalizing budget constraints to optimize reasoning models specifically for resource-constrained edge devices.
- Paper: Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought, Violet Xiang et al. (2025). Investigates Meta-CoT and internal deliberation mechanics to structure how models allocate thinking tokens across varying problem complexities.
- Paper: Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers, Ying Fan et al. (2026). Addresses chain-of-thought decoding latency by replacing explicit token streams with recurrent latent refinements under aligned step supervision.
- Paper: Understanding R1-Zero-Like Training: A Critical Perspective, Zichen Liu et al. (2025). Analyzes how reinforcement learning optimization and prompt priors govern response length and self-reflection in reasoning models.
- Paper: Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models, Fengli Xu et al. (2025). Provides a comprehensive survey synthesizing reinforcement learning, process supervision, and test-time compute scaling strategies in large reasoning models.
- Paper: Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning, Heng Wang et al. (2026). Investigates key-value cache eviction methods to reduce the physical hardware memory footprint generated by extensive reasoning traces.
