Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning
Mukul SinghAnanya SinghaArjun RadhakrishnaSumit Gulwani
Demonstrates that language models mirror human stages of competence by expanding intermediate reasoning steps during training and discarding them upon mastery, establishing reasoning token dynamics as a practical metric to track convergence and guide early stopping.
Modern artificial intelligence increasingly relies on generating intermediate "reasoning tokens" to solve complex tasks, but training these models is computationally intensive and can lead to issues like catastrophic forgetting. The article addresses how explicit reasoning evolves during model training and how tracking this progression can diagnose learning stages and optimize training efficiency. The primary objective is to demonstrate that language models follow cognitive learning trajectories—specifically mirroring the Four Stages of Competence—and to evaluate whether reasoning token dynamics can serve as reliable metrics for model convergence and early stopping.
The authors conducted reinforcement learning experiments across diverse tasks, including low-resource code generation, GSM8K mathematical problems, and CommonsenseQA logical question answering. They evaluated multiple small and large models—such as Phi-4, DeepSeek distillations, and GPT variants—using standard reinforcement learning algorithms over 100 training iterations. The team tracked reasoning token length, final answer correctness, the factual accuracy and completeness of the intermediate reasoning steps, and model performance when reasoning was disabled.
The investigation revealed four key findings. First, reasoning behavior follows a consistent four-phase arc across domains and models: an initial reasoning phase where attempts begin without accuracy gains, an acquisition phase where reasoning and accuracy rise together, a learning phase where task accuracy peaks alongside reasoning length, and a final phase where accuracy plateaus while reasoning token length declines. Second, explicit reasoning functions as temporary scaffolding; once fully trained, models maintain high accuracy even when reasoning tokens are turned off. Third, intermediate reasoning correctness actually decreases in later stages while final task accuracy remains high, indicating internal task automation. Fourth, as reasoning condenses past its peak, models experience rapid performance drops on held-out tasks, directly linking the reasoning contraction phase to catastrophic forgetting.
These findings imply that continuous training beyond peak reasoning increases compute costs without improving accuracy and heightens the risk of losing general capabilities. Consequently, organizations can leverage reasoning token length as a real-time signal for early stopping and convergence detection, reducing unnecessary training spend while preserving generalist performance. The article recommends that engineering teams monitor reasoning length and disable reasoning generation during late-stage deployment when internal competence is achieved.
Decision-makers should note certain limitations: reasoning metrics remain surface-level text generation artifacts rather than proof of human-like cognition, and experiments focused on structured domains like math and code. Further validation is needed in ambiguous or unstructured environments before applying these metrics as universal training heuristics.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). It establishes the foundational paradigm of chain-of-thought prompting that the source investigates across stages of model fine-tuning.
- Paper: STaR: Bootstrapping Reasoning With Reasoning, Eric Zelikman et al. (2022). It introduces the self-taught bootstrapping framework where language models learn by generating their own intermediate reasoning rationales.
- Paper: Why think step by step? Reasoning emerges from the locality of experience, Ben Prystawski et al. (2023). It provides the foundational theoretical and statistical motivation for why intermediate step-by-step reasoning improves inference in language models.
- Paper: ReFT: Reasoning with Reinforced Fine-Tuning, Luong Quoc Trung et al. (2024). It details how reinforcement learning and fine-tuning optimize reasoning paths during training, providing essential context for training dynamic analysis.
- Paper: Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters, Boshi Wang et al. (2023). It investigates the underlying mechanics and role of reasoning tokens in model predictions, serving as foundational background for reasoning token dynamics.
- Paper: Fractured Chain-of-Thought Reasoning, Baohao Liao et al. (2026). It directly builds upon the insight that full reasoning chains are redundant by exploring fractured and truncated intermediate reasoning stages during inference.
- Paper: TokenSkip: Controllable Chain-of-Thought Compression in LLMs, Heming Xia et al. (2025). It applies the concept of reasoning token redundancy by introducing a controllable mechanism to skip unnecessary reasoning tokens during generation.
- Paper: CoT-Valve: Length-Compressible Chain-of-Thought Tuning, Xinyin Ma et al. (2025). It extends the finding of reasoning token compression by providing a parameter-space tuning framework that elastically modulates reasoning chain length.
- Paper: Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models, Xingyu Chen et al. (2025). It investigates the practical consequences of overthinking and explores self-training methods to shrink excessive reasoning token usage on simple tasks.
- Paper: Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills, Agamdeep Singh et al.. It takes the notion of internalizing and omitting intermediate reasoning to the agent setting by amortizing reasoning chains into distilled static skills.
- Paper: Bridging the Gap Between Latent and Explicit Reasoning with Looped Transformers, Ying Fan et al. (2026). It explores transitioning explicit token-by-token reasoning chains into compact latent recurrent representations.
- Paper: Evaluating Mathematical Reasoning Beyond Accuracy, Shijie Xia et al. (2025). It provides evaluation metrics that assess the intermediate validity and token redundancy of reasoning paths beyond final-answer accuracy.
