Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning

Mukul SinghAnanya SinghaArjun RadhakrishnaSumit Gulwani

article2025arXiv0 citations

Demonstrates that language models mirror human stages of competence by expanding intermediate reasoning steps during training and discarding them upon mastery, establishing reasoning token dynamics as a practical metric to track convergence and guide early stopping.

Listen

Modern artificial intelligence increasingly relies on generating intermediate "reasoning tokens" to solve complex tasks, but training these models is computationally intensive and can lead to issues like catastrophic forgetting. The article addresses how explicit reasoning evolves during model training and how tracking this progression can diagnose learning stages and optimize training efficiency. The primary objective is to demonstrate that language models follow cognitive learning trajectories—specifically mirroring the Four Stages of Competence—and to evaluate whether reasoning token dynamics can serve as reliable metrics for model convergence and early stopping.

The authors conducted reinforcement learning experiments across diverse tasks, including low-resource code generation, GSM8K mathematical problems, and CommonsenseQA logical question answering. They evaluated multiple small and large models—such as Phi-4, DeepSeek distillations, and GPT variants—using standard reinforcement learning algorithms over 100 training iterations. The team tracked reasoning token length, final answer correctness, the factual accuracy and completeness of the intermediate reasoning steps, and model performance when reasoning was disabled.

The investigation revealed four key findings. First, reasoning behavior follows a consistent four-phase arc across domains and models: an initial reasoning phase where attempts begin without accuracy gains, an acquisition phase where reasoning and accuracy rise together, a learning phase where task accuracy peaks alongside reasoning length, and a final phase where accuracy plateaus while reasoning token length declines. Second, explicit reasoning functions as temporary scaffolding; once fully trained, models maintain high accuracy even when reasoning tokens are turned off. Third, intermediate reasoning correctness actually decreases in later stages while final task accuracy remains high, indicating internal task automation. Fourth, as reasoning condenses past its peak, models experience rapid performance drops on held-out tasks, directly linking the reasoning contraction phase to catastrophic forgetting.

These findings imply that continuous training beyond peak reasoning increases compute costs without improving accuracy and heightens the risk of losing general capabilities. Consequently, organizations can leverage reasoning token length as a real-time signal for early stopping and convergence detection, reducing unnecessary training spend while preserving generalist performance. The article recommends that engineering teams monitor reasoning length and disable reasoning generation during late-stage deployment when internal competence is achieved.

Decision-makers should note certain limitations: reasoning metrics remain surface-level text generation artifacts rather than proof of human-like cognition, and experiments focused on structured domains like math and code. Further validation is needed in ambiguous or unstructured environments before applying these metrics as universal training heuristics.

arXiv: 2511.21743
Cover for Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning

Abstract

We analyze reasoning in language models during task-specific fine-tuning and draws parallel between reasoning tokens--intermediate steps generated while solving problem and the human working memory. Drawing from cognitive science, we align training dynamics with the Four Stages of Competence: models initially produce incorrect outputs without reasoning, then begin reasoning (but still fail), eventually reason effectively, and finally solve tasks without explicit reasoning. We find that reasoning token length expands as performance improves, peaks at the stage of conscious competence, then declines as the model internalizes the task. Notably, after training, models retain performance even when reasoning is removed--suggesting it scaffolded learning but is no longer needed. This progression offers actionable insights: reasoning token dynamics can serve as a signal for diagnosing training stage, identifying convergence, and guiding early stopping. We propose metrics to track this trajectory and argue that reasoning behavior is valuable for understanding and optimizing reasoning model training.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 3 Reasoning and Memory
  • 3.1 Four stages of competence
  • 3.2 Reasoning as scaffold to distill knowledge
  • 3.3 Learning stages
  • 4 Experimental Setup
  • 4.1 Benchmarks
  • 4.2 Metrics
  • 4.3 Models and Configuration
  • 5 Results
  • 5.1 RQ1: Stages in Learning
  • 5.2 RQ2: Reasoning as Scaffolding
  • 5.3 RQ3: Tracking Training Stages
  • 6 Conclusion
  • 7 Limitations
  • References

Knowls

  1. Knowl 1 — Four-Stage Competence Acquisition Model for Reasoning Language Models

    model/method

    A cognitive framework mapping the "Four Stages of Competence" to the reinforcement learning dynamics of reasoning language models, where generated intermediate reasoning tokens represent temporary working memory and trained model parameters represent long-term memory:

    1. Reasoning Phase (Unconscious to Conscious Incompetence): The model begins to allocate inference-time budget toward generating reasoning tokens to solve the task, but task performance (accuracy) remains low and unimproved.
    2. Acquisition Phase (Conscious Incompetence to Competence): Task performance starts improving substantially, driven directly by the increasing utilization and length of explicit reasoning traces.
    3. Learning/Distillation Phase (Conscious Competence): Reasoning token length reaches its maximum peak while task performance reaches its upper plateau. The model achieves high accuracy but requires maximum deliberative effort.
    4. Condensation Phase (Unconscious Competence): Task performance remains stable at high accuracy while reasoning token length contracts significantly. The inferential steps that were previously explicitly generated in working memory become internalized directly into the model's parameter weights.
  2. Knowl 2 — Non-Monotonic Trajectory of Reasoning Token Length During RL Training

    empirical result

    Across reinforcement learning training of language models (including Phi-4-13B, Phi-4-reasoning, DeepSeek-distill-7B, DeepSeek-distill-23B, GPT-4o-mini, GPT-o4-mini, and GPT-4o) on mathematical reasoning (GSM8K), logical QA (CommonsenseQA), and low-resource code generation (Rust, OCaml, PHP), reasoning token length exhibits a consistent, domain-agnostic, non-monotonic curve as training steps progress:

    • Token Expansion: In the initial training steps (0 to ~3000 steps), the average number of generated reasoning tokens increases rapidly alongside rising task accuracy.
    • Token Peak: Reasoning token length reaches an empirical peak (e.g., ~35–40 tokens on benchmark scales) at intermediate training steps (around 4000–5000 steps), coinciding with peak task accuracy (~80%–90%).
    • Token Contraction: In later training steps (5000 to >8000 steps), task accuracy remains plateaued while reasoning token length monotonically declines (condensing to ~5–15 tokens). The model continues to solve tasks correctly while generating increasingly brief reasoning traces.
  3. Knowl 3 — Scaffolding Effect and Inference Without Reasoning in Competent Models

    empirical result

    During reinforcement learning on task-specific domains, explicit reasoning tokens act as temporary scaffolding for distilling task-solving logic into model parameters:

    • In early training stages, models fail to solve complex tasks if reasoning tokens are disabled at inference time.
    • Once training progresses into the late condensation phase (unconscious competence), models retain their task accuracy even when reasoning token generation is completely turned off and the model operates as a direct completion model.
    • This demonstrates that explicit deliberation is essential during initial skill acquisition but becomes redundant once the underlying skill has been consolidated into the model weights.
  4. Knowl 4 — Catastrophic Forgetting Triggered by the Reasoning Condensation Phase

    empirical result

    Tracking out-of-domain task performance during single-domain reinforcement learning reveals a direct link between primary-task reasoning condensation and catastrophic forgetting on held-out tasks:

    • During the reasoning expansion and peak conscious competence phases on the primary task (steps 0 to ~5000), performance on held-out tasks remains largely intact (e.g., maintaining ~70%–80% accuracy).
    • As soon as the primary-task reasoning length peaks and enters the condensation phase (steps >5000), accuracy on held-out domains declines sharply (dropping below 55% by step 8000).
    • Extended training that further compresses primary-task reasoning traces accelerates the loss of out-of-domain knowledge, indicating that the peak of reasoning length can serve as an operational early-stopping criterion to prevent catastrophic forgetting.
  5. Knowl 5 — Reasoning Trace Degradation Relative to Final Answer Accuracy During RL Training

    empirical result

    Evaluating checkpoints every 50 training iterations for reasoning correctness (absence of factually incorrect statements), reasoning completeness (inclusion of all intermediate inferential steps without omission), and final answer correctness demonstrates a divergence in late-stage training:

    • In early-to-mid training iterations (steps 1000 to ~4000), the proportion of responses with both correct/complete reasoning traces and correct final answers grows steadily.
    • In late-stage training (steps >5000), overall task accuracy continues to remain high or increase, but the proportion of fully correct and complete reasoning traces decreases.
    • The model increasingly produces correct final answers accompanied by abbreviated, skipped, or formally incomplete reasoning traces, reflecting that intermediate reasoning steps are being executed implicitly within internal representations rather than fully expressed in the output token sequence.
  6. Knowl 6 — Reinforcement Learning Experimental Framework for Tracking Reasoning Evolution

    experimental setup

    The experimental environment designed to observe the evolution of reasoning tokens across competence stages:

    • Training Algorithms: Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), and REINFORCE.
    • Training Compute & Schedule: 100 iterations with 32 rollouts per iteration, conducted on an 8 ×\times NVIDIA H100 GPU cluster, averaging 2 hours and 35 minutes of training time per model.
    • Evaluated Models: DeepSeek-distill-7B, DeepSeek-distill-23B, Phi-4-13B, Phi-4-reasoning, GPT-4o-mini, GPT-o4-mini, and GPT-4o.
    • Task Domains & Rewards:
      1. Code Generation: Multiple-choice questions (MCQ) covering code generation, repair, and comprehension in low-resource languages (Rust, OCaml, PHP). Reward is a binary scalar based on final answer correctness via test execution.
      2. Mathematical Reasoning: GSM8K word problems with boolean correctness rewards on the final answer.
      3. Logical & Commonsense Reasoning: CommonsenseQA benchmark featuring MCQ and word-based logical questions with boolean correctness rewards on the final answer.
    • Diagnostics: Evaluated every 50 steps across four metrics: reasoning token count, final answer accuracy with reasoning, reasoning trace factual correctness and completeness, and final answer accuracy without reasoning (reasoning tokens suppressed).
  7. Knowl 7 — Limitations of Reasoning Token Dynamics as Cognitive Abstractions

    limitation

    The interpretation of language model reasoning tokens through cognitive competence frameworks is subject to several key constraints:

    • Surface-Level Proxies: Reasoning token length and syntactic correctness are surface-level textual artifacts of autoregressive decoding; brevity or omission of tokens in late training does not demonstrate biological or human-like cognitive internalization.
    • Domain Specificity: Experimental evaluations were restricted to structured, rule-governed domains (elementary mathematics, logical QA, and low-resource coding MCQ); dynamics in tasks requiring social reasoning, contextual ambiguity, or real-world embodied grounding remain unverified.
    • Learning Environment Sensitivity: The stability of the four competence stages may degrade in complex training regimes such as noisy supervision, multi-task curriculum training, or wide-distribution instruction tuning where reasoning traces are not distinctly separated from final outputs.

Coverage note — None was omitted; all primary conceptual contributions, empirical training dynamics, scaffolding observations, forgetting analyses, experimental settings, and stated limitations are fully covered.

References

  1. 1.J. Ahn, R. Verma, R. Lou, D. Liu, R. Zhang, and W. Yin. Large language models for mathematical reasoning: Progresses and challenges, 2024.
  2. 2.F. Cassano, J. Gouwar, D. Nguyen, S. Nguyen, L. Phipps-Costin, D. Pinckney, M.-H. Yee, Y. Zi, C. J. Anderson, M. Q. Feldman, et al. Multipl-e: a scalable and polyglot approach to benchmarking neural code generation. IEEE Transactions on Software Engineering, 49(7):3675–3691, 2023.
  3. 3.X. Chen, H. Dai, Y. Li, X. Gao, and L. Song. Learning to stop while learning to predict, 2020.
  4. 4.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman. Training verifiers to solve math word problems, 2021.
  5. 5.DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.
  6. 6.S. S. Ghosal, S. Chakraborty, A. Reddy, Y. Lu, M. Wang, D. Manocha, F. Huang, M. Ghavamzadeh, and A. S. Bedi. Does thinking more always help? understanding test-time scaling in reasoning models, 2025.
  7. 7.S. Kambhampati, K. Stechly, K. Valmeekam, L. Saldyt, S. Bhambri, V. Palod, A. Gundawar, S. R. Samineni, D. Kalwar, and U. Biswas. Stop anthropomorphizing intermediate tokens as reasoning/thinking traces!, 2025.
  8. 8.Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago, et al. Competition-level code generation with alphacode. Science, 378(6624):1092–1097, 2022.
  9. 9.Z. Liu, P. Wang, R. Xu, S. Ma, C. Ruan, P. Li, Y. Liu, and Y. Wu. Inference-time scaling for generalist reward modeling, 2025.
  10. 10.A. Lozhkov, R. Li, L. Ben Allal, F. Cassano, J. Lamy-Poirier, N. Tazi, A. Tang, D. Pykhtar, et al. StarCoder 2 and The Stack v2: The Next Generation. arXiv preprint arXiv:2402.19173, 2024.
  11. 11.S. M. Mousavi, E. Cecchinato, L. Hornikova, and G. Riccardi. Garbage in, reasoning out? why benchmark scores are unreliable and what to do about it, 2025.
  12. 12.K. Oberauer. Working memory and attention – a conceptual analysis and review. Journal of Cognition, 2(1):36, 2019.
  13. 13.OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kondrich, A. Mishchenko, A. Applebaum, A. Jiang, A. Nair, B. Zoph, B. Ghorbani, B. Rossen, B. Sokolowsky, B. Barak, B. McGrew, B. Minaiev, B. Hao, B. Baker, B. Houghton, B. McKinzie, B. Eastman, C. Lugaresi, C. Bassin, C. Hudson, C. M. Li, C. de Bourcy, C. Voss, C. Shen, C. Zhang, C. Koch, C. Orsinger, C. Hesse, C. Fischer, C. Chan, D. Roberts, D. Kappler, D. Levy, D. Selsam, D. Dohan, D. Farhi, D. Mely, D. Robinson, D. Tsipras, D. Li, D. Oprica, E. Freeman, E. Zhang, E. Wong, E. Proehl, E. Cheung, E. Mitchell, E. Wallace, E. Ritter, E. Mays, F. Wang, F. P. Such, F. Raso, F. Leoni, F. Tsimpourlas, F. Song, F. von Lohmann, F. Sulit, G. Salmon, G. Parascandolo, G. Chabot, G. Zhao, G. Brockman, G. Leclerc, H. Salman, H. Bao, H. Sheng, H. Andrin, H. Bagherinezhad, H. Ren, H. Lightman, H. W. Chung, I. Kivlichan, I. O’Connell, I. Osband, I. C. Gilaberte, I. Akkaya, I. Kostrikov, I. Sutskever, I. Kofman, J. Pachocki, J. Lennon, J. Wei, J. Harb, J. Twore, J. Feng, J. Yu, J. Weng, J. Tang, J. Yu, J. Q. Candela, J. Palermo, J. Parish, J. Heidecke, J. Hallman, J. Rizzo, J. Gordon, J. Uesato, J. Ward, J. Huizinga, J. Wang, K. Chen, K. Xiao, K. Singhal, K. Nguyen, K. Cobbe, K. Shi, K. Wood, K. Rimbach, K. Gu-Lemberg, K. Liu, K. Lu, K. Stone, K. Yu, L. Ahmad, L. Yang, L. Liu, L. Maksin, L. Ho, L. Fedus, L. Weng, L. Li, L. McCallum, L. Held, L. Kuhn, L. Kondraciuk, L. Kaiser, L. Metz, M. Boyd, M. Trebacz, M. Joglekar, M. Chen, M. Tintor, M. Meyer, M. Jones, M. Kaufer, M. Schwarzer, M. Shah, M. Yatbaz, M. Y. Guan, M. Xu, M. Yan, M. Glaese, M. Chen, M. Lampe, M. Malek, M. Wang, M. Fradin, M. McClay, M. Pavlov, M. Wang, M. Wang, M. Murati, M. Bavarian, M. Rohaninejad, N. McAleese, N. Chowdhury, N. Chowdhury, N. Ryder, N. Tezak, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, P. Chao, P. Ashbourne, P. Izmailov, P. Zhokhov, R. Dias, R. Arora, R. Lin, R. G. Lopes, R. Gaon, R. Miyara, R. Leike, R. Hwang, R. Garg, R. Brown, R. James, R. Shu, R. Cheu, R. Greene, S. Jain, S. Altman, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Hernandez, S. Baker, S. McKinney, S. Yan, S. Zhao, S. Hu, S. Santurkar, S. R. Chaudhuri, S. Zhang, S. Fu, S. Papay, S. Lin, S. Balaji, S. Sanjeev, S. Sidor, T. Broda, A. Clark, T. Wang, T. Gordon, T. Sanders, T. Patwardhan, T. Sottiaux, T. Degry, T. Dimson, T. Zheng, T. Garipov, T. Stasi, T. Bansal, T. Creech, T. Peterson, T. Eloundou, V. Qi, V. Kosaraju, V. Monaco, V. Pong, V. Fomenko, W. Zheng, W. Zhou, W. McCabe, W. Zaremba, Y. Dubois, Y. Lu, Y. Chen, Y. Cha, Y. Bai, Y. He, Y. Zhang, Y. Wang, Z. Shao, and Z. Li. Openai o1 system card, 2024.
  14. 14.J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms, 2017.
  15. 15.Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo. Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024.
  16. 16.A. Talmor, J. Herzig, N. Lourie, and J. Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge, 2019.
  17. 17.A. Talmor, J. Herzig, N. Lourie, and J. Berant. CommonsenseQA: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics.
  18. 18.X. Wang, L. Caccia, O. Ostapenko, X. Yuan, W. Y. Wang, and A. Sordoni. Guiding language model reasoning with planning tokens, 2024.
  19. 19.Z. Wang, G. Cuenca, S. Zhou, F. F. Xu, and G. Neubig. Mconala: a benchmark for code generation from multiple natural languages. arXiv preprint arXiv:2203.08388, 2022.
  20. 20.J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models, 2023.
  21. 21.Wikipedia contributors. Four stages of competence. https://en.wikipedia.org/wiki/Four_stages_of_competence. Accessed: 2025-07-29.
  22. 22.T. Wilson and J. Schooler. Thinking too much: Introspection can reduce the quality of preferences and decisions. Journal of personality and social psychology, 60:181–92, 03 1991.
  23. 23.B. Workshop, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100, 2022.
  24. 24.F. Xu, Q. Hao, Z. Zong, J. Wang, Y. Zhang, J. Wang, X. Lan, J. Gong, T. Ouyang, F. Meng, C. Shao, Y. Yan, Q. Yang, Y. Song, S. Ren, X. Hu, Y. Li, J. Feng, C. Gao, and Y. Li. Towards large reasoning models: A survey of reinforced reasoning with large language models, 2025.
  25. 25.N. Yax, H. Anlló, and S. Palminteri. Studying and improving reasoning in humans and machines. Communications Psychology, 2:51, 2024.
  26. 26.J. Zhang, J. Kim, B. O’Donoghue, and S. Boyd. Sample efficient reinforcement learning with reinforce, 2020.

Citation

MLA
Singh, M., et al. “Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning”. arXiv, 2025, http://arxiv.org/abs/2511.21743v1.
APA
Singh, M., Singha, A., Radhakrishna, A., & Gulwani, S. (2025). Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning. arXiv. http://arxiv.org/abs/2511.21743v1
Chicago
Singh, M., A. Singha, A. Radhakrishna, and S. Gulwani. 2025. “Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning”. arXiv. http://arxiv.org/abs/2511.21743v1.
Harvard
Singh, M. et al. (2025) “Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2511.21743v1.
Vancouver
1. Singh M, Singha A, Radhakrishna A, Gulwani S (2025) Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning. arXiv

BibTeX

@article{singh2025scaling,
  title = {Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning},
  author = {Singh, Mukul and Singha, Ananya and Radhakrishna, Arjun and Gulwani, Sumit},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2511.21743v1},
  eprint = {2511.21743}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/