From Growing to Looping: A Unified View of Iterative Computation in LLMs

Ferdinand KaplEmmanouil AngelisKaitlin MaileJohannes von OswaldStefan Bauer

article2026arXiv6 citations

Establishes a mechanistic link between layer looping and depth growth in language models, showing that applying inference-time looping to depth-grown architectures can double reasoning accuracy without loop-specific training.

Listen

Scaling large language models typically requires vast compute budgets and larger parameter counts, yet multi-step reasoning capabilities do not always scale cleanly with size alone. Generating longer text outputs is a common workaround, but it increases inference costs through token generation rather than improving the model's internal processing depth. Two architectural alternatives have emerged to address this: depth growing (training shallow networks and progressively duplicating middle layers to cut training compute) and looping (recurrently reusing a block of layers with tied weights to save parameters). While both approaches boost multi-step reasoning, it remained unclear whether their advantages stem from the same underlying computational behavior.

The article demonstrates that looped and depth-grown models share a common computational mechanism: iterative refinement across depth. By conducting extensive experiments across 22 benchmarks on 360-million and 1.7-billion parameter language models, the authors evaluate how these architectures trade off unique parameters, training compute, and inference compute, while tracking their internal layer dynamics.

The analysis reveals five primary findings. First, depth growing achieves equal or superior reasoning accuracy compared to standard baseline models while requiring approximately 20% less pre-training compute. Second, looped models significantly improve reasoning when unique parameter counts are constrained, and partially looped models with unique outer layers remain competitive with full-sized baselines under matched inference budgets. Third, mechanistic tests show both architectures share identical depth-usage patterns: they rely heavily on late layers, show slower residual growth, and develop four-layer periodic update cycles that counteract standard transformer degradation. Fourth, depth-grown models can be executed in loops during inference without any loop-specific training, boosting reasoning benchmark accuracy by up to 2×. Finally, depth-grown models adapt substantially better during supervised fine-tuning, in-context learning, and cooldown training on high-quality mathematical datasets.

These findings provide immediate practical implications for artificial intelligence development and resource allocation. Organizations can lower pre-training costs by adopting depth-growth schedules without sacrificing general language capabilities. At inference time, practitioners can scale reasoning dynamically by looping middle blocks without retraining complete networks. Furthermore, maintaining unique initial and final layers around a recurrent middle block eliminates the brittleness associated with fully recurrent architectures, preserving robust performance across tasks.

For engineering teams seeking to optimize reasoning models under computational constraints, the article supports a clear strategy: train models using depth-growth methods first, enrich them with high-quality math data during training cooldowns, and loop the middle layers during inference when deeper reasoning is required. To establish broader confidence across production environments, future investigations should test these architectures on larger model scales beyond 1.7 billion parameters and explore hyperparameter tuning tailored specifically to grown and looped training schedules.

arXiv: 2602.16490
  • Book: Scaling Latent Reasoning via Looped Language Models, Rui-Jie Zhu et al. (2025). Its large-scale LoopLM experiments establish how shared Transformer blocks can perform latent iterative computation, the looping mechanism the source compares with depth growth.
  • Paper: Hierarchical Reasoning Model, Guan Wang et al. (2025). HRM provides a concrete recurrent reasoning architecture, making its iterative computation framework useful for understanding the source’s analysis of looped models.
  • Book: Less is More: Recursive Reasoning with Tiny Networks, Alexia Jolicoeur-Martineau (2025). TRM develops recursive reuse of network computation for reasoning, offering context for the source’s account of how repeated depth can support stronger reasoning.

No sufficiently relevant recommendations were found.

Cover for From Growing to Looping: A Unified View of Iterative Computation in LLMs

Abstract

Looping, reusing a block of layers across depth, and depth growing, training shallow-to-deep models by duplicating middle layers, have both been linked to stronger reasoning, but their relationship remains unclear. We provide a mechanistic unification: looped and depth-grown models exhibit convergent depth-wise signatures, including increased reliance on late layers and recurring patterns aligned with the looped or grown block. These shared signatures support the view that their gains stem from a common form of iterative computation. Building on this connection, we show that the two techniques are adaptable and composable: applying inference-time looping to the middle blocks of a depth-grown model improves accuracy on some reasoning primitives by up to 2×2\times, despite the model never being trained to loop. Both approaches also adapt better than the baseline when given more in-context examples or additional supervised fine-tuning data. Additionally, depth-grown models achieve the largest reasoning gains when using higher-quality, math-heavy cooldown mixtures, which can be further boosted by adapting a middle block to loop. Overall, our results position depth growth and looping as complementary, practical methods for inducing and scaling iterative computation to improve reasoning.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 The Inductive Bias of Looped and Depth-Grown Models
  • 3.1 Notation of Looped and Depth-Grown Models
  • 3.2 Trade-offs of Looped and Depth-Grown Models
  • 4 The relationship of Looping and Growing
  • 4.1 Mechanistic Analysis
  • 4.2 Inference Scaling
  • 5 Adaptability of Looped and Grown Models
  • 5.1 In-Context Learning & Supervised Fine-Tuning
  • 5.2 High-Quality Cooldown Mixtures
  • 5.3 Retrofitted Recurrence
  • 6 Conclusion
  • References
  • A Details of Growing
  • B Additional Mechanistic Analysis
  • B.1 Swapping Interventions
  • B.2 Repeat Block experiments
  • B.3 Sublayer Usage Experiments
  • C Additional Adaptability and Inference Scaling Results
  • C.1 Supervised Fine Tuning
  • C.2 In-Context Learning
  • C.3 Inference Scaling
  • D Tasks and Benchmarks Overview
  • D.1 Reasoning Primitives
  • E Experimental Protocol for Depth and Mechanistic Analysis

Knowls

  1. Knowl 1 — Looping and depth growth create repeated computation in different ways

    model/method

    The study compares standard transformers, looped transformers, and depth-grown transformers while holding architecture class (including width and attention-head count) fixed. A standard model has untied parameters at every layer. A Loop(L×kL\times k) model has LL unique layers, with the same sequence of layers applied kk times for effective depth LkLk; a partially looped Loop(ee-L×kL\times k-dd) model also has unique encoder and decoder layers around the repeated middle block. Depth-grown models instead keep layers untied at training and inference: they are built by progressively inserting duplicated layers at mid-depth. The experiments use block size B=4B=4: MIDAS duplicates a middle block, while LIDAS duplicates the layer-wise middle. At each growth step, the copied layers begin with copied weights and AdamW optimizer state and then continue training. For final depth LfinalL_{\mathrm{final}}, the number of growth stages is k=Lfinal/Bk=L_{\mathrm{final}}/B; with total training steps TT, stage ii receives Ti=iα∑j=1kjαTT_i=\frac{i^\alpha}{\sum_{j=1}^{k}j^\alpha}T steps. The experiments use PROP-1 (α=1\alpha=1), T=170,000T=170{,}000 steps, and a continuous learning-rate schedule across stages. Growing changes the training path, not the final untied architecture.

  2. Knowl 2 — Looping and growth offer different compute–reasoning trade-offs

    empirical result

    The authors compare SmolLM-v1-based models at 360M and 1.7B parameters, trained on approximately 200B and 400B tokens, respectively, across 22 benchmarks covering language modeling, knowledge, math word problems, and reasoning primitives. On reasoning primitives, the 360M full-depth baseline scores 30.04% accuracy, MIDAS 28.42%, and LIDAS 31.58%; a 16-unique-layer standard model scores 21.62%, while Loop(16×2), with the same 32-layer effective depth, scores 31.26%. At 1.7B, the full-depth baseline scores 34.62%, MIDAS 40.88%, and LIDAS 48.02%; a 12-layer standard model scores 35.42%, compared with 41.22% for Loop(12×2) and 39.88% for Loop(4-4×4-4). Across the evaluated variants, looping generally improves reasoning over standard models with the same number of unique layers, especially on reasoning-heavy categories. With matched effective depth, looped models often trail the full baseline on knowledge and language measures, but can exceed it on reasoning primitives when they retain roughly half the baseline's unique layers. Depth-grown models match broad baseline performance across most categories while using about 80% of the pre-training compute and improving reasoning. These comparisons use the study's training setup; the authors note that they did not tune baseline training hyperparameters specifically for looping or growth.

  3. Knowl 3 — Looped and grown models rely more on late layers

    empirical result

    In a mechanistic comparison of 1.7B models at matched effective depth, the authors examine a baseline, LIDAS, Loop(4×6), and Loop(4-4×4-4). Three depth-utilization diagnostics indicate stronger late-layer use in LIDAS and both looped models than in the baseline: their depth scores are higher, their later-layer top-five vocabulary predictions overlap less with the final-layer predictions, and their Tuned Lens early-exit accuracy on Variable Assignment Math continues improving later into the network. The baseline's early-exit accuracy plateaus sooner. Together, these measures indicate that later layers make more indispensable changes to outputs in the grown and looped models; the results are diagnostic evidence of a shared depth-wise pattern, not a proof that the models implement identical computations.

  4. Knowl 4 — Residual and attention activity recurs with the four-layer block

    empirical result

    For the 1.7B baseline, LIDAS, Loop(4×6), and Loop(4-4×4-4), the grown and looped models show slower growth in residual-stream L2 norm than the baseline. Their relative attention-sublayer contributions also display a recurring four-layer pattern, aligned with the four-layer LIDAS growth block and the four-layer recurrent blocks. In a local residual-stream intervention, the contribution from one earlier layer is removed from a later layer's input without propagating that change through the rest of the network. The resulting future-layer effects show, in LIDAS and Loop(4-4×4-4), a layer about every four layers that is sensitive to many preceding layers, consistent with an aggregation-like role. The shared periodicity and norm patterns support the authors' account that looping and growth encourage repeated computation across depth.

  5. Knowl 5 — Unique encoder and decoder layers restore robustness to layer swaps

    empirical result

    The authors test intervention robustness by swapping individual layers, and also consecutive two-layer blocks, in the middle of a 1.7B model and measuring performance on Lambada and Variable Assignment Math. A single-layer swap causes substantially greater degradation in the fully looped Loop(4×6) model than in the baseline, LIDAS, or Loop(4-4×4-4); the two-layer-swap experiment shows the same general pattern. Loop(4-4×4-4) has four unique encoder layers, a four-layer block repeated four times, and four unique decoder layers. Its robustness is similar to the baseline and LIDAS, unlike the fully looped model. The authors hypothesize that middle-layer growth permits flexible, unique encoding and decoding around a repeated middle computation, whereas fully tied computation is more order-sensitive.

  6. Knowl 6 — Inference-time repetition can raise reasoning accuracy in untrained-to-loop grown models

    empirical result

    Without further training, the authors insert additional inference-time repetitions of contiguous four-layer blocks at different positions in MIDAS and LIDAS, whose growth structure uses four-layer blocks. On the Copy Real Words reasoning primitive, repeating a middle block raises accuracy by as much as 2× relative to the original network; the baseline rarely benefits comparably. Additional reasoning-primitive results show the same general advantage for grown models, although the size of the gain varies by task. One extra repetition often gives the largest improvement, two repetitions often give the best accuracy, and further repetitions usually add no benefit and can reduce accuracy. Thus, depth-grown models can use extra latent computation at inference even though they were not trained with tied weights.

  7. Knowl 7 — Looped and grown models make better use of in-context examples

    empirical result

    The authors vary the number of in-context examples for reasoning-primitive tasks up to the context-length limit, comparing a baseline with LIDAS, Loop(4×6), and Loop(4-4×4-4). The looped and grown models generally improve more as examples are added, whereas the baseline often changes little. With enough examples, Loop(4-4×4-4) sometimes exceeds LIDAS. The trend is not universal: some reasoning primitives do not improve with additional examples. The result therefore supports greater in-context adaptability on the tested tasks, rather than a guarantee that more examples help every task.

  8. Knowl 8 — Looped and grown models learn reasoning tasks from fewer fine-tuning examples

    empirical result

    For supervised fine-tuning, the authors use Variable Assignment Code at depths d=1d=1 and d=2d=2, corresponding to one and two assignment hops. Fine-tuning sets contain 64, 128, 256, or 512 examples, with equal numbers from the two depth variants, and plotted results summarize three random seeds. At 64 examples, LIDAS and Loop(4-4×4-4) are already above chance, while the baseline remains near random even with 128 examples. As the data set grows, the looped and grown models outperform the baseline. At larger data-set sizes on depth 2, Loop(4×6) matches LIDAS and Loop(4-4×4-4) reaches higher accuracy. Additional experiments find that removing weight tying during fine-tuning does not improve the looped variants: the untied fine-tuning variants underperform their tied counterparts.

  9. Knowl 9 — A math-heavy Nemotron cooldown benefits grown models most

    empirical result

    The authors increase math data to 20% during the final 30,000 training steps (15% of pre-training), comparing FineMath-4+ (FMT) with Nemotron-CC-Math-4+ (NMT); the original mixture has 6% OpenWebMath. The following accuracies are ordered as Math Word Problems, Reasoning Primitives, and GSM8K. For the 360M baseline, the original scores are 3.69%, 30.04%, and 1.06%; after FMT cooldown they are 13.46%, 30.10%, and 2.12%, and after NMT they are 21.15%, 35.70%, and 3.41%. MIDAS scores 4.39%, 28.42%, and 1.44% originally; 14.68%, 32.84%, and 3.49% after FMT; and 23.70%, 38.42%, and 4.09% after NMT. LIDAS scores 4.36%, 31.58%, and 1.36% originally; 16.54%, 40.12%, and 3.94% after FMT; and 25.64%, 42.20%, and 6.07% after NMT. NMT produces larger gains than FMT on these reported reasoning benchmarks, and LIDAS reaches the strongest post-cooldown results among the three models.

  10. Knowl 10 — Retrofitting a middle block to loop during cooldown further improves reasoning

    empirical result

    At 1.7B parameters, the authors train with a 20% Nemotron-CC-Math-4+ cooldown and, in selected variants, loop one four-layer block for one additional repetition during that cooldown. The reported accuracies are ordered as Math Word Problems, Reasoning Primitives, and GSM8K. The baseline scores 13.61%, 34.62%, and 2.35% before cooldown; NMT cooldown raises these to 32.99%, 39.20%, and 10.61%. Retrofitting layers 8–11 gives 34.75%, 40.90%, and 12.59%; retrofitting layers 12–15 gives 34.02%, 39.84%, and 11.75%. LIDAS scores 18.18%, 48.02%, and 5.61% before cooldown, then 35.54%, 52.60%, and 13.34% after NMT; looping layers 8–11 gives 38.84%, 54.84%, and 17.36%, while looping layers 12–15 gives 38.14%, 56.60%, and 15.47%. Loop(4-4×4-4) scores 6.91%, 39.88%, and 2.65% before cooldown, then 27.00%, 41.94%, and 6.52% after NMT; looping its layers 4–7 gives 27.74%, 43.62%, and 7.35%. The third block (layers 8–11) is the strongest overall choice for the baseline and LIDAS, especially on GSM8K, although LIDAS's layers 12–15 variant scores higher on Reasoning Primitives. After adapting a block to loop, LIDAS also makes the strongest use of further inference-time repetitions in the reported Variable Assignment Basic experiment.

Coverage note — Detailed scores for the other benchmark categories and supplementary intervention, fine-tuning, and inference-scaling plots are omitted because they mainly provide task-by-task corroboration of the findings summarized here.

References

  1. 1.L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blazquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíček, A. P. Lajarín, V. Srivastav, et al. Smollm2: When smol goes big–data-centric training of a small language model. arXiv preprint arXiv:2502.02737, 2025.
  2. 2.S. Bae, A. Fisch, H. Harutyunyan, Z. Ji, S. Kim, and T. Schuster. Relaxed recursive transformers: Effective parameter sharing with layer-wise loRA. In The Thirteenth International Conference on Learning Representations, 2025.
  3. 3.S. Bae, Y. Kim, R. Bayat, S. Kim, J. Ha, T. Schuster, A. Fisch, H. Harutyunyan, Z. Ji, A. Courville, and S.-Y. Yun. Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
  4. 4.N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023.
  5. 5.L. Ben Allal, A. Lozhkov, G. Penedo, T. Wolf, and L. von Werra. SmolLM-corpus, July 2024. URL https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.
  6. 6.C. Blakeney, M. Paul, B. W. Larsen, S. Owen, and J. Frankle. Does your data spark joy? performance gains from domain upsampling at the end of training. In First Conference on Language Modeling, 2024.
  7. 7.K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  8. 8.R. Csordas, K. Irie, J. Schmidhuber, C. Potts, and C. D. Manning. MOEUT: Mixture-of-experts universal transformers. Advances in Neural Information Processing Systems, 37:28589–28614, 2024.
  9. 9.R. Csordas, C. D. Manning, and C. Potts. Do language models use their depth efficiently? In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
  10. 10.R. Dabre and A. Fujita. Recurrent stacking of layers for compact neural machine translation models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6292–6299, 2019.
  11. 11.M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser. Universal transformers. In International Conference on Learning Representations, 2019.
  12. 12.W. Du, T. Luo, Z. Qiu, Z. Huang, Y. Shen, R. Cheng, Y. Guo, and J. Fu. Stacking your transformers: A closer look at model growth for efficient llm pre-training. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 10491–10540. Curran Associates, Inc., 2024. doi: 10.52202/079017-0336.
  13. 13.Y. Fan, Y. Du, K. Ramchandran, and K. Lee. Looped transformers for length generalization. In The Thirteenth International Conference on Learning Representations, 2025.
  14. 14.L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou. The language model evaluation harness, 07 2024. URL https://zenodo.org/records/12608602.
  15. 15.J. Geiping, S. M. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein. Scaling up test-time compute with latent reasoning: A recurrent depth approach. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
  16. 16.L. Gong, D. He, Z. Li, T. Qin, L. Wang, and T. Liu. Efficient training of BERT by progressively stacking. In International conference on machine learning, pages 2337–2346. PMLR, 2019.
  17. 17.J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, M. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, J. Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  18. 18.A. Jolicoeur-Martineau. Less is more: Recursive reasoning with tiny networks. arXiv preprint arXiv:2510.04871, 2025.
  19. 19.Julich Supercomputing Centre. JUWELS Cluster and Booster: Exascale Pathfinder with Modular Supercomputing Architecture at Juelich Supercomputing Centre. Journal of large-scale research facilities, 7(A138), 2021. doi: 10.17815/jlsrf-7-183.
  20. 20.F. Kapl, E. Angelis, T. Hoppe, K. Maile, J. von Oswald, N. Scherrer, and S. Bauer. Do depth-grown models overcome the curse of depth? an in-depth analysis. arXiv preprint arXiv:2512.08819, 2025.
  21. 21.J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  22. 22.S. Kim, D. Kim, C. Park, W. Lee, W. Song, Y. Kim, H. Kim, Y. Kim, H. Lee, J. Kim, C. Ahn, S. Yang, S. Lee, H. Park, G. Gim, M. Cha, H. Lee, and S. Kim. SOLAR 10.7B: Scaling large language models with simple yet effective depth up-scaling. In Y. Yang, A. Davani, A. Sil, and A. Kumar, editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track), pages 23–35, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.naacl-industry.3.
  23. 23.Y. Koishekenov, A. Lipani, and N. Cancedda. Encode, think, decode: Scaling test-time reasoning with recursive latent thoughts. arXiv preprint arXiv:2510.07358, 2025.
  24. 24.R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi. MAWPS: A math word problem repository. In Proceedings of the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 1152–1157, 2016.
  25. 25.Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations, 2020.
  26. 26.R. K. Mahabadi, S. Satheesh, S. Prabhumoye, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-cc-math: A 133 billion-token-scale high quality math pretraining dataset. arXiv preprint arXiv:2508.15096, 2025.
  27. 27.S. McLeish, A. Li, J. Kirchenbauer, D. S. Kalra, B. R. Bartoldson, B. Kailkhura, A. Schwarzschild, J. Geiping, T. Goldstein, and M. Goldblum. Teaching pretrained language models to think deeper with retrofitted recurrence. arXiv preprint arXiv:2511.07384, 2025.
  28. 28.S.-y. Miao, C.-C. Liang, and K.-Y. Su. A diverse corpus for evaluating and developing English math word problem solvers. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 975–984, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.92.
  29. 29.D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernandez. The LAMBADA dataset: Word prediction requiring a broad discourse context. In K. Erk and N. A. Smith, editors, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1525–1534, Berlin, Germany, Aug. 2016. Association for Computational Linguistics. doi: 10.18653/v1/P16-1144.
  30. 30.A. Patel, S. Bhattamishra, and N. Goyal. Are NLP models really able to solve simple math word problems? In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.168.
  31. 31.S. J. Reddi, S. Miryoosefi, S. Karp, S. Krishnan, S. Kale, S. Kim, and S. Kumar. Efficient training of language models using few-shot learning. In International Conference on Machine Learning, pages 14553–14568. PMLR, 2023.
  32. 32.N. Saunshi, S. Karp, S. Krishnan, S. Miryoosefi, S. Jakkam Reddi, and S. Kumar. On the inductive bias of stacking towards improving reasoning. Advances in Neural Information Processing Systems, 37:71437–71464, 2024.
  33. 33.N. Saunshi, N. Dikkala, Z. Li, S. Kumar, and S. J. Reddi. Reasoning with latent thoughts: On the power of looped transformers. In The Thirteenth International Conference on Learning Representations, 2025.
  34. 34.W. Sun, X. Song, P. Li, L. Yin, Y. Zheng, and S. Liu. The curse of depth in large language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
  35. 35.S. Takase and S. Kiyono. Lessons on parameter sharing across layers in transformers. In Proceedings of The Fourth Workshop on Simple and Efficient Natural Language Processing (SustaiNLP), pages 78–90, 2023.
  36. 36.E. P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jordan, N. Lambert, D. Schwenk, O. Tafjord, T. Anderson, D. Atkinson, F. Brahman, C. Clark, P. Dasigi, N. Dziri, A. Ettinger, M. Guerquin, D. Heineman, H. Ivison, P. W. Koh, J. Liu, S. Malik, W. Merrill, L. J. V. Miranda, J. Morrison, T. Murray, C. Nam, J. Poznanski, V. Pyatkin, A. Rangapur, M. Schmitz, S. Skjonsberg, D. Wadden, C. Wilhelm, M. Wilson, L. Zettlemoyer, A. Farhadi, N. A. Smith, and H. Hajishirzi. 2 OLMo 2 furious (COLM’s version). In Second Conference on Language Modeling, 2025.
  37. 37.G. Wang, J. Li, Y. Sun, X. Chen, C. Liu, Y. Wu, M. Lu, S. Song, and Y. A. Yadkori. Hierarchical reasoning model. arXiv preprint arXiv:2506.21734, 2025.
  38. 38.P. Wang, R. Panda, L. T. Hennigen, P. Greengard, L. Karlinsky, R. Feris, D. D. Cox, Z. Wang, and Y. Kim. Learning to grow pretrained models for efficient transformer training. In The Eleventh International Conference on Learning Representations, 2023.
  39. 39.J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. Survey Certification.
  40. 40.J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 24824–24837. Curran Associates, Inc., 2022.
  41. 41.F. Xu, Q. Hao, Z. Zong, J. Wang, Y. Zhang, J. Wang, X. Lan, J. Gong, T. Ouyang, F. Meng, et al. Towards large reasoning models: A survey of reinforced reasoning with large language models. arXiv preprint arXiv:2501.09686, 2025.
  42. 42.L. Yang, K. Lee, R. D. Nowak, and D. Papailiopoulos. Looped transformers are better at learning learning algorithms. In The Twelfth International Conference on Learning Representations, 2024.
  43. 43.Y. Yao, Z. Zhang, J. Li, and Y. Wang. Masked structural growth for 2x faster language model pre-training. In The Twelfth International Conference on Learning Representations, 2024.
  44. 44.R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. HellaSwag: Can a machine really finish your sentence? In A. Korhonen, D. Traum, and L. Marquez, editors, ` Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/v1/P19-1472.
  45. 45.R. Zhu, H. Zhang, T. Shi, C. Wang, T. Zhou, and Z. Qin. The 4th dimension for scaling model size. arXiv preprint arXiv:2506.18233, 2025.
  46. 46.R.-J. Zhu, Z. Wang, K. Hua, T. Zhang, Z. Li, H. Que, B. Wei, Z. Wen, F. Yin, H. Xing, et al. Scaling latent reasoning via looped language models. arXiv preprint arXiv:2510.25741, 2025.

Citation

MLA
Kapl, F., et al. “From Growing to Looping: A Unified View of Iterative Computation in LLMs”. arXiv, 2026, http://arxiv.org/abs/2602.16490v1.
APA
Kapl, F., Angelis, E., Maile, K., Oswald, J. von ., & Bauer, S. (2026). From Growing to Looping: A Unified View of Iterative Computation in LLMs. arXiv. http://arxiv.org/abs/2602.16490v1
Chicago
Kapl, F., E. Angelis, K. Maile, J. von . Oswald, and S. Bauer. 2026. “From Growing to Looping: A Unified View of Iterative Computation in LLMs”. arXiv. http://arxiv.org/abs/2602.16490v1.
Harvard
Kapl, F. et al. (2026) “From Growing to Looping: A Unified View of Iterative Computation in LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2602.16490v1.
Vancouver
1. Kapl F, Angelis E, Maile K, Oswald J von, Bauer S (2026) From Growing to Looping: A Unified View of Iterative Computation in LLMs. arXiv

BibTeX

@article{kapl2026from,
  title = {From Growing to Looping: A Unified View of Iterative Computation in LLMs},
  author = {Kapl, Ferdinand and Angelis, Emmanouil and Maile, Kaitlin and Oswald, Johannes von and Bauer, Stefan},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2602.16490v1},
  eprint = {2602.16490}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/