From Growing to Looping: A Unified View of Iterative Computation in LLMs
Ferdinand KaplEmmanouil AngelisKaitlin MaileJohannes von OswaldStefan Bauer
Establishes a mechanistic link between layer looping and depth growth in language models, showing that applying inference-time looping to depth-grown architectures can double reasoning accuracy without loop-specific training.
Scaling large language models typically requires vast compute budgets and larger parameter counts, yet multi-step reasoning capabilities do not always scale cleanly with size alone. Generating longer text outputs is a common workaround, but it increases inference costs through token generation rather than improving the model's internal processing depth. Two architectural alternatives have emerged to address this: depth growing (training shallow networks and progressively duplicating middle layers to cut training compute) and looping (recurrently reusing a block of layers with tied weights to save parameters). While both approaches boost multi-step reasoning, it remained unclear whether their advantages stem from the same underlying computational behavior.
The article demonstrates that looped and depth-grown models share a common computational mechanism: iterative refinement across depth. By conducting extensive experiments across 22 benchmarks on 360-million and 1.7-billion parameter language models, the authors evaluate how these architectures trade off unique parameters, training compute, and inference compute, while tracking their internal layer dynamics.
The analysis reveals five primary findings. First, depth growing achieves equal or superior reasoning accuracy compared to standard baseline models while requiring approximately 20% less pre-training compute. Second, looped models significantly improve reasoning when unique parameter counts are constrained, and partially looped models with unique outer layers remain competitive with full-sized baselines under matched inference budgets. Third, mechanistic tests show both architectures share identical depth-usage patterns: they rely heavily on late layers, show slower residual growth, and develop four-layer periodic update cycles that counteract standard transformer degradation. Fourth, depth-grown models can be executed in loops during inference without any loop-specific training, boosting reasoning benchmark accuracy by up to 2×. Finally, depth-grown models adapt substantially better during supervised fine-tuning, in-context learning, and cooldown training on high-quality mathematical datasets.
These findings provide immediate practical implications for artificial intelligence development and resource allocation. Organizations can lower pre-training costs by adopting depth-growth schedules without sacrificing general language capabilities. At inference time, practitioners can scale reasoning dynamically by looping middle blocks without retraining complete networks. Furthermore, maintaining unique initial and final layers around a recurrent middle block eliminates the brittleness associated with fully recurrent architectures, preserving robust performance across tasks.
For engineering teams seeking to optimize reasoning models under computational constraints, the article supports a clear strategy: train models using depth-growth methods first, enrich them with high-quality math data during training cooldowns, and loop the middle layers during inference when deeper reasoning is required. To establish broader confidence across production environments, future investigations should test these architectures on larger model scales beyond 1.7 billion parameters and explore hyperparameter tuning tailored specifically to grown and looped training schedules.
- Book: Scaling Latent Reasoning via Looped Language Models, Rui-Jie Zhu et al. (2025). Its large-scale LoopLM experiments establish how shared Transformer blocks can perform latent iterative computation, the looping mechanism the source compares with depth growth.
- Paper: Hierarchical Reasoning Model, Guan Wang et al. (2025). HRM provides a concrete recurrent reasoning architecture, making its iterative computation framework useful for understanding the source’s analysis of looped models.
- Book: Less is More: Recursive Reasoning with Tiny Networks, Alexia Jolicoeur-Martineau (2025). TRM develops recursive reuse of network computation for reasoning, offering context for the source’s account of how repeated depth can support stronger reasoning.
No sufficiently relevant recommendations were found.
