Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task
Alicia CurthRachel LawrenceSushrut KarmalkarNiranjani Prasad
Demonstrates that transformers adaptively allocate layer depth to match reasoning difficulty in multi-hop tasks, using logit lens and causal patching to reveal how computational depth scales with relationship chain length.
Recent advances in artificial intelligence rely heavily on stacking deeper network layers, yet recent research questions whether large language models use their computational depth efficiently or if later layers perform redundant work. Understanding whether deep models adapt their internal depth to problem difficulty is critical for optimizing model architectures, reducing inference costs, and improving reasoning capabilities.
The article investigates whether transformer models use their computational depth adaptively by allocating more layers and processing stages to harder tasks. Specifically, the authors evaluate how internal predictions and information integration evolve across layers as reasoning difficulty increases.
To conduct this evaluation, the researchers used a controlled multi-hop relational reasoning benchmark based on family stories, where task difficulty was strictly controlled by the number of relationship hops required (ranging from 2 to 10 hops, and up to 15 hops for testing length limits). The task was framed as predicting a single target relation token at the end of the story. The study examined five open-weight model families (GPT-2, Pythia, Phi, Qwen, and LLaMA) with parameter counts ranging from 117M to 14B. The authors analyzed internal hidden states using two primary methods: logit lens, which decodes predictions at intermediate layers using the model's output head, and causal patching, an intervention technique that tracks when and where token information flows through the network layers.
The investigation produced several key findings. First, across pretrained models, semantic family relations become decodable roughly two-thirds of the way through the network, after which final layers appear to diffuse and recalibrate output probabilities. Second, larger pretrained models (such as Phi-4, Qwen-7B, and LLaMA-3.1-8B) show clear evidence of adaptive depth use, identifying answers several layers earlier for simpler tasks compared to harder tasks, whereas smaller models do not exhibit this flexibility. Third, causal patching shows that longer relationship chains trigger earlier cross-token information mixing across intermediate layers. Fourth, finetuning models specifically on the reasoning task yields distinct behaviors depending on the training regime: full-model finetuning shows the strongest adaptive depth use and earlier internal decodability, but catastrophic forgetting ruins general language abilities (perplexity over 1000 on standard language benchmarks) and degrades performance on longer, unseen chains. In contrast, parameter-efficient LoRA finetuning preserves general language skills and generalizes well to 11–15 hops, but does not decode earlier on simpler tasks.
These findings demonstrate that deeper networks possess the capacity to reserve depth for harder reasoning steps, challenging earlier assertions that later layers are largely unneeded. However, maintaining broad general-purpose language capability creates a natural trade-off that limits how freely a network reorganizes its early layers for specialized multi-hop computations. For systems engineered to perform complex, deterministic reasoning, specialized full finetuning enables deep iterative processing, whereas general-purpose assistant models must balance this capacity against baseline capabilities.
Organizations developing reasoning models should consider adaptive inference architectures and parameter-efficient tuning like LoRA when general language competence and length generalization are required. For domain-specific pipelines with fixed problem lengths where peak reasoning accuracy is paramount, full finetuning can unlock earlier and more iterative depth utilization. Further work should explore whether dynamic inference approaches—such as recurrent architectures or early-exit mechanisms—can systematically exploit this adaptive depth behavior without sacrificing generalization.
- Paper: Do Large Language Models Latently Perform Multi-Hop Reasoning?, Sohee Yang et al. (2024). It provides foundational methodology and empirical evidence on how transformers execute latent multi-hop relational steps across internal layers via causal patching and logit lens projections.
- Paper: Dissecting Recall of Factual Associations in Auto-Regressive Language Models, Mor Geva et al. (2023). It details how factual associations and multi-step relations are retrieved across transformer layers using representation patching and vocabulary projections.
- Paper: Mixture-of-Depths: Dynamically allocating compute in transformer-based language models, David Raposo et al. (2024). It introduces dynamic allocation of transformer depth and compute across tokens, establishing the conceptual baseline for evaluating whether standard transformers use depth adaptively.
- Paper: Layer by Layer: Uncovering Hidden Representations in Language Models, Oscar Skean et al. (2025). It examines how representations evolve layer-by-layer across transformer depth, providing essential context on intermediate layer utility and internal readout dynamics.
- Book: Scaling Latent Reasoning via Looped Language Models, Rui-Jie Zhu et al. (2025). It demonstrates how iterative latent computation over multi-hop relational questions scales with input difficulty in looped transformer architectures.
- Paper: Exploring Length Generalization in Large Language Models, Cem Anil et al. (2022). It establishes key benchmarks and behavioral limits of transformer depth and reasoning when handling tasks with increasing chain lengths.
- Paper: Quantifying Attention Flow in Transformers, Samira Abnar et al. (2020). It establishes methods for tracking and quantifying information flow across transformer layers through residual connections and attention paths.
- Paper: Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers, Sajad Movahedi et al. (2026). It extends the study of adaptive depth by designing looped transformer architectures that dynamically adjust effective reasoning depth according to problem difficulty.
- Paper: DeepLoop: Depth Scaling for Looped Transformers, Shuzhen Li et al. (2026). It investigates architectural scaling rules for stabilizing deeper recurrent passes in looped transformers when solving multi-step reasoning tasks.
- Paper: Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation, Amr Hegazy et al. (2026). It explores recurrent modulation mechanisms that enable transformers to vary expressive depth dynamically and perform early exits during inference.
- Paper: Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning, Benhao Huang et al. (2026). It formalizes adaptive test-time computation and iterative depth scaling in latent dynamical systems for multi-step reasoning.
- Paper: The Topological Trouble With Transformers, Michael C. Mozer et al. (2026). It analyzes the fundamental topological limits of feedforward transformer depth in maintaining multi-step state tracking and evaluates recurrent alternatives.
