Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task

Alicia CurthRachel LawrenceSushrut KarmalkarNiranjani Prasad

article2026arXiv0 citations

Demonstrates that transformers adaptively allocate layer depth to match reasoning difficulty in multi-hop tasks, using logit lens and causal patching to reveal how computational depth scales with relationship chain length.

Listen

Recent advances in artificial intelligence rely heavily on stacking deeper network layers, yet recent research questions whether large language models use their computational depth efficiently or if later layers perform redundant work. Understanding whether deep models adapt their internal depth to problem difficulty is critical for optimizing model architectures, reducing inference costs, and improving reasoning capabilities.

The article investigates whether transformer models use their computational depth adaptively by allocating more layers and processing stages to harder tasks. Specifically, the authors evaluate how internal predictions and information integration evolve across layers as reasoning difficulty increases.

To conduct this evaluation, the researchers used a controlled multi-hop relational reasoning benchmark based on family stories, where task difficulty was strictly controlled by the number of relationship hops required (ranging from 2 to 10 hops, and up to 15 hops for testing length limits). The task was framed as predicting a single target relation token at the end of the story. The study examined five open-weight model families (GPT-2, Pythia, Phi, Qwen, and LLaMA) with parameter counts ranging from 117M to 14B. The authors analyzed internal hidden states using two primary methods: logit lens, which decodes predictions at intermediate layers using the model's output head, and causal patching, an intervention technique that tracks when and where token information flows through the network layers.

The investigation produced several key findings. First, across pretrained models, semantic family relations become decodable roughly two-thirds of the way through the network, after which final layers appear to diffuse and recalibrate output probabilities. Second, larger pretrained models (such as Phi-4, Qwen-7B, and LLaMA-3.1-8B) show clear evidence of adaptive depth use, identifying answers several layers earlier for simpler tasks compared to harder tasks, whereas smaller models do not exhibit this flexibility. Third, causal patching shows that longer relationship chains trigger earlier cross-token information mixing across intermediate layers. Fourth, finetuning models specifically on the reasoning task yields distinct behaviors depending on the training regime: full-model finetuning shows the strongest adaptive depth use and earlier internal decodability, but catastrophic forgetting ruins general language abilities (perplexity over 1000 on standard language benchmarks) and degrades performance on longer, unseen chains. In contrast, parameter-efficient LoRA finetuning preserves general language skills and generalizes well to 11–15 hops, but does not decode earlier on simpler tasks.

These findings demonstrate that deeper networks possess the capacity to reserve depth for harder reasoning steps, challenging earlier assertions that later layers are largely unneeded. However, maintaining broad general-purpose language capability creates a natural trade-off that limits how freely a network reorganizes its early layers for specialized multi-hop computations. For systems engineered to perform complex, deterministic reasoning, specialized full finetuning enables deep iterative processing, whereas general-purpose assistant models must balance this capacity against baseline capabilities.

Organizations developing reasoning models should consider adaptive inference architectures and parameter-efficient tuning like LoRA when general language competence and length generalization are required. For domain-specific pipelines with fixed problem lengths where peak reasoning accuracy is paramount, full finetuning can unlock earlier and more iterative depth utilization. Further work should explore whether dynamic inference approaches—such as recurrent architectures or early-exit mechanisms—can systematically exploit this adaptive depth behavior without sacrificing generalization.

arXiv: 2604.12426
Cover for Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task

Abstract

We investigate whether transformers use their depth adaptively across tasks of increasing difficulty. Using a controlled multi-hop relational reasoning task based on family stories, where difficulty is determined by the number of relationship hops that must be composed, we monitor (i) how predictions evolve across layers via early readouts (the logit lens) and (ii) how task-relevant information is integrated across tokens via causal patching. For pretrained models, we find some limited evidence for adaptive depth use: some larger models need fewer layers to arrive at plausible answers for easier tasks, and models generally use more layers to integrate information across tokens as chain length increases. For models finetuned on the task, we find clearer and more consistent evidence of adaptive depth use, with the effect being stronger for less constrained finetuning regimes that do not preserve general language modeling abilities.

Citation

MLA
Curth, A., et al. “Do Transformers Use Their Depth Adaptively? Evidence from a Relational Reasoning Task”. arXiv, 2026, http://arxiv.org/abs/2604.12426v1.
APA
Curth, A., Lawrence, R., Karmalkar, S., & Prasad, N. (2026). Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task. arXiv. http://arxiv.org/abs/2604.12426v1
Chicago
Curth, A., R. Lawrence, S. Karmalkar, and N. Prasad. 2026. “Do Transformers Use Their Depth Adaptively? Evidence from a Relational Reasoning Task”. arXiv. http://arxiv.org/abs/2604.12426v1.
Harvard
Curth, A. et al. (2026) “Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2604.12426v1.
Vancouver
1. Curth A, Lawrence R, Karmalkar S, Prasad N (2026) Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task. arXiv

BibTeX

@article{curth2026transformers,
  title = {Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task},
  author = {Curth, Alicia and Lawrence, Rachel and Karmalkar, Sushrut and Prasad, Niranjani},
  year = {2026},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2604.12426v1},
  eprint = {2604.12426}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission