The Impact of Depth on Compositional Generalization in Transformer Language Models
Jackson PettySjoerd van SteenkisteIshita DasguptaFei ShaDan GarretteTal Linzen
Demonstrates that deeper transformer models improve compositional generalization over wider models of equal parameter size but yield rapidly diminishing returns, proving that practitioners can adopt shallower architectures to lower latency without sacrificing performance.
Modern language models must interpret unfamiliar sentences by combining known words and grammatical structures in novel ways, an ability known as compositional generalization. While transformer-based models often struggle with compositional tasks, theory and past experiments suggest that deeper models—those with more layers—perform better. However, prior research routinely confounded model depth with total parameter count, making it unclear whether improvements stemmed from depth itself or simply larger overall model size.
To resolve this question, the article evaluates whether increasing transformer depth directly enhances compositional generalization when total parameter count is held constant. The researchers constructed three size classes of causal decoder-only transformer models (41 million, 134 million, and 374 million parameters) and varied depth by trading layer count against feed-forward width. All models were pretrained on 131 billion tokens from the Colossal Clean Crawled Corpus and then fine-tuned on four compositional benchmark tasks: COGS, variable-free COGS (COGS-vf), GeoQuery, and an English passivization task.
The analysis produced four critical findings. First, while deeper models achieve better pretraining perplexity and stronger compositional generalization, the performance gains diminish rapidly after adding just a few initial layers. Second, across several benchmarks, compositional performance saturates quickly—leveling off at approximately 4 to 6 layers for COGS and 2 to 4 layers for COGS-vf and GeoQuery. Third, the benefits of depth apply almost entirely to simpler lexical substitutions, whereas models of all depths fail to generalize to novel syntactic structures. Fourth, deeper models generalize better even after explicitly controlling for pretraining perplexity and in-distribution fine-tuning accuracy, proving that depth provides a distinct architectural advantage for compositionality.
These results have substantial operational implications for artificial intelligence engineering and computational resource allocation. Because transformer latency scales approximately linearly with the number of layers, deeper models incur significant runtime penalties during both training and inference. The rapid diminishing returns of depth mean that conventional, deep transformer designs spend substantial compute budgets on marginal accuracy gains. Consequently, engineering teams facing a fixed parameter or hardware budget can build shallower, wider architectures that achieve comparable accuracy at significantly reduced latency and operating costs.
The article recommends designing transformers that are shallower than standard configurations when aiming to minimize training time or inference costs. Teams operating under a fixed training time budget can train shallower models on larger volumes of data, potentially outperforming deeper alternatives. Future development should further investigate how pretraining corpus mixtures (such as adding code) affect compositional inductive biases, whether wide-and-shallow configurations scale effectively in few-shot in-context learning environments, and how attention heads interact with reduced layer depth.
Confidence in the findings is high for supervised fine-tuning across the evaluated parameter ranges. However, decision-makers should note that the study evaluated English-only text, parameter scales up to 374 million, and fine-tuning regimes rather than modern large-scale in-context prompting. Model architects should also exercise caution when making models excessively deep and narrow, as performance degrades sharply once the feed-forward dimension drops below the contextual embedding dimension.
- Paper: Evaluating the Impact of Model Scale for Compositional Generalization in Semantic Parsing, Linlu Qiu et al. (2022). This study establishes how scaling overall parameter count fails to resolve compositional generalization bottlenecks in semantic parsing, directly motivating the source paper's need to isolate architectural depth from total model size.
- Paper: Break It Down: Evidence for Structural Compositionality in Neural Networks, Michael A. Lepori et al. (2023). It provides foundational empirical evidence for modular, structural compositionality in neural networks across language and vision tasks, underpinning the evaluation of compositional generalization in transformers.
- Paper: The Paradox of the Compositionality of Natural Language: A Neural Machine Translation Case Study, Verna Dankers et al. (2022). It analyzes the fundamental tensions between local systematicity and global context in transformer compositionality, offering critical conceptual context for benchmark evaluations like COGS.
- Paper: Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models, Hong Liu et al. (2023). It demonstrates that pretraining loss alone does not determine downstream task performance due to implicit architectural biases, directly framing the source paper's methodology of controlling for pretraining perplexity.
- Paper: BERT Rediscovers the Classical NLP Pipeline, Ian Tenney et al. (2019). It reveals the hierarchical layer-by-layer progression of linguistic processing in transformers, explaining why minimum depth thresholds exist for resolving syntactic and semantic structures.
- Paper: Do Deep Nets Really Need to be Deep?, Lei Jimmy Ba et al. (2014). This foundational work examines whether neural networks truly require depth or can be matched by wider, shallower architectures, setting up the depth-versus-width trade-off investigated in the source.
- Paper: Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task, Alicia Curth et al. (2026). It builds upon findings of depth saturation and diminishing returns by investigating whether and how transformers dynamically leverage their internal depth across varying reasoning difficulties.
- Paper: From Growing to Looping: A Unified View of Iterative Computation in LLMs, Ferdinand Kapl et al. (2026). It explores architectural methods like depth growing and recurrent layer looping to achieve the benefits of computational depth without incurring the parameter overhead of deeper networks.
- Paper: Layer by Layer: Uncovering Hidden Representations in Language Models, Oscar Skean et al. (2025). It extends the analysis of layer-wise representation dynamics, demonstrating that downstream performance often saturates at intermediate layers rather than deep final layers.
- Paper: Mixture-of-Depths: Dynamically allocating compute in transformer-based language models, David Raposo et al. (2024). It translates insights about depth efficiency into practice by dynamically allocating transformer layer depth on a per-token basis to optimize compute budgets.
- Paper: Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models, Amartya Roy et al. (2025). It examines how architectural choices beyond depth, specifically encoder versus decoder paradigms, impact multi-hop compositional deduction under distribution shifts.
