The Impact of Depth on Compositional Generalization in Transformer Language Models

Jackson PettySjoerd van SteenkisteIshita DasguptaFei ShaDan GarretteTal Linzen

article2024NAACL47 citations

Demonstrates that deeper transformer models improve compositional generalization over wider models of equal parameter size but yield rapidly diminishing returns, proving that practitioners can adopt shallower architectures to lower latency without sacrificing performance.

Listen

Modern language models must interpret unfamiliar sentences by combining known words and grammatical structures in novel ways, an ability known as compositional generalization. While transformer-based models often struggle with compositional tasks, theory and past experiments suggest that deeper models—those with more layers—perform better. However, prior research routinely confounded model depth with total parameter count, making it unclear whether improvements stemmed from depth itself or simply larger overall model size.

To resolve this question, the article evaluates whether increasing transformer depth directly enhances compositional generalization when total parameter count is held constant. The researchers constructed three size classes of causal decoder-only transformer models (41 million, 134 million, and 374 million parameters) and varied depth by trading layer count against feed-forward width. All models were pretrained on 131 billion tokens from the Colossal Clean Crawled Corpus and then fine-tuned on four compositional benchmark tasks: COGS, variable-free COGS (COGS-vf), GeoQuery, and an English passivization task.

The analysis produced four critical findings. First, while deeper models achieve better pretraining perplexity and stronger compositional generalization, the performance gains diminish rapidly after adding just a few initial layers. Second, across several benchmarks, compositional performance saturates quickly—leveling off at approximately 4 to 6 layers for COGS and 2 to 4 layers for COGS-vf and GeoQuery. Third, the benefits of depth apply almost entirely to simpler lexical substitutions, whereas models of all depths fail to generalize to novel syntactic structures. Fourth, deeper models generalize better even after explicitly controlling for pretraining perplexity and in-distribution fine-tuning accuracy, proving that depth provides a distinct architectural advantage for compositionality.

These results have substantial operational implications for artificial intelligence engineering and computational resource allocation. Because transformer latency scales approximately linearly with the number of layers, deeper models incur significant runtime penalties during both training and inference. The rapid diminishing returns of depth mean that conventional, deep transformer designs spend substantial compute budgets on marginal accuracy gains. Consequently, engineering teams facing a fixed parameter or hardware budget can build shallower, wider architectures that achieve comparable accuracy at significantly reduced latency and operating costs.

The article recommends designing transformers that are shallower than standard configurations when aiming to minimize training time or inference costs. Teams operating under a fixed training time budget can train shallower models on larger volumes of data, potentially outperforming deeper alternatives. Future development should further investigate how pretraining corpus mixtures (such as adding code) affect compositional inductive biases, whether wide-and-shallow configurations scale effectively in few-shot in-context learning environments, and how attention heads interact with reduced layer depth.

Confidence in the findings is high for supervised fine-tuning across the evaluated parameter ranges. However, decision-makers should note that the study evaluated English-only text, parameter scales up to 374 million, and fine-tuning regimes rather than modern large-scale in-context prompting. Model architects should also exercise caution when making models excessively deep and narrow, as performance degrades sharply once the feed-forward dimension drops below the contextual embedding dimension.

arXiv: 2310.19956
Cover for The Impact of Depth on Compositional Generalization in Transformer Language Models

Abstract

To process novel sentences, language models (LMs) must generalize compositionally -- combine familiar elements in new ways. What aspects of a model's structure promote compositional generalization? Focusing on transformers, we test the hypothesis, motivated by theoretical and empirical work, that deeper transformers generalize more compositionally. Simply adding layers increases the total number of parameters; to address this confound between depth and size, we construct three classes of models which trade off depth for width such that the total number of parameters is kept constant (41M, 134M and 374M parameters). We pretrain all models as LMs and fine-tune them on tasks that test for compositional generalization. We report three main conclusions: (1) after fine-tuning, deeper models generalize more compositionally than shallower models do, but the benefit of additional layers diminishes rapidly; (2) within each family, deeper models show better language modeling performance, but returns are similarly diminishing; (3) the benefits of depth for compositional generalization cannot be attributed solely to better performance on language modeling. Because model latency is approximately linear in the number of layers, these results lead us to the recommendation that, with a given total parameter budget, transformers can be made shallower than is typical without sacrificing performance.

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Constructing Families of Models with Equal Numbers of Parameters
  • 2.2 Datasets and Training
  • 2.2.1 Language Modeling
  • 2.2.2 Compositional Generalization
  • 3 Results
  • 3.1 Language Modeling
  • 3.2 Compositional Generalization
  • 3.3 Depth Effects are Independent between Upstream and Downstream Tasks
  • 4 Training and Inference Latency
  • 5 Analysis of Feed-Forward Transforms
  • 6 Related Work
  • 7 Conclusion
  • 8 Limitations
  • 9 Ethics Statement
  • References
  • A Design and Result Tables
  • B Annotated Transformer Layer

Knowls

  1. Knowl 1 — Parameter-matched transformer families trade feed-forward width for depth

    model/method

    The study compares causal decoder-only transformer language models whose total parameter counts are held approximately fixed within three size classes: 41M, 134M, and 374M. Within each class, the contextual embedding dimension dmodeld_{\mathrm{model}} and attention dimension dattnd_{\mathrm{attn}} are fixed, while the number of layers nlayersn_{\mathrm{layers}} is varied and the feed-forward dimension dffd_{\mathrm{ff}} is adjusted to compensate. The attention head counts are 8, 8, and 64, respectively; the corresponding dmodel=dattnd_{\mathrm{model}}=d_{\mathrm{attn}} values are 512, 768, and 1024.

    The parameter accounting used for this tradeoff is

    M(dff)=2dmodeldff+4dmodeldattn=βdff+A,N(nlayers,dff)=nlayersM(dff)+2dmodelV,M(d_{\mathrm{ff}})=2d_{\mathrm{model}}d_{\mathrm{ff}}+4d_{\mathrm{model}}d_{\mathrm{attn}}=\beta d_{\mathrm{ff}}+A, \qquad N(n_{\mathrm{layers}},d_{\mathrm{ff}})=n_{\mathrm{layers}}M(d_{\mathrm{ff}})+2d_{\mathrm{model}}V,

    where MM is the parameter count per transformer layer, NN is the model parameter count, VV is the vocabulary size, β=2dmodel\beta=2d_{\mathrm{model}}, and A=4dmodeldattnA=4d_{\mathrm{model}}d_{\mathrm{attn}}. For a baseline with nlayers0n^0_{\mathrm{layers}} layers and feed-forward dimension dff0d^0_{\mathrm{ff}}, adding a signed integer kk layers and reducing the feed-forward dimension by w(k)w(k) preserves the baseline count according to

    w(k)=knlayers0+k(dff0+Aβ),N(nlayers0+k,dff0−w(k))=N(nlayers0,dff0).w(k)=\frac{k}{n^0_{\mathrm{layers}}+k}\left(d^0_{\mathrm{ff}}+\frac{A}{\beta}\right), \qquad N(n^0_{\mathrm{layers}}+k,d^0_{\mathrm{ff}}-w(k))=N(n^0_{\mathrm{layers}},d^0_{\mathrm{ff}}).

    Here the dimensions and layer counts are positive integers, VV is the tokenizer vocabulary size, and nlayers0+k>0n^0_{\mathrm{layers}}+k>0. Feed-forward dimensions derived from the formula were rounded to the nearest integer, so the realized parameter matching is approximate. The tested layer-count/feed-forward-dimension pairs were: 41M—1/4779, 2/2048, 3/1138, 4/682, 5/409, 6/227, 7/97; 134M—1/36k, 2/17k, 4/8193, 6/5121, 8/3584, 12/2048, 16/1280, 21/731, 26/393, 32/128; 374M—1/99k, 2/49k, 4/24k, 6/15k, 8/11k, 12/6998, 16/4907, 24/2816, 32/1770. The feed-forward ratio is defined as dmodel/dffd_{\mathrm{model}}/d_{\mathrm{ff}}.

  2. Knowl 2 — Pretraining and compositional fine-tuning protocol

    experimental setup

    All model variants were pretrained on the Colossal Clean Crawled Corpus (C4) using a 1024-token context and batches of 128 sequences, or approximately 131,000 tokens per step. Each model was trained for 1 million steps, corresponding to roughly 131 billion tokens. The pretrained models were then fine-tuned separately on the training portion of each of four tasks for 10,000 steps with batch size 128. Validation loss continued to decline during fine-tuning, so reported end-of-run results use the final checkpoint rather than early stopping.

    The tasks pair natural-language inputs with formal outputs or transformations: COGS tests semantic parsing with lexical and structural generalization splits; variable-free COGS (COGS-vf) uses a simplified semantic representation without numbered variables; GeoQuery uses the Standard split of geography questions paired with SQL-style queries; and English Passivization pairs active sentences with passive counterparts. Compositional generalization is evaluated on the held-out out-of-distribution split using full-sequence exact-match accuracy. C4 validation perplexity is used to measure language-modeling performance. The per-depth score listing reports one run per condition; some aggregate plots in the study report means over five runs.

  3. Knowl 3 — Depth improves C4 perplexity with rapidly diminishing returns

    empirical result

    Within each approximately parameter-matched family, deeper models generally achieved lower C4 validation perplexity after the same amount of pretraining. The improvement was largest among the shallowest models and diminished quickly as layers were added: the change from one to two layers was substantial, while gains from two to four layers were much smaller. The lowest perplexities occurred at 5 layers for the 41M family, 12 layers for 134M, and 24 layers for 374M; still deeper variants could become worse.

    The ratio between the one-layer model's perplexity and the best perplexity in its family was 1.59 for 41M models, 1.86 for 134M models, and 1.99 for 374M models. Perplexity began to deteriorate in the tested families when dffd_{\mathrm{ff}} fell below dmodeld_{\mathrm{model}}, or equivalently when the feed-forward ratio dmodel/dffd_{\mathrm{model}}/d_{\mathrm{ff}} exceeded 1. The depth at which this occurred depended on model size; larger models tolerated more extreme feed-forward ratios.

  4. Knowl 4 — Best reported scores vary by task and do not always occur at maximum depth

    empirical result

    The per-condition results show that deeper models often attain stronger compositional generalization, but the best reported depth depends on the task. The following are the best values within each parameter class in the single-run score listing; perplexity is lower-is-better, and task scores are out-of-distribution exact-match accuracy in percent, higher-is-better.

    • 41M: best C4 validation perplexity 28.8 at 4 and 5 layers; COGS 72.3% at 7 layers; COGS-vf 83.0% at 7 layers; GeoQuery Standard 79.6% at 3 layers; English Passivization 89.9% at 5 layers.
    • 134M: best C4 validation perplexity 18.1 at 12 layers; COGS 75.7% at 32 layers; COGS-vf 84.8% at 21 layers; GeoQuery Standard 82.9% at 12 layers; English Passivization 98.4% at 26 layers.
    • 374M: best C4 validation perplexity 14.4 at 24 layers; COGS 78.8% at 32 layers; COGS-vf 85.1% at 16 layers; GeoQuery Standard 84.6% at 32 layers; English Passivization 90.2% at 32 layers.

    Across the benchmark curves, deeper models generally outperform shallower models within the same size class, though COGS, COGS-vf, and GeoQuery show some non-monotonicity. English Passivization is especially variable: the 41M and 134M families mostly improve with depth, while the 374M family fluctuates more even though its deepest model exceeds its shallowest.

  5. Knowl 5 — Depth saturates at task-dependent thresholds and mainly helps lexical generalization

    empirical result

    For COGS, COGS-vf, and GeoQuery, most gains in out-of-distribution accuracy arrived within a small number of layers, after which accuracy remained relatively stable as depth increased. The approximate saturation ranges were 4–6 layers for COGS and 2–4 layers for both COGS-vf and GeoQuery; these ranges were low and similar across model sizes. English Passivization was too variable to identify a comparable size-independent threshold.

    A separate analysis of COGS and COGS-vf distinguishes lexical generalization—using a familiar word in a familiar syntactic context not seen in training—from structural generalization, which requires constructing a novel syntactic structure from familiar parts. Increasing depth improved lexical generalization but did not meaningfully improve structural generalization; even the deepest, largest models systematically struggled on structural cases. COGS-vf saturated with fewer layers than COGS despite expressing the same linguistic phenomena. The authors suggest that output-representation complexity, in addition to the generalization phenomenon itself, may affect where performance saturates.

  6. Knowl 6 — The depth advantage persists when pretraining perplexity is matched

    empirical result

    To test whether deeper models generalize better merely because they start fine-tuning with better language-modeling performance, the study compared pretrained checkpoints with equal C4 validation perplexity within a size class. The final perplexity of the shallowest reference model was used as the target; checkpoints from deeper models were selected when they reached that same perplexity, and those checkpoints were then fine-tuned on the compositional tasks. This comparison was repeated with successively deeper reference models.

    Deeper models still achieved better compositional generalization after this pretraining-perplexity match. One reported demonstration uses 134M models on COGS and finds the advantage through comparisons extending to six layers. The analysis generally considers models deeper than the reference, because shallower models typically do not reach the reference model's final perplexity during pretraining.

  7. Knowl 7 — The depth advantage persists at matched in-distribution fine-tuning loss

    empirical result

    The study also compared out-of-distribution generalization at fine-tuning checkpoints where models had the same loss on the in-distribution portion of COGS. At an in-distribution loss of 0.0002, deeper models had higher out-of-distribution accuracy than shallower models. Thus, the observed generalization advantage was not explained solely by deeper models fitting the fine-tuning training distribution better. This comparison was reported for a single run per condition.

  8. Knowl 8 — Depth raises latency, supporting shallower models under fixed parameter budgets

    empirical result

    Training latency increased strongly and approximately linearly with layer count among equal-parameter models. In a measurement of 374M models run on the same accelerator, the 32-layer model took about twice as long per training step as the 4-layer model; similar relative trends were observed for other size classes. The study also notes a per-layer inference cost. Two factors contribute: deeper variants incur a somewhat higher floating-point operation count under the chosen width tradeoff, and computations in each layer depend sequentially on the preceding layer's output, limiting parallelism across layers.

    With performance gains from depth often saturating after a few layers, the authors recommend considering shallower models when total parameters are fixed. Such models may reach acceptable performance faster for a fixed data volume, or process more data for a fixed training-time budget; their lower per-layer inference cost may also reduce serving or on-device computation. These are efficiency recommendations based on the measured latency and task results, not claims that shallow models always match the best deep model.

  9. Knowl 9 — Deep narrow feed-forward projections use their available rank unevenly

    empirical result

    The study examined the input and output affine projections in feed-forward blocks, focusing on the 134M family. For each projection, it computed the ordered singular values and normalized them by their sum, then measured the fraction of that normalized singular-value mass captured by the best rank-ii approximation. The available matrix rank is k=min⁡(dmodel,dff)k=\min(d_{\mathrm{model}},d_{\mathrm{ff}}).

    As depth increased and the feed-forward ratio dmodel/dffd_{\mathrm{model}}/d_{\mathrm{ff}} grew, both input and output projections showed increasingly concentrated singular-value distributions rather than making even use of their available rank. The authors characterize the transforms as becoming increasingly close to rank-deficient. This analysis accompanies the observed performance deterioration at extreme feed-forward ratios; it does not establish that spectral concentration is the cause of the deterioration.

  10. Knowl 10 — The findings are limited to the tested architecture, data, and adaptation regime

    limitation

    The experiments vary depth by narrowing feed-forward blocks while holding attention dimensions fixed; they do not establish how changing the number or function of attention heads affects compositional generalization. The study also does not test alternative ways of controlling parameter count, such as weight sharing across repeated layers, which could avoid some depth-induced narrowing.

    The models were pretrained on English natural-language data and evaluated on English compositional tasks, so the results do not determine how pretraining-corpus composition or language affects the depth relationship. The authors note that lexical overlap between C4 and the evaluation tasks may influence generalization. Finally, the study uses fine-tuning and does not test in-context learning, which would require scales beyond those trained here.

Coverage note — The exhaustive per-depth score cells are not reproduced individually; the knowls preserve the reported family-wise extrema and the broader depth trends. No other substantial contributed analysis was omitted.

References

  1. 1.Jason Ross Brown, Yiren Zhao, Ilia Shumailov, and Robert D Mullins. 2022. Wide attention is the way forward for Transformers? In Workshop: All Things Attention: Bridging Different Perspectives on Attention.
  2. 2.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Advances in Neural Information Processing Systems 33, volume 33, page 1877–1901. Curran Associates, Inc.
  3. 3.Róbert Csordás, Kazuki Irie, and Juergen Schmidhuber. 2021. The devil is in the detail: Simple tricks improve systematic generalization of transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 619–634, Stroudsburg, PA, USA. Association for Computational Linguistics.
  4. 4.Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. 2018. Universal Transformers. In International Conference on Learning Representations.
  5. 5.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), page 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  6. 6.Jerry A Fodor and Zenon W Pylyshyn. 1988. Connectionism and cognitive architecture: a critical analysis. Cognition, 28(1-2):3–71.
  7. 7.Google, Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, Eric Chu, Jonathan H Clark, Laurent El Shafey, Yanping Huang, Kathy Meier-Hellstern, Gaurav Mishra, Erica Moreira, Mark Omernick, Kevin Robinson, Sebastian Ruder, Yi Tay, Kefan Xiao, Yuanzhong Xu, Yujing Zhang, Gustavo Hernandez Abrego, Junwhan Ahn, Jacob Austin, Paul Barham, Jan Botha, James Bradbury, Siddhartha Brahma, Kevin Brooks, Michele Catasta, Yong Cheng, Colin Cherry, Christopher A Choquette-Choo, Aakanksha Chowdhery, Clément Crepy, Shachi Dave, Mostafa Dehghani, Sunipa Dev, Jacob Devlin, Mark Díaz, Nan Du, Ethan Dyer, Vlad Feinberg, Fangxiaoyu Feng, Vlad Fienber, Markus Freitag, Xavier Garcia, Lucas Gonzalez, Guy Gur-Ari, Steven Hand, Hadi Hashemi, Le Hou, Joshua Howland, Andrea Hu, Jeffrey Hui, Jeremy Hurwitz, Michael Isard, Abe Ittycheriah, Matthew Jagielski, Wenhao Jia, Kathleen Kenealy, Maxim Krikun, Sneha Kudugunta, Chang Lan, Katherine Lee, Benjamin Lee, Eric Li, Li, Music, Wei Li, Yaguang Li, Jian Li, Hyeontaek Lim, Hanzhao Lin, Zhongtao Liu, Frederick Liu, Marcello Maggioni, Aroma Mahendru, Joshua Maynez, Vedant Misra, Maysam Moussalem, Zachary Nado, John Nham, Eric Ni, Andrew Nystrom, Alicia Parrish, Marie Pellat, Martin Polacek, Alex Polozov, Reiner Pope, Siyuan Qiao, Emily Reif, Bryan Richter, Parker Riley, Alex Castro Ros, Aurko Roy, Brennan Saeta, Rajkumar Samuel, Renee Shelby, Ambrose Slone, Daniel Smilkov, David R So, Daniel Sohn, Simon Tokumine, Dasha Valter, Vijay Vasudevan, Kiran Vodrahalli, Xuezhi Wang, Pidong Wang, Zirui Wang, Tao Wang, John Wieting, Yuhuai Wu, Kelvin Xu, Yunhan Xu, Linting Xue, Pengcheng Yin, Jiahui Yu, Qiao Zhang, Steven Zheng, Ce Zheng, Weikang Zhou, Denny Zhou, Slav Petrov, and Yonghui Wu. 2023. PaLM 2 Technical Report. Technical report, Google.
  8. 8.Jonathan Gordon, David Lopez-Paz, Marco Baroni, and Diane Bouchacourt. 2019. Permutation Equivariant Models for Compositional Generalization in Language. In ICLR 2020 (OpenReview).
  9. 9.Andrey Gromov, Kushal Tirumala, Hassan Shapourian, Paolo Glorioso, and Daniel A. Roberts. 2024. The unreasonable ineffectiveness of the deeper layers. Technical Report MIT-CTP/5694, Center for Theoretical Physics, Massachusetts Institute of Technology, Cambridge, MA.
  10. 10.Manish Gupta and Puneet Agrawal. 2022. Compression of deep learning models for text: A survey. ACM Trans. Knowl. Discov. Data, 16(4).
  11. 11.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W Rae, Oriol Vinyals, and Laurent Sifre. 2022. Training Compute-Optimal Large Language Models.
  12. 12.Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R. Bowman. 2019. Do attention heads in bert track syntactic dependencies?
  13. 13.Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models.
  14. 14.Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, Dmitry Tsarkov, Xiao Wang, Marc van Zee, and Olivier Bousquet. 2020. Measuring Compositional Generalization: A Comprehensive Method on Realistic Data.
  15. 15.Najoung Kim and Tal Linzen. 2020. COGS: A Compositional Generalization Challenge Based on Semantic Interpretation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), page 9087–9105, Online. Association for Computational Linguistics.
  16. 16.Najoung Kim, Tal Linzen, and Paul Smolensky. 2022. Uncontrolled lexical exposure leads to overestimation of compositional generalization in pretrained models.
  17. 17.Brenden Lake and Marco Baroni. 2018. Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks. In Proceedings of the 35th International Conference on Machine Learning, pages 2873–2882. PMLR.
  18. 18.William Merrill, Ashish Sabharwal, and Noah A Smith. 2021. Saturated Transformers are Constant-Depth Threshold Circuits. Transactions of the Association for Computational Linguistics, pages 843–856.
  19. 19.Paul Michel, Omer Levy, and Graham Neubig. 2019. Are sixteen heads really better than one? In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc.
  20. 20.Richard Montague. 1970. Universal grammar. Theoria, 36(3):373–398.
  21. 21.Aaron Mueller, Robert Frank, Tal Linzen, Luheng Wang, and Sebastian Schuster. 2022. Coloring the Blank Slate: Pre-training Imparts a Hierarchical Inductive Bias to Sequence-to-sequence Models. In Findings of the Association for Computational Linguistics: ACL 2022, page 1352–1368, Dublin, Ireland. Association for Computational Linguistics.
  22. 22.Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023. Scaling data-constrained language models. In 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
  23. 23.Karl Mulligan, Robert Frank, and Tal Linzen. 2021. Structure Here, Bias There: Hierarchical Generalization by Jointly Learning Syntactic Transformations. In Proceedings of the Society for Computation in Linguistics 2021, pages 125–135.
  24. 24.Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D Manning. 2023. Characterizing Intrinsic Compositionality in Transformers with Tree Projections. In The Eleventh International Conference on Learning Representations.
  25. 25.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. 2022. In-context learning and induction heads.
  26. 26.Santiago Ontanon, Joshua Ainslie, Zachary Fisher, and Vaclav Cvicek. 2022. Making transformers solve compositional tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3591–3607, Stroudsburg, PA, USA. Association for Computational Linguistics.
  27. 27.OpenAI. 2023. GPT-4 Technical Report.
  28. 28.Isabel Papadimitriou and Dan Jurafsky. 2023. Injecting structural hints: Using language models to study inductive biases in language learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8402–8413, Singapore. Association for Computational Linguistics.
  29. 29.Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2022. Efficiently scaling transformer inference.
  30. 30.Linlu Qiu, Peter Shaw, Panupong Pasupat, Pawel Nowak, Tal Linzen, Fei Sha, and Kristina Toutanova. 2022a. Improving compositional generalization with latent structure and data augmentation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4341–4362, Seattle, United States. Association for Computational Linguistics.
  31. 31.Linlu Qiu, Peter Shaw, Panupong Pasupat, Tianze Shi, Jonathan Herzig, Emily Pitler, Fei Sha, and Kristina Toutanova. 2022b. Evaluating the impact of model scale for compositional generalization in semantic parsing. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9157–9179, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  32. 32.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research: JMLR, 21(2020):1–67.
  33. 33.Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl-Dickstein. 2017. On the Expressive Power of Deep Neural Networks. In International Conference on Machine Learning, pages 2847–2854. PMLR.
  34. 34.Adam Roberts, Hyung Won Chung, Anselm Levskaya, Gaurav Mishra, James Bradbury, Daniel Andor, Sharan Narang, Brian Lester, Colin Gaffney, Afroz Mohiuddin, Curtis Hawthorne, Aitor Lewkowycz, Alex Salcianu, Marc van Zee, Jacob Austin, Sebastian Goodman, Livio Baldini Soares, Haitang Hu, Sasha Tsvyashchenko, Aakanksha Chowdhery, Jasmijn Bastings, Jannis Bulian, Xavier Garcia, Jianmo Ni, Andrew Chen, Kathleen Kenealy, Jonathan H. Clark, Stephan Lee, Dan Garrette, James Lee-Thorp, Colin Raffel, Noam Shazeer, Marvin Ritter, Maarten Bosma, Alexandre Passos, Jeremy Maitin-Shepard, Noah Fiedel, Mark Omernick, Brennan Saeta, Ryan Sepassi, Alexander Spiridonov, Joshua Newlan, and Andrea Gesmundo. 2022. Scaling up models and data with t5x and seqio.
  35. 35.Tal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani, Dara Bahri, Vinh Tran, Yi Tay, and Donald Metzler. 2022. Confident adaptive language modeling. In Advances in Neural Information Processing Systems, volume 35, pages 17456–17472. Curran Associates, Inc.
  36. 36.Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Boyd-Graber, and Lijuan Wang. 2022. Prompting GPT-3 to be reliable. In The Eleventh International Conference on Learning Representations.
  37. 37.Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in NLP. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3645–3650, Florence, Italy. Association for Computational Linguistics.
  38. 38.Yi Tay, Mostafa Dehghani, Jinfeng Rao, William Fedus, Samira Abnar, Hyung Won Chung, Sharan Narang, Dani Yogatama, Ashish Vaswani, and Donald Metzler. 2021. Scale efficiently: Insights from pre-training and fine-tuning Transformers. In International Conference on Learning Representations.
  39. 39.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6000–6010, Red Hook, NY, USA. Curran Associates Inc.
  40. 40.Andreas Veit, Michael Wilber, and Serge Belongie. 2016. Residual networks behave like ensembles of relatively shallow networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS’16, page 550–558, Red Hook, NY, USA. Curran Associates Inc.
  41. 41.Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797–5808, Florence, Italy. Association for Computational Linguistics.
  42. 42.Thomas Wang, Adam Roberts, Daniel Hesslow, Teven Le Scao, Hyung Won Chung, Iz Beltagy, Julien Launay, and Colin Raffel. 2022. What language model architecture and pretraining objective works best for zero-shot generalization? In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 22964–22984. PMLR.
  43. 43.John M Zelle and Raymond J Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the thirteenth national conference on Artificial intelligence - Volume 2, AAAI’96, pages 1050–1055. AAAI Press.
  44. 44.Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020. BERT loses patience: Fast and robust inference with early exit. In Advances in Neural Information Processing Systems, volume 33, pages 18330–18341. Curran Associates, Inc.

Citation

MLA
Petty, J., et al. “The Impact of Depth on Compositional Generalization in Transformer Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 7239–52, https://doi.org/10.18653/v1/2024.naacl-long.402.
APA
Petty, J., Steenkiste, S. van ., Dasgupta, I., Sha, F., Garrette, D., & Linzen, T. (2024). The Impact of Depth on Compositional Generalization in Transformer Language Models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7239–7252. https://doi.org/10.18653/v1/2024.naacl-long.402
Chicago
Petty, J., S. van . Steenkiste, I. Dasgupta, F. Sha, D. Garrette, and T. Linzen. 2024. “The Impact of Depth on Compositional Generalization in Transformer Language Models”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 7239–52. https://doi.org/10.18653/v1/2024.naacl-long.402.
Harvard
Petty, J. et al. (2024) “The Impact of Depth on Compositional Generalization in Transformer Language Models”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7239–7252. Available at: https://doi.org/10.18653/v1/2024.naacl-long.402.
Vancouver
1. Petty J, Steenkiste S van, Dasgupta I, Sha F, Garrette D, Linzen T (2024) The Impact of Depth on Compositional Generalization in Transformer Language Models. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 7239–7252

BibTeX

@inproceedings{petty-etal-2024-impact,
    title = "The Impact of Depth on Compositional Generalization in Transformer Language Models",
    author = "Petty, Jackson  and
      van Steenkiste, Sjoerd  and
      Dasgupta, Ishita  and
      Sha, Fei  and
      Garrette, Dan  and
      Linzen, Tal",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.402/",
    doi = "10.18653/v1/2024.naacl-long.402",
    pages = "7239--7252"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/