Mechanistic Design and Scaling of Hybrid Architectures

Michael PoliArmin W. ThomasEric NguyenPragaash PonnusamyBjörn DeiserothKristian KerstingTaiji SuzukiBrian L. HieStefano ErmonChristopher Ré

article2024ICML67 citations

Proposes a mechanistic architecture design framework that uses small-scale synthetic token manipulation tasks to predict large-scale compute-optimal performance, yielding hybrid models combining attention, convolutions, and recurrences that outperform standard Transformers and Mamba across models up to 7B parameters.

Listen

Developing modern deep learning architectures is increasingly constrained by immense computational costs, lengthy prototyping cycles, and the combinatorial explosion of potential layer combinations. While standard models rely almost exclusively on uniform attention-based Transformer recipes, emerging signal processing primitives—such as gated convolutions and linear recurrences—offer better computational efficiency and fast inference. However, evaluating these alternatives traditionally requires training massive models from scratch. The article addresses this critical bottleneck by proposing a systematic framework to rapidly design, test, and predict the large-scale performance of new sequence modeling architectures.

The main objective of the article is to introduce Mechanistic Architecture Design (MAD), a lightweight evaluation pipeline using small synthetic unit tests, and to demonstrate that these proxy tests accurately predict large-scale compute-optimal scaling laws across emerging hybrid architectures. To validate this framework, the authors conducted an extensive scaling law study by training over 500 language models ranging from 70 million to 7 billion parameters across diverse computational budgets on standard pretraining datasets.

The high-level approach evaluates candidate architectures—composed of linear recurrences, convolutions, attention, and mixture-of-experts mechanisms—on six isolated token manipulation tasks: in-context recall, fuzzy recall, noisy recall, selective copying, compression, and memorization. These synthetic unit tests run in minutes while strictly controlling for parameter count and recurrent state dimension. Promising candidates identified by the testing pipeline are then scaled up across multiple fixed-compute budgets to construct empirical scaling laws, map compute-optimal frontiers, and analyze memory state trade-offs.

The analysis yields four key findings. First, aggregate scores on small synthetic proxy tasks exhibit a strong rank-correlation with compute-optimal language modeling perplexity at scale, allowing reliable architectural filtering at a fraction of standard evaluation costs. Second, hybrid "striped" architectures that interleave specialized recurrent or convolutional layers with attention layers consistently outperform pure architectures, achieving up to a 20% reduction in perplexity for the same compute budget and an average 8.1% accuracy gain on synthetic tests. Third, hybrid architectures achieve an optimal attention-to-alternative layer ratio of approximately 25% across compute budgets, while proving significantly more robust when smaller models are overtrained for longer durations. Fourth, the state-optimal scaling analysis reveals a consistent power-law relationship between model state size and perplexity, demonstrating that hybrid designs balance memory footprint and compute efficiency far better than standard Transformers.

These findings have direct practical implications for reducing the development timelines, hardware costs, and operational risks associated with large-scale artificial intelligence models. Hybrid architectures allow organizations to deploy models that achieve superior accuracy while requiring lower inference memory and serving costs. Furthermore, because hybrids maintain higher performance when trained on massive token volumes outside the theoretical compute-optimal frontier, engineering teams can prioritize smaller, high-throughput models for deployment without suffering the severe quality degradation typical of standard Transformers.

Organizations should adopt mechanistic proxy benchmarks to rapidly prototype and screen novel architectural building blocks before committing large training budgets. When architecting foundation models, engineering teams should shift away from pure Transformer recipes toward striped hybrid topologies incorporating sparse channel and sequence experts. However, stakeholders should note that the proxy framework was primarily validated on two-block prototype models and standard autoregressive text modeling; teams exploring highly complex multi-primitive topologies or non-language domains should perform intermediate pilot scale-ups before executing full-scale production training runs.

arXiv: 2403.17844

No sufficiently relevant recommendations were found.

Cover for Mechanistic Design and Scaling of Hybrid Architectures

Abstract

The development of deep learning architectures is a resource-demanding process, due to a vast design space, long prototyping times, and high compute costs associated with at-scale model training and evaluation. We set out to simplify this process by grounding it in an end-to-end mechanistic architecture design (MAD) pipeline, encompassing small-scale capability unit tests predictive of scaling laws. Through a suite of synthetic token manipulation tasks such as compression and recall, designed to probe capabilities, we identify and test new hybrid architectures constructed from a variety of computational primitives. We experimentally validate the resulting architectures via an extensive compute-optimal and a new state-optimal scaling law analysis, training over 500 language models between 70M to 7B parameters. Surprisingly, we find MAD synthetics to correlate with compute-optimal perplexity, enabling accurate evaluation of new architectures via isolated proxy tasks. The new architectures found via MAD, based on simple ideas such as hybridization and sparsity, outperform state-of-the-art Transformer, convolutional, and recurrent architectures (Transformer++, Hyena, Mamba) in scaling, both at compute-optimal budgets and in overtrained regimes. Overall, these results provide evidence that performance on curated synthetic tasks can be predictive of scaling laws, and that an optimal architecture should leverage specialized layers via a hybrid topology.

Table of Contents

  • 1. Introduction
  • 2. Background: Architecture Design
  • 2.1. Computational primitives
  • 2.2. State, cache, and memory
  • 2.3. Topology
  • 3. Mechanistic Architecture Design
  • 3.1. Synthetic tasks to probe model skills
  • 3.1.1. IN-CONTEXT RECALL
  • 3.1.2. FUZZY IN-CONTEXT RECALL
  • 3.1.3. NOISY IN-CONTEXT RECALL
  • 3.1.4. SELECTIVE COPYING
  • 3.1.5. COMPRESSION
  • 3.1.6. MEMORIZATION
  • 3.2. MAD Protocol
  • 3.3. Candidate architecture designs
  • 3.4. Results
  • 4. Scaling Analysis
  • 4.1. Compute-optimal frontier for new architectures
  • 4.2. State-optimal scaling
  • 5. Connecting MAD to scaling metrics
  • 6. Conclusion
  • Impact Statement
  • Acknowledgments
  • References
  • A. Additional Related Work
  • B. Mechanistic Architecture Design
  • B.1. Tasks
  • B.1.1. IN-CONTEXT RECALL
  • B.1.2. FUZZY IN-CONTEXT RECALL
  • B.1.3. NOISY IN-CONTEXT RECALL
  • B.1.4. SELECTIVE COPYING
  • B.1.5. COMPRESSION
  • B.1.6. MEMORIZATION
  • B.2. Manipulating Task Difficulty
  • B.3. Architectures
  • B.3.1. CHANNEL-MIXING LAYERS
  • B.3.2. SEQUENCE-MIXING LAYERS
  • B.4. Training
  • B.5. Results
  • B.5.1. TASK PERFORMANCES
  • B.5.2. PERFORMANCE ON INDIVIDUAL TASKS
  • B.6. Extensions and Limitations of MAD
  • C. Scaling Laws
  • C.1. Training Details
  • C.2. Model architectures
  • C.3. Model sizes and training hyperparameters
  • C.4. FLOP calculation
  • C.4.1. TRANSFORMER ++
  • C.4.2. HYENA
  • C.4.3. MULTI-HEAD HYENA
  • C.4.4. STRIPEDHYENA
  • C.4.5. MAMBA
  • C.4.6. STRIPEDMAMBA
  • C.4.7. STRIPEDHYENA-MOE
  • C.4.8. STRIPEDHYENA EXPERTS + MOE
  • D. Scaling Laws
  • D.1. Training Details
  • D.2. Model architectures
  • D.3. Model sizes and training hyperparameters
  • D.4. FLOP calculation
  • D.4.1. TRANSFORMER ++
  • D.4.2. HYENA
  • D.4.3. MULTI-HEAD HYENA
  • D.4.4. STRIPEDHYENA
  • D.4.5. MAMBA
  • D.4.6. STRIPEDMAMBA
  • D.4.7. STRIPEDHYENA-MOE
  • D.4.8. STRIPEDHYENA EXPERTS + MOE
  • D.5. State-optimal scaling
  • E. Extended Scaling Results
  • E.1. Optimal hybridization topologies
  • E.2. Byte-level scaling laws
  • Scaling Rates for Compute-Suboptimal Model Sizes

Knowls

  1. Knowl 1 — MAD uses synthetic token tasks as architecture capability tests

    model/method

    Mechanistic architecture design (MAD) is a rapid prototyping method that evaluates small models on synthetic token-manipulation tasks intended to isolate capabilities relevant to sequence modeling. Its six tasks are: (1) multi-query in-context recall, retrieving values for keys already encountered in a sequence when the key–value mapping changes between examples; (2) fuzzy recall, where keys and values comprise variable-length spans of adjacent tokens; (3) noisy recall, which adds irrelevant tokens from a separate vocabulary among key–value pairs; (4) selective copying, which requires copying relevant input tokens in their original order while ignoring noise; (5) compression, which requires collecting a random sequence into one token representation from which an MLP decoder reconstructs the sequence; and (6) memorization, which requires learning a fixed key–value mapping from training examples. The tasks probe distinct abilities, including prompt-based recall, noise filtering, ordered memory, information aggregation, and stored factual associations.

  2. Knowl 2 — MAD controls task difficulty and model-state capacity

    experimental setup

    MAD scores average model performance across task settings that vary one difficulty factor at a time, including sequence length, vocabulary size, and training-set size; noisy recall and selective copying also vary noise fraction and number of copied tokens, respectively. Models are assessed on independent evaluation data. In the main protocol, architectures are small two-block models with four layers and width 128. Fixed-state sequence mixers are normalized to a common total state dimension of 4,096, including architectures with sparse expert layers, to reduce capacity differences as an explanation for performance gaps. Runs sweep learning rates {0.0001,0.0005,0.001}\{0.0001, 0.0005, 0.001\} and weight decay {0,0.1}\{0, 0.1\}; reported scores use the best run. Training uses AdamW, batch size 128, 200 epochs, cosine learning-rate decay, and no dropout.

  3. Knowl 3 — Hyena experts sparsely route sequence mixing

    model/method

    The Hyena experts layer applies input-dependent sparse routing to sequence-mixing computation. For an input sequence U∈RT×DU\in\mathbb{R}^{T\times D}, where TT is sequence length and DD is width, a learned router with weights Wg∈RD×EW_g\in\mathbb{R}^{D\times E} scores EE Hyena experts at each token position. It selects the top KK experts, normalizes their selected scores with softmax to obtain weights st,ks_{t,k}, and forms each output as a weighted sum of those experts’ outputs: yt=∑k=1Kst,k Hyenak(U)ty_t=\sum_{k=1}^{K}s_{t,k}\,\mathrm{Hyena}_k(U)_t. Thus, only a subset of the sequence mixers contributes at each position. The evaluated expert configuration uses eight total experts and two active experts.

  4. Knowl 4 — Striped hybrids improve aggregate MAD performance

    empirical result

    Across the tested synthetic tasks, architectures that sequentially interleave different sequence-mixing primitives in striped topologies outperform non-striped architectures on the aggregate MAD measure. The reported average accuracy gain is 8.1%. The comparison evaluates 21 architectures built from attention, Hyena, Mamba, gated linear attention, and channel mixers; the result supports using complementary sequence-mixing blocks together rather than relying on a single primitive throughout.

  5. Knowl 5 — Heads and sparse channel mixers provide additional MAD gains

    empirical result

    In the MAD comparisons, fixed-state architectures that organize their state using heads outperform corresponding architectures without heads, with a reported average accuracy gain of 2.3% across the synthetic tasks. Adding sparsely activated mixture-of-experts channel mixers instead of SwiGLU channel mixers yields a further reported average accuracy gain of 1.7%. Task-level results also show specialization among primitives: Mamba performs particularly well on compression, while Hyena performs particularly well on fuzzy recall. These comparisons motivate combining specialized components, while the head comparison also entails increased computation in the layer.

  6. Knowl 6 — Compute-optimal scaling is estimated with fixed-FLOP model sweeps

    experimental setup

    The scaling analysis trains more than 500 language models, spanning approximately 70 million to 7 billion parameters, and compares architectures on The Pile. For an architecture ξ\xi, let NN be parameter count, DD the number of training tokens, cξ(N,D)c_\xi(N,D) the training cost in FLOPs, and Lξ(N,D)L_\xi(N,D) the resulting loss. At compute budget CC, the compute-optimal loss is ℓξ(C)=min⁡(N,D): cξ(N,D)=CLξ(N,D)\ell_\xi(C)=\min_{(N,D):\,c_\xi(N,D)=C} L_\xi(N,D). The study estimates this frontier using IsoFLOP sweeps: it trains different model-size and token-count allocations at fixed compute budgets, then fits the observed losses to estimate the minimum. The tested budgets range from 4×10184\times10^{18} to 2×10212\times10^{21} FLOPs.

  7. Knowl 7 — Hybrid architectures improve compute-optimal scaling

    empirical result

    On The Pile, the new hybrid architectures improve compute-optimal pretraining perplexity over the tested Transformer++, Hyena, and Mamba baselines; the paper reports reductions of up to 20% at the same compute budget. In the scaling sweeps, the compute-optimal striped hybrid uses 25% attention across the tested IsoFLOP groups; when state-optimality is also considered, the reported optimal attention ratio shifts to 10%. The fitted allocation trends indicate that striped hybrids favor relatively more training data and smaller models than Transformer models at a given compute budget.

  8. Knowl 8 — Hybrids are more robust to training beyond the compute-optimal frontier

    empirical result

    When compute is held fixed but allocated to a smaller-than-optimal model trained on more tokens, the tested hybrid architectures have flatter IsoFLOP perplexity curves and smaller suboptimality gaps than Transformer++ models. This makes the hybrids more robust in the evaluated overtraining regime, a setting relevant to training smaller models for lower inference cost. The reported gap increases with hybridization ratio across the IsoFLOP groups, so the robustness result is conditional on the allocation and hybridization choices rather than a claim that every hybrid is equally tolerant.

  9. Knowl 9 — Compute-optimal perplexity follows a state-size scaling relation

    empirical result

    The state-optimal analysis estimates a power-law relation between compute-optimal perplexity P∗P^* and total state dimension MM: P∗∝McP^*\propto M^c, with c≈−0.28c\approx-0.28 in the experimental setup. The exponent is reported as consistent across the tested model architectures, while the architecture class shifts the curve’s offset. Here, total state includes fixed recurrent state and dynamic attention state. The results indicate that a target perplexity, up to saturation of compute-optimal scaling, can be pursued with fixed-state architectures by paying additional FLOPs and training longer to use the available state; input-varying recurrences, multihead designs, and striped hybrids show favorable trade-offs between perplexity and state size.

  10. Knowl 10 — MAD scores predict compute-optimal perplexity within the tested scope

    empirical result

    For the tested architecture variants, aggregate MAD accuracy is reported to be linearly correlated with compute-optimal perplexity on The Pile across compute budgets: better aggregate synthetic-task performance accompanies lower compute-optimal perplexity. The relationship is especially useful for iterating on related architectures, such as successive Hyena-based designs, and the paper reports particularly strong rank correlation among models within the same architecture class. This is evidence that small, shallow synthetic-task evaluations can predict some scaling outcomes, not proof that MAD predicts all domains or architectures. The evaluation does not establish transfer to more complex topologies requiring many blocks or more than two sequence-mixing primitives.

Coverage note — The supplementary byte-level and DNA scaling analyses, full task-by-task performance curves, and detailed FLOP calculators are omitted; they are auxiliary domain-specific or implementation analyses rather than necessary components of the central MAD and language-model scaling results.

References

  1. 1.Akyürek, E., Wang, B., Kim, Y., and Andreas, J. In-context language learning: Architectures and algorithms. arXiv preprint arXiv:2401.12973, 2024.
  2. 2.Arora, S., Eyuboglu, S., Timalsina, A., Johnson, I., Poli, M., Zou, J., Rudra, A., and Ré, C. Zoology: Measuring and improving recall in efficient language models. arXiv preprint arXiv:2312.04927, 2023.
  3. 3.Bhattamishra, S., Patel, A., Blunsom, P., and Kanade, V. Understanding in-context learning in transformers and llms by learning to learn discrete functions. arXiv preprint arXiv:2310.03016, 2023.
  4. 4.Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al. Deepseek llm: Scaling open-source language models with longtermism. arXiv preprint arXiv:2401.02954, 2024.
  5. 5.Dupont, E., Doucet, A., and Teh, Y. W. Augmented neural odes. Advances in neural information processing systems, 32, 2019.
  6. 6.Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al. A mathematical framework for transformer circuits. Transformer Circuits Thread, 1, 2021.
  7. 7.Fathi, M., Pilault, J., Bacon, P.-L., Pal, C., Firat, O., and Goroshin, R. Block-state transformer. arXiv preprint arXiv:2306.09539, 2023.
  8. 8.Fu, D. Y., Dao, T., Saab, K. K., Thomas, A. W., Rudra, A., and Ré, C. Hungry hungry hippos: Towards language modeling with state space models. arXiv preprint arXiv:2212.14052, 2022.
  9. 9.Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020.
  10. 10.Geva, M., Bastings, J., Filippova, K., and Globerson, A. Dissecting recall of factual associations in auto-regressive language models. arXiv preprint arXiv:2304.14767, 2023.
  11. 11.Gu, A. and Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023.
  12. 12.Gu, A., Goel, K., and Ré, C. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021.
  13. 13.Hewitt, J., Hahn, M., Ganguli, S., Liang, P., and Manning, C. D. Rnns can generate bounded hierarchical languages with optimal memory. arXiv preprint arXiv:2010.07515, 2020.
  14. 14.Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022.
  15. 15.Hua, W., Dai, Z., Liu, H., and Le, Q. Transformer quality in linear time. In International Conference on Machine Learning, pp. 9099–9117. PMLR, 2022.
  16. 16.Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al. Mistral 7b. arXiv preprint arXiv:2310.06825, 2023.
  17. 17.Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020.
  18. 18.Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F. Transformers are rnns: Fast autoregressive transformers with linear attention. In International conference on machine learning, pp. 5156–5165. PMLR, 2020.
  19. 19.Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z. Gshard: Scaling giant models with conditional computation and automatic sharding. arXiv preprint arXiv:2006.16668, 2020.
  20. 20.Lin, Z., Feng, M., Santos, C. N. d., Yu, M., Xiang, B., Zhou, B., and Bengio, Y. A structured self-attentive sentence embedding. arXiv preprint arXiv:1703.03130, 2017.
  21. 21.Ma, X., Zhou, C., Kong, X., He, J., Gui, L., Neubig, G., May, J., and Zettlemoyer, L. Mega: moving average equipped gated attention. arXiv preprint arXiv:2209.10655, 2022.
  22. 22.Massaroli, S., Poli, M., Park, J., Yamashita, A., and Asama, H. Dissecting neural odes. Advances in Neural Information Processing Systems, 33:3952–3963, 2020.
  23. 23.Massaroli, S., Poli, M., Fu, D. Y., Kumbong, H., Parnichkun, R. N., Timalsina, A., Romero, D. W., McIntyre, Q., Chen, B., Rudra, A., et al. Laughing hyena distillery: Extracting compact recurrences from convolutions. arXiv preprint arXiv:2310.18780, 2023.
  24. 24.Nanda, N., Rajamanoharan, S., Kramár, J., and Shah, R. Fact finding: Attempting to reverse-engineer factual recall on the neuron level. Alignment Forum, 2023.
  25. 25.Nguyen, E., Poli, M., Faizi, M., Thomas, A., Birch-Sykes, C., Wornow, M., Patel, A., Rabideau, C., Massaroli, S., Bengio, Y., et al. Hyenadna: Long-range genomic sequence modeling at single nucleotide resolution. arXiv preprint arXiv:2306.15794, 2023.
  26. 26.Olsson, C., Elhage, N., Nanda, N., Joseph, N., Das-Sarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al. In-context learning and induction heads. arXiv preprint arXiv:2209.11895, 2022.
  27. 27.Peng, B., Alcaide, E., Anthony, Q., Albalak, A., Arcadinho, S., Cao, H., Cheng, X., Chung, M., Grella, M., GV, K. K., et al. Rwkv: Reinventing rnns for the transformer era. arXiv preprint arXiv:2305.13048, 2023.
  28. 28.Poli, M., Massaroli, S., Nguyen, E., Fu, D. Y., Dao, T., Baccus, S., Bengio, Y., Ermon, S., and Ré, C. Hyena hierarchy: Towards larger convolutional language models. arXiv preprint arXiv:2302.10866, 2023.
  29. 29.Ren, L., Liu, Y., Wang, S., Xu, Y., Zhu, C., and Zhai, C. X. Sparse modular activation for efficient sequence modeling. Advances in Neural Information Processing Systems, 36, 2024.
  30. 30.Sardana, N. and Frankle, J. Beyond chinchilla-optimal: Accounting for inference in language model scaling laws. arXiv preprint arXiv:2401.00448, 2023.
  31. 31.Schlag, I., Irie, K., and Schmidhuber, J. Linear transformers are secretly fast weight programmers. In International Conference on Machine Learning, pp. 9355–9366. PMLR, 2021.
  32. 32.Shazeer, N. Glu variants improve transformer. arXiv preprint arXiv:2002.05202, 2020.
  33. 33.Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017.
  34. 34.So, D. R., Manke, W., Liu, H., Dai, Z., Shazeer, N., and Le, Q. V. Primer: Searching for efficient transformers for language modeling. arXiv preprint arXiv:2109.08668, 2021.
  35. 35.Stanić, A., Ashley, D., Serikov, O., Kirsch, L., Faccio, F., Schmidhuber, J., Hofmann, T., and Schlag, I. The languini kitchen: Enabling language modelling research at different scales of compute. arXiv preprint arXiv:2309.11197, 2023.
  36. 36.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  37. 37.Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  38. 38.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  39. 39.Weiss, G., Goldberg, Y., and Yahav, E. On the practical computatifonal power of finite precision rnns for language recognition. arXiv preprint arXiv:1805.04908, 2018.
  40. 40.White, C., Safari, M., Sukthanker, R., Ru, B., Elsken, T., Zela, A., Dey, D., and Hutter, F. Neural architecture search: Insights from 1000 papers. arXiv preprint arXiv:2301.08727, 2023.
  41. 41.Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y. Gated linear attention transformers with hardware-efficient training. arXiv preprint arXiv:2312.06635, 2023.
  42. 42.Zhang, B. and Sennrich, R. Root mean square layer normalization. Advances in Neural Information Processing Systems, 32, 2019.
  43. 43.Zhang, M., Saab, K. K., Poli, M., Dao, T., Goel, K., and Ré, C. Effectively modeling time series with simple discrete state spaces. arXiv preprint arXiv:2303.09489, 2023.

Citation

MLA
Poli, M., et al. “Mechanistic Design and Scaling of Hybrid Architectures”. arXiv, 2024, http://arxiv.org/abs/2403.17844v2.
APA
Poli, M., Thomas, A. W., Nguyen, E., Ponnusamy, P., Deiseroth, B., Kersting, K., Suzuki, T., Hie, B., Ermon, S., Ré, C., Zhang, C., & Massaroli, S. (2024). Mechanistic Design and Scaling of Hybrid Architectures. arXiv. http://arxiv.org/abs/2403.17844v2
Chicago
Poli, M., A. W. Thomas, E. Nguyen, et al. 2024. “Mechanistic Design and Scaling of Hybrid Architectures”. arXiv. http://arxiv.org/abs/2403.17844v2.
Harvard
Poli, M. et al. (2024) “Mechanistic Design and Scaling of Hybrid Architectures”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.17844v2.
Vancouver
1. Poli M, Thomas AW, Nguyen E, et al (2024) Mechanistic Design and Scaling of Hybrid Architectures. arXiv

BibTeX

@article{poli2024mechanistic,
  title = {Mechanistic Design and Scaling of Hybrid Architectures},
  author = {Poli, Michael and Thomas, Armin W and Nguyen, Eric and Ponnusamy, Pragaash and Deiseroth, Björn and Kersting, Kristian and Suzuki, Taiji and Hie, Brian and Ermon, Stefano and Ré, Christopher and Zhang, Ce and Massaroli, Stefano},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.17844v2},
  eprint = {2403.17844}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/