keyword
Universal Transformers
A Universal Transformer is a sequence-processing neural network architecture that combines self-attention mechanisms with recurrent depth by repeatedly applying a single, weight-shared layer across multiple computational steps. Unlike standard Transformer models that pass inputs through a fixed sequence of distinct layers, a Universal Transformer updates the vector representations of all tokens in parallel using the same recurrent transition block at each step. This recurrent parameter-sharing design allows the network to iteratively refine representations over an arbitrary number of steps, enabling variable computation depth and granting the architecture theoretical Turing completeness under standard precision assumptions. Additionally, these models often incorporate adaptive computation mechanisms that allow individual token positions to dynamically halt processing once sufficient computation has occurred.
7 items

The Impact of Depth on Compositional Generalization in Transformer Language Models
Jackson Petty, Sjoerd van Steenkiste, Ishita Dasgupta, Fei Sha, Dan Garrette, Tal Linzen
Why you should read this
Demonstrates that deeper transformer models improve compositional generalization over wider models of equal parameter size but yield rapidly diminishing returns, proving that practitioners can adopt shallower architectures to lower latency without sacrificing performance.
To process novel sentences, language models (LMs) must generalize compositionally -- combine familiar elements in new ways. What aspects of a model's structure promote compositional generalization? Focusing on transformers, we test the hypothesis, motivated by theoretical and empirical work, that deeper transformers generalize more compositionally. Simply adding layers increases the total number of parameters; to address this confound between depth and size, we construct three classes of models which trade off depth for width such that the total number of parameters is kept constant (41M, 134M and 374M parameters). We pretrain all models as LMs and fine-tune them on tasks that test for compositional generalization. We report three main conclusions: (1) after fine-tuning, deeper models generalize more compositionally than shallower models do, but the benefit of additional layers diminishes rapidly; (2) within each family, deeper models show better language modeling performance, but returns are similarly diminishing; (3) the benefits of depth for compositional generalization cannot be attributed solely to better performance on language modeling. Because model latency is approximately linear in the number of layers, these results lead us to the recommendation that, with a given total parameter budget, transformers can be made shallower than is typical without sacrificing performance.
Added
2026-10-04

Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri, Donald Metzler
Why you should read this
Systematizes dozens of efficient Transformer variants into a clear taxonomy of computational and memory trade-offs, helping researchers choose the optimal architecture for resource-constrained deep learning applications.
Transformer model architectures have garnered immense interest lately due to their effectiveness across a range of domains like language, vision and reinforcement learning. In the field of natural language processing for example, Transformers have become an indispensable staple in the modern deep learning stack. Recently, a dizzying number of "X-former" models have been proposed - Reformer, Linformer, Performer, Longformer, to name a few - which improve upon the original Transformer architecture, many of which make improvements around computational and memory efficiency. With the aim of helping the avid researcher navigate this flurry, this paper characterizes a large and thoughtful selection of recent efficiency-flavored "X-former" models, providing an organized and comprehensive overview of existing work and models across multiple domains.
Added
2026-09-24

Recurrent Looped Transformer
Yifan Zhang
Why you should read this
Introduces an architecture that combines unbounded temporal depth with hardware-aware execution and consistent reinforcement learning algorithms.
We propose Recurrent Looped Transformer (RLT), built around three principles: latent reasoning with unbounded temporal depth, model-hardware co-design for efficient execution, and model-RL algorithm co-design for consistent policy optimization. A causal encoder constructs key-value memory, while a recurrent decoder carries its final hidden state and layerwise sliding-window attention cache across every prompt and response token. This creates a latent computation path with no fixed architectural depth limit as the sequence extends, using a fixed number of blocks per token.
Added
2026-09-14

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation
Amr Hegazy, Amr Alanwar, Mostafa Elhoushi
Why you should read this
Introduces a gated recurrent transformer architecture that uses state-conditioned modulation across iterated shared layers to match standard language model performance while reducing parameters and peak decoding memory by roughly sixty percent.
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce Gated Recurrent Transformer, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer Gated Recurrent Transformer matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality: at large scale, 63% fewer parameters and 59% less peak decoding memory for a 10% increase in compiled generation latency.
Added
2026-08-29


DeepLoop: Depth Scaling for Looped Transformers
Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient . The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from to as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets and for unrolled depth . On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.
Added
2026-07-31

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus, Thomas Hofmann, Valentina Boeva, T. Konstantin Rusch, Antonio Orvieto
Why you should read this
Presents FPRM, a Transformer-based Fixed-Point Reasoning Model that leverages fixed-point convergence to dynamically adapt its computational effort to task difficulty in looped architectures, addressing signal propagation issues through architectural modifications.
Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by looping determines the quality of the solution these models find. Like deep architectures, looped architectures are prone to a signal propagation problem induced by depth as the halting decision is postponed. In this paper, we address this signal propagation issue using pre-norm layers and residual scaling. Building on these architectural modifications, we propose FPRM, a Transformer-based Fixed-Point Reasoning Model that uses fixed-point convergence as an end-to-end halting mechanism in a looped architecture. We show that fixed-point halting allows FPRM to adapt its compute to task difficulty. FPRM is effective on common reasoning benchmarks, namely Sudoku, Maze, state-tracking, and ARC-AGI.
Added
2026-06-21


Synthesizer: Rethinking Self-Attention for Transformer Models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, Che Zheng
Why you should read this
Proposes Synthesizer, a novel model that achieves competitive or superior performance to traditional Transformers with significantly increased speed and efficiency by synthesizing attention weights without direct token-token interactions, challenging the fundamental reliance on dot product self-attention.
The dot product self-attention is known to be central and indispensable to state-of-the-art Transformer models. But is it really required? This paper investigates the true importance and contribution of the dot product-based self-attention mechanism on the performance of Transformer models. Via extensive experiments, we find that (1) random alignment matrices surprisingly perform quite competitively and (2) learning attention weights from token-token (query-key) interactions is useful but not that important after all. To this end, we propose \textsc{Synthesizer}, a model that learns synthetic attention weights without token-token interactions. In our experiments, we first show that simple Synthesizers achieve highly competitive performance when compared against vanilla Transformer models across a range of tasks, including machine translation, language modeling, text generation and GLUE/SuperGLUE benchmarks. When composed with dot product attention, we find that Synthesizers consistently outperform Transformers. Moreover, we conduct additional comparisons of Synthesizers against Dynamic Convolutions, showing that simple Random Synthesizer is not only faster but also improves perplexity by a relative . Finally, we show that simple factorized Synthesizers can outperform Linformers on encoding only tasks.
Added
2026-02-13

