Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Universal Transformers

A Universal Transformer is a sequence-processing neural network architecture that combines self-attention mechanisms with recurrent depth by repeatedly applying a single, weight-shared layer across multiple computational steps. Unlike standard Transformer models that pass inputs through a fixed sequence of distinct layers, a Universal Transformer updates the vector representations of all tokens in parallel using the same recurrent transition block at each step. This recurrent parameter-sharing design allows the network to iteratively refine representations over an arbitrary number of steps, enabling variable computation depth and granting the architecture theoretical Turing completeness under standard precision assumptions. Additionally, these models often incorporate adaptive computation mechanisms that allow individual token positions to dynamically halt processing once sufficient computation has occurred.

7 items

The Impact of Depth on Compositional Generalization in Transformer Language Models

The Impact of Depth on Compositional Generalization in Transformer Language Models

Jackson Petty, Sjoerd van Steenkiste, Ishita Dasgupta, Fei Sha, Dan Garrette, Tal Linzen

OrganizationsGoogleNew York University

Why you should read this

Demonstrates that deeper transformer models improve compositional generalization over wider models of equal parameter size but yield rapidly diminishing returns, proving that practitioners can adopt shallower architectures to lower latency without sacrificing performance.

To process novel sentences, language models (LMs) must generalize compositionally -- combine familiar elements in new ways. What aspects of a model's structure promote compositional generalization? Focusing on transformers, we test the hypothesis, motivated by theoretical and empirical work, that deeper transformers generalize more compositionally. Simply adding layers increases the total number of parameters; to address this confound between depth and size, we construct three classes of models which trade off depth for width such that the total number of parameters is kept constant (41M, 134M and 374M parameters). We pretrain all models as LMs and fine-tune them on tasks that test for compositional generalization. We report three main conclusions: (1) after fine-tuning, deeper models generalize more compositionally than shallower models do, but the benefit of additional layers diminishes rapidly; (2) within each family, deeper models show better language modeling performance, but returns are similarly diminishing; (3) the benefits of depth for compositional generalization cannot be attributed solely to better performance on language modeling. Because model latency is approximately linear in the number of layers, these results lead us to the recommendation that, with a given total parameter budget, transformers can be made shallower than is typical without sacrificing performance.

Added

2026-10-04

Recurrent Looped Transformer

Recurrent Looped Transformer

Yifan Zhang

Why you should read this

Introduces an architecture that combines unbounded temporal depth with hardware-aware execution and consistent reinforcement learning algorithms.

We propose Recurrent Looped Transformer (RLT), built around three principles: latent reasoning with unbounded temporal depth, model-hardware co-design for efficient execution, and model-RL algorithm co-design for consistent policy optimization. A causal encoder constructs key-value memory, while a recurrent decoder carries its final hidden state and layerwise sliding-window attention cache across every prompt and response token. This creates a latent computation path with no fixed architectural depth limit as the sequence extends, using a fixed number of blocks per token.

Added

2026-09-14

License: Apache 2.0
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Amr Hegazy, Amr Alanwar, Mostafa Elhoushi

OrganizationsCerebras SystemsGerman University in CairoTechnical University of Munich

Why you should read this

Introduces a gated recurrent transformer architecture that uses state-conditioned modulation across iterated shared layers to match standard language model performance while reducing parameters and peak decoding memory by roughly sixty percent.

Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce Gated Recurrent Transformer, a recurrent depth transformer where fixed-depth prelude and coda blocks bracket a single shared core iterated R times. Inspired by gated recurrent neural networks, we employ a lightweight projection and an elementwise update gate---conditioned on the hidden state, the fixed prelude output, and noise resampled at every step---to modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer Gated Recurrent Transformer matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads MoR and heavy-tail depth sampling in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality: at large scale, 63% fewer parameters and 59% less peak decoding memory for a 10% increase in compiled generation latency.

Added

2026-08-29

Creative Commons License
DeepLoop: Depth Scaling for Looped Transformers

DeepLoop: Depth Scaling for Looped Transformers

Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang

OrganizationsPrinceton UniversityUniversity of California, Los Angeles

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formalize this tied-depth effect through a first-order perturbation bound controlled by a visit-alignment coefficient κR\kappa_R. The bound recovers the DeepNorm exponent when visits decorrelate, but in the conservative aligned regime it requires the exponent to increase from 1/41/4 to 1/21/2 as loop count grows at fixed physical depth. The resulting method, \textbf{DeepLoop}, keeps the Post-LN DeepNorm architecture and sets α=(2N)1/2\alpha=(2N)^{1/2} and β=(8N)−1/2\beta=(8N)^{-1/2} for unrolled depth NN. On GPT-style looped language models at GPT-2 small and GPT-2 medium scale, DeepLoop is neutral when no physical block is revisited and improves validation loss and downstream accuracy once recurrent depth is activated. These results show that stable recurrent depth requires residual scaling rules that account for parameter visits, not only nominal layer count.

Added

2026-07-31

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin, Alexander Theus, Thomas Hofmann, Valentina Boeva, T. Konstantin Rusch, Antonio Orvieto

OrganizationsELLIS Institute TübingenETH ZurichLiquid AIMax Planck Institute for Intelligent SystemsSwiss Institute of BioinformaticsTübingen AI CenterUniversité Paris Cité

Why you should read this

Presents FPRM, a Transformer-based Fixed-Point Reasoning Model that leverages fixed-point convergence to dynamically adapt its computational effort to task difficulty in looped architectures, addressing signal propagation issues through architectural modifications.

Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by looping determines the quality of the solution these models find. Like deep architectures, looped architectures are prone to a signal propagation problem induced by depth as the halting decision is postponed. In this paper, we address this signal propagation issue using pre-norm layers and residual scaling. Building on these architectural modifications, we propose FPRM, a Transformer-based Fixed-Point Reasoning Model that uses fixed-point convergence as an end-to-end halting mechanism in a looped architecture. We show that fixed-point halting allows FPRM to adapt its compute to task difficulty. FPRM is effective on common reasoning benchmarks, namely Sudoku, Maze, state-tracking, and ARC-AGI.

Added

2026-06-21

Creative Commons License
Synthesizer: Rethinking Self-Attention for Transformer Models

Synthesizer: Rethinking Self-Attention for Transformer Models

Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, Che Zheng

OrganizationsGoogle

Why you should read this

Proposes Synthesizer, a novel model that achieves competitive or superior performance to traditional Transformers with significantly increased speed and efficiency by synthesizing attention weights without direct token-token interactions, challenging the fundamental reliance on dot product self-attention.

The dot product self-attention is known to be central and indispensable to state-of-the-art Transformer models. But is it really required? This paper investigates the true importance and contribution of the dot product-based self-attention mechanism on the performance of Transformer models. Via extensive experiments, we find that (1) random alignment matrices surprisingly perform quite competitively and (2) learning attention weights from token-token (query-key) interactions is useful but not that important after all. To this end, we propose \textsc{Synthesizer}, a model that learns synthetic attention weights without token-token interactions. In our experiments, we first show that simple Synthesizers achieve highly competitive performance when compared against vanilla Transformer models across a range of tasks, including machine translation, language modeling, text generation and GLUE/SuperGLUE benchmarks. When composed with dot product attention, we find that Synthesizers consistently outperform Transformers. Moreover, we conduct additional comparisons of Synthesizers against Dynamic Convolutions, showing that simple Random Synthesizer is not only 60%60\% faster but also improves perplexity by a relative 3.5%3.5\%. Finally, we show that simple factorized Synthesizers can outperform Linformers on encoding only tasks.

Added

2026-02-13

Creative Commons License