Built independently by an author, for readers. Read the story and support ChapterPal

keyword

linear RNNs

Linear RNNs are a class of recurrent neural network architectures in which the transition from one hidden state to the next is governed by linear transformations rather than non-linear activation functions. By removing non-linearities from the recurrent step, these models can compute representations across entire sequences in parallel during training through efficient associative parallel scans or convolutions, overcoming the sequential training bottleneck of traditional recurrent networks. Non-linear expressive capacity is instead incorporated outside the recurrent loop through external gating mechanisms, feed-forward layers, or non-linear output projections. During autoregressive generation, linear RNNs maintain constant memory per step and linear computational complexity with respect to sequence length, providing an efficient framework that closely connects structured state-space models and linear attention architectures.

5 items

Gated Linear Attention Transformers with Hardware-Efficient Training

Gated Linear Attention Transformers with Hardware-Efficient Training

Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, Yoon Kim

OrganizationsMassachusetts Institute of TechnologyMIT-IBM Watson AI Lab

Why you should read this

Presents an I/O-aware chunkwise parallel algorithm and data-dependent gating mechanism for linear attention that outperforms FlashAttention-2 in training speed while matching standard Transformers and Mamba on language modeling benchmarks.

Transformers with linear attention allow for efficient parallel training but can simultaneously be formulated as an RNN with 2D (matrix-valued) hidden states, thus enjoying linear-time inference complexity. However, linear attention generally underperforms ordinary softmax attention. Moreover, current implementations of linear attention lack I/O-awareness and are thus slower than highly optimized implementations of softmax attention. This work describes a hardware-efficient algorithm for linear attention that trades off memory movement against parallelizability. The resulting implementation, dubbed FLASHLINEARATTENTION, is faster than FLASHATTENTION-2 (Dao, 2023) as a standalone layer even on short sequence lengths (e.g., 1K). We then generalize this algorithm to a more expressive variant of linear attention with data-dependent gates. When used as a replacement for the standard attention layer in Transformers, the resulting gated linear attention (GLA) Transformer is found to perform competitively against the LLaMA-architecture Transformer (Touvron et al., 2023) as well recent linear-time-inference baselines such as RetNet (Sun et al., 2023a) and Mamba (Gu & Dao, 2023) on moderate-scale language modeling experiments. GLA Transformer is especially effective at length generalization, enabling a model trained on 2K to generalize to sequences longer than 20K without significant perplexity degradations. For training speed, the GLA Transformer has higher throughput than a similarly-sized Mamba model.

Added

2026-09-28

Resurrecting Recurrent Neural Networks for Long Sequences

Resurrecting Recurrent Neural Networks for Long Sequences

Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Çaglar Gülçehre, Razvan Pascanu, Soham De

OrganizationsETH ZurichGoogle

Why you should read this

Introduces the Linear Recurrent Unit to demonstrate that recurrent neural networks, when designed with linear diagonal recurrences and proper signal propagation, can match both the training speed and long-range modeling accuracy of deep state-space models.

Recurrent Neural Networks (RNNs) offer fast inference on long sequences but are hard to optimize and slow to train. Deep state-space models (SSMs) have recently been shown to perform remarkably well on long sequence modeling tasks, and have the added benefits of fast parallelizable training and RNN-like fast inference. However, while SSMs are superficially similar to RNNs, there are important differences that make it unclear where their performance boost over RNNs comes from. In this paper, we show that careful design of deep RNNs using standard signal propagation arguments can recover the impressive performance of deep SSMs on long-range reasoning tasks, while also matching their training speed. To achieve this, we analyze and ablate a series of changes to standard RNNs including linearizing and diagonalizing the recurrence, using better parameterizations and initializations, and ensuring proper normalization of the forward pass. Our results provide new insights on the origins of the impressive performance of deep SSMs, while also introducing an RNN block called the Linear Recurrent Unit that matches both their performance on the Long Range Arena benchmark and their computational efficiency.

Added

2026-09-28

Mamba-3: Improved Sequence Modeling using State Space Principles

Mamba-3: Improved Sequence Modeling using State Space Principles

Aakash Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang, Aviv Bick, J. Zico Kolter, Tri Dao, Albert Gu

OrganizationsCarnegie Mellon UniversityCartesia AIPrinceton UniversityTogether AI

Why you should read this

Develops Mamba-3, a novel state-space model that significantly advances the performance-efficiency Pareto frontier for sequence modeling by integrating an expressive recurrence, complex-valued state updates, and MIMO formulation to achieve superior accuracy and efficiency over prior art.

Scaling inference-time compute has emerged as an important driver of LLM performance, making inference efficiency a central focus of model design alongside model quality. While the current Transformer-based models deliver strong model quality, their quadratic compute and linear memory make inference expensive. This has spurred the development of sub-quadratic models with reduced linear compute and constant memory requirements. However, many recent linear models trade off model quality and capability for algorithmic efficiency, failing on tasks such as state tracking. Moreover, their theoretically linear inference remains hardware-inefficient in practice. Guided by an inference-first perspective, we introduce three core methodological improvements inspired by the state space model (SSM) viewpoint of linear models. We combine: (1) a more expressive recurrence derived from SSM discretization, (2) a complex-valued state update rule that enables richer state tracking, and (3) a multi-input, multi-output (MIMO) formulation for better model performance without increasing decode latency. Together with architectural refinements, our Mamba-3 model achieves significant gains across retrieval, state-tracking, and downstream language modeling tasks. At the 1.5B scale, Mamba-3 improves average downstream accuracy by 0.6 percentage points compared to the next best model (Gated DeltaNet), with Mamba-3's MIMO variant further improving accuracy by another 1.2 points for a total 1.8 point gain. Across state-size experiments, Mamba-3 achieves comparable perplexity to Mamba-2 despite using half of its predecessor's state size. Our evaluations demonstrate Mamba-3's ability to advance the performance-efficiency Pareto frontier.

Added

2026-05-07

Creative Commons License