Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Deep Linear Recurrent Unit

A Deep Linear Recurrent Unit is a recurrent neural network architecture designed for sequence modeling that removes nonlinear activation functions from its internal recurrent state transitions in favor of purely linear operations. By replacing traditional nonlinear recurrent gates with linear, diagonalized recurrence mechanisms often parameterized in the complex domain, it enables parallelized training across long sequences using associative scans while maintaining fast, step-by-step inference. Non-linear expressive capacity is preserved by interleaving feedforward layers throughout the deep network rather than embedding them directly within the recurrence loop. Coupled with stable parameter initialization and normalization techniques that control the magnitude of recurrent eigenvalues, this design mitigates vanishing and exploding gradients and allows the model to process long-range dependencies efficiently.

1 item

Resurrecting Recurrent Neural Networks for Long Sequences

Resurrecting Recurrent Neural Networks for Long Sequences

Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Çaglar Gülçehre, Razvan Pascanu, Soham De

OrganizationsETH ZurichGoogle

Why you should read this

Introduces the Linear Recurrent Unit to demonstrate that recurrent neural networks, when designed with linear diagonal recurrences and proper signal propagation, can match both the training speed and long-range modeling accuracy of deep state-space models.

Recurrent Neural Networks (RNNs) offer fast inference on long sequences but are hard to optimize and slow to train. Deep state-space models (SSMs) have recently been shown to perform remarkably well on long sequence modeling tasks, and have the added benefits of fast parallelizable training and RNN-like fast inference. However, while SSMs are superficially similar to RNNs, there are important differences that make it unclear where their performance boost over RNNs comes from. In this paper, we show that careful design of deep RNNs using standard signal propagation arguments can recover the impressive performance of deep SSMs on long-range reasoning tasks, while also matching their training speed. To achieve this, we analyze and ablate a series of changes to standard RNNs including linearizing and diagonalizing the recurrence, using better parameterizations and initializations, and ensuring proper normalization of the forward pass. Our results provide new insights on the origins of the impressive performance of deep SSMs, while also introducing an RNN block called the Linear Recurrent Unit that matches both their performance on the Long Range Arena benchmark and their computational efficiency.

Added

2026-09-28