Resurrecting Recurrent Neural Networks for Long Sequences
Antonio OrvietoSamuel L. SmithAlbert GuAnushan FernandoÇaglar GülçehreRazvan PascanuSoham De
Introduces the Linear Recurrent Unit to demonstrate that recurrent neural networks, when designed with linear diagonal recurrences and proper signal propagation, can match both the training speed and long-range modeling accuracy of deep state-space models.
Modern sequence modeling increasingly demands architectures that can process very long contexts efficiently. While standard Recurrent Neural Networks (RNNs) offer fast, linear-time inference, they historically suffer from vanishing and exploding gradients and slow, sequential training. Attention-based Transformers resolve optimization bottlenecks but incur quadratic computational and memory costs, making them expensive on long sequences. Recent state-space models (SSMs) such as S4 overcome these challenges through continuous-time differential equations, specialized mathematical initializations, and discretization frameworks. The article investigates whether standard deep RNNs can match the high performance and computational efficiency of deep SSMs on long-sequence reasoning without relying on complex continuous-time theory.
The article evaluates a systematic redesign of vanilla deep RNNs, demonstrating that standard signal propagation, parameterization, and normalization techniques can recover SSM-level capability. The authors conducted rigorous empirical experiments across all six benchmarks of the Long Range Arena (LRA), which test long-range sequence understanding across tasks such as sequential image classification, mathematical list operations, text processing, document retrieval, and visual reasoning up to 16,000 steps. They also supported their empirical ablations with theoretical spectral analysis, random matrix theory, and dynamical systems theory, tracking both accuracy and wall-clock training speed.
The key findings reveal several insights for model design. First, eliminating nonlinearities from the recurrent state update (linear recurrences) significantly improves test accuracy over standard hyperbolic tangent and rectified linear activations, while non-recurrent feed-forward layers sufficiently preserve expressivity. Second, diagonalizing the recurrence with complex values enables associative parallel scans, delivering training speeds up to 29 times faster than standard RNNs and matching state-of-the-art SSM speeds. Third, using a stable exponential parameterization allows eigenvalues to be safely initialized close to the unit circle boundary, mitigating vanishing gradients and boosting accuracy on difficult tasks like Pathfinder to above 93%. Fourth, introducing a dedicated input-scaling normalization factor prevents forward-pass activation blow-up when eigenvalues approach magnitude one; combining this normalization with a restricted eigenvalue phase at initialization enables the model to reach 94.2% accuracy on PathX, the hardest LRA task.
These findings imply that the breakthrough performance of recent deep state-space models does not stem fundamentally from differential equation discretization or structured polynomial initializations. Instead, their success arises from linear recurrences, complex diagonal parameterizations, eigenvalue stability, and forward-pass normalization. By implementing these principles directly, the authors introduce the Linear Recurrent Unit (LRU), a simpler recurrent block that achieves equivalent accuracy and efficiency without the theoretical overhead or parameter sharing of continuous-time SSMs.
For future architectural development, engineering teams should consider adopting Linear Recurrent Units as streamlined, drop-in alternatives to attention mechanisms and SSMs for long-sequence tasks. Practitioners should leverage parallel associative scans for accelerated training and enforce stable exponential parameterization with normalization when long-range context is required. Although empirical confidence is high across the evaluated Long Range Arena benchmarks, future work should validate the scalability and generalization of LRUs on broader real-world applications, such as large-scale natural language generation, audio modeling, and production-scale time-series forecasting.
- Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). Introduces the structured state-space model (S4) and sets the benchmark performance on the Long Range Arena that the source directly analyzes and seeks to match using simplified linear RNNs.
- Paper: On the difficulty of training recurrent neural networks, Razvan Pascanu et al. (2012). Provides the foundational dynamical systems analysis of vanishing and exploding gradients in recurrent neural networks that motivates the source's signal propagation and initialization techniques.
- Paper: A Simple Way to Initialize Recurrent Networks of Rectified Linear Units, Quoc V. Le et al. (2015). Demonstrates how recurrent weight initialization to the identity matrix allows linear and rectified activations to retain long-range memory, a core principle refined by the source's Linear Recurrent Unit.
- Paper: Layer Normalization, Jimmy Lei Ba et al. (2016). Establishes normalization techniques across hidden features in recurrent networks, providing the basis for the forward-pass normalization strategies used in the source.
- Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Presents the foundational gating and constant error carousel mechanism designed to overcome the long-sequence gradient failure modes discussed throughout the source.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). Extends linear time-invariant state-space and recurrent formulations by introducing input-dependent selective state updates with hardware-efficient parallel scans.
- Paper: Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Tri Dao et al. (2024). Unifies linear recurrent state-space models and attention through structured state space duality to achieve hardware-efficient training at scale.
- Paper: Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling, Liliang Ren et al. (2025). Combines selective linear recurrent state-space layers with sliding window attention to enable infinite-context sequence modeling.
- Paper: Test-time regression: a unifying framework for designing sequence models with associative memory, Ke Alexander Wang et al. (2025). Provides a unifying theoretical framework showing how linear state-space models and attention variants perform test-time regression for associative memory.
- Paper: Dynamic Linear Attention, Xin Wang et al. (2026). Builds upon efficient linear recurrence and state-space architectures by implementing dynamic, information-aware memory state merging.
