Built independently by an author, for readers. Read the story and support ChapterPal

keyword

sequence modelling

Sequence modeling is a branch of machine learning focused on processing, analyzing, and generating ordered data where the relative positions and contextual relationships of individual elements determine overall meaning. Unlike methods designed for independent and identically distributed data, sequence modeling explicitly accounts for sequential and temporal dependencies, making it essential for tasks such as natural language processing, speech recognition, time-series forecasting, and biological sequence analysis. Modern sequence models, including recurrent neural networks, state-space models, and transformers, process ordered inputs to perform tasks such as predicting subsequent tokens, classifying entire sequences, or translating one sequence into another while managing the computational challenges of capturing both local and long-range dependencies.

1 item

Resurrecting Recurrent Neural Networks for Long Sequences

Resurrecting Recurrent Neural Networks for Long Sequences

Antonio Orvieto, Samuel L. Smith, Albert Gu, Anushan Fernando, Çaglar Gülçehre, Razvan Pascanu, Soham De

OrganizationsETH ZurichGoogle

Why you should read this

Introduces the Linear Recurrent Unit to demonstrate that recurrent neural networks, when designed with linear diagonal recurrences and proper signal propagation, can match both the training speed and long-range modeling accuracy of deep state-space models.

Recurrent Neural Networks (RNNs) offer fast inference on long sequences but are hard to optimize and slow to train. Deep state-space models (SSMs) have recently been shown to perform remarkably well on long sequence modeling tasks, and have the added benefits of fast parallelizable training and RNN-like fast inference. However, while SSMs are superficially similar to RNNs, there are important differences that make it unclear where their performance boost over RNNs comes from. In this paper, we show that careful design of deep RNNs using standard signal propagation arguments can recover the impressive performance of deep SSMs on long-range reasoning tasks, while also matching their training speed. To achieve this, we analyze and ablate a series of changes to standard RNNs including linearizing and diagonalizing the recurrence, using better parameterizations and initializations, and ensuring proper normalization of the forward pass. Our results provide new insights on the origins of the impressive performance of deep SSMs, while also introducing an RNN block called the Linear Recurrent Unit that matches both their performance on the Long Range Arena benchmark and their computational efficiency.

Added

2026-09-28