Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Mamba models

Mamba models are a class of deep learning architectures designed for sequence modeling that utilize selective state-space mechanisms to process data efficiently. Developed as an alternative to attention-based Transformer models, Mamba models scale linearly with sequence length rather than quadratically, enabling faster processing and lower memory consumption on long sequences. They operate by making state-space parameters dependent on the input data, which allows the network to dynamically retain relevant information and discard irrelevant details at each step of a sequence. Through a hardware-aware parallel scanning algorithm, Mamba models combine the parallel training capabilities of feedforward architectures with the fast, constant-time autoregressive inference of recurrent neural networks, making them effective for diverse domains including natural language processing, genomics, audio analysis, and computer vision.

3 items

The Illusion of State in State-Space Models

The Illusion of State in State-Space Models

William Merrill, Jackson Petty, Ashish Sabharwal

OrganizationsAllen Institute for AINew York University

Why you should read this

Proves that popular state-space models like S4 and Mamba share the same fundamental expressive limitations as transformers for sequential state tracking, while identifying a minimal architectural modification to overcome this barrier.

State-space models (SSMs) have emerged as a potential alternative to transformers. One theoretical weakness of transformers is that they cannot express certain kinds of sequential computation and state tracking (Merrill & Sabharwal, 2023a), which SSMs are explicitly designed to address via their close architectural similarity to recurrent neural networks. But do SSMs truly have an advantage (over transformers) in expressive power for state tracking? Surprisingly, the answer is no. Our analysis reveals that the expressive power of S4, Mamba, and related SSMs is limited very similarly to transformers (within TC⁰), meaning these SSMs cannot solve simple state-tracking problems like permutation composition and consequently are provably unable to accurately track chess moves with certain notation, evaluate code, or track entities in a long narrative. To supplement our formal analysis, we report experiments showing that S4 and Mamba indeed struggle with state tracking. Thus, despite their recurrent formulation, the “state” in common SSMs is an illusion: S4, Mamba, and related models have similar expressiveness limitations to non-recurrent models like transformers, which may fundamentally limit their ability to solve real-world state-tracking problems. Moreover, we show that only a minimal change allows SSMs to express and learn state tracking, motivating the development of new, more expressive SSM architectures.

Added

2026-10-01

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Albert Gu, Tri Dao

OrganizationsCarnegie Mellon UniversityPrinceton University

Why you should read this

Introduces Mamba, a selective state space architecture that achieves linear-time sequence scaling and five times higher inference throughput while matching or outperforming standard Transformers across language, audio, and genomics benchmarks.

Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers' computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5×\times higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pretraining and downstream evaluation.

Added

2026-09-24