keyword
Mamba models
Mamba models are a class of deep learning architectures designed for sequence modeling that utilize selective state-space mechanisms to process data efficiently. Developed as an alternative to attention-based Transformer models, Mamba models scale linearly with sequence length rather than quadratically, enabling faster processing and lower memory consumption on long sequences. They operate by making state-space parameters dependent on the input data, which allows the network to dynamically retain relevant information and discard irrelevant details at each step of a sequence. Through a hardware-aware parallel scanning algorithm, Mamba models combine the parallel training capabilities of feedforward architectures with the fast, constant-time autoregressive inference of recurrent neural networks, making them effective for diverse domains including natural language processing, genomics, audio analysis, and computer vision.
3 items

The Hidden Attention of Mamba Models
Ameen Ali, Itamar Zimerman, Lior Wolf
Why you should read this
Reformulates Mamba's selective state-space layers as implicit attention mechanisms, enabling direct theoretical comparisons with transformers and the extraction of attention maps for explainability.
The Mamba layer offers an efficient selective state-space model (SSM) that is highly effective in modeling multiple domains, including NLP, long-range sequence processing, and computer vision. Selective SSMs are viewed as dual models, in which one trains in parallel on the entire sequence via an IO-aware parallel scan, and deploys in an autoregressive manner. We add a third view and show that such models can be viewed as attention-driven models. This new perspective enables us to empirically and theoretically compare the underlying mechanisms to that of the attention in transformers and allows us to peer inside the inner workings of the Mamba model with explainability methods. Our code is publicly available¹.
Added
2026-10-01

The Illusion of State in State-Space Models
William Merrill, Jackson Petty, Ashish Sabharwal
Why you should read this
Proves that popular state-space models like S4 and Mamba share the same fundamental expressive limitations as transformers for sequential state tracking, while identifying a minimal architectural modification to overcome this barrier.
State-space models (SSMs) have emerged as a potential alternative to transformers. One theoretical weakness of transformers is that they cannot express certain kinds of sequential computation and state tracking (Merrill & Sabharwal, 2023a), which SSMs are explicitly designed to address via their close architectural similarity to recurrent neural networks. But do SSMs truly have an advantage (over transformers) in expressive power for state tracking? Surprisingly, the answer is no. Our analysis reveals that the expressive power of S4, Mamba, and related SSMs is limited very similarly to transformers (within TC⁰), meaning these SSMs cannot solve simple state-tracking problems like permutation composition and consequently are provably unable to accurately track chess moves with certain notation, evaluate code, or track entities in a long narrative. To supplement our formal analysis, we report experiments showing that S4 and Mamba indeed struggle with state tracking. Thus, despite their recurrent formulation, the “state” in common SSMs is an illusion: S4, Mamba, and related models have similar expressiveness limitations to non-recurrent models like transformers, which may fundamentally limit their ability to solve real-world state-tracking problems. Moreover, we show that only a minimal change allows SSMs to express and learn state tracking, motivating the development of new, more expressive SSM architectures.
Added
2026-10-01

Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert Gu, Tri Dao
Why you should read this
Introduces Mamba, a selective state space architecture that achieves linear-time sequence scaling and five times higher inference throughput while matching or outperforming standard Transformers across language, audio, and genomics benchmarks.
Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers' computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5 higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pretraining and downstream evaluation.
Added
2026-09-24
