Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert GuTri Dao
Introduces Mamba, a selective state space architecture that achieves linear-time sequence scaling and five times higher inference throughput while matching or outperforming standard Transformers across language, audio, and genomics benchmarks.
Modern foundation models rely almost exclusively on the Transformer architecture, which processes sequences using attention mechanisms. While effective at capturing complex patterns, attention scales quadratically with sequence length and requires storing past context in memory during text generation. This creates severe computational bottlenecks, high deployment costs, and practical limits on processing long sequences. Earlier subquadratic alternatives, such as structured state space models, offered linear scaling but struggled with information-dense and discrete data like language because their fixed, time-invariant dynamics prevented them from selectively focusing on or ignoring specific information based on content.
The article introduces and evaluates Mamba, a sequence modeling architecture based on selective state space models. The objective is to demonstrate that introducing input-dependent selectivity into state space models achieves Transformer-level quality while maintaining linear compute and memory scaling during training and constant time per step during generation.
To accomplish this, the researchers made state space parameters dynamic functions of the input, enabling the model to filter irrelevant noise or retain key context indefinitely. Because this dynamic behavior rules out efficient convolution algorithms, the authors designed a hardware-aware parallel scan that computes recurrences directly in fast processor memory (GPU SRAM) rather than slow main memory (HBM) and recomputes states during training to minimize memory traffic. They integrated this mechanism into a unified neural network block that replaces separate attention and multi-layer perceptron blocks. The architecture was tested across synthetic reasoning tasks, natural language processing up to 3 billion parameters, genomics modeling with sequences up to 1 million elements, and audio generation.
The evaluation yielded several key findings. First, Mamba matched or exceeded the performance of modern Transformer architectures across model sizes while scaling linearly with context length. In language modeling, a 3-billion-parameter Mamba model matched the pretraining and downstream evaluation performance of standard Transformers twice its size. Second, during text generation, Mamba delivered up to five times higher inference throughput than comparable Transformers because it does not require a key-value memory cache. Third, the hardware-aware scan ran up to three times faster than previous methods on modern GPUs and surpassed FlashAttention-2 speed at sequence lengths above 2,000 tokens. Fourth, in long-context domains like DNA and audio pretraining, Mamba continually improved as sequence lengths scaled up to 1 million tokens, whereas prior models plateaued or degraded due to an inability to discard context noise.
These findings suggest that the long-standing trade-off between modeling quality and sequence efficiency can be resolved without attention mechanisms. For technical organizations, deploying selective state space models could dramatically lower inference serving costs, reduce hardware memory footprints, and enable cost-effective analysis of massive contexts in fields like genomics, audio, and document processing. Furthermore, Mamba challenges the assumption that Transformer-style self-attention is essential for high-quality language modeling.
Organizations handling large-scale sequence processing should pilot Mamba implementations for latency-sensitive and long-context applications, particularly in document retrieval, DNA sequence analysis, and audio generation. Engineering teams should also evaluate the open-source codebase to assess integration within existing deployment pipelines. However, further research is required to evaluate Mamba at frontier model scales beyond 7 billion parameters and to thoroughly assess its compatibility with common ecosystem tooling such as instruction fine-tuning, quantization, reinforcement learning from human feedback, and in-context prompting.
While confidence in the experimental findings is high for model scales up to 3 billion parameters, readers should note that tests were confined to these smaller model sizes. Additionally, ablations reveal that while selectivity excels on discrete data like text and DNA, raw continuous signals (such as dense audio waveforms) still benefit from linear time-invariant modeling. System architects should account for this trade-off when selecting model backbones across different data modalities.
- Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). Introduces the structured state space architecture (S4) and continuous-time state-space parameterizations that Mamba builds upon by making the parameters input-dependent.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). Pioneers the hardware-aware, memory-hierarchy-conscious kernel design principles that directly inspired Mamba's hardware-aware selective scan algorithm.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). Establishes the foundational linear attention formulation that reveals the dual recurrent and attention viewpoints for subquadratic sequence models.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Presents the standard Transformer architecture whose quadratic attention bottleneck Mamba specifically aims to replace with selective state spaces.
- Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). Provides key theoretical foundations and kernel approximations for scaling sequence attention mechanisms linearly with sequence length.
- Paper: An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling, Shaojie Bai et al. (2018). Systematically evaluates convolutional versus recurrent designs on long-range sequence modeling tasks, establishing benchmarks targeted by subquadratic models.
- Paper: Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Tri Dao et al. (2024). Unifies Mamba's selective state space mechanism with attention architectures through structured state space duality (SSD), leading to the Mamba-2 architecture.
- Paper: Mamba-3: Improved Sequence Modeling using State Space Principles, Aakash Lahoti et al. (2026). Extends the Mamba sequence modeling line to Mamba-3 by incorporating exponential-trapezoidal discretization, complex-valued state updates, and MIMO inference.
- Paper: VMamba: Visual State Space Model, Yue Liu et al. (2024). Adapts Mamba's selective state-space formulation to 2D computer vision tasks using 2D Selective Scan (SS2D) modules.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). Develops a bidirectional selective state-space visual backbone directly derived from Mamba's linear-time sequence architecture.
- Paper: Test-time regression: a unifying framework for designing sequence models with associative memory, Ke Alexander Wang et al. (2025). Formulates a unified test-time regression framework that theoretically explains why architectures like Mamba excel at associative recall.
- Paper: Dynamic Linear Attention, Xin Wang et al. (2026). Builds on Mamba-based backbones to introduce dynamic, information-aware state merging for long-context sequence modeling.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). Provides a comprehensive architectural survey categorizing modern subquadratic models, including Mamba and its descendants.
