Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Tri DaoAlbert Gu
Establishes a theoretical connection between Transformers and state-space models via structured semiseparable matrices, introducing the Mamba-2 architecture that processes sequences two to eight times faster than Mamba while matching Transformer language modeling performance.
Modern language models rely heavily on attention-based architectures, but these systems demand increasing computational power and memory as text sequences grow longer. Alternative sequence modeling techniques, such as structured state-space approaches, scale much more efficiently by processing data linearly rather than quadratically. However, these alternative architectures have historically developed independently from mainstream models, making them harder to optimize on modern hardware accelerators that are heavily specialized for dense matrix operations.
The article establishes a unified theoretical framework connecting state-space models and attention mechanisms through structured matrix representations. Using this bridge, the authors introduce a refined sequence model, along with hardware-optimized training algorithms and systems techniques designed to match or exceed mainstream model capabilities.
To evaluate this framework, the authors implemented block-decomposed matrix multiplication algorithms and trained a family of models ranging from 125 million to 2.7 billion parameters on standard benchmark text corpora. They benchmarked processing speed across various sequence lengths on standard accelerator hardware and evaluated downstream accuracy across synthetic associative recall and zero-shot reasoning tasks against competitive baselines.
The investigation produced several key findings regarding efficiency and task performance. First, the new algorithm computes sequence transformations 2 to 8 times faster than previous selective state-space implementations and scales to much larger hidden state dimensions without computational slowdown. Second, processing speed surpasses optimized standard attention methods at sequence lengths of 2,000 tokens and is approximately 6 times faster at 16,000 tokens. Third, language models built on this framework achieve equal or superior perplexity and zero-shot task accuracy compared to standard architectures trained under identical token budgets; for instance, a 2.7 billion parameter model outperformed existing open-source baselines that were more than double its size. Finally, hybrid architectures that incorporate a small fraction of attention layers—around 10%—yield higher accuracy than either pure attention models or pure state-space models.
These findings indicate that organizations can significantly reduce training and inference compute costs on long-context workloads without sacrificing model performance. The shared mathematical formulation also allows developers to port mature scaling techniques, such as multi-device parallelization strategies, directly to state-space architectures.
Engineering teams deploying large language models should consider evaluating these hybrid architectures for long-context tasks to balance throughput and memory constraints. For future initiatives, practitioners should explore integrating advanced matrix decompositions and testing these models on larger scaling regimes beyond 3 billion parameters.
The primary limitation of this evaluation is that tests were conducted up to moderate model scales and standard pretraining datasets; behavior at frontier scales of tens or hundreds of billions of parameters remains unverified. Readers can place high confidence in the demonstrated speedups and core benchmarks, though careful pilot evaluation is recommended when adapting non-standard head patterns to specific production environments.
- Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). This foundational work introduces structured state-space models (S4) and normal plus low-rank parameterizations that provide the essential theoretical grounding for the State Space Duality framework.
- Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). Reading this paper provides necessary context on framing linear attention as recurrent state updates, a perspective central to establishing the duality between Transformers and SSMs.
- Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). This work establishes kernel-based linear-complexity attention approximations that directly inform how linear attention connects to continuous and recurrent state spaces.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). This landmark paper defines the standard multi-head self-attention mechanism whose mathematical equivalence to structured state space models is formalized in the source.
- Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). This paper introduces multi-query attention, establishing key-value sharing principles leveraged in designing hardware-efficient Mamba-2 architectures.
- Paper: Mamba-3: Improved Sequence Modeling using State Space Principles, Aakash Lahoti et al. (2026). This work directly builds upon Mamba-2 by incorporating exponential-trapezoidal discretization, complex-valued state updates, and MIMO designs to advance state-space architectures.
- Paper: Dynamic Linear Attention, Xin Wang et al. (2026). This work directly adopts Mamba-2 as a foundational linear-attention backbone and improves its multi-state memory retention via dynamic token-merging rules.
- Paper: Test-time regression: a unifying framework for designing sequence models with associative memory, Ke Alexander Wang et al. (2025). This paper expands on the duality between state-space models and attention by casting both mechanisms within a broader test-time regression framework for associative memory.
- Paper: Titans: Learning to Memorize at Test Time, Ali Behrouz et al. (2024). This paper benchmarks against Mamba-2 and introduces learnable test-time neural memory modules to address historical compression limitations in recurrent state-space models.
- Paper: VMamba: Visual State Space Model, Yue Liu et al. (2024). This work extends 1D selective state-space modeling concepts into visual backbones via a 2D selective scan module for computer vision tasks.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). This paper applies selective state-space sequence modeling bidirectionally to replace self-attention in high-resolution visual representation learning.
- Paper: Fast Weight Attention for Continual Learning, Yifan Zhang et al. (2026). This work advances beyond selective SSMs and linear attention by proposing explicit read-after-write causal state updates for continual learning.
- Paper: End-to-End Test-Time Training for Long Context, Arnuv Tandon et al. (2025). This paper evaluates long-context scaling limitations of Mamba-2 and proposes an end-to-end test-time training alternative to recurrent sequence modeling.
- Paper: Speed Always Wins: A Survey on Efficient Architectures for Large Language Models, Weigao Sun et al. (2025). This comprehensive survey categorizes the paradigm shift toward sub-quadratic sequence models, extensively analyzing Mamba and linear state-space models.
