Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Tri DaoAlbert Gu

article2024ICML1,938 citations

Establishes a theoretical connection between Transformers and state-space models via structured semiseparable matrices, introducing the Mamba-2 architecture that processes sequences two to eight times faster than Mamba while matching Transformer language modeling performance.

Listen

Modern language models rely heavily on attention-based architectures, but these systems demand increasing computational power and memory as text sequences grow longer. Alternative sequence modeling techniques, such as structured state-space approaches, scale much more efficiently by processing data linearly rather than quadratically. However, these alternative architectures have historically developed independently from mainstream models, making them harder to optimize on modern hardware accelerators that are heavily specialized for dense matrix operations.

The article establishes a unified theoretical framework connecting state-space models and attention mechanisms through structured matrix representations. Using this bridge, the authors introduce a refined sequence model, along with hardware-optimized training algorithms and systems techniques designed to match or exceed mainstream model capabilities.

To evaluate this framework, the authors implemented block-decomposed matrix multiplication algorithms and trained a family of models ranging from 125 million to 2.7 billion parameters on standard benchmark text corpora. They benchmarked processing speed across various sequence lengths on standard accelerator hardware and evaluated downstream accuracy across synthetic associative recall and zero-shot reasoning tasks against competitive baselines.

The investigation produced several key findings regarding efficiency and task performance. First, the new algorithm computes sequence transformations 2 to 8 times faster than previous selective state-space implementations and scales to much larger hidden state dimensions without computational slowdown. Second, processing speed surpasses optimized standard attention methods at sequence lengths of 2,000 tokens and is approximately 6 times faster at 16,000 tokens. Third, language models built on this framework achieve equal or superior perplexity and zero-shot task accuracy compared to standard architectures trained under identical token budgets; for instance, a 2.7 billion parameter model outperformed existing open-source baselines that were more than double its size. Finally, hybrid architectures that incorporate a small fraction of attention layers—around 10%—yield higher accuracy than either pure attention models or pure state-space models.

These findings indicate that organizations can significantly reduce training and inference compute costs on long-context workloads without sacrificing model performance. The shared mathematical formulation also allows developers to port mature scaling techniques, such as multi-device parallelization strategies, directly to state-space architectures.

Engineering teams deploying large language models should consider evaluating these hybrid architectures for long-context tasks to balance throughput and memory constraints. For future initiatives, practitioners should explore integrating advanced matrix decompositions and testing these models on larger scaling regimes beyond 3 billion parameters.

The primary limitation of this evaluation is that tests were conducted up to moderate model scales and standard pretraining datasets; behavior at frontier scales of tens or hundreds of billions of parameters remains unverified. Readers can place high confidence in the demonstrated speedups and core benchmarks, though careful pilot evaluation is recommended when adapting non-standard head patterns to specific production environments.

  • Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). This foundational work introduces structured state-space models (S4) and normal plus low-rank parameterizations that provide the essential theoretical grounding for the State Space Duality framework.
  • Paper: Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention, Angelos Katharopoulos et al. (2020). Reading this paper provides necessary context on framing linear attention as recurrent state updates, a perspective central to establishing the duality between Transformers and SSMs.
  • Paper: Rethinking Attention with Performers, Krzysztof Choromanski et al. (2021). This work establishes kernel-based linear-complexity attention approximations that directly inform how linear attention connects to continuous and recurrent state spaces.
  • Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). This landmark paper defines the standard multi-head self-attention mechanism whose mathematical equivalence to structured state space models is formalized in the source.
  • Paper: Fast Transformer Decoding: One Write-Head is All You Need, Noam Shazeer (2019). This paper introduces multi-query attention, establishing key-value sharing principles leveraged in designing hardware-efficient Mamba-2 architectures.
Cover for Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Abstract

While Transformers have been the main architecture behind deep learning's success in language modeling, state-space models (SSMs) such as Mamba have recently been shown to match or outperform Transformers at small to medium scale. We show that these families of models are actually quite closely related, and develop a rich framework of theoretical connections between SSMs and variants of attention, connected through various decompositions of a well-studied class of structured semiseparable matrices. Our state space duality (SSD) framework allows us to design a new architecture (Mamba-2) whose core layer is an a refinement of Mamba's selective SSM that is 2-8X faster, while continuing to be competitive with Transformers on language modeling.

Table of Contents

  • 1 Introduction
  • 2 Background and Overview
  • 2.1 Structured State Space Models
  • 2.2 Attention
  • 2.3 Structured Matrices
  • 2.4 Overview: Structured State Space Duality
  • 2.5 Notation
  • 3 State Space Models are Structured Matrices
  • 3.1 The Matrix Transformation Form of State Space Models
  • 3.2 Semiseparable Matrices
  • 3.2.1 The Sequentially Semiseparable (SSS) Representation
  • 3.2.2 1-Semiseparable Matrices: the Scalar SSM Recurrence
  • 3.3 State Space Models are Semiseparable Matrices
  • 3.4 Computing State Space Models through Structured Matrix Algorithms
  • 3.4.1 The Linear (Recurrent) Mode
  • 3.4.2 The Quadratic (Naive) Mode
  • 3.4.3 Summary
  • 4 Structured Masked Attention: Generalizing Linear Attention with Structured Matrices
  • 4.1 The Attention Framework
  • 4.1.1 Attention
  • 4.1.2 Self-Attention
  • 4.1.3 Kernel Attention
  • 4.1.4 Masked (Kernel) Attention
  • 4.2 Linear Attention
  • 4.2.1 A Tensor Contraction Proof of Linear Attention
  • 4.3 Structured Masked Attention
  • 4.3.1 Summary: The Dual Forms of Masked Attention
  • 5 State Space Duality
  • 5.1 Scalar-Identity Structured State Space Models
  • 5.2 1-Semiseparable Structured Masked Attention
  • 5.3 Structured State-Space Duality (SSD)
  • 6 A Hardware-Efficient Algorithm for SSD Models
  • 6.1 Diagonal Blocks
  • 6.2 Low-Rank Blocks
  • 6.3 Computational Cost
  • 7 The Mamba-2 Architecture
  • 7.1 Block Design
  • 7.2 Multihead Patterns for Sequence Transformations
  • 7.3 Other SSD Extensions from Linear Attention
  • 8 Systems Optimization for SSMs
  • 8.1 Tensor Parallel
  • 8.2 Sequence Parallelism
  • 8.3 Variable Length
  • 9 Empirical Validation
  • 9.1 Synthetics: Associative Recall
  • 9.2 Language Modeling
  • 9.2.1 Scaling Laws
  • 9.2.2 Downstream Evaluations
  • 9.2.3 Hybrid Models: Combining SSD Layer with MLP and Attention
  • 9.3 Speed Benchmarks
  • 9.4 Architecture Ablations
  • 9.4.1 Block Design
  • 9.4.2 Head Structure
  • 9.4.3 Attention Kernel Approximations
  • 10 Related Work and Discussion
  • 10.1 State Space Models
  • 10.2 Structured Matrices
  • 10.3 (Linear) Attention
  • 10.4 Related Models
  • 11 Conclusion
  • References
  • A Glossary
  • B Efficient Algorithms for the Scalar SSM Scan (1-SS Multiplication)
  • B.1 Problem Definition
  • B.2 Classical Algorithms
  • B.2.1 Sequential Recurrence
  • B.2.2 Parallel Associative Scan
  • B.3 Efficient Algorithms via Structured Matrix Decompositions
  • B.3.1 Dilated Mode
  • B.3.2 State-Passing (Chunkwise) Mode
  • B.3.3 Fully Recurrent Mode
  • B.3.4 (Parallel) Block Decomposition Mode
  • B.3.5 Associative Scan Mode
  • C Theory Details
  • C.1 Extras: Closure Properties of SSMs
  • C.2 Autoregressive Masked Attention is Semiseparable-Structured Attention
  • D Experimental Details
  • D.1 MQAR Details
  • D.2 Scaling Law Details
  • D.3 Downstream Evaluation Details
  • D.4 Ablation Details

Knowls

  1. Knowl 1 — Structured State Space Duality

    theoretical result

    Structured State Space Duality (SSD) establishes an exact mathematical equivalence between a scalar-identity structured state-space model (SSM) and a 1-semiseparable structured masked attention (1-SS SMA) mechanism.

    Let X∈RT×PX \in \mathbb{R}^{T \times P} denote an input sequence of length TT and head dimension PP, A=(a0,a1,…,aT−1)∈[0,1]TA = (a_0, a_1, \dots, a_{T-1}) \in [0, 1]^T be an input-dependent scalar decay sequence, and B,C∈RT×NB, C \in \mathbb{R}^{T \times N} be input and output projection matrices with state expansion dimension NN.

    The Linear (Recurrent) Form evolves an implicit recurrent state ht∈RN×Ph_t \in \mathbb{R}^{N \times P}: ht=atht−1+BtXt⊤h_t = a_t h_{t-1} + B_t X_t^\top Yt=Ct⊤htY_t = C_t^\top h_t with initial condition h−1=0h_{-1} = 0. This computation scales with time complexity O(TNP)O(TNP) and constant memory footprint per step during autoregressive generation.

    The Dual (Quadratic Attention) Form computes the sequence transformation as a masked kernel attention operation without softmax: Y=MX=(L∘(CB⊤))XY = M X = (L \circ (C B^\top)) X where ∘\circ denotes the Hadamard (elementwise) product, CB⊤∈RT×TC B^\top \in \mathbb{R}^{T \times T} is the attention Gram matrix, and L∈RT×TL \in \mathbb{R}^{T \times T} is a 1-semiseparable causal mask defined by: Lji={∏s=i+1jas=aj:i×if j≥i0if j<iL_{ji} = \begin{cases} \prod_{s=i+1}^j a_s = a_{j:i}^\times & \text{if } j \ge i \\ 0 & \text{if } j < i \end{cases}

    This duality shows that linear-time selective state space models and quadratic-time attention are dual representations of matrix multiplication by semiseparable structured matrices.

  2. Knowl 2 — Hardware-Efficient Block-Decomposition SSD Algorithm

    algorithm

    The SSD algorithm evaluates the state space dual transformation Y=SSD(X,A,B,C)Y = \text{SSD}(X, A, B, C) by partitioning the sequence length TT into chunks of size QQ (where typically Q=N=PQ = N = P, matching chunk size, state dimension, and head dimension). It combines intra-chunk quadratic attention computations with an inter-chunk linear scalar SSM recurrence, allowing the workload to be dominated by batched matrix multiplications (BMMs) executed on hardware tensor cores.

    Input: Input sequence X∈RT×PX \in \mathbb{R}^{T \times P}, decay parameters A∈RTA \in \mathbb{R}^T, expansion parameters B∈RT×NB \in \mathbb{R}^{T \times N}, contraction parameters C∈RT×NC \in \mathbb{R}^{T \times N}, chunk size QQ
    Output: Transformed sequence Y∈RT×PY \in \mathbb{R}^{T \times P}, final recurrent state hfinal∈RN×Ph_{\text{final}} \in \mathbb{R}^{N \times P}
    Partition inputs into T/QT/Q chunks of length QQ: X(c),A(c),B(c),C(c)X^{(c)}, A^{(c)}, B^{(c)}, C^{(c)} for c=0,…,T/Q−1c = 0, \dots, T/Q - 1
    Compute intra-chunk cumulative products of AA for each chunk cc
    for each chunk c=0,…,T/Q−1c = 0, \dots, T/Q - 1 in parallel do
        Construct 1-SS chunk mask L(c)∈RQ×QL^{(c)} \in \mathbb{R}^{Q \times Q} where Lji(c)=∏s=i+1jas(c)L^{(c)}_{ji} = \prod_{s=i+1}^j a^{(c)}_s for j≥ij \ge i
        Compute intra-chunk diagonal outputs: Ydiag(c)=(L(c)∘(C(c)(B(c))⊤))X(c)Y_{\text{diag}}^{(c)} = (L^{(c)} \circ (C^{(c)} (B^{(c)})^\top)) X^{(c)}
        Compute right factor chunk state: S(c)=(B(c))⊤(decayed X(c))∈RN×PS^{(c)} = (B^{(c)})^\top (\text{decayed } X^{(c)}) \in \mathbb{R}^{N \times P}
    end for
    Initialize h0=0∈RN×Ph_0 = 0 \in \mathbb{R}^{N \times P}
    for c=0c = 0 to T/Q−1T/Q - 1 do
        Compute total chunk state: hc+1=(∏t=0Q−1at(c))hc+S(c)h_{c+1} = (\prod_{t=0}^{Q-1} a^{(c)}_t) h_c + S^{(c)}
    end for
    hfinal=hT/Qh_{\text{final}} = h_{T/Q}
    for each chunk c=0,…,T/Q−1c = 0, \dots, T/Q - 1 in parallel do
        Compute inter-chunk output from incoming state hch_c: Yoff(c)=C(c)(decayed hc)Y_{\text{off}}^{(c)} = C^{(c)} (\text{decayed } h_c)
        Combine outputs: Y(c)=Ydiag(c)+Yoff(c)Y^{(c)} = Y_{\text{diag}}^{(c)} + Y_{\text{off}}^{(c)}
    end for
    Assemble chunks Y=[Y(0),…,Y(T/Q−1)]Y = [Y^{(0)}, \dots, Y^{(T/Q-1)}]
    return Y,hfinalY, h_{\text{final}}

    For N=P=QN = P = Q, this algorithm requires O(TN2)O(TN^2) training FLOPs, O(TN)O(TN) inference FLOPs, O(TN)O(TN) training memory, and O(N2)O(N^2) inference memory.

  3. Knowl 3 — Equivalence of Structured SSMs and Semiseparable Matrices

    theoretical result

    Every discrete linear state space model (SSM) mapping an input sequence x∈RTx \in \mathbb{R}^T to y∈RTy \in \mathbb{R}^T via the recurrence: ht=Atht−1+Btxth_t = A_t h_{t-1} + B_t x_t yt=Ct⊤hty_t = C_t^\top h_t with state size NN, transition matrices At∈RN×NA_t \in \mathbb{R}^{N \times N}, input vectors Bt∈RN×1B_t \in \mathbb{R}^{N \times 1}, and output vectors Ct∈RN×1C_t \in \mathbb{R}^{N \times 1} is mathematically identical to a matrix multiplication y=Mxy = M x, where M∈RT×TM \in \mathbb{R}^{T \times T} is a lower-triangular NN-semiseparable matrix in sequentially semiseparable (SSS) form: Mji=Cj⊤AjAj−1⋯Ai+1Bi(j≥i)M_{ji} = C_j^\top A_j A_{j-1} \cdots A_{i+1} B_i \quad (j \ge i)

    Conversely, every lower-triangular NN-semiseparable matrix (defined as having every submatrix on or below the diagonal possessing rank at most NN) admits an NN-SSS representation. Consequently, any general SSM of state size NN on sequence length TT can be computed in O(TN)O(TN) time and space.

  4. Knowl 4 — Structured Masked Attention

    definition

    Structured Masked Attention (SMA) generalizes linear and masked kernel attention by expressing sequence transformations as a 4-way tensor contraction over queries Q∈RT×NQ \in \mathbb{R}^{T \times N}, keys K∈RS×NK \in \mathbb{R}^{S \times N}, values V∈RS×PV \in \mathbb{R}^{S \times P}, and a structured mask matrix L∈RT×SL \in \mathbb{R}^{T \times S}: Y=contract(TN,SN,SP,TS→TP)(Q,K,V,L)Y = \text{contract}(TN, SN, SP, TS \to TP)(Q, K, V, L)

    SMA admits two dual execution modes depending on the contraction order:

    1. Quadratic Attention Mode: Computes Gram matrix G=QK⊤G = Q K^\top, applies mask M=G∘LM = G \circ L, and computes Y=MVY = M V. This involves O(T2N)O(T^2 N) operations when S=TS = T.
    2. Linear Recurrent Mode: Expands input into state features Z=V⊗KZ = V \otimes K, applies the structured matrix transformation H=LZH = L Z, and contracts with queries Y=QHY = Q H. When LL has subquadratic (or linear) matrix-vector multiplication algorithms, this executes with O(TNP)O(T N P) complexity.
  5. Knowl 5 — Autoregressive Masked Attention Characterization Theorem

    theoretical result

    Let L∈RT×TL \in \mathbb{R}^{T \times T} be a causal attention mask defining an autoregressive linear transformation y=Lxy = L x of bounded order kk, such that each output yty_t depends only on the current input and the preceding kk outputs: yt=μtxt+ℓt,1yt−1+⋯+ℓt,kyt−ky_t = \mu_t x_t + \ell_{t,1} y_{t-1} + \dots + \ell_{t,k} y_{t-k} Then the matrix LL is necessarily an order-(k+1)(k+1) semiseparable matrix and can be realized as a state space model of order k+1k+1.

    This proves that any causal masked kernel attention mechanism that admits an efficient, constant-step-time autoregressive recurrence must be an instance of semiseparable structured masked attention.

  6. Knowl 6 — Mamba-2 Block Architecture

    model/method

    The Mamba-2 block modifies the Mamba-1 layer to improve scalability, hardware efficiency, and compatibility with tensor parallelism:

    1. Parallel Parameter Projections: The layer input u∈RL×du \in \mathbb{R}^{L \times d} is projected at the beginning of the block in parallel to produce the sequence mixer input X∈RL×edX \in \mathbb{R}^{L \times e d}, multiplicative gate z∈RL×edz \in \mathbb{R}^{L \times e d}, decay parameters AA (or Δ\Delta), and SSM projection matrices B,CB, C, eliminating sequential linear dependencies.
    2. Short 1D Convolution: A depthwise 1D convolution of small kernel width is applied along the sequence dimension independently per channel to XX.
    3. Core SSD Layer: The SSD sequence transformation computes Y=SSD(Xc,A,B,C)Y = \text{SSD}(X_c, A, B, C) across heads.
    4. Gating and Normalization: The mixer output is gated elementwise by the SiLU-activated branch zz: Yg=Y⋅SiLU(z)Y_g = Y \cdot \text{SiLU}(z), followed by a normalization layer (GroupNorm or RMSNorm) placed before the final linear projection W(o)∈Red×dW^{(o)} \in \mathbb{R}^{e d \times d}.
  7. Knowl 7 — Multihead and Multi-Value Patterns for SSMs

    definition

    For a multihead sequence transformation with model dimension DD, head dimension PP, number of heads H=D/PH = D/P, and state expansion dimension NN, parameter-sharing configurations across heads define specific head patterns:

    • Multi-Head SSM (MHS) (analogous to Multi-Head Attention, MHA): X∈RT×H×PX \in \mathbb{R}^{T \times H \times P}, A∈RT×HA \in \mathbb{R}^{T \times H}, B∈RT×H×NB \in \mathbb{R}^{T \times H \times N}, C∈RT×H×NC \in \mathbb{R}^{T \times H \times N}.
    • Multi-Contract SSM (MCS) (analogous to Multi-Query Attention, MQA): X∈RT×1×PX \in \mathbb{R}^{T \times 1 \times P}, A∈RT×HA \in \mathbb{R}^{T \times H}, B∈RT×1×NB \in \mathbb{R}^{T \times 1 \times N}, C∈RT×H×NC \in \mathbb{R}^{T \times H \times N}.
    • Multi-Expand SSM (MES) (analogous to Multi-Key Attention, MKA): X∈RT×1×PX \in \mathbb{R}^{T \times 1 \times P}, A∈RT×HA \in \mathbb{R}^{T \times H}, B∈RT×H×NB \in \mathbb{R}^{T \times H \times N}, C∈RT×1×NC \in \mathbb{R}^{T \times 1 \times N}.
    • Multi-Input SSM (MIS) (analogous to Multi-Value Attention, MVA): X∈RT×H×PX \in \mathbb{R}^{T \times H \times P}, A∈RT×HA \in \mathbb{R}^{T \times H}, B∈RT×1×NB \in \mathbb{R}^{T \times 1 \times N}, C∈RT×1×NC \in \mathbb{R}^{T \times 1 \times N}.
    • Grouped-Value Attention (GVA) (or Grouped-Input SSM, GIS): BB and CC share GG heads (1<G<H1 < G < H), while XX has HH heads.

    In language modeling ablations, the Multi-Value / Multi-Input pattern consistently yields lower perplexity than MQA or MKA when controlling for parameter count and state size.

  8. Knowl 8 — Tensor, Sequence, and Variable-Length Parallelism in Mamba-2

    model/method

    Mamba-2 enables distributed systems optimizations adapted from Transformer training frameworks:

    • Tensor Parallelism (TP): By generating projections for (X,z,Δ,B,C)(X, z, \Delta, B, C) directly from block input uu and applying GroupNorm with the number of groups divisible by the TP degree, input and output projection matrices are sharded across GPUs with only a single all-reduce communication step per block (matching the communication frequency of Transformer MLP/attention blocks).
    • Sequence / Context Parallelism (CP): Long sequences are partitioned across GPUs along the sequence length axis. Each GPU computes its local chunk SSD transformation and passes only the boundary recurrent state h∈RN×Ph \in \mathbb{R}^{N \times P} sequentially to the next GPU, requiring communication bandwidth that scales linearly with worker count (O(1)O(1) transfer size per boundary).
    • Variable-Length Processing: Sequences of different lengths are concatenated into a single flat sequence without zero-padding by setting the transition scalar at=0a_t = 0 at document boundaries, resetting the recurrent state between independent sequences.
  9. Knowl 9 — Hybrid Architectures Combining SSD, Attention, and MLP Layers

    empirical result

    Combining SSD layers with a small proportion of softmax attention layers and gated MLP layers achieves superior language modeling perplexity compared to pure SSM or pure Transformer architectures.

    In experiments on a 48-layer, 350M parameter model trained on 7B tokens from the Pile:

    • A pure Mamba-2 model (0 attention layers) achieves a validation perplexity of 8.60.
    • A pure Transformer++ baseline achieves a perplexity of 8.68.
    • Hybrid models allocating approximately 10% of layers to attention achieve the lowest perplexity (e.g., 6 attention layers achieve 8.26 perplexity).

    At the 2.7B parameter scale (64 layers trained on 300B tokens):

    • Transformer++ (32 attention, 32 MLP): 6.13 Pile ppl, 60.2% downstream zero-shot average.
    • Mamba-2 (64 SSD layers): 6.09 Pile ppl, 60.2% downstream zero-shot average.
    • Mamba-2-Attention (58 SSD, 6 attention): 5.95 Pile ppl, 61.0% downstream zero-shot average.
    • Mamba-2-MLP-Attention (28 SSD, 4 attention, 32 MLP): 6.00 Pile ppl, 60.7% downstream zero-shot average.
  10. Knowl 10 — Downstream Zero-Shot Evaluations of Mamba-2

    data/table

    Mamba-2 models evaluated zero-shot across standard NLP benchmarks match or outperform Mamba-1 and open-source Transformer baselines (e.g., Pythia) trained on 300B tokens of the Pile dataset using the GPT-NeoX tokenizer.

    Model Pile PPL ↓\downarrow LAMBADA Acc ↑\uparrow HellaSwag Acc ↑\uparrow PIQA Acc ↑\uparrow ARC-E Acc ↑\uparrow ARC-C Acc ↑\uparrow WinoGrande Acc ↑\uparrow OpenBookQA Acc ↑\uparrow Average Acc ↑\uparrow
    Pythia-1B 7.82 56.1 47.2 70.7 57.0 27.1 53.5 31.4 49.0
    Mamba-790M 7.33 62.7 55.1 72.1 61.2 29.5 56.1 34.2 53.0
    Mamba-2-780M 7.26 61.7 54.9 72.0 61.0 28.5 60.2 36.2 53.5
    Pythia-1.4B 7.51 61.7 52.1 71.0 60.5 28.5 57.2 30.8 51.7
    RWKV4-1.5B 7.70 56.4 52.5 72.4 60.5 29.4 54.6 34.0 51.4
    Mamba-1.4B 6.80 65.0 59.1 74.2 65.5 32.8 61.5 36.4 56.4
    Mamba-2-1.3B 6.66 65.7 59.9 73.2 64.3 33.3 60.9 37.8 56.4
    Pythia-2.8B 6.73 64.7 59.3 74.0 64.1 32.9 59.7 35.2 55.7
    RWKV4-3B 7.00 63.9 59.6 73.7 67.8 33.1 59.6 37.0 56.4
    Mamba-2.8B 6.22 69.2 66.1 75.2 69.7 36.3 63.5 39.6 59.9
    Mamba-2-2.7B 6.09 69.7 66.6 76.4 69.6 36.4 64.0 38.8 60.2
    Pythia-6.9B 6.51 67.1 64.0 75.2 67.3 35.5 61.3 38.0 58.3

    Mamba-2-2.7B achieves an average downstream accuracy of 60.2% and a validation perplexity of 6.09 on the Pile, outperforming Mamba-2.8B (59.9% / 6.22) and Pythia-6.9B (58.3% / 6.51).

  11. Knowl 11 — Associative Recall Scaling on MQAR

    empirical result

    On the synthetic Multi-Query Associative Recall (MQAR) task across sequence lengths T∈{256,512,1024}T \in \{256, 512, 1024\} with T/4T/4 key-value pairs and vocabulary size 8192, Mamba-2 substantially outperforms Mamba-1.

    Key empirical findings include:

    1. At controlled state dimension N=16N = 16, Mamba-2 achieves markedly higher accuracy than Mamba-1 across all model dimensions D∈{32,64,128,256}D \in \{32, 64, 128, 256\}.
    2. Increasing Mamba-2 state capacity from N=16N = 16 to N=64N = 64 and N=256N = 256 yields monotonic improvements in recall accuracy, approaching 1.01.0 accuracy at sequence lengths 512 and 1024, matching or exceeding softmax attention baselines.
  12. Knowl 12 — Computational Speed Benchmarks of SSD

    empirical result

    Benchmarked on an NVIDIA A100 80GB PCIe GPU, the SSD kernel demonstrates substantial wall-clock speed improvements over prior SSM scan implementations and attention:

    • vs. Mamba-1 Scan: SSD is 2×2\times to 8×8\times faster than the hardware-aware fused selective scan kernel of Mamba-1 at large state expansion (N=64N = 64), because SSD leverages GPU Tensor Cores via matrix-matrix multiplication rather than memory-bandwidth-bound associative scans.
    • vs. FlashAttention-2: Due to linear O(T)O(T) asymptotic sequence scaling, SSD surpasses FlashAttention-2 speed starting at sequence length 2K2\text{K}, achieving approximately 6×6\times higher throughput at sequence length 16K16\text{K}.

Coverage note — Ablations on 1-SS matrix factorization algorithms (Appendix B) and secondary kernel activation functions (cosFormer, PRF/Performer, ReBased) were omitted as they are auxiliary derivations or yielded negative empirical results.

References

  1. 1.Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”. In: arXiv preprint arXiv:2305.13245 (2023).
  2. 2.Yaroslav Aksenov, Nikita Balagansky, Sofia Maria Lo Cicero Vaina, Boris Shaposhnikov, Alexey Gorbatovski, and Daniil Gavrilov. “Linear Transformers with Learnable Kernel Functions are Better In-Context Models”. In: arXiv preprint arXiv:2402.10644 (2024).
  3. 3.Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas. “In-Context Language Learning: Architectures and Algorithms”. In: The International Conference on Machine Learning (ICML). 2024.
  4. 4.Ameen Ali, Itamar Zimerman, and Lior Wolf. The Hidden Attention of Mamba Models. 2024. arXiv: 2403.01590 [cs.LG].
  5. 5.Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christo­pher Ré. “Zoology: Measuring and Improving Recall in Efficient Language Models”. In: The International Conference on Learning Representations (ICLR). 2024.
  6. 6.Simran Arora, Sabri Eyuboglu, Michael Zhang, Aman Timalsina, Silas Alberti, Dylan Zinsley, James Zou, Atri Rudra, and Christopher Ré. “Simple Linear Attention Language Models Balance the Recall-Throughput Tradeoff”. In: The International Conference on Machine Learning (ICML). 2024.
  7. 7.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. “Neural Machine Translation by Jointly Learning to Align and Translate”. In: The International Conference on Learning Representations (ICLR). 2015.
  8. 8.George A Baker, George A Baker Jr, Peter Graves-Morris, and Susan S Baker. Pade Approximants: Encyclopedia of Mathematics and It’s Applications, Vol. 59 George A. Baker, Jr., Peter Graves-Morris. Vol. 59. Cambridge University Press, 1996.
  9. 9.Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Gün­ter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. “XLSTM: Extended Long Short-Term Memory”. In: arXiv preprint arXiv:2405.04517 (2024).
  10. 10.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mo­hammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. “ Pythia: A Suite for Analyzing Large Language Models across Training and Scaling”. In: The International Conference on Machine Learning (ICML). PMLR. 2023, pp. 2397–2430.
  11. 11.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. “ PIQA: Reasoning about Physical Commonsense in Natural Language”. In: Proceedings of the AAAI conference on Artificial Intelligence. Vol. 34. 2020.
  12. 12.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. “ Gpt-NeoX-20B: An Open-source Autoregressive Language Model”. In: arXiv preprint arXiv:2204.06745 (2022).
  13. 13.Guy E Blelloch. “ Prefix Sums and Their Applications”. In: (1990).
  14. 14.Aleksandar Botev, Soham De, Samuel L Smith, Anushan Fernando, George-Cristian Muraru, Ruba Haroun, Leonard Berrada, Razvan Pascanu, Pier Giuseppe Sessa, Robert Dadashi, et al. “ RecurrentGemma: Moving Past Transform­ers for Efficient Open Language Models”. In: arXiv preprint arXiv:2404.07839 (2024).
  15. 15.George EP Box, Gwilym M Jenkins, Gregory C Reinsel, and Greta M Ljung. Time Series Analysis: Forecasting and Control. John Wiley & Sons, 2015.
  16. 16.James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher. “ Quasi-recurrent Neural Networks”. In: arXiv preprint arXiv:1611.01576 (2016).
  17. 17.William Brandon, Aniruddha Nrusimha, Kevin Qian, Zachary Ankner, Tian Jin, Zhiye Song, and Jonathan Ragan­Kelley. “ Striped attention: Faster ring attention for causal transformers”. In: arXiv preprint arXiv:2311.09431 (2023).
  18. 18.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakan­tan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. “ Language Models are Few-shot Learners”. In: Advances in Neural Information Processing Systems (NeurIPS) 33 (2020), pp. 1877–1901.
  19. 19.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. “ Rethinking Attention with Performers”. In: The International Conference on Learning Representations (ICLR). 2021.
  20. 20.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. “ PaLM: Scaling Language Modeling with Path­ways”. In: Journal of Machine Learning Research 24.240 (2023), pp. 1–113. url: http://jmlr.org/papers/v24/22-1144.html.
  21. 21.Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. “ Empirical Evaluation of Gated Recur­rent Neural Networks on Sequence Modeling”. In: arXiv preprint arXiv:1412.3555 (2014).
  22. 22.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. “ Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”. In: arXiv preprint arXiv:1803.05457 (2018).
  23. 23.Tri Dao. “ FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”. In: The International Conference on Learning Representations (ICLR). 2024.
  24. 24.Tri Dao, Beidi Chen, Nimit S Sohoni, Arjun Desai, Michael Poli, Jessica Grogan, Alexander Liu, Aniruddh Rao, Atri Rudra, and Christopher Ré. “ Monarch: Expressive structured matrices for efficient and accurate training”. In: International Conference on Machine Learning. PMLR. 2022, pp. 4690–4721.
  25. 25.Tri Dao, Daniel Y Fu, Khaled K Saab, Armin W Thomas, Atri Rudra, and Christopher Ré. “ Hungry Hungry Hippos: Towards Language Modeling with State Space Models”. In: The International Conference on Learning Representations (ICLR). 2023.
  26. 26.Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré. “ Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations”. In: The International Conference on Machine Learning (ICML). 2019.
  27. 27.Tri Dao, Nimit Sohoni, Albert Gu, Matthew Eichhorn, Amit Blonder, Megan Leszczynski, Atri Rudra, and Christo­pher Ré. “ Kaleidoscope: An Efficient, Learnable Representation for All Structured Linear Maps”. In: The Interna­tional Conference on Learning Representations (ICLR). 2020.
  28. 28.Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. “ Vision Transformers Need Registers”. In: The International Conference on Learning Representations (ICLR). 2024.
  29. 29.Soham De, Samuel L Smith, Anushan Fernando, Aleksandar Botev, George Cristian-Muraru, Albert Gu, Ruba Haroun, Leonard Berrada, Yutian Chen, Srivatsan Srinivasan, et al. “ Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models”. In: arXiv preprint arXiv:2402.19427 (2024).
  30. 30.Christopher De Sa, Albert Gu, Rohan Puttagunta, Christopher Ré, and Atri Rudra. “ A Two-Pronged Progress in Structured Dense Matrix Vector Multiplication”. In: Proceedings of the Twenty-Ninth Annual ACM-SIAM Symposium on Discrete Algorithms. SIAM. 2018, pp. 1060–1079.
  31. 31.Hantian Ding, Zijian Wang, Giovanni Paolini, Varun Kumar, Anoop Deoras, Dan Roth, and Stefano Soatto. “ Fewer truncations improve language modeling”. In: arXiv preprint arXiv:2404.10830 (2024).
  32. 32.Yuli Eidelman and Israel Gohberg. “ On a new class of structured matrices”. In: Integral Equations and Operator Theory 34.3 (1999), pp. 293–324.
  33. 33.Dan Fu, Simran Arora, Jessica Grogan, Isys Johnson, Evan Sabri Eyuboglu, Armin Thomas, Benjamin Spector, Michael Poli, Atri Rudra, and Christopher Ré. “ Monarch mixer: A simple sub-quadratic gemm-based architecture”. In: Advances in Neural Information Processing Systems 36 (2024).
  34. 34.Daniel Y Fu, Elliot L Epstein, Eric Nguyen, Armin W Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré. “ Simple Hardware-efficient Long Convolutions for Sequence Modeling”. In: The International Conference on Machine Learning (ICML) (2023).
  35. 35.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. “ The Pile: An 800GB Dataset of Diverse Text for Language Modeling”. In: arXiv preprint arXiv:2101.00027 (2020).
  36. 36.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A Framework for Few-shot Language Model Evaluation. Version v0.0.1. Sept. 2021. doi: 10.5281/zenodo.5371628. url: https://doi.org/10.5281/zenodo.5371628.
  37. 37.Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge. “ Zamba: A Compact 7B SSM Hybrid Model”. In: arXiv preprint arXiv:2405.16712 (2024).
  38. 38.Riccardo Grazzi, Julien Siems, Simon Schrodi, Thomas Brox, and Frank Hutter. “ Is Mamba Capable of In-Context Learning?” In: arXiv preprint arXiv:2402.03170 (2024).
  39. 39.Albert Gu. “ Modeling Sequences with Structured State Spaces”. PhD thesis. Stanford University, 2023.
  40. 40.Albert Gu and Tri Dao. “ Mamba: Linear-Time Sequence Modeling with Selective State Spaces”. In: arXiv preprint arXiv:2312.00752 (2023).
  41. 41.Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré. “ HIPPO: Recurrent Memory with Optimal Polynomial Projections”. In: Advances in Neural Information Processing Systems (NeurIPS). 2020.
  42. 42.Albert Gu, Karan Goel, and Christopher Ré. “ Efficiently Modeling Long Sequences with Structured State Spaces”. In: The International Conference on Learning Representations (ICLR). 2022.
  43. 43.Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré. “ On the Parameterization and Initialization of Diagonal State Space Models”. In: Advances in Neural Information Processing Systems (NeurIPS). 2022.
  44. 44.Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. “ Combining Recurrent, Convolutional, and Continuous-time Models with the Linear State Space Layer”. In: Advances in Neural Information Processing Systems (NeurIPS). 2021.
  45. 45.Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré. “ How to Train Your HIPPO: State Space Models with Generalized Basis Projections”. In: The International Conference on Learning Representations (ICLR). 2023.
  46. 46.Ankit Gupta, Albert Gu, and Jonathan Berant. “ Diagonal State Spaces are as Effective as Structured State Spaces”. In: Advances in Neural Information Processing Systems 35 (2022), pp. 22982–22994.
  47. 47.Dan Hendrycks and Kevin Gimpel. “ Gaussian Error Linear Units (GELUs)”. In: arXiv preprint arXiv:1606.08415 (2016).
  48. 48.W Daniel Hillis and Guy L Steele Jr. “ Data Parallel Algorithms”. In: Communications of the ACM 29.12 (1986), pp. 1170–1183.
  49. 49.Sepp Hochreiter and Jürgen Schmidhuber. “ Long Short-Term Memory”. In: Neural Computation 9.8 (1997), pp. 1735–1780.
  50. 50.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. “ An Empirical Analysis of Compute-Optimal Large Language Model Training”. In: Advances in Neural Information Processing Systems (NeurIPS) 35 (2022), pp. 30016–30030.
  51. 51.Samy Jelassi, David Brandfonbrener, Sham M Kakade, and Eran Malach. “ Repeat After Me: Transformers Are Better Than State Space Models at Copying”. In: The International Conference on Machine Learning (ICML). 2024.
  52. 52.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. “ Transformers are RNNs: Fast Au­toregressive Transformers with Linear Attention”. In: International Conference on Machine Learning. PMLR. 2020, pp. 5156–5165.
  53. 53.Tobias Katsch. “ GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling”. In: arXiv preprint arXiv:2311.01927 (2023).
  54. 54.Shiva Kaul. “ Linear Dynamical Systems as a Core Computational Primitive”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 16808–16820.
  55. 55.Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. “ Reducing activation recomputation in large transformer models”. In: Proceedings of Machine Learning and Systems 5 (2023).
  56. 56.James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. “ Fnet: Mixing tokens with fourier trans­forms”. In: arXiv preprint arXiv:2105.03824 (2021).
  57. 57.Tao Lei. “ When Attention Meets Fast Recurrence: Training Language Models with Reduced Compute”. In: Pro­ceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021, pp. 7633–7648.
  58. 58.Tao Lei, Yu Zhang, Sida I Wang, Hui Dai, and Yoav Artzi. “ Simple Recurrent Units for Highly Parallelizable Recur­rence”. In: arXiv preprint arXiv:1709.02755 (2017).
  59. 59.Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen, and Debadeepta Dey. “ What Makes Convolutional Models Great on Long Sequence Modeling?” In: The International Conference on Learning Representations (ICLR). 2023.
  60. 60.Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. “ Jamba: A Hybrid Transformer-Mamba Language Model”. In: arXiv preprint arXiv:2403.19887 (2024).
  61. 61.Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel. “ World Model on Million-Length Video And Language With RingAttention”. In: arXiv preprint arXiv:2402.08268 (2024).
  62. 62.Hao Liu, Matei Zaharia, and Pieter Abbeel. “ Ring attention with blockwise transformers for near-infinite context”. In: arXiv preprint arXiv:2310.01889 (2023).
  63. 63.Chris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto, Jakob Foerster, Satinder Singh, and Feryal Behbahani. “ Structured State Space Models for In-Context Reinforcement Learning”. In: Advances in Neural Information Pro­cessing Systems (NeurIPS). 2023.
  64. 64.Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. “ Mega: Moving Average Equipped Gated Attention”. In: The International Conference on Learning Representations (ICLR). 2023.
  65. 65.Eric Martin and Chris Cundy. “ Parallelizing Linear Recurrent Neural Nets Over Sequence Length”. In: The Inter­national Conference on Learning Representations (ICLR). 2018.
  66. 66.Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. “ Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering”. In: arXiv preprint arXiv:1809.02789 (2018).
  67. 67.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. “ In-context Learning and Induction Heads”. In: Trans­former Circuits Thread (2022). https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html.
  68. 68.Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De. “ Resurrecting Recurrent Neural Networks for Long Sequences”. In: The International Conference on Machine Learning (ICML). 2023.
  69. 69.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. “ The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context”. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguis­tics. 2016, pp. 1525–1534.
  70. 70.Jongho Park, Jaeseung Park, Zheyang Xiong, Nayoung Lee, Jaewoong Cho, Samet Oymak, Kangwook Lee, and Dimitris Papailiopoulos. “ Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks”. In: The International Conference on Machine Learning (ICML). 2024.
  71. 71.Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al. “ RWKV: Reinventing RNNs for the Transformer Era”. In: arXiv preprint arXiv:2305.13048 (2023).
  72. 72.Bo Peng, Daniel Goldstein, Quentin Anthony, Alon Albalak, Eric Alcaide, Stella Biderman, Eugene Cheah, Teddy Ferdinan, Haowen Hou, Przemysław Kazienko, et al. “ Eagle and Finch: RWKV with matrix-valued states and dy­namic recurrence”. In: arXiv preprint arXiv:2404.05892 (2024).
  73. 73.Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong. “ Random Feature Attention”. In: The International Conference on Learning Representations (ICLR). 2021.
  74. 74.Clément Pernet. “ Computing with Quasiseparable Matrices”. In: Proceedings of the ACM on International Sympo­sium on Symbolic and Algebraic Computation. 2016, pp. 389–396.
  75. 75.Clément Pernet, Hippolyte Signargout, and Gilles Villard. “ Exact computations with quasiseparable matrices”. In: arXiv preprint arXiv:2302.04515 (2023).
  76. 76.Clément Pernet and Arne Storjohann. “ Time and space efficient generators for quasiseparable matrices”. In: Journal of Symbolic Computation 85 (2018), pp. 224–246.
  77. 77.Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. “ Hyena Hierarchy: Towards Larger Convolutional Language Models”. In: The International Conference on Machine Learning (ICML). 2023.
  78. 78.Hadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pavan Kumar Anasosalu Vasu, Cem Koc, Vaishaal Shankar, and Oncel Tuzel. “ Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum”. In: arXiv preprint arXiv:2405.13226 (2024).
  79. 79.Ofir Press, Noah Smith, and Mike Lewis. “ Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation”. In: International Conference on Learning Representations. 2022.
  80. 80.Zhen Qin, Xiaodong Han, Weixuan Sun, Bowen He, Dong Li, Dongxu Li, Yuchao Dai, Lingpeng Kong, and Yiran Zhong. “ Toeplitz Neural Network for Sequence Modeling”. In: The International Conference on Learning Represen­tations (ICLR). 2023.
  81. 81.Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong. “ The devil in linear transformer”. In: arXiv preprint arXiv:2210.10340 (2022).
  82. 82.Zhen Qin, Dong Li, Weigao Sun, Weixuan Sun, Xuyang Shen, Xiaodong Han, Yunshen Wei, Baohong Lv, Xiao Luo, Yu Qiao, et al. “ TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer”. In: arXiv preprint arXiv:2307.14995 (2023).
  83. 83.Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong. “ CosFormer: Rethinking Softmax in Attention”. In: The International Conference on Learning Representations (ICLR). 2022.
  84. 84.Zhen Qin, Songlin Yang, Weixuan Sun, Xuyang Shen, Dong Li, Weigao Sun, and Yiran Zhong. “ HGRN2: Gated Linear RNNs with State Expansion”. In: arXiv preprint arXiv:2404.07904 (2024).
  85. 85.Zhen Qin, Songlin Yang, and Yiran Zhong. “ Hierarchically Gated Recurrent Neural Network for Sequence Model­ing”. In: Advances in Neural Information Processing Systems 36 (2023).
  86. 86.Ali Rahimi and Benjamin Recht. “ Random Features for Large-Scale Kernel Machines”. In: Advances in Neural In­formation Processing Systems (NeurIPS) 20 (2007).
  87. 87.Prajit Ramachandran, Barret Zoph, and Quoc V Le. “ Swish: A Self-gated Activation Function”. In: arXiv preprint arXiv:1710.05941 7.1 (2017), p. 5.
  88. 88.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. “ Winogrande: An Adversarial Winograd Schema Challenge at Scale”. In: Communications of the ACM 64.9 (2021), pp. 99–106.
  89. 89.Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. “ Linear Transformers are Secretly Fast Weight Programmers”. In: The International Conference on Machine Learning (ICML). PMLR. 2021, pp. 9355–9366.
  90. 90.Noam Shazeer. “ Fast Transformer Decoding: One Write-head is All You Need”. In: arXiv preprint arXiv:1911.02150 (2019).
  91. 91.Sam Shleifer, Jason Weston, and Myle Ott. “ NormFormer: Improved Transformer Pretraining with Extra Normal­ization”. In: arXiv preprint arXiv:2110.09456 (2021).
  92. 92.Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. “ Megatron­LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”. In: arXiv preprint arXiv:1909.08053 (2019).
  93. 93.Jimmy TH Smith, Andrew Warrington, and Scott W Linderman. “ Simplified State Space Layers for Sequence Mod­eling”. In: The International Conference on Learning Representations (ICLR). 2023.
  94. 94.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. “ Roformer: Enhanced Transformer with Rotary Position Embedding”. In: arXiv preprint arXiv:2104.09864 (2021).
  95. 95.Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. “ Reten­tive network: A successor to transformer for large language models”. In: arXiv preprint arXiv:2307.08621 (2023).
  96. 96.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. “ Efficient Transformers: A Survey”. In: ACM Comput­ing Surveys 55.6 (2022), pp. 1–28.
  97. 97.Chameleon Team. “ Chameleon: Mixed-Modal Early-Fusion Foundation Models”. In: arXiv preprint arXiv:2405.09818 (2024).
  98. 98.Anna Thomas, Albert Gu, Tri Dao, Atri Rudra, and Christopher Ré. “ Learning Compressed Transforms with Low Displacement Rank”. In: Advances in Neural Information Processing Systems (NeurIPS). 2018, pp. 9052–9060.
  99. 99.Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. “ MLP-Mixer: An All-MLP Architecture for Vision”. In: Advances in Neural Information Processing Systems (NeurIPS) 34 (2021), pp. 24261–24272.
  100. 100.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. “ Llama: Open and Efficient Foundation Language Models”. In: arXiv preprint arXiv:2302.13971 (2023).
  101. 101.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. “ Llama 2: Open foundation and fine-tuned chat models”. In: arXiv preprint arXiv:2307.09288 (2023).
  102. 102.Raf Vandebril, M Van Barel, Gene Golub, and Nicola Mastronardi. “ A bibliography on semiseparable matrices”. In: Calcolo 42 (2005), pp. 249–270.
  103. 103.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. “ Attention Is All You Need”. In: Advances in Neural Information Processing Systems (NeurIPS). 2017.
  104. 104.Shida Wang and Beichen Xue. “ State-space Models with Layer-wise Nonlinearity are Universal Approximators with Exponential Decaying Memory”. In: arXiv preprint arXiv:2309.13414 (2023).
  105. 105.Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. “ Linformer: Self-attention with Linear Com­plexity”. In: arXiv preprint arXiv:2006.04768 (2020).
  106. 106.Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. “ Efficient Streaming Language Models with Attention Sinks”. In: The International Conference on Learning Representations (ICLR). 2024.
  107. 107.Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. “ Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention”. In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 35. 2021.
  108. 108.Songlin Yang, Bailin Wang, Yikang Shen, Rameswar Panda, and Yoon Kim. “ Gated Linear Attention Transformers with Hardware-Efficient Training”. In: The International Conference on Machine Learning (ICML). 2024.
  109. 109.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. “ HellaSwag: Can a Machine Really Finish Your Sentence?” In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
  110. 110.Jinle Zeng, Min Li, Zhihua Wu, Jiaqi Liu, Yuang Liu, Dianhai Yu, and Yanjun Ma. “ Boosting distributed training performance of the unpadded bert model”. In: arXiv preprint arXiv:2208.08124 (2022).
  111. 111.Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind. “ An Attention Free Transformer”. In: arXiv preprint arXiv:2105.14103 (2021).
  112. 112.Yujia Zhai, Chengquan Jiang, Leyuan Wang, Xiaoying Jia, Shang Zhang, Zizhong Chen, Xin Liu, and Yibo Zhu. “ Bytetransformer: A high-performance transformer boosted for variable-length inputs”. In: 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE. 2023, pp. 344–355.
  113. 113.Michael Zhang, Kush Bhatia, Hermann Kumbong, and Christopher Ré. “ The Hedgehog & the Porcupine: Expres­sive Linear Attentions with Softmax Mimicry”. In: The International Conference on Learning Representations (ICLR). 2024.
  114. 114.Lin Zheng, Chong Wang, and Lingpeng Kong. “ Linear complexity randomized self-attention mechanism”. In: In­ternational Conference on Machine Learning. PMLR. 2022, pp. 27011–27041.

Citation

MLA
Dao, T., and A. Gu. “Transformers Are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality”. arXiv, 2024, http://arxiv.org/abs/2405.21060v1.
APA
Dao, T., & Gu, A. (2024). Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. arXiv. http://arxiv.org/abs/2405.21060v1
Chicago
Dao, T., and A. Gu. 2024. “Transformers Are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality”. arXiv. http://arxiv.org/abs/2405.21060v1.
Harvard
Dao, T. and Gu, A. (2024) “Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2405.21060v1.
Vancouver
1. Dao T, Gu A (2024) Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. arXiv

BibTeX

@article{dao2024transformers,
  title = {Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality},
  author = {Dao, Tri and Gu, Albert},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2405.21060v1},
  eprint = {2405.21060}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/