Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Albert GuTri Dao

article2023arXiv9,138 citationsOutstanding Paper Award, COLM 2024

Introduces Mamba, a selective state space architecture that achieves linear-time sequence scaling and five times higher inference throughput while matching or outperforming standard Transformers across language, audio, and genomics benchmarks.

Listen

Modern foundation models rely almost exclusively on the Transformer architecture, which processes sequences using attention mechanisms. While effective at capturing complex patterns, attention scales quadratically with sequence length and requires storing past context in memory during text generation. This creates severe computational bottlenecks, high deployment costs, and practical limits on processing long sequences. Earlier subquadratic alternatives, such as structured state space models, offered linear scaling but struggled with information-dense and discrete data like language because their fixed, time-invariant dynamics prevented them from selectively focusing on or ignoring specific information based on content.

The article introduces and evaluates Mamba, a sequence modeling architecture based on selective state space models. The objective is to demonstrate that introducing input-dependent selectivity into state space models achieves Transformer-level quality while maintaining linear compute and memory scaling during training and constant time per step during generation.

To accomplish this, the researchers made state space parameters dynamic functions of the input, enabling the model to filter irrelevant noise or retain key context indefinitely. Because this dynamic behavior rules out efficient convolution algorithms, the authors designed a hardware-aware parallel scan that computes recurrences directly in fast processor memory (GPU SRAM) rather than slow main memory (HBM) and recomputes states during training to minimize memory traffic. They integrated this mechanism into a unified neural network block that replaces separate attention and multi-layer perceptron blocks. The architecture was tested across synthetic reasoning tasks, natural language processing up to 3 billion parameters, genomics modeling with sequences up to 1 million elements, and audio generation.

The evaluation yielded several key findings. First, Mamba matched or exceeded the performance of modern Transformer architectures across model sizes while scaling linearly with context length. In language modeling, a 3-billion-parameter Mamba model matched the pretraining and downstream evaluation performance of standard Transformers twice its size. Second, during text generation, Mamba delivered up to five times higher inference throughput than comparable Transformers because it does not require a key-value memory cache. Third, the hardware-aware scan ran up to three times faster than previous methods on modern GPUs and surpassed FlashAttention-2 speed at sequence lengths above 2,000 tokens. Fourth, in long-context domains like DNA and audio pretraining, Mamba continually improved as sequence lengths scaled up to 1 million tokens, whereas prior models plateaued or degraded due to an inability to discard context noise.

These findings suggest that the long-standing trade-off between modeling quality and sequence efficiency can be resolved without attention mechanisms. For technical organizations, deploying selective state space models could dramatically lower inference serving costs, reduce hardware memory footprints, and enable cost-effective analysis of massive contexts in fields like genomics, audio, and document processing. Furthermore, Mamba challenges the assumption that Transformer-style self-attention is essential for high-quality language modeling.

Organizations handling large-scale sequence processing should pilot Mamba implementations for latency-sensitive and long-context applications, particularly in document retrieval, DNA sequence analysis, and audio generation. Engineering teams should also evaluate the open-source codebase to assess integration within existing deployment pipelines. However, further research is required to evaluate Mamba at frontier model scales beyond 7 billion parameters and to thoroughly assess its compatibility with common ecosystem tooling such as instruction fine-tuning, quantization, reinforcement learning from human feedback, and in-context prompting.

While confidence in the experimental findings is high for model scales up to 3 billion parameters, readers should note that tests were confined to these smaller model sizes. Additionally, ablations reveal that while selectivity excels on discrete data like text and DNA, raw continuous signals (such as dense audio waveforms) still benefit from linear time-invariant modeling. System architects should account for this trade-off when selecting model backbones across different data modalities.

Cover for Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Abstract

Foundation models, now powering most of the exciting applications in deep learning, are almost universally based on the Transformer architecture and its core attention module. Many subquadratic-time architectures such as linear attention, gated convolution and recurrent models, and structured state space models (SSMs) have been developed to address Transformers' computational inefficiency on long sequences, but they have not performed as well as attention on important modalities such as language. We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements. First, simply letting the SSM parameters be functions of the input addresses their weakness with discrete modalities, allowing the model to selectively propagate or forget information along the sequence length dimension depending on the current token. Second, even though this change prevents the use of efficient convolutions, we design a hardware-aware parallel algorithm in recurrent mode. We integrate these selective SSMs into a simplified end-to-end neural network architecture without attention or even MLP blocks (Mamba). Mamba enjoys fast inference (5×\times higher throughput than Transformers) and linear scaling in sequence length, and its performance improves on real data up to million-length sequences. As a general sequence model backbone, Mamba achieves state-of-the-art performance across several modalities such as language, audio, and genomics. On language modeling, our Mamba-3B model outperforms Transformers of the same size and matches Transformers twice its size, both in pretraining and downstream evaluation.

Table of Contents

  • 1 Introduction
  • 2 State Space Models
  • 3 Selective State Space Models
  • 3.1 Motivation: Selection as a Means of Compression
  • 3.2 Improving SSMs with Selection
  • 3.3 Efficient Implementation of Selective SSMs
  • 3.3.1 Motivation of Prior Models
  • 3.3.2 Overview of Selective Scan: Hardware-Aware State Expansion
  • 3.4 A Simplified SSM Architecture
  • 3.5 Properties of Selection Mechanisms
  • 3.5.1 Connection to Gating Mechanisms
  • 3.5.2 Interpretation of Selection Mechanisms
  • 3.6 Additional Model Details
  • 4 Empirical Evaluation
  • 4.1 Synthetic Tasks
  • 4.1.1 Selective Copying
  • 4.1.2 Induction Heads
  • 4.2 Language Modeling
  • 4.2.1 Scaling Laws
  • 4.2.2 Downstream Evaluations
  • 4.3 DNA Modeling
  • 4.3.1 Scaling: Model Size
  • 4.3.2 Scaling: Context Length
  • 4.3.3 Synthetic Species Classification
  • 4.4 Audio Modeling and Generation
  • 4.4.1 Long-Context Autoregressive Pretraining
  • 4.4.2 Autoregressive Speech Generation
  • 4.5 Speed and Memory Benchmarks
  • 4.6 Model Ablations
  • 4.6.1 Architecture
  • 4.6.2 Selective SSM
  • 5 Discussion
  • 6 Conclusion
  • References
  • A Discussion: Selection Mechanism
  • B Related Work
  • B.1 S4 Variants and Derivatives
  • B.2 SSM Architectures
  • B.3 Relationship to RNNs
  • B.4 Linear Attention
  • B.5 Long Context Models
  • C Mechanics of Selective SSMs
  • D Hardware-aware Algorithm For Selective SSMs
  • E Experimental Details and Additional Results
  • E.1 Synthetic Tasks
  • E.2 Language Modeling
  • E.2.1 Scaling Law Details
  • E.2.2 Additional Scaling Law Ablations
  • E.2.3 Downstream Evaluation Details
  • E.3 DNA Modeling
  • E.3.1 Pretraining Details
  • E.3.2 Scaling: Model Size Details
  • E.3.3 Scaling: Context Length Details
  • E.3.4 Species (Great Apes) Classification
  • E.4 Audio Details
  • E.4.1 YouTubeMix Audio Pretraining
  • E.4.2 SC09 Speech Generation
  • E.5 Efficiency Benchmark

Knowls

  1. Knowl 1 — Selective State Space Layer (S6)

    algorithm

    The selective state space model (S6) generalizes linear time-invariant (LTI) structured state space models by making the continuous parameters Δ\Delta, BB, and CC input-dependent functions of the current token xtx_t, transforming the sequence transformation into a time-varying recurrence.

    Given an input sequence x∈RB×L×Dx \in \mathbb{R}^{B \times L \times D} (batch size BB, sequence length LL, model dimension DD) and state expansion dimension NN:

    1. The system matrix A∈RD×NA \in \mathbb{R}^{D \times N} is maintained as a learned parameter (often initialized via structured diagonal real matrices).
    2. Intermediate parameters are computed via linear projections from the input: B=sB(x)=LinearN(x)∈RB×L×NB = s_B(x) = \text{Linear}_N(x) \in \mathbb{R}^{B \times L \times N}, C=sC(x)=LinearN(x)∈RB×L×NC = s_C(x) = \text{Linear}_N(x) \in \mathbb{R}^{B \times L \times N}, and Δ=τΔ(Parameter+sΔ(x))∈RB×L×D\Delta = \tau_\Delta(\text{Parameter} + s_\Delta(x)) \in \mathbb{R}^{B \times L \times D}, where sΔ(x)=BroadcastD(Linear1(x))s_\Delta(x) = \text{Broadcast}_D(\text{Linear}_1(x)) or a low-rank projection LinearD(LinearR(x))\text{Linear}_D(\text{Linear}_R(x)), and τΔ=softplus\tau_\Delta = \text{softplus}.
    3. The parameters are discretized using the Zero-Order Hold (ZOH) rule: Aˉ=exp⁡(ΔA),Bˉ=(ΔA)−1(exp⁡(ΔA)−I)⋅(ΔB)\bar{A} = \exp(\Delta A), \quad \bar{B} = (\Delta A)^{-1}(\exp(\Delta A) - I) \cdot (\Delta B)
    4. The model computes the time-varying state recurrence: ht=Aˉtht−1+Bˉtxt,yt=Cthth_t = \bar{A}_t h_{t-1} + \bar{B}_t x_t, \quad y_t = C_t h_t
    Input: x of shape (B, L, D)
    Output: y of shape (B, L, D)
    A: parameter of shape (D, N)
    B = Linear_N(x) of shape (B, L, N)
    C = Linear_N(x) of shape (B, L, N)
    Delta = softplus(Delta_param + Broadcast_D(Linear_1(x))) of shape (B, L, D)
    for each batch b, channel d, timestep t:
        A_bar[b, t, d, :] = exp(Delta[b, t, d] * A[d, :])
        B_bar[b, t, d, :] = (exp(Delta[b, t, d] * A[d, :]) - 1) / A[d, :] * B[b, t, :]
        h[b, t, d, :] = A_bar[b, t, d, :] * h[b, t - 1, d, :] + B_bar[b, t, d, :] * x[b, t, d]
        y[b, t, d] = sum(h[b, t, d, :] * C[b, t, :])
    return y
  2. Knowl 2 — Hardware-Aware State Expansion and Fused Parallel Scan

    model/method

    Because the selective SSM has time-varying parameters (Δt,Bt,Ct)(\Delta_t, B_t, C_t), it cannot be computed as a global convolution via Fast Fourier Transforms. Computing recurrence naively requires materializing the expanded hidden state h∈RB×L×D×Nh \in \mathbb{R}^{B \times L \times D \times N} in GPU High-Bandwidth Memory (HBM), creating a memory I/O bottleneck scaling as O(BLDN)O(B L D N).

    The hardware-aware selective scan avoids this by exploiting GPU SRAM:

    1. Kernel Fusion: The parameter projections, discretization (Aˉ,Bˉ)(\bar{A}, \bar{B}), associative parallel scan, and output contraction with CC are fused into a single GPU kernel. The parameters (Δ,A,B,C)(\Delta, A, B, C) of size O(BLD+DN)O(B L D + D N) are loaded directly from HBM into on-chip SRAM, discretization and parallel prefix scans are executed entirely within SRAM, and only the contracted output y∈RB×L×Dy \in \mathbb{R}^{B \times L \times D} is written back to HBM. This reduces memory I/O by a factor of O(N)O(N). For sequences exceeding SRAM capacity, sequences are split into chunks and processed sequentially across chunk boundaries.
    2. Activation Recomputation: To prevent storing intermediate states of size (B,L,D,N)(B, L, D, N) for backpropagation, the forward states hth_t are discarded after the forward pass. During the backward pass, the intermediate states are recomputed in SRAM from the inputs (Δ,A,B,C)(\Delta, A, B, C) loaded from HBM, eliminating the need to read O(BLDN)O(B L D N) activation data from HBM.
  3. Knowl 3 — Mamba Block Architecture

    model/method

    The Mamba architecture simplifies deep sequence models by unifying state space layers and multi-layer perceptron (MLP) blocks into a single homogenous repeating block, eliminating self-attention and separate MLP sub-layers.

    Given an input dimension DD and an expansion factor E=2E = 2:

    1. The input passes through a normalization layer (LayerNorm/RMSNorm).
    2. The input is projected into two parallel branches of dimension ED=2DE D = 2D via linear projections (2ED22 E D^2 parameters total).
    3. In the primary branch, the representation passes through:
      • A 1D short temporal convolution with kernel size dconv=4d_{\text{conv}} = 4.
      • A SiLU/Swish\text{SiLU} / \text{Swish} activation function.
      • A Selective SSM layer (S6) with state expansion dimension N=16N = 16.
    4. In the gating branch, the representation passes through a SiLU\text{SiLU} activation function.
    5. The output of the Selective SSM is gated by multiplying elementwise with the output of the gating branch (forming a SwiGLU-style gating structure).
    6. The result is projected back to dimension DD through a linear projection (ED2E D^2 parameters).

    Each block uses ≈3ED2+O(DN)≈6D2\approx 3 E D^2 + O(D N) \approx 6 D^2 parameters. Stacking two consecutive Mamba blocks matches the 12D212 D^2 parameter count of a standard Transformer's interleaved Multi-Head Attention (MHA) and MLP layer pair.

  4. Knowl 4 — Equivalence of Zero-Order Hold Selective SSM to Gated Recurrent Units

    theoretical result

    Classical gating in Recurrent Neural Networks is an exact instance of a selective State Space Model discretized via Zero-Order Hold (ZOH).

    Theorem: When the state dimension is scalar (N=1N = 1), system matrices are constant A=−1A = -1 and B=1B = 1, the input projection is sΔ(x)=Linear(x)s_\Delta(x) = \text{Linear}(x), and the step-size transformation is τΔ=softplus\tau_\Delta = \text{softplus}, the continuous leaky integrator h˙(t)=−h(t)+x(t)\dot{h}(t) = -h(t) + x(t) under ZOH discretization yields the discrete recurrence: gt=σ(Linear(xt))g_t = \sigma(\text{Linear}(x_t)) ht=(1−gt)ht−1+gtxth_t = (1 - g_t) h_{t-1} + g_t x_t where σ(z)=(1+e−z)−1\sigma(z) = (1 + e^{-z})^{-1} is the standard sigmoid function.

    Mechanistically, the parameter Δt\Delta_t acts as a generalized gating control: large Δt→∞\Delta_t \to \infty forces gt→1g_t \to 1, which resets the hidden state and focuses entirely on the current token xtx_t; small Δt→0\Delta_t \to 0 forces gt→0g_t \to 0, causing the model to persist its previous hidden state ht−1h_{t-1} and ignore xtx_t.

  5. Knowl 5 — Synthetic Task Generalization: Selective Copying and Induction Heads

    empirical result

    Linear time-invariant (LTI) SSMs and convolutional models fail on synthetic tasks that require content-aware memory filtering and dynamic spacing, whereas selective SSMs solve them and extrapolate over extreme context lengths:

    • Selective Copying: Evaluated on sequences of length 4096 with vocabulary size 16 containing 16 data tokens separated by random spacing of noise tokens. Time-invariant architectures fail (S4 achieves 18.3% accuracy; H3 with S4 achieves 57.0%; H3 with Hyena achieves 30.1%). Equipping these architectures with selection (S6) solves the task: S6 without architectural gating achieves 97.0%, H3-S6 achieves 99.7%, and Mamba (Mamba block with S6) achieves 99.8% accuracy.
    • Induction Heads Extrapolation: Models trained on sequence length L=256L = 256 with vocabulary size 16 were evaluated on test sequences up to L=220=1,048,576L = 2^{20} = 1,048,576 (a 4000×4000\times extrapolation factor). Multi-head attention (MHA with Absolute, RoPE, and xPos encodings) degrades sharply beyond 2×2\times training length and runs out of memory beyond 214=16,3842^{14} = 16,384. LTI SSMs (H3, Hyena) collapse to <10%<10\% accuracy. Mamba achieves 100.0%100.0\% accuracy across all sequence lengths from 26=642^6 = 64 up to 220=1,048,5762^{20} = 1,048,576.
  6. Knowl 6 — Language Modeling Scaling Laws and Zero-Shot Downstream Evaluation

    data/table

    Pretraining autoregressive language models on the Pile dataset following Chinchilla compute-optimal scaling laws demonstrates that Mamba matches or exceeds the performance of modern Transformer recipes (Transformer++ with RoPE, SwiGLU, and RMSNorm) across all model sizes (125M to 1.3B parameters). Zero-shot downstream benchmarks on models trained up to 300B tokens show Mamba-3B outperforms open-source models of the same parameter scale and matches models twice its size.

    Model Pile ppl ↓\downarrow LAMBADA ppl ↓\downarrow LAMBADA acc ↑\uparrow HellaSwag acc ↑\uparrow PIQA acc ↑\uparrow Arc-E acc ↑\uparrow Arc-C acc ↑\uparrow WinoGrande acc ↑\uparrow Average acc ↑\uparrow
    Pythia-160M 29.64 38.10 33.0 30.2 61.4 43.2 24.1 51.9 40.6
    Mamba-130M 10.56 16.07 44.3 35.3 64.5 48.0 24.3 51.9 44.7
    Pythia-410M 9.95 10.84 51.4 40.6 66.9 52.1 24.6 53.8 48.2
    Mamba-370M 8.28 8.14 55.6 46.5 69.5 55.1 28.0 55.3 50.0
    Pythia-1.4B 7.51 6.08 61.7 52.1 71.0 60.5 28.5 57.2 55.2
    RWKV-1.5B 7.70 7.04 56.4 52.5 72.4 60.5 29.4 54.6 54.3
    Mamba-1.4B 6.80 5.04 64.9 59.1 74.2 65.5 32.8 61.5 59.7
    Pythia-2.8B 6.73 5.04 64.7 59.3 74.0 64.1 32.9 59.7 59.1
    RWKV-3B 7.00 5.24 63.9 59.6 73.7 67.8 33.1 59.6 59.6
    Mamba-2.8B 6.22 4.23 69.2 66.1 75.2 69.7 36.3 63.5 63.3
    Pythia-6.9B 6.51 4.45 67.1 64.0 75.2 67.3 35.5 61.3 61.7
    RWKV-7.4B 6.31 4.38 67.2 65.5 76.1 67.8 37.5 61.0 62.5
  7. Knowl 7 — Long-Context DNA Pretraining and Species Classification

    empirical result

    When evaluated on genomic sequence modeling using the HG38 human genome dataset (4.5B base pairs):

    1. Model Size Scaling: At sequence length 1024, Mamba achieves better pretraining perplexity than HyenaDNA and Transformer++, matching baseline perplexity with 3×3\times to 4×4\times fewer parameters at the 40M parameter scale.
    2. Context Length Scaling: When varying context length from 210=1,0242^{10} = 1,024 to 220=1,048,5762^{20} = 1,048,576 base pairs at constant total training tokens ( ≈330B\,\approx 330\text{B} tokens), Mamba's pretraining perplexity improves monotonically with longer context (decreasing from ≈2.89\approx 2.89 at 1k1\text{k} to ≈2.76\approx 2.76 at 1M1\text{M} for Mamba-7M), whereas HyenaDNA's perplexity degrades as context length increases beyond 16k16\text{k} (rising to ≈3.01\approx 3.01 at 1M1\text{M}).
    3. Great Apes Downstream Classification: Fine-tuning pretrained models on 5-way classification between closely related species ({human, chimpanzee, gorilla, orangutan, bonobo}, sharing 99% DNA sequence similarity) shows accuracy gains with longer context up to 1M tokens. At length 220=1,048,5762^{20} = 1,048,576, Mamba-1.4M achieves 71.67% and Mamba-7M achieves 81.31% accuracy, compared to 54.87% for HyenaDNA-1.4M and 20.0% for random guessing.
  8. Knowl 8 — Audio Waveform Autoregressive Pretraining and Generation

    empirical result

    Evaluating Mamba as a backbone within a U-Net architecture for raw 16 kHz audio modeling:

    • Long-Context Pretraining (YouTubeMix): On autoregressive next-sample prediction across context lengths from 213=8,1922^{13} = 8,192 to 220≈1062^{20} \approx 10^6 samples (up to 1 minute of raw audio), Mamba achieves lower bits-per-byte (BPB) than SaShiMi (S4+MLP), with the performance advantage widening as context length increases.
    • Speech Generation (SC09): On unconditional 1-second speech generation evaluated on negative log-likelihood (NLL), Fréchet Inception Distance (FID), Inception Score (IS), modified Inception Score (mIS), and AM score:
      • Mamba-6.1M achieves NLL 1.852, FID 0.94, IS 6.26, mIS 88.54, and AM 0.52, outperforming GAN and diffusion baselines including WaveGAN-19.1M (FID 2.03, IS 4.90), DiffWave-24.1M (FID 1.92, IS 5.26), and SaShiMi-5.8M (FID 1.99, IS 5.13).
      • Scaling Mamba to 24.3M parameters improves generation metrics to FID 0.67, IS 7.33, mIS 144.9, and AM 0.36.
  9. Knowl 9 — Computational Efficiency and Inference Throughput of Mamba

    empirical result

    Benchmarked on an NVIDIA A100 80GB PCIe GPU with BF16 precision:

    1. Core Scan Kernel Speed: The fused selective scan is 20×20\times to 40×40\times faster than a naive PyTorch parallel scan implementation. It scales strictly linearly with sequence length O(L)O(L) in contrast to FlashAttention-2's O(L2)O(L^2) and convolution's O(Llog⁡L)O(L \log L), outperforming FlashAttention-2 speed at sequence lengths ≥2k\ge 2\text{k} and reaching up to 7×7\times faster processing at sequence length 32k32\text{k}.
    2. Inference Throughput: Due to its fixed recurrent state which bypasses the key-value (KV) cache bottleneck, Mamba supports substantially larger generation batch sizes. For prompt length 2048 and generation length 128, Mamba achieves 4×4\times to 5×5\times higher token generation throughput than same-sized Transformers (e.g., Mamba-1.4B achieves ≈1814\approx 1814 tokens/s at batch size 128, whereas Transformer-1.3B runs out of memory beyond batch size 32 where it achieves 364 tokens/s). An untrained Mamba-6.9B achieves higher throughput ( ≈515\,\approx 515 tokens/s at batch size 128) than a 5×5\times smaller Transformer-1.3B.
    3. Training Memory: Training activation memory footprint is comparable to FlashAttention-2 with torch.compile (e.g., 38.2 GB for Mamba-125M vs. 34.5 GB for Transformer-125M at batch size 32, sequence length 2048).
  10. Knowl 10 — Continuous vs. Discrete Modality Inductive Bias Tradeoff

    limitation

    State space models exhibit a domain-dependent inductive bias tradeoff between continuous signals and discrete sequences:

    • Continuous Perceptual Modalities (e.g., Raw Audio): Uniformly sampled continuous signals benefit from linear time-invariant (LTI) dynamics and complex-valued state representations. Applying input-dependent selectivity (S6) or strictly real-valued state matrices directly to raw audio waveforms degrades autoregressive pretraining BPB compared to complex-valued LTI S4. Selectivity is only beneficial once signals are compressed/tokenized in the inner stages of a hierarchical network (such as the center stages of a U-Net).
    • Discrete Symbolic Modalities (e.g., Text, DNA): Discrete sequences require dynamic filtering and variable token spacing, making time-varying selective mechanisms essential. In these regimes, real-valued state initializations (such as S4D-Real) perform on par with or better than complex parameterizations while being more hardware-efficient.

Coverage note — None omitted; all core methodological contributions (S6 selective SSM, hardware-aware scan, Mamba architecture), theoretical results (Theorem 1), and primary empirical findings across synthetics, language modeling, DNA, audio, and computational benchmarking are represented as standalone knowls.

References

  1. 1.Martin Arjovsky, Amar Shah, and Yoshua Bengio. “Unitary Evolution Recurrent Neural Networks”. In: The International Conference on Machine Learning (ICML). 2016, pp. 1120–1128.
  2. 2.Žiga Avsec, Vikram Agarwal, Daniel Visentin, Joseph R Ledsam, Agnieszka Grabska-Barwinska, Kyle R Taylor, Yannis Assael, John Jumper, Pushmeet Kohli, and David R Kelley. “Effective Gene Expression Prediction from Sequence by Integrating Long-range Interactions”. In: Nature Methods 18.10 (2021), pp. 1196–1203.
  3. 3.Jimmy Ba, Geoffrey E Hinton, Volodymyr Mnih, Joel Z Leibo, and Catalin Ionescu. “Using Fast Weights to Attend to the Recent Past”. In: Advances in Neural Information Processing Systems (NeurIPS) 29 (2016).
  4. 4.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. “Layer Normalization”. In: arXiv preprint arXiv:1607.06450 (2016).
  5. 5.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. “Neural Machine Translation by Jointly Learning to Align and Translate”. In: The International Conference on Learning Representations (ICLR). 2015.
  6. 6.David Balduzzi and Muhammad Ghifary. “Strongly-typed Recurrent Neural Networks”. In: International Conference on Machine Learning. PMLR. 2016, pp. 1292–1300.
  7. 7.Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. “Pythia: A Suite for Analyzing Large Language Models across Training and Scaling”. In: The International Conference on Machine Learning (ICML). PMLR. 2023, pp. 2397–2430.
  8. 8.Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. “PIQA: Reasoning about Physical Commonsense in Natural Language”. In: Proceedings of the AAAI conference on Artificial Intelligence. Vol. 34. 2020.
  9. 9.Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, et al. “Gpt-NeoX-20B: An Open-source Autoregressive Language Model”. In: arXiv preprint arXiv:2204.06745 (2022).
  10. 10.Guy E Blelloch. “Prefix Sums and Their Applications”. In: (1990).
  11. 11.James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher. “Quasi-recurrent Neural Networks”. In: arXiv preprint arXiv:1611.01576 (2016).
  12. 12.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. “Language Models are Few-shot Learners”. In: Advances in Neural Information Processing Systems (NeurIPS) 33 (2020), pp. 1877–1901.
  13. 13.Aydar Bulatov, Yuri Kuratov, and Mikhail S Burtsev. “Scaling Transformer to 1M tokens and Beyond with RMT”. In: arXiv preprint arXiv:2304.11062 (2023).
  14. 14.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. “Generating Long Sequences with Sparse Transformers”. In: arXiv preprint arXiv:1904.10509 (2019).
  15. 15.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. “Rethinking Attention with Performers”. In: The International Conference on Learning Representations (ICLR). 2021.
  16. 16.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. “PaLM: Scaling Language Modeling with Pathways”. In: Journal of Machine Learning Research 24.240 (2023), pp. 1–113. url: http://jmlr.org/papers/v24/22-1144.html.
  17. 17.Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. “Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling”. In: arXiv preprint arXiv:1412.3555 (2014).
  18. 18.Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. “Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”. In: arXiv preprint arXiv:1803.05457 (2018).
  19. 19.Tri Dao. “FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning”. In: The International Conference on Learning Representations (ICLR). 2024.
  20. 20.Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”. In: Advances in Neural Information Processing Systems (NeurIPS). 2022.
  21. 21.Tri Dao, Daniel Y Fu, Khaled K Saab, Armin W Thomas, Atri Rudra, and Christopher Ré. “Hungry Hungry Hippos: Towards Language Modeling with State Space Models”. In: The International Conference on Learning Representations (ICLR). 2023.
  22. 22.Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. “Language Modeling with Gated Convolutional Networks”. In: The International Conference on Machine Learning (ICML). PMLR. 2017, pp. 933–941.
  23. 23.DeepSound. SampleRNN. https://github.com/deepsound-project/samplernn-pytorch. 2017.
  24. 24.Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, and Furu Wei. “LongNet: Scaling Transformers to 1,000,000,000 Tokens”. In: arXiv preprint arXiv:2307.02486 (2023).
  25. 25.Chris Donahue, Julian McAuley, and Miller Puckette. “Adversarial Audio Synthesis”. In: The International Conference on Learning Representations (ICLR). 2019.
  26. 26.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”. In: The International Conference on Learning Representations (ICLR). 2020.
  27. 27.Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. “A Mathematical Framework for Transformer Circuits”. In: Transformer Circuits Thread (2021). https://transformer-circuits.pub/2021/framework/index.html.
  28. 28.Mahan Fathi, Jonathan Pilault, Pierre-Luc Bacon, Christopher Pal, Orhan Firat, and Ross Goroshin. “Block-State Transformer”. In: arXiv preprint arXiv:2306.09539 (2023).
  29. 29.Yassir Fathullah, Chunyang Wu, Yuan Shangguan, Junteng Jia, Wenhan Xiong, Jay Mahadeokar, Chunxi Liu, Yangyang Shi, Ozlem Kalinli, Mike Seltzer, and Mark J. F. Gales. “Multi-Head State Space Model for Speech Recognition”. In: Proc. INTERSPEECH 2023. 2023, pp. 241–245. doi: 10.21437/Interspeech.2023-1036.
  30. 30.Karl J Friston, Lee Harrison, and Will Penny. “Dynamic Causal Modelling”. In: Neuroimage 19.4 (2003), pp. 1273–1302.
  31. 31.Daniel Y Fu, Elliot L Epstein, Eric Nguyen, Armin W Thomas, Michael Zhang, Tri Dao, Atri Rudra, and Christopher Ré. “Simple Hardware-efficient Long Convolutions for Sequence Modeling”. In: The International Conference on Machine Learning (ICML) (2023).
  32. 32.Ken-ichi Funahashi and Yuichi Nakamura. “Approximation of Dynamical Systems by Continuous Time Recurrent Neural Networks”. In: Neural Networks 6.6 (1993), pp. 801–806.
  33. 33.Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. “The Pile: An 800GB Dataset of Diverse Text for Language Modeling”. In: arXiv preprint arXiv:2101.00027 (2020).
  34. 34.Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. A Framework for Few-shot Language Model Evaluation. Version v0.0.1. Sept. 2021. doi: 10.5281/zenodo.5371628. url: https://doi.org/10.5281/zenodo.5371628.
  35. 35.Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré. “It’s Raw! Audio Generation with State-Space Models”. In: The International Conference on Machine Learning (ICML). 2022.
  36. 36.Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Ré. “HIPPO: Recurrent Memory with Optimal Polynomial Projections”. In: Advances in Neural Information Processing Systems (NeurIPS). 2020.
  37. 37.Albert Gu, Karan Goel, and Christopher Ré. “Efficiently Modeling Long Sequences with Structured State Spaces”. In: The International Conference on Learning Representations (ICLR). 2022.
  38. 38.Albert Gu, Caglar Gulcehre, Tom Le Paine, Matt Hoffman, and Razvan Pascanu. “Improving the Gating Mechanism of Recurrent Neural Networks”. In: The International Conference on Machine Learning (ICML). 2020.
  39. 39.Albert Gu, Ankit Gupta, Karan Goel, and Christopher Ré. “On the Parameterization and Initialization of Diagonal State Space Models”. In: Advances in Neural Information Processing Systems (NeurIPS). 2022.
  40. 40.Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. “Combining Recurrent, Convolutional, and Continuous-time Models with the Linear State Space Layer”. In: Advances in Neural Information Processing Systems (NeurIPS). 2021.
  41. 41.Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré. “How to Train Your HIPPO: State Space Models with Generalized Basis Projections”. In: The International Conference on Learning Representations (ICLR). 2023.
  42. 42.Ankit Gupta, Albert Gu, and Jonathan Berant. “Diagonal State Spaces are as Effective as Structured State Spaces”. In: Advances in Neural Information Processing Systems 35 (2022), pp. 22982–22994.
  43. 43.Ankit Gupta, Harsh Mehta, and Jonathan Berant. “Simplifying and Understanding State Space Models with Diagonal Linear RNNs”. In: arXiv preprint arXiv:2212.00768 (2022).
  44. 44.David Ha, Andrew Dai, and Quoc V. Le. “HyperNetworks”. In: The International Conference on Learning Representations (ICLR). 2017.
  45. 45.Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. “Dream to Control: Learning Behaviors by Latent Imagination”. In: The International Conference on Learning Representations (ICLR). 2020.
  46. 46.Ramin Hasani, Mathias Lechner, Tsun-Hsuan Wang, Makram Chahine, Alexander Amini, and Daniela Rus. “Liquid Structural State-Space Models”. In: The International Conference on Learning Representations (ICLR). 2023.
  47. 47.Mikael Henaff, Arthur Szlam, and Yann LeCun. “Recurrent Orthogonal Networks and Long-Memory Tasks”. In: The International Conference on Machine Learning (ICML). 2016.
  48. 48.Dan Hendrycks and Kevin Gimpel. “Gaussian Error Linear Units (GELUs)”. In: arXiv preprint arXiv:1606.08415 (2016).
  49. 49.Sepp Hochreiter. “Untersuchungen zu dynamischen neuronalen Netzen”. In: Diploma, Technische Universität München 91.1 (1991), p. 31.
  50. 50.Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al. Gradient Flow in Recurrent Nets: The Difficulty of Learning Long-term Dependencies. 2001.
  51. 51.Sepp Hochreiter and Jürgen Schmidhuber. “Long Short-Term Memory”. In: Neural Computation 9.8 (1997), pp. 1735–1780.
  52. 52.Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. “An Empirical Analysis of Compute-Optimal Large Language Model Training”. In: Advances in Neural Information Processing Systems (NeurIPS) 35 (2022), pp. 30016–30030.
  53. 53.Weizhe Hua, Zihang Dai, Hanxiao Liu, and Quoc Le. “Transformer Quality in Linear Time”. In: The International Conference on Machine Learning (ICML). PMLR. 2022, pp. 9099–9117.
  54. 54.Hassan Ismail Fawaz, Germain Forestier, Jonathan Weber, Lhassane Idoumghar, and Pierre-Alain Muller. “Deep Learning for Time Series Classification: A Review”. In: Data Mining and Knowledge Discovery 33.4 (2019), pp. 917–963.
  55. 55.Andrei Ivanov, Nikoli Dryden, Tal Ben-Nun, Shigang Li, and Torsten Hoefler. “Data Movement is All You Need: A Case Study on Optimizing Transformers”. In: Proceedings of Machine Learning and Systems 3 (2021), pp. 711–732.
  56. 56.Li Jing, Caglar Gulcehre, John Peurifoy, Yichen Shen, Max Tegmark, Marin Soljacic, and Yoshua Bengio. “Gated Orthogonal Recurrent Units: On Learning to Forget”. In: Neural Computation 31.4 (2019), pp. 765–783.
  57. 57.Rudolph Emil Kalman. “A New Approach to Linear Filtering and Prediction Problems”. In: (1960).
  58. 58.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. “Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention”. In: International Conference on Machine Learning. PMLR. 2020, pp. 5156–5165.
  59. 59.Shiva Kaul. “Linear Dynamical Systems as a Core Computational Primitive”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 16808–16820.
  60. 60.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. “DiffWave: A Versatile Diffusion Model for Audio Synthesis”. In: International Conference on Learning Representations. 2021.
  61. 61.Chrysoula Kosma, Giannis Nikolentzos, and Michalis Vazirgiannis. “Time-Parameterized Convolutional Neural Networks for Irregularly Sampled Time Series”. In: arXiv preprint arXiv:2308.03210 (2023).
  62. 62.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. “ImageNet Classification with Deep Convolutional Neural Networks”. In: Advances in Neural Information Processing Systems (NeurIPS) 25 (2012).
  63. 63.Tao Lei. “When Attention Meets Fast Recurrence: Training Language Models with Reduced Compute”. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021, pp. 7633–7648.
  64. 64.Tao Lei, Yu Zhang, Sida I Wang, Hui Dai, and Yoav Artzi. “Simple Recurrent Units for Highly Parallelizable Recurrence”. In: arXiv preprint arXiv:1709.02755 (2017).
  65. 65.Mario Lezcano-Casado and David Martínez-Rubio. “Cheap Orthogonal Constraints in Neural Networks: A Simple Parametrization of the Orthogonal and Unitary Group”. In: The International Conference on Machine Learning (ICML). 2019.
  66. 66.Yuhong Li, Tianle Cai, Yi Zhang, Deming Chen, and Debadeepta Dey. “What Makes Convolutional Models Great on Long Sequence Modeling?” In: The International Conference on Learning Representations (ICLR). 2023.
  67. 67.Vasileios Lioutas and Yuhong Guo. “Time-aware Large Kernel Convolutions”. In: The International Conference on Machine Learning (ICML). PMLR. 2020, pp. 6172–6183.
  68. 68.Chris Lu, Yannick Schroecker, Albert Gu, Emilio Parisotto, Jakob Foerster, Satinder Singh, and Feryal Behbahani. “Structured State Space Models for In-Context Reinforcement Learning”. In: Advances in Neural Information Processing Systems (NeurIPS). 2023.
  69. 69.Shahar Lutati, Itamar Zimerman, and Lior Wolf. “Focus Your Attention (with Adaptive IIR Filters)”. In: arXiv preprint arXiv:2305.14952 (2023).
  70. 70.Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. “Mega: Moving Average Equipped Gated Attention”. In: The International Conference on Learning Representations (ICLR). 2023.
  71. 71.Eric Martin and Chris Cundy. “Parallelizing Linear Recurrent Neural Nets Over Sequence Length”. In: The International Conference on Learning Representations (ICLR). 2018.
  72. 72.Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio. “SampleRNN: An Unconditional End-to-End Neural Audio Generation Model”. In: The International Conference on Learning Representations (ICLR). 2017.
  73. 73.Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur. “Long Range Language Modeling via Gated State Spaces”. In: The International Conference on Learning Representations (ICLR). 2023.
  74. 74.Zakaria Mhammedi, Andrew Hellicar, Ashfaqur Rahman, and James Bailey. “Efficient Orthogonal Parametrisation of Recurrent Neural Networks using Householder Reflections”. In: International Conference on Machine Learning. PMLR. 2017, pp. 2401–2409.
  75. 75.Eric Nguyen, Karan Goel, Albert Gu, Gordon Downs, Preey Shah, Tri Dao, Stephen Baccus, and Christopher Ré. “S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces”. In: Advances in Neural Information Processing Systems (NeurIPS). 2022.
  76. 76.Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas, Callum Birch-Sykes, Michael Wornow, Aman Patel, Clayton Rabideau, Stefano Massaroli, Yoshua Bengio, et al. “HyenaDNA: Long-range Genomic Sequence Modeling at Single Nucleotide Resolution”. In: Advances in Neural Information Processing Systems (NeurIPS). 2023.
  77. 77.Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. “In-context Learning and Induction Heads”. In: Transformer Circuits Thread (2022). https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html.
  78. 78.Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. “WaveNet: A Generative Model for Raw Audio”. In: arXiv preprint arXiv:1609.03499 (2016).
  79. 79.Antonio Orvieto, Samuel L Smith, Albert Gu, Anushan Fernando, Caglar Gulcehre, Razvan Pascanu, and Soham De. “Resurrecting Recurrent Neural Networks for Long Sequences”. In: The International Conference on Machine Learning (ICML). 2023.
  80. 80.Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. “The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context”. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics. 2016, pp. 1525–1534.
  81. 81.Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. “On the Difficulty of Training Recurrent Neural Networks”. In: International Conference on Machine Learning. 2013, pp. 1310–1318.
  82. 82.Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al. “RWKV: Reinventing RNNs for the Transformer Era”. In: arXiv preprint arXiv:2305.13048 (2023).
  83. 83.Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong. “Random Feature Attention”. In: The International Conference on Learning Representations (ICLR). 2021.
  84. 84.Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. “Hyena Hierarchy: Towards Larger Convolutional Language Models”. In: The International Conference on Machine Learning (ICML). 2023.
  85. 85.Zhen Qin, Xiaodong Han, Weixuan Sun, Bowen He, Dong Li, Dongxu Li, Yuchao Dai, Lingpeng Kong, and Yiran Zhong. “Toeplitz Neural Network for Sequence Modeling”. In: The International Conference on Learning Representations (ICLR). 2023.
  86. 86.Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong. “The devil in linear transformer”. In: arXiv preprint arXiv:2210.10340 (2022).
  87. 87.Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong. “CosFormer: Rethinking Softmax in Attention”. In: The International Conference on Learning Representations (ICLR). 2022.
  88. 88.Ali Rahimi and Benjamin Recht. “Random Features for Large-Scale Kernel Machines”. In: Advances in Neural Information Processing Systems (NeurIPS) 20 (2007).
  89. 89.Prajit Ramachandran, Barret Zoph, and Quoc V Le. “Swish: A Self-gated Activation Function”. In: arXiv preprint arXiv:1710.05941 7.1 (2017), p. 5.
  90. 90.David W Romero, Anna Kuzina, Erik J Bekkers, Jakub M Tomczak, and Mark Hoogendoorn. “CKConv: Continuous Kernel Convolution For Sequential Data”. In: arXiv preprint arXiv:2102.02611 (2021).
  91. 91.Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. “Winogrande: An Adversarial Winograd Schema Challenge at Scale”. In: Communications of the ACM 64.9 (2021), pp. 99–106.
  92. 92.George Saon, Ankit Gupta, and Xiaodong Cui. “Diagonal State Space Augmented Transformers for Speech Recognition”. In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE. 2023, pp. 1–5.
  93. 93.Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. “Linear Transformers are Secretly Fast Weight Programmers”. In: The International Conference on Machine Learning (ICML). PMLR. 2021, pp. 9355–9366.
  94. 94.Jürgen Schmidhuber. “Learning to control fast-weight memories: An alternative to dynamic recurrent networks”. In: Neural Computation 4.1 (1992), pp. 131–139.
  95. 95.Noam Shazeer. “GLU Variants Improve Transformer”. In: arXiv preprint arXiv:2002.05202 (2020).
  96. 96.Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. “Large Language Models can be Easily Distracted by Irrelevant Context”. In: The International Conference on Machine Learning (ICML). PMLR. 2023, pp. 31210–31227.
  97. 97.Jiaxin Shi, Ke Alexander Wang, and Emily Fox. “Sequence Modeling with Multiresolution Convolutional Memory”. In: The International Conference on Machine Learning (ICML). PMLR. 2023, pp. 31312–31327.
  98. 98.Jimmy TH Smith, Andrew Warrington, and Scott W Linderman. “Simplified State Space Layers for Sequence Modeling”. In: The International Conference on Learning Representations (ICLR). 2023.
  99. 99.Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. “Roformer: Enhanced Transformer with Rotary Position Embedding”. In: arXiv preprint arXiv:2104.09864 (2021).
  100. 100.Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. “Retentive network: A successor to transformer for large language models”. In: arXiv preprint arXiv:2307.08621 (2023).
  101. 101.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. “Sequence to Sequence Learning with Neural Networks”. In: Advances in Neural Information Processing Systems (NeurIPS) 27 (2014).
  102. 102.Corentin Tallec and Yann Ollivier. “Can Recurrent Neural Networks Warp Time?” In: The International Conference on Learning Representations (ICLR). 2018.
  103. 103.Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler. “Long Range Arena: A Benchmark for Efficient Transformers”. In: International Conference on Learning Representations (ICLR). 2021.
  104. 104.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. “Efficient Transformers: A Survey”. In: ACM Computing Surveys 55.6 (2022), pp. 1–28.
  105. 105.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. “Llama: Open and Efficient Foundation Language Models”. In: arXiv preprint arXiv:2302.13971 (2023).
  106. 106.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. “Attention Is All You Need”. In: Advances in Neural Information Processing Systems (NeurIPS). 2017.
  107. 107.Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal. “On Orthogonality and Learning Recurrent Networks with Long Term Dependencies”. In: International Conference on Machine Learning. PMLR. 2017, pp. 3570–3578.
  108. 108.Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar, and Raffay Hamid. “Selective Structured State-Spaces for Long-form Video Understanding”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023, pp. 6387–6397.
  109. 109.Pete Warden. “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition”. In: ArXiv abs/1804.03209 (2018).
  110. 110.Samuel Williams, Andrew Waterman, and David Patterson. “Roofline: An Insightful Visual Performance Model for Multicore Architectures”. In: Communications of the ACM 52.4 (2009), pp. 65–76.
  111. 111.Brandon Yang, Gabriel Bender, Quoc V Le, and Jiquan Ngiam. “CondConv: Conditionally Parameterized Convolutions for Efficient Inference”. In: Advances in Neural Information Processing Systems (NeurIPS) 32 (2019).
  112. 112.Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. “HellaSwag: Can a Machine Really Finish Your Sentence?” In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
  113. 113.Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind. “An Attention Free Transformer”. In: arXiv preprint arXiv:2105.14103 (2021).
  114. 114.Michael Zhang, Khaled K Saab, Michael Poli, Tri Dao, Karan Goel, and Christopher Ré. “Effectively Modeling Time Series with Simple Discrete State Spaces”. In: The International Conference on Learning Representations (ICLR). 2023.
  115. 115.Lin Zheng, Chong Wang, and Lingpeng Kong. “Linear complexity randomized self-attention mechanism”. In: International Conference on Machine Learning. PMLR. 2022, pp. 27011–27041.
  116. 116.Simiao Zuo, Xiaodong Liu, Jian Jiao, Denis Charles, Eren Manavoglu, Tuo Zhao, and Jianfeng Gao. “Efficient Long Sequence Modeling via State Space Augmented Transformer”. In: arXiv preprint arXiv:2212.08136 (2022).

Citation

MLA
Gu, A., and T. Dao. “Mamba: Linear-Time Sequence Modeling with Selective State Spaces”. arXiv, 2023, http://arxiv.org/abs/2312.00752v2.
APA
Gu, A., & Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv. http://arxiv.org/abs/2312.00752v2
Chicago
Gu, A., and T. Dao. 2023. “Mamba: Linear-Time Sequence Modeling with Selective State Spaces”. arXiv. http://arxiv.org/abs/2312.00752v2.
Harvard
Gu, A. and Dao, T. (2023) “Mamba: Linear-Time Sequence Modeling with Selective State Spaces”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.00752v2.
Vancouver
1. Gu A, Dao T (2023) Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv

BibTeX

@article{gu2023mamba,
  title = {Mamba: Linear-Time Sequence Modeling with Selective State Spaces},
  author = {Gu, Albert and Dao, Tri},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.00752v2},
  eprint = {2312.00752}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/