Hungry Hungry Hippos: Towards Language Modeling with State Space Models
Daniel Y. FuTri DaoKhaled Kamal SaabArmin W. ThomasAtri RudraChristopher Ré
Introduces the H3 state space architecture and FlashConv algorithm, enabling sub-quadratic language models scaled up to 2.7 billion parameters that surpass Transformers on SuperGLUE benchmarks while generating text 2.4 times faster.
Modern natural language processing and artificial intelligence rely heavily on Transformer models driven by attention mechanisms. However, standard attention scales quadratically with sequence length, creating severe computational and memory bottlenecks when handling long inputs. State space models offer a mathematically efficient alternative that scales near-linearly, but they have historically lagged behind Transformers in language modeling accuracy and suffered from poor hardware utilization on modern accelerators.
The article evaluates the root causes of this expressivity gap and demonstrates new model architectures and computational algorithms that enable state space models to match or exceed Transformer quality while dramatically improving execution speed.
To conduct this evaluation, the authors analyzed state space model failures on synthetic language tasks designed to test key-value associative recall and token comparison capabilities. Building on these diagnostic insights, they introduced a novel architecture named Hungry Hungry Hippo (H3), which stacks shift and diagonal state space layers with multiplicative gating to capture sequential dependencies. To eliminate hardware memory bottlenecks, they designed FlashConv, an algorithm combining fused block matrix-multiplication operations for fast Fourier transform convolutions with a chunked state-passing method for processing long sequences. The authors trained and tested pure H3 and hybrid models—retaining just two standard attention layers—on benchmark corpora including OpenWebText, the 800-gigabyte Pile dataset across scales from 125 million to 2.7 billion parameters, SuperGLUE benchmarks, and non-text modalities such as audio, electroencephalography, and functional magnetic resonance imaging.
The investigation produced four primary findings. First, diagnostic testing revealed that standard state space models struggle to remember tokens appearing after specific events and compare items across a sequence, but the H3 layer fully resolves these synthetic benchmarks. Second, pure H3 closed the language modeling perplexity gap on OpenWebText to within 0.4 points of Transformers, while a hybrid H3-attention model outperformed Transformers by 1.0 perplexity point. Third, when scaled up to 2.7 billion parameters on the Pile, hybrid models achieved lower perplexity and superior zero-shot and few-shot accuracy over Transformer baselines across a majority of SuperGLUE tasks. Fourth, the FlashConv implementation delivered a 2x speedup on long-range sequence benchmarks, accelerated long-sequence training by 4x to 8x over standard attention, and enabled 2.4x higher text generation throughput during inference.
These findings demonstrate that quadratic self-attention is not strictly necessary throughout deep networks to achieve state-of-the-art language understanding. Deploying hybrid architectures drastically lowers the operational compute costs, hardware memory footprints, and latency of running foundation models, unlocking practical real-time inference and the processing of ultra-long contexts such as continuous audio and medical sensor data.
Organizations developing large language and sequence models should explore hybrid state space architectures to reduce serving costs and accelerate data throughput. Decision-makers should consider piloting hybrid models for high-throughput text generation and long-context applications where standard Transformers encounter hardware memory limits. Further exploration is recommended to optimize layer placement and investigate whether similar state space mechanisms can fully replace attention across other multimodal pipelines.
Readers should note that while hybrid models demonstrated robust gains, the authors did not extensively tune training hyperparameters specifically for state space models, relying instead on standard GPT configurations. Additionally, pure H3 models exhibited performance degradation on certain generative zero-shot evaluation formats without prompt examples. Confidence in the reported scalability and inference speedups remains high based on comprehensive evaluations spanning multiple parameter scales and diverse sequential data domains.
- Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). Introduces the structured state space (S4) framework and continuous-time HiPPO parameterization that H3 diagnoses and adapts specifically for language modeling.
- Paper: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness, Tri Dao et al. (2022). Establishes the IO-aware hardware optimization principles and fused kernel designs directly extended by H3's FlashConv algorithm.
- Paper: Attention Is All You Need, Ashish Vaswani et al. (2017). Presents the foundational Transformer architecture and quadratic attention mechanism that H3 benchmarks against and partially retains in hybrid configurations.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). Directly succeeds H3 by introducing input-dependent selective state spaces to solve associative recall natively without requiring hybrid attention layers.
- Paper: Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Tri Dao et al. (2024). Unifies state space models and attention mechanisms through structured matrix duality, advancing the theoretical foundations behind H3 and selective SSMs.
- Paper: Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling, Liliang Ren et al. (2025). Continues H3's hybrid sequence modeling strategy by combining selective state space blocks with localized sliding-window attention for long-context language tasks.
- Paper: Mamba-3: Improved Sequence Modeling using State Space Principles, Aakash Lahoti et al. (2026). Pushes the boundaries of state space sequence modeling further by introducing refined discretization and complex state updates building upon the H3 and Mamba lineage.
