The Illusion of State in State-Space Models
William MerrillJackson PettyAshish Sabharwal
Proves that popular state-space models like S4 and Mamba share the same fundamental expressive limitations as transformers for sequential state tracking, while identifying a minimal architectural modification to overcome this barrier.
State-space models, such as S4 and Mamba, have recently gained significant attention as potential alternatives to transformer architectures for foundation models. A central motivation for these architectures is their recurrent structure, which many believed would overcome the inability of transformers to perform sequential reasoning and track state changes over time. Because core capabilities such as tracking entities in long narratives, executing computer code, and tracking game states like chess depend fundamentally on sequential state tracking, determining whether these models genuinely offer superior state-tracking power is a critical question for machine learning system design.
The article evaluates the mathematical expressiveness and empirical capabilities of popular state-space models for state-tracking tasks. Specifically, it demonstrates whether standard state-space architectures can solve inherently sequential problems that standard recurrent neural networks can naturally handle.
To assess these capabilities, the analysis uses circuit complexity theory alongside formal algebraic language theory to classify the computational limits of fixed-depth state-space models using standard floating-point precision. The theoretical findings are evaluated through controlled empirical experiments comparing transformers, classic recurrent neural networks, and several state-space model variants on sequence-tagging tasks involving algebraic permutation composition across varying sequence lengths.
The findings establish that the recurrent "state" in common state-space models is an illusion with respect to computational expressiveness. First, non-gated models such as S4 and selective diagonal models such as Mamba belong to the same complexity class as transformers (L-uniform TC0), meaning they are provably unable to solve inherently sequential state-tracking problems such as permutation composition. Second, empirical tests confirm that both transformers and standard state-space models fail to maintain state across long sequences without scaling the number of layers linearly with sequence length. In contrast, standard recurrent neural networks solve these tasks across arbitrary lengths using a single layer. Third, the article demonstrates that modifying state-space models to use input-dependent, non-diagonal transition matrices restores full state-tracking capabilities, allowing a single layer to solve permutation tasks while preserving efficient parallel training.
These results demonstrate that switching from transformers to current state-space architectures will not resolve fundamental state-tracking limitations in tasks like multi-step reasoning, program execution, or narrative tracking. Deploying current state-space models in domains that require exact sequential state updates carries the operational risk of silent tracking failures unless input sequences are short or model depth is significantly increased. However, the success of input-dependent transition matrices indicates that the architectural design space can be expanded to achieve both expressiveness and computational efficiency.
Organizations evaluating alternative architectures should avoid selecting current state-space models solely under the assumption that they provide superior state tracking over transformers. Research and engineering teams should instead pilot more expressive variants, such as input-dependent state-space architectures, while monitoring training stability and hardware-level parallelism trade-offs before committing to large-scale deployment.
These conclusions assume standard fixed-depth architectures operating with realistic precision constraints and rely on established computational complexity separations. While confidence in the formal proofs and synthetic empirical results is high, further validation on full-scale language pretraining benchmarks is required to determine whether enhanced state-space architectures retain their advantages in broader real-world settings.
- Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). Read the S4 paper first to understand the structured state-space architecture whose expressive limits this work analyzes.
- Paper: Mamba: Linear-Time Sequence Modeling with Selective State Spaces, Albert Gu et al. (2023). Mamba introduces the selective SSM whose ability to track information is directly tested and bounded in the source.
- Paper: Tighter Bounds on the Expressivity of Transformer Encoders, David Chiang et al. (2023). Its formal bounds on transformer expressivity provide the circuit-complexity framework that the source extends to SSMs.
- Paper: Hungry Hungry Hippos: Towards Language Modeling with State Space Models, Daniel Y. Fu et al. (2023). H3’s diagnostic tests of SSM memory and sequence comparison establish relevant precedents for the source’s state-tracking analysis.
- Paper: Mamba-3: Improved Sequence Modeling using State Space Principles, Aakash Lahoti et al. (2026). Mamba-3 develops newer SSM mechanisms for state tracking, making it a direct architectural continuation to assess in light of the source’s limits.
- Paper: Learning to (Learn at Test Time): RNNs with Expressive Hidden States, Yu Sun 0020 et al. (2025). TTT replaces a fixed recurrent state with an expressive learned model, pursuing the stronger sequence memory the source motivates.
- Paper: The Topological Trouble With Transformers, Michael C. Mozer et al. (2026). This later analysis revisits state-tracking barriers and recurrent alternatives, extending the source’s critique toward broader architecture design.
