S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces
Eric NguyenKaran GoelAlbert GuGordon W. DownsPreey ShahTri DaoStephen BaccusChristopher Ré
Extends deep state space models to multidimensional signals with S4ND, enabling continuous-signal modeling of images and videos that matches or exceeds standard vision architectures while supporting zero-shot resolution adaptation and faster training via progressive resizing.
Modern computer vision models typically treat visual data as discrete pixels and local patches rather than underlying continuous signals. While effective on benchmark tasks, these discrete architectures struggle to adapt across varying resolutions and sample rates without extensive retraining. The article addresses this fundamental limitation by asking how deep state space models (SSMs)—which have demonstrated strong performance on continuous, one-dimensional sequence data such as audio—can be effectively generalized to handle multidimensional visual signals in images and videos.
The main objective of the article is to introduce and evaluate S4ND, a multidimensional state space layer designed to process visual data as continuous signals in one, two, and three dimensions. The authors aim to demonstrate that S4ND can serve as a drop-in replacement for standard self-attention and convolutional layers in leading vision backbones, matching or exceeding their performance while providing continuous-signal capabilities such as zero-shot resolution adaptation.
To evaluate this framework, the authors integrated S4ND into established vision architectures, replacing self-attention layers in Vision Transformers (ViT) and two-dimensional convolutional layers in ConvNeXt. The models were tested on standard large-scale and benchmark datasets: ImageNet-1k (1.3 million images) for large-scale image classification, HMDB-51 for 3D video activity classification, and CIFAR-10 and Celeb-A for resolution-transfer experiments. A key technical element introduced is a low-pass bandlimiting mechanism that filters out frequencies above the Nyquist cutoff, mitigating visual aliasing artifacts when moving across different spatial scales.
The experimental findings show clear advantages. First, S4ND improves or matches state-of-the-art architectures: replacing self-attention in a base ViT with S4ND boosted top-1 ImageNet accuracy by 1.5% (reaching 80.4%), while matching ConvNeXt performance (82.2%). Second, in 3D video classification on HMDB-51, inflating a pretrained 2D S4ND backbone to 3D outperformed the inflated ConvNeXt baseline by 4.0% top-1 accuracy (achieving 62.1% using only RGB data). Third, S4ND demonstrated strong zero-shot generalization to unseen resolutions: when trained on 8x8 images and tested on 32x32 images on CIFAR-10, S4ND outperformed a standard two-dimensional convolutional network by over 40% accuracy. Finally, using progressive resizing during training allowed S4ND to train 21.8% faster while remaining within approximately 1% accuracy of models trained purely at full resolution.
These results demonstrate that visual models do not need to rely on rigid discrete tokenization or strictly local convolution windows to achieve top-tier performance. By modeling continuous underlying signals, S4ND provides global spatial and temporal context at every layer while naturally handling multi-resolution workflows. This reduces the risk of performance collapse when deploying visual systems in operational environments where sensor resolutions, camera distances, or frame rates vary dynamically.
For practical implementation, engineering and research teams should evaluate S4ND layers when building vision systems requiring multi-resolution resilience, progressive training acceleration, or multimodal continuous data processing across audio, image, and video. Future efforts should prioritize optimizing hardware-level GPU implementations, such as fusing memory operations, to eliminate current runtime bottlenecks. Additionally, further research should test S4ND on larger video benchmarks and explore its zero-shot capabilities across non-uniform and varied temporal frame rates.
- Paper: Efficiently Modeling Long Sequences with Structured State Spaces, Albert Gu et al. (2022). This paper introduces the foundational 1D Structured State Space (S4) sequence model and continuous-time HiPPO parameterization that S4ND directly generalizes to multidimensional visual data.
- Paper: Swin Transformer: Hierarchical Vision Transformer using Shifted Windows, Ze Liu et al. (2021). This work establishes hierarchical vision backbone principles and benchmarks that S4ND compares against and aims to improve with continuous global modeling.
- Paper: Is Space-Time Attention All You Need for Video Understanding?, Gedas Bertasius et al. (2021). This paper establishes pure space-time attention mechanisms for video understanding, which S4ND replaces with multidimensional continuous state space layers.
- Paper: ViViT: A Video Vision Transformer, Anurag Arnab et al. (2021). This study details factorized spatio-temporal token processing in vision transformers, framing the baseline approaches S4ND seeks to transcend via continuous signals.
- Paper: Non-local Neural Networks, Xiaolong Wang et al. (2018). This foundational paper presents non-local operations for capturing long-range spatio-temporal dependencies in visual signals, motivating S4ND's global receptive field design.
- Paper: Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification, Saining Xie et al. (2017). This work investigates speed-accuracy trade-offs and dimension factorizations in 3D spatio-temporal video networks that inform S4ND's multidimensional architecture.
- Paper: Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model, Lianghui Zhu et al. (2024). This work extends state space modeling in visual backbones by introducing bidirectional selective state space models (Vision Mamba) as an alternative to multidimensional S4 formulations.
- Paper: VMamba: Visual State Space Model, Yue Liu et al. (2024). This paper builds on visual state space architectures by proposing a 2D selective scan mechanism to traverse spatial feature maps with linear complexity.
