S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces

Eric NguyenKaran GoelAlbert GuGordon W. DownsPreey ShahTri DaoStephen BaccusChristopher Ré

article2022NeurIPS267 citations

Extends deep state space models to multidimensional signals with S4ND, enabling continuous-signal modeling of images and videos that matches or exceeds standard vision architectures while supporting zero-shot resolution adaptation and faster training via progressive resizing.

Listen

Modern computer vision models typically treat visual data as discrete pixels and local patches rather than underlying continuous signals. While effective on benchmark tasks, these discrete architectures struggle to adapt across varying resolutions and sample rates without extensive retraining. The article addresses this fundamental limitation by asking how deep state space models (SSMs)—which have demonstrated strong performance on continuous, one-dimensional sequence data such as audio—can be effectively generalized to handle multidimensional visual signals in images and videos.

The main objective of the article is to introduce and evaluate S4ND, a multidimensional state space layer designed to process visual data as continuous signals in one, two, and three dimensions. The authors aim to demonstrate that S4ND can serve as a drop-in replacement for standard self-attention and convolutional layers in leading vision backbones, matching or exceeding their performance while providing continuous-signal capabilities such as zero-shot resolution adaptation.

To evaluate this framework, the authors integrated S4ND into established vision architectures, replacing self-attention layers in Vision Transformers (ViT) and two-dimensional convolutional layers in ConvNeXt. The models were tested on standard large-scale and benchmark datasets: ImageNet-1k (1.3 million images) for large-scale image classification, HMDB-51 for 3D video activity classification, and CIFAR-10 and Celeb-A for resolution-transfer experiments. A key technical element introduced is a low-pass bandlimiting mechanism that filters out frequencies above the Nyquist cutoff, mitigating visual aliasing artifacts when moving across different spatial scales.

The experimental findings show clear advantages. First, S4ND improves or matches state-of-the-art architectures: replacing self-attention in a base ViT with S4ND boosted top-1 ImageNet accuracy by 1.5% (reaching 80.4%), while matching ConvNeXt performance (82.2%). Second, in 3D video classification on HMDB-51, inflating a pretrained 2D S4ND backbone to 3D outperformed the inflated ConvNeXt baseline by 4.0% top-1 accuracy (achieving 62.1% using only RGB data). Third, S4ND demonstrated strong zero-shot generalization to unseen resolutions: when trained on 8x8 images and tested on 32x32 images on CIFAR-10, S4ND outperformed a standard two-dimensional convolutional network by over 40% accuracy. Finally, using progressive resizing during training allowed S4ND to train 21.8% faster while remaining within approximately 1% accuracy of models trained purely at full resolution.

These results demonstrate that visual models do not need to rely on rigid discrete tokenization or strictly local convolution windows to achieve top-tier performance. By modeling continuous underlying signals, S4ND provides global spatial and temporal context at every layer while naturally handling multi-resolution workflows. This reduces the risk of performance collapse when deploying visual systems in operational environments where sensor resolutions, camera distances, or frame rates vary dynamically.

For practical implementation, engineering and research teams should evaluate S4ND layers when building vision systems requiring multi-resolution resilience, progressive training acceleration, or multimodal continuous data processing across audio, image, and video. Future efforts should prioritize optimizing hardware-level GPU implementations, such as fusing memory operations, to eliminate current runtime bottlenecks. Additionally, further research should test S4ND on larger video benchmarks and explore its zero-shot capabilities across non-uniform and varied temporal frame rates.

arXiv: 2210.06583
Cover for S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces

Abstract

Visual data such as images and videos are typically modeled as discretizations of inherently continuous, multidimensional signals. Existing continuous-signal models attempt to exploit this fact by modeling the underlying signals of visual (e.g., image) data directly. However, these models have not yet been able to achieve competitive performance on practical vision tasks such as large-scale image and video classification. Building on a recent line of work on deep state space models (SSMs), we propose S4ND, a new multidimensional SSM layer that extends the continuous-signal modeling ability of SSMs to multidimensional data including images and videos. We show that S4ND can model large-scale visual data in 1D, 2D, and 3D as continuous multidimensional signals and demonstrates strong performance by simply swapping Conv2D and self-attention layers with S4ND layers in existing state-of-the-art models. On ImageNet-1k, S4ND exceeds the performance of a Vision Transformer baseline by 1.5% when training with a 1D sequence of patches, and matches ConvNeXt when modeling images in 2D. For videos, S4ND improves on an inflated 3D ConvNeXt in activity classification on HMDB-51 by 4%. S4ND implicitly learns global, continuous convolutional kernels that are resolution invariant by construction, providing an inductive bias that enables generalization across multiple resolutions. By developing a simple bandlimiting modification to S4 to overcome aliasing, S4ND achieves strong zero-shot (unseen at training time) resolution performance, outperforming a baseline Conv2D by 40% on CIFAR-10 when trained on 8 × 8 and tested on 32 × 32 images. When trained with progressive resizing, S4ND comes within ~1% of a high-resolution model while training 22% faster.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 Method
  • 4.1 S4ND
  • 4.2 Resolution Change and Bandlimiting
  • 5 Experiments
  • 5.1 S4ND in 1D & 2D: Large-scale Image Classification
  • 5.2 S4ND in 3D: Video Classification
  • 5.3 Continuous-signal Capabilities for Images
  • 6 Discussion
  • Acknowledgments
  • References
  • Checklist

Knowls

  1. Knowl 1 — Multidimensional State Space Model PDE Formulation

    definition

    The multidimensional state space model (SSM) extends 1D continuous-time linear ODE systems to multidimensional continuous domains (such as 2D spatial grids or 3D spatiotemporal volumes). For a 2D scalar continuous input signal u(t(1),t(2))∈Cu(t^{(1)}, t^{(2)}) \in \mathbb{C} and output signal y(t(1),t(2))∈Cy(t^{(1)}, t^{(2)}) \in \mathbb{C}, the 2D SSM is defined by the linear system of partial differential equations (PDEs) with boundary condition x(0,0)=0x(0, 0) = 0:

    ∂∂t(1)x(t(1),t(2))=(A(1)x(1)(t(1),t(2)),x(2)(t(1),t(2)))+B(1)u(t(1),t(2))\frac{\partial}{\partial t^{(1)}} x(t^{(1)}, t^{(2)}) = \left(A^{(1)} x^{(1)}(t^{(1)}, t^{(2)}), x^{(2)}(t^{(1)}, t^{(2)})\right) + B^{(1)} u(t^{(1)}, t^{(2)})

    ∂∂t(2)x(t(1),t(2))=(x(1)(t(1),t(2)),A(2)x(2)(t(1),t(2)))+B(2)u(t(1),t(2))\frac{\partial}{\partial t^{(2)}} x(t^{(1)}, t^{(2)}) = \left(x^{(1)}(t^{(1)}, t^{(2)}), A^{(2)} x^{(2)}(t^{(1)}, t^{(2)})\right) + B^{(2)} u(t^{(1)}, t^{(2)})

    y(t(1),t(2))=⟨C,x(t(1),t(2))⟩y(t^{(1)}, t^{(2)}) = \langle C, x(t^{(1)}, t^{(2)}) \rangle

    where t(1),t(2)∈Rt^{(1)}, t^{(2)} \in \mathbb{R} denote continuous coordinates along each axis, x(t(1),t(2))∈CN(1)×N(2)x(t^{(1)}, t^{(2)}) \in \mathbb{C}^{N^{(1)} \times N^{(2)}} is the multidimensional state tensor with component state maps x(τ):R2→CN(τ)x^{(\tau)} : \mathbb{R}^2 \to \mathbb{C}^{N^{(\tau)}}, A(τ)∈CN(τ)×N(τ)A^{(\tau)} \in \mathbb{C}^{N^{(\tau)} \times N^{(\tau)}} and B(τ)∈CN(τ)×1B^{(\tau)} \in \mathbb{C}^{N^{(\tau)} \times 1} are state transition and input matrices for axis τ∈{1,2}\tau \in \{1, 2\}, and C∈CN(1)×N(2)C \in \mathbb{C}^{N^{(1)} \times N^{(2)}} is the output state projection tensor.

  2. Knowl 2 — Equivalence of Multidimensional SSMs to Multidimensional Convolutions

    theoretical result

    The 2D linear time-invariant state space PDE system with parameters (A(1),B(1))(A^{(1)}, B^{(1)}), (A(2),B(2))(A^{(2)}, B^{(2)}), and CC is mathematically equivalent to a continuous 2D convolution y(t(1),t(2))=(K∗u)(t(1),t(2))y(t^{(1)}, t^{(2)}) = (K * u)(t^{(1)}, t^{(2)}) governed by the multidimensional continuous kernel:

    K(t(1),t(2))=⟨C,(et(1)A(1)B(1))⊗(et(2)A(2)B(2))⟩K(t^{(1)}, t^{(2)}) = \langle C, (e^{t^{(1)} A^{(1)}} B^{(1)}) \otimes (e^{t^{(2)} A^{(2)}} B^{(2)}) \rangle

    where ⊗\otimes denotes the outer (tensor) product. This kernel constitutes a linear combination of N(1)×N(2)N^{(1)} \times N^{(2)} continuous basis kernels:

    {Kn(1)(1)(t(1))⊗Kn(2)(2)(t(2)):n(1)∈[N(1)],n(2)∈[N(2)]}\{K_{n^{(1)}}^{(1)}(t^{(1)}) \otimes K_{n^{(2)}}^{(2)}(t^{(2)}) : n^{(1)} \in [N^{(1)}], n^{(2)} \in [N^{(2)}]\}

    where each Kn(τ)(τ)(t(τ))=(et(τ)A(τ)B(τ))n(τ)K_{n^{(\tau)}}^{(\tau)}(t^{(\tau)}) = (e^{t^{(\tau)} A^{(\tau)}} B^{(\tau)})_{n^{(\tau)}} represents the standard 1D continuous basis kernel defined along axis τ\tau.

  3. Knowl 3 — Low-Rank Tensor Factorization of S4ND Convolution Kernels

    model/method

    To eliminate exponential parameter and computational scaling across dimensions (N(1)×N(2)×⋯×N(d)N^{(1)} \times N^{(2)} \times \dots \times N^{(d)}), the state output projection tensor C∈CN(1)×⋯×N(d)C \in \mathbb{C}^{N^{(1)} \times \dots \times N^{(d)}} is factorized into a rank-rr tensor decomposition:

    C=∑i=1rCi(1)⊗Ci(2)⊗⋯⊗Ci(d)C = \sum_{i=1}^r C_i^{(1)} \otimes C_i^{(2)} \otimes \dots \otimes C_i^{(d)}

    where each Ci(τ)∈CN(τ)C_i^{(\tau)} \in \mathbb{C}^{N^{(\tau)}}. Under this decomposition, the continuous multidimensional convolution kernel factors directly into a sum of 1D continuous state space kernels:

    K(t(1),…,t(d))=∑i=1r⨂τ=1d(Ci(τ)et(τ)A(τ)B(τ))=∑i=1rKi(1)(t(1))⊗⋯⊗Ki(d)(t(d))K(t^{(1)}, \dots, t^{(d)}) = \sum_{i=1}^r \bigotimes_{\tau=1}^d \left( C_i^{(\tau)} e^{t^{(\tau)} A^{(\tau)}} B^{(\tau)} \right) = \sum_{i=1}^r K_i^{(1)}(t^{(1)}) \otimes \dots \otimes K_i^{(d)}(t^{(d)})

    Setting rank r=1r=1 corresponds to maintaining an independent 1D S4 layer per dimension. To apply the multidimensional convolution over discrete grids, 1D kernels of length L(τ)L^{(\tau)} are computed independently via 1D S4 at step size Δ(τ)\Delta^{(\tau)}, multiplied via an outer product to form a global kernel spanning the full multidimensional field, and applied depthwise using the fast Fourier transform (FFT).

  4. Knowl 4 — S4ND Low-Pass Bandlimiting for Resolution Invariance

    model/method

    When transferring continuous convolution kernels across different spatial or temporal sampling rates, changing the step size Δ\Delta can introduce aliasing if high-frequency components exceed the Nyquist cutoff frequency.

    For a diagonal state matrix A=diag(a1,…,aN)A = \mathrm{diag}(a_1, \dots, a_N), the nn-th continuous basis kernel is Kn(t)=etanBnK_n(t) = e^{t a_n} B_n, with oscillation frequency determined by the imaginary part Im(an)\mathrm{Im}(a_n). S4ND applies a low-pass bandlimiting filter by masking the coefficients of CC:

    Cn←0if an⋅Δ≥12αC_n \leftarrow 0 \quad \text{if } a_n \cdot \Delta \ge \frac{1}{2\alpha}

    where α>0\alpha > 0 is a cutoff hyperparameter. In the theoretical case of undamped sinusoidal basis functions, α=1.0\alpha = 1.0 matches the exact Nyquist cutoff frequency; empirically, α\alpha is set lower (e.g., α≈0.4–0.5\alpha \approx 0.4\text{--}0.5) to compensate for exponential decay factors eRe(an)e^{\mathrm{Re}(a_n)} and finite-state truncation approximations.

  5. Knowl 5 — Drop-in Image Classification Performance of S4ND on ImageNet, CIFAR-10, and Celeb-A

    empirical result

    S4ND serves as a drop-in replacement for 1D multi-head self-attention in Vision Transformers (ViT) and 2D local depthwise convolutions in ConvNeXt architectures without requiring positional encodings in ViT.

    Model Dataset Parameters Top-1 Accuracy (%)
    ViT-B ImageNet-1k 88.0M 78.9
    S4ND-ViT-B ImageNet-1k 88.8M 80.4
    ConvNeXt-T ImageNet-1k 28.4M 82.1
    S4ND-ConvNeXt-T ImageNet-1k 30.0M 82.2
    Conv2D-ISO CIFAR-10 2.2M 93.7
    S4ND-ISO CIFAR-10 5.3M 94.1
    ConvNeXt-M Celeb-A 9.2M 91.0
    S4ND-ConvNeXt-M Celeb-A 9.6M 91.3

    On ImageNet-1k trained from scratch for 300 epochs, S4ND-ViT-B achieves an 80.4% top-1 accuracy, outperforming the baseline ViT-B by 1.5 percentage points. S4ND-ConvNeXt-T achieves 82.2% top-1 accuracy, matching standard 2D ConvNeXt-T (82.1%).

  6. Knowl 6 — 3D Spatiotemporal Video Action Recognition on HMDB-51

    empirical result

    A 2D S4ND model pretrained on ImageNet can be inflated into a 3D spatiotemporal model by freezing or reusing spatial parameters (A(1),B(1),C(1))(A^{(1)}, B^{(1)}, C^{(1)}) and (A(2),B(2),C(2))(A^{(2)}, B^{(2)}, C^{(2)}) and initializing independent 1D temporal SSM parameters (A(3),B(3),C(3))(A^{(3)}, B^{(3)}, C^{(3)}) from scratch. Tested on HMDB-51 using 30-frame RGB clips (224×224224 \times 224, 2 seconds duration) without optical flow:

    Model Parameters Modality Top-1 Accuracy (%)
    Inception-I3D 25.0M Optical Flow 61.9
    Inception-I3D 25.0M RGB 49.8
    ConvNeXt-I3D 28.5M RGB 58.1
    ConvNeXt-S3D 27.9M RGB 58.6
    S4ND-ConvNeXt-3D 31.4M RGB 62.1

    S4ND-ConvNeXt-3D reaches 62.1% accuracy on RGB inputs, exceeding ConvNeXt-I3D by 4.0 percentage points, separable ConvNeXt-S3D by 3.5 percentage points, and Inception-I3D trained on optical flow (61.9%).

  7. Knowl 7 — Effect of Initial Temporal Kernel Length on 3D Action Classification

    empirical result

    In 3D S4ND video models, temporal convolution kernels span the entire temporal sequence (e.g., all 30 frames), while the initial receptive scale is governed by the SSM discretization step size Δ(3)\Delta^{(3)}, where the expected initial kernel timescale is proportional to 1/Δ(3)1/\Delta^{(3)}.

    Evaluating S4ND-ConvNeXt-3D on HMDB-51 activity classification across different initial temporal kernel length configurations:

    Initial Kernel Length (1/Δ1/\Delta) Top-1 Accuracy (%)
    20.0 53.74
    4.0 58.33
    2.0 60.30
    1.0 62.07

    Initializing temporal kernels with an expected length of 1.01.0 frame yields the highest accuracy (62.07%), allowing the continuous parameterization to progressively adapt and capture long-range temporal dependencies across the full clip duration.

  8. Knowl 8 — Zero-Shot Cross-Resolution Generalization in 2D Vision Models

    empirical result

    S4ND enables zero-shot testing on image resolutions unseen during training by adjusting the continuous kernel sampling parameter Δ\Delta to match the test grid size. Results are averaged over 2 random seeds:

    Resolution CIFAR-10 Accuracy (%) Celeb-A Accuracy (%)
    Train Test S4ND Conv2D FlexNet-16 S4ND Conv2D
    base base 93.10±0.2293.10 \pm 0.22 91.9±0.291.9 \pm 0.2 92.2±0.192.2 \pm 0.1 91.75±0.0091.75 \pm 0.00 91.44±0.0391.44 \pm 0.03
    mid mid 88.80±0.1288.80 \pm 0.12 87.2±0.187.2 \pm 0.1 86.5±2.086.5 \pm 2.0 91.63±0.0491.63 \pm 0.04 91.09±0.0891.09 \pm 0.08
    mid base 88.77±0.0388.77 \pm 0.03 73.1±0.373.1 \pm 0.3 82.7±2.082.7 \pm 2.0 90.14±0.3890.14 \pm 0.38 80.52±0.0880.52 \pm 0.08
    low low 78.17±0.1378.17 \pm 0.13 76.0±0.276.0 \pm 0.2 – 90.95±0.0290.95 \pm 0.02 90.37±0.0490.37 \pm 0.04
    low mid 78.86±0.2278.86 \pm 0.22 57.4±0.357.4 \pm 0.3 – 84.44±1.0484.44 \pm 1.04 80.45±0.1180.45 \pm 0.11
    low base 73.71±0.4773.71 \pm 0.47 33.1±1.333.1 \pm 1.3 – 84.73±0.5484.73 \pm 0.54 80.59±0.1480.59 \pm 0.14

    (CIFAR-10 resolution dimensions: base =32×32= 32 \times 32, mid =16×16= 16 \times 16, low =8×8= 8 \times 8. Celeb-A resolution dimensions: base =160×160= 160 \times 160, mid =128×128= 128 \times 128, low =64×64= 64 \times 64).

    When trained on 8×88 \times 8 and tested on 32×3232 \times 32 CIFAR-10, S4ND achieves 73.71% accuracy, exceeding Conv2D (33.1%) by 40.61 percentage points. When trained on mid (16×1616 \times 16) and evaluated on base (32×3232 \times 32), S4ND achieves 88.77%, outperforming FlexNet-16 (82.7%) by 6.07 percentage points and Conv2D (73.1%) by 15.67 percentage points.

  9. Knowl 9 — Progressively Resized Multi-Resolution Training Efficiency

    empirical result

    Progressive resizing trains vision models across multi-stage resolution schedules (resetting the learning rate schedule at each stage transition). S4ND adapts to resolution transitions while maintaining accuracy close to models trained exclusively at the high base resolution:

    Dataset Model Epoch Schedule Train Resolution Val Acc @ Base (%) Step-Time Speedup
    CIFAR-10 Conv2D – base 91.90 0%
    CIFAR-10 S4ND – base 93.40 0%
    CIFAR-10 Conv2D 80→2080 \to 20 low →\to base 90.94 51.7%
    CIFAR-10 S4ND 80→2080 \to 20 low →\to base 92.32 21.8%
    Celeb-A Conv2D – base 91.44 0%
    Celeb-A S4ND – base 91.75 0%
    Celeb-A Conv2D 16→416 \to 4 low →\to mid 80.89 76.7%
    Celeb-A S4ND 16→416 \to 4 low →\to mid 88.57 57.3%

    On CIFAR-10, an 80-epoch low (8×88 \times 8) followed by 20-epoch base (32×3232 \times 32) schedule reaches 92.32% validation accuracy (within 1.08% of full base training) with a 21.8% step-time speedup. On Celeb-A, combining progressive training with zero-shot generalization (low →\to mid training evaluated at base resolution) achieves 88.57% accuracy with a 57.3% speedup, beating the Conv2D baseline (80.89%) by 7.68 percentage points.

  10. Knowl 10 — Memory I/O Training Bottlenecks in Multidimensional S4ND

    limitation

    In 2D image architectures, S4ND-ConvNeXt trains approximately 2×2\times slower than the baseline ConvNeXt using local 2D convolutions. This runtime penalty is caused by GPU memory bandwidth constraints: 65% of S4ND execution time is consumed by reading and writing intermediate state tensors to GPU memory during fast Fourier transforms (FFT), pointwise frequency-domain multiplications, and inverse fast Fourier transforms (iFFT). Fusing these operators into a single memory-efficient kernel is projected to provide a 2–3×2\text{--}3\times execution speedup.

Coverage note — Standard 1D continuous state space foundations (such as 1D HiPPO formulas and orthogonal polynomial basis derivations) were omitted as they constitute prior work reviewed in the preliminaries rather than new contributions of this paper.

References

  1. 1.Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems, 34, 2021.
  2. 2.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lu\v{c}i'c, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6836–6846, 2021.
  3. 3.Hangbo Bao, Li Dong, and Furu Wei. Beit: Bert pre-training of image transformers. arXiv preprint arXiv:2106.08254, 2021.
  4. 4.Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021.
  5. 5.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
  6. 6.Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. Electra: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555, 2020.
  7. 7.Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 702–703, 2020.
  8. 8.Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European Conference on Computer Vision (ECCV), pages 720–736, 2018.
  9. 9.Tri Dao, Daniel Y Fu, Stefano Ermon, Atri Rudra, and Christopher R'e. Flashattention: Fast and memory-efficient exact attention with io-awareness. arXiv preprint arXiv:2205.14135, 2022.
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  11. 11.Arjun D Desai, Beliz Gunel, Batu M Ozturkler, Harris Beg, Shreyas Vasanawala, Brian A Hargreaves, Christopher R'e, John M Pauly, and Akshay S Chaudhari. Vortex: Physics-driven data augmentations for consistency training for robust accelerated mri reconstruction. arXiv preprint arXiv:2111.02549, 2021.
  12. 12.Mingyu Ding, Bin Xiao, Noel Codella, Ping Luo, Jingdong Wang, and Lu Yuan. Davit: Dual attention vision transformers. arXiv preprint arXiv:2204.03645, 2022.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  14. 14.fastai. Training imagenet in 3 hours for 25 minutes. https://www.fast.ai/2018/04/30/dawnbench-fastai/, 2018.
  15. 15.Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6202–6211, 2019.
  16. 16.Marc Finzi, Samuel Stanton, Pavel Izmailov, and Andrew Gordon Wilson. Generalizing convolutional neural networks for equivariance to lie groups on arbitrary continuous data. In ICML, 2020.
  17. 17.Karan Goel, Albert Gu, Chris Donahue, and Christopher R'e. It’s raw! audio generation with state-space models. arXiv preprint arXiv:2202.09729, 2022.
  18. 18.Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision, pages 5842–5850, 2017.
  19. 19.Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher R'e. Hippo: Recurrent memory with optimal polynomial projections. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedings.neurips.cc/paper/2020/hash/102f0bb6efb3a6128a3c750dd16729be-Abstract.html.
  20. 20.Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher R'e. Combining recurrent, convolutional, and continuous-time models with the linear state space layer. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  21. 21.Albert Gu, Karan Goel, and Christopher R'e. Efficiently modeling long sequences with structured state spaces. In The International Conference on Learning Representations (ICLR), 2022.
  22. 22.Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher R'e. How to train your hippo: State space models with generalized orthogonal basis projections. arXiv preprint arXiv:2206.12037, 2022.
  23. 23.Ankit Gupta. Diagonal state spaces are as effective as structured state spaces. arXiv preprint arXiv:2203.14343, 2022.
  24. 24.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Learning spatio-temporal features with 3d residual networks for action recognition. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 3154–3160, 2017.
  25. 25.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  26. 26.Dan Hendrycks, Norman Mu, Ekin D Cubuk, Barret Zoph, Justin Gilmer, and Balaji Lakshminarayanan. Augmix: A simple data processing method to improve robustness and uncertainty. arXiv preprint arXiv:1912.02781, 2019.
  27. 27.Sepp Hochreiter and J"urgen Schmidhuber. Long short-term memory. Neural computation, 9(8):1735–1780, 1997.
  28. 28.Elad Hoffer, Berry Weinstein, Itay Hubara, Tal Ben-Nun, Torsten Hoefler, and Daniel Soudry. Mix & match: training convnets with mixed image sizes for improved accuracy, speed and scale resiliency. arXiv preprint arXiv:1908.08986, 2019.
  29. 29.Elad Hoffer, Tal Ben-Nun, Itay Hubara, Niv Giladi, Torsten Hoefler, and Daniel Soudry. Augment your batch: Improving generalization through instance repetition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8129–8138, 2020.
  30. 30.Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. Deep networks with stochastic depth. In European conference on computer vision, pages 646–661. Springer, 2016.
  31. 31.Md Mohaiminul Islam and Gedas Bertasius. Long movie clip classification with state-space video models. arXiv preprint arXiv:2204.01692, 2022.
  32. 32.Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 3d convolutional neural networks for human action recognition. IEEE transactions on pattern analysis and machine intelligence, 35(1):221–231, 2012.
  33. 33.Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. Large-scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 1725–1732, 2014.
  34. 34.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017.
  35. 35.Dan Kondratyuk, Liangzhe Yuan, Yandong Li, Li Zhang, Mingxing Tan, Matthew Brown, and Boqing Gong. Movinets: Mobile video networks for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16020–16030, 2021.
  36. 36.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. Citeseer, 2009.
  37. 37.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM, 60:84 – 90, 2012.
  38. 38.Hildegard Kuehne, Hueihan Jhuang, Est'\i{}baliz Garrote, Tomaso Poggio, and Thomas Serre. Hmdb: a large video database for human motion recognition. In 2011 International conference on computer vision, pages 2556–2563. IEEE, 2011.
  39. 39.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. arXiv preprint arXiv:2111.09883, 2021.
  40. 40.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021.
  41. 41.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021.
  42. 42.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  43. 43.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In ICCV, pages 3730–3738. IEEE Computer Society, 2015. ISBN 978-1-4673-8391-2. URL http://dblp.uni-trier.de/db/conf/iccv/iccv2015.html#LiuLWT15.
  44. 44.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
  45. 45.Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. In European conference on computer vision, pages 405–421. Springer, 2020.
  46. 46.Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfreund, Carl Vondrick, et al. Moments in time dataset: one million videos for event understanding. IEEE transactions on pattern analysis and machine intelligence, 42(2):502–508, 2019.
  47. 47.Bruno A Olshausen and David J Field. Natural image statistics and efficient coding. Network: computation in neural systems, 7(2):333, 1996.
  48. 48.Boris T Polyak and Anatoli B Juditsky. Acceleration of stochastic approximation by averaging. SIAM journal on control and optimization, 30(4):838–855, 1992.
  49. 49.Zhaofan Qiu, Ting Yao, and Tao Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In proceedings of the IEEE International Conference on Computer Vision, pages 5533–5541, 2017.
  50. 50.David W. Romero, Robert-Jan Bruintjes, Jakub M. Tomczak, Erik J. Bekkers, Mark Hoogendoorn, and Jan C. van Gemert. Flexconv: Continuous kernel convolutions with differentiable kernel sizes. ArXiv, abs/2110.08059, 2021.
  51. 51.David W Romero, Anna Kuzina, Erik J Bekkers, Jakub M Tomczak, and Mark Hoogendoorn. Ckconv: Continuous kernel convolution for sequential data. arXiv preprint arXiv:2102.02611, 2021.
  52. 52.Daniel Ruderman and William Bialek. Statistics of natural images: Scaling in the woods. Advances in neural information processing systems, 6, 1993.
  53. 53.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115:211–252, 2015.
  54. 54.Kristof Sch"utt, Pieter-Jan Kindermans, Huziel Enoc Sauceda Felix, Stefan Chmiela, Alexandre Tkatchenko, and Klaus-Robert M"uller. Schnet: A continuous-filter convolutional neural network for modeling quantum interactions. In NeurIPS, 2017.
  55. 55.Gunnar A Sigurdsson, G"ul Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In European Conference on Computer Vision, pages 510–526. Springer, 2016.
  56. 56.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  57. 57.David St\v{r}el'ak and Ji\v{r}'\i{} Filipovi\v{c}. Performance analysis and autotuning setup of the cufft library. In Proceedings of the 2nd Workshop on AutotuniNg and aDaptivity AppRoaches for Energy efficient HPC Systems, pages 1–6, 2018.
  58. 58.Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1–9, 2015.
  59. 59.Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2818–2826, 2016.
  60. 60.Mingxing Tan and Quoc Le. Efficientnetv2: Smaller models and faster training. In International Conference on Machine Learning, pages 10096–10106. PMLR, 2021.
  61. 61.Hugo Touvron, Andrea Vedaldi, Matthijs Douze, and Herv'e J'egou. Fixing the train-test resolution discrepancy. Advances in neural information processing systems, 32, 2019.
  62. 62.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herv'e J'egou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, pages 10347–10357. PMLR, 2021.
  63. 63.Hugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve, and Herv'e J'egou. Going deeper with image transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 32–42, 2021.
  64. 64.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 4489–4497, 2015.
  65. 65.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 6450–6459, 2018.
  66. 66.Ross Wightman. Pytorch image models. https://github.com/rwightman/pytorch-image-models, 2019.
  67. 67.Chao-Yuan Wu and Philipp Krahenbuhl. Towards long-form video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1884–1894, 2021.
  68. 68.Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. Rethinking spatiotemporal feature learning for video understanding. arXiv preprint arXiv:1712.04851, 1(2):5, 2017.
  69. 69.Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. Tokens-to-token vit: Training vision transformers from scratch on imagenet. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 558–567, 2021.
  70. 70.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023–6032, 2019.
  71. 71.Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. CoRR, abs/2106.04560, 2021.
  72. 72.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. arXiv preprint arXiv:1710.09412, 2017.
  73. 73.Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 13001–13008, 2020.

Citation

MLA
Nguyen, E., et al. “S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 2846–61, https://proceedings.neurips.cc/paper_files/paper/2022/file/13388efc819c09564c66ab2dc8463809-Paper-Conference.pdf.
APA
Nguyen, E., Goel, K., Gu, A., Downs, G., Shah, P., Dao, T., Baccus, S., & Ré, C. (2022). S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces. Advances in Neural Information Processing Systems, 35, 2846–2861. https://proceedings.neurips.cc/paper_files/paper/2022/file/13388efc819c09564c66ab2dc8463809-Paper-Conference.pdf
Chicago
Nguyen, E., K. Goel, A. Gu, et al. 2022. “S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces”. Advances in Neural Information Processing Systems 35: 2846–61. https://proceedings.neurips.cc/paper_files/paper/2022/file/13388efc819c09564c66ab2dc8463809-Paper-Conference.pdf.
Harvard
Nguyen, E. et al. (2022) “S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 2846–2861. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/13388efc819c09564c66ab2dc8463809-Paper-Conference.pdf.
Vancouver
1. Nguyen E, Goel K, Gu A, Downs G, Shah P, Dao T, Baccus S, Ré C (2022) S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 2846–2861

BibTeX

@inproceedings{nguyen2022s4nd,
  title = {S4ND: Modeling Images and Videos as Multidimensional Signals with State Spaces},
  author = {Nguyen, Eric and Goel, Karan and Gu, Albert and Downs, Gordon and Shah, Preey and Dao, Tri and Baccus, Stephen and Ré, Christopher},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {2846-2861},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/13388efc819c09564c66ab2dc8463809-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors