Selective Structured State-Spaces for Long-Form Video Understanding

Jue WangWentao ZhuPichao WangXiang YuLinda LiuMohamed OmarRaffay Hamid

article2023CVPR196 citations

Proposes a selective structured state-space architecture and masked contrastive pre-training strategy that adaptively discards uninformative visual tokens, cutting memory consumption by 23% while improving long-form video understanding accuracy by up to 9.6%.

Listen

Analyzing long-form video content—such as movies and multi-step instructional guides lasting several minutes—requires deep neural networks to capture complex relationships across long stretches of time and space. While standard vision models struggle with severe computational bottlenecks and high memory usage over extended sequences, newer Structured State-Space Sequence models (known as S4) offer efficient, linear computation. However, baseline S4 approaches treat every visual token as equally important, allowing irrelevant background frames to dilute performance across different video analysis tasks.

The article demonstrates a novel framework called Selective S4 (S5) that adaptively filters out redundant visual information, paired with a training strategy named Long-Short Masked Contrastive Learning (LSMCL). The main objective is to establish an architecture that simultaneously boosts accuracy on long-form video benchmarks while substantially lowering computational and hardware overhead.

To evaluate this framework, the authors conducted empirical experiments across three established benchmark datasets: the Long-form Video Understanding dataset (comprising approximately 30,000 clips spanning nine distinct tasks), COIN (11,827 instructional videos across 180 tasks), and Breakfast (1,712 procedural cooking videos). The S5 method incorporates a lightweight, linear mask generator that leverages sequence context without performing expensive self-attention computations. In addition, the LSMCL pretraining randomly masks long and short video segments to teach the model to handle missing tokens and predict broader context from shorter inputs.

The experimental findings show clear improvements across all benchmarks. First, the S5 architecture outperforms the previous state-of-the-art S4 baseline by up to 9.6% in classification accuracy on the LVU dataset, while reducing graphics memory footprint by 23% to 25% with no loss in processing throughput. Second, S5 achieved top-tier performance on the COIN and Breakfast datasets, delivering accuracies of 90.81% and 90.70%, respectively. Third, ablation testing revealed that optimal token selection occurs at a 50% masking ratio, successfully halving the processed tokens without degrading model accuracy. Finally, the LSMCL pretraining effectively enabled shorter video clips to achieve performance levels comparable to unmasked models fed 66% more frames.

These results demonstrate that long-form video modeling does not require processing every visual element equally to achieve high accuracy. By dynamically selecting only the most informative tokens and applying robust pretraining, organizations can deploy high-performing video intelligence systems with lower infrastructure costs, lower memory constraints, and faster deployment timelines across diverse tasks.

For practical implementation, practitioners should prioritize lightweight linear token-selection modules over complex transformer selectors and configure masking ratios around 50% to maximize efficiency gains. When planning data ingestion pipelines, teams can reduce frame counts by utilizing LSMCL pretraining to maintain strong predictive performance. Future investigations should test the framework on untrimmed full-length video archives and multi-modal settings incorporating speech and audio before executing large-scale production deployments.

The conclusions are supported by thorough comparative evaluations on standardized datasets. Readers should note that performance gains taper off if token masking exceeds 50% or if video sequences already have very low redundancy, indicating that token reduction parameters must be calibrated based on the underlying video density.

arXiv: 2303.14526
Cover for Selective Structured State-Spaces for Long-Form Video Understanding

Abstract

Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all image-tokens equally as done by S4 model can adversely affect its efficiency and accuracy. To address this limitation, we present a novel Selective S4 (i.e., S5) model that employs a lightweight mask generator to adaptively select informative image tokens resulting in more efficient and accurate modeling of long-term spatiotemporal dependencies in videos. Unlike previous mask-based token reduction methods used in transformers, our S5 model avoids the dense self-attention calculation by making use of the guidance of the momentum-updated S4 model. This enables our model to efficiently discard less informative tokens and adapt to various long-form video understanding tasks more effectively. However, as is the case for most token reduction methods, the informative image tokens could be dropped incorrectly. To improve the robustness and the temporal horizon of our model, we propose a novel long-short masked contrastive learning (LSMCL) approach that enables our model to predict longer temporal context using shorter input videos. We present extensive comparative results using three challenging long-form video understanding datasets (LVU, COIN and Breakfast), demonstrating that our approach consistently outperforms the previous state-of-the-art S4 model by up to 9.6% accuracy while reducing its memory footprint by 23%.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Preliminaries
  • 3.1.1 S4 Model
  • 3.1.2 ViS4mer Model
  • 3.2. S4 Model in Long-form Video Understanding
  • 3.3. Adaptive Token in Long-form Videos
  • 3.4. Long-Short Mask Contrastive Learning
  • 4. Experiments
  • 4.1. Dataset
  • 4.2. Implementation Details
  • 4.3. Ablation Study
  • 4.4. Comparison with the State-Of-The-Arts
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Selective Structured State-Space (S5) Architecture for Long-Form Video

    model/method

    The Selective Structured State-Space (S5) model is an architecture designed for long-form video understanding that adaptively filters uninformative spatiotemporal tokens without relying on quadratic-complexity self-attention.

    Standard vision transformers and Structured State-Space (S4) video models treat all spatiotemporal tokens equally across long temporal horizons. In contrast, S5 introduces an adaptive mask generator module prior to the S4 sequence modeling layers:

    1. Given a sequence of S⋅TS \cdot T spatiotemporal tokens X∈RST×DX \in \mathbb{R}^{ST \times D} from TT frames with SS spatial tokens each and feature dimension DD.
    2. Global sequence-context representations are dynamically extracted by a momentum-updated S4 teacher model (maintained as an exponential moving average of the primary S4 parameters).
    3. The momentum S4 features xs4x_{s_4} are passed to a lightweight mask generator (such as a single linear layer), which predicts categorical probability distributions over the S⋅TS \cdot T tokens.
    4. The top-KK most informative tokens are selected using a differentiable Gumbel-Softmax straight-through sampling mechanism, discarding the remaining (S⋅T−K)(S \cdot T - K) tokens.
    5. The selected KK tokens are then fed into downstream S4 blocks for long-term temporal reasoning.

    By using linear-complexity S4 features rather than all-to-all self-attention to guide token selection, S5 scales efficiently to long video inputs while mitigating the noise from redundant background or task-irrelevant tokens.

  2. Knowl 2 — Long-Short Masked Contrastive Learning (LSMCL) Pre-Training

    model/method

    Long-Short Masked Contrastive Learning (LSMCL) is a self-supervised pre-training framework designed to enhance the robustness and temporal reasoning horizon of state-space video models (such as S5) when operating on partially dropped or mis-selected token sequences.

    In LSMCL:

    1. Multi-Scale Sampling: From each raw video sequence, a long clip xLx_L and a short clip xSx_S are extracted using different frame sampling strides τL\tau_L and τS\tau_S with τS<τL\tau_S < \tau_L. Crucially, the temporal span of xLx_L subsumes that of xSx_S, ensuring shared semantic overlap rather than unrelated clip semantics.
    2. Random Token Masking: Both xLx_L and xSx_S are subjected to independent random spatiotemporal token masking Rmask(x,η)\mathcal{R}_{\text{mask}}(x, \eta) with masking ratio η\eta (typically 50%50\%). This simulates potential token dropouts and mis-predictions that may occur during downstream adaptive token selection.
    3. Contrastive Objective: The masked clips are passed through an online query encoder fqf_q and a momentum-updated key encoder fkf_k (whose parameters θk\theta_k update via θk←mθk+(1−m)θq\theta_k \leftarrow m\theta_k + (1-m)\theta_q with momentum coefficient m∈[0,1]m \in [0, 1]). The network optimizes a symmetrized InfoNCE contrastive loss aligning representations of the short clip to the long clip and vice versa.

    Through this formulation, the state-space backbone learns discretization parameters (such as the step size Δ\Delta) and state-space transitions that predict long-range context from short, partially observed token subsets, allowing models pre-trained with LSMCL on shorter inputs to match the accuracy of models trained on significantly longer unmasked inputs.

  3. Knowl 3 — Differentiable Token Mask Generation Formulation

    equation

    In the Selective Structured State-Space (S5) model, adaptive token selection is formulated as a categorical sampling problem over all S⋅TS \cdot T spatiotemporal video tokens.

    Let X∈RST×DX \in \mathbb{R}^{ST \times D} denote the matrix of S⋅TS \cdot T input tokens of dimension DD. Let xs4x_{s_4} denote the sequence features generated by the momentum-updated S4 model. A mask generator network MG(⋅)MG(\cdot) outputs a normalized categorical distribution p(c∣xs4)∈[0,1]p(c \mid x_{s_4}) \in [0, 1] across token indices c∈{C1,…,CST}c \in \{C_1, \dots, C_{ST}\} such that ∑c=C1CSTp(c∣xs4)=1\sum_{c=C_1}^{C_{ST}} p(c \mid x_{s_4}) = 1.

    To select the top-KK tokens while allowing end-to-end gradient propagation during training, the Gumbel-Softmax with Straight-Through estimator is employed. Independent Gumbel noise g∈R1×STg \in \mathbb{R}^{1 \times ST} is generated as:

    g=−log⁡(−log⁡(u+ϵ)+ϵ),u∼Uniform(0,1)g = -\log(-\log(u + \epsilon) + \epsilon), \quad u \sim \text{Uniform}(0, 1)

    where ϵ>0\epsilon > 0 is a small constant for numerical stability.

    During the forward pass, the top-KK tokens are sampled without replacement from the perturbed distribution p+gp + g, producing one-hot indicator vectors ck∈{0,1}STc^k \in \{0, 1\}^{ST} for k∈{1,…,K}k \in \{1, \dots, K\}, yielding selected tokens:

    xink=X⊤ckx_{\text{in}}^k = X^\top c^k

    During the backward pass, the gradient GG with respect to the mask generator parameters is estimated using the temperature-scaled softmax distribution:

    G≈∇MGexp⁡((log⁡p(c∣xs4)+g(c))/ρ)∑c′=C1CSTexp⁡((log⁡p(c′∣xs4)+g(c′))/ρ)G \approx \nabla_{MG} \frac{\exp\left((\log p(c \mid x_{s_4}) + g(c)) / \rho\right)}{\sum_{c'=C_1}^{C_{ST}} \exp\left((\log p(c' \mid x_{s_4}) + g(c')) / \rho\right)}

    where ρ>0\rho > 0 is a temperature hyperparameter controlling distribution sharpness.

  4. Knowl 4 — Long-Short Masked Contrastive Loss Formulation

    equation

    The Long-Short Masked Contrastive Learning (LSMCL) objective optimizes a symmetrized InfoNCE loss over randomly masked multi-stride video clips.

    Given a short clip xSx_S and a long clip xLx_L sampled from the same video with strides τS<τL\tau_S < \tau_L such that the temporal span of xLx_L encompasses xSx_S, independent random masking Rmask(⋅,η)\mathcal{R}_{\text{mask}}(\cdot, \eta) with masking ratio η∈[0,1]\eta \in [0, 1] is applied to both clips. Let fqf_q denote the online query encoder and fkf_k denote the momentum target key encoder. The representations for sample ii in a batch are:

    qi=fq(Rmask(xSi,η)),ki=fk(Rmask(xLi,η))q^i = f_q(\mathcal{R}_{\text{mask}}(x_S^i, \eta)), \quad k^i = f_k(\mathcal{R}_{\text{mask}}(x_L^i, \eta))

    The contrastive loss for predicting the long-clip representation from the short-clip query across a batch of BB video pairs is:

    LLSMCL=∑i=1B−log⁡exp⁡(qi⊤ki/ρ)exp⁡(qi⊤ki/ρ)+∑j≠iexp⁡(qi⊤kj/ρ)\mathcal{L}_{\text{LSMCL}} = \sum_{i=1}^B -\log \frac{\exp\left({q^i}^\top k^i / \rho\right)}{\exp\left({q^i}^\top k^i / \rho\right) + \sum_{j \neq i} \exp\left({q^i}^\top k^j / \rho\right)}

    where ρ>0\rho > 0 is the contrastive temperature hyperparameter. The total pre-training loss is symmetrized by swapping the clip inputs to fqf_q and fkf_k, evaluating LLSMCL\mathcal{L}_{\text{LSMCL}} with q=fq(Rmask(xL,η))q = f_q(\mathcal{R}_{\text{mask}}(x_L, \eta)) and k=fk(Rmask(xS,η))k = f_k(\mathcal{R}_{\text{mask}}(x_S, \eta)), and averaging the two directional losses.

  5. Knowl 5 — Performance Comparison on the Long-Form Video Understanding (LVU) Benchmark

    data/table

    The Long-form Video Understanding (LVU) benchmark comprises nine diverse tasks grouped into Content Understanding (relationship, speaking style, scene/place), Metadata Prediction (director, genre, writer, movie release year), and User Engagement (YouTube like ratio, YouTube popularity). Accuracy (%) is reported for classification tasks (higher is better ↑\uparrow) and Mean Squared Error (MSE) is reported for regression tasks (lower is better ↓\downarrow). GPU memory is measured in gigabytes (GB).

    Model Relation (↑\uparrow) Speak (↑\uparrow) Scene (↑\uparrow) Director (↑\uparrow) Genre (↑\uparrow) Writer (↑\uparrow) Year (↑\uparrow) Like (↓\downarrow) View (↓\downarrow) GPU Usage (GB) (↓\downarrow)
    Obj. T4mer 54.76 33.17 52.94 47.66 52.74 36.30 37.76 0.30 3.68 N/A
    Performer 50.00 38.80 60.46 58.87 49.45 48.21 41.25 0.31 3.93 5.93
    Orthoformer 50.00 38.30 66.27 55.14 55.79 47.02 43.35 0.29 3.86 5.56
    VideoBERT 52.80 37.90 54.90 47.30 51.90 38.50 36.10 0.32 4.46 N/A
    LST 52.38 37.31 62.79 56.07 52.70 42.26 39.16 0.31 3.83 41.38
    ViS4mer 57.14 40.79 67.44 62.61 54.71 48.80 44.75 0.26 3.63 5.15
    Ours (60 frames) 61.98 41.75 69.88 66.40 58.80 50.60 47.70 0.25 3.51 3.85
    Ours (60 frames + LSMCL) 61.98 41.75 72.53 66.40 61.34 50.60 47.70 0.24 3.51 3.85
    Ours (100 frames) 66.71 41.78 73.28 66.64 63.65 50.60 47.85 0.25 3.51 3.95
    Ours (100 frames + LSMCL) 67.11 42.12 73.49 67.32 65.41 51.27 47.95 0.24 3.51 3.95

    The selective structured state-space model (S5) with 60 frames outperforms the previous state-of-the-art S4 video model (ViS4mer) across all 9 tasks while reducing GPU memory usage from 5.15 GB to 3.85 GB (a 25% reduction). Adding LSMCL pre-training and extending input frames to 100 further improves accuracy across all tasks (e.g., reaching 67.11% on Relationship, 65.41% on Genre, and 73.49% on Scene).

  6. Knowl 6 — Procedural Activity Classification on COIN and Breakfast Datasets

    data/table

    The performance of the S5 model is evaluated on two long-duration procedural action benchmarks: COIN (11,827 videos across 180 procedural action categories, average duration 2.36 minutes) and Breakfast (1,712 videos across 10 complex cooking activities, average duration 2.7 minutes). Top-1 classification accuracy (%) is compared alongside pre-training datasets and sample counts.

    Dataset Method Pre-Training Dataset (Samples) Accuracy (%)
    COIN TSN Kinetics-400 (306K) 73.40
    COIN D-Sprv. HowTo100M (136M) 90.00
    COIN ViS4mer Kinetics-600 (495K) 88.41
    COIN Ours (S5) Kinetics-600 (495K) 90.42
    COIN Ours (S5 + LSMCL) Kinetics-600 (495K) 90.81
    Breakfast VideoGraph Kinetics-400 (306K) 69.50
    Breakfast Timeception Kinetics-400 (306K) 71.30
    Breakfast GHRM Kinetics-400 (306K) 75.50
    Breakfast D-Sprv. HowTo100M (136M) 89.90
    Breakfast ViS4mer Kinetics-600 (495K) 85.10
    Breakfast Ours (S5) Kinetics-600 (495K) 90.14
    Breakfast Ours (S5 + LSMCL) Kinetics-600 (495K) 90.70

    On COIN, S5 achieves 90.42% accuracy without LSMCL and 90.81% with LSMCL, improving by +2.40% over ViS4mer (88.41%) and surpassing D-Sprv. (90.00%) despite D-Sprv. using the 136-million-sample HowTo100M pre-training dataset. On Breakfast, S5 + LSMCL reaches 90.70% accuracy, outperforming ViS4mer (85.10%) by +5.60% and establishing a new state of the art.

  7. Knowl 7 — Mask Generator Design and Memory Footprint Trade-offs

    data/table

    The choice of mask generator architecture in S5 affects both task performance and computational cost. On the LVU benchmark (60 frames per clip, 50% token masking ratio), token selection guided by momentum S4 features (xs4x_{s_4}) is evaluated across linear and transformer-based mask generators against ViT feature inputs and unmasked/random baselines.

    Mask Generator Relation (↑\uparrow) Speak (↑\uparrow) Scene (↑\uparrow) Director (↑\uparrow) Genre (↑\uparrow) Writer (↑\uparrow) Year (↑\uparrow) Like (↓\downarrow) View (↓\downarrow) Avg. Gain vs ViT
    No Mask (ViS4mer) 57.14 40.79 67.44 62.61 54.71 48.80 44.75 0.26 3.63 –
    Random Mask 54.81 38.22 67.44 63.60 54.97 47.00 42.70 0.25 4.00 –
    Single TX (ViT feats) 57.85 40.79 68.66 63.98 55.12 48.85 43.46 0.26 3.82 –
    Single TX (S4 feats) 60.54 41.21 69.83 66.43 57.55 49.47 44.15 0.25 3.51 +3.4%
    Stacked TXs (ViT feats) 59.51 41.21 69.83 64.91 55.12 51.83 47.55 0.25 3.42 –
    Stacked TXs (S4 feats) 61.98 41.75 70.94 67.34 59.16 51.83 47.55 0.24 3.42 +2.5%
    Linear (ViT feats) 54.81 40.28 67.44 63.90 54.97 48.17 42.77 0.26 3.95 –
    Linear (S4 feats) 61.98 41.75 69.88 66.40 58.80 50.60 47.70 0.25 3.51 +6.7%

    Using S4 features from the momentum-updated S4 model consistently improves token selection across all architectures (e.g., a +6.7% average boost for a linear mask generator over raw ViT features). A single linear layer guided by S4 features achieves competitive accuracy with deeper transformer mask generators while maintaining lower GPU memory consumption (3.85 GB vs ~5.5-6.0 GB) and matching ViS4mer inference throughput (~30 samples/s).

  8. Knowl 8 — Impact of Token Masking Ratio and Sequence Length on S5 Performance

    empirical result

    Empirical analysis of the Selective S4 (S5) model on the LVU dataset reveals distinct operational dynamics across token masking ratios and sequence lengths:

    1. Masking Ratio Dynamics: As the adaptive masking ratio increases from 0% (unmasked baseline) up to 50%, downstream task performance steadily improves, confirming that discarding uninformative background and redundant spatiotemporal tokens enhances representation learning. However, beyond a 50% masking ratio, performance degrades sharply because critical task-relevant tokens are forced to be discarded. Thus, a 50% token masking ratio provides an optimal trade-off.
    2. Sequence Length Scaling: When increasing the input sequence length from 60 to 120 frames at 1 fps, standard S4 models (ViS4mer) show plateauing or declining accuracy across several LVU tasks due to the accumulation of redundant spatiotemporal tokens. In contrast, S5 monotonically improves with longer input lengths, showing that adaptive token selection successfully filters redundant tokens and allows state-space models to benefit from extended temporal context.
    3. LSMCL Temporal Expansion: In Long-Short Masked Contrastive Learning (LSMCL), increasing the long-to-short sampling stride ratio τL/τS\tau_L / \tau_S from 1.0 to 2.0 systematically boosts downstream performance. S5 pre-trained with LSMCL on 60 input frames matches the performance gain (~6% average improvement) of an un-pre-trained S5 model evaluated on 100 input frames (a 66% frame increase), effectively broadening the model's temporal horizon.
  9. Knowl 9 — Implementation and Architectural Setup for S5

    experimental setup

    The experimental configuration and model pipeline for the Selective Structured State-Space (S5) framework are structured as follows:

    1. Video Tokenization: Frames are sampled at 1 fps and resized to 224×224224 \times 224 pixels, partitioned into 16×1616 \times 16 non-overlapping patches (S=196S = 196 spatial tokens per frame). Separate learnable spatial positional encodings ese_s and temporal positional encodings ete_t are added to the projected patch embeddings zst∈RDz_s^t \in \mathbb{R}^D to form input token representations xst=zst+es+etx_s^t = z_s^t + e_s + e_t.
    2. Feature Backbones: ViT-L pre-trained on ImageNet-21K is used as the spatial feature extractor for the LVU dataset (with default input length T=60T = 60 frames). Swin-B pre-trained on Kinetics-600 is used as the spatial feature extractor for the COIN and Breakfast datasets (with default input length T=64T = 64 frames).
    3. Model Architecture: The backbone contains three stacked state-space blocks. The first block incorporates the S5 adaptive token selection module (a lightweight linear mask generator operating on momentum-updated S4 features) with a default 50% token selection ratio (K=0.5⋅S⋅TK = 0.5 \cdot S \cdot T). The subsequent two blocks are standard S4 layers with layer normalization, MLP, skip connections, and pooling.
    4. Pre-Training Configuration: In LSMCL pre-training, independent global random masking is applied at a 50% ratio to long and short clips with momentum coefficient m∈[0,1]m \in [0, 1] for the target encoder fkf_k.

Coverage note — Standard background formulations of the continuous state-space model, HiPPO initialization, and bilinear discretization are omitted as they are established prior work.

References

  1. 1.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lućić, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 6836–6846, October 2021.
  2. 2.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  3. 3.Moez Baccouche, Franck Mamalet, Christian Wolf, Christophe Garcia, and Atilla Baskurt. Sequential deep learning for human action recognition. In International workshop on human behavior understanding, pages 29–39. Springer, 2011.
  4. 4.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In Marina Meila and Tong Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 813–824. PMLR, 2021.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020.
  6. 6.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. arXiv preprint arXiv:2006.09882, 2020.
  7. 7.Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. A short note about kinetics-600. arXiv preprint arXiv:1808.01340, 2018.
  8. 8.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  9. 9.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. arXiv preprint arXiv:2011.10566, 2020.
  10. 10.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised visual transformers. arXiv preprint arXiv:2104.02057, 2021.
  11. 11.Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. Rethinking attention with performers. arXiv preprint arXiv:2009.14794, 2020.
  12. 12.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019.
  13. 13.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  14. 14.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  15. 15.Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 6824–6835, October 2021.
  16. 16.Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. Masked autoencoders as spatiotemporal learners. arXiv preprint arXiv:2205.09113, 2022.
  17. 17.Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He. A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3299–3309, 2021.
  18. 18.Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross B. Girshick, and Kaiming He. A large-scale study on unsupervised spatiotemporal representation learning. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021, pages 3299–3309. Computer Vision Foundation / IEEE, 2021.
  19. 19.Christoph Feichtenhofer, Axel Pinz, and Richard P Wildes. Spatiotemporal multiplier networks for video action recognition. In CVPR, 2017.
  20. 20.Christoph Feichtenhofer, Axel Pinz, and Richard P Wildes. Temporal residual networks for dynamic scene recognition. In CVPR, 2017.
  21. 21.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.
  22. 22.Albert Gu, Tri Dao, Stefano Ermon, Atri Rudra, and Christopher Re. Hippo: Recurrent memory with optimal polynomial projections. Advances in Neural Information Processing Systems, 33:1474–1487, 2020.
  23. 23.Albert Gu, Karan Goel, and Christopher Re. Efficiently modeling long sequences with structured state spaces. arXiv preprint arXiv:2111.00396, 2021.
  24. 24.Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Re. Combining recurrent, convolutional, and continuous-time models with linear state space layers. Advances in neural information processing systems, 34:572–585, 2021.
  25. 25.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022.
  26. 26.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020.
  27. 27.Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders. Timeception for complex action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 254–263, 2019.
  28. 28.Noureldien Hussein, Efstratios Gavves, and Arnold WM Smeulders. Videograph: Recognizing minutes-long human activities in videos. arXiv preprint arXiv:1905.05143, 2019.
  29. 29.Md Mohaiminul Islam and Gedas Bertasius. Long movie clip classification with state-space video models. In Proceedings of the European Conference on Computer Vision (ECCV), 2022.
  30. 30.Eric Jang, Shixiang Gu, and Ben Poole. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144, 2016.
  31. 31.Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Francois Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning, pages 5156–5165. PMLR, 2020.
  32. 32.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. The Kinetics-400 dataset is licensed under the Creative Commons Attribution-NonCommercial 4.0 International License.
  33. 33.Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451, 2020.
  34. 34.Bruno Korbar, Du Tran, and Lorenzo Torresani. Scsampler: Sampling salient clips from video for efficient action recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6232–6242, 2019.
  35. 35.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems, 25:1097–1105, 2012.
  36. 36.Hilde Kuehne, Ali Arslan, and Thomas Serre. The language of actions: Recovering the syntax and semantics of goal-directed human activities. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 780–787, 2014.
  37. 37.Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. Not all patches are what you need: Expediting vision transformers via token reorganizations. arXiv preprint arXiv:2202.07800, 2022.
  38. 38.Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, and Lorenzo Torresani. Vx2text: End-to-end learning of video-based text generation from multimodal inputs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7005–7015, 2021.
  39. 39.Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani. Learning to recognize procedural activities with distant supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13853–13863, 2022.
  40. 40.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030, 2021.
  41. 41.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021.
  42. 42.Lingchen Meng, Hengduo Li, Bor-Chun Chen, Shiyi Lan, Zuxuan Wu, Yu-Gang Jiang, and Ser-Nam Lim. Adavit: Adaptive vision transformers for efficient image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12309–12318, 2022.
  43. 43.Yue Meng, Chung-Ching Lin, Rameswar Panda, Prasanna Sattigeri, Leonid Karlinsky, Aude Oliva, Kate Saenko, and Rogerio Feris. Ar-net: Adaptive frame resolution for efficient action recognition. In European Conference on Computer Vision, pages 86–104. Springer, 2020.
  44. 44.Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2630–2640, 2019.
  45. 45.Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. Video transformer network. arXiv preprint arXiv:2102.00719, 2021.
  46. 46.Eric Nguyen, Karan Goel, Albert Gu, Gordon W Downs, Preey Shah, Tri Dao, Stephen A Baccus, and Christopher Re. S4nd: Modeling images and videos as multidimensional signals using state spaces. Advances in neural information processing systems, 2022.
  47. 47.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  48. 48.Zizheng Pan, Bohan Zhuang, Jing Liu, Haoyu He, and Jianfei Cai. Scalable vision transformers with hierarchical pooling. In Proceedings of the IEEE/cvf international conference on computer vision, pages 377–386, 2021.
  49. 49.Mandela Patrick, Dylan Campbell, Yuki Asano, Ishan Misra, Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Joao F Henriques. Keeping your eye on the ball: Trajectory attention in video transformers. Advances in neural information processing systems, 34:12493–12506, 2021.
  50. 50.Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. Dynamicvit: Efficient vision transformers with dynamic token sparsification. Advances in neural information processing systems, 34:13937–13949, 2021.
  51. 51.Adria` Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altche, Michal Valko, et al. Broaden your views for self-supervised video learning. arXiv preprint arXiv:2103.16559, 2021.
  52. 52.Karen Simonyan and Andrew Zisserman. Two-stream convolutional networks for action recognition in videos. arXiv preprint arXiv:1406.2199, 2014.
  53. 53.Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7464–7473, 2019.
  54. 54.Yuchong Sun, Bei Liu, Hongwei Xue, Ruihua Sone, Huan Yang, and Jianlong Fu. Long-form video-language pretraining with multimodal temporal contrastive learning. Advances in neural information processing systems, 2022.
  55. 55.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.
  56. 56.Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou. Coin: A large-scale dataset for comprehensive instructional video analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1207–1216, 2019.
  57. 57.Yansong Tang, Jiwen Lu, and Jie Zhou. Comprehensive instructional video analysis: The coin dataset and performance evaluation. IEEE transactions on pattern analysis and machine intelligence, 43(9):3138–3153, 2020.
  58. 58.Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. arXiv preprint arXiv:2203.12602, 2022.
  59. 59.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 4489–4497, 2015.
  60. 60.Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. Video classification with channel-separated convolutional networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5552–5561, 2019.
  61. 61.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 6450–6459, 2018.
  62. 62.Arnold Tustin. A method of analysing the behaviour of linear systems in terms of time series. Journal of the Institution of Electrical Engineers-Part IIA: Automatic Regulators and Servo Mechanisms, 94(1):130–142, 1947.
  63. 63.Vivek Veeriah, Naifan Zhuang, and Guo-Jun Qi. Differential recurrent neural networks for action recognition. In Proceedings of the IEEE international conference on computer vision, pages 4041–4049, 2015.
  64. 64.Jue Wang, Gedas Bertasius, Du Tran, and Lorenzo Torresani. Long-short temporal contrastive learning of video transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14010–14020, 2022.
  65. 65.Jue Wang and Lorenzo Torresani. Deformable video transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14053–14062, 2022.
  66. 66.Junke Wang, Xitong Yang, Hengduo Li, Zuxuan Wu, and Yu-Gang Jiang. Efficient video transformers with spatialtemporal token selection. arXiv preprint arXiv:2111.11591, 2021.
  67. 67.Chao-Yuan Wu and Philipp Krahenb uhl. Towards Long-Form Video Understanding. In CVPR, 2021.
  68. 68.Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13587–13597, 2022.
  69. 69.Zuxuan Wu, Caiming Xiong, Chih-Yao Ma, Richard Socher, and Larry S Davis. Adaframe: Adaptive frame selection for fast video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1278–1287, 2019.
  70. 70.Hongxu Yin, Arash Vahdat, Jose Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. Adavit: Adaptive tokens for efficient vision transformer. arXiv preprint arXiv:2112.07658, 2021.
  71. 71.Hongxu Yin, Arash Vahdat, Jose M Alvarez, Arun Mallya, Jan Kautz, and Pavlo Molchanov. A-vit: Adaptive tokens for efficient vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10809–10818, 2022.
  72. 72.Pengfei Zhang, Cuiling Lan, Junliang Xing, Wenjun Zeng, Jianru Xue, and Nanning Zheng. View adaptive recurrent neural networks for high performance human action recognition from skeleton data. In Proceedings of the IEEE international conference on computer vision, pages 2117–2126, 2017.
  73. 73.Jiaming Zhou, Kun-Yu Lin, Haoxin Li, and Wei-Shi Zheng. Graph-based high-order relation modeling for long-term action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8984–8993, 2021.

Citation

MLA
Wang, J., et al. “Selective Structured State-Spaces for Long-Form Video Understanding”. arXiv, 2023, http://arxiv.org/abs/2303.14526v1.
APA
Wang, J., Zhu, W., Wang, P., Yu, X., Liu, L., Omar, M., & Hamid, R. (2023). Selective Structured State-Spaces for Long-Form Video Understanding. arXiv. http://arxiv.org/abs/2303.14526v1
Chicago
Wang, J., W. Zhu, P. Wang, et al. 2023. “Selective Structured State-Spaces for Long-Form Video Understanding”. arXiv. http://arxiv.org/abs/2303.14526v1.
Harvard
Wang, J. et al. (2023) “Selective Structured State-Spaces for Long-Form Video Understanding”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.14526v1.
Vancouver
1. Wang J, Zhu W, Wang P, Yu X, Liu L, Omar M, Hamid R (2023) Selective Structured State-Spaces for Long-Form Video Understanding. arXiv

BibTeX

@article{wang2023selective,
  title = {Selective Structured State-Spaces for Long-Form Video Understanding},
  author = {Wang, Jue and Zhu, Wentao and Wang, Pichao and Yu, Xiang and Liu, Linda and Omar, Mohamed and Hamid, Raffay},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.14526v1},
  eprint = {2303.14526}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE