keyword
spatiotemporal representation learning
Spatiotemporal representation learning is a machine learning process focused on automatically discovering and encoding both spatial patterns and temporal dynamics from structured data into unified numerical features. In applications such as video analysis, physical simulations, and sensor monitoring, spatial dimensions represent physical layout, visual appearance, or localized structural information, while temporal dimensions capture motion, evolution, and chronological dependencies across time. By jointly modeling how visual or spatial configurations change across sequential observations, neural networks and related algorithms can capture complex interactions that static representations miss. These learned multidimensional representations serve as foundational models that facilitate downstream tasks including action recognition, trajectory forecasting, event detection, and long-form video comprehension.
2 items

Selective Structured State-Spaces for Long-Form Video Understanding
Jue Wang, Wentao Zhu, Pichao Wang, Xiang Yu, Linda Liu, Mohamed Omar, Raffay Hamid
Why you should read this
Proposes a selective structured state-space architecture and masked contrastive pre-training strategy that adaptively discards uninformative visual tokens, cutting memory consumption by 23% while improving long-form video understanding accuracy by up to 9.6%.
Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all image-tokens equally as done by S4 model can adversely affect its efficiency and accuracy. To address this limitation, we present a novel Selective S4 (i.e., S5) model that employs a lightweight mask generator to adaptively select informative image tokens resulting in more efficient and accurate modeling of long-term spatiotemporal dependencies in videos. Unlike previous mask-based token reduction methods used in transformers, our S5 model avoids the dense self-attention calculation by making use of the guidance of the momentum-updated S4 model. This enables our model to efficiently discard less informative tokens and adapt to various long-form video understanding tasks more effectively. However, as is the case for most token reduction methods, the informative image tokens could be dropped incorrectly. To improve the robustness and the temporal horizon of our model, we propose a novel long-short masked contrastive learning (LSMCL) approach that enables our model to predict longer temporal context using shorter input videos. We present extensive comparative results using three challenging long-form video understanding datasets (LVU, COIN and Breakfast), demonstrating that our approach consistently outperforms the previous state-of-the-art S4 model by up to 9.6% accuracy while reducing its memory footprint by 23%.
Added
2026-09-26

SVFormer: Semi-supervised Video Transformer for Action Recognition
Zhen Xing, Qi Dai, Han Hu, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang
Why you should read this
Presents SVFormer, a semi-supervised video transformer framework combining an exponential moving average teacher with specialized tube-level token mixing and temporal warping augmentations to substantially improve action recognition performance under extremely limited supervision.
Semi-supervised action recognition is a challenging but critical task due to the high cost of video annotations. Existing approaches mainly use convolutional neural networks, yet current revolutionary vision transformer models have been less explored. In this paper, we investigate the use of transformer models under the SSL setting for action recognition. To this end, we introduce SVFormer, which adopts a steady pseudo-labeling framework (i.e., EMA-Teacher) to cope with unlabeled video samples. While a wide range of data augmentations have been shown effective for semi-supervised image classification, they generally produce limited results for video recognition. We therefore introduce a novel augmentation strategy, Tube TokenMix, tailored for video data where video clips are mixed via a mask with consistent masked tokens over the temporal axis. In addition, we propose a temporal warping augmentation to cover the complex temporal variation in videos, which stretches selected frames to various temporal durations in the clip. Extensive experiments on three datasets Kinetics-400, UCF-101, and HMDB-51 verify the advantage of SVFormer. In particular, SVFormer outperforms the state-of-the-art by 31.5% with fewer training epochs under the 1% labeling rate of Kinetics-400. Our method can hopefully serve as a strong benchmark and encourage future search on semi-supervised action recognition with Transformer networks. Code is released at https://github.com/ChenHsing/SVFormer.
Added
2026-09-26
