keyword
space-time interest points
Space-time interest points are local visual features in video sequences that exhibit significant variations in both space and time, identifying salient locations where dynamic visual events occur. Extending the concept of two-dimensional spatial image interest points into the spatio-temporal domain, these features detect specific coordinates in video data characterized by distinctive spatial structures alongside sudden or complex movements. By capturing local motion patterns and visual structures across varying spatial and temporal scales, space-time interest points provide a compact, sparse representation of video data that is widely used in computer vision for tasks such as human action recognition, event detection, and video classification without requiring prior segmentation or tracking.
3 items

Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks
Zhaofan Qiu, Ting Yao, Tao Mei
Why you should read this
Proposes Pseudo-3D Residual Networks (P3D ResNet), an architecture that factorizes 3D convolutions into separable 2D spatial and 1D temporal filters within residual blocks, drastically cutting computation and memory costs while outperforming standard 3D CNNs on video classification benchmarks.
Convolutional Neural Networks (CNN) have been regarded as a powerful class of models for image recognition problems. Nevertheless, it is not trivial when utilizing a CNN for learning spatio-temporal video representation. A few studies have shown that performing 3D convolutions is a rewarding approach to capture both spatial and temporal dimensions in videos. However, the development of a very deep 3D CNN from scratch results in expensive computational cost and memory demand. A valid question is why not recycle off-the-shelf 2D networks for a 3D CNN. In this paper, we devise multiple variants of bottleneck building blocks in a residual learning framework by simulating convolutions with convolutional filters on spatial domain (equivalent to 2D CNN) plus convolutions to construct temporal connections on adjacent feature maps in time. Furthermore, we propose a new architecture, named Pseudo-3D Residual Net (P3D ResNet), that exploits all the variants of blocks but composes each in different placement of ResNet, following the philosophy that enhancing structural diversity with going deep could improve the power of neural networks. Our P3D ResNet achieves clear improvements on Sports-1M video classification dataset against 3D CNN and frame-based 2D CNN by 5.3% and 1.8%, respectively. We further examine the generalization performance of video representation produced by our pre-trained P3D ResNet on five different benchmarks and three different tasks, demonstrating superior performances over several state-of-the-art techniques.
Added
2026-09-18

Action recognition by dense trajectories
Heng Wang, Alexander Klaser, Cordelia Schmid, Cheng-Lin Liu
Why you should read this
Proposes tracking densely sampled points using optical flow alongside motion boundary histogram descriptors to effectively capture complex video motion and improve action recognition across challenging benchmarks.
Feature trajectories have shown to be efficient for representing videos. Typically, they are extracted using the KLT tracker or matching SIFT descriptors between frames. However, the quality as well as quantity of these trajectories is often not sufficient. Inspired by the recent success of dense sampling in image classification, we propose an approach to describe videos by dense trajectories. We sample dense points from each frame and track them based on displacement information from a dense optical flow field. Given a state-of-the-art optical flow algorithm, our trajectories are robust to fast irregular motions as well as shot boundaries. Additionally, dense trajectories cover the motion information in videos well. We, also, investigate how to design descriptors to encode the trajectory information. We introduce a novel descriptor based on motion boundary histograms, which is robust to camera motion. This descriptor consistently outperforms other state-of-the-art descriptors, in particular in uncontrolled realistic videos. We evaluate our video description in the context of action classification with a bag-of-features approach. Experimental results show a significant improvement over the state of the art on four datasets of varying difficulty, i.e. KTH, YouTube, Hollywood2 and UCF sports.
Added
2026-09-14

On Space-Time Interest Points
Ivan Laptev
Why you should read this
Introduces a spatio-temporal interest point detector and scale-selection mechanism that identifies salient space-time events directly from video, enabling compact representation and recognition of complex human actions without tracking or optical flow.
Local image features or interest points provide compact and abstract representations of patterns in an image. In this paper, we extend the notion of spatial interest points into the spatio-temporal domain and show how the resulting features often reflect interesting events that can be used for a compact representation of video data as well as for interpretation of spatio-temporal events. To detect spatio-temporal events, we build on the idea of the Harris and Förstner interest point operators and detect local structures in space-time where the image values have significant local variations in both space and time. We estimate the spatio-temporal extents of the detected events by maximizing a normalized spatio-temporal Laplacian operator over spatial and temporal scales. To represent the detected events, we then compute local, spatio-temporal, scale-invariant N-jets and classify each event with respect to its jet descriptor. For the problem of human motion analysis, we illustrate how a video representation in terms of local space-time features allows for detection of walking people in scenes with occlusions and dynamic cluttered backgrounds.
Added
2026-09-11
