Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Temporal Segment Networks

Temporal Segment Networks are a deep learning framework designed for video action recognition that models long-range temporal structure across entire video sequences. The framework operates by dividing an input video into a fixed number of segments of equal duration and sparsely sampling short snippets or frames from each segment. These sampled inputs are typically processed using two-stream convolutional neural networks that extract spatial features from appearance images and temporal motion features from optical flow fields. The segment-level predictions are then combined using a temporal consensus aggregation function, allowing the entire model to be trained efficiently under video-level supervision while capturing actions distributed across the whole video without processing every intermediate frame.

5 items

Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks

Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks

Zhaofan Qiu, Ting Yao, Tao Mei

OrganizationsMicrosoftUniversity of Science and Technology of China

Why you should read this

Proposes Pseudo-3D Residual Networks (P3D ResNet), an architecture that factorizes 3D convolutions into separable 2D spatial and 1D temporal filters within residual blocks, drastically cutting computation and memory costs while outperforming standard 3D CNNs on video classification benchmarks.

Convolutional Neural Networks (CNN) have been regarded as a powerful class of models for image recognition problems. Nevertheless, it is not trivial when utilizing a CNN for learning spatio-temporal video representation. A few studies have shown that performing 3D convolutions is a rewarding approach to capture both spatial and temporal dimensions in videos. However, the development of a very deep 3D CNN from scratch results in expensive computational cost and memory demand. A valid question is why not recycle off-the-shelf 2D networks for a 3D CNN. In this paper, we devise multiple variants of bottleneck building blocks in a residual learning framework by simulating 3×3×33\times3\times3 convolutions with 1×3×31\times3\times3 convolutional filters on spatial domain (equivalent to 2D CNN) plus 3×1×13\times1\times1 convolutions to construct temporal connections on adjacent feature maps in time. Furthermore, we propose a new architecture, named Pseudo-3D Residual Net (P3D ResNet), that exploits all the variants of blocks but composes each in different placement of ResNet, following the philosophy that enhancing structural diversity with going deep could improve the power of neural networks. Our P3D ResNet achieves clear improvements on Sports-1M video classification dataset against 3D CNN and frame-based 2D CNN by 5.3% and 1.8%, respectively. We further examine the generalization performance of video representation produced by our pre-trained P3D ResNet on five different benchmarks and three different tasks, demonstrating superior performances over several state-of-the-art techniques.

Added

2026-09-18

TSM: Temporal Shift Module for Efficient Video Understanding

TSM: Temporal Shift Module for Efficient Video Understanding

Ji Lin, Chuang Gan, Song Han

OrganizationsMassachusetts Institute of TechnologyMIT-IBM Watson AI Lab

Why you should read this

Introduces the Temporal Shift Module, an approach that shifts feature channels across neighboring frames to equip standard 2D CNNs with 3D-level temporal modeling at zero extra parameters and computational cost, making real-time video recognition feasible on edge devices.

The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture temporal relationships; 3D CNN based methods can achieve good performance but are computationally intensive, making it expensive to deploy. In this paper, we propose a generic and effective Temporal Shift Module (TSM) that enjoys both high efficiency and high performance. Specifically, it can achieve the performance of 3D CNN but maintain 2D CNN's complexity. TSM shifts part of the channels along the temporal dimension; thus facilitate information exchanged among neighboring frames. It can be inserted into 2D CNNs to achieve temporal modeling at zero computation and zero parameters. We also extended TSM to online setting, which enables real-time low-latency online video recognition and video object detection. TSM is accurate and efficient: it ranks the first place on the Something-Something leaderboard upon publication; on Jetson Nano and Galaxy Note8, it achieves a low latency of 13ms and 35ms for online video recognition. The code is available at: this https URL.

Added

2026-09-16

Temporal Segment Networks: Towards Good Practices for Deep Action Recognition

Temporal Segment Networks: Towards Good Practices for Deep Action Recognition

Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, Luc Van Gool

OrganizationsETH ZurichShenzhen Institute of Advanced Technology, Chinese Academy of SciencesThe Chinese University of Hong Kong

Why you should read this

Proposes Temporal Segment Networks, a framework combining sparse temporal sampling and video-level supervision to efficiently model long-range temporal structure and establish practical guidelines for deep action recognition in videos.

Deep convolutional networks have achieved great success for visual recognition in still images. However, for action recognition in videos, the advantage over traditional methods is not so evident. This paper aims to discover the principles to design effective ConvNet architectures for action recognition in videos and learn these models given limited training samples. Our first contribution is temporal segment network (TSN), a novel framework for video-based action recognition. which is based on the idea of long-range temporal structure modeling. It combines a sparse temporal sampling strategy and video-level supervision to enable efficient and effective learning using the whole action video. The other contribution is our study on a series of good practices in learning ConvNets on video data with the help of temporal segment network. Our approach obtains the state-the-of-art performance on the datasets of HMDB51 ( 69.4%69.4\%) and UCF101 (94.2%94.2\%). We also visualize the learned ConvNet models, which qualitatively demonstrates the effectiveness of temporal segment network and the proposed good practices.

Added

2026-09-10

Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset

João Carreira, Andrew Zisserman

OrganizationsGoogleUniversity of Oxford

Why you should read this

Introduces the large-scale Kinetics dataset and Inflated 3D ConvNet (I3D) architecture, demonstrating how inflating successful 2D image models into 3D video networks sets a new benchmark in human action recognition.

The paucity of videos in current action classification datasets (UCF-101 and HMDB-51) has made it difficult to identify good video architectures, as most methods obtain similar performance on existing small-scale benchmarks. This paper re-evaluates state-of-the-art architectures in light of the new Kinetics Human Action Video dataset. Kinetics has two orders of magnitude more data, with 400 human action classes and over 400 clips per class, and is collected from realistic, challenging YouTube videos. We provide an analysis on how current architectures fare on the task of action classification on this dataset and how much performance improves on the smaller benchmark datasets after pre-training on Kinetics. We also introduce a new Two-Stream Inflated 3D ConvNet (I3D) that is based on 2D ConvNet inflation: filters and pooling kernels of very deep image classification ConvNets are expanded into 3D, making it possible to learn seamless spatio-temporal feature extractors from video while leveraging successful ImageNet architecture designs and even their parameters. We show that, after pre-training on Kinetics, I3D models considerably improve upon the state-of-the-art in action classification, reaching 80.9% on HMDB-51 and 98.0% on UCF-101.

Added

2026-09-07