Built independently by an author, for readers. Read the story and support ChapterPal

keyword

temporal receptive field

A temporal receptive field is the span of time or sequence of consecutive frames across an input signal that directly influences the activation of a specific unit in a neural network. Analogous to a spatial receptive field in image processing, which measures the two-dimensional region influencing a feature, the temporal receptive field determines how much temporal context or motion history a model can perceive at any given layer. In architectures designed for sequential or video data, such as temporal convolutional networks or three-dimensional convolutional neural networks, operations like temporal convolutions, pooling, and channel shifting progressively expand this span across deeper layers. A larger temporal receptive field enables the network to capture complex, long-range temporal dynamics and dependencies across extended durations, whereas a smaller receptive field focuses on fine-grained, short-term motion cues.

2 items

TSM: Temporal Shift Module for Efficient Video Understanding

TSM: Temporal Shift Module for Efficient Video Understanding

Ji Lin, Chuang Gan, Song Han

OrganizationsMassachusetts Institute of TechnologyMIT-IBM Watson AI Lab

Why you should read this

Introduces the Temporal Shift Module, an approach that shifts feature channels across neighboring frames to equip standard 2D CNNs with 3D-level temporal modeling at zero extra parameters and computational cost, making real-time video recognition feasible on edge devices.

The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture temporal relationships; 3D CNN based methods can achieve good performance but are computationally intensive, making it expensive to deploy. In this paper, we propose a generic and effective Temporal Shift Module (TSM) that enjoys both high efficiency and high performance. Specifically, it can achieve the performance of 3D CNN but maintain 2D CNN's complexity. TSM shifts part of the channels along the temporal dimension; thus facilitate information exchanged among neighboring frames. It can be inserted into 2D CNNs to achieve temporal modeling at zero computation and zero parameters. We also extended TSM to online setting, which enables real-time low-latency online video recognition and video object detection. TSM is accurate and efficient: it ranks the first place on the Something-Something leaderboard upon publication; on Jetson Nano and Galaxy Note8, it achieves a low latency of 13ms and 35ms for online video recognition. The code is available at: this https URL.

Added

2026-09-16

Convolutional Two-Stream Network Fusion for Video Action Recognition

Convolutional Two-Stream Network Fusion for Video Action Recognition

Christoph Feichtenhofer, Axel Pinz, Andrew Zisserman

OrganizationsGraz University of TechnologyUniversity of Oxford

Why you should read this

Proposes a two-stream ConvNet fusion architecture that combines appearance and motion features at convolutional layers rather than the softmax classifier, substantially reducing parameter counts while achieving state-of-the-art accuracy in video action recognition.

Recent applications of Convolutional Neural Networks (ConvNets) for human action recognition in videos have proposed different solutions for incorporating the appearance and motion information. We study a number of ways of fusing ConvNet towers both spatially and temporally in order to best take advantage of this spatio-temporal information. We make the following findings: (i) that rather than fusing at the softmax layer, a spatial and temporal network can be fused at a convolution layer without loss of performance, but with a substantial saving in parameters; (ii) that it is better to fuse such networks spatially at the last convolutional layer than earlier, and that additionally fusing at the class prediction layer can boost accuracy; finally (iii) that pooling of abstract convolutional features over spatiotemporal neighbourhoods further boosts performance. Based on these studies we propose a new ConvNet architecture for spatiotemporal fusion of video snippets, and evaluate its performance on standard benchmarks where this architecture achieves state-of-the-art results.

Added

2026-09-14