Built independently by an author, for readers. Read the story and support ChapterPal

keyword

efficient video understanding

Efficient video understanding is a subfield of computer vision and artificial intelligence focused on analyzing, recognizing, and reasoning about video content while minimizing computational costs, memory usage, latency, and model parameter sizes. Because video data inherently combines dense spatial and temporal dimensions that make standard deep neural networks resource-intensive, efficient video understanding employs strategies such as temporal redundancy reduction, adaptive frame selection, lightweight spatio-temporal operations, multiscale feature representations, and parameter-efficient adaptation. These techniques enable models to achieve high accuracy on tasks such as action recognition, temporal localization, and video question answering while operating effectively in real-time or on resource-constrained hardware without the heavy overhead of dense, exhaustive frame-by-frame processing.

4 items

M-LLM Based Video Frame Selection for Efficient Video Understanding

M-LLM Based Video Frame Selection for Efficient Video Understanding

Kai Hu, Feng Gao, Xiaohan Nie, Peng Zhou, Son Tran, Tal Neiman, Lingyun Wang, Mubarak Shah, Raffay Hamid, Bing Yin, Trishul Chilimbi

OrganizationsAmazonCarnegie Mellon UniversityUniversity of Central Florida

Why you should read this

Develops a lightweight, plug-and-play video frame selector trained with spatial and temporal pseudo-labels to replace uniform sampling, boosting question-answering accuracy and efficiency for frozen multimodal LLMs on long- and medium-context video benchmarks.

Recent advances in Multi-Modal Large Language Models (M-LLMs) show promising results in video reasoning. Popular Multi-Modal Large Language Model (M-LLM) frameworks usually apply naive uniform sampling to reduce the number of video frames that are fed into an M-LLM, particularly for long context videos. However, it could lose crucial context in certain periods of a video, so that the downstream M-LLM may not have sufficient visual information to answer a question. To attack this pain point, we propose a light-weight M-LLM-based frame selection method that adaptively select frames that are more relevant to users' queries. In order to train the proposed frame selector, we introduce two supervision signals (i) Spatial signal, where single frame importance score by prompting an M-LLM; (ii) Temporal signal, in which multiple frames selection by prompting Large Language Model (LLM) using the captions of all frame candidates. The selected frames are then digested by a frozen downstream video M-LLM for visual reasoning and question answering. Empirical results show that the proposed M-LLM video frame selector improves the performances various downstream video Large Language Model (video-LLM) across medium (ActivityNet, NExT-QA) and long (EgoSchema, LongVideoBench) context video question answering benchmarks.

Added

2026-09-26

ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Junting Pan, Ziyi Lin, Xiatian Zhu, Jing Shao, Hongsheng Li

OrganizationsCentre for Perceptual and Interactive Intelligence (CPII)The Chinese University of Hong KongUniversity of Surrey

Why you should read this

Proposes a lightweight Spatio-Temporal Adapter that enables frozen pre-trained image vision transformers to perform video action recognition by updating only about eight percent of parameters while matching or exceeding full fine-tuning performance.

Capitalizing on large pre-trained models for various downstream tasks of interest have recently emerged with promising performance. Due to the ever-growing model size, the standard full fine-tuning based task adaptation strategy becomes prohibitively costly in terms of model training and storage. This has led to a new research direction in parameter-efficient transfer learning. However, existing attempts typically focus on downstream tasks from the same modality (e.g., image understanding) of the pre-trained model. This creates a limit because in some specific modalities, (e.g., video understanding) such a strong pre-trained model with sufficient knowledge is less or not available. In this work, we investigate such a novel cross-modality transfer learning setting, namely parameter-efficient image-to-video transfer learning. To solve this problem, we propose a new Spatio-Temporal Adapter (ST-Adapter) for parameter-efficient fine-tuning per video task. With a built-in spatio-temporal reasoning capability in a compact design, ST-Adapter enables a pre-trained image model without temporal knowledge to reason about dynamic video content at a small (~8%) per-task parameter cost, requiring approximately 20 times fewer updated parameters compared to previous work. Extensive experiments on video action recognition tasks show that our ST-Adapter can match or even outperform the strong full fine-tuning strategy and state-of-the-art video models, whilst enjoying the advantage of parameter efficiency. Code and model are available at https://github.com/linziyi96/st-adapter

Added

2026-09-26

Multiscale Vision Transformers

Multiscale Vision Transformers

Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, Christoph Feichtenhofer

OrganizationsMetaUniversity of California Berkeley

Why you should read this

Develops Multiscale Vision Transformers, a hierarchical architecture that incorporates multiscale feature pyramids into visual attention to achieve superior video and image recognition performance with up to ten times less computation and without requiring massive external pre-training.

We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multiscale Transformers have several channel-resolution scale stages. Starting from the input resolution and a small channel dimension, the stages hierarchically expand the channel capacity while reducing the spatial resolution. This creates a multiscale pyramid of features with early layers operating at high spatial resolution to model simple low-level visual information, and deeper layers at spatially coarse, but complex, high-dimensional features. We evaluate this fundamental architectural prior for modeling the dense nature of visual signals for a variety of video recognition tasks where it outperforms concurrent vision transformers that rely on large scale external pre-training and are 5-10x more costly in computation and parameters. We further remove the temporal dimension and apply our model for image classification where it outperforms prior work on vision transformers. Code is available at: this https URL

Added

2026-09-24

TSM: Temporal Shift Module for Efficient Video Understanding

TSM: Temporal Shift Module for Efficient Video Understanding

Ji Lin, Chuang Gan, Song Han

OrganizationsMassachusetts Institute of TechnologyMIT-IBM Watson AI Lab

Why you should read this

Introduces the Temporal Shift Module, an approach that shifts feature channels across neighboring frames to equip standard 2D CNNs with 3D-level temporal modeling at zero extra parameters and computational cost, making real-time video recognition feasible on edge devices.

The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture temporal relationships; 3D CNN based methods can achieve good performance but are computationally intensive, making it expensive to deploy. In this paper, we propose a generic and effective Temporal Shift Module (TSM) that enjoys both high efficiency and high performance. Specifically, it can achieve the performance of 3D CNN but maintain 2D CNN's complexity. TSM shifts part of the channels along the temporal dimension; thus facilitate information exchanged among neighboring frames. It can be inserted into 2D CNNs to achieve temporal modeling at zero computation and zero parameters. We also extended TSM to online setting, which enables real-time low-latency online video recognition and video object detection. TSM is accurate and efficient: it ranks the first place on the Something-Something leaderboard upon publication; on Jetson Nano and Galaxy Note8, it achieves a low latency of 13ms and 35ms for online video recognition. The code is available at: this https URL.

Added

2026-09-16