Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video concept spotting

Video concept spotting is a computer vision and video understanding technique that identifies and localizes specific semantic concepts, actions, or objects across the timeline of a video sequence. By evaluating the alignment between visual frames and textual concept representations, this process captures temporal saliency to determine which moments within a video are most relevant to target semantic categories. Instead of treating every frame uniformly, video concept spotting isolates key segments and suppresses background or redundant information, thereby producing more discriminative video representations for downstream tasks such as action recognition, video retrieval, and video question answering.

1 item

Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models

Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models

Wenhao Wu, Xiaohan Wang, Haipeng Luo, Jingdong Wang, Yi Yang, Wanli Ouyang

OrganizationsBaiduShanghai Artificial Intelligence LaboratoryUniversity of Chinese Academy of SciencesUniversity of SydneyZhejiang University

Why you should read this

Proposes a bidirectional framework called BIKE that transfers pre-trained vision-language knowledge into video recognition by retrieving complementary textual attributes and using category concepts to capture frame-level temporal saliency.

Vision-language models (VLMs) pre-trained on large-scale image-text pairs have demonstrated impressive transferability on various visual tasks. Transferring knowledge from such powerful VLMs is a promising direction for building effective video recognition models. However, current exploration in this field is still limited. We believe that the greatest value of pre-trained VLMs lies in building a bridge between visual and textual domains. In this paper, we propose a novel framework called BIKE, which utilizes the cross-modal bridge to explore bidirectional knowledge: i) We introduce the Video Attribute Association mechanism, which leverages the Video-to-Text knowledge to generate textual auxiliary attributes for complementing video recognition. ii) We also present a Temporal Concept Spotting mechanism that uses the Text-to-Video expertise to capture temporal saliency in a parameter-free manner, leading to enhanced video representation. Extensive studies on six popular video datasets, including Kinetics-400 & 600, UCF-101, HMDB-51, ActivityNet and Charades, show that our method achieves state-of-the-art performance in various recognition scenarios, such as general, zero-shot, and few-shot video recognition. Our best model achieves a state-of-the-art accuracy of 88.6% on the challenging Kinetics-400 using the released CLIP model. The code is available at https://github.com/whwu95/BIKE.

Added

2026-09-26