Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video referring

Video referring is a computer vision and multimodal artificial intelligence task that involves identifying, localizing, and tracking specific target entities—such as objects, actions, or events—within a video sequence based on a natural language description or visual prompt. Unlike static image referring, video referring requires models to understand and integrate spatial details with temporal dynamics across continuous frames, maintaining target consistency despite camera movement, object occlusion, perspective shifts, and temporal boundaries. Broadly encompassing subtasks such as video referring expression comprehension, referring video object segmentation, and spatio-temporal grounding, it enables systems to output bounding boxes, pixel-level masks, or trajectories that align a specific linguistic query with its visual manifestation over time.

1 item

Adaptive Keyframe Sampling for Long Video Understanding

Adaptive Keyframe Sampling for Long Video Understanding

Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, Qixiang Ye

OrganizationsUniversity at BuffaloUniversity of Chinese Academy of Sciences

Why you should read this

Proposes a plug-and-play keyframe selection algorithm that balances prompt relevance with temporal coverage to improve long video question-answering accuracy in multimodal large language models without exceeding token limits.

Multimodal large language models (MLLMs) have enabled open-world visual understanding by injecting visual input as extra tokens into large language models (LLMs) as contexts. However, when the visual input changes from a single image to a long video, the above paradigm encounters difficulty because the vast amount of video tokens has significantly exceeded the maximal capacity of MLLMs. Therefore, existing video-based MLLMs are mostly established upon sampling a small portion of tokens from input data, which can cause key information to be lost and thus produce incorrect answers. This paper presents a simple yet effective algorithm named Adaptive Keyframe Sampling (AKS). It inserts a plug-and-play module known as keyframe selection, which aims to maximize the useful information with a fixed number of video tokens. We formulate keyframe selection as an optimization involving (1) the relevance between the keyframes and the prompt, and (2) the coverage of the keyframes over the video, and present an adaptive algorithm to approximate the best solution. Experiments on two long video understanding benchmarks validate that AKS improves video QA accuracy (beyond strong baselines) upon selecting informative keyframes. Our study reveals the importance of information pre-filtering in video-based MLLMs. Our codes are available at https://github.com/ncTimTang/AKS

Added

2026-09-26