keyword
video referring
Video referring is a computer vision and multimodal artificial intelligence task that involves identifying, localizing, and tracking specific target entities—such as objects, actions, or events—within a video sequence based on a natural language description or visual prompt. Unlike static image referring, video referring requires models to understand and integrate spatial details with temporal dynamics across continuous frames, maintaining target consistency despite camera movement, object occlusion, perspective shifts, and temporal boundaries. Broadly encompassing subtasks such as video referring expression comprehension, referring video object segmentation, and spatio-temporal grounding, it enables systems to output bounding boxes, pixel-level masks, or trajectories that align a specific linguistic query with its visual manifestation over time.
1 item

