Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video-language understanding

Video-language understanding is an artificial intelligence domain focused on enabling computational systems to jointly interpret, align, and reason over multimodal content combining dynamic visual video streams and natural language. Unlike static image-text processing, this discipline requires models to analyze both spatial visual features within individual frames and temporal dynamics across sequences of frames, capturing evolving human actions, object trajectories, and causal event progressions over time. Common tasks and capabilities within video-language understanding include text-to-video retrieval, video question answering, automated video captioning and summarization, temporal event grounding, and open-ended multimodal reasoning, allowing intelligent systems to comprehend, index, and interact with complex video narratives based on natural language queries.

3 items

Revealing Single Frame Bias for Video-and-Language Learning

Revealing Single Frame Bias for Video-and-Language Learning

Jie Lei, Tamara L. Berg, Mohit Bansal

OrganizationsDepartment of Computer ScienceUniversity of North Carolina at Chapel Hill

Why you should read this

Reveals a pervasive static appearance bias in standard video-and-language benchmarks by showing that single-frame training paired with inference-time frame ensembling outperforms multi-frame methods, while proposing two new action-focused retrieval tasks to properly evaluate temporal reasoning.

Training an effective video-and-language model intuitively requires multiple frames as model inputs. However, it is unclear whether using multiple frames is beneficial to downstream tasks, and if yes, whether the performance gain is worth the drastically-increased computation and memory costs resulting from using more frames. In this work, we explore single-frame models for video-and-language learning. On a diverse set of video-and-language tasks (including text-to-video retrieval and video question answering), we show the surprising result that, with large-scale pre-training and a proper frame ensemble strategy at inference time, a single-frame trained model that does not consider temporal information can achieve better performance than existing methods that use multiple frames for training. This result reveals the existence of a strong “static appearance bias” in popular video-and-language datasets. Therefore, to allow for a more comprehensive evaluation of video-and-language models, we propose two new retrieval tasks based on existing fine-grained action recognition datasets that encourage temporal modeling. Our code is available at https://github.com/jayleicn/singularity.

Added

2026-10-01

AVA: Towards Agentic Video Analytics with Vision Language Models

AVA: Towards Agentic Video Analytics with Vision Language Models

Yuxuan Yan, Shiqi Jiang, Ting Cao, Yifan Yang, Qianqian Yang, Yuanchao Shu, Yuqing Yang, Lili Qiu

OrganizationsMicrosoftTsinghua UniversityZhejiang University

Why you should read this

Introduces AVA, an agentic video analytics system that builds real-time Event Knowledge Graphs to enable vision language models to accurately reason over ultra-long video streams exceeding ten hours.

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their adaptability in open-ended analytical scenarios. The recent emergence of Vision Language Models (VLMs) as transformative technologies offers significant potential for enabling open-ended video understanding, reasoning, and analytics. Nevertheless, their limited context windows present challenges when processing ultra-long video content, which is prevalent in real-world applications. To address this, we introduce AVA, a VLM-powered system designed for open-ended, advanced video analytics. AVA incorporates two key innovations: (1) the near real-time construction of Event Knowledge Graphs (EKGs) for efficient indexing of long or continuous video streams, and (2) an agentic retrieval-generation mechanism that leverages EKGs to handle complex and diverse queries. Comprehensive evaluations on public benchmarks, LVBench and VideoMME-Long, demonstrate that AVA achieves state-of-the-art performance, attaining 62.3% and 64.1% accuracy, respectively-significantly surpassing existing VLM and video Retrieval-Augmented Generation (RAG) systems. Furthermore, to evaluate video analytics in ultra-long and open-world video scenarios, we introduce a new benchmark, AVA-100. This benchmark comprises 8 videos, each exceeding 10 hours in duration, along with 120 manually annotated, diverse, and complex question-answer pairs. On AVA-100, AVA achieves top-tier performance with an accuracy of 75.8%. The source code of AVA is available at this https URL. The AVA-100 benchmark can be accessed at this https URL.

Added

2026-09-30

HourVideo: 1-Hour Video-Language Understanding

HourVideo: 1-Hour Video-Language Understanding

Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, Li Fei-Fei

OrganizationsStanford University

Why you should read this

Presents a benchmark of nearly thirteen thousand questions across five hundred hour-long egocentric videos to expose the critical performance gap between state-of-the-art multimodal models and human-level long-form video comprehension.

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at hourvideo.stanford.edu.

Added

2026-09-26