Built independently by an author, for readers. Read the story and support ChapterPal

keyword

audio-visual language model

An audio-visual language model is a multimodal artificial intelligence system capable of jointly processing, integrating, and reasoning over auditory signals, visual content, and natural language text. Built typically by connecting specialized audio and visual neural encoders to a large language model backbone, these systems align spatial, temporal, and acoustic representations into a shared multimodal space. This integrated architecture enables the model to understand dynamic video streams and spoken or ambient sounds, allowing it to perform complex cross-modal tasks such as video question answering, audio-visual event localization, multimodal summarization, and conversational reasoning across multimedia inputs.

1 item

HourVideo: 1-Hour Video-Language Understanding

HourVideo: 1-Hour Video-Language Understanding

Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, Li Fei-Fei

OrganizationsStanford University

Why you should read this

Presents a benchmark of nearly thirteen thousand questions across five hundred hour-long egocentric videos to expose the critical performance gap between state-of-the-art multimodal models and human-level long-form video comprehension.

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at hourvideo.stanford.edu.

Added

2026-09-26