Built independently by an author, for readers. Read the story and support ChapterPal

keyword

HourVideo

HourVideo is a multimodal benchmark dataset designed to evaluate the video-language understanding capabilities of artificial intelligence models on long-duration video content. Based on extended egocentric video footage lasting up to two hours, the benchmark assesses how effectively multimodal systems can process, retain, and analyze visual information across broad temporal contexts. It incorporates thousands of multiple-choice questions spanning diverse evaluation tasks, including video summarization, visual tracking and perception, spatial and causal reasoning, and environmental navigation. By probing long-form visual comprehension and temporal reasoning, HourVideo establishes a standardized framework for measuring the performance of long-context multimodal models in comparison to human-level comprehension.

1 item

HourVideo: 1-Hour Video-Language Understanding

HourVideo: 1-Hour Video-Language Understanding

Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, Li Fei-Fei

OrganizationsStanford University

Why you should read this

Presents a benchmark of nearly thirteen thousand questions across five hundred hour-long egocentric videos to expose the critical performance gap between state-of-the-art multimodal models and human-level long-form video comprehension.

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at hourvideo.stanford.edu.

Added

2026-09-26