Built independently by an author, for readers. Read the story and support ChapterPal

keyword

long-context multimodal model

A long-context multimodal model is an artificial intelligence system designed to process, analyze, and reason over vast amounts of diverse data formats, such as text, audio, images, and extended video, within a significantly expanded context window. Unlike conventional multimodal architectures that are constrained to brief media clips or limited token lengths, these models scale context capacities to handle hundreds of thousands or millions of tokens simultaneously. This capability allows the system to sustain long-range temporal understanding, retrieve specific details across hour-long video feeds or extensive document libraries, and perform complex cross-modal reasoning over unified, large-scale inputs without relying on external segmentation.

1 item

HourVideo: 1-Hour Video-Language Understanding

HourVideo: 1-Hour Video-Language Understanding

Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, Li Fei-Fei

OrganizationsStanford University

Why you should read this

Presents a benchmark of nearly thirteen thousand questions across five hundred hour-long egocentric videos to expose the critical performance gap between state-of-the-art multimodal models and human-level long-form video comprehension.

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0% vs. 37.3%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at hourvideo.stanford.edu.

Added

2026-09-26