LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Haoning WuDongxu LiBei ChenJunnan Li
Establishes LongVideoBench, a challenging benchmark of hour-long interleaved video-subtitle inputs and a novel referring reasoning task that exposes substantial performance gaps in leading large multimodal models.
Artificial intelligence models are rapidly expanding their context windows to ingest large volumes of information, yet standard benchmarks primarily evaluate text-only inputs. Existing video benchmarks suffer from single-frame bias, where models achieve top scores by processing just a handful of frames or global summaries rather than analyzing extensive visual content over time. Consequently, practitioners lack reliable methods to measure whether vision-language models can genuinely process, retrieve, and reason across extended multimodal inputs such as hour-long subtitled videos.
The article introduces and evaluates LongVideoBench, a benchmark designed to assess the long-context multimodal understanding of large multimodal models. Its primary objective is to evaluate how effectively proprietary and open-source models retrieve fine-grained details and reason across extended, interleaved video and subtitle sequences ranging from seconds up to one hour in length.
To construct this evaluation, the authors compiled 3,763 diverse web videos across 10 categories, pairing them with aligned subtitles across four duration brackets up to 60 minutes. They developed a novel question-answering paradigm termed referring reasoning, creating 6,678 human-annotated multiple-choice questions divided into perception tasks (grounded in single video moments) and relation tasks (requiring temporal or sequential reasoning across multiple moments). The benchmark was then used to evaluate 22 models—comprising proprietary leaders, open-source long-context systems, image-based models, and video-specific architectures—under zero-shot conditions.
The evaluation revealed several key findings regarding model capabilities. First, unlike prior benchmarks, performance on LongVideoBench strictly improves as models process more frames; leading proprietary models such as GPT-4o and Gemini-1.5-Pro improved by over 10 percentage points on videos exceeding three minutes when scaled from 16 to 256 frames. Second, a substantial capability gap exists between proprietary and open-source systems: top proprietary models achieved accuracy between 60% and 67%, while open-source models scored around 40% to 53% and suffered performance degradations when pushed beyond 16 frames. Third, reasoning across temporal relationships proved significantly harder than single-scene perception, with sequence ordering tasks showing the lowest overall scores. Finally, model accuracy was unevenly distributed across video duration, dropping when target queries were located near the beginning or middle of videos rather than near the end.
These results demonstrate that existing open-source and specialized video models are not yet equipped for real-world long-context video comprehension, largely because prior training datasets over-indexed on short clips and high-level summaries. For decision-makers and developers, relying on current open-source systems for complex video analysis introduces operational risk and high error rates, whereas proprietary services offer viable but imperfect accuracy at higher computational and API costs.
Organizations developing or deploying multimodal systems should prioritize training regimens that emphasize fine-grained multimodal retrieval and temporal reasoning rather than simple global captioning. For production deployments requiring hour-long video analysis, practitioners should deploy high-capacity models capable of processing dense frame inputs while monitoring for retrieval failures at earlier timestamps. While the benchmark provides high confidence through human-verified annotations and consistent test-validation results, current findings are limited to vision and text inputs under one hour and do not yet evaluate the audio modality.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). MVBench establishes foundational benchmarks and task designs for dynamic temporal video understanding that motivate the long-context interleaved evaluations in LongVideoBench.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench introduces rigorous, fine-grained multi-modal evaluation taxonomies and objective multiple-choice question-answering methodologies that LongVideoBench adapts for long-form video reasoning.
- Paper: Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, Muhammad Maaz et al. (2023). Video-ChatGPT pioneered instruction-tuned video-language understanding and standard video QA evaluation protocols that LongVideoBench scales to hour-long contexts.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). Video-LLaVA details the alignment of unified visual and linguistic representations, providing foundational context on how multimodal models process temporal video inputs.
- Paper: LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding, Yushi Bai et al. (2023). LongBench formulates the multi-task long-context evaluation paradigm that LongVideoBench extends from pure text to multimodal video-language sequences.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). MME provides a comprehensive, standardized evaluation framework for multimodal LLMs across fine-grained subtasks that informs the structural design of LongVideoBench.
- Paper: Adaptive Keyframe Sampling for Long Video Understanding, Xi Tang et al. (2025). This work introduces an adaptive keyframe sampling technique specifically validated on LongVideoBench to overcome long-video context constraints.
- Paper: M-LLM Based Video Frame Selection for Efficient Video Understanding, Kai Hu 0010 et al. (2025). This paper presents a plug-and-play multimodal frame selector and directly benchmarks its long-context QA improvements on LongVideoBench.
- Paper: TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning, Xiangyu Zeng 0004 et al. (2025). TimeSuite extends long-form video understanding by integrating grounded temporal tuning and timestamp supervision to enhance model reasoning over long benchmarks.
- Paper: Streaming Video Question-Answering with In-context Video KV-Cache Retrieval, Shangzhe Di et al. (2025). ReKV develops an in-context key-value cache retrieval mechanism for continuous video question-answering, addressing the long-horizon processing bottlenecks evaluated by LongVideoBench.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Video-MME expands comprehensive multimodal video benchmarking to incorporate full audio tracks alongside subtitles across broader temporal scales.
- Paper: HourVideo: 1-Hour Video-Language Understanding, Keshigeyan Chandrasegaran et al. (2024). HourVideo builds on long-duration video understanding evaluation by providing an hour-long benchmark tailored to first-person egocentric activities.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Qwen3-VL develops advanced long-context multimodal architectures capable of processing extensive video sequences evaluated in benchmarks like LongVideoBench.
