MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Kunchang LiYali WangYinan HeYizhuo LiYi WangYi LiuZun WangJilan XuGuo ChenPing Luo
Presents MVBench, a comprehensive benchmark comprising 20 dynamic video tasks designed to evaluate temporal reasoning in multimodal language models beyond static image perception, alongside VideoChat2, a progressive baseline that outperforms prior methods by over 15%.
Recent advances in artificial intelligence have produced multi-modal large language models capable of processing visual and textual data together. However, existing evaluation benchmarks primarily test static image comprehension, neglecting how well models understand changes over time in dynamic video. Current video benchmarks remain narrow, expensive to build, or vulnerable to subjective automated scoring. This creates a critical assessment gap for organizations seeking to deploy artificial intelligence systems in dynamic, real-world environments.
The article addresses this gap by introducing MVBench, a comprehensive benchmark designed to evaluate temporal video understanding across diverse tasks, and VideoChat2, an open-source baseline model engineered specifically for video comprehension.
The researchers developed MVBench by transforming static image understanding tasks into 20 distinct dynamic video tasks spanning perception and cognition, such as action sequence tracking, object trajectory, and counterfactual inference. They established an automated pipeline using eleven public video datasets to generate multiple-choice question-and-answer pairs, ensuring standardized evaluation without costly manual labeling or biased scoring. To advance model performance, the authors also developed VideoChat2 using a 2-million-sample instruction dataset across 34 sources and a three-stage progressive training strategy that aligns visual features with language representations.
Evaluation results show that existing multi-modal models perform poorly on dynamic video tasks. Many models scored near random guessing, with leading systems achieving around 35.5% accuracy—barely outperforming text-only baselines without video input. In contrast, the standard VideoChat2 model achieved 51.1% accuracy, outperforming existing models by more than 15 percentage points. An upgraded version based on the Mistral architecture reached 60.4% accuracy, exceeding commercial proprietary models such as GPT-4V (43.5%) by 16.9 percentage points. Ablation analyses revealed that performance gains are heavily driven by the quality of the visual encoder and instruction-tuning data rather than language model size alone.
These findings demonstrate that static image perception does not translate to dynamic temporal reasoning. For technical leaders and decision-makers, relying on standard vision-language models for video-based operational tasks presents substantial reliability risks. Building effective video artificial intelligence requires specialized temporal architectures and diverse video training rather than merely scaling traditional language models.
Organizations should adopt comprehensive temporal benchmarks like MVBench before deploying computer vision models in operational environments requiring sequence or event reasoning. Model developers should prioritize dedicated video pre-training backbones and diverse instruction tuning. Future development must focus on improving fine-grained localization, counting, and integrating additional modalities such as audio and subtitles, where current models continue to face limitations.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). It introduces the standardized fine-grained multiple-choice evaluation paradigm for multimodal large language models that MVBench directly builds upon and extends from static images to dynamic video.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). It provides the foundational diagnostic evaluation taxonomy for multimodal large language models, motivating MVBench's aim to address the neglect of dynamic temporal understanding in static benchmarks.
- Paper: Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, Hang Zhang et al. (2023). It presents one of the initial frameworks for adapting instruction-tuned large language models to video comprehension, establishing key baselines and architectural context for VideoChat2.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). It demonstrates unified visual representation and instruction tuning for video-language modeling, providing foundational baseline methodology assessed in video benchmarks.
- Paper: The “Something Something” Video Database for Learning and Evaluating Visual Common Sense, Raghav Goyal et al. (2017). It establishes essential video dataset tasks requiring genuine temporal reasoning that cannot be solved via isolated static frames, directly underpinning MVBench's task conversion rationale.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). It extends multimodal video benchmarking beyond short-clip temporal QA by introducing comprehensive evaluations across diverse durations, full-length videos, and multimodal audio-subtitle tracks.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It advances multimodal architecture and training by unifying single-image, multi-image, and video understanding within a single scalable framework evaluated on benchmarks like MVBench.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). It implements dynamic resolution and 3D spatial-temporal positional encodings to substantially improve multimodal model performance on extended video understanding benchmarks.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). It scales open-source multimodal models across vision encoders, data curation, and test-time reasoning to overcome the video understanding deficiencies identified in diagnostic suites like MVBench.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). It introduces native multimodal pretraining recipes to further enhance open-source model capabilities on complex visual and temporal reasoning tasks.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). It integrates multi-dimensional rotary embeddings and explicit temporal timestamps to address key spatiotemporal modeling bottlenecks identified by video benchmarks.
