Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Video-MME benchmark

The Video-MME benchmark is a comprehensive evaluation dataset designed to assess the capabilities of multimodal large language models in video analysis and understanding. It measures how effectively artificial intelligence systems can perceive, interpret, and reason about dynamic visual content through expert-annotated question-answering tasks. The benchmark distinguishes itself by covering a wide variety of video genres and visual domains across diverse temporal lengths, ranging from brief clips of several seconds to long-form videos up to an hour. In addition to sequential video frames, it incorporates complementary modalities such as audio streams and subtitles to provide a full-spectrum assessment of temporal reasoning and cross-modal comprehension.

1 item

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He, Xing Sun

OrganizationsEast China Normal UniversityInstitute of Automation, Chinese Academy of SciencesNanjing UniversityPeking UniversityState Key Laboratory of Cognitive IntelligenceThe Chinese University of Hong KongUniversity of Hong KongXiamen University

Why you should read this

Introduces Video-MME, a pioneering full-spectrum benchmark that evaluates multi-modal large language models across diverse domains and durations up to one hour, exposing current limitations in long-range temporal reasoning while demonstrating the significant impact of integrated subtitle and audio data on video comprehension.

In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual data is still insufficiently explored, highlighting the lack of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, and reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models with an average accuracy of 75%, compared to 71.9% for GPT-4o. The results also demonstrate that Video-MME is a universal benchmark that applies to both image and video MLLMs. Further analysis indicates that subtitle and audio information could significantly enhance video understanding. Besides, a decline in MLLM performance is observed as video duration increases for all models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data, shedding light on future MLLM development. Project page: https://video-mme.github.io.

Added

2026-05-13