Built independently by an author, for readers. Read the story and support ChapterPal

keyword

video comprehension

Video comprehension is the ability of an artificial intelligence system to interpret, analyze, and reason about the dynamic content of sequential visual media over time. Moving beyond the recognition of static images, it involves modeling temporal dependencies across video frames to track objects, recognize actions, and understand narrative flow and causal relationships. Advanced video comprehension often integrates multimodal inputs, combining visual signals with synchronized audio and text transcripts to perform complex cognitive tasks such as video question answering, event summarization, and contextual reasoning across varying durations of footage.

1 item

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He, Xing Sun

OrganizationsEast China Normal UniversityInstitute of Automation, Chinese Academy of SciencesNanjing UniversityPeking UniversityState Key Laboratory of Cognitive IntelligenceThe Chinese University of Hong KongUniversity of Hong KongXiamen University

Why you should read this

Introduces Video-MME, a pioneering full-spectrum benchmark that evaluates multi-modal large language models across diverse domains and durations up to one hour, exposing current limitations in long-range temporal reasoning while demonstrating the significant impact of integrated subtitle and audio data on video comprehension.

In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual data is still insufficiently explored, highlighting the lack of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, and reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models with an average accuracy of 75%, compared to 71.9% for GPT-4o. The results also demonstrate that Video-MME is a universal benchmark that applies to both image and video MLLMs. Further analysis indicates that subtitle and audio information could significantly enhance video understanding. Besides, a decline in MLLM performance is observed as video duration increases for all models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data, shedding light on future MLLM development. Project page: https://video-mme.github.io.

Added

2026-05-13