Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Multi-modal LLMs

Multi-modal large language models are artificial intelligence systems that extend traditional text-based language models to process, interpret, and generate content across multiple data modalities, including text, images, video, and audio. These architectures typically combine modality-specific neural network encoders with a large language model backbone acting as a central reasoning engine. By mapping diverse sensory inputs into a unified representation space, multi-modal large language models can perform complex cross-modal tasks such as visual question answering, document parsing, audio transcription, and temporal video comprehension while retaining general conversational and reasoning capabilities.

1 item

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Chaoyou Fu, Yuhan Dai, Yongdong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Caifeng Shan, Ran He, Xing Sun

OrganizationsEast China Normal UniversityInstitute of Automation, Chinese Academy of SciencesNanjing UniversityPeking UniversityState Key Laboratory of Cognitive IntelligenceThe Chinese University of Hong KongUniversity of Hong KongXiamen University

Why you should read this

Introduces Video-MME, a pioneering full-spectrum benchmark that evaluates multi-modal large language models across diverse domains and durations up to one hour, exposing current limitations in long-range temporal reasoning while demonstrating the significant impact of integrated subtitle and audio data on video comprehension.

In the quest for artificial general intelligence, Multi-modal Large Language Models (MLLMs) have emerged as a focal point in recent advancements. However, the predominant focus remains on developing their capabilities in static image understanding. The potential of MLLMs to process sequential visual data is still insufficiently explored, highlighting the lack of a comprehensive, high-quality assessment of their performance. In this paper, we introduce Video-MME, the first-ever full-spectrum, Multi-Modal Evaluation benchmark of MLLMs in Video analysis. Our work distinguishes from existing benchmarks through four key features: 1) Diversity in video types, spanning 6 primary visual domains with 30 subfields to ensure broad scenario generalizability; 2) Duration in temporal dimension, encompassing both short-, medium-, and long-term videos, ranging from 11 seconds to 1 hour, for robust contextual dynamics; 3) Breadth in data modalities, integrating multi-modal inputs besides video frames, including subtitles and audios, to unveil the all-round capabilities of MLLMs; 4) Quality in annotations, utilizing rigorous manual labeling by expert annotators to facilitate precise and reliable model assessment. With Video-MME, we extensively evaluate various state-of-the-art MLLMs, and reveal that Gemini 1.5 Pro is the best-performing commercial model, significantly outperforming the open-source models with an average accuracy of 75%, compared to 71.9% for GPT-4o. The results also demonstrate that Video-MME is a universal benchmark that applies to both image and video MLLMs. Further analysis indicates that subtitle and audio information could significantly enhance video understanding. Besides, a decline in MLLM performance is observed as video duration increases for all models. Our dataset along with these findings underscores the need for further improvements in handling longer sequences and multi-modal data, shedding light on future MLLM development. Project page: https://video-mme.github.io.

Added

2026-05-13