MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Kaining YingFanqing MengJin WangZhiqian LiHan LinYue YangHao ZhangWenbo ZhangYuqi LinShuo Liu
Presents MMT-Bench, an extensive evaluation benchmark spanning over 31,000 questions across 162 vision-language tasks, to expose key performance limits in advanced multimodal models and map their progress toward general visual intelligence.
Artificial intelligence systems that integrate vision and text are advancing rapidly, driving interest in multimodal artificial general intelligence capable of performing diverse human-level tasks across various domains. However, standard evaluation benchmarks remain constrained by narrow task coverage and simple visual questions, failing to effectively measure progress toward broad, expert-level visual intelligence. This creates a critical need for rigorous testing environments that assess how well models generalize across diverse image modalities and complex reasoning scenarios.
The article introduces MMT-Bench, a comprehensive benchmark designed to evaluate large vision-language models across massive multimodal tasks requiring expert knowledge, spatial localization, fine-grained perception, and deliberate reasoning. It aims to quantify current system performance, uncover task interdependencies, and identify specific domains where existing models succeed or struggle.
To establish this benchmark, the researchers curated 31,325 multiple-choice visual questions spanning 32 core meta-tasks, 162 subtasks, and 13 distinct visual input types such as natural scenes, medical images, depth maps, and graphical user interfaces. Using this dataset, the authors evaluated 32 leading closed-source and open-source models under standardized protocols. In addition, the study mapped the relationships among tasks by constructing task vectors through parameter-efficient fine-tuning, grouping tasks into clusters to analyze in-domain and out-of-domain performance trends.
The evaluation produced several critical findings regarding current model capabilities. First, the benchmark presents a significant challenge to state-of-the-art models: the top-performing closed-source system, GPT-4o, achieved an overall accuracy of only 65.5%, which fell to 59.5% when standard visual recognition tasks were excluded. Second, leading open-source models demonstrated strong competitiveness, with InternVL-Chat-v1.2 scoring 63.4% and outperforming several proprietary systems like GPT-4V (61.1%) and GeminiProVision (61.6%). Third, detailed error analyses showed that model failures are predominantly driven by perception errors (51% to 77% across top models) and complex reasoning deficits. Finally, task mapping revealed that while models excel at high-level recognition and captioning, they consistently underperform in out-of-domain areas requiring fine-grained spatial localization, coordinate detection, and user interface navigation.
These results demonstrate that high performance on conventional recognition tasks does not translate to robust spatial reasoning or operational execution in specialized environments. For organizations deploying these technologies, the findings indicate substantial operational risks when using current foundation models for tasks requiring precise pixel-level localization, document structuring, or graphical interface interaction. Notably, instruction tuning does not uniformly improve generalization, as un-tuned baseline models outperformed several instruction-tuned counterparts on structured closed-set evaluations.
To improve multimodal performance, developers should expand training regimens to include visual referring data, normalized coordinates, and multi-image sequence inputs. Organizations evaluating or deploying vision-language models should prioritize addressing core perception and spatial reasoning limitations rather than relying exclusively on high-level recognition benchmarks. Future work should focus on incorporating a broader range of multimodal tasks and developing training techniques that mitigate performance trade-offs across distinct task domains.
The benchmark's findings should be interpreted with awareness of its structured multiple-choice design and curated data sources, which may introduce domain weighting biases. While the evaluation demonstrates high internal consistency and reliability across tested models, stakeholders should exercise caution when extrapolating these scores to unconstrained real-world environments with unrepresented demographic contexts or visual modalities.
- Paper: MMBench: Is Your Multi-modal Model an All-around Player?, Yuanzhan Liu et al. (2023). MMBench establishes the foundational fine-grained taxonomy and robust circular evaluation strategy for multiple-choice vision-language benchmarking that MMT-Bench builds upon and scales.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). MMMU introduces massive multi-discipline expert-level evaluation for multimodal models, establishing the standard for AGI-oriented multimodal benchmarks that MMT-Bench broadens to diverse meta-tasks.
- Paper: MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, Chaoyou Fu et al. (2023). MME provides the initial systematic framework for evaluating multimodal LLMs across distinct perception and cognition subtasks, setting the precedent for comprehensive multi-task assessment.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). MM-Vet conceptualizes the systematic evaluation of integrated multimodal capabilities across combined subskills, directly informing MMT-Bench's meta-task structuring.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). MathVista formalizes deliberate visual reasoning and problem-solving benchmarks across diverse multimodal contexts, which MMT-Bench incorporates within its comprehensive task scope.
- Paper: Measuring Massive Multitask Language Understanding, Dan Hendrycks et al. (2020). MMLU is the classic massive multitask multiple-choice evaluation paradigm that MMT-Bench adapts and extends into the multimodal domain.
- Paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark, Kunchang Li et al. (2023). MVBench demonstrates how to evaluate multimodal models across diverse dynamic tasks, informing MMT-Bench's inclusion of scenarios like vehicle driving and embodied navigation.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Qwen-VL represents a seminal general-purpose open-source LVLM that defined foundational capabilities in visual understanding, localization, and text reading assessed by MMT-Bench.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). InternVL3 represents a next-generation open-source multimodal foundation model developed to tackle the multidisciplinary challenges and comprehensive task maps highlighted by benchmarks like MMT-Bench.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). Qwen3-VL develops advanced native long-context and multi-dimensional spatial-temporal reasoning architectures in response to broad multitask visual intelligence evaluations.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Video-MME extends comprehensive multi-task multimodal evaluation from static images and scenarios into full-length, multi-domain dynamic video understanding.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision builds on comprehensive visual benchmarking insights by creating a unified multimodal framework capable of transferring across single-image, multi-image, and video tasks.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). InternVL 2.5 applies scaling and test-time reasoning techniques to enhance open-source LVLM performance on comprehensive multimodal reasoning benchmarks.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). M3CoT deepens multimodal benchmarking by focusing specifically on multi-step chain-of-thought visual reasoning across complex scientific and mathematical domains.
- Paper: Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces, Jihan Yang et al. (2025). VSI-Bench extends multimodal evaluation by probing how vision-language models remember, recall, and reason through 3D physical spaces in video environments.
- Paper: MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark, S. Sakshi et al. (2025). MMAU generalizes the massive multitask evaluation philosophy to auditory understanding and expert-level reasoning across speech, sound, and music.
