MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Pan LuHritik BansalTony XiaJiacheng LiuChun-yue LiHannaneh HajishirziHao ChengKai-Wei ChangMichel GalleyJianfeng Gao
Presents MathVista, a 6,141-example benchmark for assessing mathematical reasoning in visual contexts, revealing that leading multimodal models like GPT-4V still trail human performance on complex visual problem-solving.
Artificial intelligence systems have advanced rapidly in solving complex text-based problems, but real-world decision-making often requires interpreting visual data such as charts, plots, and diagrams alongside rigorous numerical computation. Prior benchmarks have largely evaluated mathematical problem-solving through text alone, leaving the multimodal reasoning capabilities of current AI models unmeasured in visually rich environments. The article presents MathVista, a consolidated benchmark designed to systematically evaluate how effectively leading foundation models perform mathematical reasoning within diverse visual contexts.
The researchers constructed a standardized evaluation framework comprising 6,141 multimodal questions drawn from 28 existing datasets and three newly introduced datasets covering logic puzzles, function plots, and academic figures. The benchmark spans seven mathematical reasoning categories—such as arithmetic, geometry, algebra, and statistics—across 19 visual context formats. Using this suite, the authors tested 12 prominent artificial intelligence models under text-only, tool-augmented, and multimodal configurations, establishing both automated and human performance baselines to quantify current technological capabilities.
The findings show that GPT-4V achieved the highest overall performance among all tested models with an accuracy of 49.9%, outperforming the second-best model, Multimodal Bard, by 15.1 percentage points. Despite this lead, GPT-4V still lagged behind the human benchmark of 60.3% by 10.4 percentage points. Open-source multimodal models trailed significantly, achieving overall accuracies between 19.8% and 26.1%. A qualitative evaluation of model outputs revealed that approximately half of Multimodal Bard's flawed explanations involved hallucinations or incorrect fact extraction from visual inputs, while tool-augmented language models showed improvements primarily when supplied with high-accuracy optical character recognition and visual captions.
These results indicate that while advanced multimodal foundation models possess strong foundational capabilities—even outperforming humans in specialized narrow areas like algebraic function plots and geometry—they remain prone to perceptual errors and calculation mistakes in complex, multistep workflows. For organizations considering the operational deployment of multimodal AI assistants in scientific analysis, finance, or education, these limitations present compliance, safety, and reliability risks if models operate without human oversight.
The authors recommend that developers focus future efforts on improving the visual perception of complex figures, mitigating multimodal hallucinations, and integrating self-verification and programmatic reasoning tools to boost calculation reliability. Additionally, the benchmark itself faces constraints, including uneven annotation depth across heterogeneous data sources and potential data noise. Because a measurable performance gap remains between automated models and human accuracy, stakeholders should treat current vision-language models as assistive tools requiring active validation rather than autonomous reasoning engines.
- Paper: Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering, Pan Lu et al. (2022). Introduces the ScienceQA benchmark and multimodal chain-of-thought reasoning, establishing the foundation for evaluating complex scientific problem-solving with vision-language models.
- Paper: Measuring Mathematical Problem Solving With the MATH Dataset, Dan Hendrycks et al. (2021). Presents the competition-level MATH benchmark that MathVista incorporates and extends into visual and multimodal contexts.
- Paper: Training Verifiers to Solve Math Word Problems, Karl Cobbe et al. (2021). Establishes the GSM8K benchmark and self-verification paradigms for multi-step mathematical reasoning that MathVista evaluates across vision-language foundation models.
- Paper: Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Jason Wei et al. (2022). Demonstrates chain-of-thought prompting for eliciting step-by-step reasoning in foundation models, a core prompting technique leveraged throughout MathVista.
- Paper: Solving Quantitative Reasoning Problems with Language Models, Aitor Lewkowycz et al. (2022). Introduces Minerva, providing essential baseline methodology and benchmarks for evaluating large language models on complex quantitative and scientific reasoning.
- Paper: GPT-4 Technical Report, OpenAI (2023). Details the architecture and multimodal capabilities of GPT-4 and GPT-4V, which serves as the primary frontier model evaluated and analyzed in MathVista.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). Defines the foundational Visual Question Answering (VQA) problem formulation and dataset construction upon which multimodal evaluation benchmarks build.
- Paper: CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning, Justin Johnson et al. (2016). Pioneers diagnostic visual question answering designed to evaluate multi-step compositionality and elementary visual reasoning without statistical shortcuts.
- Paper: Visual Instruction Tuning, Haotian Liu et al. (2023). Introduces the LLaVA framework for visual instruction tuning, establishing the standard open-source paradigm for multimodal chat and reasoning evaluated by MathVista.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Establishes the TextVQA benchmark, highlighting the necessity of optical character recognition and visual text interpretation for complex multimodal QA.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). Develops the unified LLaVA-OneVision architecture and directly uses MathVista as a benchmark to assess open-source progress in mathematical visual reasoning.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). Advances high-resolution multimodal perception with dynamic resolution in Qwen2-VL, directly addressing visual perception bottlenecks identified in MathVista.
- Paper: Improved Baselines with Visual Instruction Tuning, Haotian Liu et al. (2024). Builds improved baselines for visual instruction tuning (LLaVA-1.5) to enhance fine-grained perception and multimodal reasoning capability.
- Paper: Innovator-VL: A Multimodal Large Language Model for Scientific Discovery, Zichen Wen et al. (2026). Extends visual reasoning evaluation to expert-level multimodal scientific discovery and complex technical domains.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). Expands comprehensive multimodal foundation model evaluation from static images and mathematical figures to dynamic video sequences.
- Paper: DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models, Zhihong Shao et al. (2024). Explores reinforcement learning algorithms like GRPO to substantially boost open foundation models on complex mathematical reasoning benchmarks.
- Paper: DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI et al. (2025). Incentivizes deliberate System 2 reasoning and long-chain self-verification in large models through large-scale reinforcement learning.
