MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Pan LuHritik BansalTony XiaJiacheng LiuChun-yue LiHannaneh HajishirziHao ChengKai-Wei ChangMichel GalleyJianfeng Gao

article2023ICLR1,888 citations

Presents MathVista, a 6,141-example benchmark for assessing mathematical reasoning in visual contexts, revealing that leading multimodal models like GPT-4V still trail human performance on complex visual problem-solving.

Listen

Artificial intelligence systems have advanced rapidly in solving complex text-based problems, but real-world decision-making often requires interpreting visual data such as charts, plots, and diagrams alongside rigorous numerical computation. Prior benchmarks have largely evaluated mathematical problem-solving through text alone, leaving the multimodal reasoning capabilities of current AI models unmeasured in visually rich environments. The article presents MathVista, a consolidated benchmark designed to systematically evaluate how effectively leading foundation models perform mathematical reasoning within diverse visual contexts.

The researchers constructed a standardized evaluation framework comprising 6,141 multimodal questions drawn from 28 existing datasets and three newly introduced datasets covering logic puzzles, function plots, and academic figures. The benchmark spans seven mathematical reasoning categoriessuch as arithmetic, geometry, algebra, and statisticsacross 19 visual context formats. Using this suite, the authors tested 12 prominent artificial intelligence models under text-only, tool-augmented, and multimodal configurations, establishing both automated and human performance baselines to quantify current technological capabilities.

The findings show that GPT-4V achieved the highest overall performance among all tested models with an accuracy of 49.9%, outperforming the second-best model, Multimodal Bard, by 15.1 percentage points. Despite this lead, GPT-4V still lagged behind the human benchmark of 60.3% by 10.4 percentage points. Open-source multimodal models trailed significantly, achieving overall accuracies between 19.8% and 26.1%. A qualitative evaluation of model outputs revealed that approximately half of Multimodal Bard's flawed explanations involved hallucinations or incorrect fact extraction from visual inputs, while tool-augmented language models showed improvements primarily when supplied with high-accuracy optical character recognition and visual captions.

These results indicate that while advanced multimodal foundation models possess strong foundational capabilitieseven outperforming humans in specialized narrow areas like algebraic function plots and geometrythey remain prone to perceptual errors and calculation mistakes in complex, multistep workflows. For organizations considering the operational deployment of multimodal AI assistants in scientific analysis, finance, or education, these limitations present compliance, safety, and reliability risks if models operate without human oversight.

The authors recommend that developers focus future efforts on improving the visual perception of complex figures, mitigating multimodal hallucinations, and integrating self-verification and programmatic reasoning tools to boost calculation reliability. Additionally, the benchmark itself faces constraints, including uneven annotation depth across heterogeneous data sources and potential data noise. Because a measurable performance gap remains between automated models and human accuracy, stakeholders should treat current vision-language models as assistive tools requiring active validation rather than autonomous reasoning engines.

Cover for MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Abstract

Large Language Models (LLMs) and Large Multimodal Models (LMMs) exhibit impressive problem-solving skills in many tasks and domains, but their ability in mathematical reasoning in visual contexts has not been systematically studied. To bridge this gap, we present MathVista, a benchmark designed to combine challenges from diverse mathematical and visual tasks. It consists of 6,141 examples, derived from 28 existing multimodal datasets involving mathematics and 3 newly created datasets (i.e., IQTest, FunctionQA, and PaperQA). Completing these tasks requires fine-grained, deep visual understanding and compositional reasoning, which all state-of-the-art foundation models find challenging. With MathVista, we have conducted a comprehensive, quantitative evaluation of 12 prominent foundation models. The best-performing GPT-4V model achieves an overall accuracy of 49.9%, substantially outperforming Bard, the second-best performer, by 15.1%. Our in-depth analysis reveals that the superiority of GPT-4V is mainly attributed to its enhanced visual perception and mathematical reasoning. However, GPT-4V still falls short of human performance by 10.4%, as it often struggles to understand complex figures and perform rigorous reasoning. This significant gap underscores the critical role that MathVista will play in the development of general-purpose AI agents capable of tackling mathematically intensive and visually rich real-world tasks. We further explore the new ability of self-verification, the application of self-consistency, and the interactive chatbot capabilities of GPT-4V, highlighting its promising potential for future research. The project is available at this https URL.

Citation

MLA
Lu, P., et al. “MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts”. arXiv, 2023, http://arxiv.org/abs/2310.02255v3.
APA
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., & Gao, J. (2023). MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. arXiv. http://arxiv.org/abs/2310.02255v3
Chicago
Lu, P., H. Bansal, T. Xia, et al. 2023. “MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts”. arXiv. http://arxiv.org/abs/2310.02255v3.
Harvard
Lu, P. et al. (2023) “MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2310.02255v3.
Vancouver
1. Lu P, Bansal H, Xia T, Liu J, Li C, Hajishirzi H, Cheng H, Chang K-W, Galley M, Gao J (2023) MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts. arXiv

BibTeX

@article{lu2023mathvista,
  title = {MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts},
  author = {Lu, Pan and Bansal, Hritik and Xia, Tony and Liu, Jiacheng and Li, Chunyuan and Hajishirzi, Hannaneh and Cheng, Hao and Chang, Kai-Wei and Galley, Michel and Gao, Jianfeng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2310.02255v3},
  eprint = {2310.02255}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/