MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering
Fangyu LiuFrancesco PiccinnoSyrine KricheneChenxi PangKenton LeeMandar JoshiYasemin AltunNigel CollierJulian Martin Eisenschlos
Proposes a pretraining framework combining chart derendering and math reasoning tasks that improves visual language models, boosting performance on ChartQA and PlotQA by up to 20% while generalizing effectively to diverse visual document tasks.
Visual language data, such as charts, plots, diagrams, and infographics, are essential for communicating complex information across documents, reports, and websites. Despite advances in vision-language artificial intelligence, existing models struggle to interpret these visual assets because standard natural-image training fails to capture layout organization, numerical extraction, and mathematical reasoning. Addressing this limitation is critical for automating data analysis and building reliable multimodal systems.
The article demonstrates a pretraining method called MATCHA (Math reasoning and Chart derendering pretraining) designed to improve an image-to-text model's ability to process visual language. The approach continually trains an existing base vision model, Pix2Struct, using a balanced mixture of tasks: 40% chart derendering (converting visual plots back into underlying data tables or rendering code), 40% mathematical reasoning (solving text-based arithmetic and comparison problems converted into images), and 20% webpage screenshot parsing.
The findings show that MATCHA establishes new state-of-the-art performance across several standard benchmarks. On chart question answering benchmarks, MATCHA achieves an overall accuracy of 64.2% on ChartQA and 91.5% on PlotQA, outperforming previous top models that lack access to underlying source tables by 8.2% and up to 19% respectively. In fact, MATCHA surpasses baseline systems that were provided with the actual underlying data tables. The model also sets a new benchmark on chart summarization and transfers effectively beyond plots, improving average performance across other document and interface understanding tasks by 2.3%.
These results demonstrate that visual layout deconstruction and mathematical reasoning are the core capabilities needed for visual language understanding. By training a relatively compact 300-million-parameter model on these targeted tasks, MATCHA significantly outperforms substantially larger general-purpose vision models (such as the 17-billion-parameter PaLI) on chart reasoning while reducing computational demands and eliminating reliance on brittle optical character recognition pipelines.
For practical implementation, organizations should adopt chart derendering and numerical reasoning objectives when building automated chart-processing and document-intelligence workflows. Future work should explore hybrid approaches—such as pairing chart-to-table translation modules with external calculators or symbolic program executors—to solve high-precision arithmetic where pure neural computation still makes errors.
Key limitations include difficulty with highly complex calculations requiring exact numerical precision, weaker performance on visual plot attributes (such as identifying colors or shapes) compared to massive web-scale vision models, and reliance on single-run evaluations due to computational constraints. Readers should maintain moderate confidence in the reported metrics while exercising caution in deployment scenarios that require strict arithmetic precision.
- Paper: ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning, Ahmed Masry et al. (2022). Introduces the ChartQA benchmark for visual and mathematical reasoning over real-world charts, which serves as a primary evaluation target and design motivation for MatCha.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). Establishes visual question answering on complex document layouts and diagrams, providing essential task foundations for MatCha's document transfer evaluations.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Pioneers text reading and reasoning within visual question answering, formalizing the core optical and textual extraction principles underpinning MatCha's derendering objectives.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). Presents a unified transformer baseline for vision-and-language tasks, establishing foundational cross-modal pretraining concepts built upon by models like Pix2Struct and MatCha.
- Paper: Are NLP Models really able to Solve Simple Math Word Problems?, Arkil Patel et al. (2021). Demonstrates the necessity of robust mathematical reasoning beyond superficial statistical shortcuts, motivating MatCha's targeted numerical reasoning pretraining tasks.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). Extends the evaluation of multimodal foundation models by benchmarking mathematical reasoning across visual contexts including charts, plots, and scientific diagrams.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Applies localized visual chain-of-thought grounding and multi-turn reasoning to interpret dense visual data like charts and documents.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). Generalizes multimodal reasoning evaluation to multi-domain, multi-step chain-of-thought tasks requiring complex mathematics and visual comprehension.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Incorporates high-resolution text reading and visual localization into generalist vision-language foundation models.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). Assesses integrated multimodal capabilities including arithmetic calculation and optical character recognition across complex visual inputs.
- Paper: Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning, Pan Lu et al. (2023). Builds on structured mathematical problem solving by using dynamic prompt learning for semi-structured tabular and visual math tasks.
