Multimodal Table Understanding
Mingyu ZhengXinwei FengQingyi SiQiaoqiao SheZheng LinWenbin JiangWeiping Wang
Presents the MMTab dataset and Table-LLaVA, an open-source multimodal model that directly interprets table images without requiring text serialization, outperforming existing vision-language baselines across 23 tabular benchmarks.
Tables are critical for presenting structured information across finance, scientific research, and reporting. While modern artificial intelligence has advanced table processing, current approaches rely heavily on converting tables into raw text sequences such as Markdown or HTML. In practice, obtaining clean textual representations from sources like scanned documents and webpage screenshots is difficult, whereas visual table images are easily accessible. Furthermore, standard language models analyze tables in a one-directional textual flow, which fails to capture two-dimensional spatial layouts, colored elements, and hierarchical structures. Existing multimodal large language models, which process both images and text, also struggle significantly with table images because they lack specialized tabular training.
The article addresses this gap by defining the multimodal table understanding problem, in which models must answer diverse user requests directly from table images. The primary objective is to build a foundational, open-source dataset to support this task and develop a versatile tabular multimodal model capable of interpreting diverse table structures and answering complex questions without requiring text conversion.
To accomplish this, the authors constructed MMTab, a comprehensive dataset derived from 14 public sources spanning 8 domains. MMTab contains 105,000 rendered table images across web page, Excel, and Markdown visual styles, paired with diverse instructions formatted as input-request and output-response pairs. It includes 150,000 pre-training samples, 232,000 instruction-tuning samples covering 14 tasks, and a test suite of 49,000 samples across 17 held-in and 7 held-out benchmarks. Using this dataset, the authors developed Table-LLaVA through a two-stage training strategy: first pre-training the model on visual-to-text table recognition to ground basic layout perception, followed by instruction tuning on downstream tabular and structural tasks.
Evaluation showed that Table-LLaVA significantly outperforms existing open-source multimodal models across nearly all benchmarks. On academic table tasks such as question answering, fact verification, and text generation, Table-LLaVA achieved substantial performance gains, while older open-source multimodal models scored near zero. On basic structure understanding tasks—such as detecting table dimensions, locating cells, and identifying merged regions—Table-LLaVA outperformed existing open-source models, which struggled to grasp basic layout geometry. In comparative testing against GPT-4V across 14 benchmarks, Table-LLaVA achieved competitive or superior results compared to the low-resolution baseline, though high-resolution GPT-4V retained an advantage on large, complex tables. Ablation studies confirmed that table recognition pre-training and structure understanding tasks were critical drivers of overall accuracy and generalizability, while also providing modest performance benefits on non-tabular visual tasks.
These findings indicate that directly processing visual tables is a viable and effective alternative to fragile optical character recognition (OCR) pipelines, reducing the risk of cascading errors in enterprise workflows. By demonstrating that tabular training enhances broader visual and reasoning performance, the article shows that structural perception is an essential capability for general-purpose artificial intelligence models.
Organizations handling high volumes of tabular images should consider adopting end-to-end multimodal architectures rather than multi-step text-extraction pipelines. Next steps should focus on scaling the model to support higher input image resolutions, which is necessary to preserve legibility in large, dense spreadsheets and corporate filings. Future work should also extend datasets and training to support multilingual tables, multi-table document reasoning, and imperfect real-world imagery such as distorted, handwritten, or low-resolution scans.
- Paper: SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables, Xinyuan Lu et al. (2023). Its scientific-table claim-verification benchmark provides a concrete predecessor to the source’s evaluation of fact verification from table images.
- Paper: ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding, Xingyu Fu et al. (2025). Building on direct visual understanding of structured tables and charts, ReFocus adds iterative image editing so models can focus their visual reasoning on relevant regions.
