Multimodal Table Understanding

Mingyu ZhengXinwei FengQingyi SiQiaoqiao SheZheng LinWenbin JiangWeiping Wang

article2024ACL49 citations

Presents the MMTab dataset and Table-LLaVA, an open-source multimodal model that directly interprets table images without requiring text serialization, outperforming existing vision-language baselines across 23 tabular benchmarks.

Listen

Tables are critical for presenting structured information across finance, scientific research, and reporting. While modern artificial intelligence has advanced table processing, current approaches rely heavily on converting tables into raw text sequences such as Markdown or HTML. In practice, obtaining clean textual representations from sources like scanned documents and webpage screenshots is difficult, whereas visual table images are easily accessible. Furthermore, standard language models analyze tables in a one-directional textual flow, which fails to capture two-dimensional spatial layouts, colored elements, and hierarchical structures. Existing multimodal large language models, which process both images and text, also struggle significantly with table images because they lack specialized tabular training.

The article addresses this gap by defining the multimodal table understanding problem, in which models must answer diverse user requests directly from table images. The primary objective is to build a foundational, open-source dataset to support this task and develop a versatile tabular multimodal model capable of interpreting diverse table structures and answering complex questions without requiring text conversion.

To accomplish this, the authors constructed MMTab, a comprehensive dataset derived from 14 public sources spanning 8 domains. MMTab contains 105,000 rendered table images across web page, Excel, and Markdown visual styles, paired with diverse instructions formatted as input-request and output-response pairs. It includes 150,000 pre-training samples, 232,000 instruction-tuning samples covering 14 tasks, and a test suite of 49,000 samples across 17 held-in and 7 held-out benchmarks. Using this dataset, the authors developed Table-LLaVA through a two-stage training strategy: first pre-training the model on visual-to-text table recognition to ground basic layout perception, followed by instruction tuning on downstream tabular and structural tasks.

Evaluation showed that Table-LLaVA significantly outperforms existing open-source multimodal models across nearly all benchmarks. On academic table tasks such as question answering, fact verification, and text generation, Table-LLaVA achieved substantial performance gains, while older open-source multimodal models scored near zero. On basic structure understanding tasks—such as detecting table dimensions, locating cells, and identifying merged regions—Table-LLaVA outperformed existing open-source models, which struggled to grasp basic layout geometry. In comparative testing against GPT-4V across 14 benchmarks, Table-LLaVA achieved competitive or superior results compared to the low-resolution baseline, though high-resolution GPT-4V retained an advantage on large, complex tables. Ablation studies confirmed that table recognition pre-training and structure understanding tasks were critical drivers of overall accuracy and generalizability, while also providing modest performance benefits on non-tabular visual tasks.

These findings indicate that directly processing visual tables is a viable and effective alternative to fragile optical character recognition (OCR) pipelines, reducing the risk of cascading errors in enterprise workflows. By demonstrating that tabular training enhances broader visual and reasoning performance, the article shows that structural perception is an essential capability for general-purpose artificial intelligence models.

Organizations handling high volumes of tabular images should consider adopting end-to-end multimodal architectures rather than multi-step text-extraction pipelines. Next steps should focus on scaling the model to support higher input image resolutions, which is necessary to preserve legibility in large, dense spreadsheets and corporate filings. Future work should also extend datasets and training to support multilingual tables, multi-table document reasoning, and imperfect real-world imagery such as distorted, handwritten, or low-resolution scans.

arXiv: 2406.08100SpursGoZmy/Table-LLaVA
Cover for Multimodal Table Understanding

Abstract

Although great progress has been made by previous table understanding methods including recent approaches based on large language models (LLMs), they rely heavily on the premise that given tables must be converted into a certain text sequence (such as Markdown or HTML) to serve as model input. However, it is difficult to access such high-quality textual table representations in some real-world scenarios, and table images are much more accessible. Therefore, how to directly understand tables using intuitive visual information is a crucial and urgent challenge for developing more practical applications. In this paper, we propose a new problem, multimodal table understanding, where the model needs to generate correct responses to various table-related requests based on the given table image. To facilitate both the model training and evaluation, we construct a large-scale dataset named MMTab, which covers a wide spectrum of table images, instructions and tasks. On this basis, we develop Table-LLaVA, a generalist tabular multimodal large language model (MLLM), which significantly outperforms recent open-source MLLM baselines on 23 benchmarks under held-in and held-out settings. The code and data is available at https://github.com/SpursGoZmy/Table-LLaVA.

Citation

MLA
Zheng, M., et al. “Multimodal Table Understanding”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 9102–24, https://doi.org/10.18653/v1/2024.acl-long.493.
APA
Zheng, M., Feng, X., Si, Q., She, Q., Lin, Z., Jiang, W., & Wang, W. (2024). Multimodal Table Understanding. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9102–9124. https://doi.org/10.18653/v1/2024.acl-long.493
Chicago
Zheng, M., X. Feng, Q. Si, et al. 2024. “Multimodal Table Understanding”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 9102–24. https://doi.org/10.18653/v1/2024.acl-long.493.
Harvard
Zheng, M. et al. (2024) “Multimodal Table Understanding”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 9102–9124. Available at: https://doi.org/10.18653/v1/2024.acl-long.493.
Vancouver
1. Zheng M, Feng X, Si Q, She Q, Lin Z, Jiang W, Wang W (2024) Multimodal Table Understanding. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 9102–9124

BibTeX

@inproceedings{zheng-etal-2024-multimodal,
    title = "Multimodal Table Understanding",
    author = "Zheng, Mingyu  and
      Feng, Xinwei  and
      Si, Qingyi  and
      She, Qiaoqiao  and
      Lin, Zheng  and
      Jiang, Wenbin  and
      Wang, Weiping",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.493/",
    doi = "10.18653/v1/2024.acl-long.493",
    pages = "9102--9124"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/