Chart-to-Text: A Large-Scale Benchmark for Chart Summarization
Shankar KantharajRixie Tiffany Ko LeongXiang LinAhmed MasryMegh ThakkarEnamul HoqueShafiq R. Joty
Presents a large-scale benchmark of over 44,000 diverse charts alongside state-of-the-art neural baselines to evaluate automated chart summarization from both raw images and underlying data tables.
Visual data representations such as bar, line, and pie charts are essential across modern organizations for communicating insights and supporting decision-making. However, extracting key takeaways directly from charts requires significant cognitive effort, and charts often lack clear explanatory captions. Generating natural language summaries automatically can assist report writers, improve document indexing, and enhance accessibility for individuals who rely on screen readers. Existing methods remain limited because they rely heavily on rigid templates, lack large-scale benchmark data, or focus narrowly on raw tables without addressing the unique visual structures and high-level trends conveyed by charts.
The article introduces a large-scale benchmark and establishes baseline models to advance automated chart summarization. It evaluates text generation across two practical scenarios: one where the underlying structured data table is available, and a more realistic and difficult scenario where models must interpret chart images directly without access to raw data.
To conduct this evaluation, the authors compiled a comprehensive benchmark of 44,096 charts and natural language descriptions drawn from two major sources: Statista (34,811 charts with structured tables) and Pew Research (9,285 charts without raw tables). They established baseline architectures across three distinct technical categories: pure image captioning models, structured data-to-text models, and combined vision-text pipelines that extract visual data using optical character recognition (OCR) and deep learning. The models were evaluated using both standard automatic language metrics and human evaluations on factual correctness, coherence, and fluency.
The findings show that pretrained transformer language models—specifically T5 and BART—substantially outperform standard image captioning and non-pretrained models. When structured data tables are available, these models achieve strong fluency and content selection scores, reaching a BLEU quality score of 37.01. However, when models must rely solely on chart images through OCR pipelines, performance drops significantly, resulting in a BLEU score of 10.49 on complex charts. Furthermore, both automatic and human evaluations revealed that current models frequently suffer from hallucinations and factual errors, particularly struggling to accurately interpret visual trends, complex patterns, and relationships between numbers and visual marks.
These results demonstrate that while current language models can generate fluent summaries, automated chart-to-text systems are not yet sufficiently reliable for unmonitored decision-making. In high-stakes environments, publishing unedited model outputs introduces significant compliance and operational risks by potentially spreading incorrect facts or misleading trend analyses. The findings also emphasize that pretraining models on general table data yields only marginal improvements, underscoring that chart interpretation requires fundamentally distinct visual and logical reasoning capabilities.
Organizations and researchers pursuing automated visualization reporting should avoid deploying pure image captioning architectures and focus instead on hybrid vision-language systems with structured intermediate extraction. Before adopting these systems in production, organizations should establish robust human-in-the-loop verification processes to catch factual distortions. Future technical development should prioritize improving chart-specific data extraction, developing graph representations to map visual marks to data points, and expanding benchmarks to cover more visual styles and chart formats.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Reading this paper provides essential foundational background on integrating OCR systems and copy mechanisms into vision-language architectures to read and reason about visual text.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). This work establishes the foundational task and methodology for document visual question answering over complex graphical layouts and tables.
- Paper: On Faithfulness and Factuality in Abstractive Summarization, Joshua Maynez et al. (2020). This study introduces foundational definitions and analysis of factual errors and hallucinations in neural text generation that motivate Chart-to-Text's factuality evaluation.
- Paper: Text Summarization with Pretrained Encoders, Yang Liu et al. (2019). This paper establishes the core techniques for adapting pretrained transformer language models to abstractive text summarization.
- Paper: Don’t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization, Shashi Narayan et al. (2018). This foundational work formalizes abstractive summarization benchmarks and evaluation protocols that underpin modern data-to-text generation tasks.
- Paper: Show, Attend and Tell: Neural Image Caption Generation with Visual Attention, Kelvin Xu et al. (2015). This seminal work establishes the foundational visual attention framework used by baseline image captioning architectures evaluated in the benchmark.
- Paper: MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering, Fangyu Liu et al. (2023). MatCha directly extends and builds upon the Chart-to-Text benchmark by designing specialized pretraining objectives for chart derendering and numerical reasoning to improve chart summarization and QA.
- Paper: ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning, Ahmed Masry et al. (2022). ChartQA extends visual chart reasoning to open-vocabulary question answering over complex real-world charts collected from similar statistical sources.
- Paper: MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, Pan Lu et al. (2023). MathVista expands multimodal numerical and visual reasoning evaluations to a broader range of foundation models across plots, charts, and mathematical figures.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Visual CoT advances multimodal reasoning over dense text and charts by introducing fine-grained visual localization and chain-of-thought methods.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). This work deepens the analysis of tabular data understanding in large language models by evaluating their structural robustness and reasoning pathways.
- Paper: TableBench: A Comprehensive and Complex Benchmark for Table Question Answering, Xianjie Wu et al. (2025). TableBench broadens tabular analysis evaluation across fact checking, numerical reasoning, and chart visualization capabilities in modern language models.
- Paper: Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning, Pan Lu et al. (2023). This paper advances semi-structured table and math reasoning through dynamic prompt learning using policy gradients.
- Paper: Hallucination Augmented Contrastive Learning for Multimodal Large Language Model, Chaoya Jiang et al. (2024). HACL addresses the visual hallucination issues identified in chart-to-text generation by introducing contrastive learning on counterfactual image-text pairs.
- Paper: MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities, Weihao Yu et al. (2023). MM-Vet introduces an integrated evaluation benchmark testing large multimodal models on combinations of OCR, arithmetic, and visual generation capabilities.
- Paper: Repairing the Cracked Foundation: A Survey of Obstacles in Evaluation Practices for Generated Text, Sebastian Gehrmann et al. (2023). This survey analyzes systemic obstacles in NLG evaluation practices, offering rigorous guidelines to overcome the metric and human-evaluation limitations documented in generation benchmarks.
