TableBench: A Comprehensive and Complex Benchmark for Table Question Answering

Xianjie WuJian Yang 0030Linzheng ChaiGe Zhang 0009Jiaheng LiuXeron DuDi LiangDaixin ShuXianfu ChengTianzhen Sun

article2025AAAI130 citations

Introduces TableBench, a challenging table question-answering benchmark spanning 18 domains across fact checking, numerical reasoning, data analysis, and visualization, revealing that even leading large language models still lag significantly behind human performance in complex real-world tabular tasks.

Listen

Organizations increasingly rely on artificial intelligence language models to automate the interpretation, analysis, and visualization of tabular data. However, existing academic benchmarks predominantly focus on straightforward lookup tasks and simple arithmetic, failing to mirror the complex, multi-step analytical reasoning demanded in professional environments. Consequently, enterprise leaders face uncertainty regarding the actual readiness, reliability, and deployment risks of applying these models to critical operational and financial workflows.

The article aims to evaluate the practical tabular reasoning capabilities of modern language models and bridge the gap between academic evaluations and real-world requirements. To achieve this, the authors introduce TableBench, a specialized benchmark covering 18 analytical capabilities across four core categories—fact checking, numerical reasoning, data analysis, and chart visualization. The evaluation tested 34 leading open-source and commercial language models against human benchmarks using 886 manually curated questions across 20 industries, alongside an accompanying instruction dataset containing nearly 20,000 training examples.

The findings reveal a substantial gap between automated models and human capabilities on complex tabular tasks. Humans achieved an overall baseline score of 85.91%, whereas the top-performing commercial system, GPT-4-Turbo, achieved an overall score of only 51.32%. While most models performed reliably on basic fact checking (with top scores reaching around 75%), performance deteriorated significantly on advanced numerical calculations, statistical data analysis, and chart generation. In visualization tasks specifically, smaller models failed almost entirely, and even top-tier commercial models generated successfully executable code only 78.67% of the time in single attempts. However, fine-tuning open-source models using domain-targeted training datasets showed high data efficiency; fine-tuned 7-billion to 8-billion parameter models achieved performance on par with GPT-3.5 at a fraction of the operating cost using fewer than 4,000 training samples.

These results indicate that enterprises should exercise caution before deploying language models for fully autonomous reporting, financial analysis, or automated data visualization. The significant error rate in code generation and intermediate calculations presents direct operational, compliance, and decision-making risks if left unmonitored. While commercial proprietary models maintain a performance lead on complex reasoning, targeted fine-tuning of open-source architectures provides a viable, cost-effective alternative for routine tabular workloads.

Organizations should implement human-in-the-loop validation for high-stakes tabular queries and incorporate automated execution and error-correction loops to improve code-based reasoning reliability. Further work should focus on integrating multi-turn code correction mechanisms and extending benchmarks to evaluate complex table structures and image-based tabular documents. The benchmark's findings offer high confidence for textual and semi-structured tabular text, though readers should note the evaluation excluded visual image formats and unusually complex structural table layouts.

arXiv: 2408.09174
Cover for TableBench: A Comprehensive and Complex Benchmark for Table Question Answering

Citation

MLA
Wu, X., et al. “TableBench: A Comprehensive and Complex Benchmark for Table Question Answering”. arXiv, 2024, http://arxiv.org/abs/2408.09174v2.
APA
Wu, X., Yang, J., Chai, L., Zhang, G., Liu, J., Du, X., Liang, D., Shu, D., Cheng, X., Sun, T., Niu, G., Li, T., & Li, Z. (2024). TableBench: A Comprehensive and Complex Benchmark for Table Question Answering. arXiv. http://arxiv.org/abs/2408.09174v2
Chicago
Wu, X., J. Yang, L. Chai, et al. 2024. “TableBench: A Comprehensive and Complex Benchmark for Table Question Answering”. arXiv. http://arxiv.org/abs/2408.09174v2.
Harvard
Wu, X. et al. (2024) “TableBench: A Comprehensive and Complex Benchmark for Table Question Answering”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2408.09174v2.
Vancouver
1. Wu X, Yang J, Chai L, et al (2024) TableBench: A Comprehensive and Complex Benchmark for Table Question Answering. arXiv

BibTeX

@article{wu2024tablebench,
  title = {TableBench: A Comprehensive and Complex Benchmark for Table Question Answering},
  author = {Wu, Xianjie and Yang, Jian and Chai, Linzheng and Zhang, Ge and Liu, Jiaheng and Du, Xinrun and Liang, Di and Shu, Daixin and Cheng, Xianfu and Sun, Tianzhen and Niu, Guanglin and Li, Tongliang and Li, Zhoujun},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2408.09174v2},
  eprint = {2408.09174}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF