TableBench: A Comprehensive and Complex Benchmark for Table Question Answering
Xianjie WuJian Yang 0030Linzheng ChaiGe Zhang 0009Jiaheng LiuXeron DuDi LiangDaixin ShuXianfu ChengTianzhen Sun
Introduces TableBench, a challenging table question-answering benchmark spanning 18 domains across fact checking, numerical reasoning, data analysis, and visualization, revealing that even leading large language models still lag significantly behind human performance in complex real-world tabular tasks.
Organizations increasingly rely on artificial intelligence language models to automate the interpretation, analysis, and visualization of tabular data. However, existing academic benchmarks predominantly focus on straightforward lookup tasks and simple arithmetic, failing to mirror the complex, multi-step analytical reasoning demanded in professional environments. Consequently, enterprise leaders face uncertainty regarding the actual readiness, reliability, and deployment risks of applying these models to critical operational and financial workflows.
The article aims to evaluate the practical tabular reasoning capabilities of modern language models and bridge the gap between academic evaluations and real-world requirements. To achieve this, the authors introduce TableBench, a specialized benchmark covering 18 analytical capabilities across four core categories—fact checking, numerical reasoning, data analysis, and chart visualization. The evaluation tested 34 leading open-source and commercial language models against human benchmarks using 886 manually curated questions across 20 industries, alongside an accompanying instruction dataset containing nearly 20,000 training examples.
The findings reveal a substantial gap between automated models and human capabilities on complex tabular tasks. Humans achieved an overall baseline score of 85.91%, whereas the top-performing commercial system, GPT-4-Turbo, achieved an overall score of only 51.32%. While most models performed reliably on basic fact checking (with top scores reaching around 75%), performance deteriorated significantly on advanced numerical calculations, statistical data analysis, and chart generation. In visualization tasks specifically, smaller models failed almost entirely, and even top-tier commercial models generated successfully executable code only 78.67% of the time in single attempts. However, fine-tuning open-source models using domain-targeted training datasets showed high data efficiency; fine-tuned 7-billion to 8-billion parameter models achieved performance on par with GPT-3.5 at a fraction of the operating cost using fewer than 4,000 training samples.
These results indicate that enterprises should exercise caution before deploying language models for fully autonomous reporting, financial analysis, or automated data visualization. The significant error rate in code generation and intermediate calculations presents direct operational, compliance, and decision-making risks if left unmonitored. While commercial proprietary models maintain a performance lead on complex reasoning, targeted fine-tuning of open-source architectures provides a viable, cost-effective alternative for routine tabular workloads.
Organizations should implement human-in-the-loop validation for high-stakes tabular queries and incorporate automated execution and error-correction loops to improve code-based reasoning reliability. Further work should focus on integrating multi-turn code correction mechanisms and extending benchmarks to evaluate complex table structures and image-based tabular documents. The benchmark's findings offer high confidence for textual and semi-structured tabular text, though readers should note the evaluation excluded visual image formats and unusually complex structural table layouts.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). Provides fundamental insights into how large language models process, reason over, and struggle with tabular structures and representations.
- Paper: Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning, Pan Lu et al. (2023). Introduces foundational benchmarks and prompting techniques for multi-step arithmetic and mathematical reasoning over semi-structured tables.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). Establishes core methods for prompting and enabling large language models to reason over structured data like tables and databases.
- Paper: ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning, Ahmed Masry et al. (2022). Pioneers the benchmark evaluation of combining visual data interpretation with logical and mathematical question answering.
- Paper: Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task, Tao Yu et al. (2018). Introduces the standard benchmark for semantic parsing and text-to-code execution over complex relational tabular databases.
- Paper: Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation, Jiawei Liu et al. (2023). Analyzes the functional correctness and execution reliability of model-generated code, a central mechanism TableBench uses for data analysis and chart generation.
- Paper: Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs, Wei Zhou et al. (2026). Surveys the broader downstream application of large language models in preparing, cleaning, and managing messy real-world tabular data.
