SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables
Xinyuan LuLiangming PanQian LiuPreslav NakovMin-Yen Kan
Introduces SCITAB, a benchmark of 1.2K expert-verified scientific claims derived from real research papers that tests whether language models can perform compositional and numerical reasoning directly over scientific tables.
Automated scientific fact-checking is vital for combating misinformation, preserving scientific integrity, and managing the overwhelming volume of research findings. However, existing evaluation benchmarks suffer from significant flaws: they rely heavily on crowd-sourced claims that oversimplify real scientific discourse and evaluate claims almost exclusively against unstructured text, such as paper abstracts. In real-world research, empirical claims are deeply tied to quantitative experimental data structured in tables, creating a major gap between current artificial intelligence capabilities and actual verification needs.
To address this gap, the article introduces SCITAB, a diagnostic evaluation dataset designed to test compositional reasoning and claim verification against scientific tables. The primary objective is to evaluate how effectively state-of-the-art computational models can verify complex, authentic scientific claims using structured tabular evidence.
The researchers constructed SCITAB through a human-in-the-loop collaborative process. They extracted 872 authentic claims and corresponding tables from computer science papers on arXiv. Large language models then generated candidate counter-claims and unverifiable claims, which were rigorously vetted by domain-trained human annotators. The resulting dataset contains 1,225 expert-verified claims classified as supported, refuted, or lacking sufficient information (Not Enough Info). The authors evaluated multiple model classes—including table-specific pre-trained models, open-source language models, and advanced proprietary models like GPT-4—across zero-shot and few-shot in-context learning environments.
The findings reveal that scientific table verification is exceptionally challenging for current artificial intelligence systems. First, while human annotators achieved strong performance (macro-F1 scores of 92.4% on two-class and 84.7% on three-class tasks), open-source models performed barely above random chance, peaking at only 38.1% on three-class verification. Second, GPT-4 significantly outperformed other systems, attaining a 64.8% three-class F1 score, yet it still fell nearly 20 percentage points short of human capability. Third, standard prompting strategies, including Chain-of-Thought and program-aided generation, failed to deliver meaningful gains due to frequent table grounding errors (accounting for 50% of program failures) and difficulties interpreting ambiguous academic phrasing (22% of failures). Finally, models struggled heavily with unverifiable claims, with smaller models underconfidently defaulting to 'Not Enough Info' and GPT-4 overconfidently forcing claims into supported or refuted categories.
These results indicate that current language models cannot reliably interpret structured quantitative data in specialized research domains. Relying on current artificial intelligence for automated scientific review or technical due diligence presents substantial accuracy and compliance risks. Standard techniques designed for general tabular data do not transfer well to complex scientific tables, which frequently demand multi-step arithmetic, caption context, and domain-specific knowledge.
Organizations developing or deploying automated verification systems should exercise caution and avoid fully autonomous pipelines for technical literature. Future technical initiatives must prioritize improving table grounding, integrating external domain knowledge, and refining models to handle nuanced, ambiguous statements. Additional research should also expand benchmarks beyond computer science to evaluate broader scientific domains, multimodal data, and combined text-table evidence.
- Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). This paper establishes the foundational claim verification framework of supported, refuted, and not enough info classifications that SCITAB directly adapts for scientific tabular data.
- Paper: Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning, Pan Lu et al. (2023). This work introduces multi-step mathematical reasoning benchmarks over tabular structures, serving as key context for evaluating numerical reasoning limits in table verification.
- Paper: SciBERT: A Pretrained Language Model for Scientific Text, Iz Beltagy et al. (2019). This work develops domain-specific pre-training for scientific text, providing foundational representation models essential for scientific natural language understanding.
- Paper: Annotation Artifacts in Natural Language Inference Data, Suchin Gururangan et al. (2018). This paper demonstrates how natural language inference benchmarks suffer from annotation artifacts, motivating SCITAB's use of authentic scientific claims over crowdsourced alternatives.
- Paper: Can We Automate Scientific Reviewing?, Weizhe Yuan et al. (2022). This work examines automated scientific peer review and exposes reasoning failures in NLP models, providing critical motivation for evaluating scientific claim verification.
- Paper: TableBench: A Comprehensive and Complex Benchmark for Table Question Answering, Xianjie Wu et al. (2025). This benchmark expands the evaluation of complex tabular reasoning across commercial and open-source models into broader real-world analytical, numerical, and fact-checking settings.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). This study deepens the investigation into table grounding failures by analyzing language model robustness under structural table perturbations and multi-pathway reasoning.
- Paper: Factuality of Large Language Models: A Survey, Yuxia Wang et al. (2024). This survey provides a comprehensive synthesis of factuality benchmarks, failure modes, and mitigation strategies across large language models, contextualizing tabular fact-checking challenges.
- Paper: A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery, Yu Zhang et al. (2024). This survey synthesizes scientific language models and cross-modal reasoning over scientific tables and multimodal data, continuing the agenda proposed in SCITAB.
- Paper: Towards Automating Scientific Review with Google's Paper Assistant Tool, Rajesh Jayaram et al. (2026). This work applies automated verification principles to full scientific manuscripts in conference review pipelines, extending SCITAB's findings to operational peer-review assistance.
