SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables

Xinyuan LuLiangming PanQian LiuPreslav NakovMin-Yen Kan

article2023EMNLP58 citations

Introduces SCITAB, a benchmark of 1.2K expert-verified scientific claims derived from real research papers that tests whether language models can perform compositional and numerical reasoning directly over scientific tables.

Listen

Automated scientific fact-checking is vital for combating misinformation, preserving scientific integrity, and managing the overwhelming volume of research findings. However, existing evaluation benchmarks suffer from significant flaws: they rely heavily on crowd-sourced claims that oversimplify real scientific discourse and evaluate claims almost exclusively against unstructured text, such as paper abstracts. In real-world research, empirical claims are deeply tied to quantitative experimental data structured in tables, creating a major gap between current artificial intelligence capabilities and actual verification needs.

To address this gap, the article introduces SCITAB, a diagnostic evaluation dataset designed to test compositional reasoning and claim verification against scientific tables. The primary objective is to evaluate how effectively state-of-the-art computational models can verify complex, authentic scientific claims using structured tabular evidence.

The researchers constructed SCITAB through a human-in-the-loop collaborative process. They extracted 872 authentic claims and corresponding tables from computer science papers on arXiv. Large language models then generated candidate counter-claims and unverifiable claims, which were rigorously vetted by domain-trained human annotators. The resulting dataset contains 1,225 expert-verified claims classified as supported, refuted, or lacking sufficient information (Not Enough Info). The authors evaluated multiple model classes—including table-specific pre-trained models, open-source language models, and advanced proprietary models like GPT-4—across zero-shot and few-shot in-context learning environments.

The findings reveal that scientific table verification is exceptionally challenging for current artificial intelligence systems. First, while human annotators achieved strong performance (macro-F1 scores of 92.4% on two-class and 84.7% on three-class tasks), open-source models performed barely above random chance, peaking at only 38.1% on three-class verification. Second, GPT-4 significantly outperformed other systems, attaining a 64.8% three-class F1 score, yet it still fell nearly 20 percentage points short of human capability. Third, standard prompting strategies, including Chain-of-Thought and program-aided generation, failed to deliver meaningful gains due to frequent table grounding errors (accounting for 50% of program failures) and difficulties interpreting ambiguous academic phrasing (22% of failures). Finally, models struggled heavily with unverifiable claims, with smaller models underconfidently defaulting to 'Not Enough Info' and GPT-4 overconfidently forcing claims into supported or refuted categories.

These results indicate that current language models cannot reliably interpret structured quantitative data in specialized research domains. Relying on current artificial intelligence for automated scientific review or technical due diligence presents substantial accuracy and compliance risks. Standard techniques designed for general tabular data do not transfer well to complex scientific tables, which frequently demand multi-step arithmetic, caption context, and domain-specific knowledge.

Organizations developing or deploying automated verification systems should exercise caution and avoid fully autonomous pipelines for technical literature. Future technical initiatives must prioritize improving table grounding, integrating external domain knowledge, and refining models to handle nuanced, ambiguous statements. Additional research should also expand benchmarks beyond computer science to evaluate broader scientific domains, multimodal data, and combined text-table evidence.

Cover for SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables

Abstract

Current scientific fact-checking benchmarks exhibit several shortcomings, such as biases arising from crowd-sourced claims and an over-reliance on text-based evidence. We present SCITAB, a challenging evaluation dataset consisting of 1.2K expert-verified scientific claims that 1) originate from authentic scientific publications and 2) require compositional reasoning for verification. The claims are paired with evidence-containing scientific tables annotated with labels. Through extensive evaluations, we demonstrate that SCITAB poses a significant challenge to state-of-the-art models, including table-based pretraining models and large language models. All models except GPT-4 achieved performance barely above random guessing. Popular prompting techniques, such as Chain-of-Thought, do not achieve much performance gains on SCITAB. Our analysis uncovers several unique challenges posed by SCITAB, including table grounding, claim ambiguity, and compositional reasoning. Our codes and data are publicly available at https://github.com/XinyuanLu00/SciTab.

Table of Contents

  • 1 Introduction
  • 2 The SCITAB Dataset
  • 2.1 Data Preparation
  • 2.2 Automatic Claim Generation
  • 2.3 Manual Claim Verification
  • 3 Data Analysis
  • 3.1 Reasoning Analysis
  • 3.2 Refuted and NEI Claims Analysis
  • 4 Experiment
  • 4.1 Models
  • 4.2 Main Results
  • 4.3 Error Analysis
  • 5 Related Work
  • 6 Conclusion and Future Work
  • Ethics Statement
  • Limitations
  • Acknowledgements
  • References
  • A Claim Extraction Procedure
  • A.1 Claim Definition
  • A.2 Claim Extraction Interface
  • B Manual Claim Verification Procedure
  • B.1 Annotator Training Process
  • B.2 NEI Claim Verification Interface
  • B.3 Refuted Claim Verification Interface
  • B.4 Annotation Post-Survey
  • C Analysis of Refuted Reasons in the Sci-Fact dataset
  • D Discussions on Human-Machine Collaboration
  • E Case Study for Refuted Claims
  • F Error Cases for InstructGPT
  • G Error Cases for Program-of-Thoughts
  • H Prompts
  • H.1 Zero-shot Prompts
  • H.2 Few-shot Prompts
  • H.3 Chain-of-Thought Prompts
  • H.4 Program-of-Thoughts Prompts

Knowls

  1. Knowl 1 — SCITAB Benchmark for Scientific Table Fact-Checking

    data/table

    SCITAB is an expert-annotated evaluation dataset designed to benchmark compositional reasoning and claim verification over scientific tables crawled from arXiv.org (derived from the SciGen dataset). It comprises 1,225 scientific claims paired with evidence tables and captions across three veracity classes: Supported, Refuted, and Not Enough Info (NEI).

    Statistics TabFact FEVEROUS SEM-TAB-FACTS SCITAB
    Domain Wiki Tables Wiki Tables Scientific Articles Scientific Articles
    Annotator Amazon Mechanical Turk Amazon Mechanical Turk Amazon Mechanical Turk CS Graduate Experts
    Max. Reasoning Hops 7 2 1 11
    Supported Claims (%) 54% 56% 58% 37%
    Refuted Claims (%) 46% 39% 38% 34%
    NEI Claims (%) — 5% 4% 29%
    Total # of Claims 117,854 87,026 5,715 1,225
    Avg. Claims per Table 7.11 0.07 5.27 6.16

    Compared to prior table verification datasets, SCITAB uses expert-verified real-world scientific findings rather than crowdsourced claims, features significantly deeper reasoning paths (up to 11 hops), and maintains a substantially higher proportion of unverifiable (NEI) claims.

  2. Knowl 2 — Scientific Table Fact-Checking Task Formulation and Evaluation Protocol

    definition

    Scientific table-based fact-checking is formulated as follows: a scientific table TT consists of a caption/title PP and a structured matrix of cell entries ({Ti,j∣i≤RT,j≤CT})(\{T_{i,j} \mid i \le R_T, j \le C_T\}) with RTR_T rows and CTC_T columns, where Ti,jT_{i,j} denotes the text or numeric token sequence in the (i,j)(i,j)-th cell. Given a claim CC stating a finding with respect to TT, a fact-checking model FF maps the pair (T,C)(T, C) to a predicted veracity label Y∈{Supported,Refuted,Not Enough Info}Y \in \{\text{Supported}, \text{Refuted}, \text{Not Enough Info}\}.

    The evaluation protocol assesses models under two main regimes evaluated with Macro-F1F_1:

    • Zero-shot: The model receives the linearized table T~\tilde{T}, caption PP, claim CC, and a multi-choice prompt QQ without in-domain demonstrations.
    • In-context (few-shot): The model is provided 3 hold-out demonstration examples sampled from a dedicated hold-out set of 5 tables and 25 claims, leaving the remaining data as unseen test instances.

    Evaluation is reported across a 2-class setting (Supported vs. Refuted, excluding NEI) and a 3-class setting (Supported vs. Refuted vs. NEI).

  3. Knowl 3 — Human–Model Collaborative Pipeline for Scientific Claim Dataset Construction

    model/method

    SCITAB is constructed using a three-stage human-in-the-loop and model-assisted generation workflow:

    1. Data Preparation and Extraction: From 1,301 table descriptions in the computer science SciGen corpus, check-worthy candidate claims are filtered based on academic writing standards (highlighting key data and commenting on key data). A domain expert filtered 872 authentic true claims verifiable from tables.
    2. Automatic Claim Generation via LLM:
      • Refuted Claims: InstructGPT (text-davinci-003) is prompted with 5 in-context examples and temperature 0.70.7 to apply minimal edits to true claims to flip their meaning. Minimal editing prevents verification models from exploiting superficial lexical distribution shifts.
      • Not Enough Info (NEI) Claims: InstructGPT is prompted to generate 5 free-formed relevant scientific claims directly from the table. Ungrounded and partially ungrounded generations are retained as candidate NEI claims.
    3. Expert Manual Claim Verification: 12 trained computer science university students manually review all pairs (c,c′)(c, c') and candidate NEI claims to verify that supported claims require no external paper text, refuted claims are genuinely contradicted by the table, and NEI claims cannot be resolved solely from the table. Quality control via double annotation achieved substantial Cohen's Kappa agreement (κ=0.630\kappa = 0.630 for refuted claims; κ=0.719\kappa = 0.719 for NEI claims).
  4. Knowl 4 — Zero-Shot and In-Context Benchmark Performance on SCITAB

    data/table

    Macro-F1F_1 performance across table-specialized, text-pretrained, open-source, and closed-source language models evaluated on SCITAB under zero-shot and 3-shot in-context learning conditions:

    Models # of Para. Zero-shot In-Context
    2-class 3-class 2-class 3-class
    I. Table-based LLMs
    TAPAS-large (TabFact fine-tuned) 340M 50.30 — — —
    TAPEX-large (TabFact fine-tuned) 400M 56.06 — — —
    TAPEX-Zero-large 780M 48.28 29.72 42.44 23.47
    TAPEX-Zero-XL 3B 49.77 34.30 42.12 25.62
    II. Encoder–Decoder LLMs
    Flan-T5-base 250M 47.38 26.56 44.82 24.09
    Flan-T5-large 780M 51.58 32.55 49.62 27.30
    Flan-T5-XL 3B 52.41 38.05 48.05 29.21
    Flan-T5-XXL 11B 59.60 34.91 60.48 34.04
    III. Open source LLMs
    Alpaca-7B 7B 37.22 27.59 40.46 28.95
    Vicuna-7B 7B 63.62 32.47 50.35 34.26
    Vicuna-13B 13B 41.82 29.63 55.11 35.16
    LLaMA-7B 7B 49.05 32.26 45.24 27.17
    LLaMA-13B 13B 53.97 37.18 44.39 32.66
    IV. Closed source LLMs
    InstructGPT 175B 68.44 41.41 68.10 41.58
    InstructGPT + Chain-of-Thought 175B — — 68.46 42.60
    Program-of-Thoughts (PoT) 175B — — 63.79 —
    GPT-4 — 78.22 64.80 77.98 63.21
    GPT-4 + Chain-of-Thought — — — 76.85 62.77
    Human Expert Upper Bound — — — 92.40 84.73

    Key empirical findings include:

    1. All open-source models perform near random guessing (random baseline is 50.0 for 2-class and 33.3 for 3-class).
    2. Table-specialized pre-trained models (TAPAS, TAPEX) underperform pure-text instruction-tuned models (e.g., Flan-T5), driven by differences between Wikipedia tables and complex scientific tables with multi-level row/column headers.
    3. Prompting methods like Chain-of-Thought (CoT) and Program-of-Thoughts (PoT) fail to yield significant gains over standard zero-shot/in-context prompting.
    4. GPT-4 leads all models but still trails human performance by 14.42 points in 2-class and 19.93 points in 3-class Macro-F1F_1.
  5. Knowl 5 — Atomic Reasoning Types and Reasoning Depth Distribution in SCITAB

    data/table

    Analysis of 476 atomic reasoning steps identified across 100 manually annotated SCITAB instances reveals 14 distinct reasoning operations:

    Atomic Reasoning Function Operational Definition Proportion (%)
    Simple lookup Retrieve the value for a specific cell 20.6
    Comparison Compare two numbers 19.5
    Closed-domain knowledge Extract information from table caption or article context 12.1
    Open-domain knowledge Integrate expert domain knowledge not explicit in the table 5.3
    Commonsense knowledge Apply general world knowledge required for verification 5.3
    Subtract Compute the difference between two numbers 5.3
    Divide Compute the quotient of two numbers 5.3
    Rank Determine the relative rank order across a set of numbers 5.3
    Different / Same Check identity or difference between two values 5.3
    Add Compute the sum of two numbers 4.0
    Max / Min Identify the extremum value in a set 3.1
    Col / Rowname Resolve column or row identifiers from table headers 3.1
    Trend same / different Compare trend directions across columns or rows 2.9
    Set check Verify whether a value belongs to a specific set 2.9

    The distribution of required reasoning steps (depth) has a mean of 4.76 steps and a maximum of 11 steps per claim. Shallow claims (1–2 reasoning steps) account for only 14% of the dataset, while 86% of claims require 3 or more compositional reasoning steps (with 3 steps at 15%, 4 steps at 18%, 5 steps at 20%, 6 steps at 15%, 7 steps at 7%, 8 steps at 5%, 9 steps at 3%, 10 steps at 2%, and 11 steps at 1%).

  6. Knowl 6 — Taxonomy and Distribution of Refuted and NEI Causes in SCITAB

    data/table

    Manual failure mode analysis of 60 refuted claims and 60 Not Enough Info (NEI) claims in SCITAB:

    Label Class Error / Unverifiability Reason Proportion (%)
    Refuted The calculation result is wrong 41.7
    The approximation word is wrong (e.g., incorrect degree of difference) 33.3
    The claim is partially right (contains a half-truth across conditions) 10.0
    The values in the claim do not match the table cells 8.3
    The operation type is wrong (e.g., claimed superior instead of inferior) 6.7
    Not Enough Info (NEI) The claim lacks sufficient matching evidence in the table 33.3
    The claim lacks required open-domain background knowledge 25.0
    The claim lacks closed-domain context from paper text 15.0
    The claim refers to another table in the paper 11.7
    The claim contains ambiguous/vague pronouns without referents 8.3
    The claim omits specific necessary qualifier information 6.7

    In contrast to benchmarks like Sci-Fact—where 85% of refuted claims are generated via simple token negation (adding 'not')—SCITAB refuted claims involve nuanced numerical calculation errors, invalid approximation modifiers, and partial truths.

  7. Knowl 7 — Error Taxonomy in Program-of-Thoughts (PoT) Table Verification

    data/table

    Evaluation of 50 randomly selected claims where Program-of-Thoughts (PoT)—which translates reasoning into executable Python code—failed to predict the correct veracity label:

    Error Category Description Proportion (%)
    I. Grounding errors Incorrect entity linking or cell-variable association from table 50
    II. Ambiguity errors Inability to represent fuzzy/approximate qualifiers (e.g., `comparable`) 22
    III. Calculation errors Float precision / arithmetic rounding mismatches in Python 20
    IV. Program errors Code bugs, incorrect arguments, missing variables, or invalid logic 8

    Programmatic verification underperforms direct prompting because symbolic programs struggle to ground unstructured table cells accurately (50%) and cannot map ambiguous natural language qualifiers into hard boolean logic (22%).

  8. Knowl 8 — Confidence Asymmetry Between InstructGPT and GPT-4 on SCITAB

    empirical result

    Confusion matrix analysis of zero-shot 3-class fact verification reveals contrasting failure modes between InstructGPT and GPT-4:

    1. InstructGPT Underconfidence: InstructGPT defaults to a conservative bias, frequently classifying verifiable Supported (26.8% of all predictions) and Refuted (23.6% of all predictions) claims as Not Enough Info (NEI). True Supported accuracy is only 9.1% and true Refuted accuracy is only 5.4%.
    2. GPT-4 Overconfidence: GPT-4 exhibits the opposite pattern, over-predicting certainty and incorrectly categorizing genuinely unverifiable NEI claims as Supported (10.3%) or Refuted (8.5%), while achieving high recall on true Supported (32.1%) and Refuted (25.2%) instances.
    3. Lexical Negation Blindness: Both LLMs frequently misclassify Refuted claims containing negation as Supported due to superficial entity and term matching, and misclassify Supported claims requiring arithmetic comparisons as Refuted or NEI.
  9. Knowl 9 — Limitations of the SCITAB Dataset

    limitation

    The authors state four primary limitations of SCITAB:

    1. Morphological and Language Scope: The methodology and claim annotations are restricted exclusively to English scientific literature.
    2. Evidence Modality Restriction: The benchmark exclusively isolates table-based evidence and captions, omitting multi-modal integration with full scientific article text and figures.
    3. Domain and Reasoning Bias: Claims are sourced from the computer science section of arXiv (via SciGen), biasing the dataset towards numerical and algorithmic benchmarking metrics rather than non-computational or qualitative scientific reasoning.
    4. Graph Annotation Completeness: Full reasoning graphs and step counts are analyzed on a diagnostic sample of 100 claims rather than fully annotated across all 1,225 claims in the benchmark.

Coverage note — None was omitted; all key contributions—dataset statistics, task formulation, construction pipeline, reasoning analysis, baseline evaluations, error analyses, and limitations—are fully represented.

References

  1. 1.Mubashara Akhtar, Oana Cocarascu, and Elena Simperl. 2022. Pubhealthtab: A public health table-based dataset for evidence-based fact checking. In Findings of the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), pages 1–16.
  2. 2.Rami Aly, Zhijiang Guo, Michael Sejr Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal. 2021. FEVEROUS: fact extraction and verification over unstructured and structured information. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks.
  3. 3.Emily M. Bender and Batya Friedman. 2018. Data statements for natural language processing: Toward mitigating system bias and enabling better science. Transactions of the Association for Computational Linguistics (TACL), 6:587–604.
  4. 4.Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2022. Program of thoughts prompting: Disentangling computation from reasoning for numerical reasoning tasks. CoRR, abs/2211.12588.
  5. 5.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. Tabfact: A large-scale dataset for table-based fact verification. In Proceedings of the 8th International Conference on Learning Representations (ICLR).
  6. 6.Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.
  7. 7.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. CoRR, abs/2210.11416.
  8. 8.Jacob Cohen. 1960. A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20:37 – 46.
  9. 9.Thomas Diggelmann, Jordan L. Boyd-Graber, Jannis Bulian, Massimiliano Ciaramita, and Markus Leippold. 2020. CLIMATE-FEVER: A dataset for verification of real-world climate claims. CoRR, abs/2012.00614.
  10. 10.Max Glockner, Ieva Staliūnaitė, James Thorne, Gisela Vallejo, Andreas Vlachos, and Iryna Gurevych. 2023. Ambifc: Fact-checking ambiguous claims with evidence. CoRR, abs/2104.00640.
  11. 11.Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, and Xiaoyong Du. 2022. PASTA: table-operations aware fact verification via sentence-table cloze pre-training. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4971–4983.
  12. 12.Zhijiang Guo, Michael Sejr Schlichtkrull, and Andreas Vlachos. 2022. A survey on automated fact-checking. Transactions of the Association for Computational Linguistics (TACL), 10:178–206.
  13. 13.Vivek Gupta, Maitrey Mehta, Pegah Nokhiz, and Vivek Srikumar. 2020. INFOTABS: inference on tables as semi-structured data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2309–2324.
  14. 14.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. 2020. Tapas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 4320–4333.
  15. 15.Yichen Jiang, Shikha Bordia, Zheng Zhong, Charles Dognin, Maneesh Kumar Singh, and Mohit Bansal. 2020. Hover: A dataset for many-hop fact extraction and claim verification. In Findings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), volume EMNLP 2020, pages 3441–3460.
  16. 16.W.Y. Lee, L. Ho, and M.E.T. Ng. 2009. Research Writing: A Workbook for Graduate Students. Prentice Hall.
  17. 17.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 7871–7880.
  18. 18.Alisa Liu, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2022a. WANLI: worker and AI collaboration for natural language inference dataset creation. In Findings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6826–6847.
  19. 19.Alisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr, Peter West, Alexander Koller, Swabha Swayamdipta, Noah A. Smith, and Yejin Choi. 2023a. We’re afraid language models aren’t modeling ambiguity. CoRR, abs/2304.14399.
  20. 20.Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022b. TAPEX: table pre-training via learning a neural SQL executor. In Proceedings of the 10th International Conference on Learning Representations (ICLR).
  21. 21.Qian Liu, Fan Zhou, Zhengbao Jiang, Longxu Dou, and Min Lin. 2023b. From zero to hero: Examining the power of symbolic tasks in instruction tuning. CoRR, abs/2304.07995.
  22. 22.Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023. Chameleon: Plug-and-play compositional reasoning with large language models. CoRR, abs/2304.09842.
  23. 23.Isabelle Mohr, Amelie Wührl, and Roman Klinger. 2022. Covert: A corpus of fact-checked biomedical COVID-19 tweets. In Proceedings of the 13th Language Resources and Evaluation Conference (LREC), pages 244–257.
  24. 24.Nafise Sadat Moosavi, Andreas Rücklé, Dan Roth, and Iryna Gurevych. 2021. Scigen: a dataset for reasoning-aware text generation from scientific tables. In Proceedings of the Neural Information Processing Systems (NeurIPS) Track on Datasets and Benchmarks.
  25. 25.OpenAI. 2023. GPT-4 technical report. CoRR, abs/2303.08774.
  26. 26.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS).
  27. 27.Liangming Pan, Xiaobao Wu, Xinyuan Lu, Anh Tuan Luu, William Yang Wang, Min-Yen Kan, and Preslav Nakov. 2023. Fact-checking complex claims with program-guided reasoning. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), pages 6981–7004.
  28. 28.Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis. 2022. Measuring and narrowing the compositionality gap in language models. CoRR, abs/2210.03350.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR), 21:140:1–140:67.
  30. 30.Arkadiy Saakyan, Tuhin Chakrabarty, and Smaranda Muresan. 2021. Covid-fact: Fact extraction and verification of real-world claims on COVID-19 pandemic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), pages 2116–2129.
  31. 31.Mourad Sarrouti, Asma Ben Abacha, Yassine Mrabet, and Dina Demner-Fushman. 2021. Evidence-based fact-checking of health-related claims. In Findings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3499–3512.
  32. 32.Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language models can teach themselves to use tools. CoRR, abs/2302.04761.
  33. 33.Tal Schuster, Darsh J. Shah, Yun Jie Serene Yeo, Daniel Filizzola, Enrico Santus, and Regina Barzilay. 2019. Towards debiasing fact verification models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3417–3423.
  34. 34.Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca.
  35. 35.Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971.
  36. 36.Juraj Vladika and Florian Matthes. 2023. Scientific fact-checking: A survey of resources and approaches. In Findings of the 61st Association for Computational Linguistics (ACL), pages 6215–6230.
  37. 37.David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or fiction: Verifying scientific claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7534–7550.
  38. 38.David Wadden, Kyle Lo, Bailey Kuehl, Arman Cohan, Iz Beltagy, Lucy Lu Wang, and Hannaneh Hajishirzi. 2022. Scifact-open: Towards open-domain scientific claim verification. In Findings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4719–4734.
  39. 39.Gengyu Wang, Kate Harwood, Lawrence Chillrud, Amith Ananthram, Melanie Subbiah, and Kathleen R. McKeown. 2023. Check-covid: Fact-checking COVID-19 news claims with scientific evidence. In Findings of the 61st Association for Computational Linguistics (ACL), pages 14114–14127.
  40. 40.Nancy Xin Ru Wang, Diwakar Mahajan, Marina Danilevsky, and Sara Rosenthal. 2021. Semeval-2021 task 9: Fact verification and evidence finding for tabular data in scientific documents (SEM-TAB-FACTS). In Proceedings of the 15th International Workshop on Semantic Evaluation (SemEval@ACL/IJCNLP), pages 317–326.
  41. 41.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed H. Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. CoRR, abs/2201.11903.
  42. 42.Xiaoyu Yang, Feng Nie, Yufei Feng, Quan Liu, Zhigang Chen, and Xiaodan Zhu. 2020. Program enhanced fact verification with verbalization and graph attention network. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7810–7825.
  43. 43.Yunhu Ye, Binyuan Hui, Min Yang, Binhua Li, Fei Huang, and Yongbin Li. 2023. Large language models are versatile decomposers: Decomposing evidence and questions for table-based reasoning. In Proceedings of the 46th International ACM Conference on Research and Development in Information Retrieval (SIGIR), pages 174–184.
  44. 44.Pengcheng Yin, Zhengdong Lu, Hang Li, and Ben Kao. 2016. Neural enquirer: Learning to query tables in natural language. In Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI), pages 2308–2314.
  45. 45.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir R. Radev. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3911–3921.
  46. 46.Wanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan, Ming Zhou, Ming Gong, Linjun Shou, Daxin Jiang, Jiahai Wang, and Jian Yin. 2020. Logicalfactchecker: Leveraging logical operations for fact checking with graph module network. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), pages 6053–6065.
  47. 47.Yuxuan Zhou, Xien Liu, Kaiyin Zhou, and Ji Wu. 2022. Table-based fact verification with self-adaptive mixture of experts. In Findings of the 60th Association for Computational Linguistics (ACL), pages 139–149.

Citation

MLA
Lu, X., et al. “SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 7787–813, https://doi.org/10.18653/v1/2023.emnlp-main.483.
APA
Lu, X., Pan, L., Liu, Q., Nakov, P., & Kan, M.-Y. (2023). SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7787–7813. https://doi.org/10.18653/v1/2023.emnlp-main.483
Chicago
Lu, X., L. Pan, Q. Liu, P. Nakov, and M.-Y. Kan. 2023. “SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 7787–7813. https://doi.org/10.18653/v1/2023.emnlp-main.483.
Harvard
Lu, X. et al. (2023) “SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7787–7813. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.483.
Vancouver
1. Lu X, Pan L, Liu Q, Nakov P, Kan M-Y (2023) SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7787–7813

BibTeX

@inproceedings{lu-etal-2023-scitab,
    title = "{SCITAB}: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables",
    author = "Lu, Xinyuan  and
      Pan, Liangming  and
      Liu, Qian  and
      Nakov, Preslav  and
      Kan, Min-Yen",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.483/",
    doi = "10.18653/v1/2023.emnlp-main.483",
    pages = "7787--7813"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/