PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training
Zihui GuJu FanNan TangPreslav NakovXiaoman ZhaoXiaoyong Du
Presents PASTA, a table-based fact verification framework that pre-trains language models on 1.2 million synthesized sentence-table cloze tasks covering common operations like aggregation and comparison, setting state-of-the-art results on TabFact and SEM-TAB-FACTS.
Misinformation and disinformation online pose growing operational, reputational, and decision-making risks across journalism, public policy, and corporate governance. While verifying assertions against unstructured text is challenging, many false statements can be directly validated or refuted using structured tables. However, standard language models often struggle to interpret tabular data because they lack the ability to perform basic logical and arithmetic operations, such as comparing entries or aggregating columns.
The article aims to introduce and evaluate PASTA, a novel framework designed to enhance language models with table-operation awareness for automated fact verification without relying on complex, error-prone logical program generators.
The authors constructed an automated pre-training dataset of roughly 1.2 million fill-in-the-blank questions derived from 20,000 Wikipedia tables. These questions target six core table operations: filtering, aggregation, superlatives, comparisons, ordinal rankings, and uniqueness checks. Using the DeBERTaV3 language model architecture, the system was trained to predict masked operation-specific words and values. During downstream deployment, a select-then-rank pre-processing strategy was applied to prioritize the most relevant table rows and columns within the model's limited input window. The framework was evaluated across standard benchmarks containing simple and multi-row complex statements, including scientific domain data.
The evaluation yielded three primary findings. First, PASTA established a new state-of-the-art benchmark on the primary testing dataset (TabFact), achieving 89.3% overall accuracy and outperforming the previous state-of-the-art by 4.7 percentage points on complex, multi-operation statements (85.6% versus 80.9%). Second, the model substantially closed the gap with human accuracy on a held-out test set, trailing human performance by only 1.5 percentage points (90.6% versus 92.1%). Third, the model demonstrated strong cross-domain transferability on scientific literature tables (SEM-TAB-FACTS), exceeding baseline models by 5.2 points (84.1% micro-F1).
These results indicate that pre-training language models on structured, operation-aware tasks significantly enhances their symbolic reasoning capabilities without requiring expensive manual annotations or brittle semantic parsing. Organizations deploying automated fact-checking or tabular data analysis can achieve higher accuracy and reliability with relatively modest training data sizes, lowering development timelines and computational overhead.
Decision-makers seeking to implement automated table verification should consider adopting targeted operation-aware pre-training over traditional random masking approaches, while implementing row-ranking heuristics to manage large tabular inputs. Before deploying such systems in production, practitioners should conduct pilot testing specifically focused on complex claims that combine multiple operational steps.
Confidence in these findings is high for single-table verification scenarios supported by clean data structures. However, stakeholders should exercise caution, as the system's performance declines when processing very large tables or statements that require chains of multiple distinct operations. Additionally, the current framework is constrained to verifying claims against a single table at a time and assumes the underlying reference tables are accurate and unbiased.
- Paper: Successive Prompting for Decomposing Complex Questions, Dheeru Dua et al. (2022). Its table-derived multi-step reasoning data provides a useful precursor to PASTA’s operation-focused synthetic pre-training.
- Paper: SCITAB: A Challenging Benchmark for Compositional Reasoning and Claim Verification on Scientific Tables, Xinyuan Lu et al. (2023). SCITAB carries table-based claim verification into scientific papers and probes the compositional reasoning challenges that PASTA’s scientific-table transfer results make especially relevant.
