RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations
Yilun ZhaoChen ZhaoLinyong NanZhenting QiWenlin ZhangXiangru TangBoyu MiDragomir Radev
Presents RobuT, a human-annotated diagnostic benchmark spanning 138,149 examples, demonstrating that state-of-the-art table question answering models fail under realistic perturbations and providing an LLM-based adversarial training framework to remedy these vulnerabilities.
Modern natural language processing systems increasingly rely on automated models to query and extract insights from structured tabular data. While these table question answering systems achieve strong results on standard benchmarks, existing evaluations measure performance only on data formatted identically to training examples. In real-world enterprise deployments, tables and user queries vary naturally through synonymous column headers, shuffled rows, or paraphrased wording. The article addresses the critical risk that deployed systems may appear accurate under benchmark conditions but fail unpredictably when exposed to minor, realistic variations.
The main objective of the article is to systematically evaluate the robustness of state-of-the-art table question answering systems against realistic perturbations and to establish an automated, cost-effective framework to remediate identified vulnerabilities.
To conduct this evaluation, the authors created ROBUT, a comprehensive diagnostic benchmark derived from three widely used tabular datasets. ROBUT comprises 138,149 human-annotated test pairs covering ten perturbation types across table headers (synonym and abbreviation replacements), table contents (row and column shuffling, column extensions, masking, and additions), and natural language questions (word-level and sentence-level paraphrasing). The authors evaluated prominent specialized models—including TAPAS, TableFormer, TAPEX, and OmniTab—alongside large language models such as GPT-3 under few-shot settings. They subsequently developed LETA, a framework using large language model prompts to generate synthetic adversarial training data to improve the resilience of smaller, specialized models.
The findings reveal substantial performance drops across all specialized models when exposed to minor variations. First, specialized systems experienced severe accuracy drops across the benchmark, frequently losing 10 to nearly 40 percentage points on complex structural tasks such as column extension. Second, architectural modifications offered only narrow defenses; for instance, TableFormer proved resilient against row and column order changes but remained vulnerable to linguistic modifications in headers and questions. Third, few-shot large language models demonstrated significantly higher overall robustness, with GPT-3 maintaining robustness accuracy rates between 80% and 97% across most single-perturbation categories. Finally, fine-tuning specialized models using synthetic adversarial data generated by the LETA framework restored post-perturbation accuracy by up to 5.7 percentage points, markedly outperforming traditional rule-based data augmentation while costing a fraction of human annotation.
These results indicate that current specialized systems rely heavily on superficial dataset shortcuts rather than genuine tabular reasoning, posing operational and decision-making risks if deployed without safeguards. While large foundation models demonstrate superior resilience, their operational costs and latency can be prohibitive. The LETA framework demonstrates that organizations can capture the robustness benefits of large models by using them offline to generate adversarial training samples for smaller, cost-effective specialized models. However, this process incurs a standard trade-off: enhancing adversarial robustness slightly reduces accuracy on original, clean data (typically by 1 to 4 percentage points).
Organizations deploying tabular question answering tools should immediately avoid relying solely on standard validation accuracy and adopt comprehensive adversarial testing before deployment. Where low latency and hosting costs necessitate smaller models, teams should implement automated adversarial augmentation during model training. Future technical work should focus on improving synthetic generation prompts for complex table structures and extending diagnostic evaluations to adjacent tabular tasks, such as automated fact-checking and data-to-text generation.
Confidence in these findings is high regarding standard question-answering formats across open-domain web tables. However, users should note key limitations: the benchmark does not modify underlying numerical cell values to avoid altering ground-truth answers, and large language model data generation occasionally exhibits hallucinations or slight shifts in semantic meaning that require ongoing monitoring.
- Paper: TableFormer: Robust Transformer Modeling for Table-Text Encoding, Jingfeng Yang et al. (2022). Read this first to understand TableFormer, one of the table-QA architectures RobuT evaluates and whose permutation robustness provides a key point of comparison.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). This later study carries the robustness question into large language models, testing how table-structure changes affect their reasoning and proposing a normalization method to address those vulnerabilities.
