RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations

Yilun ZhaoChen ZhaoLinyong NanZhenting QiWenlin ZhangXiangru TangBoyu MiDragomir Radev

article2023ACL63 citations

Presents RobuT, a human-annotated diagnostic benchmark spanning 138,149 examples, demonstrating that state-of-the-art table question answering models fail under realistic perturbations and providing an LLM-based adversarial training framework to remedy these vulnerabilities.

Listen

Modern natural language processing systems increasingly rely on automated models to query and extract insights from structured tabular data. While these table question answering systems achieve strong results on standard benchmarks, existing evaluations measure performance only on data formatted identically to training examples. In real-world enterprise deployments, tables and user queries vary naturally through synonymous column headers, shuffled rows, or paraphrased wording. The article addresses the critical risk that deployed systems may appear accurate under benchmark conditions but fail unpredictably when exposed to minor, realistic variations.

The main objective of the article is to systematically evaluate the robustness of state-of-the-art table question answering systems against realistic perturbations and to establish an automated, cost-effective framework to remediate identified vulnerabilities.

To conduct this evaluation, the authors created ROBUT, a comprehensive diagnostic benchmark derived from three widely used tabular datasets. ROBUT comprises 138,149 human-annotated test pairs covering ten perturbation types across table headers (synonym and abbreviation replacements), table contents (row and column shuffling, column extensions, masking, and additions), and natural language questions (word-level and sentence-level paraphrasing). The authors evaluated prominent specialized models—including TAPAS, TableFormer, TAPEX, and OmniTab—alongside large language models such as GPT-3 under few-shot settings. They subsequently developed LETA, a framework using large language model prompts to generate synthetic adversarial training data to improve the resilience of smaller, specialized models.

The findings reveal substantial performance drops across all specialized models when exposed to minor variations. First, specialized systems experienced severe accuracy drops across the benchmark, frequently losing 10 to nearly 40 percentage points on complex structural tasks such as column extension. Second, architectural modifications offered only narrow defenses; for instance, TableFormer proved resilient against row and column order changes but remained vulnerable to linguistic modifications in headers and questions. Third, few-shot large language models demonstrated significantly higher overall robustness, with GPT-3 maintaining robustness accuracy rates between 80% and 97% across most single-perturbation categories. Finally, fine-tuning specialized models using synthetic adversarial data generated by the LETA framework restored post-perturbation accuracy by up to 5.7 percentage points, markedly outperforming traditional rule-based data augmentation while costing a fraction of human annotation.

These results indicate that current specialized systems rely heavily on superficial dataset shortcuts rather than genuine tabular reasoning, posing operational and decision-making risks if deployed without safeguards. While large foundation models demonstrate superior resilience, their operational costs and latency can be prohibitive. The LETA framework demonstrates that organizations can capture the robustness benefits of large models by using them offline to generate adversarial training samples for smaller, cost-effective specialized models. However, this process incurs a standard trade-off: enhancing adversarial robustness slightly reduces accuracy on original, clean data (typically by 1 to 4 percentage points).

Organizations deploying tabular question answering tools should immediately avoid relying solely on standard validation accuracy and adopt comprehensive adversarial testing before deployment. Where low latency and hosting costs necessitate smaller models, teams should implement automated adversarial augmentation during model training. Future technical work should focus on improving synthetic generation prompts for complex table structures and extending diagnostic evaluations to adjacent tabular tasks, such as automated fact-checking and data-to-text generation.

Confidence in these findings is high regarding standard question-answering formats across open-domain web tables. However, users should note key limitations: the benchmark does not modify underlying numerical cell values to avoid altering ground-truth answers, and large language model data generation occasionally exhibits hallucinations or slight shifts in semantic meaning that require ongoing monitoring.

Cover for RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations

Abstract

Despite significant progress having been made in question answering on tabular data (Table QA), it’s unclear whether, and to what extent existing Table QA models are robust to task-specific perturbations, e.g., replacing key question entities or shuffling table columns. To systematically study the robustness of Table QA models, we propose a benchmark called RobuT, which builds upon existing Table QA datasets (WTQ, WikiSQL-Weak, and SQA) and includes human-annotated adversarial perturbations in terms of table header, table content, and question. Our results indicate that both state-of-the-art Table QA models and large language models (e.g., GPT-3) with few-shot learning falter in these adversarial sets. We propose to address this problem by using large language models to generate adversarial examples to enhance training, which significantly improves the robustness of Table QA models. Our data and code is publicly available at https://github.com/yilunzhao/RobuT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 ROBUT Benchmark
  • 3.1 Table Header Perturbation
  • 3.2 Table Content Perturbation
  • 3.3 NLQ Perturbation
  • 3.4 Mix Perturbation
  • 4 Diagnostic Experiments
  • 4.1 Experimental Setup
  • 4.2 Diagnostic Results
  • 5 LETA Framework
  • 5.1 Table Header Augmentation
  • 5.2 Table Content Augmentation
  • 5.3 NLQ Augmentation
  • 6 Adversarial Training Experiments
  • 6.1 Experiment Setup
  • 6.2 Results
  • 6.3 Analysis
  • 7 Conclusion
  • Acknowledgements
  • Limitations
  • Ethical Consideration
  • References
  • A Appendix

Knowls

  1. Knowl 1 — ROBUT measures Table QA robustness with paired adversarial examples

    definition

    ROBUT is a diagnostic benchmark for question answering over tables. It augments the development data of WTQ, WIKISQL-WEAK, and SQA with human-annotated, answer-preserving perturbations, pairing each original example with a perturbed version. The benchmark contains 138,149 such pairs: 39,471 from WTQ, 83,816 from WIKISQL-WEAK, and 14,862 from SQA. The source datasets cover complex, simple, and conversational table QA, respectively. Their reported source statistics are: WTQ, 2,108 tables and 22,033 examples; WIKISQL-WEAK, 24,241 tables and 80,654 examples; SQA, 982 tables and 6,066 examples.

  2. Knowl 2 — ROBUT covers ten answer-preserving perturbation types

    model/method

    ROBUT organizes its perturbations across table headers, table contents, natural-language questions, and mixtures of perturbations. Header perturbations replace column names with context-appropriate synonyms or abbreviations. Table-content perturbations shuffle row order, shuffle column order, extend a compound column into semantically equivalent columns, mask columns inferable from other columns, or add semantically related columns. Question perturbations paraphrase at the word level or sentence level. A mix example combines two or three perturbations. The perturbations are designed to retain the original answer; questions whose answers depend on absolute row or column position are excluded from the corresponding shuffling sets. The benchmark does not alter original cell values.

  3. Knowl 3 — Human annotation emphasizes semantic validity and model-sensitive question paraphrases

    experimental setup

    ROBUT annotation was guided by three criteria: cover multiple diagnostic angles, use linguistically correct and varied changes, and preserve semantic association with the original context. For question-level attacks, annotators received the table, original question, and the answer predicted by a fine-tuned TaBERT-Small model. They rewrote the question at the word or sentence level while retaining its meaning; a rewrite was kept as an adversarial example when the model's prediction changed, with the original predicted answer retained as the target. Word-level changes focus on reasoning-operation, header, or cell-value indicators; sentence-level changes alter sentence structure without adding meaning-changing noise. For column addition, a TAPAS-based retriever finds three relevant tables from Web Data Commons and annotators select suitable columns. Header changes, column extensions, and masking were likewise checked for context-appropriate meaning preservation.

  4. Knowl 4 — Robustness Accuracy measures answer retention conditional on baseline correctness

    equation

    The benchmark reports Exact Match Accuracy on original and perturbed examples, and Robustness Accuracy (R-ACC), which is the fraction of originally correct examples that remain correct after perturbation. For a paired evaluation set of NN examples, let ciprec_i^{\mathrm{pre}} and cipostc_i^{\mathrm{post}} be binary indicators that the model answers pair ii correctly before and after perturbation. Then

    R-ACC=∑i=1Nciprecipost∑i=1Ncipre.\mathrm{R\text{-}ACC}=\frac{\sum_{i=1}^{N} c_i^{\mathrm{pre}}c_i^{\mathrm{post}}}{\sum_{i=1}^{N} c_i^{\mathrm{pre}}}.

    The denominator is the number of examples answered correctly before perturbation; thus R-ACC is not the same as post-perturbation accuracy. For SQA, the paper reports average accuracy over sequential questions. In the diagnostic experiments, fine-tuned models used the Large versions of the evaluated architectures and were trained for 20 epochs with batch size 128; each official training set was split 8:2 into train and validation portions, and the checkpoint with lowest validation loss was selected. GPT-3 was evaluated with two-shot prompting at temperature 0.7.

  5. Knowl 5 — Existing Table QA systems lose accuracy under every tested perturbation family

    empirical result

    The evaluation of TAPAS, TableFormer, TAPEX, OmniTab, and few-shot GPT-3 found accuracy degradation on the ROBUT adversarial sets across the tested perturbation types. The largest and most consistent weaknesses include column extension and mixed perturbations: on ROBUT-WIKISQL, column extension lowered accuracy by 37.9, 35.2, 38.8, and 37.4 percentage points for TAPAS, TableFormer, TAPEX, and OmniTab, respectively, and by 24.7 points for GPT-3. On ROBUT-WTQ, mixed perturbations reduced accuracy by 11.3–12.5 points for the four fine-tuned models and 6.8 points for GPT-3. TableFormer was especially robust to row and column order shuffling, but its advantage did not extend to most other perturbations; the authors conclude that architecture choices may protect against particular attacks rather than robustness failures in general.

  6. Knowl 6 — Few-shot GPT-3 has high conditional robustness on most WTQ perturbations

    empirical result

    For ROBUT-WTQ, text-davinci-003 was evaluated on 200 randomly sampled examples per perturbation type. Its R-ACC was 90.7 for header synonym replacement, 93.8 for abbreviation replacement, 90.2 for row shuffling, 93.3 for column shuffling, 81.4 for column extension, 97.0 for column masking, 85.6 for column addition, 93.7 for word-level question paraphrases, 94.2 for sentence-level paraphrases, and 83.2 for mixed perturbations. These conditional robustness scores generally exceeded those of fine-tuned Table QA systems, although GPT-3's post-perturbation accuracy was not always high. Among the three GPT-series models tested on the same sampled sets, gpt-3.5-turbo achieved the highest R-ACC on mixed perturbations (84.9, compared with 83.2 for text-davinci-003 and 82.5 for text-davinci-002).

  7. Knowl 7 — LETA uses prompted language models to generate adversarial training data

    model/method

    LETA (LLM-Enhanced Table QA Augmentation) augments the original Table QA training data with generated adversarial examples, then fine-tunes the target Table QA model on the combined data. It prompts GPT-3 (text-davinci-003) or CodeX (code-davinci-002) with task-specific demonstrations, repeating generation three times to obtain diverse examples. Header synonym and abbreviation prompts use 10 human-annotated demonstrations, each showing a header and the first two table rows. Column extension and masking prompts use eight demonstrations with original and modified columns and explanations. For column addition, a dense retriever supplies the three most relevant tables, and CodeX proposes one or two associated columns. Row and column shuffling use heuristics. Question prompts cover paraphrase categories involving reasoning-operation indicators, header indicators, cell-value indicators, sentence simplification, and interrogative transformation; each category uses five to eight demonstrations with explanations. The comparison system, RTA, uses rule-based generation for header and table-content changes and BERT-Attack for question paraphrases.

  8. Knowl 8 — LETA improves WTQ adversarial accuracy over both baseline and RTA training

    data/table

    The following table reports WTQ accuracy before augmentation and after training with RTA or LETA. Values are percentages; parenthetical values are changes from the corresponding unaugmented model. LETA gives higher adversarial accuracy than RTA for both TAPAS and TAPEX in nearly every category, while retaining more accuracy on the original development set than RTA.

    Evaluation setTAPAS baselineTAPAS + RTATAPAS + LETATAPEX baselineTAPEX + RTATAPEX + LETA
    Development set48.345.3 (-3.0)46.5 (-1.8)57.353.6 (-3.7)55.3 (-2.0)
    Header synonym replacement38.540.8 (+2.3)42.4 (+3.9)48.451.0 (+2.6)52.5 (+4.1)
    Header abbreviation replacement35.138.9 (+3.8)40.7 (+5.6)44.348.7 (+4.4)50.0 (+5.7)
    Row order shuffling40.642.3 (+1.7)42.2 (+1.6)45.748.1 (+2.4)48.2 (+2.5)
    Column order shuffling42.543.8 (+1.3)43.6 (+1.1)48.550.1 (+1.6)50.1 (+1.6)
    Column extension42.544.2 (+1.7)46.3 (+3.8)47.850.0 (+2.2)51.3 (+3.5)
    Column masking45.245.4 (+0.2)45.6 (+0.4)54.454.3 (-0.1)54.6 (+0.2)
    Column adding47.147.6 (+0.5)47.9 (+0.8)50.453.1 (+2.7)54.2 (+3.8)
    Word-level question paraphrase38.641.0 (+2.4)43.1 (+4.5)49.251.0 (+1.8)52.4 (+3.2)
    Sentence-level question paraphrase41.141.7 (+0.6)43.6 (+2.5)49.550.7 (+1.2)52.9 (+3.4)
    Mixed perturbations32.033.1 (+1.1)35.2 (+3.2)39.541.0 (+1.5)42.3 (+2.8)

    The results show that robustness gains usually come with lower accuracy on the unperturbed development set, but LETA incurs a smaller decrease than RTA for both model families.

  9. Knowl 9 — LETA's generation quality approaches human annotation for headers and questions, but not table content

    data/table

    Two evaluators compared 100 human-created and 100 LETA-created examples per perturbation type. They rated examples from 1 to 5 and selected which example was better; the table gives the percentage with average rating at least 4, the percentage of comparisons in which the example was selected as better (ties may count), and estimated annotation cost in dollars per 100 examples. LETA approaches human quality for header and question perturbations at much lower cost, but has a pronounced quality gap on table-content perturbations.

    PerturbationRating ≥ 4: humanRating ≥ 4: LETABetter selection: humanBetter selection: LETACost: human ($/100)Cost: LETA ($/100)
    Header synonym95.590.0695260.01.5
    Header abbreviation90.582.5764160.01.5
    Column extension90.063.59022100.06.0
    Column masking91.569.0852760.06.0
    Column adding92.070.0833530.08.5
    Word-level question paraphrase96.090.0705680.01.5
    Sentence-level question paraphrase94.092.0745080.01.5

    The authors also identify recurring LETA question-generation errors: changing the original meaning, failing to follow the prompt's requested paraphrase, omitting information, and hallucinating details.

  10. Knowl 10 — ROBUT and LETA leave cell-value attacks and other table tasks open

    limitation

    The study focuses on Table QA and does not establish robustness for other table-reasoning tasks, such as fact checking or table-to-text generation. ROBUT also omits perturbations that change original cell values, because those changes may alter the answer and require more annotation effort. LETA's automatic generation is less successful for table-content perturbations than for header and question paraphrases, and its generated questions can change meaning, omit information, mismatch the requested transformation, or introduce unsupported details.

Coverage note — No substantial contributed component was deliberately omitted; detailed prompt examples and individual error-case examples were not repeated because the knowls capture their methods, measured quality, and recurring failure modes.

References

  1. 1.Rami Aly, Zhijiang Guo, Michael Sejr Schlichtkrull, James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Oana Cocarascu, and Arpit Mittal. 2021. FEVEROUS: Fact extraction and VERification over unstructured and structured information. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1).
  2. 2.Stephen H Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, et al. 2022. Promptsource: An integrated development environment and repository for natural language prompts. arXiv preprint arXiv:2202.01279.
  3. 3.Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp. 2020. Beat the AI: Investigating adversarial human annotation for reading comprehension. Transactions of the Association for Computational Linguistics, 8:662–678.
  4. 4.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  5. 5.Kai-Wei Chang, He He, Robin Jia, and Sameer Singh. 2021. Robustness and adversarial examples in natural language processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: Tutorial Abstracts, pages 22–26, Punta Cana, Dominican Republic & Online. Association for Computational Linguistics.
  6. 6.Shuaichen Chang, Jun Wang, Mingwen Dong, Lin Pan, Henghui Zhu, Alexander Hanbo Li, Wuwei Lan, Sheng Zhang, Jiarong Jiang, Joseph Lilien, et al. 2023. Dr.spider: A diagnostic evaluation benchmark towards text-to-SQL robustness. In International Conference on Learning Representations.
  7. 7.Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
  8. 8.Wenhu Chen. 2022. Large language models are few(1)-shot table reasoners.
  9. 9.Wenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen, and William Yang Wang. 2020a. Logical natural language generation from open-domain tables. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7929–7942, Online. Association for Computational Linguistics.
  10. 10.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020b. Tabfact: A large-scale dataset for table-based fact verification. In International Conference on Learning Representations.
  11. 11.Zhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, and Dongmei Zhang. 2022. HiTab: A hierarchical table dataset for question answering and natural language generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1094–1110, Dublin, Ireland. Association for Computational Linguistics.
  12. 12.Minseok Cho, Reinald Kim Amplayo, Seung won Hwang, and Jonghyuck Park. 2018. Adversarial tableqa: Attention supervision for question answering on tables. In Asian Conference on Machine Learning.
  13. 13.Julian Eisenschlos, Syrine Krichene, and Thomas Müller. 2020. Understanding tables with intermediate pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 281–296, Online. Association for Computational Linguistics.
  14. 14.Yujian Gan, Xinyun Chen, Qiuping Huang, Matthew Purver, John R. Woodward, Jinxia Xie, and Pengsheng Huang. 2021. Towards robustness of text-to-SQL models against synonym substitution. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2505–2515, Online. Association for Computational Linguistics.
  15. 15.Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré. 2021. Robustness gym: Unifying the NLP evaluation landscape. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Demonstrations, pages 42–55, Online. Association for Computational Linguistics.
  16. 16.Jiaqi Guo, Ziliang Si, Yu Wang, Qian Liu, Ming Fan, Jian-Guang Lou, Zijiang Yang, and Ting Liu. 2021. Chase: A large-scale and pragmatic Chinese dataset for cross-database context-dependent text-to-SQL. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2316–2331, Online. Association for Computational Linguistics.
  17. 17.Vivek Gupta, Maitrey Mehta, Pegah Nokhiz, and Vivek Srikumar. 2020. INFOTABS: Inference on tables as semi-structured data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2309–2324, Online. Association for Computational Linguistics.
  18. 18.Vivek Gupta, Shuo Zhang, Alakananda Vempala, Yujie He, Temma Choji, and Vivek Srikumar. 2022. Right for the right reason: Evidence extraction for trustworthy tabular reasoning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3268–3283, Dublin, Ireland. Association for Computational Linguistics.
  19. 19.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4320–4333, Online. Association for Computational Linguistics.
  20. 20.Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1821–1831, Vancouver, Canada. Association for Computational Linguistics.
  21. 21.Zhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, and Weizhu Chen. 2022. OmniTab: Pretraining with natural and synthetic data for few-shot table-based question answering. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 932–942, Seattle, United States. Association for Computational Linguistics.
  22. 22.Oliver Lehmberg, Dominique Ritze, Robert Meusel, and Christian Bizer. 2016. A large public corpus of web tables containing time and context metadata. In Proceedings of the 25th International Conference Companion on World Wide Web, WWW ’16 Companion, page 75–76, Republic and Canton of Geneva, CHE. International World Wide Web Conferences Steering Committee.
  23. 23.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7871–7880, Online. Association for Computational Linguistics.
  24. 24.Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu. 2020. BERT-ATTACK: Adversarial attack against BERT using BERT. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6193–6202, Online. Association for Computational Linguistics.
  25. 25.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2021. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. CoRR, abs/2107.13586.
  26. 26.Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022. TAPEX: Table pre-training via learning a neural SQL executor. In International Conference on Learning Representations.
  27. 27.Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1470–1480, Beijing, China. Association for Computational Linguistics.
  28. 28.Xinyu Pi, Bing Wang, Yan Gao, Jiaqi Guo, Zhoujun Li, and Jian-Guang Lou. 2022. Towards robustness of text-to-SQL models against natural and realistic adversarial table perturbation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2007–2022, Dublin, Ireland. Association for Computational Linguistics.
  29. 29.Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilic, Daniel Hesslow, Roman ´ Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100.
  30. 30.Torsten Scholak, Nathan Schucher, and Dzmitry Bahdanau. 2021. PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9895–9901, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Alane Suhr, Ming-Wei Chang, Peter Shaw, and Kenton Lee. 2020. Exploring unexplored generalization challenges for cross-database semantic parsing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8372–8388, Online. Association for Computational Linguistics.
  32. 32.Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. 2019. Robustness may be at odds with accuracy. In International Conference on Learning Representations.
  33. 33.Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. 2020a. RAT-SQL: Relation-aware schema encoding and linking for text-to-SQL parsers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7567–7578, Online. Association for Computational Linguistics.
  34. 34.Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li. 2021. Adversarial GLUE: A multi-task benchmark for robustness evaluation of language models. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
  35. 35.Lijie Wang, Ao Zhang, Kun Wu, Ke Sun, Zhenghua Li, Hua Wu, Min Zhang, and Haifeng Wang. 2020b. DuSQL: A large-scale and pragmatic Chinese text-to-SQL dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6923–6935, Online. Association for Computational Linguistics.
  36. 36.Tianlu Wang, Rohit Sridhar, Diyi Yang, and Xuezhi Wang. 2022a. Identifying and mitigating spurious correlations for improving robustness in NLP models. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1719–1729, Seattle, United States. Association for Computational Linguistics.
  37. 37.Xuezhi Wang, Haohan Wang, and Diyi Yang. 2022b. Measure and improve robustness in NLP models: A survey. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4569–4586, Seattle, United States. Association for Computational Linguistics.
  38. 38.Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022a. Emergent abilities of large language models. Transactions on Machine Learning Research. Survey Certification.
  39. 39.Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022b. Chain of thought prompting elicits reasoning in large language models.
  40. 40.Jingfeng Yang, Aditya Gupta, Shyam Upadhyay, Luheng He, Rahul Goel, and Shachi Paul. 2022. TableFormer: Robust transformer modeling for table-text encoding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 528–537, Dublin, Ireland. Association for Computational Linguistics.
  41. 41.Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. TaBERT: Pretraining for joint understanding of textual and tabular data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8413–8426, Online. Association for Computational Linguistics.
  42. 42.Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, Zilin Zhang, and Dragomir Radev. 2018. Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921, Brussels, Belgium. Association for Computational Linguistics.
  43. 43.Tao Yu, Rui Zhang, Michihiro Yasunaga, Yi Chern Tan, Xi Victoria Lin, Suyi Li, Heyang Er, Irene Li, Bo Pang, Tao Chen, Emily Ji, Shreya Dixit, David Proctor, Sungrok Shim, Jonathan Kraft, Vincent Zhang, Caiming Xiong, Richard Socher, and Dragomir Radev. 2019. SParC: Cross-domain semantic parsing in context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4511–4523, Florence, Italy. Association for Computational Linguistics.
  44. 44.Jichuan Zeng, Xi Victoria Lin, Steven C.H. Hoi, Richard Socher, Caiming Xiong, Michael Lyu, and Irwin King. 2020. Photon: A robust cross-domain text-to-SQL system. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 204–214, Online. Association for Computational Linguistics.
  45. 45.Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan. 2019. Theoretically principled trade-off between robustness and accuracy. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pages 7472–7482. PMLR.
  46. 46.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. Opt: Open pre-trained transformer language models.
  47. 47.Chen Zhao, Yu Su, Adam Pauls, and Emmanouil Antonios Platanios. 2022a. Bridging the generalization gap in text-to-SQL parsing with schema expansion. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5568–5578, Dublin, Ireland. Association for Computational Linguistics.
  48. 48.Yilun Zhao, Linyong Nan, Zhenting Qi, Rui Zhang, and Dragomir Radev. 2022b. ReasTAP: Injecting table reasoning skills during pre-training via synthetic reasoning examples. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 9006–9018, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  49. 49.Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. CoRR, abs/1709.00103.
  50. 50.Yi Zhu, Yiwei Zhou, and Menglin Xia. 2020. Generating semantically valid adversarial questions for tableqa.

Citation

MLA
Zhao, Y., et al. “RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 6064–81, https://doi.org/10.18653/v1/2023.acl-long.334.
APA
Zhao, Y., Zhao, C., Nan, L., Qi, Z., Zhang, W., Tang, X., Mi, B., & Radev, D. (2023). RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6064–6081. https://doi.org/10.18653/v1/2023.acl-long.334
Chicago
Zhao, Y., C. Zhao, L. Nan, et al. 2023. “RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 6064–81. https://doi.org/10.18653/v1/2023.acl-long.334.
Harvard
Zhao, Y. et al. (2023) “RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 6064–6081. Available at: https://doi.org/10.18653/v1/2023.acl-long.334.
Vancouver
1. Zhao Y, Zhao C, Nan L, Qi Z, Zhang W, Tang X, Mi B, Radev D (2023) RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 6064–6081

BibTeX

@inproceedings{zhao-etal-2023-robut,
    title = "{R}obu{T}: A Systematic Study of Table {QA} Robustness Against Human-Annotated Adversarial Perturbations",
    author = "Zhao, Yilun  and
      Zhao, Chen  and
      Nan, Linyong  and
      Qi, Zhenting  and
      Zhang, Wenlin  and
      Tang, Xiangru  and
      Mi, Boyu  and
      Radev, Dragomir",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.334/",
    doi = "10.18653/v1/2023.acl-long.334",
    pages = "6064--6081"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/