PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training

Zihui GuJu FanNan TangPreslav NakovXiaoman ZhaoXiaoyong Du

article2022EMNLP67 citations

Presents PASTA, a table-based fact verification framework that pre-trains language models on 1.2 million synthesized sentence-table cloze tasks covering common operations like aggregation and comparison, setting state-of-the-art results on TabFact and SEM-TAB-FACTS.

Listen

Misinformation and disinformation online pose growing operational, reputational, and decision-making risks across journalism, public policy, and corporate governance. While verifying assertions against unstructured text is challenging, many false statements can be directly validated or refuted using structured tables. However, standard language models often struggle to interpret tabular data because they lack the ability to perform basic logical and arithmetic operations, such as comparing entries or aggregating columns.

The article aims to introduce and evaluate PASTA, a novel framework designed to enhance language models with table-operation awareness for automated fact verification without relying on complex, error-prone logical program generators.

The authors constructed an automated pre-training dataset of roughly 1.2 million fill-in-the-blank questions derived from 20,000 Wikipedia tables. These questions target six core table operations: filtering, aggregation, superlatives, comparisons, ordinal rankings, and uniqueness checks. Using the DeBERTaV3 language model architecture, the system was trained to predict masked operation-specific words and values. During downstream deployment, a select-then-rank pre-processing strategy was applied to prioritize the most relevant table rows and columns within the model's limited input window. The framework was evaluated across standard benchmarks containing simple and multi-row complex statements, including scientific domain data.

The evaluation yielded three primary findings. First, PASTA established a new state-of-the-art benchmark on the primary testing dataset (TabFact), achieving 89.3% overall accuracy and outperforming the previous state-of-the-art by 4.7 percentage points on complex, multi-operation statements (85.6% versus 80.9%). Second, the model substantially closed the gap with human accuracy on a held-out test set, trailing human performance by only 1.5 percentage points (90.6% versus 92.1%). Third, the model demonstrated strong cross-domain transferability on scientific literature tables (SEM-TAB-FACTS), exceeding baseline models by 5.2 points (84.1% micro-F1).

These results indicate that pre-training language models on structured, operation-aware tasks significantly enhances their symbolic reasoning capabilities without requiring expensive manual annotations or brittle semantic parsing. Organizations deploying automated fact-checking or tabular data analysis can achieve higher accuracy and reliability with relatively modest training data sizes, lowering development timelines and computational overhead.

Decision-makers seeking to implement automated table verification should consider adopting targeted operation-aware pre-training over traditional random masking approaches, while implementing row-ranking heuristics to manage large tabular inputs. Before deploying such systems in production, practitioners should conduct pilot testing specifically focused on complex claims that combine multiple operational steps.

Confidence in these findings is high for single-table verification scenarios supported by clean data structures. However, stakeholders should exercise caution, as the system's performance declines when processing very large tables or statements that require chains of multiple distinct operations. Additionally, the current framework is constrained to verifying claims against a single table at a time and assumes the underlying reference tables are accurate and unbiased.

Cover for PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training

Abstract

Fact verification has attracted a lot of research attention recently, e.g., in journalism, marketing, and policymaking, as misinformation and disinformation online can sway one's opinion and affect one's actions. While fact-checking is a hard task in general, in many cases, false statements can be easily debunked based on analytics over tables with reliable information. Hence, table-based fact verification has recently emerged as an important and growing research area. Yet, progress has been limited due to the lack of datasets that can be used to pre-train language models (LMs) to be aware of common table operations, such as aggregating a column or comparing tuples. To bridge this gap, in this paper we introduce PASTA, a novel state-of-the-art framework for table-based fact verification via pre-training with synthesized sentence–table cloze questions. In particular, we design six types of common sentence–table cloze tasks, including Filter, Aggregation, Superlative, Comparative, Ordinal, and Unique, based on which we synthesize a large corpus consisting of 1.2 million sentence–table pairs from WikiTables. PASTA uses a recent pre-trained LM, DeBERTaV3, and further pre-trains it on our corpus. Our experimental results show that PASTA achieves new state-of-the-art performance on two table-based fact verification benchmarks: TabFact and SEM-TAB-FACTS. In particular, on the complex set of TabFact, which contains multiple operations, PASTA largely outperforms the previous state of the art by 4.7 points (85.6% vs. 80.9%), and the gap between PASTA and human performance on the small TabFact test set is narrowed to just 1.5 points (90.6% vs. 92.1%).

Table of Contents

  • 1 Introduction
  • 2 Preliminaries
  • 2.1 Problem Formulation
  • 2.2 DeBERTa for Sentence-Table Encoding
  • 3 Our PASTA Model
  • 3.1 Sentence-Table Cloze Pre-training
  • 3.2 Pre-training Corpus Generation
  • 3.3 Fine-tuning with Select-then-Rank
  • 4 Experimental Setup
  • 4.1 Datasets
  • 4.2 Baselines
  • 4.3 Implementation Details
  • 5 Experiments and Results
  • 5.1 Overall Performance
  • 5.2 Impact of Operation Aware Pre-training
  • 5.3 Impact of Select-then-Rank
  • 5.4 Error Analysis
  • 6 Related Work
  • 7 Conclusion and Future Work
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Details of Pre-training Corpus
  • B Trigger Words Definition

Knowls

  1. Knowl 1 — PASTA pre-trains a table-aware encoder before binary fact verification

    model/method

    PASTA adapts DeBERTaV3 to verify whether a natural-language statement is entailed or refuted by a table. It first pre-trains the encoder on synthetic sentence–table cloze examples, then fine-tunes it for binary verification without generating an explicit logical form. To encode a table, PASTA serializes its headers after a [Header] marker and each row after a [Row] marker; | separates cells. It concatenates this serialization with the statement and feeds the joint input to DeBERTaV3.

    The implementation starts from the public DeBERTaV3-Large checkpoint. Pre-training uses Adam for up to 400,000 steps, with batch size 16 and learning rate 1×10−61\times10^{-6}. Fine-tuning runs for up to 300,000 steps, with batch size 8 and learning rate 5×10−65\times10^{-6}.

  2. Knowl 2 — The cloze objective predicts only table-grounded operation tokens

    equation

    For a sentence S=(x1,…,xn)S=(x_1,\ldots,x_n) and table TT, PASTA masks the operation-aware token span in the sentence, producing the corrupted sentence S~\tilde S. Let MM be the set of masked token positions; xix_i is the original token at position ii, and pθp_\theta is the conditional token distribution of the language model with parameters θ\theta. PASTA trains the model to recover the original span from the corrupted sentence and table:

    LPASTA=−∑i∈Mlog⁡pθ(xi∣S~,T).\mathcal{L}_{PASTA}=-\sum_{i\in M}\log p_\theta(x_i\mid\tilde S,T).

    Unlike ordinary random-token masking, this objective targets sentence tokens that express an operation and can be predicted by reasoning over the table. A generated example masks one atomic operation type; combinations of operations are left for downstream fine-tuning. The answer span may be a table cell or an operation result, such as a comparison word or calculated value.

  3. Knowl 3 — PASTA covers six atomic table-operation types

    definition

    PASTA’s synthetic cloze tasks cover six operation types: Filter retrieves a value from rows satisfying a condition; Aggregation computes a quantity such as a sum or average; Superlative identifies an extremum, such as the highest or lowest value; Comparative determines a comparison between values or against a threshold; Ordinal identifies an item at a ranked position, such as the second highest; and Unique counts distinct values. The operation-aware span is the sentence text corresponding to the operation and its table-grounded answer.

  4. Knowl 4 — A template-and-execution pipeline creates the pre-training corpus

    model/method

    PASTA selects well-formed relational tables from WikiTables that have headers and at least one numeric column, and excludes tables with more than 500 cells. Although this filtering yields about 580,000 tables, the authors randomly select 20,000 for corpus generation. They create sentence–table examples from 50 manually designed natural-language/SQL template pairs: table headers and cell values fill template slots, the SQL instance is executed on the table, and its result fills the corresponding answer slot in the sentence. This execution step makes each generated sentence a correct description of its table. Depending on table size, the pipeline generates up to 100 sentences per table.

    To make template language more natural, a fixed pre-trained language model selects among candidate words for context-sensitive template slots. For example, it can select older rather than higher when the populated column concerns age. The resulting corpus contains 1,293,488 sentence–table examples: Filter 77,609 (6%; mean answer length 3.2 tokens), Superlative 349,241 (27%; 2.6), Aggregation 388,046 (30%; 1.3), Comparative 349,241 (27%; 1.0), Ordinal 103,479 (8%; 2.3), and Unique 25,872 (2%; 1.0). The overall mean answer length is 1.8 tokens.

  5. Knowl 5 — Select-then-rank fits tables to the encoder and prioritizes relevant rows

    model/method

    For downstream verification, PASTA addresses the 512-token input limit with two table transformations. First, column-wise selection keeps only columns containing entities linked to the statement; it does not discard rows, because operations such as aggregation may require a whole column. Second, row-wise ranking reorders the selected table’s rows by descending lexical overlap with the statement. The relevance score for a row is the number of shared tokens between that row and the statement after stop-word removal. PASTA then serializes the reordered table with its table markers. The authors’ rationale is that placing relevant cells nearer to the statement may help DeBERTa capture their relationship.

    On TabFact, the full procedure scores 89.2% validation, 89.3% test, 96.7% simple-test, and 85.6% complex-test accuracy. Without column selection, the respective scores are 88.9%, 89.0%, 95.8%, and 85.5%; without row ranking, they are 88.2%, 88.4%, 95.1%, and 84.7%. Thus, both components contribute in this ablation, with the larger observed drop occurring when row ranking is removed.

  6. Knowl 6 — PASTA sets new TabFact accuracy results, including on complex statements

    empirical result

    On TabFact, PASTA achieves 89.2±0.4% validation accuracy, 89.3±0.3% test accuracy, 96.7±0.2% simple-test accuracy, 85.6±0.3% complex-test accuracy, and 90.6±0.2% accuracy on the small test set. The reported values are means with standard deviations over five runs. The complex set contains statements involving multiple rows and table operations. PASTA’s 85.6% on that set exceeds the previous best listed result, SaMoE’s 80.9%, by 4.7 percentage points. On the small test set, PASTA’s 90.6% is 1.5 points below the reported human accuracy of 92.1%.

    The DeBERTaV3 baseline scores 86.1±0.2% validation, 86.2±0.1% test, 92.8±0.2% simple-test, 82.9±0.1% complex-test, and 86.5±0.3% small-test accuracy. PASTA therefore improves test accuracy by 3.1 points over this baseline, including 2.7 points on the complex test set. The authors report that PASTA leads the compared systems on all listed TabFact splits.

  7. Knowl 7 — PASTA transfers to scientific tables in SEM-TAB-FACTS

    empirical result

    On SEM-TAB-FACTS, a benchmark of statements and tables from scientific articles, PASTA obtains 84.23% validation and 84.10% test micro-F1. DeBERTaV3 obtains 81.85% validation and 78.92% test micro-F1; LKA, the strongest other listed test result, obtains 80.34% validation and 78.54% test micro-F1. PASTA’s test score is 5.18 points above DeBERTaV3 and 5.56 points above LKA. The result shows transfer to this scientific-table benchmark even though PASTA’s synthetic pre-training tables were drawn from Wikipedia.

  8. Knowl 8 — Operation-aware masking beats random MLM masking on TabFact

    empirical result

    A controlled TabFact comparison tests operation-aware PASTA masking against random masked-language-model (MLM) masking. Both variants start from DeBERTaV3, are pre-trained for 140,000 steps, and are evaluated without table pre-processing; results are means with standard deviations over five runs. MLM randomly masks 15% of tokens in a sentence–table pair, whereas PASTA masks the operation-aware sentence span. PASTA scores 87.0±0.1% validation, 87.9±0.2% test, 94.0±0.2% simple-test, and 84.9±0.2% complex-test accuracy; MLM scores 84.8±0.2%, 84.9±0.2%, 92.5±0.1%, and 81.2±0.2%, respectively. The test-set gains are 3.0 points overall and 3.7 points on complex statements.

    The authors also evaluate cloze completion on six held-out sets of 1,000 examples each, with no tables shared with training. After 400,000 pre-training steps, PASTA correctly completes more than 60% of examples for every operation type. Comparative cloze performance develops earliest, while Aggregation and Filter are the last types to be mastered.

  9. Knowl 9 — PASTA improves verification accuracy across all six operation subsets

    empirical result

    The authors form six non-overlapping TabFact test subsets by operation trigger words, with 200 statement–table pairs per subset, and compare PASTA with DeBERTaV3. PASTA’s versus DeBERTaV3’s accuracy is: Filter, 90.7% versus 88.0%; Superlative, 86.7% versus 85.9%; Aggregation, 84.5% versus 81.0%; Comparative, 86.2% versus 85.2%; Ordinal, 86.9% versus 83.8%; and Unique, 79.1% versus 74.2%. PASTA is higher in every subset, with the largest absolute gain on Unique (+4.9 points), followed by Aggregation (+3.5) and Ordinal (+3.1). The authors attribute these improvements to operation-aware pre-training.

  10. Knowl 10 — Template diversity, large tables, multiple operations, and multi-table evidence remain limitations

    limitation

    The authors identify two scope limitations: synthetic statements are generated from human-designed templates and may lack linguistic diversity, and the verification setup uses one table per statement rather than combining evidence across multiple tables. Their TabFact error analysis also indicates residual difficulty with larger tables and operation-rich statements. PASTA’s error subset has an average of 97.5 cells per table, compared with 89.0 cells across the full test set; this is lower than DeBERTaV3’s error-subset average of 107.4 cells but still above the test-set average. The reported proportion of multiple-operation statements is 16.5% in PASTA’s error subset versus 11.3% in the full test set. The authors consequently identify more diverse controlled pre-training data, larger tables, and more complex operations as areas for further work.

Coverage note — The appendix’s exhaustive list of template pairs, candidate-word sets, and trigger words is omitted as implementation detail; their roles are summarized in the operation taxonomy and corpus-generation knowls.

References

  1. 1.Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey. 2013. Methods for exploring and mining tables on wikipedia. In Proceedings of the ACM SIGKDD Workshop on Interactive Data Exploration and Analytics, IDEA@KDD 2013, Chicago, Illinois, USA, August 11, 2013, pages 18–26. ACM.
  2. 2.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020a. Tabfact: A large-scale dataset for table-based fact verification. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  3. 3.Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, and Brian D. Davison. 2020b. Table search using a deep contextualized language model. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, China, July 25-30, 2020, pages 589–598. ACM.
  4. 4.Zhoujun Cheng, Haoyu Dong, Fan Cheng, Ran Jia, Pengfei Wu, Shi Han, and Dongmei Zhang. 2021. FORTAP: using formulae for numerical-reasoning-aware table pretraining. CoRR, abs/2109.07323.
  5. 5.Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: pre-training text encoders as discriminators rather than generators. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics.
  7. 7.Julian Martin Eisenschlos, Syrine Krichene, and Thomas Müller. 2020. Understanding tables with intermediate pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 281–296. Association for Computational Linguistics.
  8. 8.Devansh Gautam, Kshitij Gupta, and Manish Shrivastava. 2021. Volta at semeval-2021 task 9: Statement verification and evidence finding with tables using TAPAS and transfer learning. In Proceedings of the 15th International Workshop on Semantic Evaluation, SemEval@ACL/IJCNLP 2021, Virtual Event / Bangkok, Thailand, August 5-6, 2021, pages 1262–1270. Association for Computational Linguistics.
  9. 9.Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021a. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. CoRR, abs/2111.09543.
  10. 10.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021b. Deberta: decoding-enhanced bert with disentangled attention. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  11. 11.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. 2020. Tapas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 4320–4333. Association for Computational Linguistics.
  12. 12.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  13. 13.Qian Liu, Bei Chen, Jiaqi Guo, Zeqi Lin, and Jian-Guang Lou. 2021. TAPEX: table pre-training via learning a neural SQL executor. CoRR, abs/2107.07653.
  14. 14.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized BERT pretraining approach. CoRR, abs/1907.11692.
  15. 15.Thomas Müller, Julian Eisenschlos, and Syrine Krichene. 2021. TAPAS at semeval-2021 task 9: Reasoning over tables with intermediate pre-training. In Proceedings of the 15th International Workshop on Semantic Evaluation, SemEval@ACL/IJCNLP 2021, Virtual Event / Bangkok, Thailand, August 5-6, 2021, pages 423–430. Association for Computational Linguistics.
  16. 16.Myle Ott, Yejin Choi, Claire Cardie, and Jeffrey T. Hancock. 2011. Finding deceptive opinion spam by any stretch of the imagination. In The 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Proceedings of the Conference, 19-24 June, 2011, Portland, Oregon, USA, pages 309–319. The Association for Computer Linguistics.
  17. 17.Kashyap Popat, Subhabrata Mukherjee, Jannik Strötgen, and Gerhard Weikum. 2017. Where the truth lies: Explaining the credibility of emerging claims on the web and social media. In WWW, pages 1003–1012. ACM.
  18. 18.Michael Sejr Schlichtkrull, Vladimir Karpukhin, Barlas Oguz, Mike Lewis, Wen-tau Yih, and Sebastian Riedel. 2021. Joint verification and reranking for open fact checking over tables. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 6787–6799. Association for Computational Linguistics.
  19. 19.Shaden Shaar, Nikolay Babulkov, Giovanni Da San Martino, and Preslav Nakov. 2020. That is a known lie: Detecting previously fact-checked claims. In ACL, pages 3607–3618. Association for Computational Linguistics.
  20. 20.Tianze Shi, Chen Zhao, Jordan L. Boyd-Graber, Hal Daumé III, and Lillian Lee. 2020. On the potential of lexico-logical alignments for semantic parsing to SQL queries. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 1849–1864. Association for Computational Linguistics.
  21. 21.Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. 2017. Fake news detection on social media: A data mining perspective. SIGKDD Explor., 19(1):22–36.
  22. 22.James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018. FEVER: a large-scale dataset for fact extraction and verification. In NAACL-HLT, pages 809–819. Association for Computational Linguistics.
  23. 23.Fei Wang, Kexuan Sun, Jay Pujara, Pedro A. Szekely, and Muhao Chen. 2021a. Table-based fact verification with salience-aware learning. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 4025–4036. Association for Computational Linguistics.
  24. 24.Nancy Xin Ru Wang, Diwakar Mahajan, Marina Danilevsky, and Sara Rosenthal. 2021b. Semeval-2021 task 9: Fact verification and evidence finding for tabular data in scientific documents (SEM-TAB-FACTS). In Proceedings of the 15th International Workshop on Semantic Evaluation, SemEval@ACL/IJCNLP 2021, Virtual Event / Bangkok, Thailand, August 5-6, 2021, pages 317–326. Association for Computational Linguistics.
  25. 25.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  26. 26.Xiaoyu Yang, Feng Nie, Yufei Feng, Quan Liu, Zhigang Chen, and Xiaodan Zhu. 2020. Program enhanced fact verification with verbalization and graph attention network. In Proceedings of the 2020 Conference on Empirical Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 7810–7825. Association for Computational Linguistics.
  27. 27.Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. Tabert: Pretraining for joint understanding of textual and tabular data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 8413–8426. Association for Computational Linguistics.
  28. 28.Seunghyun Yoon, Kunwoo Park, Joongbo Shin, Hongjun Lim, Seungpil Won, Meeyoung Cha, and Kyomin Jung. 2019. Detecting incongruity between news headline and body text via a deep hierarchical encoder. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pages 791–800. AAAI Press.
  29. 29.Hongzhi Zhang, Yingyao Wang, Sirui Wang, Xuezhi Cao, Fuzheng Zhang, and Zhongyuan Wang. 2020. Table fact verification with structure-aware transformer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 1624–1629. Association for Computational Linguistics.
  30. 30.Guangzhen Zhao and Peng Yang. 2022. Table-based fact verification with self-labeled keypoint alignment. In Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12-17, 2022, pages 1401–1411. International Committee on Computational Linguistics.
  31. 31.Wanjun Zhong, Duyu Tang, Zhangyin Feng, Nan Duan, Ming Zhou, Ming Gong, Linjun Shou, Daxin Jiang, Jiahai Wang, and Jian Yin. 2020. Logicalfactchecker: Leveraging logical operations for fact checking with graph module network. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 6053–6065. Association for Computational Linguistics.
  32. 32.Yuxuan Zhou, Xien Liu, Kaiyin Zhou, and Ji Wu. 2022. Table-based fact verification with self-adaptive mixture of experts. In Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 139–149. Association for Computational Linguistics.

Citation

MLA
Gu, Z., et al. “PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 4971–83, https://doi.org/10.18653/v1/2022.emnlp-main.331.
APA
Gu, Z., Fan, J., Tang, N., Nakov, P., Zhao, X., & Du, X. (2022). PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4971–4983. https://doi.org/10.18653/v1/2022.emnlp-main.331
Chicago
Gu, Z., J. Fan, N. Tang, P. Nakov, X. Zhao, and X. Du. 2022. “PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 4971–83. https://doi.org/10.18653/v1/2022.emnlp-main.331.
Harvard
Gu, Z. et al. (2022) “PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 4971–4983. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.331.
Vancouver
1. Gu Z, Fan J, Tang N, Nakov P, Zhao X, Du X (2022) PASTA: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 4971–4983

BibTeX

@inproceedings{gu-etal-2022-pasta,
    title = "{PASTA}: Table-Operations Aware Fact Verification via Sentence-Table Cloze Pre-training",
    author = "Gu, Zihui  and
      Fan, Ju  and
      Tang, Nan  and
      Nakov, Preslav  and
      Zhao, Xiaoman  and
      Du, Xiaoyong",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.331/",
    doi = "10.18653/v1/2022.emnlp-main.331",
    pages = "4971--4983"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/