OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering

Zhengbao JiangYi MaoPengcheng HeGraham NeubigWeizhu Chen

article2022NAACL60 citations

Proposes a pretraining framework that pairs natural text-table alignment with synthetic SQL-derived questions, establishing a new state-of-the-art on WikiTableQuestions while dramatically reducing annotation requirements for few-shot table question answering.

Listen

Tables present vast amounts of critical data across web pages, business reports, and scientific literature, but extracting insights from them remains challenging. Effective table-based question answering systems must connect free-form natural language questions to tabular structures and execute complex multi-step reasoning, such as aggregation, sorting, and comparison. Building such systems has historically required complex multi-stage architectures and extensive manual data annotation, creating high development costs and deployment friction.

The article demonstrates an end-to-end table question answering system, named OmniTab, designed to answer complex tabular questions with minimal human annotation. The primary objective is to evaluate whether combining natural text retrieved from the web with programmatically generated synthetic data can effectively train models under low-resource and few-shot conditions.

The researchers developed an omnivorous pretraining framework utilizing two complementary data streams derived from approximately 500,000 Wikipedia tables. The first stream retrieves relevant sentences from Wikipedia documents and masks cell mentions to teach the model how to align natural phrasing with table contents. The second stream converts structured database queries into synthetic natural language questions using a query-to-text model refined by verification-based self-training, which teaches the model formal reasoning skills. The combined system was evaluated across standard benchmarks, including WikiTableQuestions and WikiSQL, primarily focusing on few-shot scenarios ranging from 16 to 1,024 labeled examples.

The experimental findings show substantial performance gains over existing baseline methods. When trained on only 128 labeled examples, OmniTab achieved a 41.4% accuracy on WikiTableQuestions, outperforming the best baseline by an absolute margin of 16.2 percentage points (a relative improvement of approximately 64%). Using natural or synthetic data independently in this 128-shot setting improved performance by 13.2% and 12.3% respectively, confirming that the two data types provide distinct, complementary benefits. Even in full-data settings with all 11,000 training examples, OmniTab achieved a state-of-the-art accuracy of 62.8%, improving by 2.7 percentage points over prior models. Furthermore, cross-topic evaluations demonstrated that the model maintains robust performance when transferring across different subject domains, including sports, politics, culture, and people.

These results indicate that organizations can build high-performing tabular question answering systems without investing in large, costly manual annotation efforts. By pairing dense retrieval mechanisms with verified synthetic question generation, developers can bridge the gap between human language variation and structured multi-step reasoning. This capability substantially reduces the time and expense required to deploy natural language interfaces over enterprise databases and tabular records.

Organizations aiming to build tabular question answering capabilities should adopt hybrid pretraining strategies that combine natural retrieval with synthetic reasoning data. Teams should prioritize salient mention masking over random token masking, and implement verification-based filtering when generating synthetic training data. Because dense retrieval depends heavily on similarity thresholds and the synthetic generation model still exhibits a performance gap compared to fully supervised data, future initiatives should conduct small-scale pilot testing and investigate improved retrieval and synthesis architectures before full-scale deployment.

arXiv: 2207.03637jzbjyb/OmniTab
Cover for OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering

Abstract

The information in tables can be an important complement to text, making table-based question answering (QA) systems of great value. The intrinsic complexity of handling tables often adds an extra burden to both model design and data annotation. In this paper, we aim to develop a simple table-based QA model with minimal annotation effort. Motivated by the fact that table-based QA requires both alignment between questions and tables and the ability to perform complicated reasoning over multiple table elements, we propose an omnivorous pretraining approach that consumes both natural and synthetic data to endow models with these respective abilities. Specifically, given freely available tables, we leverage retrieval to pair them with relevant natural sentences for mask-based pretraining, and synthesize NL questions by converting SQL sampled from tables for pretraining with a QA loss. We perform extensive experiments in both few-shot and full settings, and the results clearly demonstrate the superiority of our model OmniTab, with the best multitasking approach achieving an absolute gain of 16.2% and 2.7% in 128-shot and full settings respectively, also establishing a new state-of-the-art on WikiTableQuestions. Detailed ablations and analyses reveal different characteristics of natural and synthetic data, shedding light on future directions in omnivorous pretraining.

Table of Contents

  • 1 Introduction
  • 2 End2End Table-based QA
  • 3 OmniTab: Pretraining with Natural and Synthetic Data
  • 3.1 NL-Table Alignment Through Retrieval
  • 3.1.1 Retrieval Protocol
  • 3.1.2 Learning Objective
  • 3.2 Synthetic Questions Converted from SQL
  • 3.3 Combining Natural and Synthetic Data
  • 4 Experiments
  • 4.1 Experimental Settings
  • 4.2 Overall Results
  • 4.3 Ablation Study
  • 4.4 Analysis
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — OmniTab uses a unified end-to-end model for table question answering

    model/method

    OmniTab adapts TAPEX, which is based on BART, to answer a natural-language question qq using a table TT. The table is linearized in row-major order: the header is prefixed with col:, each data row with its row index, and cells in a row are separated by |. The question is prepended to this representation, and an encoder-decoder model generates the answer sequence; multiple answers are joined with commas. OmniTab first pretrains this model using both retrieved natural sentences paired with tables and synthetic questions paired with tables, then fine-tunes it on a limited set of annotated questions. The two pretraining sources target complementary skills: connecting natural-language expressions to table elements and learning table-based reasoning.

  2. Knowl 2 — Natural sentence–table pairs are retrieved from the table’s own Wikipedia page

    model/method

    OmniTab constructs natural-language/table pretraining pairs from Wikipedia tables and sentences on the same page, restricting retrieval to the page containing each table to reduce noise. It compares three ways to select a sentence. In longest-common-substring matching, each sentence is checked against table cells; a matched substring is retained as a mention only if it is not a stopword, contains an alphanumeric character, is a complete word, and exceeds 70% of the cell’s length. The sentence with the most mentions is paired with the table. As an alternative, BM25 retrieves a sentence using the linearized table as a bag-of-tokens query, after which the substring procedure identifies mentions. The resulting corpus contains approximately 0.5 million table–sentence pairs.

  3. Knowl 3 — Dense retrieval aligns sentence phrases directly with table cells

    model/method

    To match expressions that differ lexically between a table and its associated sentence, OmniTab represents each table cell and each candidate sentence phrase with a dense vector. BART supplies token representations; candidate phrases are detected by named-entity recognition, and each phrase or cell is represented by the mean of its token vectors, normalized to unit length. For every cell–phrase pair, the dot product gives a similarity, forming a matrix whose rows are cells and columns are phrases. OmniTab sparsifies this matrix in two steps: for each cell, it keeps only the highest-scoring phrase, then for each remaining phrase, it keeps only its highest-scoring cell. The sum of the retained similarities ranks candidate sentences, and the retained score for each phrase ranks mentions to mask. In the reported setup, phrases scoring above 0.60.6 are treated as mentions; this threshold balances match quality against the number of available matches.

  4. Knowl 4 — Masked mention recovery trains cross-format alignment

    equation

    For a retrieved sentence ss paired with a table TT, OmniTab uses both random masking and salient-mention masking. Random masking hides tokens in the sentence or cells in the table; salient masking hides sentence phrases identified as aligned with table elements. The model receives the masked sentence s∗s^* and masked table T∗T^* and is trained to reconstruct the original inputs, with loss applied only at masked positions. This focuses learning on recovering information missing from one of the two formats and, in particular, on sentence phrases that correspond to table content.

    Lmask=−log⁡Pmask(s,T∣s∗,T∗)\mathcal{L}_{\mathrm{mask}}=-\log P_{\mathrm{mask}}(s,T\mid s^*,T^*)

    Here, PmaskP_{\mathrm{mask}} denotes the model probability of the original sentence and table evaluated only at masked positions.

  5. Knowl 5 — SQL queries are converted into synthetic natural-language QA examples

    model/method

    OmniTab generates synthetic table-based QA examples by sampling SQL queries from templates and converting each query into a natural-language question. The SQL templates encode operations such as filtering, aggregation, comparisons, and nested queries; executing each SQL query on its table supplies the target answer. The SQL-to-question converter is a BART model fine-tuned on a limited set of SQL–natural-language pairs. For pretraining, the generated question qq and its table TT are input to the QA model, which generates the executed answer aa using the standard sequence-generation loss:

    LQA=−log⁡P(a∣q,T)\mathcal{L}_{\mathrm{QA}}=-\log P(a\mid q,T)

    In this equation, P(a∣q,T)P(a\mid q,T) is the model probability of the answer sequence given the question and table. The experiments use approximately 0.5 million sampled SQL queries to produce synthetic questions.

  6. Knowl 6 — QA-based verification selects SQL-to-question self-training data

    algorithm

    OmniTab improves its SQL-to-question model by selecting generated questions according to whether a table-QA model can use them to recover the SQL execution answer. The inputs are labeled SQL–question pairs, unlabeled SQL queries with their tables and execution answers, and an OmniTab QA verifier trained with the same annotation budget as the target setting. The procedure is:

    Input: Labeled SQL-question pairs, unlabeled SQL queries with tables and execution answers, QA verifier, selection size K
    Output: Self-trained SQL-to-question model
    Fine-tune BART on the labeled SQL-question pairs
    For each unlabeled SQL query, generate 50 candidate questions using beam search
    For each candidate, score the probability that the QA verifier assigns to the SQL execution answer given the candidate and its table
    For each SQL query, retain the candidate with the highest QA-verification score
    Rank the retained SQL-question pairs by their verification scores
    Select the top K pairs and combine them with the labeled pairs
    Fine-tune BART on the combined labeled and selected pairs
    Return the fine-tuned model

    The paper’s final model adds approximately 10,000 SQL–question pairs through this process. Selecting candidates by the SQL-to-question model’s own generation probability did not provide the same benefit; the verification score uses the QA model’s ability to assess whether a question elicits the correct answer.

  7. Knowl 7 — Multitask pretraining combines masked recovery, natural questions, and SQL execution

    equation

    OmniTab combines the natural sentence–table masking objective with QA training on synthetic natural-language questions. It also retains TAPEX-style pretraining on SQL queries, since SQL-based pretraining had been shown to provide reasoning capability. The combined objective is an unweighted sum:

    Ltotal=Lmask+LQA+LSQL,LSQL=−log⁡P(a∣o,T)\mathcal{L}_{\mathrm{total}}=\mathcal{L}_{\mathrm{mask}}+\mathcal{L}_{\mathrm{QA}}+\mathcal{L}_{\mathrm{SQL}}, \qquad \mathcal{L}_{\mathrm{SQL}}=-\log P(a\mid o,T)

    Here, TT is a table, aa is its answer, and oo is an SQL query executed over TT. The masking loss reconstructs masked natural sentence/table content; the QA loss predicts answers from generated natural-language questions and tables; the SQL loss predicts answers from SQL queries and tables.

  8. Knowl 8 — OmniTab improves few-shot and full-data accuracy on WTQ and WikiSQL

    empirical result

    On WikiTableQuestions (WTQ), OmniTab was evaluated with 16, 128, or 1,024 annotated examples and in the full-data setting, using the official answer-accuracy evaluator. The full model’s WTQ accuracies were 26.8%, 41.4%, 51.9%, and 62.8%, respectively. For comparison, the strongest baseline in those settings achieved 15.7%, 25.2%, 45.5%, and 60.1%; OmniTab’s absolute gains were therefore 11.1, 16.2, 6.4, and 2.7 percentage points. At 128 shots, natural-data-only pretraining reached 38.4% and synthetic-data-only pretraining reached 37.5%; combining the data reached 41.4%. The reported WTQ accuracies for TAPEX further trained on SQL data for a comparable number of steps were 15.7%, 25.2%, 44.6%, and 60.1%; the natural-only model scored 22.8%, 38.4%, 49.8%, and 61.3%, and the synthetic-only model scored 21.5%, 37.5%, 48.8%, and 61.3%.

    On WikiSQL, TAPEX with continued SQL training scored 43.4%, 63.6%, 75.6%, and 88.1% at 16, 128, 1,024, and full data; OmniTab scored 63.6%, 75.6%, 82.9%, and 88.7%. The experiments used BART-large/TAPEX-large, approximately 0.5 million natural pairs and 0.5 million SQL-derived examples, and a learning rate of 2×10−52\times10^{-5}. Pretraining used batch size 512 for five epochs; fine-tuning used batch size 96 for 50 epochs.

  9. Knowl 9 — Ablations support dense retrieval, salient masking, and QA-based selection

    empirical result

    In WTQ 128-shot experiments, the natural-data retrieval variants scored 34.2% with a title-based heuristic, 35.5% with string matching that chose the sentence with the fewest mentions, 36.7% with string matching that chose the most mentions, and 36.4% with BM25. Dense retrieval scored 35.9%, 36.8%, and 38.4% at thresholds 0.50.5, 0.70.7, and 0.60.6, respectively. Thus, dense retrieval at 0.60.6 performed best among these tested choices, while its results varied with the threshold. With dense retrieval at 0.60.6, removing salient-mention masking reduced accuracy from 38.4% to 33.6%; removing random masking reduced it to 37.8%.

    For synthetic pretraining with 128 annotated SQL–question pairs, WTQ accuracy at 16, 128, and 1,024 QA fine-tuning examples was 26.5%, 35.0%, and 45.5% without self-training. Generation-probability selection scored 24.0%, 35.8%, and 44.3%; QA-verification selection using BART scored 28.9%, 36.4%, and 45.4%; and QA-verification selection using OmniTab scored 30.8%, 37.5%, and 47.3%. Selecting questions least likely to elicit the correct answer with OmniTab scored 15.5%, 27.4%, and 41.9%. The results show that QA-based verification with the stronger OmniTab verifier was the most effective tested self-training strategy.

  10. Knowl 10 — Natural and synthetic pretraining target different error patterns, and gains persist under topic shift

    empirical result

    An analysis of WTQ development examples where natural-data and synthetic-data models differed found 309 cases favoring natural pretraining and 315 favoring synthetic pretraining. The natural-favoring cases averaged 1.8 question tokens aligned with table content, compared with 1.0 in synthetic-favoring cases. The average question/SQL lengths were 10.6/11.4 tokens for natural-favoring cases and 10.8/12.9 for synthetic-favoring cases. Synthetic-favoring examples more often involved reasoning-rich SQL patterns, including nested filtering and counting, supporting the interpretation that natural data contributes more to cross-format alignment while synthetic data emphasizes reasoning.

    Under WTQ topical distribution shift, a model fine-tuned on 128 examples from one of five topics was tested on each of the five topics. OmniTab outperformed TAPEX in all 25 train/test combinations; the absolute accuracy gains ranged from 10.3 to 21.5 percentage points.

Coverage note — The exact 5-by-5 topical-shift gain matrix and detailed per-keyword frequency plot are omitted; their main findings are captured by the across-combination gain range and the qualitative alignment-versus-reasoning analysis.

References

  1. 1.Saneem A. Chemmengath, Vishwajeet Kumar, Samarth Bharadwaj, Jaydeep Sen, Mustafa Canim, Soumen Chakrabarti, Alfio Gliozzo, and Karthik Sankaranarayanan. 2021. Topic transferable table question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 4159–4172. Association for Computational Linguistics.
  2. 2.Xiang Deng, Ahmed Hassan Awadallah, Christopher Meek, Oleksandr Polozov, Huan Sun, and Matthew Richardson. 2021. Structure-grounded pretraining for text-to-sql. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 1337–1350. Association for Computational Linguistics.
  3. 3.Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2020. TURL: table understanding through representation learning. Proc. VLDB Endow., 14(3):307–319.
  4. 4.Angela Fan, Mike Lewis, and Yann N. Dauphin. 2018. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 889–898. Association for Computational Linguistics.
  5. 5.Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: revisit exact lexical match in information retrieval with contextualized inverted list. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 3030–3042. Association for Computational Linguistics.
  6. 6.Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020. REALM: retrieval-augmented language model pre-training. CoRR, abs/2002.08909.
  7. 7.Jonathan Herzig, Thomas Müller, Syrine Krichene, and Julian Eisenschlos. 2021. Open domain question answering over tables via dense retrieval. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 512–519. Association for Computational Linguistics.
  8. 8.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. 2020. Tapas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 4320–4333. Association for Computational Linguistics.
  9. 9.Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 1821–1831. Association for Computational Linguistics.
  10. 10.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick S. H. Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 6769–6781. Association for Computational Linguistics.
  11. 11.Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, SIGIR 2020, Virtual Event, China, July 25-30, 2020, pages 39–48. ACM.
  12. 12.Jayant Krishnamurthy, Pradeep Dasigi, and Matt Gardner. 2017. Neural semantic parsing with type constraints for semi-structured tables. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, pages 1516–1526. Association for Computational Linguistics.
  13. 13.Mike Lewis, Marjan Ghazvininejad, Gargi Ghosh, Armen Aghajanyan, Sida I. Wang, and Luke Zettlemoyer. 2020a. Pre-training via paraphrasing. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  14. 14.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020b. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7871–7880. Association for Computational Linguistics.
  15. 15.Chen Liang, Mohammad Norouzi, Jonathan Berant, Quoc V. Le, and Ni Lao. 2018. Memory augmented policy optimization for program synthesis and semantic parsing. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pages 10015–10027.
  16. 16.Qian Liu, Bei Chen, Jiaqi Guo, Zeqi Lin, and Jian-Guang Lou. 2021. TAPEX: table pre-training via learning a neural SQL executor. CoRR, abs/2107.07653.
  17. 17.Kaixin Ma, Hao Cheng, Xiaodong Liu, Eric Nyberg, and Jianfeng Gao. 2021. Open domain question answering over virtual documents: A unified approach for data and text. CoRR, abs/2110.08417.
  18. 18.Barlas Oguz, Xilun Chen, Vladimir Karpukhin, ˘ Stanislav Peshterliev, Dmytro Okhonko, M. Schlichtkrull, Sonal Gupta, Yashar Mehdad, and Scott Yih. 2020. Unik-qa: Unified representations of structured and unstructured knowledge for open-domain question answering.
  19. 19.Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, ACL 2015, July 26-31, 2015, Beijing, China, Volume 1: Long Papers, pages 1470–1480. The Association for Computer Linguistics.
  20. 20.Ori Ram, Yuval Kirstain, Jonathan Berant, Amir Globerson, and Omer Levy. 2021. Few-shot question answering by pretraining span selection. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 3066–3079. Association for Computational Linguistics.
  21. 21.Peng Shi, Patrick Ng, Zhiguo Wang, Henghui Zhu, Alexander Hanbo Li, Jun Wang, Cícero Nogueira dos Santos, and Bing Xiang. 2021. Learning contextual representations for semantic parsing with generation-augmented pre-training. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 13806–13814. AAAI Press.
  22. 22.Tianze Shi, Chen Zhao, Jordan L. Boyd-Graber, Hal Daumé III, and Lillian Lee. 2020. On the potential of lexico-logical alignments for semantic parsing to SQL queries. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, volume EMNLP 2020 of Findings of ACL, pages 1849–1864. Association for Computational Linguistics.
  23. 23.Bailin Wang, Ivan Titov, and Mirella Lapata. 2019. Learning semantic parsers from denotations with latent structured alignments and abstract programs. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 3772–3783. Association for Computational Linguistics.
  24. 24.Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang. 2021. TUTA: tree-based transformers for generally structured table pre-training. In KDD ’21: The 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, Singapore, August 14-18, 2021, pages 1780–1790. ACM.
  25. 25.Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. Tabert: Pretraining for joint understanding of textual and tabular data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 8413–8426. Association for Computational Linguistics.
  26. 26.Tao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir R. Radev, Richard Socher, and Caiming Xiong. 2021. Grappa: Grammar-augmented pre-training for table semantic parsing. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  27. 27.Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. CoRR, abs/1709.00103.

Citation

MLA
Jiang, Z., et al. “OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 932–42, https://doi.org/10.18653/v1/2022.naacl-main.68.
APA
Jiang, Z., Mao, Y., He, P., Neubig, G., & Chen, W. (2022). OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 932–942. https://doi.org/10.18653/v1/2022.naacl-main.68
Chicago
Jiang, Z., Y. Mao, P. He, G. Neubig, and W. Chen. 2022. “OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 932–42. https://doi.org/10.18653/v1/2022.naacl-main.68.
Harvard
Jiang, Z. et al. (2022) “OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 932–942. Available at: https://doi.org/10.18653/v1/2022.naacl-main.68.
Vancouver
1. Jiang Z, Mao Y, He P, Neubig G, Chen W (2022) OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 932–942

BibTeX

@inproceedings{jiang-etal-2022-omnitab,
    title = "{O}mni{T}ab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering",
    author = "Jiang, Zhengbao  and
      Mao, Yi  and
      He, Pengcheng  and
      Neubig, Graham  and
      Chen, Weizhu",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.68/",
    doi = "10.18653/v1/2022.naacl-main.68",
    pages = "932--942"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/