OmniTab: Pretraining with Natural and Synthetic Data for Few-shot Table-based Question Answering
Zhengbao JiangYi MaoPengcheng HeGraham NeubigWeizhu Chen
Proposes a pretraining framework that pairs natural text-table alignment with synthetic SQL-derived questions, establishing a new state-of-the-art on WikiTableQuestions while dramatically reducing annotation requirements for few-shot table question answering.
Tables present vast amounts of critical data across web pages, business reports, and scientific literature, but extracting insights from them remains challenging. Effective table-based question answering systems must connect free-form natural language questions to tabular structures and execute complex multi-step reasoning, such as aggregation, sorting, and comparison. Building such systems has historically required complex multi-stage architectures and extensive manual data annotation, creating high development costs and deployment friction.
The article demonstrates an end-to-end table question answering system, named OmniTab, designed to answer complex tabular questions with minimal human annotation. The primary objective is to evaluate whether combining natural text retrieved from the web with programmatically generated synthetic data can effectively train models under low-resource and few-shot conditions.
The researchers developed an omnivorous pretraining framework utilizing two complementary data streams derived from approximately 500,000 Wikipedia tables. The first stream retrieves relevant sentences from Wikipedia documents and masks cell mentions to teach the model how to align natural phrasing with table contents. The second stream converts structured database queries into synthetic natural language questions using a query-to-text model refined by verification-based self-training, which teaches the model formal reasoning skills. The combined system was evaluated across standard benchmarks, including WikiTableQuestions and WikiSQL, primarily focusing on few-shot scenarios ranging from 16 to 1,024 labeled examples.
The experimental findings show substantial performance gains over existing baseline methods. When trained on only 128 labeled examples, OmniTab achieved a 41.4% accuracy on WikiTableQuestions, outperforming the best baseline by an absolute margin of 16.2 percentage points (a relative improvement of approximately 64%). Using natural or synthetic data independently in this 128-shot setting improved performance by 13.2% and 12.3% respectively, confirming that the two data types provide distinct, complementary benefits. Even in full-data settings with all 11,000 training examples, OmniTab achieved a state-of-the-art accuracy of 62.8%, improving by 2.7 percentage points over prior models. Furthermore, cross-topic evaluations demonstrated that the model maintains robust performance when transferring across different subject domains, including sports, politics, culture, and people.
These results indicate that organizations can build high-performing tabular question answering systems without investing in large, costly manual annotation efforts. By pairing dense retrieval mechanisms with verified synthetic question generation, developers can bridge the gap between human language variation and structured multi-step reasoning. This capability substantially reduces the time and expense required to deploy natural language interfaces over enterprise databases and tabular records.
Organizations aiming to build tabular question answering capabilities should adopt hybrid pretraining strategies that combine natural retrieval with synthetic reasoning data. Teams should prioritize salient mention masking over random token masking, and implement verification-based filtering when generating synthetic training data. Because dense retrieval depends heavily on similarity thresholds and the synthetic generation model still exhibits a performance gap compared to fully supervised data, future initiatives should conduct small-scale pilot testing and investigate improved retrieval and synthesis architectures before full-scale deployment.
- Paper: Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning, Victor Zhong et al. (2017). Seq2SQL introduces the WikiSQL benchmark and an early neural approach to translating natural-language questions into executable table queries, giving essential context for OmniTab’s WikiSQL evaluation.
- Paper: RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations, Yilun Zhao et al. (2023). RobuT directly evaluates OmniTab under realistic table and question perturbations, extending its benchmark results by showing where the model’s few-shot performance remains vulnerable.
