Text-to-Table: A New Way of Information Extraction
Xueqing WuJiacheng ZhangHang Li
Proposes a sequence-to-sequence formulation for end-to-end information extraction that automatically generates complex, multi-row tables directly from unstructured text without requiring predefined schemas.
Modern organizations face significant challenges in extracting structured information from unstructured, long-form documents. Traditional information extraction approaches rely heavily on predefined schemas and manual template designs, limiting their utility to narrow domains and simple outputs such as entity names or single relationship pairs.
The article introduces and evaluates a new task known as text-to-table, which transforms unstructured text directly into well-formulated data tables without requiring manually defined schemas. It evaluates whether sequence-to-sequence neural network models can effectively perform this end-to-end extraction across diverse domains.
The authors formulate table extraction as a sequence translation task by fine-tuning pre-trained language models. They enhance the baseline framework with two novel structural techniques: table constraints, which enforce consistent column counts across rows, and table relation embeddings, which maintain alignments between table cells and their corresponding row and column headers. The approach was evaluated on four public datasets covering sports game summaries (Rotowire), restaurant descriptions (E2E), and Wikipedia text (WikiTableText and WikiBio), comparing the model against standard named entity recognition and relation extraction systems.
The evaluation yielded several key findings. First, sequence-to-sequence language models consistently outperformed traditional relation extraction and entity recognition baselines, raising cell-level extraction accuracy by several percentage points across all datasets. Second, the proposed table constraints and relation embeddings eliminated structural table formatting errors, reducing the structural failure rate to zero compared to up to 7.4% in the unconstrained sequence baseline. Third, structural enhancements provided the greatest benefit on complex, multi-column tables, improving non-header cell accuracy from 81.96% to 82.53% on the complex sports dataset. Finally, scaling to a larger language model improved overall performance across all datasets, achieving up to 86.83% cell accuracy on challenging sports player tables.
These results demonstrate that end-to-end table generation provides a scalable, automated alternative to labor-intensive schema creation, reducing the setup costs and engineering timelines required for enterprise document extraction. While standard models struggle with table integrity on large structures, adding structural constraints and header alignments ensures reliable, well-formed tabular outputs for downstream analytics and knowledge graphs.
For operational deployment, organizations should adopt structurally constrained sequence-to-sequence models when converting complex documents into tabular formats, utilizing larger pre-trained architectures where computational budgets allow. When dealing with simpler two-column tables, standard sequence models remain sufficient.
Decision-makers should note that the current findings are bounded by existing parallel text-and-table datasets, where tables are largely extracted from texts of limited length. Performance remains lower on tasks that demand open-domain background knowledge, multi-sentence reasoning, and heavy filtering of redundant text, indicating that additional research is required before fully autonomous extraction can be deployed on complex analytical domains.
- Paper: Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, Colin Raffel et al. (2020). It establishes the unified sequence-to-sequence text-to-text pre-training framework that forms the architectural foundation for generative structured extraction.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). It introduces foundational deep bidirectional transformer representations and fine-tuning paradigms that serve as the baseline for neural information extraction.
- Paper: Pre-trained models for natural language processing: A survey, Xipeng Qiu et al. (2020). It provides a comprehensive survey of pre-trained language model architectures and adaptation mechanisms essential for understanding modern generative NLP pipelines.
- Paper: Recent Advances in Natural Language Processing via Large Pre-trained Language Models: A Survey, Bonan Min et al. (2021). It details advances in casting diverse natural language processing and structured extraction tasks as generative sequence problems using pre-trained language models.
- Paper: A Survey on Deep Learning for Named Entity Recognition, Jing Li et al. (2018). It surveys traditional deep learning approaches for named entity recognition and information extraction against which end-to-end text-to-table translation is benchmarked.
- Paper: Generative Knowledge Graph Construction: A Review, Hongbin Ye et al. (2022). It synthesizes generative end-to-end information extraction paradigms across structured knowledge graph and relational table construction.
- Paper: Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning, Saibo Geng et al. (2023). It generalizes decoding constraints by using formal grammars to enforce syntactically valid structured outputs from generative models without fine-tuning.
- Paper: MultiTabQA: Generating Tabular Answers for Multi-Table Question Answering, Vaishali Pal et al. (2023). It extends generative tabular sequence-to-sequence modeling to the multi-table question answering domain to directly synthesize complex structured answer tables.
- Paper: HyTrel: Hypergraph-enhanced Tabular Data Representation Learning, Pei Chen et al. (2023). It introduces hypergraph-enhanced structural representations to improve tabular understanding and model invariance beyond linear sequence translation.
- Paper: Rethinking Tabular Data Understanding with Large Language Models, Tianyang Liu et al. (2024). It explores how structural perturbations affect tabular reasoning in large language models, providing crucial insights into model robustness over generated tables.
- Paper: OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition, Jianqiang Wan et al. (2024). It expands end-to-end table extraction to multimodal document images by unifying text spotting, key information extraction, and table recognition.
- Paper: Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph Construction, Bowen Zhang et al. (2024). It applies large language models to open schema-free extraction and post-processing canonicalization for automated structured knowledge graph construction.
