Text-to-Table: A New Way of Information Extraction

Xueqing WuJiacheng ZhangHang Li

article2022ACL91 citations

Proposes a sequence-to-sequence formulation for end-to-end information extraction that automatically generates complex, multi-row tables directly from unstructured text without requiring predefined schemas.

Listen

Modern organizations face significant challenges in extracting structured information from unstructured, long-form documents. Traditional information extraction approaches rely heavily on predefined schemas and manual template designs, limiting their utility to narrow domains and simple outputs such as entity names or single relationship pairs.

The article introduces and evaluates a new task known as text-to-table, which transforms unstructured text directly into well-formulated data tables without requiring manually defined schemas. It evaluates whether sequence-to-sequence neural network models can effectively perform this end-to-end extraction across diverse domains.

The authors formulate table extraction as a sequence translation task by fine-tuning pre-trained language models. They enhance the baseline framework with two novel structural techniques: table constraints, which enforce consistent column counts across rows, and table relation embeddings, which maintain alignments between table cells and their corresponding row and column headers. The approach was evaluated on four public datasets covering sports game summaries (Rotowire), restaurant descriptions (E2E), and Wikipedia text (WikiTableText and WikiBio), comparing the model against standard named entity recognition and relation extraction systems.

The evaluation yielded several key findings. First, sequence-to-sequence language models consistently outperformed traditional relation extraction and entity recognition baselines, raising cell-level extraction accuracy by several percentage points across all datasets. Second, the proposed table constraints and relation embeddings eliminated structural table formatting errors, reducing the structural failure rate to zero compared to up to 7.4% in the unconstrained sequence baseline. Third, structural enhancements provided the greatest benefit on complex, multi-column tables, improving non-header cell accuracy from 81.96% to 82.53% on the complex sports dataset. Finally, scaling to a larger language model improved overall performance across all datasets, achieving up to 86.83% cell accuracy on challenging sports player tables.

These results demonstrate that end-to-end table generation provides a scalable, automated alternative to labor-intensive schema creation, reducing the setup costs and engineering timelines required for enterprise document extraction. While standard models struggle with table integrity on large structures, adding structural constraints and header alignments ensures reliable, well-formed tabular outputs for downstream analytics and knowledge graphs.

For operational deployment, organizations should adopt structurally constrained sequence-to-sequence models when converting complex documents into tabular formats, utilizing larger pre-trained architectures where computational budgets allow. When dealing with simpler two-column tables, standard sequence models remain sufficient.

Decision-makers should note that the current findings are bounded by existing parallel text-and-table datasets, where tables are largely extracted from texts of limited length. Performance remains lower on tasks that demand open-domain background knowledge, multi-sentence reasoning, and heavy filtering of redundant text, indicating that additional research is required before fully autonomous extraction can be deployed on complex analytical domains.

arXiv: 2109.02707
Cover for Text-to-Table: A New Way of Information Extraction

Abstract

We study a new problem setting of information extraction (IE), referred to as text-to-table. In text-to-table, given a text, one creates a table or several tables expressing the main content of the text, while the model is learned from text-table pair data. The problem setting differs from those of the existing methods for IE. First, the extraction can be carried out from long texts to large tables with complex structures. Second, the extraction is entirely data-driven, and there is no need to explicitly define the schemas. As far as we know, there has been no previous work that studies the problem. In this work, we formalize text-to-table as a sequence-to-sequence (seq2seq) problem. We first employ a seq2seq model fine-tuned from a pre-trained language model to perform the task. We also develop a new method within the seq2seq approach, exploiting two additional techniques in table generation: table constraint and table relation embeddings. We consider text-to-table as an inverse problem of the well-studied table-to-text, and make use of four existing table-to-text datasets in our experiments on text-to-table. Experimental results show that the vanilla seq2seq model can outperform the baseline methods of using relation extraction and named entity extraction. The results also show that our method can further boost the performances of the vanilla seq2seq model. We further discuss the main challenges of the proposed task. The code and data are available at https://github.com/shirley-wu/text_to_table.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Formulation
  • 4 Our Method
  • 4.1 Vanilla Seq2Seq
  • 4.2 Techniques
  • Table Constraint
  • Table Relation Embeddings
  • 5 Experiments
  • 5.1 Datasets
  • 5.2 Procedure
  • 5.3 Results on Rotowire
  • 5.4 Results on E2E, WikiTableText and WikiBio
  • 5.5 Additional Study
  • 5.6 Discussions
  • 6 Conclusion
  • References
  • A Hyper-parameters
  • B Table Constraint Algorithm
  • C Our Method with Multiple Tables
  • D Information Extraction Baselines
  • D.1 Relation Extraction
  • D.2 Named Entity Recognition
  • E Detailed Cases for Challenges

Knowls

  1. Knowl 1 — Task Formulation and Sequence Representation for Text-to-Table Extraction

    definition

    Text-to-table is an information extraction framework where a model converts unstructured text x=(x1,…,x∣x∣)x = (x_1, \dots, x_{|x|}) into one or more structured tables TT without explicit, predefined extraction schemas.

    For a target table TT containing nrn_r rows and ncn_c columns, cell (i,j)(i, j) contains a sequence of words ti,jt_{i,j}. A table can contain column headers t1,jt_{1,j} (j=2,…,ncj=2,\dots,n_c), row headers ti,1t_{i,1} (i=2,…,nri=2,\dots,n_r), and non-header content cells ti,jt_{i,j} (i,j≥2i,j \ge 2).

    To formalize this task within a sequence-to-sequence framework, special tokens ⟨s⟩\langle s \rangle (cell separator) and ⟨n⟩\langle n \rangle (new-line delimiter) are introduced. A row tit_i is serialized as: ti=⟨s⟩,ti,1,⟨s⟩,…,⟨s⟩,ti,nc,⟨s⟩t_i = \langle s \rangle, t_{i,1}, \langle s \rangle, \dots, \langle s \rangle, t_{i,n_c}, \langle s \rangle

    The full table sequence tt is serialized by concatenating rows with new-line tokens: t=t1,⟨n⟩,t2,⟨n⟩,…,⟨n⟩,tnrt = t_1, \langle n \rangle, t_2, \langle n \rangle, \dots, \langle n \rangle, t_{n_r}

    When multiple tables are generated from a single text, table captions act as delimiters between the serialized tables: Caption1⟨n⟩t(1)⟨n⟩Caption2⟨n⟩t(2)…\text{Caption}_1 \langle n \rangle t^{(1)} \langle n \rangle \text{Caption}_2 \langle n \rangle t^{(2)} \dots

  2. Knowl 2 — Table Relation Embeddings in Transformer Decoder Self-Attention

    model/method

    To enhance alignment between table content cells and their corresponding row and column headers, Table Relation Embeddings (TRE) modify the self-attention mechanism within the Transformer decoder.

    In standard decoder self-attention at step tt over representations z=(z1,…,zt)z = (z_1, \dots, z_t) with zi∈Rdz_i \in \mathbb{R}^d, query, key, and value projection matrices WQ,WK,WV∈Rd×dkW^Q, W^K, W^V \in \mathbb{R}^{d \times d_k}, and output projection WO∈Rdk×dW^O \in \mathbb{R}^{d_k \times d}: hi=(∑j=1iαij(zjWV))WOh_i = \left( \sum_{j=1}^i \alpha_{ij} (z_j W^V) \right) W^O αij=exp⁡(eij)∑j′=1iexp⁡(eij′),eij=(ziWQ)(zjWK)Tdk\alpha_{ij} = \frac{\exp(e_{ij})}{\sum_{j'=1}^i \exp(e_{ij'})}, \quad e_{ij} = \frac{(z_i W^Q)(z_j W^K)^T}{\sqrt{d_k}}

    With Table Relation Embeddings, learnable relation vectors rijK,rijV∈Rdkr_{ij}^K, r_{ij}^V \in \mathbb{R}^{d_k} are incorporated into the attention scores and value representations: hi=(∑j=1iαij(zjWV+rijV))WOh_i = \left( \sum_{j=1}^i \alpha_{ij} (z_j W^V + r_{ij}^V) \right) W^O eij=(ziWQ)(zjWK+rijK)Tdke_{ij} = \frac{(z_i W^Q)(z_j W^K + r_{ij}^K)^T}{\sqrt{d_k}}

    For a token at output position ii in a non-header cell:

    • If token position jj is part of its row header, rijK=τrKr_{ij}^K = \tau_r^K and rijV=τrVr_{ij}^V = \tau_r^V (learned row relation embeddings).
    • If token position jj is part of its column header, rijK=τcKr_{ij}^K = \tau_c^K and rijV=τcVr_{ij}^V = \tau_c^V (learned column relation embeddings).
    • Otherwise, rijK=0r_{ij}^K = 0 and rijV=0r_{ij}^V = 0.

    During autoregressive inference, the row and column headers for position ii are identified dynamically by parsing the generated prefix using ⟨s⟩\langle s \rangle and ⟨n⟩\langle n \rangle tokens. Relation parameters are shared across multiple tables.

  3. Knowl 3 — Table Constraint Decoding Algorithm

    algorithm

    Table Constraint (TC) enforces the structural consistency of generated tables during autoregressive decoding by masking logits so that every generated row contains the exact number of cells established by the first row.

    Input: Input text tokens x=[x1,…,x∣x∣]x = [x_1, \dots, x_{|x|}], seq2seq model seq2seq\text{seq2seq}, decode function Decode\text{Decode}
    Output: Generated token sequence y=[y1,…,y∣y∣]y = [y_1, \dots, y_{|y|}]
    y←[]y \leftarrow []
    repeat
        p(⋅)←seq2seq(x,y)p(\cdot) \leftarrow \text{seq2seq}(x, y)
        if last token y∣y∣≠⟨s⟩y_{|y|} \neq \langle s \rangle then
            p(⟨n⟩)←0p(\langle n \rangle) \leftarrow 0
            p(⟨eos⟩)←0p(\langle eos \rangle) \leftarrow 0
        y.append(Decode(p))y\text{.append}(\text{Decode}(p))
    until y∣y∣=⟨n⟩y_{|y|} = \langle n \rangle or y∣y∣=⟨eos⟩y_{|y|} = \langle eos \rangle
    if y∣y∣=⟨eos⟩y_{|y|} = \langle eos \rangle then
        return yy
    nc←number of cells in the first rown_c \leftarrow \text{number of cells in the first row}
    repeat
        repeat
            p(⋅)←seq2seq(x,y)p(\cdot) \leftarrow \text{seq2seq}(x, y)
            if current row has ncn_c columns then
                p(t)←0,∀t∉{⟨eos⟩,⟨n⟩}p(t) \leftarrow 0, \forall t \notin \{\langle eos \rangle, \langle n \rangle\}
            else
                p(⟨n⟩)←0p(\langle n \rangle) \leftarrow 0
                p(⟨eos⟩)←0p(\langle eos \rangle) \leftarrow 0
            y.append(Decode(p))y\text{.append}(\text{Decode}(p))
        until y∣y∣=⟨n⟩y_{|y|} = \langle n \rangle or y∣y∣=⟨eos⟩y_{|y|} = \langle eos \rangle
    until y∣y∣=⟨eos⟩y_{|y|} = \langle eos \rangle
    return yy

    When generating multi-table outputs with captions, constraints are only applied when the current line begins with a cell separator ⟨s⟩\langle s \rangle, allowing caption text to generate without column constraints.

  4. Knowl 4 — Set-Based Evaluation Metrics for Text-to-Table Generation

    equation

    Text-to-table extraction accuracy is evaluated independently of row and column ordering by computing set-level precision, recall, and F1 over generated table elements. For a predicted element set Y\mathcal{Y} and a ground-truth set Y∗\mathcal{Y}^*:

    P=1∣Y∣∑y∈Ymax⁡y∗∈Y∗O(y,y∗),R=1∣Y∗∣∑y∗∈Y∗max⁡y∈YO(y,y∗),F1=21P+1RP = \frac{1}{|\mathcal{Y}|} \sum_{y \in \mathcal{Y}} \max_{y^* \in \mathcal{Y}^*} O(y, y^*), \quad R = \frac{1}{|\mathcal{Y}^*|} \sum_{y^* \in \mathcal{Y}^*} \max_{y \in \mathcal{Y}} O(y, y^*), \quad F_1 = \frac{2}{\frac{1}{P} + \frac{1}{R}}

    where O(y,y∗)∈[0,1]O(y, y^*) \in [0, 1] represents string similarity calculated via exact match, chrF (character nn-gram F-score), or rescaled BERTScore.

    For a non-header cell cc, similarity incorporates alignment with its row header hr(c)h_r(c) and column header hc(c)h_c(c): O(c,c∗)=O(hr(c),hr(c∗))×O(hc(c),hc(c∗))×O(content(c),content(c∗))O(c, c^*) = O(h_r(c), h_r(c^*)) \times O(h_c(c), h_c(c^*)) \times O(\text{content}(c), \text{content}(c^*))

    For tables lacking column headers, only the row header factor is retained. Empty cells are omitted from scoring. The error rate measures the percentage of generated outputs that cannot be decoded into rectangular tables.

  5. Knowl 5 — Text-to-Table Benchmark Datasets

    experimental setup

    Four data-to-text benchmark datasets were adapted for the reverse text-to-table extraction task by filtering out table cells whose content does not appear in the input text:

    1. Rotowire: Long sports reports about basketball games mapping to two tables (Team scores and Player statistics) containing row and column headers.
    2. E2E: Short restaurant descriptions mapping to two-column tables with row headers.
    3. WikiTableText: Open-domain short descriptions from Wikipedia mapping to two-column tables with row headers.
    4. WikiBio: Wikipedia biography introductions mapping to two-column infobox tables with row headers.
    Dataset Train Valid Test # Tokens # Rows # Columns
    Rotowire 3.4k 727 728 351.05 Team: 2.71 / Player: 7.26 Team: 4.84 / Player: 8.75
    E2E 42.1k 4.7k 4.7k 24.90 4.58 2.00
    WikiTableText 10.0k 1.3k 2.0k 19.59 4.26 2.00
    WikiBio 582.7k 72.8k 72.7k 122.30 4.20 2.00

    In Rotowire, non-empty cells average 6.56 (85.40% non-empty) for Team tables and 22.63 (43.93% non-empty) for Player tables.

  6. Knowl 6 — Extraction Performance of Text-to-Table Models vs. IE Baselines

    empirical result

    Vanilla sequence-to-sequence (seq2seq) and the seq2seq model with Table Constraint and Table Relation Embeddings (Our Method) fine-tuned on BART-base outperform standard Information Extraction baselines based on BERT-base (Named Entity Recognition for two-column datasets and PURE Relation Extraction for Rotowire).

    Dataset Model Row Header F1 Non-header Cell F1 Err. Rate
    Exact ChrF BERT Exact ChrF BERT (%)
    Rotowire (Team) Sent-level RE 85.28 87.12 93.65 77.17 79.10 87.48 0.00
    Doc-level RE 84.90 86.73 93.44 75.66 77.89 87.82 0.00
    Vanilla seq2seq 94.71 94.93 97.35 82.97 84.43 90.62 0.49
    Our Method 94.97 95.20 97.51 83.36 84.76 90.80 0.00
    Rotowire (Player) Sent-level RE 89.05 93.00 90.98 79.59 83.42 85.35 0.00
    Doc-level RE 89.26 93.28 91.19 80.76 84.64 86.50 0.00
    Vanilla seq2seq 92.16 93.89 93.60 81.96 84.19 88.66 7.40
    Our Method 92.31 94.00 93.71 82.53 84.74 88.97 0.00
    E2E NER 91.23 92.40 95.34 90.80 90.97 92.20 0.00
    Vanilla seq2seq 99.62 99.69 99.88 97.87 97.99 98.56 0.00
    Our Method 99.63 99.69 99.88 97.88 98.00 98.57 0.00
    WikiTableText NER 59.72 70.98 94.36 52.23 59.62 73.40 0.00
    Vanilla seq2seq 78.15 84.00 95.60 59.26 69.12 80.69 0.41
    Our Method 78.16 83.96 95.68 59.14 68.95 80.74 0.00
    WikiBio NER 63.99 71.19 81.03 56.51 62.52 61.95 0.00
    Vanilla seq2seq 80.53 84.98 92.61 68.98 77.16 76.54 0.00
    Our Method 80.52 84.96 92.60 69.02 77.16 76.56 0.00

    The seq2seq approach outperforms the pipeline NER and RE baselines across all four datasets. On the Rotowire player tables (featuring 8.75 average columns), vanilla seq2seq exhibits a 7.40% table formatting error rate, whereas Our Method achieves a 0.00% error rate.

  7. Knowl 7 — Ablation Study on Pre-training, Table Constraint, and Relation Embeddings

    empirical result

    An ablation study evaluated the relative contributions of pre-trained language modeling (Pre), Table Constraint (TC), and Table Relation Embeddings (TRE) using Exact Match F1 for non-header cells:

    Pre TC TRE Rotowire/Team Rotowire/Player E2E WikiTableText WikiBio
    28.05 7.75 94.45 46.37 67.51
    ✓ ✓ 30.61 10.67 95.53 47.13 67.43
    ✓ 82.97 81.96 97.87 59.26 68.98
    ✓ ✓ 83.09 82.24 97.88 59.29 68.98
    ✓ ✓ 83.30 82.50 97.87 59.12 69.02
    ✓ ✓ ✓ 83.36 82.53 97.88 59.14 69.02

    Key findings:

    • Pre-training provides the largest overall performance boost across all datasets (e.g. improving Rotowire/Player F1 from 7.75 to 81.96).
    • TC and TRE yield statistically significant improvements on multi-column structured tables (Rotowire Team and Player datasets, p<0.01p < 0.01 for the complete method vs vanilla seq2seq on Player).
    • On two-column datasets (E2E, WikiTableText, WikiBio), formatting is straightforward, resulting in minimal differences from TC or TRE additions.
  8. Knowl 8 — Effect of Language Model Parameter Scale on Text-to-Table Performance

    empirical result

    Evaluating vanilla seq2seq and the proposed method with BART-base (139M parameters, 12 layers) versus BART-large (406M parameters, 24 layers) reveals consistent performance gains across all benchmarks:

    Method Rotowire/Team Rotowire/Player E2E WikiTableText WikiBio
    Vanilla seq2seq (BART base) 82.97 81.96 97.87 59.26 68.98
    Our method (BART base) 83.36 82.53 97.88 59.14 69.02
    Vanilla seq2seq (BART large) 86.31 86.59 97.94 62.71 69.66
    Our method (BART large) 86.31 86.83 97.90 62.41 69.71

    Scaling to BART-large produces the largest benefits on complex multi-column documents (Rotowire Player non-header cell exact-match F1 rises from 82.53% to 86.83%) and open-domain text (WikiTableText F1 rises from 59.14% to 62.41%).

  9. Knowl 9 — Core Challenges and Limitations in Text-to-Table Information Extraction

    limitation

    Error analysis across the four benchmark datasets identifies five principal challenges in text-to-table extraction:

    1. Text Diversity: Expressing entities through synonyms or metonyms (e.g. referring to the "Knicks" as "New York" or "Celtics" as "Boston") requires implicit coreference and entity normalization.
    2. Text Redundancy: Long source documents containing mostly irrelevant narrative (e.g. Wikipedia biography texts) require summarization and selective filtering to extract concise tables.
    3. Large Multi-Column Tables: Generation of tables with many columns requires maintaining alignment across long sequences; without structural constraints, autoregressive decoders frequently misalign or omit columns.
    4. Background Knowledge: Open-domain entity attributes (such as identifying French electoral constituencies) require world knowledge not present in the input text.
    5. Numerical and Relational Reasoning: Information implicit in text requires mathematical inference rather than span copying (e.g. deducing that an opening lead of 31-14 implies the opponent's score is 14).

Coverage note — None; all primary technical contributions, formulations, algorithms, evaluation metrics, empirical results, ablations, and limitation analyses from the paper are represented.

References

  1. 1.Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  2. 2.Michele Banko, Michael J. Cafarella, Stephen Soderland, Matthew Broadhead, and Oren Etzioni. 2007. Open information extraction from the web. In IJCAI 2007, Proceedings of the 20th International Joint Conference on Artificial Intelligence, Hyderabad, India, January 6-12, 2007, pages 2670–2676.
  3. 3.Junwei Bao, Duyu Tang, Nan Duan, Zhao Yan, Yuanhua Lv, Ming Zhou, and Tiejun Zhao. 2018. Table-to-text: Describing table region with natural language. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32.
  4. 4.Lingzhen Chen and Alessandro Moschitti. 2018. Learning to progressively recognize new named entities with sequence to sequence models. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2181–2191.
  5. 5.Wenhu Chen, Jianshu Chen, Yu Su, Zhiyu Chen, and William Yang Wang. 2020. Logical natural language generation from open-domain tables. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7929–7942.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186.
  7. 7.Xinya Du, Alexander M. Rush, and Claire Cardie. 2021. GRIT: generative role-filler transformers for document-level event entity extraction. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, EACL 2021, Online, April 19 - 23, 2021, pages 634–644. Association for Computational Linguistics.
  8. 8.Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017. Creating training corpora for nlg micro-planning. In 55th annual meeting of the Association for Computational Linguistics (ACL).
  9. 9.Li Gong, Josep M Crego, and Jean Senellart. 2019. Enhanced transformer model for data-to-text generation. In Proceedings of the 3rd Workshop on Neural Generation and Translation, pages 148–156.
  10. 10.Anwen Hu, Zhicheng Dou, Jian-Yun Nie, and Ji-Rong Wen. 2020. Leveraging multi-token entities in document-level named entity recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7961–7968.
  11. 11.Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao, Kun Wang, Jun Xie, and Yue Zhang. 2020. What have we achieved on text summarization? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 446–469.
  12. 12.Kung-Hsiang Huang, Sam Tang, and Nanyun Peng. 2021. Document-level entity-based extraction as template generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 5257–5269. Association for Computational Linguistics.
  13. 13.Zhiheng Huang, Wei Xu, and Kai Yu. 2015. Bidirectional lstm-crf models for sequence tagging. arXiv preprint arXiv:1508.01991.
  14. 14.Juraj Juraska, Panagiotis Karagiannis, Kevin Bowden, and Marilyn Walker. 2018. A deep ensemble model with slot alignment for sequence-to-sequence natural language generation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 152–162.
  15. 15.Rik Koncel-Kedziorski, Dhanush Bekal, Yi Luan, Mirella Lapata, and Hannaneh Hajishirzi. 2019. Text generation from knowledge graphs with graph transformers. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2284–2293.
  16. 16.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 260–270.
  17. 17.Rémi Lebret, David Grangier, and Michael Auli. 2016. Neural text generation from structured data with application to the biography domain. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1203–1213.
  18. 18.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7871–7880. Association for Computational Linguistics.
  19. 19.Sha Li, Heng Ji, and Jiawei Han. 2021. Document-level event argument extraction by conditional generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 894–908. Association for Computational Linguistics.
  20. 20.Ying Lin, Heng Ji, Fei Huang, and Lingfei Wu. 2020. A joint neural model for information extraction with global features. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7999–8009.
  21. 21.Tianyu Liu, Kexiang Wang, Lei Sha, Baobao Chang, and Zhifang Sui. 2018. Table-to-text generation by structure-aware seq2seq learning. In Thirty-Second AAAI Conference on Artificial Intelligence.
  22. 22.Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. Multilingual denoising pre-training for neural machine translation. Transactions of the Association for Computational Linguistics, 8:726–742.
  23. 23.Yaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han, Jialong Tang, Annan Li, Le Sun, Meng Liao, and Shaoyi Chen. 2021. Text2event: Controllable sequence-to-structure generation for end-to-end event extraction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 2795–2806. Association for Computational Linguistics.
  24. 24.Yi Luan, Dave Wadden, Luheng He, Amy Shah, Mari Ostendorf, and Hannaneh Hajishirzi. 2019. A general framework for information extraction using dynamic span graphs. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3036–3046.
  25. 25.Ling Luo, Zhihao Yang, Pei Yang, Yin Zhang, Lei Wang, Hongfei Lin, and Jian Wang. 2018. An attention-based bilstm-crf approach to document-level chemical named entity recognition. Bioinformatics, 34(8):1381–1388.
  26. 26.Xuezhe Ma and Eduard Hovy. 2016. End-to-end sequence labeling via bi-directional lstm-cnns-crf. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1064–1074.
  27. 27.Diego Marcheggiani and Laura Perez-Beltrachini. 2018. Deep graph convolutional encoders for structured data to text generation. In Proceedings of the 11th International Conference on Natural Language Generation, pages 1–9.
  28. 28.Mausam, Michael Schmitz, Stephen Soderland, Robert Bart, and Oren Etzioni. 2012. Open language learning for information extraction. In Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, EMNLP-CoNLL 2012, July 12-14, 2012, Jeju Island, Korea, pages 523–534. ACL.
  29. 29.Guoshun Nan, Zhijiang Guo, Ivan Sekulic, and Wei Lu. 2020. Reasoning with latent structure refinement for document-level relation extraction. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1546–1557.
  30. 30.Linyong Nan, Dragomir R. Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani. 2021. DART: open-domain structured data record to text generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 432–447. Association for Computational Linguistics.
  31. 31.Tapas Nayak and Hwee Tou Ng. 2020. Effective modeling of encoder-decoder architecture for joint entity and relation extraction. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 8528–8535.
  32. 32.Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017. The e2e dataset: New challenges for end-to-end generation. In Proceedings of the 18th Annual SIGdial Meeting on Discourse and Dialogue, pages 201–206.
  33. 33.Giovanni Paolini, Ben Athiwaratkun, Jason Krone, Jie Ma, Alessandro Achille, Rishita Anubhai, Cícero Nogueira dos Santos, Bing Xiang, and Stefano Soatto. 2021. Structured prediction as translation between augmented natural languages. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  34. 34.Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. Totto: A controlled table-to-text generation dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1173–1186.
  35. 35.Maja Popovic. 2015. chrf: character n-gram f-score for automatic MT evaluation. In Proceedings of the Tenth Workshop on Statistical Machine Translation, WMT@EMNLP 2015, 17-18 September 2015, Lisbon, Portugal, pages 392–395. The Association for Computer Linguistics.
  36. 36.Ratish Puduppully, Li Dong, and Mirella Lapata. 2019a. Data-to-text generation with content selection and planning. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 6908–6915.
  37. 37.Ratish Puduppully, Li Dong, and Mirella Lapata. 2019b. Data-to-text generation with entity modeling. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2023–2035.
  38. 38.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.
  39. 39.Xiaoyu Shen, Ernie Chang, Hui Su, Cheng Niu, and Dietrich Klakow. 2020. Neural data-to-text generation via jointly learning the segmentation and correspondence. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7155–7165.
  40. 40.Gabriel Stanovsky, Julian Michael, Luke Zettlemoyer, and Ido Dagan. 2018. Supervised open information extraction. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 885–895. Association for Computational Linguistics.
  41. 41.Emma Strubell, Patrick Verga, David Belanger, and Andrew McCallum. 2017. Fast and accurate entity recognition with iterated dilated convolutions. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2670–2680.
  42. 42.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. Advances in Neural Information Processing Systems, 27:3104–3112.
  43. 43.Craig Thomson, Ehud Reiter, and Somayajulu Sripada. 2020. Sportsett: Basketball-a robust and maintainable data-set for natural language generation. In Proceedings of the Workshop on Intelligent Information Processing and Natural Language Generation, pages 32–40.
  44. 44.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems, 30:5998–6008.
  45. 45.David Wadden, Ulme Wennberg, Yi Luan, and Hannaneh Hajishirzi. 2019. Entity, relation, and event extraction with contextualized span representations. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5784–5789.
  46. 46.Sam Wiseman, Stuart M Shieber, and Alexander M Rush. 2017. Challenges in data-to-document generation. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2253–2263.
  47. 47.Fei Wu and Daniel S. Weld. 2010. Open information extraction using wikipedia. In ACL 2010, Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, July 11-16, 2010, Uppsala, Sweden, pages 118–127. The Association for Computer Linguistics.
  48. 48.Guohai Xu, Chengyu Wang, and Xiaofeng He. 2018. Improving clinical named entity recognition with global neural attention. In Asia-Pacific Web (APWeb) and Web-Age Information Management (WAIM) Joint International Conference on Web and Big Data, pages 264–279. Springer.
  49. 49.Hang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, and Xipeng Qiu. 2021. A unified generative framework for various NER subtasks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 5808–5822. Association for Computational Linguistics.
  50. 50.Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, and Maosong Sun. 2019. Docred: A large-scale document-level relation extraction dataset. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 764–777.
  51. 51.Xiangrong Zeng, Daojian Zeng, Shizhu He, Kang Liu, and Jun Zhao. 2018. Extracting relational facts by an end-to-end neural model with copy mechanism. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 506–514.
  52. 52.Junlang Zhan and Hai Zhao. 2020. Span model for open information extraction on accurate corpus. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 9523–9530. AAAI Press.
  53. 53.Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. Bertscore: Evaluating text generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  54. 54.Tongtao Zhang, Heng Ji, and Avirup Sil. 2019. Joint entity and event extraction with generative adversarial imitation learning. Data Intelligence, 1(2):99–120.
  55. 55.Suncong Zheng, Feng Wang, Hongyun Bao, Yuexing Hao, Peng Zhou, and Bo Xu. 2017. Joint extraction of entities and relations based on a novel tagging scheme. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1227–1236.
  56. 56.Zexuan Zhong and Danqi Chen. 2021. A frustratingly easy approach for entity and relation extraction. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 50–61, Online. Association for Computational Linguistics.

Citation

MLA
Wu, X., et al. “Text-to-Table: A New Way of Information Extraction”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 2518–33, https://doi.org/10.18653/v1/2022.acl-long.180.
APA
Wu, X., Zhang, J., & Li, H. (2022). Text-to-Table: A New Way of Information Extraction. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2518–2533. https://doi.org/10.18653/v1/2022.acl-long.180
Chicago
Wu, X., J. Zhang, and H. Li. 2022. “Text-to-Table: A New Way of Information Extraction”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2518–33. https://doi.org/10.18653/v1/2022.acl-long.180.
Harvard
Wu, X., Zhang, J. and Li, H. (2022) “Text-to-Table: A New Way of Information Extraction”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2518–2533. Available at: https://doi.org/10.18653/v1/2022.acl-long.180.
Vancouver
1. Wu X, Zhang J, Li H (2022) Text-to-Table: A New Way of Information Extraction. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2518–2533

BibTeX

@inproceedings{wu-etal-2022-text-table,
    title = "Text-to-Table: A New Way of Information Extraction",
    author = "Wu, Xueqing  and
      Zhang, Jiacheng  and
      Li, Hang",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.180/",
    doi = "10.18653/v1/2022.acl-long.180",
    pages = "2518--2533"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/