TableFormer: Robust Transformer Modeling for Table-Text Encoding

Jingfeng YangAditya GuptaShyam UpadhyayLuheng HeRahul GoelShachi Paul

article2022ACL140 citations

Proposes TableFormer, a table-text encoding architecture that replaces standard positional embeddings with learnable structural attention biases to achieve strict invariance to row and column perturbations while outperforming existing models on SQA, WTQ, and TabFact benchmarks.

Listen

Modern digital assistants and enterprise search engines increasingly rely on machine learning models to extract answers and verify facts from tabular data found across the web and internal databases. Standard transformer-based architectures process these semi-structured tables by flattening them into sequential text strings and assigning absolute row, column, and global position markers. This conventional approach inadvertently introduces arbitrary ordering biases. As a result, existing systems are brittle: simply reordering rows or columns without altering the underlying data causes accuracy to drop significantly, leading to unreliable answers in real-world applications.

The article demonstrates and evaluates a novel neural architecture, TABLEFORMER, designed to robustly understand tables and paired text without relying on artificial ordering markers. The primary objective is to make table-based question answering and fact verification invariant to row and column permutations while enhancing overall accuracy through explicit structural awareness.

To overcome the vulnerability of previous systems, the authors removed global position numbers as well as row and column identity numbers entirely, switching instead to per-cell positional numbering. In their place, the architecture introduces 13 learnable, task-independent attention bias numbers that explicitly describe the structural relationships between tokens (such as being in the same row, same column, cell-to-header, or cell-to-sentence). The authors evaluated this approach against standard baselines across three established benchmark datasets: Sequential Question Answering (SQA), WikiTableQuestions (WTQ), and TABFACT for fact verification. They also constructed stress-test datasets by randomly shuffling rows and columns to measure prediction consistency under input perturbations.

The evaluation yielded several key findings regarding accuracy and stability. First, TABLEFORMER achieved state-of-the-art results on the SQA benchmark (reaching 72.4% cell selection accuracy and 75.9% denotation accuracy with intermediate pre-training) and outperformed baselines on WTQ and TABFACT across all test configurations. Second, baseline models suffered a 4% to 6% absolute drop in accuracy when exposed to row and column shuffling, whereas TABLEFORMER maintained stable performance with virtually zero prediction variation (0.1% compared to 10%–15% in standard baselines). Third, the architecture achieved these gains with fewer parameters by eliminating bulky row and column embedding matrices in exchange for a negligible number of attention bias values. Fourth, ablation studies revealed that structural biases—particularly the "same row" bias—are critical, and that soft learnable biases significantly outperform rigid attention masking.

These findings indicate that table understanding systems do not need sequential ordering assumptions to model tabular logic effectively. By eliminating spurious positional biases, organizations can deploy more reliable table reasoning models that prevent silent prediction failures caused by arbitrary data presentation. Furthermore, the results show that architectural invariance is superior to data augmentation, as training baselines on shuffled data mitigated prediction shifts only partially while still lagging behind TABLEFORMER in absolute accuracy.

Organizations developing search, analytics, or automated question-answering systems over tabular data should adopt structural relative attention mechanisms rather than flattening tables into plain text sequences. Practitioners aiming to implement this design should account for an approximate 20% increase in training time and scope compute resources accordingly, especially for long tables. Additionally, because the architecture is strictly order-invariant, developers addressing niche queries that explicitly ask about presentation order (such as identifying the top-listed row) will need to reintroduce positional signals for those specific use cases.

The results provide high confidence in the model's robustness and accuracy across standard question answering and fact-checking tasks. However, users should remain mindful of boundary conditions, particularly the modest training overhead and the rare edge cases (accounting for roughly 0.2% of benchmark queries) where presentation order is semantically meaningful to the question.

arXiv: 2203.00274google-research/tapas
Cover for TableFormer: Robust Transformer Modeling for Table-Text Encoding

Abstract

Understanding tables is an important aspect of natural language understanding. Existing models for table understanding require linearization of the table structure, where row or column order is encoded as an unwanted bias. Such spurious biases make the model vulnerable to row and column order perturbations. Additionally, prior work has not thoroughly modeled the table structures or table-text alignments, hindering the table-text understanding ability. In this work, we propose a robust and structurally aware table-text encoding architecture TableFormer, where tabular structural biases are incorporated completely through learnable attention biases. TableFormer is (1) strictly invariant to row and column orders, and, (2) could understand tables better due to its tabular inductive biases. Our evaluations showed that TableFormer outperforms strong baselines in all settings on SQA, WTQ and TabFact table reasoning datasets, and achieves state-of-the-art performance on SQA, especially when facing answer-invariant row and column order perturbations (6% improvement over the best baseline), because previous SOTA models’ performance drops by 4% - 6% when facing such perturbations while TableFormer is not affected.

Table of Contents

  • Abstract
  • 1 Introduction
  • 2 Preliminaries: TAPAS for Table Encoding
  • 3 TABLEFORMER: Robust Structural Table Encoding
  • 4 Experimental Setup
  • 4.1 Datasets and Evaluation
  • 4.2 Baselines
  • 4.3 Perturbing Tables as Augmented Data
  • 5 Experiments and Results
  • 5.1 Main Results
  • 5.2 Perturbation Results
  • 5.3 Model Size Comparison
  • 5.4 Analysis of TABLEFORMER Submodules
  • 5.5 Comparison of TABLEFORMER and Perturbed Data Augmentation
  • 5.6 Attention Bias Ablation Study
  • 5.7 Limitations of TABLEFORMER
  • 6 Other Related Work
  • 7 Conclusion
  • Acknowledgments
  • Ethical Considerations
  • References

Knowls

  1. Knowl 1 — TABLEFORMER replaces ordered table embeddings with relation-biased attention

    model/method

    TABLEFORMER encodes a tokenized question–table sequence using the same token, segment, and numerical-rank embeddings as TAPAS, but removes row-ID and column-ID embeddings. It also replaces global token positions with per-cell positions: the position index restarts at the beginning of each cell, preserving word order within a cell without encoding the order of cells. Table and text structure instead affects attention through learned scalar biases.

    For token positions ii and jj, let qiq_i and kjk_j be their query and key vectors in one attention head, let vjv_j be the value vector at position jj, and let dkd_k be the key-vector dimension. Let xix_i and xjx_j denote the table-text elements associated with those positions, let ϕ(xi,xj)\phi(x_i,x_j) identify their structural relation, and let brb_r be the learned scalar for relation type rr. The attention score and output are:

    aij=qi⊤kjdk+bϕ(xi,xj),αij=exp⁡(aij)∑j′=1nexp⁡(aij′),oi=∑j=1nαijvj,a_{ij}=\frac{q_i^\top k_j}{\sqrt{d_k}}+b_{\phi(x_i,x_j)},\qquad \alpha_{ij}=\frac{\exp(a_{ij})}{\sum_{j'=1}^{n}\exp(a_{ij'})},\qquad o_i=\sum_{j=1}^{n}\alpha_{ij}v_j,

    where nn is the number of sequence positions, αij\alpha_{ij} is the attention weight from position ii to position jj, and oio_i is the resulting attention output. Each attention head in each Transformer layer has one learned scalar for each of the 13 relation types. The structural biases are added after scaling the query–key dot product; TABLEFORMER does not hard-mask all relations outside selected rows or columns.

  2. Knowl 2 — The 13 attention relations encode table structure and table-text alignment

    definition

    TABLEFORMER assigns each pair of sequence elements one of 13 task-independent attention-bias types: same row; same column; same cell; header to column cell; cell to column header; header to sentence; cell to sentence; sentence to header; sentence to cell; sentence to sentence; header to same header; header to other header; and others. The row and column relations convey shared table structure without ordered row or column IDs. The header–cell relations connect cells to their column schema, while the sentence–header and sentence–cell relations support grounding between text and the table. The header-to-header relations encode schema relations, and same-cell bias supports cell-content processing. The others category covers relations not explicitly represented by the other 12 types; it allows attention between cells in different rows or columns rather than excluding those connections.

  3. Knowl 3 — TABLEFORMER is invariant to row and column permutations that preserve table meaning

    theoretical result

    For a table-text input whose meaning is unchanged by reordering its rows or columns, TABLEFORMER is designed to return the same prediction after those permutations. It does not encode global cell order or ordered row and column IDs; instead, its per-cell positions preserve only within-cell token order, and its attention relations describe properties such as shared row, shared column, and cell-to-header association. A row or column permutation therefore changes sequence placement without changing the structural relations that determine attention. The paper reports a small evaluation exception for SQA: some tables contain identical columns while the gold answer selects cells in only one of them, so a model invariant to column order can select either column. This accounts for up to a 0.1 percentage-point performance difference in that evaluation, rather than a substantive prediction change caused by order.

  4. Knowl 4 — TABLEFORMER improves SQA accuracy and largely removes perturbation sensitivity

    empirical result

    On the SQA test set, TABLEFORMER outperforms its corresponding TAPAS model in every reported model-size and pretraining setting. After random row and column reordering at inference, TAPAS loses accuracy while TABLEFORMER's scores are unchanged or differ by at most 0.1 percentage point. The following are medians over five runs; ALL and SEQ are cell-selection accuracy, ALLd is denotation accuracy, and VP is the percentage of example predictions that change correctness status after perturbation.

    Model ALL before SEQ before ALLd before ALL after VP
    TAPASBASE 61.1 31.3 – 57.4 14.0%
    TABLEFORMERBASE 66.7 39.7 – 66.7 0.2%
    TAPASLARGE 66.8 39.9 – 60.5 15.1%
    TABLEFORMERLARGE 70.3 44.8 – 70.3 0.1%
    TAPASBASE inter 67.5 38.8 – 61.0 14.3%
    TABLEFORMERBASE inter 69.4 43.5 – 69.3 0.1%
    TAPASLARGE inter 70.6 43.9 – 66.1 10.8%
    TABLEFORMERLARGE inter 72.4 47.5 75.9 72.3 0.1%

    The strongest reported TABLEFORMER result is 75.9 ALLd for the large model with intermediate pretraining. The TAPAS ALL scores fall by 4.2–6.5 points between the unperturbed and perturbed test conditions in these rows, whereas TABLEFORMER remains effectively stable.

  5. Knowl 5 — TABLEFORMER improves TABFACT accuracy and maintains it on reordered test tables

    empirical result

    On TABFACT, TABLEFORMER has higher binary-classification accuracy than its matching TAPAS baseline in each reported model-size and pretraining setting. Reordering test-table rows and columns leaves TABLEFORMER's scores unchanged across the reported aggregate test and test splits; TAPAS scores decline. Entries are accuracy percentages, with five-run medians. The test-simple, test-complex, and test-small values are for the perturbed test condition.

    Model Test before Test after Simple after Complex after Small after
    TAPASBASE 72.3 71.2 83.4 65.2 72.5
    TABLEFORMERBASE 75.0 75.0 88.2 68.5 77.1
    TAPASLARGE 74.5 73.7 86.0 67.7 76.1
    TABLEFORMERLARGE 77.0 77.0 90.2 70.5 80.3
    TAPASBASE inter 77.9 76.8 89.5 70.5 79.7
    TABLEFORMERBASE inter 79.2 79.2 91.6 73.1 81.7
    TAPASLARGE inter 80.6 79.2 91.7 73.0 83.0
    TABLEFORMERLARGE inter 81.6 81.6 93.3 75.9 84.6

    For example, the large intermediate-pretraining TABLEFORMER scores 81.6 on the aggregate test both before and after perturbation, while the corresponding TAPAS score changes from 80.6 to 79.2. The perturbed simple, complex, and small split scores for TABLEFORMER also match their unperturbed counterparts.

  6. Knowl 6 — TABLEFORMER achieves the best reported WTQ result in the evaluated setups

    empirical result

    On WikiTableQuestions (WTQ), TABLEFORMER exceeds its corresponding TAPAS model on test denotation accuracy in each reported model-size and pretraining setting. With large-scale intermediate pretraining that includes further pretraining on SQA, TABLEFORMER reaches 52.6 test accuracy, above the listed prior results of 48.8 and 51.5. Entries are denotation accuracy percentages and are medians over five runs; dashes indicate values not reported.

    Model Development Test
    Herzig et al. (2020) – 48.8
    Eisenschlos et al. (2021) – 51.5
    TAPASBASE 23.6 24.1
    TABLEFORMERBASE 34.4 34.8
    TAPASLARGE 40.8 41.7
    TABLEFORMERLARGE 42.5 43.9
    TAPASBASE inter-sqa 44.8 45.1
    TABLEFORMERBASE inter-sqa 46.7 46.5
    TAPASLARGE inter-sqa 49.9 50.4
    TABLEFORMERLARGE inter-sqa 51.3 52.6

    The inter-sqa setting is the setting with intermediate pretraining followed by additional pretraining on SQA before WTQ fine-tuning.

  7. Knowl 7 — Attention-bias ablations identify same-row structure as especially important

    empirical result

    On the SQA development set, removing the same-row bias causes the largest individual drop among the tested relation ablations: ALL cell-selection accuracy falls from 62.1 to 32.1 and SEQ accuracy from 38.4 to 2.8. Removing all three column-related biases together also causes a large decline, while removing same-column bias alone has little effect, consistent with the cell-to-header and header-to-cell relations supplying overlapping column information. For each ablation, the removed relation is assigned to the others category. Scores are percentages.

    TABLEFORMER variant ALL SEQ
    All biases 62.1 38.4
    Without same row 32.1 2.8
    Without same column 62.1 37.7
    Without same cell 61.8 38.4
    Without cell to column header 60.7 36.6
    Without cell to sentence 60.5 36.4
    Without header to column cell 60.5 35.8
    Without header to other header 60.6 35.8
    Without header to same header 61.0 36.9
    Without header to sentence 61.1 36.3
    Without sentence to cell 60.8 36.2
    Without sentence to header 61.0 37.3
    Without sentence to sentence 60.0 35.3
    Without all column-related biases 54.5 29.3

    Additional SQA development-set comparisons show that soft learned biases outperform hard attention masking and that the bias should be added after query–key scaling. In the following table, rc-gp means row IDs, column IDs, and global positions; c-gp means column IDs and global positions; gp means global positions only; and pcp means per-cell positions. SO adds the bias before scaling, while SAT masks selected attention scores.

    Model rc-gp c-gp gp pcp
    TAPASBASE 57.6 47.4 46.4 29.1
    TAPASBASE-SAT 45.2 – – –
    TABLEFORMERBASE-SO 60.0 60.2 59.8 60.7
    TABLEFORMERBASE 62.2 61.5 61.7 61.9

    TABLEFORMER's pcp score remains close to its rc-gp score as ordered IDs are removed, whereas TAPAS degrades substantially. The authors interpret this as evidence that structural attention biases supply useful table-structure information without ordered IDs.

  8. Knowl 8 — Perturbed-table training improves TAPAS robustness but does not match TABLEFORMER

    empirical result

    The authors trained TAPASBASE with one, two, four, eight, or sixteen randomly reordered versions of each training table as augmentation. At each epoch, a version generated with a different random seed was used cyclically, and answer-cell positions were updated to match the reordered table. On SQA, augmentation improves TAPASBASE over its unaugmented baseline up to eight versions on SEQ accuracy, but does not eliminate prediction changes after test-time perturbation. TABLEFORMERBASE remains more accurate and has much lower prediction variation. Scores are percentages and are medians over five runs.

    Model ALL SEQ ALL after perturbation VP
    TAPASBASE 61.1 31.3 57.4 14.0%
    TAPASBASE 1p 63.4 34.6 63.4 9.9%
    TAPASBASE 2p 64.6 35.6 64.5 8.4%
    TAPASBASE 4p 65.1 37.0 65.0 8.1%
    TAPASBASE 8p 65.1 37.3 64.3 7.2%
    TAPASBASE 16p 62.4 33.6 62.2 7.0%
    TABLEFORMERBASE 66.7 39.7 66.7 0.1%

    Here, pp is the number of perturbed versions added per table, and VP is the percentage of examples whose predictions change correctness status after perturbation. Sixteen versions yield the lowest TAPASBASE VP among the augmentation settings but lower unperturbed accuracy than eight versions; even that VP, 7.0%, remains far above TABLEFORMERBASE's 0.1%.

  9. Knowl 9 — Evaluation uses standard and order-perturbed table reasoning benchmarks

    experimental setup

    The evaluation covers WikiTableQuestions (WTQ) and Sequential QA (SQA) for table question answering, and TABFACT for table-text entailment. SQA contains 6,066 question sequences, averaging 2.9 questions per sequence; TABFACT contains 118,000 crowd-written statements labeled entailed or not entailed. All models are pretrained on a Wikipedia table-text dataset, optionally receive intermediate synthetic-data pretraining, and are then fine-tuned on the target task. For WTQ, the inter-sqa setting adds SQA pretraining after intermediate pretraining.

    For SQA and TABFACT, the authors also evaluate test sets in which each table's rows and columns are randomly reordered while preserving the table content and the association between column headers and columns. SQA reports cell-selection accuracy for all questions (ALL), cell-selection accuracy for all sequences (SEQ), denotation accuracy (ALLd), and prediction variation (VP); WTQ reports denotation accuracy; TABFACT reports binary-classification accuracy. VP is the fraction of examples whose prediction changes from correct to incorrect or incorrect to correct after perturbation: VP=(t2f+f2t)/(t2t+t2f+f2t+f2f)VP=(t2f+f2t)/(t2t+t2f+f2t+f2f), where each count records the corresponding before-to-after correctness transition. Reported results are medians of five runs.

  10. Knowl 10 — TABLEFORMER uses fewer parameters than TAPAS but adds training cost

    empirical result

    TABLEFORMER adds 13 scalar parameters per attention head per layer and removes the row- and column-ID embedding tables. The reported parameter counts are approximately 110M for TAPASBASE and 340M for TAPASLARGE; TABLEFORMERBASE is calculated as 110M−2×512×768+12×12×13110\text{M}-2\times512\times768+12\times12\times13, or about 110M−0.8M+0.002M110\text{M}-0.8\text{M}+0.002\text{M}. TABLEFORMERLARGE is calculated as 340M−2×512×1024+24×16×13340\text{M}-2\times512\times1024+24\times16\times13, or about 340M−1.0M+0.005M340\text{M}-1.0\text{M}+0.005\text{M}. Thus the deleted embedding parameters outweigh the added attention scalars. The paper separately reports that TABLEFORMER training takes around 20% longer.

  11. Knowl 11 — Order invariance limits questions that depend on absolute table position

    limitation

    Because TABLEFORMER removes row and column order information, it cannot directly answer questions whose correct response depends on absolute row or column order, such as asking which item is at the top or listed first. In a manual review of 1,800 SQA questions, the authors found four questions (0.2%) whose answers depend on row order: three ask which item is at the top and one asks which is listed first. The paper suggests that order information could potentially be added back for such cases, but does not report an evaluated solution. TABLEFORMER also incurs around 20% more training time, which the authors note may be undesirable for very long tables and may require a scoped approach.

Coverage note — No substantial contributed material was deliberately omitted; background and related-work comparisons are excluded because they are not part of the paper's contribution.

References

  1. 1.Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020. ETC: Encoding long and structured inputs in transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 268–284, Online. Association for Computational Linguistics.
  2. 2.Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150.
  3. 3.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. Tabfact : A large-scale dataset for table-based fact verification. In International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia.
  4. 4.Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019. Transformer-XL: Attentive language models beyond a fixed-length context. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 2978–2988, Florence, Italy. Association for Computational Linguistics.
  5. 5.Xiang Deng, Huan Sun, Alyssa Lees, You Wu, and Cong Yu. 2021. TURL: Table Understanding through Representation Learning. In VLDB.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  7. 7.Julian Eisenschlos, Maharshi Gor, Thomas Müller, and William Cohen. 2021. MATE: Multi-view attention for table transformer efficiency. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7606–7619, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  8. 8.Julian Eisenschlos, Syrine Krichene, and Thomas Müller. 2020. Understanding tables with intermediate pre-training. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 281–296, Online. Association for Computational Linguistics.
  9. 9.Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Eisenschlos. 2020. TaPas: Weakly supervised table parsing via pre-training. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4320–4333, Online. Association for Computational Linguistics.
  10. 10.Mohit Iyyer, Wen-tau Yih, and Ming-Wei Chang. 2017. Search-based neural structured learning for sequential question answering. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1821–1831, Vancouver, Canada. Association for Computational Linguistics.
  11. 11.Qian Liu, Bei Chen, Jiaqi Guo, Zeqi Lin, and Jianguang Lou. 2021. Tapex: Table pre-training via learning a neural sql executor. arXiv preprint arXiv:2107.07653.
  12. 12.Thomas Mueller, Francesco Piccinno, Peter Shaw, Massimo Nicosia, and Yasemin Altun. 2019. Answering conversational questions on structured data without logical forms. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 5902–5910, Hong Kong, China. Association for Computational Linguistics.
  13. 13.Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1470–1480, Beijing, China. Association for Computational Linguistics.
  14. 14.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-attention with relative position representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), pages 464–468, New Orleans, Louisiana. Association for Computational Linguistics.
  15. 15.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008.
  16. 16.Bailin Wang, Richard Shin, Xiaodong Liu, Oleksandr Polozov, and Matthew Richardson. 2020. RAT-SQL: Relation-aware schema encoding and linking for text-to-SQL parsers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7567–7578, Online. Association for Computational Linguistics.
  17. 17.Zhiruo Wang, Haoyu Dong, Ran Jia, Jia Li, Zhiyi Fu, Shi Han, and Dongmei Zhang. 2021. TUTA: Tree-based Transformers for Generally Structured Table Pre-training. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pages 1780–1790.
  18. 18.Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020. TaBERT: Pretraining for joint understanding of textual and tabular data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8413–8426, Online. Association for Computational Linguistics.
  19. 19.Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do Transformers Really Perform Bad for Graph Representation? arXiv preprint arXiv:2106.05234.
  20. 20.Tao Yu, Chien-Sheng Wu, Xi Victoria Lin, Bailin Wang, Yi Chern Tan, Xinyi Yang, Dragomir Radev, Richard Socher, and Caiming Xiong. 2020. GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing. arXiv preprint arXiv:2009.13845.
  21. 21.Hongzhi Zhang, Yingyao Wang, Sirui Wang, Xuezhi Cao, Fuzheng Zhang, and Zhongyuan Wang. 2020. Table fact verification with structure-aware transformer. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1624–1629, Online. Association for Computational Linguistics.

Citation

MLA
Yang, J., et al. “TableFormer: Robust Transformer Modeling for Table-Text Encoding”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 528–37, https://doi.org/10.18653/v1/2022.acl-long.40.
APA
Yang, J., Gupta, A., Upadhyay, S., He, L., Goel, R., & Paul, S. (2022). TableFormer: Robust Transformer Modeling for Table-Text Encoding. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 528–537. https://doi.org/10.18653/v1/2022.acl-long.40
Chicago
Yang, J., A. Gupta, S. Upadhyay, L. He, R. Goel, and S. Paul. 2022. “TableFormer: Robust Transformer Modeling for Table-Text Encoding”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 528–37. https://doi.org/10.18653/v1/2022.acl-long.40.
Harvard
Yang, J. et al. (2022) “TableFormer: Robust Transformer Modeling for Table-Text Encoding”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 528–537. Available at: https://doi.org/10.18653/v1/2022.acl-long.40.
Vancouver
1. Yang J, Gupta A, Upadhyay S, He L, Goel R, Paul S (2022) TableFormer: Robust Transformer Modeling for Table-Text Encoding. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 528–537

BibTeX

@inproceedings{yang-etal-2022-tableformer,
    title = "{T}able{F}ormer: Robust Transformer Modeling for Table-Text Encoding",
    author = "Yang, Jingfeng  and
      Gupta, Aditya  and
      Upadhyay, Shyam  and
      He, Luheng  and
      Goel, Rahul  and
      Paul, Shachi",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.40/",
    doi = "10.18653/v1/2022.acl-long.40",
    pages = "528--537"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/