FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction

Chen-Yu LeeChun-Liang LiTimothy DozatVincent PerotGuolong SuNan HuaJoshua AinslieRenshen WangYasuhisa FujiiTomas Pfister

article2022ACL100 citations

Proposes a structure-aware sequence model that combines graph-convolutional token representations with spatial attention to extract key information from visually rich form documents without relying on expensive image features or large pre-training budgets.

Listen

Organizations routinely rely on extracting critical information from structured forms such as invoices, receipts, and applications to automate business workflows. However, standard language processing methods convert documents into a single stream of text, which breaks the physical layout and frequently scatters related information across separate text segments. When this serialization fails, automated systems misread relationships among key entities, driving up extraction errors and requiring costly manual intervention. The article introduces and evaluates FormNet, a structure-aware neural network architecture designed to overcome these layout-induced serialization errors without relying on computationally expensive visual processing.

The authors tackle this challenge by combining graph-based representation learning with an enhanced transformer sequence model. Before the extracted text is flattened into a sequential order, a graph convolutional network creates rich "Super-Tokens" by connecting neighboring words based on their two-dimensional layout geometry, preserving local context that might otherwise be broken apart. The system then processes these representations through a long-sequence transformer enhanced with "Rich Attention," a mechanism that directly incorporates horizontal and vertical spatial distances to penalize illogical text connections. The authors evaluate this architecture on three standard benchmarks—CORD (receipts), FUNSD (noisy forms), and Payment (invoices)—using unsupervised pre-training on 700,000 unlabeled documents followed by task-specific fine-tuning.

The findings show that FormNet achieves state-of-the-art performance across all three benchmarks while operating with notable computational efficiency. It sets new performance records with accuracy scores (F1) of 97.28% on CORD, 84.69% on FUNSD, and 92.19% on Payment. Furthermore, FormNet outperforms leading multimodal models while using a 64% smaller model footprint and over seven times less pre-training data. Ablation experiments confirm that both the graph-based Super-Tokens and Rich Attention contribute substantial, complementary gains, improving baseline performance on receipts by 5.3 percentage points and proving that direct structural modeling is far more effective than simply enlarging standard transformer architectures.

These results demonstrate that explicitly modeling geometric relationships enables high-accuracy document parsing at substantially lower operational and compute costs. By eliminating the need to process raw document images alongside text, FormNet reduces hardware requirements and inference latency, making high-volume document extraction more scalable and reliable. Organizations adopting this approach can lower error rates in automated back-office processing and achieve faster processing cycles.

For practical implementation, technical teams should consider adopting graph-enhanced structural modeling as the standard pipeline for form-based information extraction. While confidence in these results is high across standardized form datasets, the current architecture focuses strictly on textual and layout coordinates rather than image pixels and is designed for entity extraction rather than arbitrary key-value pairing. Future efforts should conduct pilot evaluations on proprietary, highly noisy layouts and explore lightweight integrations of visual features for documents containing rich pictorial cues.

arXiv: 2203.08411
  • Paper: Longformer: The Long-Document Transformer, Iz Beltagy et al. (2020). Longformer’s efficient local-and-global attention provides the long-sequence transformer foundation that FormNet adapts for document text.
Cover for FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 FormNet for Information Extraction
  • 3.1 Extended Transformer Construction
  • 3.2 Rich Attention
  • 3.3 Super-Token by Graph Learning
  • 4 Evaluation
  • 4.1 Datasets
  • 4.2 Experimental Setup
  • 4.3 Results
  • Effect of Structural Encoding in Pre-training
  • Effect of Structural Encoding in Fine-tuning
  • 4.4 Visualization
  • 5 Conclusion
  • References
  • A Implementation Details
  • B Impact of Super-Tokens by Graph Convolutional Networks
  • C Rich Attention Derivations
  • D Examples of β-Skeleton Graphs
  • E Additional Attention Visualization

Knowls

  1. Knowl 1 — FormNet combines 2D graph encoding with long-sequence attention

    model/method

    FormNet extracts entities from form documents by combining graph-based structural encoding with sequential tagging. It starts from OCR text and token bounding boxes, tokenizes the words with the BERT-multilingual vocabulary, and builds a graph over the resulting tokens. A graph convolutional network (GCN) produces a structure-aware Super-Token representation for each token before serialization. An Extended Transformer Construction (ETC) encoder then processes these representations in the dataset's serialized order, using Rich Attention to incorporate spatial relationships. The system predicts entity tags for tokens and decodes the tag sequence. FormNet uses text and layout coordinates rather than document-image features.

  2. Knowl 2 — Rich Attention adds learned spatial order and distance scores

    model/method

    Rich Attention modifies each ETC attention head's pre-softmax score using the layout order and distance of token pairs along both the horizontal and vertical axes. For tokens ii and jj, let hiℓh_i^{\ell} and hjℓh_j^{\ell} be their representations at layer ℓ\ell, and let qiq_i and kjk_j be the usual learned query and key projections. For each layout axis a∈{x,y}a\in\{x,y\}, the model predicts an order probability pija=sigmoid⁡(affine⁡([hiℓ;hjℓ]))p_{ij}^{a}=\operatorname{sigmoid}(\operatorname{affine}([h_i^{\ell};h_j^{\ell}])) and an ideal log-distance μija=affine⁡([hiℓ;hjℓ])\mu_{ij}^{a}=\operatorname{affine}([h_i^{\ell};h_j^{\ell}]). It compares the predicted order with the actual binary order oijao_{ij}^{a} using a sigmoid cross-entropy score, and penalizes disagreement between the actual log-distance dijad_{ij}^{a} and the predicted distance with sij(d,a)=−(θa)22(dija−μija)2s_{ij}^{(d,a)}=-\frac{(\theta_a)^2}{2}(d_{ij}^{a}-\mu_{ij}^{a})^2, where θa\theta_a is a learned temperature for the head. The resulting attention score is the query-key dot product plus the order and distance scores for both axes. The order and distance are computed from token positions in the document layout; distance is log-scaled. The paper also reports using query/key projections instead of the layer representations as inputs to the affine predictors for computational speed.

  3. Knowl 3 — Super-Tokens propagate local document structure through a beta-skeleton graph

    model/method

    FormNet creates one graph vertex for each OCR token. A vertex representation concatenates a word representation with spatial features: normalized Cartesian coordinates of the token bounding-box corners, height, and width. Directed edges carry relative geometric features, including distances between box centers and corresponding corners, shortest horizontal and vertical box-to-box distances, aspect ratios of the two token boxes, and the aspect ratio of their joint bounding box. Connectivity is defined by a β\beta-skeleton graph with β=1\beta=1, a sparse “ball-of-sight” construction that the paper describes as globally connected with a linearly bounded number of edges. GCN message passing over this graph aggregates neighboring token and edge information into a Super-Token for each token, providing context before text serialization can separate related words. In the implementation, each GCN layer uses a two-layer MLP with the ETC hidden size, one-head attention aggregation, skip connections, and layer normalization; the maximum number of neighbors is 8.

  4. Knowl 4 — Rich Attention has a probabilistic interpretation as feature-conditioned attention

    theoretical result

    The paper interprets attention as a posterior distribution over a latent binary relation aija_{ij} indicating whether token jj contributes to token ii's context. Under this interpretation, attention weights are obtained by applying a softmax over key tokens to log-probability terms. If paired token representations are modeled as normally distributed conditional on the relation, their log-likelihood yields the usual query-key dot-product score, up to terms that can be treated as constants or key-only biases. If an additional low-level feature, such as relative order or distance, is assumed conditionally independent of the other features given the attention relation, its conditional log-likelihood adds a separate score to attention. Under the paper's distributional assumptions, a binary feature gives a sigmoid cross-entropy score, while a log-normal distance feature gives a quadratic penalty in log-distance. Thus Rich Attention is presented as a way to add feature-conditioned likelihood terms to attention, rather than as a positional embedding lookup.

  5. Knowl 5 — Entity extraction uses BIOES tags and Viterbi decoding

    model/method

    FormNet casts form-document key information extraction as sequence tagging over tokenized OCR words. It assigns each token a label under the BIOES scheme—Begin, Inside, Outside, End, or Single—combined with the relevant entity class, so entity spans are represented in the serialized token sequence. The model produces tag logits, and the final sequence is decoded with the Viterbi algorithm. The evaluated task is direct entity extraction; the authors explicitly leave adapting FormNet to key-value-pair extraction for future work.

  6. Knowl 6 — Training and benchmark setup for FormNet

    experimental setup

    Experiments cover CORD, FUNSD, and Payment. CORD has 800 training, 100 validation, and 100 test receipts with 30 fine-grained entity types. FUNSD contains 199 forms, 9,707 entities, and 31,485 word-level annotations over four types—header, question, answer, and other—and uses the official 75/25 train/test split. Payment contains around 10,000 documents from vendors with different templates and seven labels; the paper follows the split and evaluation protocol of the prior benchmark. FormNet uses 12 GCN layers and 12 ETC layers, with a maximum ETC sequence length of 1,024; the model family uses 512 hidden units and 8 attention heads (A1), 768 and 12 (A2), or 1,024 and 16 (A3). For CORD and FUNSD, the authors pretrain from scratch with masked language modeling on around 700,000 unlabeled forms, using Adam, batch size 512, learning rate 0.0002, and 0.01 warm-up proportion. Fine-tuning uses Adam, batch size 8, learning rate 0.0001, no warm-up, and cross-entropy loss; the largest-corpus runs take about 10 hours on Tesla V100 GPUs. Payment models are trained from scratch without pretraining. CORD and FUNSD use micro-F1, while Payment uses macro-F1.

  7. Knowl 7 — FormNet sets the reported best F1 on all three benchmarks

    empirical result

    Under the reported benchmark protocols, FormNet achieves the highest listed entity-level F1 on CORD, FUNSD, and Payment. On CORD, FormNet reports precision 98.02, recall 96.55, and F1 97.28, compared with the latest listed DocFormer result of 97.25 precision, 96.74 recall, and 96.99 F1. On FUNSD, FormNet reports 85.21 precision, 84.18 recall, and 84.69 F1, compared with DocFormer's 82.29, 86.94, and 84.55. On Payment, FormNet reports 92.70 precision, 91.69 recall, and 92.19 F1, compared with the prior NeuralScoring F1 of 87.80 (the paper does not list its precision or recall). The CORD and FUNSD DocFormer comparison uses a 536-million-parameter model pretrained on 5 million documents; the corresponding FormNet entries list 345 million parameters and 0.7 million pretraining documents (9 GB) on CORD, and 217 million parameters and 0.7 million documents on FUNSD.

  8. Knowl 8 — Rich Attention and GCN each improve entity-tagging performance

    empirical result

    Ablations compare the ETC baseline, ETC with Rich Attention, ETC with the GCN Super-Tokens, and ETC with both components. On CORD using FormNet-A1, the respective precision/recall/F1 scores are 91.40/91.75/91.57, 97.28/95.19/96.03, 96.50/95.13/95.81, and 97.50/96.25/96.87. On FUNSD using A1, the scores are 69.24/62.86/65.90, 82.16/82.28/82.22, 78.83/79.93/79.37, and 84.17/84.88/84.53. On Payment using A2, they are 83.91/83.27/83.58, 92.10/91.48/91.79, 87.79/84.47/86.10, and 92.70/91.69/92.19. Both components outperform the ETC baseline on each dataset, and the combination has the highest F1 among the four configurations in each ablation.

  9. Knowl 9 — Scaling FormNet improves CORD results without relying only on transformer capacity

    empirical result

    The FormNet family scales from A1 (512 hidden units, 8 attention heads) through A2 (768 units, 12 heads) to A3 (1,024 units, 16 heads). On CORD, their F1 scores are 96.87, 97.10, and 97.28, respectively; on FUNSD, A1 and A2 score 84.53 and 84.69. A separate FUNSD capacity comparison finds F1 of 82.22 for ETC-standard plus Rich Attention at 104 million parameters, 82.92 for ETC-heavy plus Rich Attention at 187 million parameters, and 84.53 for ETC-standard plus Rich Attention and GCN at 131 million parameters. Thus, adding the GCN Super-Tokens raises F1 by more than simply increasing ETC capacity in this comparison, while using fewer parameters than the heavier ETC model.

  10. Knowl 10 — Both structural components improve masked-language-model pretraining

    empirical result

    The authors compare ETC alone, ETC with Rich Attention, ETC with GCN Super-Tokens, and ETC with both components on masked language modeling across FormNet-A1, A2, and A3. The plotted results show that each structural component improves masked-token reconstruction accuracy over ETC alone for all three model sizes, and the configuration with both Rich Attention and GCN performs best. The paper presents these results graphically without reporting exact bar values in the text.

Coverage note — The qualitative attention maps and example extraction errors are omitted because they illustrate model behavior but add no separate quantitative result or method beyond the findings captured here.

References

  1. 1.Milan Aggarwal, Hiresh Gupta, Mausoom Sarkar, and Balaji Krishnamurthy. 2020. Form2seq: A framework for higher-order form structure extraction. In EMNLP.
  2. 2.Joshua Ainslie, Santiago Ontañón, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020. Etc: Encoding long and structured data in transformers. In EMNLP.
  3. 3.Apostolos Antonacopoulos, David Bridson, Christos Papadopoulos, and Stefan Pletschacher. 2009. A realistic dataset for performance evaluation of document layout analysis. In ICDAR.
  4. 4.Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha. 2021. Docformer: End-to-end transformer for document understanding. In ICCV.
  5. 5.Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, et al. 2020. Unilmv2: Pseudo-masked language models for unified language model pre-training. In ICML.
  6. 6.Laura Chiticariu, Yunyao Li, and Frederick Reiss. 2013. Rule-based information extraction is dead! long live rule-based information extraction systems! In EMNLP.
  7. 7.Brian Davis, Bryan Morse, Scott Cohen, Brian Price, and Chris Tensmeyer. 2019. Deep visual template-free form parsing. In ICDAR.
  8. 8.Timo I Denk and Christian Reisswig. 2019. Bert-grid: Contextualized embedding for 2d document representation and understanding. arXiv preprint arXiv:1909.04948.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT.
  10. 10.Timothy Dozat. 2019. Arc-factored Biaffine Dependency Parsing. Stanford University.
  11. 11.Morris L Eaton. 1983. Multivariate statistics: a vector space approach. John Wiley & Sons, Inc., 605 Third Ave., New York, NY 10158, USA.
  12. 12.Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Gralinski. 2020. Lambert: Layout-aware (language) modeling for information extraction. arXiv preprint arXiv:2002.08087.
  13. 13.Justin Gilmer, Samuel S Schoenholz, Patrick F Riley, Oriol Vinyals, and George E Dahl. 2017. Neural message passing for quantum chemistry. In ICML.
  14. 14.Filip Gralinski, Tomasz Stanisławek, Anna Wróblewska, Dawid Lipinski, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemysław Biecek. 2020. Kleister: A novel task for information extraction involving long documents with complex layout. arXiv preprint arXiv:2003.02356.
  15. 15.Jaekyu Ha, Robert M Haralick, and Ihsin T Phillips. 1995. Recursive xy cut using bounding boxes of connected components. In ICDAR.
  16. 16.Takashi Hirano, Yuichi Okano, Yasuhiro Okada, and Fumio Yoda. 2007. Text and layout information extraction from document files of various formats based on the analysis of page description language. In ICDAR.
  17. 17.Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and CV Jawahar. 2019. Icdar2019 competition on scanned receipt ocr and information extraction. In ICDAR.
  18. 18.Wonseok Hwang, Seonghyeon Kim, Minjoon Seo, Jinyeong Yim, Seunghyun Park, Sungrae Park, Junyeop Lee, Bado Lee, and Hwalsuk Lee. 2019. Post-ocr parsing: building simple and robust parser via bio tagging. In Workshop on Document Intelligence at NeurIPS 2019.
  19. 19.Wonseok Hwang, Jinyeong Yim, Seunghyun Park, Sohee Yang, and Minjoon Seo. 2021. Spatial dependency parsing for semi-structured document information extraction. In ACL-IJCNLP (Findings).
  20. 20.Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019. Funsd: A dataset for form understanding in noisy scanned documents. In ICDAR-OST.
  21. 21.Anoop Raveendra Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul. 2018. Chargrid: Towards understanding 2d documents. In EMNLP.
  22. 22.David G Kirkpatrick and John D Radke. 1985. A framework for computational morphology. In Machine Intelligence and Pattern Recognition. Elsevier.
  23. 23.Frank Lebourgeois, Zbigniew Bublinski, and Hubert Emptoz. 1992. A fast and efficient method for extracting text paragraphs and graphics from unconstrained documents. In ICPR.
  24. 24.Chen-Yu Lee, Chun-Liang Li, Chu Wang, Renshen Wang, Yasuhisa Fujii, Siyang Qin, Ashok Popat, and Tomas Pfister. 2021. Rope: Reading order equivariant positional encoding for graph-based document information extraction. In ACL-IJCNLP.
  25. 25.Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao. 2019a. Graph convolution for multimodal information extraction from visually rich documents. In NAACL-HLT.
  26. 26.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  27. 27.Minh-Thang Luong, Thuy Dung Nguyen, and Min-Yen Kan. 2012. Logical structure recovery in scholarly articles with rich document features. In Multimedia Storage and Retrieval Innovations for Digital Library Systems. IGI Global.
  28. 28.Bodhisattwa Prasad Majumder, Navneet Potti, Sandeep Tata, James Bradley Wendt, Qi Zhao, and Marc Najork. 2020. Representation learning for information extraction from form-like documents. In ACL.
  29. 29.Simone Marinai, Marco Gori, and Giovanni Soda. 2005. Artificial neural networks for document analysis and recognition. IEEE Transactions on pattern analysis and machine intelligence.
  30. 30.Lawrence O’Gorman. 1993. The document spectrum for page layout analysis. IEEE Transactions on pattern analysis and machine intelligence.
  31. 31.Rasmus Berg Palm, Ole Winther, and Florian Laws. 2017. Cloudscan-a configuration-free invoice analysis system using recurrent neural networks. In ICDAR.
  32. 32.Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019. Cord: A consolidated receipt dataset for post-ocr parsing. In Workshop on Document Intelligence at NeurIPS 2019.
  33. 33.Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018. Image transformer. In ICML.
  34. 34.Nanyun Peng, Hoifung Poon, Chris Quirk, Kristina Toutanova, and Wen-tau Yih. 2017. Cross-sentence n-ary relation extraction with graph lstms. Transactions of the Association for Computational Linguistics (TACL).
  35. 35.Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka. 2021. Going full-tilt boogie on document understanding with text-image-layout transformer. In ICDAR.
  36. 36.Yujie Qian, Enrico Santus, Zhijing Jin, Jiang Guo, and Regina Barzilay. 2019. GraphIE: A graph-based framework for information extraction. In NAACL-HLT.
  37. 37.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research (JMLR).
  38. 38.Lev Ratinov and Dan Roth. 2009. Design challenges and misconceptions in named entity recognition. In Conference on Computational Natural Language Learning (CoNLL).
  39. 39.Daniel Schuster, Klemens Muthmann, Daniel Esser, Alexander Schill, Michael Berger, Christoph Weidling, Kamil Aliyev, and Andreas Hofmeier. 2013. Intellix–end-user trained information extraction for document archiving. In ICDAR.
  40. 40.Sofia Serrano and Noah A Smith. 2019. Is attention interpretable? arXiv preprint arXiv:1906.03731.
  41. 41.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. Self-attention with relative position representations. In NAACL-HLT.
  42. 42.Michael Shilman, Percy Liang, and Paul Viola. 2005. Learning nongenerative grammatical models for document analysis. In ICCV.
  43. 43.Anikó Simon, J-C Pret, and A Peter Johnson. 1997. A fast algorithm for bottom-up document layout analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence.
  44. 44.Linfeng Song, Yue Zhang, Zhiguo Wang, and Daniel Gildea. 2018. N-ary relation extraction using graph state lstm. arXiv preprint arXiv:1808.09101.
  45. 45.Carlos Soto and Shinjae Yoo. 2019. Visual detection with context for document layout analysis. In EMNLP-IJCNLP.
  46. 46.Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In NIPS.
  47. 47.Wilson L Taylor. 1953. “cloze procedure”: A new tool for measuring readability. Journalism quarterly.
  48. 48.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NIPS.
  49. 49.Jesse Vig. 2019. A multiscale visualization of attention in the transformer model. In ACL: System Demonstrations.
  50. 50.Renshen Wang, Yasuhisa Fujii, and Ashok C. Popat. 2022. Post-ocr paragraph recognition by graph convolutional networks. In WACV.
  51. 51.Hao Wei, Micheal Baechler, Fouad Slimane, and Rolf Ingold. 2013. Evaluation of svm, mlp and gmm classifiers for layout analysis of historical documents. In ICDAR.
  52. 52.Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al. 2021. Layoutlmv2: Multi-modal pre-training for visually-rich document understanding. In ACL-IJCNLP.
  53. 53.Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020. Layoutlm: Pre-training of text and layout for document image understanding. In KDD.
  54. 54.Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, and C Lee Giles. 2017. Learning to extract semantic structure from documents using multimodal fully convolutional neural networks. In CVPR.
  55. 55.Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao. 2020. Pick: processing key information extraction from documents using improved graph learning-convolutional networks. In ICPR.
  56. 56.Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020. Big bird: Transformers for longer sequences. In NeurIPS.
  57. 57.Kaixuan Zhang, Zejiang Shen, Jie Zhou, and Melissa Dell. 2019. Information extraction from text regions with complex tabular structure. In NeurIPS.
  58. 58.Shi-Xue Zhang, Xiaobin Zhu, Jie-Bo Hou, Chang Liu, Chun Yang, Hongfa Wang, and Xu-Cheng Yin. 2020. Deep relational reasoning graph network for arbitrary shape text detection. In CVPR.
  59. 59.Xiaohui Zhao, Endi Niu, Zhuo Wu, and Xiaoguang Wang. 2019. Cutie: Learning to understand documents with convolutional universal text information extractor. In ICDAR.

Citation

MLA
Lee, C.-Y., et al. “FormNet: Structural Encoding Beyond Sequential Modeling in Form Document Information Extraction”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 3735–54, https://doi.org/10.18653/v1/2022.acl-long.260.
APA
Lee, C.-Y., Li, C.-L., Dozat, T., Perot, V., Su, G., Hua, N., Ainslie, J., Wang, R., Fujii, Y., & Pfister, T. (2022). FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3735–3754. https://doi.org/10.18653/v1/2022.acl-long.260
Chicago
Lee, C.-Y., C.-L. Li, T. Dozat, et al. 2022. “FormNet: Structural Encoding Beyond Sequential Modeling in Form Document Information Extraction”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 3735–54. https://doi.org/10.18653/v1/2022.acl-long.260.
Harvard
Lee, C.-Y. et al. (2022) “FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3735–3754. Available at: https://doi.org/10.18653/v1/2022.acl-long.260.
Vancouver
1. Lee C-Y, Li C-L, Dozat T, Perot V, Su G, Hua N, Ainslie J, Wang R, Fujii Y, Pfister T (2022) FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 3735–3754

BibTeX

@inproceedings{lee-etal-2022-formnet,
    title = "{F}orm{N}et: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction",
    author = "Lee, Chen-Yu  and
      Li, Chun-Liang  and
      Dozat, Timothy  and
      Perot, Vincent  and
      Su, Guolong  and
      Hua, Nan  and
      Ainslie, Joshua  and
      Wang, Renshen  and
      Fujii, Yasuhisa  and
      Pfister, Tomas",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.260/",
    doi = "10.18653/v1/2022.acl-long.260",
    pages = "3735--3754"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/