LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding

Jiapeng WangLianwen JinKai Ding

article2022ACL199 citations

Proposes a dual-stream layout Transformer that decouples document visual structure from text, allowing layout representations pre-trained on a single language to transfer directly across multiple languages when paired with off-the-shelf text models.

Listen

Structured document understanding plays a critical role in digitizing and automating workflows across industries such as finance, insurance, and healthcare. However, existing artificial intelligence models typically rely on pre-training data from a single language, primarily English, or require massive and costly multilingual document collection pipelines. This dependency limits their deployment in diverse, multilingual enterprise environments where structured document data in other languages is scarce.

The article demonstrates and evaluates a framework called the Language-independent Layout Transformer. The core objective is to show that layout structural knowledge learned exclusively from monolingual English documents can be effectively decoupled and reused across multiple target languages by combining it with off-the-shelf textual language models.

The researchers developed a parallel dual-stream architecture that separately embeds text and 2D spatial layout coordinates. To facilitate language-independent interaction between modalities, they introduced a bi-directional attention complementation mechanism along with two new self-supervised pre-training tasks: key point location and cross-modal alignment identification. The model was pre-trained using 11 million scanned English documents from the IIT-CDIP archive and evaluated across eight languages on standard benchmarks spanning form understanding, receipt parsing, examination papers, and document classification under language-specific, cross-lingual zero-shot, and multi-task settings.

The findings confirm that this decoupled layout approach consistently matches or exceeds the performance of specialized state-of-the-art models. First, when fine-tuned on individual languages, the framework achieved higher accuracy than baseline and multilingual models, reaching an average F1 score of 0.8251 on semantic entity recognition across eight languages compared to 0.8056 for its primary multilingual competitor. Second, in strict zero-shot cross-lingual transfer—where the model had never encountered non-English documents during pre-training—it achieved an average entity recognition F1 score of 0.6061, substantially outperforming prior systems trained on 30 million multilingual documents. Third, the lightweight layout component adds only 6.1 million parameters, delivering high computational efficiency and demonstrating that asynchronous training prevents the layout flow from degrading off-the-shelf text models.

These results show that organizations do not need to undertake expensive, time-consuming data collection and cleaning campaigns to support multi-language document processing. Instead, teams can achieve high extraction accuracy across global document types by pairing a single, lightweight layout model with existing off-the-shelf language encoders, reducing model maintenance costs and operational risk.

Organizations should adopt modular, decoupled architectures for international document extraction workflows and explore zero-shot deployment when bootstrapping operations in new languages. Future development should investigate integrating generalized visual features beyond layout coordinates, though readers should note that current performance remains dependent on the initial accuracy of optical character recognition engines.

arXiv: 2202.13669jpWang/LiLT
Cover for LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding

Abstract

Structured document understanding has attracted considerable attention and made significant progress recently, owing to its crucial role in intelligent document processing. However, most existing related models can only deal with the document data of specific language(s) (typically English) included in the pre-training collection, which is extremely limited. To address this issue, we propose a simple yet effective Language-independent Layout Transformer (LiLT) for structured document understanding. LiLT can be pre-trained on the structured documents of a single language and then directly fine-tuned on other languages with the corresponding off-the-shelf monolingual/multilingual pre-trained textual models. Experimental results on eight languages have shown that LiLT can achieve competitive or even superior performance on diverse widely-used downstream benchmarks, which enables language-independent benefit from the pre-training of document layout structure. Code and model are publicly available at https://github.com/jpWang/LiLT.

Table of Contents

  • 1 Introduction
  • 2 LiLT
  • 2.1 Model Architecture
  • 2.1.1 Text Embedding
  • 2.1.2 Layout Embedding
  • 2.1.3 BiACM
  • 2.2 Pre-training Tasks
  • 2.2.1 Masked Visual-Language Modeling
  • 2.2.2 Key Point Location
  • 2.2.3 Cross-modal Alignment Identification
  • 2.3 Optimization Strategy
  • 3 Experiments
  • 3.1 Pre-training Setting
  • 3.2 Ablation Study
  • 3.3 Comparisons with the SOTAs
  • 3.3.1 Language-specific Fine-tuning
  • 3.3.2 Zero-shot Transfer Learning
  • 3.3.3 Multi-task Fine-tuning
  • 4 Related Work
  • 5 Conclusion
  • 6 Acknowledgement
  • References
  • Appendix
  • A Dataset Details
  • B Fine-tuning Details

Knowls

  1. Knowl 1 — Language-Independent Layout Transformer (LiLT) Architecture

    model/method

    LiLT is a dual-stream Transformer architecture designed to decouple textual and layout representations for structured document understanding (SDU). Text strings and bounding box coordinates extracted via OCR are processed through parallel text and layout streams:

    1. Text Embedding: A tokenized sequence StS_t of length NN (ordered top-left to bottom-right, padded or truncated, with special tokens [CLS] and [SEP]) is embedded as: ET=LN(Etoken+P1D)∈RN×dTE_T = \text{LN}(E_{\text{token}} + P_{1\text{D}}) \in \mathbb{R}^{N \times d_T} where EtokenE_{\text{token}} is the token embedding, P1DP_{1\text{D}} is the 1D positional embedding, LN(⋅)\text{LN}(\cdot) denotes layer normalization, and dTd_T is the text hidden dimension.

    2. Layout Embedding: Bounding boxes B=(xmin⁡,xmax⁡,ymin⁡,ymax⁡,width,height)B = (x_{\min}, x_{\max}, y_{\min}, y_{\max}, \text{width}, \text{height}) are normalized and discretized into integers in the range [0,1000][0, 1000]. Coordinate features are generated through four separate embedding layers, concatenated, and projected into a 2D positional embedding P2D∈RN×dLP_{2\text{D}} \in \mathbb{R}^{N \times d_L}: P2D=Linear(CAT(Exmin⁡,Exmax⁡,Eymin⁡,Eymax⁡,Ewidth,Eheight))P_{2\text{D}} = \text{Linear}(\text{CAT}(E_{x_{\min}}, E_{x_{\max}}, E_{y_{\min}}, E_{y_{\max}}, E_{\text{width}}, E_{\text{height}})) The layout embedding EL∈RN×dLE_L \in \mathbb{R}^{N \times d_L} is formed by adding 1D positional embeddings: EL=LN(P2D+P1D)E_L = \text{LN}(P_{2\text{D}} + P_{1\text{D}}) Special tokens [CLS], [SEP], and [PAD] use dummy boxes (0,0,0,0,0,0)(0,0,0,0,0,0), (1000,1000,1000,1000,0,0)(1000,1000,1000,1000,0,0), and (0,0,0,0,0,0)(0,0,0,0,0,0), respectively.

    In the Base configuration, the layout stream is parameterized as a 12-layer Transformer with hidden size dL=192d_L = 192, feed-forward dimension 768, and 12 attention heads (approximately 6.1M parameters). Text and layout representations are concatenated at the final layer for downstream classification heads.

  2. Knowl 2 — Bi-directional Attention Complementation Mechanism (BiACM)

    model/method

    The Bi-directional Attention Complementation Mechanism (BiACM) introduces cross-modal interaction between the text and layout Transformer streams at every layer while preserving the representation structure of pre-trained language backbones.

    For a single self-attention head with query-key projection dimension dhd_h, let αijT\alpha_{ij}^T and αijL\alpha_{ij}^L denote the dot-product attention scores between position ii and position jj in the text and layout streams, respectively: αij=(xiWQ)(xjWK)⊤dh\alpha_{ij} = \frac{(x_i W^Q)(x_j W^K)^\top}{\sqrt{d_h}}

    BiACM modifies these attention scores across streams: α~ijT=αijL+αijT\tilde{\alpha}_{ij}^T = \alpha_{ij}^L + \alpha_{ij}^T α~ijL={αijL+DETACH(αijT),during Pre-trainingαijL+αijT,during Fine-tuning\tilde{\alpha}_{ij}^L = \begin{cases} \alpha_{ij}^L + \text{DETACH}(\alpha_{ij}^T), & \text{during Pre-training} \\ \alpha_{ij}^L + \alpha_{ij}^T, & \text{during Fine-tuning} \end{cases}

    The operator DETACH(⋅)\text{DETACH}(\cdot) blocks the gradient flow into the text stream during pre-training. This prevents non-textual layout gradients from corrupting the linguistic parameter space of the pre-trained text encoder, which allows the pre-trained layout stream (LiLT) to subsequently pair with any off-the-shelf monolingual or multilingual text model during fine-tuning. The adjusted attention scores are normalized via softmax and used to weight the value vectors in their respective streams.

  3. Knowl 3 — Pre-training Objectives: MVLM, KPL, and CAI

    model/method

    LiLT is pre-trained using three self-supervised objectives designed for joint text-layout representation learning:

    1. Masked Visual-Language Modeling (MVLM): 15% of the text tokens are selected for masking (80% replaced by [MASK], 10% replaced by random tokens from the vocabulary, and 10% kept unchanged), while layout coordinates remain intact. The model predicts the original masked tokens via cross-entropy loss over the vocabulary.

    2. Key Point Location (KPL): The 2D document space is partitioned into a uniform 7×7=497 \times 7 = 49 grid of regions. 15% of bounding boxes are masked (80% replaced by (0,0,0,0,0,0)(0,0,0,0,0,0), 10% replaced by random bounding boxes from the same mini-batch, and 10% kept unchanged). Using three separate classification heads, the model predicts the grid region index containing the top-left corner, bottom-right corner, and center point of each target box using a cross-entropy loss. Predicting discrete grid regions rather than exact pixel coordinates provides tolerance against OCR bounding box noise.

    3. Cross-modal Alignment Identification (CAI): Encoded representations of token-box pairs that have been altered (misaligned) or retained (aligned) by the MVLM and KPL masking routines are classified by a binary classification head with cross-entropy loss to determine whether text and layout components are aligned.

  4. Knowl 4 — Asynchronous Pre-training Optimization Strategy

    model/method

    In standard multi-modal pre-training, end-to-end parameter updates with a uniform learning rate tightly couple the non-textual network to the evolving parameters of the text encoder, hindering transfer to new text backbones.

    To decouple layout knowledge from specific text weights during pre-training, LiLT employs an asynchronous optimization strategy:

    • The learning rate for the text stream is reduced by a factor of 1000 relative to the layout stream (which uses learning rate 2×10−52 \times 10^{-5} with Adam optimizer and linear warmup/decay over 5 epochs on the IIT-CDIP dataset).
    • Scaling down the text learning rate by a factor of 1000 outperforms full parameter freezing (which scores an average F1 of 0.7893 on FUNSD/XFUND vs. 0.7963 for a slow-down ratio of 1000) and outperforms training without slow-down (which scores 0.7840).
    • During downstream fine-tuning, the slow-down ratio is removed (using a uniform learning rate across both modalities), and the DETACH operator in BiACM is deactivated.
  5. Knowl 5 — Multilingual Form Understanding Performance on XFUND Benchmark

    empirical result

    LiLT pre-trained exclusively on 11M English documents from the IIT-CDIP Test Collection 1.0 was evaluated on the multilingual XFUND benchmark across seven non-English languages (Chinese [ZH], Japanese [JA], Spanish [ES], French [FR], Italian [IT], German [DE], Portuguese [PT]) and the English FUNSD dataset for Semantic Entity Recognition (SER) and Relation Extraction (RE).

    When combined with an off-the-shelf multilingual text model (InfoXLMBASE_{\text{BASE}}), LiLT outperforms LayoutXLMBASE_{\text{BASE}} (which required pre-training on 30M multilingual documents) in language-specific fine-tuning (fine-tuning on language XX and testing on language XX):

    Task / Model Pre-train Docs EN ZH JA ES FR IT DE PT Avg.
    SER
    XLM-RoBERTaBASE_{\text{BASE}} - 0.6670 0.8774 0.7761 0.6105 0.6743 0.6687 0.6814 0.6818 0.7047
    InfoXLMBASE_{\text{BASE}} - 0.6852 0.8868 0.7865 0.6230 0.7015 0.6751 0.7063 0.7008 0.7207
    LayoutXLMBASE_{\text{BASE}} Multi 30M 0.7940 0.8924 0.7921 0.7550 0.7902 0.8082 0.8222 0.7903 0.8056
    LiLT[InfoXLM]BASE_{\text{BASE}} EN 11M 0.8415 0.8938 0.7964 0.7911 0.7953 0.8376 0.8231 0.8220 0.8251
    RE
    XLM-RoBERTaBASE_{\text{BASE}} - 0.2659 0.5105 0.5800 0.5295 0.4965 0.5305 0.5041 0.3982 0.4769
    InfoXLMBASE_{\text{BASE}} - 0.2920 0.5214 0.6000 0.5516 0.4913 0.5281 0.5262 0.4170 0.4910
    LayoutXLMBASE_{\text{BASE}} Multi 30M 0.5483 0.7073 0.6963 0.6896 0.6353 0.6415 0.6551 0.5718 0.6432
    LiLT[InfoXLM]BASE_{\text{BASE}} EN 11M 0.6276 0.7297 0.7037 0.7195 0.6965 0.7043 0.6558 0.5874 0.6781

    LiLT achieves an average SER F1 of 0.8251 (compared to 0.8056 for LayoutXLM) and an average RE F1 of 0.6781 (compared to 0.6432 for LayoutXLM).

  6. Knowl 6 — Cross-Lingual Zero-Shot Transfer and Multitask Fine-Tuning Performance

    empirical result

    LiLT was evaluated under cross-lingual zero-shot transfer learning and multi-task learning paradigms across the 8 languages of FUNSD and XFUND:

    1. Cross-Lingual Zero-Shot Transfer: Models are fine-tuned strictly on English FUNSD and evaluated on the seven target XFUND languages without seeing non-English training documents.

      • SER Task: LiLT[InfoXLM]BASE_{\text{BASE}} achieves an average F1 score of 0.6061 (EN: 0.8415, ZH: 0.6152, JA: 0.5184, ES: 0.5101, FR: 0.5923, IT: 0.5371, DE: 0.6013, PT: 0.6325), outperforming LayoutXLMBASE_{\text{BASE}} (average F1 of 0.5561), despite LayoutXLM having seen non-English documents during pre-training.
      • RE Task: LiLT[InfoXLM]BASE_{\text{BASE}} achieves an average F1 score of 0.4930, outperforming LayoutXLMBASE_{\text{BASE}} (0.4388) and InfoXLMBASE_{\text{BASE}} (0.2423).
    2. Multitask Fine-Tuning: Models are simultaneously fine-tuned on all eight languages and evaluated on each language:

      • SER Task: LiLT[InfoXLM]BASE_{\text{BASE}} achieves an average F1 of 0.8585 vs. 0.8201 for LayoutXLMBASE_{\text{BASE}} and 0.7106 for InfoXLMBASE_{\text{BASE}}.
      • RE Task: LiLT[InfoXLM]BASE_{\text{BASE}} achieves an average F1 of 0.8125 vs. 0.7823 for LayoutXLMBASE_{\text{BASE}} and 0.6145 for InfoXLMBASE_{\text{BASE}}.
  7. Knowl 7 — Monolingual Structured Document Understanding Performance

    empirical result

    When paired with language-specific text models, LiLT matches or exceeds dedicated multi-modal architectures across English and Chinese benchmarks without using visual image feature backbones (except for RVL-CDIP classification, where image features are concatenated with [CLS] embeddings):

    • FUNSD (English SER): LiLT[EN-RoBERTa]BASE_{\text{BASE}} reaches an F1 score of 0.8841 (Precision 0.8721, Recall 0.8965), outperforming LayoutLMv2BASE_{\text{BASE}} (0.8276), DocFormerBASE_{\text{BASE}} (0.8334), and BROSBASE_{\text{BASE}} (0.8121).
    • CORD (English Receipt SER): LiLT[EN-RoBERTa]BASE_{\text{BASE}} attains an F1 score of 0.9607, surpassing LayoutLMv2BASE_{\text{BASE}} (0.9495), TILTBASE_{\text{BASE}} (0.9511), and LayoutLMBASE_{\text{BASE}} (0.9472).
    • EPHOIE (Chinese Examination Paper SER): LiLT[ZH-RoBERTa]BASE_{\text{BASE}} achieves an F1 score of 0.9797 (Precision 0.9762, Recall 0.9833), exceeding StrucTexTBASE_{\text{BASE}} (0.9795), LayoutXLMBASE_{\text{BASE}} (0.9759), and text-only Chinese RoBERTaBASE_{\text{BASE}} (0.9521).
    • RVL-CDIP (English Document Classification): LiLT[EN-RoBERTa]BASE_{\text{BASE}} achieves 95.68% accuracy, outperforming LayoutLMv2BASE_{\text{BASE}} (95.25%), LayoutXLMBASE_{\text{BASE}} (95.21%), and BROSBASE_{\text{BASE}} (95.58%).
  8. Knowl 8 — Ablation Analysis of BiACM, Pre-training Tasks, and Optimization Ratios

    empirical result

    Ablation studies on FUNSD and XFUND (using LiLTBASE_{\text{BASE}} pre-trained on a 2M subset of IIT-CDIP for 5 epochs) isolate the impact of individual design choices on downstream SER average F1:

    1. Cross-Modality Interaction Mechanisms:

      • Concatenating text and layout features at the model output without cross-attention yields 0.6751 F1 (worse than the text flow alone at 0.7207).
      • Standard co-attention (passing keys and values across streams) degrades performance to 0.6276 F1 by interfering with text representation consistency.
      • BiACM yields 0.7963 F1.
      • Removing DETACH during pre-training reduces F1 from 0.7963 to 0.7682. Keeping DETACH during fine-tuning lowers F1 to 0.7822.
    2. Pre-training Task Configurations:

      • MVLM alone: 0.7616 F1.
      • MVLM + KPL: 0.7748 F1.
      • MVLM + CAI: 0.7809 F1.
      • MVLM + KPL + CAI: 0.7963 F1.
    3. Text Stream Slow-Down Ratios:

      • Ratio 1 (no slow-down): 0.7840 F1.
      • Ratio 500: 0.7901 F1.
      • Ratio 800: 0.7947 F1.
      • Ratio 1000: 0.7963 F1 (optimal).
      • Ratio 1200: 0.7935 F1.
      • Ratio ∞\infty (parameter freezing): 0.7893 F1.

Coverage note — None was omitted; all primary architectural components (LiLT dual stream, BiACM, layout embeddings), pre-training tasks (MVLM, KPL, CAI), optimization strategies, and full empirical evaluations across eight languages and three evaluation setups are covered.

References

  1. 1.Muhammad Zeshan Afzal, Andreas Kölsch, Sheraz Ahmed, and Marcus Liwicki. 2017. Cutting the error by half: Investigation of very deep CNN and advanced training strategies for document image classification. In ICDAR, volume 1, pages 883–888.
  2. 2.Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha. 2021. DocFormer: End-to-end Transformer for document understanding. In ICCV.
  3. 3.Dario Augusto Borges Oliveira et al. 2017. Fast CNN-based document layout analysis. In ICCV Workshop, pages 1173–1180.
  4. 4.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  5. 5.Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Jianfeng Gao, Songhao Piao, Ming Zhou, et al. 2020. UniLMv2: Pseudo-masked language models for unified language model pre-training. In ICML, pages 642–652.
  6. 6.Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, He-Yan Huang, and Ming Zhou. 2021. InfoXLM: An information-theoretic framework for cross-lingual language model pre-training. In NAACL-HLT, pages 3576–3588.
  7. 7.Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In ACL, pages 8440–8451.
  8. 8.Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Shijin Wang, and Guoping Hu. 2020. Revisiting pre-trained models for Chinese natural language processing. In Findings of EMNLP, pages 657–668.
  9. 9.Arindam Das, Saikat Roy, Ujjwal Bhattacharya, and Swapan K Parui. 2018. Document image classification with intra-domain transfer learning and stacked generalization of deep convolutional neural networks. In ICPR, pages 3180–3185.
  10. 10.Tyler Dauphinee, Nikunj Patel, and Mohammad Rashidi. 2019. Modular multimodal architecture for document classification. arXiv preprint arXiv:1912.04376.
  11. 11.Timo I Denk and Christian Reisswig. 2019. BERT-grid: Contextualized embedding for 2D document representation and understanding. In Workshop on Document Intelligence at NeurIPS.
  12. 12.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional Transformers for language understanding. In NAACL-HLT, pages 4171–4186.
  13. 13.Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, and Filip Graliński. 2021. LAMBERT: Layout-aware (language) modeling using BERT for information extraction. In ICDAR.
  14. 14.Adam W Harley et al. 2015. Evaluation of deep convolutional nets for document image classification and retrieval. In ICDAR, pages 991–995.
  15. 15.Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. 2020. BROS: A pre-trained language model for understanding texts in document.
  16. 16.Guillaume Jaume et al. 2019. FUNSD: A dataset for form understanding in noisy scanned documents. In ICDAR, volume 2, pages 1–6.
  17. 17.Anoop R Katti, Christian Reisswig, Cordula Guder, Sebastian Brarda, Steffen Bickel, Johannes Höhne, and Jean Baptiste Faddoul. 2018. Chargrid: Towards understanding 2D documents. In EMNLP, pages 4459–4469.
  18. 18.Diederik P Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In ICLR.
  19. 19.Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016. Neural architectures for named entity recognition. In NAACL-HLT, pages 260–270.
  20. 20.David Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David Grossman, and Jefferson Heard. 2006. Building a test collection for complex document information processing. In ACM SIGIR, pages 665–666.
  21. 21.Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2021a. StructuralLM: Structural pre-training for form understanding. In ACL.
  22. 22.Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu. 2021b. SelfDoc: Self-supervised document representation learning. In CVPR, pages 5652–5660.
  23. 23.Yulin Li, Yuxi Qian, Yuchen Yu, Xiameng Qin, Chengquan Zhang, Yan Liu, Kun Yao, Junyu Han, Jingtuo Liu, and Errui Ding. 2021c. StrucTexT: Structured text understanding with multi-modal Transformers. In ACM-MM.
  24. 24.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017. Feature pyramid networks for object detection. In CVPR, pages 2117–2125.
  25. 25.Weihong Lin, Qifang Gao, Lei Sun, Zhuoyao Zhong, Kai Hu, Qin Ren, and Qiang Huo. 2021. ViBERT-grid: A jointly trained multi-modal 2D document representation for key information extraction from documents. In ICDAR.
  26. 26.Xiaojing Liu, Feiyu Gao, Qiong Zhang, and Huasha Zhao. 2019a. Graph convolution for multimodal information extraction from visually rich documents. In NAACL-HLT, pages 32–39.
  27. 27.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b. RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
  28. 28.Ilya Loshchilov and Frank Hutter. 2018. Decoupled weight decay regularization. In ICLR.
  29. 29.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. NeurIPS, 32:13–23.
  30. 30.Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019. CORD: A consolidated receipt dataset for post-OCR parsing. In Workshop on Document Intelligence at NeurIPS.
  31. 31.Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka. 2021. Going full-TILT boogie on document understanding with text-image-layout Transformer. In ICDAR.
  32. 32.Yujie Qian, Enrico Santus, Zhijing Jin, Jiang Guo, and Regina Barzilay. 2019. GraphIE: A graph-based framework for information extraction. In NAACL-HLT, pages 751–761.
  33. 33.Ritesh Sarkhel and Arnab Nandi. 2019. Deterministic routing between layout abstractions for multi-scale classification of visually rich documents. In IJCAI, pages 3360–3366.
  34. 34.Noah Siegel, Nicholas Lourie, Russell Power, and Waleed Ammar. 2018. Extracting scientific figures with distantly supervised neural networks. In JCDL, pages 223–232.
  35. 35.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017. Inception-v4, Inception-ResNet and the impact of residual connections on learning. In AAAI, pages 4278–4284.
  36. 36.Guozhi Tang, Lele Xie, Lianwen Jin, Jiapeng Wang, Jingdong Chen, Zhen Xu, Qianying Wang, Yaqiang Wu, and Hui Li. 2021. MatchVIE: Exploiting match relevancy between entities for visual information extraction. In IJCAI, pages 1039–1045.
  37. 37.Jiapeng Wang, Chongyu Liu, Lianwen Jin, Guozhi Tang, Jiaxin Zhang, Shuaitao Zhang, Qianying Wang, Yaqiang Wu, and Mingxiang Cai. 2021a. Towards robust visual information extraction in real world: New dataset and novel solution. In AAAI, volume 35, pages 2738–2745.
  38. 38.Jiapeng Wang, Tianwei Wang, Guozhi Tang, Lianwen Jin, Weihong Ma, Kai Ding, and Yichao Huang. 2021b. Tag, copy or predict: A unified weakly-supervised learning framework for visual information extraction using sequences. In IJCAI, pages 1082–1090.
  39. 39.Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017. Aggregated residual transformations for deep neural networks. In CVPR, pages 1492–1500.
  40. 40.Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al. 2021a. LayoutLMv2: Multi-modal pre-training for visually-rich document understanding. In ACL.
  41. 41.Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020. LayoutLM: Pre-training of text and layout for document image understanding. In ACM-SIGKDD, pages 1192–1200.
  42. 42.Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. 2021b. LayoutXLM: Multimodal pre-training for multilingual visually-rich document understanding. arXiv preprint arXiv:2104.08836.
  43. 43.Xiao Yang, Ersin Yumer, Paul Asente, Mike Kraley, Daniel Kifer, and C Lee Giles. 2017. Learning to extract semantic structure from documents using multimodal fully convolutional neural networks. In CVPR, pages 5315–5324.
  44. 44.Wenwen Yu, Ning Lu, Xianbiao Qi, Ping Gong, and Rong Xiao. 2021. PICK: Processing key information extraction from documents using improved graph learning-convolutional networks. In ICPR, pages 4363–4370.
  45. 45.Peng Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Jing Lu, Liang Qiao, Yi Niu, and Fei Wu. 2020. TRIE: End-to-end text reading and information extraction for document understanding. In ACM-MM, pages 1413–1422.

Citation

MLA
Wang, J., et al. “LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7747–57, https://doi.org/10.18653/v1/2022.acl-long.534.
APA
Wang, J., Jin, L., & Ding, K. (2022). LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7747–7757. https://doi.org/10.18653/v1/2022.acl-long.534
Chicago
Wang, J., L. Jin, and K. Ding. 2022. “LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7747–57. https://doi.org/10.18653/v1/2022.acl-long.534.
Harvard
Wang, J., Jin, L. and Ding, K. (2022) “LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7747–7757. Available at: https://doi.org/10.18653/v1/2022.acl-long.534.
Vancouver
1. Wang J, Jin L, Ding K (2022) LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7747–7757

BibTeX

@inproceedings{wang-etal-2022-lilt,
    title = "{L}i{LT}: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding",
    author = "Wang, Jiapeng  and
      Jin, Lianwen  and
      Ding, Kai",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.534/",
    doi = "10.18653/v1/2022.acl-long.534",
    pages = "7747--7757"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/