Unifying Vision, Text, and Layout for Universal Document Processing

Zineng TangZiyi YangGuoxin WangYuwei FangYang LiuChenguang ZhuMichael ZengCha ZhangMohit Bansal

article2023CVPR209 citationsHighlight

Proposes UDOP, a unified foundation model that fuses text, layout, and visual modalities into a single sequence-to-sequence framework to achieve state-of-the-art results across diverse document understanding and generation tasks.

Listen

Organizations routinely process high volumes of visually rich documents, such as financial statements, tax forms, receipts, and research reports. Automated document intelligence systems struggle because critical meaning depends on the tight interplay among textual content, visual appearance, and two-dimensional spatial layouts. Existing artificial intelligence methods typically model text and images through separate processing paths and treat spatial coordinates as basic positional markers. Furthermore, previous systems require distinct, manually engineered architectures for individual tasks, limiting their flexibility and performance across diverse business document workflows.

The article demonstrates Universal Document Processing (UDOP), a foundation model designed to unify text, visual, and layout modalities into a single sequence-to-sequence generation framework. The primary objective is to create a versatile architecture capable of performing diverse document understanding and image generation tasks within one system.

The developers evaluated UDOP by pretraining a 794-million-parameter transformer on 11 million unlabeled public scanned documents and 1.8 million labeled examples across 11 supervised datasets. At the input stage, the model integrates text tokens directly with corresponding image patch representations based on layout coordinates. It discretizes bounding box coordinates into location tokens, allowing spatial prediction to function as language generation. The architecture combines self-supervised objectives—such as layout modeling, visual text recognition, and masked image reconstruction—with diverse supervised tasks, using an image resolution curriculum from 224 up to 1024 pixels.

The evaluation yielded several key findings. First, UDOP established state-of-the-art performance across eight standard benchmark tasks, achieving first place on the Document Understanding Benchmark with an average score of 64.8 points and outperforming specialized models like LayoutLMv3. Second, the system set leading accuracy records on specific industry datasets, including a 97.58% score on the Consolidated Receipt Dataset (CORD) and 96.00% on the RVL-CDIP document classification benchmark. Third, ablation testing showed that pretraining with specialized spatial and visual objectives significantly improved accuracy over traditional text-only masked language modeling. Finally, UDOP demonstrated high-fidelity visual generation and editing, allowing users to modify text, headers, and numbers on document images while matching surrounding fonts, styles, and orientations.

These results indicate that unified foundation architectures can replace fragmented, task-specific pipelines in enterprise document processing. Consolidating classification, information extraction, and visual document editing into a single model can reduce engineering maintenance overhead, accelerate deployment timelines, and improve data extraction quality across complex document formats. The capacity to generate realistic document variations also offers a practical way to synthesize training data for rare or sensitive document types.

Organizations evaluating automated document intelligence should consider unified generative architectures for multimodal document processing instead of maintaining separate systems for optical character recognition, layout detection, and text extraction. Implementing high-resolution processing requires substantial computing resources; deploying teams should conduct initial pilot studies on target document workflows to assess the operational trade-offs among model scale, image resolution, and inference latency.

While the reported performance is strong across standard public benchmarks, the findings are bounded by the model's reliance on optical character recognition preprocessing to generate initial text inputs and bounding boxes. Operational confidence in real-world deployments will depend on validating the system against severe document degradations, private enterprise templates, and compliance requirements regarding generative document alterations.

arXiv: 2212.02623
Cover for Unifying Vision, Text, and Layout for Universal Document Processing

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Universal Document Processing
  • 3.1 A Unified Vision, Text, and Layout Encoder
  • 3.2 Vision-Text-Layout Decoder
  • 4 Unified Generative Pretraining
  • 4.1 Self-Supervised Pretraining Tasks
  • 4.2 Supervised Pretraining Tasks
  • 5 Experimental Setup
  • 5.1 Model Pretraining
  • 5.2 Downstream Evaluations
  • 6 Analysis
  • 6.1 Visualization Analysis
  • 6.2 Ablation Analysis
  • 6.3 Effectiveness of the Vision Modality
  • 7 Conclusion
  • References
  • A Appendix Overview
  • B Visualization Analysis
  • C UDOP-Dual Performance
  • D Supervised Pretraining Tasks
  • D.1 Classification
  • D.2 Layout Analysis
  • D.3 Information Extraction
  • D.4 Question Answering
  • D.5 Document NLI
  • D.6 Finetuning Experiment Setting
  • E Curriculum Learning
  • F Performance Variance
  • G Limitations and Societal Impact

Knowls

  1. Knowl 1 — Layout-Induced Vision-Text Representation

    model/method

    To capture the fine-grained spatial correspondence between textual content and image pixels in document processing, Universal Document Processing (UDOP) fuses image patch features, text token embeddings, and layout bounding boxes into a unified representation at the input stage.

    Given a document image v∈RH×W×Cv \in \mathbb{R}^{H \times W \times C}, MM extracted text tokens {si}i=1M\{s_i\}_{i=1}^M, and their bounding boxes {(x1i,y1i,x2i,y2i)}i=1M\{(x_1^i, y_1^i, x_2^i, y_2^i)\}_{i=1}^M normalized to [0,1][0, 1], the image is partitioned into N=HP×WPN = \frac{H}{P} \times \frac{W}{P} non-overlapping patches of size P×P×CP \times P \times C. Each patch is projected to a DD-dimensional vector to form visual embeddings {vj∈RD}j=1N\{v_j \in \mathbb{R}^D\}_{j=1}^N, while text tokens are mapped via vocabulary lookup to embeddings {si∈RD}i=1M\{s_i \in \mathbb{R}^D\}_{i=1}^M.

    The spatial assignment of text tokens to image patches is determined by an indicator function:

    ϕ(si,vj)={1,if the center of si’s bounding box is within image patch vj,0,otherwise.\phi(s_i, v_j) = \begin{cases} 1, & \text{if the center of } s_i\text{'s bounding box is within image patch } v_j, \\ 0, & \text{otherwise.} \end{cases}

    The joint representations fed into the encoder are constructed as follows:

    1. For each text token embedding sis_i, its joint representation is si′=si+vjs_i' = s_i + v_j, where ϕ(si,vj)=1\phi(s_i, v_j) = 1. Text tokens without physical locations (such as task prompts) are assigned coordinates (0,0,0,0)(0, 0, 0, 0) and map to a pseudo image patch.
    2. For image patches vjv_j containing no text token centers (i.e., ∀i,ϕ(si,vj)=0\forall i, \phi(s_i, v_j) = 0), the joint representation is vj′=vjv_j' = v_j. Image patches that contain text tokens are omitted from separate visual input since their features are already incorporated into si′s_i'.

    Continuous bounding box coordinates (x1i,y1i,x2i,y2i)∈[0,1]4(x_1^i, y_1^i, x_2^i, y_2^i) \in [0, 1]^4 are discretized into layout tokens by multiplying each coordinate by the layout vocabulary size (e.g., 500) and rounding to the nearest integer, producing discrete tokens (e.g., ⟨50⟩⟨100⟩⟨250⟩⟨300⟩\langle 50 \rangle \langle 100 \rangle \langle 250 \rangle \langle 300 \rangle) that can be seamlessly interleaved with text tokens.

  2. Knowl 2 — Vision-Text-Layout Transformer Architecture

    model/method

    The Vision-Text-Layout (VTL) Transformer in Universal Document Processing (UDOP) is a multimodal encoder-decoder architecture designed to jointly encode and generate vision, text, and layout modalities:

    1. Modality-Agnostic Unified Encoder: Encodes the combined sequence of layout-induced text embeddings {si′}\{s_i'\} and empty image patch embeddings {vj′}\{v_j'\}. Spatial layout is modeled by adding 2D relative attention bias between tokens based on their 2D bounding box positions rather than standard 1D position embeddings. The encoder follows the T5-large encoder configuration.
    2. Text-Layout Decoder: A unidirectional autoregressive Transformer decoder (initialized from the T5-large decoder) that cross-attends to the unified encoder representations. It generates text and discretized layout tokens sequentially, enabling unified generative text, key information extraction, and bounding box localization.
    3. Vision Decoder: An image decoder built on the Masked Autoencoder (MAE-large) architecture that cross-attends to the unified encoder outputs and character-level text token embeddings to reconstruct pixel-level image patches.

    In total, the unified encoder and text-layout decoder comprise 794M trainable parameters.

  3. Knowl 3 — Self-Supervised Generative Pretraining Objectives

    model/method

    Universal Document Processing (UDOP) pretrains on unlabeled document images and OCR-extracted text bounding boxes using four self-supervised generative objectives formatted with explicit task prompts:

    1. Joint Text-Layout Reconstruction: Text tokens are masked at a 15% masking ratio. Prompted by the context containing sentinel tokens (e.g., <text_layout_0>), the model generates target sequences containing both the masked text tokens and their discretized bounding box layout tokens (e.g., <text_layout_0> Ship Date <100><350><118><372>).
    2. Layout Modeling: Using a high masking ratio of 75%, text spans are retained in the prompt inside sentinel boundaries (e.g., <layout_0> Ship Date </layout_0>), and the target sequence requires predicting only the corresponding discretized 2D bounding box tokens (e.g., <layout_0> <100><350><118><372>).
    3. Visual Text Recognition: Text at masked locations (50% masking ratio) is replaced by its bounding box coordinate tokens in the prompt (e.g., <text_0> <100><350><118><372> </text_0>), and the target sequence generates the corresponding textual tokens (e.g., <text_0> Ship Date).
    4. Masked Image Reconstruction: Masks a portion of the document image patches and reconstructs the raw pixels of the masked regions using the vision decoder conditioned on non-masked patches and text/layout embeddings.
  4. Knowl 4 — Masked Document Image Generation with Character-Level Cross-Attention

    model/method

    UDOP extends the Masked Autoencoder (MAE) framework for visually rich documents to enable joint document image generation and localized editing:

    1. Character-Level Cross-Attention: The vision decoder attends via cross-attention to both the contextual representations from the unified VTL encoder and trainable character embeddings representing the individual characters, digits, and punctuation composing each token. These character embeddings bypass the main encoder and provide fine-grained glyph information that guides pixel rendering with linear computational complexity.
    2. Placeholder Image Decoding: Because the unified encoder processes only non-masked patches fused with text tokens, the vision decoder takes a sequence of trainable placeholder embeddings matching the full grid of target image patches. Two distinct placeholder embeddings indicate whether an image patch was masked or unmasked in the input document image.
    3. Controllable Document Editing: By masking target regions in an input document image and providing new text tokens along with their target layout bounding box tokens in the input prompt, the vision decoder synthesizes replacement patches matching the font style, size, orientation, and surrounding visual context.
  5. Knowl 5 — Multi-Task Supervised Pretraining via Sequence-to-Sequence Generation

    model/method

    UDOP unifies supervised document understanding tasks into a prompt-based sequence-to-sequence generation framework during pretraining across 1.8M examples from 11 datasets:

    1. Document Classification: Prompt: "Document Classification on {Dataset Name} {Text Tokens}" →\to Target: "{Document Class}" (using RVL-CDIP with 16 document categories).
    2. Layout Analysis: Prompt: "Layout Analysis on {Dataset Name} {Entity Name}" →\to Target: "{Entity Name} <x1><y1><x2><y2>..." predicting bounding boxes of entities such as paragraphs or titles (using PubLayNet).
    3. Information Extraction: Prompt: "Information Extraction on {Dataset Name} {Text Query}" →\to Target: "{Entity Label} {Token Bounding Boxes}" (using DocBank, Kleister Charity, PWC, DeepForm).
    4. Document Question Answering: Prompt: "Question Answering on {Dataset Name} {Question} {Document Tokens}" →\to Target: "{Answer}" (using WebSRC, VisualMRC, DocVQA, InfographicsVQA, WTQ).
    5. Document Natural Language Inference: Prompt: "Document Natural Language Inference on {Dataset Name} {Sentence Pair}" →\to Target: "Entailment" or "Not Entailment" (using TabFact).
  6. Knowl 6 — Performance of UDOP on the DUE-Benchmark

    data/table

    The Document Understanding Evaluation (DUE) benchmark evaluates models across seven diverse document understanding tasks spanning Question Answering (DocVQA, InfoVQA), Information Extraction (Kleister Charity / KLC, PWC, DeepForm), and Table QA/NLI (WikiTableQuestions / WTQ, TabFact). UDOP utilizes a single open-vocabulary generative model across all tasks.

    Model Modality DocVQA InfoVQA KLC PWC DeepForm WTQ TabFact Avg.
    Donut V 72.1 - - - - - - -
    BERTlarge_{\text{large}} T 67.5 - - - - - - -
    T5large_{\text{large}} T 70.4 36.7 74.3 25.3 74.4 33.3 58.9 50.7
    T5large_{\text{large}}+U T 76.3 37.1 76.0 27.6 82.9 38.1 76.0 56.5
    T5large_{\text{large}}+2D T+L 69.8 39.2 72.6 25.7 74.0 30.8 58.0 50.4
    T5large_{\text{large}}+2D+U T+L 81.0 46.1 75.9 26.8 83.3 43.3 78.6 59.8
    LAMBERT T+L - - 81.3 - - - - -
    StructuralLMlarge_{\text{large}} T+L 83.9 - - - - - - -
    LayoutLMv2large_{\text{large}} V+T+L 78.8 - - - - - - -
    LayoutLMv3large_{\text{large}} V+T+L 83.4 45.1 77.1 26.9 84.0 45.7 78.1 62.9
    UDOP V+T+L 84.7 47.4 82.8 28.0 85.5 47.2 78.9 64.8

    Modality markers: V = Vision, T = Text, L = Layout. When pre-trained with auxiliary supervised data following the TILT protocol, UDOP scores reach 87.8 on DocVQA and 63.0 on InfoVQA. UDOP achieves state-of-the-art results across all 7 tasks and attains the highest overall average score of 64.8, outperforming LayoutLMv3large_{\text{large}} (62.9).

  7. Knowl 7 — Performance of UDOP on FUNSD, CORD, and RVL-CDIP

    data/table

    UDOP performance was evaluated on standard Document AI benchmarks: form key information extraction on FUNSD (entity F1), receipt parsing on CORD (F1), and document image classification on RVL-CDIP (classification accuracy).

    Model Modality FUNSD (F1) CORD (F1) RVL-CDIP (Acc.)
    Donut V - 91.6 95.3
    BERTlarge_{\text{large}} T 65.63 90.25 89.92
    BROSlarge_{\text{large}} T+L 84.52 97.40 -
    StructuralLMlarge_{\text{large}} T+L 85.14 - 96.08
    LiLT T+L 88.41 96.07 95.68
    FormNet T+L 84.69 97.28 -
    LayoutLMlarge_{\text{large}} T+L 77.89 - 91.90
    SelfDoc V+T+L 83.36 - 92.81
    UniDoc V+T+L 87.93 96.86 95.05
    DocFormerlarge_{\text{large}} V+T+L 84.55 96.99 95.50
    TILTlarge_{\text{large}} V+T+L - 96.33 95.52
    LayoutLMv2large_{\text{large}} V+T+L 84.20 96.01 95.64
    LayoutLMv3large_{\text{large}} V+T+L 92.08 97.46 95.93
    UDOP V+T+L 91.62 97.58 96.00

    UDOP achieves state-of-the-art results on CORD (97.58) and RVL-CDIP (96.00) while competitive on FUNSD (91.62) using a unified sequence-to-sequence generative architecture without task-specific classification heads.

  8. Knowl 8 — Ablation Study on Pretraining Objectives

    data/table

    An ablation study on the validation sets of DocVQA (ANLS metric) and RVL-CDIP (classification accuracy) assessed the cumulative contribution of each pretraining objective in UDOP (using 224×224224 \times 224 image resolution).

    Pretrain Objectives #Pretrain Data DocVQA RVL-CDIP
    MLM 11.0M 79.7±0.479.7 \pm 0.4 95.3±0.395.3 \pm 0.3
    Joint Text-Layout 11.0M 82.8±0.182.8 \pm 0.1 95.4±0.395.4 \pm 0.3
    + Visual Text Recognition 11.0M 83.3±0.283.3 \pm 0.2 95.4±0.295.4 \pm 0.2
    + Layout Modeling 11.0M 84.0±0.384.0 \pm 0.3 95.6±0.295.6 \pm 0.2
    + Image Reconstruction 11.0M 84.4±0.284.4 \pm 0.2 96.2±0.296.2 \pm 0.2
    + Supervised 12.8M 85.0±0.2\mathbf{85.0 \pm 0.2} 96.3±0.1\mathbf{96.3 \pm 0.1}

    Replacing standard Masked Language Modeling (MLM) with Joint Text-Layout reconstruction yields a +3.1+3.1 gain on DocVQA. Successive additions of Visual Text Recognition and Layout Modeling further improve performance to 84.084.0. Adding pixel-level masked image reconstruction provides gains across both tasks (84.484.4 and 96.296.2), and multi-task supervised data further maximizes downstream accuracy (85.085.0 and 96.396.3).

  9. Knowl 9 — UDOP Pretraining Setup and Resolution Curriculum Learning

    experimental setup

    UDOP pretraining incorporates large unlabeled and multi-task supervised document corpora alongside an image resolution curriculum:

    • Model Configuration: 794M trainable parameters combining a T5-large unified encoder and text-layout decoder with an MAE-large vision decoder. Vocabulary is expanded to support special sentinel and discretized layout coordinate tokens.
    • Pretraining Datasets: Unlabeled pretraining uses 11 million scanned documents from the IIT-CDIP Test Collection 1.0. Supervised pretraining incorporates 1.8M labeled instances across 11 datasets (RVL-CDIP, PubLayNet, DocBank, Kleister Charity, PWC, DeepForm, WebSRC, VisualMRC, DocVQA, InfographicsVQA, TabFact).
    • Resolution Curriculum Learning: High document image resolution (1024×10241024 \times 1024) yields (1024/16)2=4096(1024/16)^2 = 4096 patch sequence length. To mitigate training cost, pretraining proceeds through three resolution stages: 224×224→512×512→1024×1024224 \times 224 \to 512 \times 512 \to 1024 \times 1024, training for 1 epoch at each stage. Average DUE-Benchmark performance scales with resolution: 63.9 at 224, 64.3 at 512, and 65.1 at 1024.
    • Optimization: Trained using the Adam optimizer (learning rate 5×10−55 \times 10^{-5}, β1=0.9,β2=0.98\beta_1 = 0.9, \beta_2 = 0.98, weight decay 1×10−21 \times 10^{-2}), batch size 512, and 1,000 warmup steps.
  10. Knowl 10 — Comparison Between Unified Encoder and Dual-Encoder Architectures

    data/table

    To evaluate encoder architectures for multimodal document AI, UDOP (unified VTL encoder, 794M parameters) is compared against UDOP-Dual (a two-tower variant employing separate T5-large text-layout and MAE-large vision encoders, totalling 1098M parameters).

    Model DocVQA InfoVQA KLC PWC DeepForm WTQ TabFact Avg.
    UDOP-Dual 84.4 47.1 81.9 28.0 85.2 46.7 79.5 64.6
    UDOP 84.7 47.4 82.8 28.0 85.5 47.2 78.9 64.8

    The single unified encoder in UDOP outperforms the dual-encoder architecture on 5 out of 7 tasks and achieves a higher overall DUE average (64.8 vs. 64.6) despite using 304M fewer parameters.

Coverage note — None was omitted; all primary architectural innovations, self-supervised and supervised pretraining formulations, visual generation mechanisms, experimental results across benchmarks, and ablation studies from the paper are covered.

References

  1. 1.Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, and R Manmatha. Docformer: End-to-end transformer for document understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 993–1003, 2021. 1, 7, 14, 15
  2. 2.Łukasz Borchmann, Michał Pietruszka, Tomasz Stanislawek, Dawid Jurkiewicz, Michał Turski, Karolina Szyndler, and Filip Graliński. Due: End-to-end document understanding benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), 2021. 2, 6
  3. 3.Lu Chen, Xingyu Chen, Zihan Zhao, Danyang Zhang, Jiabao Ji, Ao Luo, Yuxuan Xiong, and Kai Yu. Websrc: A dataset for web-based structural reading comprehension. arXiv preprint arXiv:2101.09465, 2021. 6, 13
  4. 4.Ting Chen, Saurabh Saxena, Lala Li, David J. Fleet, and Geoffrey Hinton. Pix2seq: A language modeling framework for object detection. In International Conference on Learning Representations, 2022. 4
  5. 5.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. Tabfact: A large-scale dataset for table-based fact verification. arXiv preprint arXiv:1909.02164, 2019. 6, 12, 13
  6. 6.Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision, pages 104–120. Springer, 2020. 3
  7. 7.Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. Unifying vision-and-language tasks via text generation. In International Conference on Machine Learning, pages 1931–1942. PMLR, 2021. 3
  8. 8.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instructionfinetuned language models. arXiv preprint arXiv:2210.11416, 2022. 3
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In NAACL, 2018. 4, 7, 14, 15
  10. 10.Łukasz Garncarek, Rafał Powalski, Tomasz Stanisławek, Bartosz Topolski, Piotr Halama, Michał Turski, and Filip Graliński. Lambert: layout-aware language modeling for information extraction. In International Conference on Document Analysis and Recognition, pages 532–547. Springer, 2021. 1, 3, 7, 14
  11. 11.Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, and Tong Sun. Unidoc: Unified pretraining framework for document understanding. Advances in Neural Information Processing Systems, 34:39–50, 2021. 1, 3, 7, 15
  12. 12.Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang. Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4583–4592, 2022. 1, 3
  13. 13.Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis. Evaluation of deep convolutional nets for document image classification and retrieval. In 2015 13th International Conference on Document Analysis and Recognition (ICDAR), pages 991–995. IEEE, 2015. 1, 2, 6, 12
  14. 14.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. arXiv preprint arXiv:2111.06377, 2021. 2, 4, 5, 6
  15. 15.Teakgyu Hong, DongHyun Kim, Mingi Ji, Wonseok Hwang, Daehyun Nam, and Sungrae Park. Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents. Proceedings of the AAAI Conference on Artificial Intelligence, 36(10):10767–10775, Jun. 2022. 1, 3, 7, 14, 15
  16. 16.Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. Layoutlmv3: Pre-training for document ai with unified text and image masking. arXiv preprint arXiv:2204.08387, 2022. 1, 3, 4, 6, 7, 14, 15
  17. 17.Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849, 2020. 1
  18. 18.Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. Funsd: A dataset for form understanding in noisy scanned documents. In 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW), volume 2, pages 1–6. IEEE, 2019. 2, 6, 12
  19. 19.Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, and Robert Stojnic. Axcell: Automatic extraction of results from machine learning papers. arXiv preprint arXiv:2004.14356, 2020. 6, 12, 13
  20. 20.Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. Ocr-free document understanding transformer. In European Conference on Computer Vision, pages 498–517. Springer, 2022. 3
  21. 21.Geewook Kim, Teakgyu Hong, Moonbin Yim, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. Donut: Document understanding transformer without ocr. arXiv preprint arXiv:2111.15664, 2021. 7, 15
  22. 22.Wonjae Kim, Bokyung Son, and Ildoo Kim. ViLT: Visionand-Language Transformer Without Convolution or Region Supervision. In ICML, 2021. 1
  23. 23.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2014. 6, 14
  24. 24.Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, and Tomas Pfister. Formnet: Structural encoding beyond sequential modeling in form document information extraction. arXiv preprint arXiv:2203.08411, 2022. 1, 7, 14, 15
  25. 25.David Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David Grossman, and Jefferson Heard. Building a test collection for complex document information processing. In Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval, pages 665–666, 2006. 6
  26. 26.Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. Structurallm: Structural pre-training for form understanding. arXiv preprint arXiv:2105.11210, 2021. 1, 7, 14, 15
  27. 27.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557, 2019. 1
  28. 28.Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. Docbank: A benchmark dataset for document layout analysis. arXiv preprint arXiv:2006.01038, 2020. 1, 6, 12
  29. 29.Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, and Hongfu Liu. Selfdoc: Self-supervised document representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5652–5660, 2021. 1, 3, 7, 15
  30. 30.Yulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin, Chengquan Zhang, Yan Liu, Kun Yao, Junyu Han, Jingtuo Liu, and Errui Ding. Structext: Structured text understanding with multimodal transformers. In Proceedings of the 29th ACM International Conference on Multimedia, pages 1912–1920, 2021. 1
  31. 31.Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. Unified-io: A unified model for vision, language, and multi-modal tasks. arXiv preprint arXiv:2206.08916, 2022. 3
  32. 32.Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706, 2022. 6, 12, 13
  33. 33.Minesh Mathew, Dimosthenis Karatzas, and CV Jawahar. Docvqa: A dataset for vqa on document images. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 2200–2209, 2021. 2, 6, 12, 13
  34. 34.Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. Cord: a consolidated receipt dataset for post-ocr parsing. In Workshop on Document Intelligence at NeurIPS 2019, 2019. 2, 6, 12
  35. 35.Panupong Pasupat and Percy Liang. Compositional semantic parsing on semi-structured tables. arXiv preprint arXiv:1508.00305, 2015. 6, 12, 13
  36. 36.Rafał Powalski, Łukasz Borchmann, Dawid Jurkiewicz, Tomasz Dwojak, Michał Pietruszka, and Gabriela Pałka. Going full-tilt boogie on document understanding with textimage-layout transformer. In International Conference on Document Analysis and Recognition, pages 732–747. Springer, 2021. 1, 3, 4, 7, 14, 15
  37. 37.Subhojeet Pramanik, Shashank Mujumdar, and Hima Patel. Towards a multi-modal, multi-task learning based pre-training framework for document representation learning. arXiv preprint arXiv:2009.14457, 2020. 1
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 3, 8
  39. 39.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 2020. 6, 7, 14
  40. 40.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100,000+ questions for machine comprehension of text. In EMNLP, 2016. 15
  41. 41.Tomasz Stanisławek, Filip Graliński, Anna Wróblewska, Dawid Lipiński, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemysław Biecek. Kleister: key information extraction datasets involving long documents with complex layouts. In International Conference on Document Analysis and Recognition, pages 564–579. Springer, 2021. 6, 12
  42. 42.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre-training of generic visuallinguistic representations. In International Conference on Learning Representations, 2019. 3
  43. 43.S Svetlichnaya. Deepform: Understand structured documents at scale, 2020. 6, 12, 13
  44. 44.Hao Tan and Mohit Bansal. Lxmert: Learning cross-modality encoder representations from transformers. In EMNLP, 2019. 1
  45. 45.Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. Visualmrc: Machine reading comprehension on document images. Proceedings of the AAAI Conference on Artificial Intelligence, 35(15):13878–13888, May 2021. 1, 6, 13
  46. 46.Zineng Tang, Jaemin Cho, Jie Lei, and Mohit Bansal. Perceiver-vl: Efficient vision-and-language modeling with iterative latent attention. arXiv preprint arXiv:2211.11701, 2022. 1
  47. 47.Zineng Tang, Jie Lei, and Mohit Bansal. Decembert: Learning from noisy instructional videos via dense captions and entropy minimization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2415–2426, 2021. 1
  48. 48.Jiapeng Wang, Lianwen Jin, and Kai Ding. Lilt: A simple yet effective language-independent layout transformer for structured document understanding. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7747–7757, 2022. 1, 7, 14, 15
  49. 49.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. In International Conference on Machine Learning, pages 23318–23340. PMLR, 2022. 3, 4
  50. 50.Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, et al. Image as a foreign language: Beit pretraining for all vision and vision-language tasks. arXiv preprint arXiv:2208.10442, 2022. 3
  51. 51.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, Oct. 2020. Association for Computational Linguistics. 6
  52. 52.Te-Lin Wu, Cheng Li, Mingyang Zhang, Tao Chen, Spurthi Amba Hombaiah, and Michael Bendersky. Lampret: Layout-aware multimodal pretraining for document understanding. arXiv preprint arXiv:2104.08405, 2021. 1
  53. 53.Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1192–1200, 2020. 1, 3, 6, 7, 15
  54. 54.Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei. Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding. arXiv preprint arXiv:2104.08836, 2021. 1
  55. 55.Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, et al. Layoutlmv2: Multi-modal pre-training for visually-rich document understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2579–2591, 2021. 1, 2, 3, 6, 7, 14, 15
  56. 56.Ziyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant, Dongdong Chen, Yu Shi, Yichong Xu, Yao Qian, Mei Gao, Yi-Ling Chen, et al. i-code: An integrative and composable multimodal learning framework. arXiv preprint arXiv:2205.01818, 2022. 3, 8
  57. 57.Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes. Publaynet: largest dataset ever for document layout analysis. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 1015–1022. IEEE, 2019. 1, 6, 12

Citation

MLA
Tang, Z., et al. “Unifying Vision, Text, and Layout for Universal Document Processing”. arXiv, 2022, http://arxiv.org/abs/2212.02623v3.
APA
Tang, Z., Yang, Z., Wang, G., Fang, Y., Liu, Y., Zhu, C., Zeng, M., Zhang, C., & Bansal, M. (2022). Unifying Vision, Text, and Layout for Universal Document Processing. arXiv. http://arxiv.org/abs/2212.02623v3
Chicago
Tang, Z., Z. Yang, G. Wang, et al. 2022. “Unifying Vision, Text, and Layout for Universal Document Processing”. arXiv. http://arxiv.org/abs/2212.02623v3.
Harvard
Tang, Z. et al. (2022) “Unifying Vision, Text, and Layout for Universal Document Processing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.02623v3.
Vancouver
1. Tang Z, Yang Z, Wang G, Fang Y, Liu Y, Zhu C, Zeng M, Zhang C, Bansal M (2022) Unifying Vision, Text, and Layout for Universal Document Processing. arXiv

BibTeX

@article{tang2022unifying,
  title = {Unifying Vision, Text, and Layout for Universal Document Processing},
  author = {Tang, Zineng and Yang, Ziyi and Wang, Guoxin and Fang, Yuwei and Liu, Yang and Zhu, Chenguang and Zeng, Michael and Zhang, Cha and Bansal, Mohit},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.02623v3},
  eprint = {2212.02623}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE