Unifying Vision, Text, and Layout for Universal Document Processing
Zineng TangZiyi YangGuoxin WangYuwei FangYang LiuChenguang ZhuMichael ZengCha ZhangMohit Bansal
Proposes UDOP, a unified foundation model that fuses text, layout, and visual modalities into a single sequence-to-sequence framework to achieve state-of-the-art results across diverse document understanding and generation tasks.
Organizations routinely process high volumes of visually rich documents, such as financial statements, tax forms, receipts, and research reports. Automated document intelligence systems struggle because critical meaning depends on the tight interplay among textual content, visual appearance, and two-dimensional spatial layouts. Existing artificial intelligence methods typically model text and images through separate processing paths and treat spatial coordinates as basic positional markers. Furthermore, previous systems require distinct, manually engineered architectures for individual tasks, limiting their flexibility and performance across diverse business document workflows.
The article demonstrates Universal Document Processing (UDOP), a foundation model designed to unify text, visual, and layout modalities into a single sequence-to-sequence generation framework. The primary objective is to create a versatile architecture capable of performing diverse document understanding and image generation tasks within one system.
The developers evaluated UDOP by pretraining a 794-million-parameter transformer on 11 million unlabeled public scanned documents and 1.8 million labeled examples across 11 supervised datasets. At the input stage, the model integrates text tokens directly with corresponding image patch representations based on layout coordinates. It discretizes bounding box coordinates into location tokens, allowing spatial prediction to function as language generation. The architecture combines self-supervised objectives—such as layout modeling, visual text recognition, and masked image reconstruction—with diverse supervised tasks, using an image resolution curriculum from 224 up to 1024 pixels.
The evaluation yielded several key findings. First, UDOP established state-of-the-art performance across eight standard benchmark tasks, achieving first place on the Document Understanding Benchmark with an average score of 64.8 points and outperforming specialized models like LayoutLMv3. Second, the system set leading accuracy records on specific industry datasets, including a 97.58% score on the Consolidated Receipt Dataset (CORD) and 96.00% on the RVL-CDIP document classification benchmark. Third, ablation testing showed that pretraining with specialized spatial and visual objectives significantly improved accuracy over traditional text-only masked language modeling. Finally, UDOP demonstrated high-fidelity visual generation and editing, allowing users to modify text, headers, and numbers on document images while matching surrounding fonts, styles, and orientations.
These results indicate that unified foundation architectures can replace fragmented, task-specific pipelines in enterprise document processing. Consolidating classification, information extraction, and visual document editing into a single model can reduce engineering maintenance overhead, accelerate deployment timelines, and improve data extraction quality across complex document formats. The capacity to generate realistic document variations also offers a practical way to synthesize training data for rare or sensitive document types.
Organizations evaluating automated document intelligence should consider unified generative architectures for multimodal document processing instead of maintaining separate systems for optical character recognition, layout detection, and text extraction. Implementing high-resolution processing requires substantial computing resources; deploying teams should conduct initial pilot studies on target document workflows to assess the operational trade-offs among model scale, image resolution, and inference latency.
While the reported performance is strong across standard public benchmarks, the findings are bounded by the model's reliance on optical character recognition preprocessing to generate initial text inputs and bounding boxes. Operational confidence in real-world deployments will depend on validating the system against severe document degradations, private enterprise templates, and compliance requirements regarding generative document alterations.
- Paper: Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks, Jiasen Lu et al. (2022). Unified-IO’s shared token vocabulary for text, images, and bounding boxes provides useful groundwork for UDOP’s sequence-to-sequence treatment of document modalities and spatial coordinates.
- Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework, Peng Wang et al. (2022). OFA’s unified sequence-to-sequence framework, which represents image patches and bounding boxes as tokens, prepares readers for UDOP’s similar generative approach to multimodal document processing.
- Paper: FormNet: Structural Encoding beyond Sequential Modeling in Form Document Information Extraction, Chen-Yu Lee et al. (2022). FormNet shows how two-dimensional geometry preserves relationships in forms that text serialization can lose, clarifying the layout problem UDOP addresses through unified modeling.
- Paper: LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding, Jiapeng Wang et al. (2022). LiLT’s separate text and layout streams illustrate an earlier strategy for structured document understanding that makes UDOP’s unified multimodal representation easier to appreciate.
- Paper: DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding, Dongsheng Wang et al. (2024). DocLLM carries layout-aware generative document modeling forward with a lightweight architecture that uses OCR-derived coordinates instead of image encoders.
- Paper: OMNIPARSER: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition, Jianqiang Wan et al. (2024). OmniParser extends unified document processing toward end-to-end text spotting, information extraction, and table recognition in a single vision-based framework.
- Paper: Qianfan-OCR: A Unified End-to-End Model for Document Intelligence, Daxiang Dong et al. (2026). Qianfan-OCR continues the unified document-intelligence direction by turning document images directly into structured outputs for parsing and comprehension.
