DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding
Dongsheng WangNatraj RamanMathieu SibueZhiqiang MaPetr BabkinSimerjot KaurYulong PeiArmineh NourbakhshXiaomo Liu
Presents DocLLM, a lightweight generative model that incorporates document layout into large language models via disentangled spatial attention and text infilling pre-training, achieving superior performance on visual document tasks without the computational expense of vision encoders.
Enterprise workflows depend heavily on visually rich documents such as invoices, legal contracts, receipts, and forms. Traditional large language models process only raw text and overlook formatting cues, while conventional vision-language models rely on heavy, computationally expensive image encoders to process document images. To bridge this gap efficiently, the article introduces DocLLM, a lightweight, layout-aware language model designed to interpret complex structured documents without relying on image encoders.
The main objective of the article is to develop and evaluate a generative model that incorporates spatial layouts solely through bounding box coordinates derived from optical character recognition (OCR). The researchers extended autoregressive decoder architectures—specifically Falcon-1B and Llama2-7B—by introducing a disentangled attention mechanism that models spatial and textual relationships independently. They also developed a block infilling pre-training objective that trains the model to reconstruct masked, coherent blocks of text using surrounding context. The model was pre-trained on a corpus of 16.7 million document pages and instruction-tuned on over 630,000 prompts across four key document intelligence tasks: key information extraction, document classification, visual question answering, and natural language inference.
The evaluation produced several notable findings. DocLLM-7B outperformed comparably sized models on 14 out of 16 datasets when evaluated on unseen splits of known datasets, demonstrating particularly strong gains in layout-intensive tasks like key information extraction and document classification. When tested on held-out datasets unseen during instruction tuning, the model improved performance over base text-only models by 15% to 60% on four out of five benchmarks. The smaller 1-billion parameter variant performed competitively against larger baselines, indicating that layout awareness provides significant architectural leverage. Additionally, ablation tests confirmed that both the disentangled spatial attention and the block infilling objective significantly improved token prediction accuracy over standard causal language modeling.
These results demonstrate that incorporating spatial coordinates alone is sufficient to achieve strong document understanding, eliminating the processing delays and infrastructure costs associated with vision backbones. The approach enables enterprise systems to process multi-page, irregularly formatted documents faster and at lower operational expense while maintaining high extraction accuracy.
Organizations handling document-heavy pipelines should consider adopting layout-aware bounding box representations over text-only or heavyweight vision approaches. Before deploying in production, teams should assess tasks requiring deep numerical or abstract reasoning, where specialized models still hold an advantage. The article notes that performance is bound by context length limits and the quality of the underlying OCR system, though the model remains robust against moderate OCR coordinate noise. Confidence in the reported extraction and layout-parsing capabilities is high across enterprise form domains.
- Paper: LiLT: A Simple yet Effective Language-Independent Layout Transformer for Structured Document Understanding, Jiapeng Wang et al. (2022). LiLT establishes how transformer models integrate textual and spatial layout information for document understanding, clarifying the layout-aware modeling problem that DocLLM reworks for LLMs.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). DocVQA defines a central document-question-answering benchmark and task that helps contextualize DocLLM’s evaluation of document intelligence.
No sufficiently relevant recommendations were found.
