DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding

Dongsheng WangNatraj RamanMathieu SibueZhiqiang MaPetr BabkinSimerjot KaurYulong PeiArmineh NourbakhshXiaomo Liu

article2024ACL185 citations

Presents DocLLM, a lightweight generative model that incorporates document layout into large language models via disentangled spatial attention and text infilling pre-training, achieving superior performance on visual document tasks without the computational expense of vision encoders.

Listen

Enterprise workflows depend heavily on visually rich documents such as invoices, legal contracts, receipts, and forms. Traditional large language models process only raw text and overlook formatting cues, while conventional vision-language models rely on heavy, computationally expensive image encoders to process document images. To bridge this gap efficiently, the article introduces DocLLM, a lightweight, layout-aware language model designed to interpret complex structured documents without relying on image encoders.

The main objective of the article is to develop and evaluate a generative model that incorporates spatial layouts solely through bounding box coordinates derived from optical character recognition (OCR). The researchers extended autoregressive decoder architectures—specifically Falcon-1B and Llama2-7B—by introducing a disentangled attention mechanism that models spatial and textual relationships independently. They also developed a block infilling pre-training objective that trains the model to reconstruct masked, coherent blocks of text using surrounding context. The model was pre-trained on a corpus of 16.7 million document pages and instruction-tuned on over 630,000 prompts across four key document intelligence tasks: key information extraction, document classification, visual question answering, and natural language inference.

The evaluation produced several notable findings. DocLLM-7B outperformed comparably sized models on 14 out of 16 datasets when evaluated on unseen splits of known datasets, demonstrating particularly strong gains in layout-intensive tasks like key information extraction and document classification. When tested on held-out datasets unseen during instruction tuning, the model improved performance over base text-only models by 15% to 60% on four out of five benchmarks. The smaller 1-billion parameter variant performed competitively against larger baselines, indicating that layout awareness provides significant architectural leverage. Additionally, ablation tests confirmed that both the disentangled spatial attention and the block infilling objective significantly improved token prediction accuracy over standard causal language modeling.

These results demonstrate that incorporating spatial coordinates alone is sufficient to achieve strong document understanding, eliminating the processing delays and infrastructure costs associated with vision backbones. The approach enables enterprise systems to process multi-page, irregularly formatted documents faster and at lower operational expense while maintaining high extraction accuracy.

Organizations handling document-heavy pipelines should consider adopting layout-aware bounding box representations over text-only or heavyweight vision approaches. Before deploying in production, teams should assess tasks requiring deep numerical or abstract reasoning, where specialized models still hold an advantage. The article notes that performance is bound by context length limits and the quality of the underlying OCR system, though the model remains robust against moderate OCR coordinate noise. Confidence in the reported extraction and layout-parsing capabilities is high across enterprise form domains.

No sufficiently relevant recommendations were found.

Cover for DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding

Abstract

Enterprise documents such as forms, receipts, reports, and other such records, often carry rich semantics at the intersection of textual and spatial modalities. The visual cues offered by their complex layouts play a crucial role in comprehending these documents effectively. In this paper, we present DocLLM, a lightweight extension to traditional large language models (LLMs) for reasoning over visual documents, taking into account both textual semantics and spatial layout. Our model differs from existing multimodal LLMs by avoiding expensive image encoders and focuses exclusively on bounding box information to incorporate the spatial layout structure. Specifically, the cross-alignment between text and spatial modalities is captured by decomposing the attention mechanism in classical transformers to a set of disentangled matrices. Furthermore, we devise a pre-training objective that learns to infill text segments. This approach allows us to address irregular layouts and heterogeneous content frequently encountered in visual documents. The pre-trained model is fine-tuned using a large-scale instruction dataset, covering four core document intelligence tasks. We demonstrate that our solution outperforms SotA LLMs on 14 out of 16 datasets across all tasks, and generalizes well to 4 out of 5 previously unseen datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 DocLLM Framework
  • 3.1 Architecture Overview
  • 3.2 Disentangled Spatial Attention
  • 3.3 Pretraining
  • 3.4 Instruction Tuning
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Evaluation Setup
  • 4.3 Results
  • 5 Discussion
  • 6 Conclusions
  • Acknowledgments
  • References
  • A Dataset Details
  • A.1 Preprocessing
  • A.2 Instruction Tuning Templates
  • A.3 Dataset Statistics
  • B Training Details
  • C Detailed Performance Analysis
  • C.1 Qualitative Examples
  • C.2 DocVQA Deep-Dive
  • C.3 GPT4V Performance Comparison
  • C.4 SotA Performance Comparison
  • D Ablation Studies
  • E Robustness to inaccurate OCR Bounding Boxes
  • F Additional Discussion

Knowls

  1. Knowl 1 — Disentangled attention combines OCR text and spatial layout

    model/method

    DocLLM extends a causal decoder language model to use OCR text tokens and their bounding boxes without an image encoder. For a sequence of TT tokens, let H∈RT×dH \in \mathbb{R}^{T\times d} be the text hidden states, and let S∈RT×dS \in \mathbb{R}^{T\times d} be spatial hidden states encoding each token’s OCR bounding box (left,top,right,bottom)(\mathrm{left},\mathrm{top},\mathrm{right},\mathrm{bottom}). The coordinate units and the precise box-embedding function are not specified. For one attention head, text and spatial query/key projections are Qt=HWt,qQ^t=HW_{t,q}, Kt=HWt,kK^t=HW_{t,k}, Qs=SWs,qQ^s=SW_{s,q}, and Ks=SWs,kK^s=SW_{s,k}, where each WW is a learned d×dd\times d matrix. The attention logits combine four interaction types:

    Ai,j=Qit(Kjt)⊤+λt,sQit(Kjs)⊤+λs,tQis(Kjt)⊤+λs,sQis(Kjs)⊤.A_{i,j}=Q^t_i(K^t_j)^\top+\lambda_{t,s}Q^t_i(K^s_j)^\top+\lambda_{s,t}Q^s_i(K^t_j)^\top+\lambda_{s,s}Q^s_i(K^s_j)^\top.

    Here i,j∈{1,…,T}i,j\in\{1,\ldots,T\} index query and key tokens, and the λ\lambda coefficients weight text-to-spatial, spatial-to-text, and spatial-to-spatial interactions; text-to-text has unit weight. The resulting logits are scaled by d\sqrt d, masked causally, and normalized with softmax. The values come from text states, Vt=HWt,vV^t=HW_{t,v}, so the next hidden states are H′=softmax(A/d)VtH'=\mathrm{softmax}(A/\sqrt d)V^t (with the causal mask applied). The same spatial states SS are reused across layers, while each layer can have its own projection matrices. This keeps text and layout representations distinct while allowing layout to influence token attention.

  2. Knowl 2 — Autoregressive infilling predicts coherent document text blocks

    model/method

    DocLLM’s self-supervised pretraining masks coherent OCR text blocks rather than isolated tokens, to provide broader context for irregularly arranged documents. Let xx be the document token sequence partitioned into KK non-overlapping contiguous blocks c1,…,cKc_1,\ldots,c_K. The method samples M≪KM\ll K blocks z1,…,zMz_1,\ldots,z_M from that partition, replaces each sampled block in the document with a mask token [M][M], and shuffles the sampled blocks for sequential reconstruction. Each block to be generated is introduced in the infilling input by a start token [S][S]; its generated text ends with [E][E]. The model is not given the number of tokens in a masked block. Because the unmasked document includes tokens both before and after each masked block, the target can be conditioned on document prefix and suffix context as well as on earlier reconstructed blocks.

    The paper’s cross-entropy objective for the sampled text tokens is

    LIF(θ)=−∑m=1M∑j=1Nmlog⁡pθ ⁣(zm,j∣x~,z<m,zm,<j).\mathcal{L}_{\mathrm{IF}}(\theta)=-\sum_{m=1}^{M}\sum_{j=1}^{N_m}\log p_\theta\!\left(z_{m,j}\mid\tilde{x},z_{<m},z_{m,<j}\right).

    Here θ\theta denotes the model parameters, x~\tilde{x} is the document with sampled blocks masked, NmN_m is the number of text tokens in block zmz_m, zm,jz_{m,j} is token jj of that block, z<mz_{<m} denotes previously reconstructed sampled blocks, and zm,<jz_{m,<j} denotes earlier tokens in the current block. The displayed loss sums over block text tokens; the generated block is also delimited by the end token. OCR block boundaries are used for pretraining but are not supplied to the model as a masked-block-length signal.

  3. Knowl 3 — Pretraining corpus, instruction data, and model configurations

    experimental setup

    DocLLM was continued from pretrained language-model weights: DocLLM-1B uses Falcon-1B and DocLLM-7B uses Llama2-7B. The variants have 1,524,963,328 and 7,853,019,136 parameters, respectively; their layer counts, attention heads, hidden sizes, and maximum context lengths are 24, 16, 1,536, and 1,024 for DocLLM-1B, and 36, 32, 4,096, and 1,024 for DocLLM-7B. Spatial attention is added to the backbone, which is then pretrained and instruction-tuned.

    Pretraining used OCR text and boxes from IIT-CDIP Test Collection 1.0 and DocBank. The combined corpus contained 5,592,245 documents, 16,792,962 pages, and 3,865,913,752 tokens; the authors ran one pretraining epoch. IIT-CDIP and DocBank were OCR-processed with Tesseract, while the other datasets used the OCR output provided by their publishers.

    Instruction tuning used OCR documents and prompt templates for four task families: visual question answering (VQA), natural-language inference (NLI), key-information extraction (KIE), and document classification (CLS). The data mix included DocVQA, WikiTableQuestions, VisualMRC, DUDE, and BuDDIE for VQA; TabFact for NLI; Kleister Charity, CORD, FUNSD, DeepForm, PWC, SROIE, VRDU ad-buy, and BuDDIE for KIE; and RVL-CDIP and BuDDIE for CLS. KIE prompts included extraction, internal classification, and multiple-choice formats; CLS used internal classification and multiple-choice formats; VQA and NLI each used one principal prompt format. The multiple-choice prompts used randomly selected subsets of key or class names where applicable. The resulting instruction data comprised 635,883 training prompts and 96,919 test prompts: 145,090/24,347 VQA, 104,360/12,720 NLI, 236,806/38,039 KIE, and 149,627/21,813 CLS.

    The models were instruction-tuned for 10 epochs (1B) or 3 epochs (7B). Both used a 1,024-token maximum context length. Pretraining and instruction-tuning learning rates were 2×10−42\times10^{-4} and 1×10−41\times10^{-4} for DocLLM-1B, and 3×10−43\times10^{-4} and 1×10−41\times10^{-4} for DocLLM-7B; the authors used cosine schedules, 1,000 pretraining warmup steps, 500 instruction-tuning warmup steps, and weight decay 0.1. DocLLM-7B training used eight 24-GB A10G GPUs; DocLLM-1B used one such GPU.

  4. Knowl 4 — Same-dataset evaluation results across 16 datasets

    data/table

    In the same-datasets, different-splits (SDDS) evaluation, DocLLM was instruction-tuned using each benchmark’s training data and evaluated on its held-out test split, or dev split when a public test set was unavailable. GPT-4+OCR and Llama2+OCR are zero-shot text-only baselines; mPLUG-DocOwl and UReader, like DocLLM, are instruction-tuned in this setting. The table reports the paper’s metrics: ANLS for VQA except where WTQ uses accuracy and VisualMRC uses CIDEr; accuracy for NLI and CLS; and F1 for KIE. Dashes indicate results not reported. DocLLM-7B outperformed the other compared models on 12 of 16 datasets, and outperformed the comparable models excluding GPT-4 on 14 of 16.

    Task Dataset GPT-4+OCR Llama2+OCR DocOwl UReader DocLLM-1B DocLLM-7B
    VQA DocVQA 82.8 47.4 62.2 65.4 61.4 69.5
    VQA WTQ (accuracy) 65.4 25.0 26.9 29.4 21.9 27.1
    VQA VisualMRC (CIDEr) 255.1 115.5 188.8 221.7 245.0 264.1
    VQA DUDE* 54.6 38.1 – – 42.6 47.2
    VQA BuDDIE 76.4 48.8 – – 84.5 86.7
    NLI TabFact 77.1 48.2 60.2 67.6 58.0 66.4
    KIE KLC 45.9 27.8 30.3 32.8 58.9 60.3
    KIE CORD 58.3 13.8 – – 66.9 67.4
    KIE FUNSD 37.0 17.8 – – 48.2 51.8
    KIE DeepForm 42.1 20.5 42.6 49.5 71.3 75.7
    KIE PWC 18.3 6.8 – – 25.7 29.06
    KIE SROIE 90.6 56.4 – – 91.0 91.9
    KIE VRDU ad-buy* 43.7 18.7 – – 87.6 88.8
    KIE BuDDIE 66.1 10.8 – – 95.4 96.0
    CLS RVL-CDIP 68.2 32.8 – – 90.9 91.8
    CLS BuDDIE 84.9 40.9 – – 98.3 99.4

    The results show particularly large gains over the text-only Llama2 baseline on layout-intensive KIE and classification datasets. DocLLM-7B also exceeds GPT-4+OCR on several KIE and CLS datasets, while GPT-4+OCR remains stronger on some VQA and NLI benchmarks.

  5. Knowl 5 — Held-out datasets test transfer across document domains

    empirical result

    For the same-tasks, different-datasets (STDD) evaluation, DocLLM was instruction-tuned on 11 of the 16 task-dataset collections used in the in-domain evaluation and tested on held-out DocVQA, Kleister Charity (KLC), and BuDDIE data. This tests transfer to new dataset domains and layouts while keeping the task families represented during training. The metrics are ANLS for VQA, F1 for KIE, and accuracy for CLS. DocLLM-7B exceeds the Llama2+OCR zero-shot baseline on four of the five held-out dataset-task pairs, and has the highest reported score on the two held-out KIE pairs. Its held-out BuDDIE classification score is substantially lower than GPT-4+OCR’s.

    Dataset Task GPT-4+OCR (ZS) Llama2+OCR (ZS) DocLLM-1B DocLLM-7B
    DocVQA VQA 82.8 47.4 53.5 63.4
    KLC KIE 45.9 27.8 40.1 49.9
    BuDDIE VQA 76.4 48.4 65.5 73.3
    BuDDIE KIE 66.1 10.8 63.0 72.6
    BuDDIE CLS 84.9 40.9 20.8 31.1

    The reported comparison with document-oriented multimodal instruction models also found DocLLM-7B exceeded mPLUG-DocOwl on DocVQA and both mPLUG-DocOwl and UReader on KLC, despite those baselines being instruction-tuned on the evaluated datasets.

  6. Knowl 6 — Spatial-attention ablation favors spatial-to-spatial interaction

    empirical result

    The authors compared spatial-attention configurations using out-of-sample next-token-prediction (NTP) accuracy during pretraining. For this ablation, 100,000 randomly sampled chunks were used for training and 1,000 unseen documents for evaluation. In the table, T and S denote text and spatial modalities; T2T, T2S, S2T, and S2S denote text-to-text, text-to-spatial, spatial-to-text, and spatial-to-spatial attention interactions. The configuration that retained text-to-text and spatial-to-spatial interactions achieved the best score, 39.12. Text-only attention scored 35.43. DocLLM’s main experiments therefore used λs,s=1\lambda_{s,s}=1 and λt,s=λs,t=0\lambda_{t,s}=\lambda_{s,t}=0.

    Mode Interactions NTP accuracy
    Additive SEmbed + TEmbed 38.16
    Disentangled T2T 35.43
    Disentangled T2S + T2T 38.08
    Disentangled S2T + T2T 38.05
    Disentangled S2S + T2T 39.12
    Disentangled T2S + S2S + T2T 39.06
    Disentangled S2T + S2S + T2T 39.07
    Disentangled T2S + S2T + S2S + T2T 39.02

    The small differences among the other disentangled configurations indicate that adding spatial information is beneficial in this pretraining comparison, but the spatial-to-spatial-only addition was the strongest tested configuration.

  7. Knowl 7 — Block infilling improves pretraining accuracy over causal prediction

    empirical result

    On the authors’ pretraining-stage NTP evaluation, causal learning without spatial features reached 32.6 accuracy; adding spatial features to causal learning raised accuracy to 36.2; and block infilling with spatial features reached 39.1. Thus the reported comparison supports gains from both spatial input and the block-infilling objective, with the combined configuration scoring highest. A separate decoder-mask comparison found only marginal differences between prefix and causal decoders across five spatial-attention configurations, with the causal decoder slightly ahead; DocLLM retained the causal decoder.

  8. Knowl 8 — DocVQA performance is strongest on layout-dependent questions

    empirical result

    DocLLM-7B’s DocVQA ANLS varied by question category. The modality label indicates the signal expected to be uniquely useful for that category: V is visual, L is layout, and T is text; it does not imply other modalities are irrelevant. The model scored highest on form and layout questions (82.2 and 72.4), where spatial organization is central. Scores were lower on image/photo and figure/diagram questions (47.8 and 41.4), which rely more on visual content that DocLLM does not encode, and on yes/no questions (43.9).

    Question category Expected modality ANLS
    Figure/Diagram V 41.4
    Form L 82.2
    Table/List L 66.2
    Layout L 72.4
    Free text T 64.6
    Image/Photo V 47.8
    Handwritten T 62.8
    Yes/No – 43.9
    Other – 56.8
  9. Knowl 9 — DocLLM-1B is stable under moderate bounding-box noise

    empirical result

    The authors tested DocLLM-1B on DocVQA after perturbing each OCR bounding-box border. A border was shifted by ϵ∼N(0,l2σ2)\epsilon\sim\mathcal{N}(0,l^2\sigma^2), where ll is the length of the box side orthogonal to that border and σ\sigma controls noise. Shifts were clipped to ±2lσ\pm2l\sigma, and σ\sigma was limited to the range 0 to 0.25 to avoid swapping box borders. DocVQA performance remained close to the unperturbed score as noise increased: 61.4 at σ=0\sigma=0, 60.9 at σ=0.125\sigma=0.125, and 60.8 at σ=0.25\sigma=0.25. This experiment supports robustness to the tested levels of coordinate perturbation, not to arbitrary OCR errors.

  10. Knowl 10 — Scope and known limitations of DocLLM

    limitation

    DocLLM uses text and OCR bounding boxes but no image encoder, so it can miss image-specific evidence and performs relatively poorly on DocVQA image/photo and figure/diagram questions. The paper also notes failures on counting and complex reasoning, particularly when numerical understanding is required. Its 1,024-token context limits long-document processing; one reported extraction failure occurred because the answer was on a fourth page outside the model’s context. The pretraining data and evaluation focus on English-language enterprise documents, so biases and weaker transfer to other document domains, such as presentations or marketing reports, remain possible. Performance may also be affected by inaccurate OCR text or bounding boxes. Finally, the authors could evaluate transfer on only five held-out dataset-task pairs because instruction-tuning experiments were costly.

Coverage note — Detailed prompt wording and qualitative examples were omitted because their task formats and main implications are captured in the method, setup, results, and limitations knowls.

References

  1. 1.Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023. Palm 2 technical report. arXiv preprint arXiv:2305.10403.
  2. 2.Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. 2023. Qwen-vl: A frontier large vision-language model with versatile abilities. CoRR, abs/2308.12966.
  3. 3.Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. 2022. Efficient training of language models to fill in the middle. arXiv preprint arXiv:2207.14255.
  4. 4.Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez, Marçal Rusiñol, Minesh Mathew, C. V. Jawahar, Ernest Valveny, and Dimosthenis Karatzas. 2019. ICDAR 2019 competition on scene text visual question answering. In 2019 International Conference on Document Analysis and Recognition, ICDAR 2019, Sydney, Australia, September 20-25, 2019, pages 1563–1570. IEEE.
  5. 5.Lukasz Borchmann, Michal Pietruszka, Tomasz Stanislawek, Dawid Jurkiewicz, Michal Turski, Karolina Szyndler, and Filip Gralinski. 2021. DUE: end-to-end document understanding benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual.
  6. 6.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners.
  7. 7.Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. Tabfact: A large-scale dataset for table-based fact verification. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed H. Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei. 2022. Scaling instruction-finetuned language models. CoRR, abs/2210.11416.
  10. 10.Lei Cui, Yiheng Xu, Tengchao Lv, and Furu Wei. 2021. Document ai: Benchmarks, models and applications. arXiv preprint arXiv:2111.08609.
  11. 11.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. CoRR, abs/2305.06500.
  12. 12.Brian L. Davis, Bryan S. Morse, Brian L. Price, Chris Tensmeyer, Curtis Wigington, and Vlad I. Morariu. 2022. End-to-end document recognition and understanding with dessurt. In Computer Vision - ECCV 2022 Workshops - Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part IV, volume 13804 of Lecture Notes in Computer Science, pages 280–296. Springer.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations.
  14. 14.Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2021. Glm: General language model pretraining with autoregressive blank infilling. arXiv preprint arXiv:2103.10360.
  15. 15.Adam W. Harley, Alex Ufkes, and Konstantinos G. Derpanis. 2015. Evaluation of deep convolutional nets for document image classification and retrieval. In 13th International Conference on Document Analysis and Recognition, ICDAR 2015, Nancy, France, August 23-26, 2015, pages 991–995. IEEE Computer Society.
  16. 16.Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2020. Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
  17. 17.Thomas Hegghammer. 2022. Ocr with tesseract, amazon textract, and google document ai: a benchmarking experiment. Journal of Computational Social Science, 5(1):861–882.
  18. 18.Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei. 2022. Layoutlmv3: Pre-training for document ai with unified text and image masking. In Proceedings of the 30th ACM International Conference on Multimedia, pages 4083–4091.
  19. 19.Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and C. V. Jawahar. 2019. Icdar2019 competition on scanned receipt ocr and information extraction. In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 1516–1520.
  20. 20.Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019. FUNSD: A dataset for form understanding in noisy scanned documents. In 2nd International Workshop on Open Services and Tools for Document Analysis, OST@ICDAR 2019, Sydney, Australia, September 22-25, 2019, pages 1–6. IEEE.
  21. 21.Marcin Kardas, Piotr Czapla, Pontus Stenetorp, Sebastian Ruder, Sebastian Riedel, Ross Taylor, and Robert Stojnic. 2020. AxCell: Automatic extraction of results from machine learning papers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8580–8594, Online. Association for Computational Linguistics.
  22. 22.Geewook Kim, Teakgyu Hong, Moonbin Yim, JeongYeon Nam, Jinyoung Park, Jinyeong Yim, Wonseok Hwang, Sangdoo Yun, Dongyoon Han, and Seunghyun Park. 2022. Ocr-free document understanding transformer. In Computer Vision – ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXVIII, page 498–517, Berlin, Heidelberg. Springer-Verlag.
  23. 23.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  24. 24.Arjun Reddy Kunduru. 2023. From data entry to intelligence: Artificial intelligence’s impact on financial system workflows. International Journal on Orange Technologies, 5(8):38–45.
  25. 25.Jordy Van Landeghem, Rubèn Tito, Lukasz Borchmann, Michal Pietruszka, Pawel Józiak, Rafal Powalski, Dawid Jurkiewicz, Mickaël Coustaty, Bertrand Anckaert, Ernest Valveny, Matthew B. Blaschko, Sien Moens, and Tomasz Stanislawek. 2023. Document understanding dataset and evaluation (DUDE). CoRR, abs/2305.08455.
  26. 26.Chen-Yu Lee, Chun-Liang Li, Timothy Dozat, Vincent Perot, Guolong Su, Nan Hua, Joshua Ainslie, Renshen Wang, Yasuhisa Fujii, and Tomas Pfister. 2022. FormNet: Structural encoding beyond sequential modeling in form document information extraction. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3735–3754, Dublin, Ireland. Association for Computational Linguistics.
  27. 27.Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023. Pix2struct: Screenshot parsing as pretraining for visual language understanding. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 18893–18912. PMLR.
  28. 28.D. Lewis, G. Agam, S. Argamon, O. Frieder, D. Grossman, and J. Heard. 2006. Building a test collection for complex document information processing. In Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’06, page 665–666, New York, NY, USA. Association for Computing Machinery.
  29. 29.Chenliang Li, Bin Bi, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. 2021. StructuralLM: Structural pre-training for form understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 6309–6318. Association for Computational Linguistics.
  30. 30.Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei. 2022. Dit: Self-supervised pre-training for document image transformer. In Proceedings of the 30th ACM International Conference on Multimedia, pages 3530–3539.
  31. 31.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. arXiv preprint arXiv:2301.12597.
  32. 32.Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. 2020. DocBank: A benchmark dataset for document layout analysis. In Proceedings of the 28th International Conference on Computational Linguistics, pages 949–960, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  33. 33.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023a. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  34. 34.Tianyang Liu, Fei Wang, and Muhao Chen. 2023b. Rethinking tabular data understanding with large language models. CoRR, abs/2312.16702.
  35. 35.Yuliang Liu, Zhang Li, Hongliang Li, Wenwen Yu, Mingxin Huang, Dezhi Peng, Mingyu Liu, Mingrui Chen, Chunyuan Li, Lianwen Jin, and Xiang Bai. 2023c. On the hidden mystery of OCR in large multimodal models. CoRR, abs/2305.07895.
  36. 36.Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang. 2023. An empirical study of catastrophic forgetting in large language models during continual fine-tuning. arXiv preprint arXiv:2308.08747.
  37. 37.Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq R. Joty, and Enamul Hoque. 2022. Chartqa: A benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 2263–2279. Association for Computational Linguistics.
  38. 38.Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and C. V. Jawahar. 2022. Infographicvqa. In IEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2022, Waikoloa, HI, USA, January 3-8, 2022, pages 2582–2591. IEEE.
  39. 39.Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar. 2021. Docvqa: A dataset for VQA on document images. In IEEE Winter Conference on Applications of Computer Vision, WACV 2021, Waikoloa, HI, USA, January 3-8, 2021, pages 2199–2208. IEEE.
  40. 40.Zihang Meng, Licheng Yu, Ning Zhang, Tamara L Berg, Babak Damavandi, Vikas Singh, and Amy Bearman. 2021. Connecting what to say with where to look by modeling human attention traces. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12679–12688.
  41. 41.Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Nick Barnes, and Ajmal Mian. 2023. A comprehensive overview of large language models. CoRR, abs/2307.06435.
  42. 42.OpenAI. 2023. Gpt-4 technical report.
  43. 43.Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In NeurIPS.
  44. 44.Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee. 2019. Cord: A consolidated receipt dataset for post-ocr parsing.
  45. 45.Panupong Pasupat and Percy Liang. 2015. Compositional semantic parsing on semi-structured tables. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, ACL 2015, July 26-31, 2015, Beijing, China, Volume 1: Long Papers, pages 1470–1480. The Association for Computer Linguistics.
  46. 46.Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116.
  47. 47.Tianxiao Shen, Hao Peng, Ruoqi Shen, Yao Fu, Zaid Harchaoui, and Yejin Choi. 2023. Film: Fill-in language models for any-order generation. arXiv preprint arXiv:2310.09930.
  48. 48.Tomasz Stanislawek, Filip Gralinski, Anna Wróblewska, Dawid Lipinski, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemyslaw Biecek. 2021. Kleister: Key information extraction datasets involving long documents with complex layouts. In 16th International Conference on Document Analysis and Recognition, ICDAR 2021, Lausanne, Switzerland, September 5-10, 2021, Proceedings, Part I, volume 12821 of Lecture Notes in Computer Science, pages 564–579. Springer.
  49. 49.Stacey Svetlichnaya. 2020. Deepform: Understand structured documents at scale.
  50. 50.Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida. 2021. Visualmrc: Machine reading comprehension on document images. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 13878–13888. AAAI Press.
  51. 51.Zineng Tang, Ziyi Yang, Guoxin Wang, Yuwei Fang, Yang Liu, Chenguang Zhu, Michael Zeng, Cha Zhang, and Mohit Bansal. 2023. Unifying vision, text, and layout for universal document processing. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 19254–19264. IEEE.
  52. 52.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  53. 53.Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pages 4566–4575. IEEE Computer Society.
  54. 54.Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuznia, Krima Doshi, Kuntal Kumar Pal, Maitreya Patel, Mehrad Moradshahi, Mihir Parmar, Mirali Purohit, Neeraj Varshney, Phani Rohitha Kaza, Pulkit Verma, Ravsehaj Singh Puri, Rushang Karia, Savan Doshi, Shailaja Keyur Sampat, Siddhartha Mishra, Sujan Reddy A, Sumanta Patro, Tanay Dixit, and Xudong Shen. 2022. Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 5085–5109, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  55. 55.Zilong Wang, Yichao Zhou, Wei Wei, Chen-Yu Lee, and Sandeep Tata. 2023. VRDU: A benchmark for visually-rich document understanding. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2023, Long Beach, CA, USA, August 6-10, 2023, pages 5184–5193. ACM.
  56. 56.Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned language models are zero-shot learners. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net.
  57. 57.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671.
  58. 58.Te-Lin Wu, Cheng Li, Mingyang Zhang, Tao Chen, Spurthi Amba Hombaiah, and Michael Bendersky. 2021. Lampret: Layout-aware multimodal pretraining for document understanding. arXiv preprint arXiv:2104.08405.
  59. 59.Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2021. LayoutLMv2: Multi-modal pre-training for visually-rich document understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2579–2591, Online. Association for Computational Linguistics.
  60. 60.Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020. Layoutlm: Pre-training of text and layout for document image understanding. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 1192–1200.
  61. 61.Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Yuhao Dan, Chenlin Zhao, Guohai Xu, Chenliang Li, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. 2023a. mplug-docowl: Modularized multimodal large language model for document understanding. CoRR, abs/2307.02499.
  62. 62.Jiabo Ye, Anwen Hu, Haiyang Xu, Qinghao Ye, Ming Yan, Guohai Xu, Chenliang Li, Junfeng Tian, Qi Qian, Ji Zhang, Qin Jin, Liang He, Xin Alex Lin, and Fei Huang. 2023b. Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model. CoRR, abs/2310.05126.
  63. 63.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang. 2023c. mplug-owl: Modularization empowers large language models with multimodality. CoRR, abs/2304.14178.
  64. 64.Yunhu Ye, Binyuan Hui, Min Yang, Binhua Li, Fei Huang, and Yongbin Li. 2023d. Large language models are versatile decomposers: Decomposing evidence and questions for table-based reasoning. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2023, Taipei, Taiwan, July 23-27, 2023, pages 174–184. ACM.
  65. 65.Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, and Yi Ma. 2024. Investigating the catastrophic forgetting in multimodal large language model fine-tuning. In Conference on Parsimony and Learning, pages 202–227. PMLR.
  66. 66.Yanzhe Zhang, Ruiyi Zhang, Jiuxiang Gu, Yufan Zhou, Nedim Lipka, Diyi Yang, and Tong Sun. 2023a. Llava: Enhanced visual instruction tuning for text-rich image understanding. CoRR, abs/2306.17107.
  67. 67.Zhenrong Zhang, Jiefeng Ma, Jun Du, Licheng Wang, and Jianshu Zhang. 2023b. Multimodal pre-training based on graph attention network for document understanding. IEEE Trans. Multim., 25:6743–6755.
  68. 68.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.
  69. 69.Ran Zmigrod, Dongsheng Wang, Mathieu Sibue, Yulong Pei, Petr Babkin, Ivan Brugere, Xiaomo Liu, Nacho Navarro, Antony Papadimitriou, William Watson, Zhiqiang Ma, Armineh Nourbakhsh, and Sameena Shah. 2024. Buddie: A business document dataset for multi-task information extraction. CoRR, abs/2404.04003.

Citation

MLA
Wang, D., et al. “DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 8529–48, https://doi.org/10.18653/v1/2024.acl-long.463.
APA
Wang, D., Raman, N., Sibue, M., Ma, Z., Babkin, P., Kaur, S., Pei, Y., Nourbakhsh, A., & Liu, X. (2024). DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8529–8548. https://doi.org/10.18653/v1/2024.acl-long.463
Chicago
Wang, D., N. Raman, M. Sibue, et al. 2024. “DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8529–48. https://doi.org/10.18653/v1/2024.acl-long.463.
Harvard
Wang, D. et al. (2024) “DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 8529–8548. Available at: https://doi.org/10.18653/v1/2024.acl-long.463.
Vancouver
1. Wang D, Raman N, Sibue M, Ma Z, Babkin P, Kaur S, Pei Y, Nourbakhsh A, Liu X (2024) DocLLM: A Layout-Aware Generative Language Model for Multimodal Document Understanding. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 8529–8548

BibTeX

@inproceedings{wang-etal-2024-docllm,
    title = "{D}oc{LLM}: A Layout-Aware Generative Language Model for Multimodal Document Understanding",
    author = "Wang, Dongsheng  and
      Raman, Natraj  and
      Sibue, Mathieu  and
      Ma, Zhiqiang  and
      Babkin, Petr  and
      Kaur, Simerjot  and
      Pei, Yulong  and
      Nourbakhsh, Armineh  and
      Liu, Xiaomo",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.463/",
    doi = "10.18653/v1/2024.acl-long.463",
    pages = "8529--8548"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/