DocVQA: A Dataset for VQA on Document Images
Minesh MathewDimosthenis KaratzasC. V. Jawahar
Establishes a large-scale visual question answering benchmark of 50,000 questions over 12,000 document images to advance multimodal models that must interpret text alongside complex visual layouts.
Traditional document analysis systems process forms, tables, and text using isolated, task-specific modules that are blind to the end-user's ultimate intent. In business operations, however, stakeholders require systems that dynamically retrieve answers to ad-hoc, natural language questions across diverse and visually complex documents. The article addresses this gap by formalizing the task of Document Visual Question Answering (DocVQA) to drive purpose-driven document comprehension.
The main objective of the article is to establish a large-scale, open-ended benchmark and evaluate how current computer vision and natural language processing models perform when answering questions directly grounded in complex document images.
To achieve this, the authors constructed a benchmark comprising 50,000 question-answer pairs defined across 12,767 document images sourced from the UCSF Industry Documents Library. These documents span five major industries over several decades (predominantly 1960–2000) and feature printed, typewritten, handwritten, and born-digital text arranged in tables, forms, and diagrams. Questions were collected and verified via a rigorous three-stage annotation pipeline, and baselines were established by evaluating heuristic rules, multimodal visual question answering models, language-based reading comprehension models, and human annotators.
The evaluation reveals four primary findings. First, a massive performance gap exists between automated models and human capability: human volunteers achieved 94.36% accuracy, whereas the best-performing model scored only 55.77% accuracy. Second, natural language processing models—specifically a large Bidirectional Encoder Representations from Transformers (BERT) model fine-tuned on reading comprehension and document questions—outperformed visual methods, achieving an Average Normalized Levenshtein Similarity (ANLS) score of 0.665 compared to 0.391 for the best visual model. Third, standard visual question answering techniques that rely on generic object detection features proved ineffective for document images, while expanding the dynamic recognition vocabulary significantly boosted visual model performance. Fourth, automated models degraded sharply on questions requiring structural reasoning, diagram comprehension, or handwritten text interpretation.
These findings imply that deploying current automated models in fully autonomous document-processing workflows carries substantial operational risk, particularly for complex forms, charts, and handwritten records. Because language models process documents as serialized, one-dimensional text streams, they lose critical spatial relationships such as layout hierarchy and table column alignment. This architectural blind spot causes systems to fail on tasks requiring spatial or structural grounding.
For future development and practical implementations, technical teams should prioritize hybrid model architectures that simultaneously encode language representations alongside two-dimensional spatial and layout features. Furthermore, organizations investing in document automation must enhance low-level optical character recognition accuracy, as upstream recognition errors directly propagate into answer extraction failures.
A key limitation of the benchmark is that questions are framed as extractive text spans, meaning the models were not evaluated on tasks requiring numerical calculations, multi-page synthesis, or abstractive reasoning. Nevertheless, the findings offer high confidence that existing architectures are insufficient for full document understanding without purpose-built multimodal integration.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). Introduces the foundational Visual Question Answering (VQA) task formulation and baseline architectures that DocVQA adapts and extends to document-specific domains.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). Introduces TextVQA and OCR-integrated reasoning architectures (LoRRA), establishing the prerequisite paradigm of reading scene text in visual question answering before DocVQA specialized it to complex document images.
- Paper: SQuAD: 100,000+ Questions for Machine Comprehension of Text, Pranav Rajpurkar et al. (2016). Defines the standard extractive reading comprehension task and span-based evaluation methodology upon which DocVQA models and comparative baselines build.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). Establishes the foundational transformer-based multimodal architecture for jointly encoding visual features and text tokens in question answering.
- Paper: Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering, Yash Goyal et al. (2016). Demonstrates critical dataset design principles to balance visual and textual cues in VQA, informing the challenge of evaluating true document understanding over language priors.
- Paper: Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond, Jinze Bai et al. (2023). Develops a versatile vision-language model explicitly benchmarked on DocVQA to demonstrate high-resolution document parsing and text grounding.
- Paper: Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution, Peng Wang et al. (2024). Advances beyond fixed-resolution document processing by introducing dynamic visual resolution to significantly boost performance on dense document understanding tasks like DocVQA.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). Evaluates unified multimodal representations directly against the DocVQA benchmark, demonstrating strong transfer learning and high accuracy on document analysis.
- Paper: ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning, Ahmed Masry et al. (2022). Extends visual question answering on structured documents from text-dense pages to data visualizations and complex chart reasoning.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). Applies large-scale multimodal pre-training and test-time scaling strategies to achieve state-of-the-art results across complex document intelligence and chart reasoning benchmarks.
- Paper: Qianfan-OCR: A Unified End-to-End Model for Document Intelligence, Daxiang Dong et al. (2026). Proposes a unified end-to-end model for full document parsing, layout analysis, and document question answering, overcoming the multi-stage pipeline limitations highlighted in DocVQA.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). Implements native multimodal pre-training with variable visual positional encodings to scale reasoning over long, dense documents.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). Generalizes multimodal evaluation from single-document reading comprehension to massive multidisciplinary college-level visual reasoning.
