DocVQA: A Dataset for VQA on Document Images

Minesh MathewDimosthenis KaratzasC. V. Jawahar

article2020WACV1,677 citations

Establishes a large-scale visual question answering benchmark of 50,000 questions over 12,000 document images to advance multimodal models that must interpret text alongside complex visual layouts.

Listen

Traditional document analysis systems process forms, tables, and text using isolated, task-specific modules that are blind to the end-user's ultimate intent. In business operations, however, stakeholders require systems that dynamically retrieve answers to ad-hoc, natural language questions across diverse and visually complex documents. The article addresses this gap by formalizing the task of Document Visual Question Answering (DocVQA) to drive purpose-driven document comprehension.

The main objective of the article is to establish a large-scale, open-ended benchmark and evaluate how current computer vision and natural language processing models perform when answering questions directly grounded in complex document images.

To achieve this, the authors constructed a benchmark comprising 50,000 question-answer pairs defined across 12,767 document images sourced from the UCSF Industry Documents Library. These documents span five major industries over several decades (predominantly 1960–2000) and feature printed, typewritten, handwritten, and born-digital text arranged in tables, forms, and diagrams. Questions were collected and verified via a rigorous three-stage annotation pipeline, and baselines were established by evaluating heuristic rules, multimodal visual question answering models, language-based reading comprehension models, and human annotators.

The evaluation reveals four primary findings. First, a massive performance gap exists between automated models and human capability: human volunteers achieved 94.36% accuracy, whereas the best-performing model scored only 55.77% accuracy. Second, natural language processing models—specifically a large Bidirectional Encoder Representations from Transformers (BERT) model fine-tuned on reading comprehension and document questions—outperformed visual methods, achieving an Average Normalized Levenshtein Similarity (ANLS) score of 0.665 compared to 0.391 for the best visual model. Third, standard visual question answering techniques that rely on generic object detection features proved ineffective for document images, while expanding the dynamic recognition vocabulary significantly boosted visual model performance. Fourth, automated models degraded sharply on questions requiring structural reasoning, diagram comprehension, or handwritten text interpretation.

These findings imply that deploying current automated models in fully autonomous document-processing workflows carries substantial operational risk, particularly for complex forms, charts, and handwritten records. Because language models process documents as serialized, one-dimensional text streams, they lose critical spatial relationships such as layout hierarchy and table column alignment. This architectural blind spot causes systems to fail on tasks requiring spatial or structural grounding.

For future development and practical implementations, technical teams should prioritize hybrid model architectures that simultaneously encode language representations alongside two-dimensional spatial and layout features. Furthermore, organizations investing in document automation must enhance low-level optical character recognition accuracy, as upstream recognition errors directly propagate into answer extraction failures.

A key limitation of the benchmark is that questions are framed as extractive text spans, meaning the models were not evaluated on tasks requiring numerical calculations, multi-page synthesis, or abstractive reasoning. Nevertheless, the findings offer high confidence that existing architectures are insufficient for full document understanding without purpose-built multimodal integration.

Cover for DocVQA: A Dataset for VQA on Document Images

Abstract

We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at this http URL

Table of Contents

  • 1 Introduction
  • 2 Related Datasets and Tasks
  • 3 DocVQA
  • 3.1 Data Collection
  • 3.2 Statistics and Analysis
  • 4 Baselines
  • 4.1 Heuristics and Upper Bounds
  • 4.2 VQA Models
  • 4.3 Reading Comprehension Models
  • 5 Experiments
  • 5.1 Evaluation Metrics
  • 5.2 Experimental setup
  • 5.3 Results
  • 6 Conclusion
  • References
  • A Screen grabs of Annotation Tool
  • B Examples of Question Types
  • C Additional Qualitative Examples

Knowls

  1. Knowl 1 — DocVQA Dataset Specification and Partitioning

    definition

    The Document Visual Question Answering (DocVQA) dataset is a benchmark designed for purpose-driven information extraction and visual reasoning on document images. It contains 50,00050{,}000 natural language question-answer pairs defined across 12,76712{,}767 document images sampled from 6,0716{,}071 industry documents in the UCSF Industry Documents Library spanning five industry categories: tobacco, food, drug, fossil fuel, and chemical. The documents span creation dates from 1900 to 2018 (with the majority between 1960 and 2000) and cover diverse layouts including letters, forms, reports, memos, tables, scientific reports, financial reports, and invoices.

    The dataset is formulated as an extractive question answering task where answers appear verbatim in the document image. The dataset is split randomly in an 80 ⁣− ⁣10 ⁣− ⁣1080\!-\!10\!-\!10 ratio:

    • Train split: 39,46339{,}463 questions across 10,19410{,}194 images
    • Validation split: 5,3495{,}349 questions across 1,2861{,}286 images
    • Test split: 5,1885{,}188 questions across 1,2871{,}287 images
  2. Knowl 2 — Three-Stage Annotation Protocol for Document VQA

    experimental setup

    To collect high-quality question-answer pairs on document images, a three-stage crowdsourced annotation pipeline is used:

    1. Question and Answer Generation: Remote annotators view a document image and create up to 10 natural language question-answer pairs whose answers can be extracted verbatim from the image text. Annotators are encouraged to provide multiple valid ground-truth answer strings when applicable.
    2. Independent Verification and Categorization: A second annotator views the document image and the questions created in the first stage (without seeing the first annotator's answers). The second annotator enters answers independently, tags each question with one or more reasoning types from a predefined taxonomy, or flags flawed/ambiguous questions for removal.
    3. Conflict Resolution and Quality Control: If none of the answers from the second stage match any of the answers from the first stage, the question-answer pair is sent to expert reviewers (the dataset authors) to edit the question/answers, provide corrections, or reject the sample.
  3. Knowl 3 — Document Reasoning and Grounding Taxonomy

    definition

    Questions in the DocVQA benchmark are categorized into nine non-mutually exclusive question types based on the structural element and reasoning required to locate and extract the answer:

    • table/list: Answering requires parsing and understanding structured tabular cells, rows, columns, or bulleted/numbered lists.
    • form: Answering requires extracting values associated with field labels or key-value pairs.
    • layout: Answering requires spatial or structural positioning cues (such as identifying page headers, footers, titles, or relative spatial directions).
    • running text: Answering requires reading and comprehending continuous sentences or paragraphs.
    • handwritten: The target answer is written by hand rather than typeset or printed.
    • photograph: The target information is grounded on embedded pictures, photographs, or pictorial stamps.
    • figure: Answering requires interpreting data plots, charts (bar, pie, line), or workflow/hierarchy diagrams.
    • yes/no: Questions that require a binary affirmative or negative verification.
    • other: Document questions that do not fit into the other eight structural categories.
  4. Knowl 4 — Linearization and Span Extraction for Document Reading Comprehension

    model/method

    To apply text-based Machine Reading Comprehension (MRC) architectures (such as BERT) to document images, document images are converted into 1D text sequences:

    1. OCR Extraction and Serialization: An Optical Character Recognition (OCR) system detects and recognizes text tokens. Tokens are sorted in top-left to bottom-right spatial reading order and concatenated into a single serialized context string separated by whitespace.
    2. Distant Supervision Span Identification: Because MRC models predict the start index ss and end index ee of an answer span within context tokens, target answer strings are mapped onto the serialized OCR string by identifying the character span of the first exact substring occurrence of the target answer in the serialized text (following the distant supervision formulation from TriviaQA).
    3. Extractive Prediction: The transformer encoder encodes the concatenated question and serialized document text, and linear classification heads predict the start and end token indices of the answer span.
  5. Knowl 5 — Heuristic Baselines and Performance Upper Bounds for DocVQA

    data/table

    Simple heuristic rules, vocabulary bounds, OCR sequence bounds, and human performance establish performance boundaries on the DocVQA validation and test sets. Evaluation uses Average Normalized Levenshtein Similarity (ANLS) and Exact Match Accuracy (Acc., in %):

    Validation Test
    Baseline ANLS Acc. ANLS Acc.
    Human - - 0.981 94.36
    Random answer 0.003 0.00 0.003 0.00
    Random OCR token 0.013 0.52 0.014 0.58
    Longest OCR token 0.002 0.05 0.003 0.07
    Majority answer 0.017 0.90 0.017 0.89
    Vocab UB - 31.31 - 33.78
    OCR substring UB - 85.64 - 87.00
    OCR subsequence UB - 76.37 - 77.00
    • Vocab UB (Upper Bound): Measures maximum accuracy achievable if the correct answer is restricted to a fixed vocabulary of all 4,3414{,}341 answers occurring more than once in the training set.
    • OCR substring UB: Measures maximum accuracy if the correct answer appears as an exact continuous substring in the space-separated, serialized OCR token output (top-left to bottom-right order).
    • OCR subsequence UB: Measures accuracy when the target answer is a subsequence of the serialized OCR tokens, preventing accidental matches with partial numbers or subwords.
    • Heuristic baselines (random or majority) fail to exceed 1% accuracy, while human performance reaches 0.981 ANLS and 94.36% exact match accuracy.
  6. Knowl 6 — Evaluation of Multimodal Scene Text VQA Models on DocVQA

    data/table

    Performance of Look, Read, Reason & Answer (LoRRA) and Multimodal Multi-Copy Mesh (M4C) on the DocVQA benchmark across different dynamic OCR vocabulary sizes and object feature configurations:

    Validation Test
    Method Objects' feature Fixed vocab. Dynamic vocab. size ANLS Acc. ANLS Acc.
    LoRRA ✓ ✓ 50 0.110 7.22 0.112 7.63
    LoRRA ✓ × 50 0.041 2.64 0.037 2.58
    LoRRA × ✓ 50 0.102 6.73 0.100 6.43
    LoRRA ✓ ✓ 150 0.101 7.09 0.102 7.22
    LoRRA ✓ ✓ 500 0.094 6.41 0.095 6.31
    M4C ✓ ✓ 50 0.292 18.34 0.306 18.75
    M4C ✓ × 50 0.216 12.44 0.219 12.15
    M4C × ✓ 50 0.294 18.75 0.310 18.92
    M4C × ✓ 150 0.352 22.66 0.360 22.35
    M4C × ✓ 300 0.367 23.99 0.375 23.90
    M4C × ✓ 500 0.385 24.73 0.391 24.81
    • LoRRA: Relies predominantly on the fixed vocabulary. Its single-token copy mechanism cannot generate multi-token answers, causing performance degradation when dynamic vocabulary is expanded or fixed vocabulary is removed.
    • M4C: Benefits substantially from scaling the dynamic OCR vocabulary from 50 to 500 tokens (improving test ANLS from 0.306 to 0.391) due to its autoregressive multi-step pointer-generator mechanism.
    • Visual Object Features: Omission of Faster R-CNN object bounding box visual features (pre-trained on Visual Genome) does not harm performance on document images (M4C test ANLS changes from 0.306 to 0.310 with dynamic vocabulary 50), demonstrating that natural scene object features are uninformative for document structure.
  7. Knowl 7 — Performance of Pretrained BERT Reading Comprehension Models on DocVQA

    data/table

    Performance of extractive BERT models trained on serialized OCR text for DocVQA validation and test sets:

    Validation Test
    Pretrained Model DocVQA finetune ANLS Acc. (%) ANLS Acc. (%)
    bert-base ✓ 0.556 45.60 0.574 47.60
    bert-large ✓ 0.594 49.28 0.610 51.08
    bert-large-squad × 0.462 36.72 0.475 38.26
    bert-large-squad ✓ 0.655 54.48 0.665 55.77
    • bert-large-squad denotes bert-large-uncased-whole-word-masking-finetuned-squad.
    • Fine-tuning a model pre-trained on general text and SQuAD 1.1 on DocVQA yields the strongest baseline performance (0.665 ANLS and 55.77% accuracy on the test set), outperforming all multimodal scene-text VQA baselines (best M4C test ANLS is 0.391).
    • Even without DocVQA fine-tuning, zero-shot bert-large-squad achieves 0.475 ANLS, surpassing fine-tuned scene text VQA architectures.
  8. Knowl 8 — Model Performance Disparities across Question Reasoning Categories

    empirical result

    Comparing fine-tuned extractive language models (bert-large-squad), multimodal scene-text VQA models (M4C), and humans across the nine DocVQA question categories reveals significant structural weaknesses in existing architectures:

    • Human Performance: Displays uniform accuracy (~94–98%) across all nine question types.
    • Extractive Language Models (BERT): Excel at categories dominated by linear or semi-linear text such as running text and structured key-value form fields, but suffer steep performance drops on figure, photograph, layout, and handwritten questions due to the loss of 2D spatial layouts in 1D serialization.
    • Multimodal VQA Models (M4C): Capable of recognizing text embedded inside graphics, pack labels, and logos (outperforming BERT on certain photograph and visual logo questions where OCR linear ordering fails), but perform poorly on complex tabular structures and long-context textual reasoning.
    • Universal Model Failures: Both BERT and M4C struggle severely with questions grounded in figures, charts, organizational hierarchies, and diagrams requiring relational, graphical, or geometric reasoning.
  9. Knowl 9 — Text and Token Density Characteristics of Document Images vs. Scene Text

    empirical result

    Comparative analysis between DocVQA, scene text VQA datasets (TextVQA, ST-VQA, VQA 2.0), and natural language reading comprehension benchmarks (SQuAD 1.1) reveals distinct characteristics:

    • OCR Token Density: DocVQA images contain an average of 182.75182.75 OCR text tokens per image, substantially higher than SQuAD 1.1 reading contexts (average 117.23117.23 tokens) and scene text VQA benchmarks (average ≤13\le 13 OCR tokens per image in ST-VQA and TextVQA).
    • Question and Answer Lengths: DocVQA has an average question length of 8.128.12 words (70.72%70.72\% of questions are unique) and an average answer length of 2.172.17 words (63.2%63.2\% of answers are unique), showing significantly higher answer length and open-ended variability than typical scene text VQA datasets.
  10. Knowl 10 — Impact of OCR Recognition and Segmentation Errors on Extractive Document VQA

    limitation

    Extractive document VQA baselines that rely on serialized OCR text or token candidate pointers suffer from two fundamental OCR failure modes:

    1. Over-segmentation and Token Splitting: When an OCR engine splits single semantic entities (such as formatted calendar dates, currency symbols, or hyphenated codes) into disjointed tokens (for example, splitting 1/7/77 into 1, /, 7, /, 77), span-based extractive language models and pointer copy decoders output partial or malformed answers.
    2. Recognition Failures on Handwritten and Degraded Text: Degraded document scans, low contrast, and unconstrained handwritten strokes lead to character misrecognitions (e.g., misreading passenger names), causing exact span matching and vocabulary-based answers to fail even when models correctly locate the visual answer region.

Coverage note — No substantial contributed material was omitted; the knowls cover the dataset specification, 3-stage annotation pipeline, question type taxonomy, serialization method for reading comprehension, heuristic and upper-bound analyses, benchmark evaluations of LoRRA, M4C, and BERT, reasoning performance breakdowns, text density statistics, and OCR-induced failure modes.

References

  1. 1.Aishwarya Agrawal, Aniruddha Kembhavi, Dhruv Batra, and Devi Parikh. C-VQA: A compositional split of the visual question answering (VQA) v1. 0 dataset. arXiv preprint arXiv:1704.08243, 2017.
  2. 2.J. Almazan, A. Gordo, A. Forn'es, and E. Valveny. Word spotting and recognition with embedded attributes. TPAMI, 2014.
  3. 3.Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. Bottom-up and top-down attention for image captioning and visual question answering, 2017.
  4. 4.Ali Furkan Biten, Ruben Tito, Andr'es Mafla, Llu'ıs Gomez, Marc al Rusinol, Minesh Mathew, C. V. Jawahar, Ernest Valveny, and Dimosthenis Karatzas. ICDAR 2019 competition on scene text visual question answering. CoRR, abs/1907.00490, 2019.
  5. 5.Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marcal Rusinol, Ernest Valveny, C.V. Jawahar, and Dimosthenis Karatzas. Scene text visual question answering. In ICCV, 2019.
  6. 6.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5.
  7. 7.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 2017.
  8. 8.J. Deng, W. Dong, R. Socher, L. Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CPVR, 2009.
  9. 9.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In ACL, 2019.
  10. 10.David Doermann, Karl Tombre, et al. Handbook of document image processing and recognition. Springer, 2014.
  11. 11.Difei Gao, Ke Li, Ruiping Wang, Shiguang Shan, and Xilin Chen. Multi-modal graph neural network for joint reasoning on vision and scene text. In CVPR, 2020.
  12. 12.Llu'ıs Gomez, Ali Furkan Biten, Rub'en Tito, Andr'es Mafla, and Dimosthenis Karatzas. Multimodal grid features and cell pointers for scene text visual question answering. arXiv preprint arXiv:2006.00923, 2020.
  13. 13.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2016.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition, 2015.
  15. 15.Sepp Hochreiter and Jurgen Schmidhuber. Long short-term memory. Neural Comput., 1997.
  16. 16.Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach. Iterative answer prediction with pointer-augmented multimodal transformers for textvqa. In CVPR, 2020.
  17. 17.Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR, pages 2901–2910, 2017.
  18. 18.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In ACL, 2017.
  19. 19.Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan. Dvqa: Understanding data visualizations via question answering. In CVPR, 2018.
  20. 20.Kushal Kafle, Robik Shrestha, Scott Cohen, Brian Price, and Christopher Kanan. Answering questions about data visualizations using efficient bimodal fusion. In WACV, 2020.
  21. 21.Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Akos K'ad'ar, Adam Trischler, and Yoshua Bengio. Figureqa: An annotated figure dataset for visual reasoning. arXiv preprint arXiv:1710.07300, 2017.
  22. 22.Isaak Kavasidis, Carmelo Pino, Simone Palazzo, Francesco Rundo, Daniela Giordano, P Messina, and Concetto Spampinato. A saliency-based convolutional neural network for table and chart detection in digitized documents. In ICIAP, 2019.
  23. 23.Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension. In CVPR, 2017.
  24. 24.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015.
  25. 25.Ivan Krasin, Tom Duerig, Neil Alldrin, Vittorio Ferrari, Sami Abu-El-Haija, Alina Kuznetsova, Hassan Rom, Jasper Uijlings, Stefan Popov, Shahab Kamali, Matteo Malloci, Jordi Pont-Tuset, Andreas Veit, Serge Belongie, Victor Gomes, Abhinav Gupta, Chen Sun, Gal Chechik, David Cai, Zheyun Feng, Dhyanesh Narayanan, and Kevin Murphy. Openimages: A public dataset for large-scale multi-label and multi-class image classification. Dataset available from https://storage.googleapis.com/openimages/web/index.html, 2017.
  26. 26.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vision, 2017.
  27. 27.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. Natural questions: a benchmark for question answering research. Transactions of the Association of Computational Linguistics, 2019.
  28. 28.Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. OCR-VQA: Visual question answering by reading text in images. In ICDAR, 2019.
  29. 29.Tri Nguyen et al. Ms marco: A human generated machine reading comprehension dataset. CoRR, abs/1611.09268, 2016.
  30. 30.Rasmus Berg Palm, Ole Winther, and Florian Laws. Cloudscan-a configuration-free invoice analysis system using recurrent neural networks. In ICDAR, 2017.
  31. 31.Jeffrey Pennington, Richard Socher, and Christopher D. Manning. Glove: Global vectors for word representation. In EMNLP, 2014.
  32. 32.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. Squad: 100, 000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250, 2016.
  33. 33.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. In NeurIPS. 2015.
  34. 34.Amanpreet Singh, Vedanuj Goswami, Vivek Natarajan, Yu Jiang, Xinlei Chen, Meet Shah, Marcus Rohrbach, Dhruv Batra, and Devi Parikh. Mmf: A multimodal framework for vision and language research. https://github.com/facebookresearch/mmf, 2020.
  35. 35.Amanpreet Singh, Vivek Natarjan, Meet Shah, Yu Jiang, Xinlei Chen, Devi Parikh, and Marcus Rohrbach. Towards vqa models that can read. In CVPR, 2019.
  36. 36.Ajeet Kumar Singh, Anand Mishra, Shashank Shekhar, and Anirban Chakraborty. From strings to things: Knowledge-enabled vqa model that can read and reason. In ICCV, 2019.
  37. 37.Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman. Newsqa: A machine comprehension dataset. CoRR, abs/1611.09830, 2016.
  38. 38.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS. 2017.
  39. 39.Xinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng, Canjie Luo, Lianwen Jin, Chee Seng Chan, Anton van den Hengel, and Liangwei Wang. On the general value of evidence, and bilingual scene-text visual question answering. In CVPR, 2020.
  40. 40.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. Huggingface’s transformers: State-of-the-art natural language processing. ArXiv, abs/1910.03771, 2019.
  41. 41.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. In NeurIPS. 2019.

Citation

MLA
Mathew, M., et al. “DocVQA: A Dataset for VQA on Document Images”. arXiv, 2020, http://arxiv.org/abs/2007.00398v3.
APA
Mathew, M., Karatzas, D., & Jawahar, C. V. (2020). DocVQA: A Dataset for VQA on Document Images. arXiv. http://arxiv.org/abs/2007.00398v3
Chicago
Mathew, M., D. Karatzas, and C. V. Jawahar. 2020. “DocVQA: A Dataset for VQA on Document Images”. arXiv. http://arxiv.org/abs/2007.00398v3.
Harvard
Mathew, M., Karatzas, D. and Jawahar, C.V. (2020) “DocVQA: A Dataset for VQA on Document Images”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2007.00398v3.
Vancouver
1. Mathew M, Karatzas D, Jawahar CV (2020) DocVQA: A Dataset for VQA on Document Images. arXiv

BibTeX

@article{mathew2020docvqa,
  title = {DocVQA: A Dataset for VQA on Document Images},
  author = {Mathew, Minesh and Karatzas, Dimosthenis and Jawahar, C. V.},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2007.00398v3},
  eprint = {2007.00398}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF