SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images

Ryota TanakaKyosuke NishidaKosuke NishidaTaku HasegawaItsumi SaitoKuniko Saito

article2023AAAI224 citationsNLP2023 優秀賞, NLP2023 言語資源賞

Introduces SlideVQA, a large-scale multi-image benchmark paired with a unified sequence-to-sequence model that challenges document visual question answering systems to perform multi-hop and numerical reasoning across multi-page slide decks.

Listen

Organizations increasingly rely on automated systems to process and interpret visually rich business materials, such as presentation slides, which integrate text, graphic layout, and data charts. However, existing automated document reading tools primarily focus on single images, leaving artificial intelligence systems unable to effectively search, synthesize, and reason across multi-image documents. Developing AI capable of cross-page information retrieval and numerical calculation is crucial for building reliable virtual assistants and enterprise search agents.

The article introduces SlideVQA, a benchmark dataset designed to evaluate multi-image visual question answering on presentation decks, and proposes M3D, an end-to-end multi-modal artificial intelligence model. The study evaluates how well automated systems can identify relevant evidence slides across multi-page presentations and generate accurate, explainable answers requiring multi-step and arithmetic reasoning.

To establish the benchmark, the researchers collected and curated 2,619 slide decks containing over 52,000 slide images and 14,500 questions covering 39 topics. The dataset features nearly 891,000 bounding-box annotations across nine visual categories and provides arithmetic expressions for questions requiring numerical calculation. The authors designed the M3D model using a sequence-to-sequence framework that jointly selects evidence pages and generates either direct answers or mathematical expressions, using an external calculator to compute final numerical values.

The evaluation revealed several key findings. First, the proposed M3D model outperformed existing state-of-the-art baselines across joint evidence selection and question answering metrics, achieving a Joint Exact Match of 28.0% compared to 24.3% for the best alternative model. Second, predicting intermediate arithmetic expressions directly enhanced performance, yielding an approximate 10% improvement in accuracy on arithmetic questions. Third, incorporating layout and visual data alongside text significantly boosted accuracy across all baseline models, confirming that slide comprehension requires multi-modal understanding. Finally, a substantial performance gap remains between automated systems and people: while human evaluators achieved an 88.6% Joint Exact Match score, the best model reached only 28.0%.

These findings indicate that current automated tools are not yet reliable enough for fully autonomous document analysis in high-stakes operational, compliance, or financial settings. Relying on current models poses risks of inaccurate data extraction and faulty multi-step reasoning. However, the study demonstrates that joint multi-modal training and explicit intermediate reasoning steps offer a viable pathway toward more explainable and dependable document processing.

The article recommends adopting unified, generative architectures rather than disjointed pipelines for multi-document comprehension tasks. For future technical development, the authors suggest exploring two-stage retrieval systems to handle larger document sets efficiently and designing architectures that are robust to variations in optical character recognition quality. Readers should note that multi-hop questions were constructed by synthetically editing single-hop questions, which may not fully match the natural phrasing of human inquiries, and computational demands currently limit the model's scalability to very large slide collections.

arXiv: 2301.04883
Cover for SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images

Abstract

Visual question answering on document images that contain textual, visual, and layout information, called document VQA, has received much attention recently. Although many datasets have been proposed for developing document VQA systems, most of the existing datasets focus on understanding the content relationships within a single image and not across multiple images. In this study, we propose a new multi-image document VQA dataset, SlideVQA, containing 2.6k+ slide decks composed of 52k+ slide images and 14.5k questions about a slide deck. SlideVQA requires complex reasoning, including single-hop, multi-hop, and numerical reasoning, and also provides annotated arithmetic expressions of numerical answers for enhancing the ability of numerical reasoning. Moreover, we developed a new end-to-end document VQA model that treats evidence selection and question answering in a unified sequence-to-sequence format. Experiments on SlideVQA show that our model outperformed existing state-of-the-art QA models, but that it still has a large gap behind human performance. We believe that our dataset will facilitate research on document VQA.

Table of Contents

  • Introduction
  • Related Work
  • The SlideVQA Task and Dataset Task Overview and Formulation
  • Dataset Collection
  • Statistics and Analysis
  • Our Model
  • Multi-modal Task-Specific Input
  • Multi-modal Encoder-Decoder
  • Experiments
  • Experimental Setup
  • Implementation Details
  • Experimental Results and Analysis
  • Discussion and Limitations
  • Conclusion
  • References

Knowls

  1. Knowl 1 — SlideVQA Task Formulation

    definition

    The SlideVQA task is a multi-image document visual question answering benchmark. Given a natural language question qq and an ordered slide deck I={I1,I2,…,IK}\mathbf{I} = \{I_1, I_2, \dots, I_K\} containing K=20K = 20 slide images, a system must select a relevant subset of evidence slide images I^={I^1,…,I^K′}⊆I\hat{\mathbf{I}} = \{\hat{I}_1, \dots, \hat{I}_{K'}\} \subseteq \mathbf{I} and generate the correct answer yy.

    The overall problem is structured into two core subtasks:

    1. Evidence Selection (Subtask 1): Given question qq and slide deck I\mathbf{I}, identify the subset of evidence slide images I^\hat{\mathbf{I}} from which the answer is derived.
    2. Question Answering (Subtask 2): Given question qq and slide deck I\mathbf{I} (or the evidence subset I^\hat{\mathbf{I}}), generate the answer string yy.

    Answers fall into three distinct types:

    • Single-span: A contiguous sequence of tokens extracted directly from a single document image in reading order.
    • Multi-span: An answer composed of multiple disparate spans extracted from across the images.
    • Non-span: An answer not directly extractable as an image text span, encompassing numerical quantities computed via arithmetic operations or descriptions of visual appearances.
  2. Knowl 2 — SlideVQA Dataset Composition and Characteristics

    data/table

    The SlideVQA dataset contains 14,484 question-answer pairs collected from 2,619 English slide decks (truncated to the first 20 slides each, totaling 52,480 slide images) across 39 topic domains from SlideShare. The data is partitioned at the slide-deck level into 10,617 training questions, 1,652 validation questions, and 2,215 test questions, ensuring that all slides and questions for any given deck reside exclusively within the same split.

    Key structural and reasoning characteristics include:

    • Semantic Region Annotations: 890,945 bounding boxes labeled across 9 semantic categories: Title, Page-text, Obj-text (text inside figures, tables, or diagrams), Caption, Other-text (footnotes, dates, URLs), Diagram, Table, Image, and Figure.
    • Reasoning Complexity: 49.3% of the questions require multi-hop reasoning across multiple slides or numerical reasoning. The dataset comprises 12,466 single-hop QA pairs and 2,018 multi-hop QA pairs (synthesized by identifying bridge entities in single-hop questions and substituting them with referring expressions from linked slides).
    • Arithmetic Expression Annotations: 1,732 numerical questions are annotated with explicit ground-truth arithmetic expressions using basic arithmetic operators (++, −-, ∗*, //) (e.g., 30−2830 - 28 for answer 22).
    • Answer Distribution: Single-span answers constitute 67.6%, multi-span answers 16.0%, and non-span answers 16.4%.
    • OCR Context Volume: The average OCR token count across a 20-slide deck is 1,488.88 tokens extracted via Google Cloud Vision API.
  3. Knowl 3 — M3D Multi-Modal Multi-Image Document VQA Architecture

    model/method

    The Multi-Modal Multi-image Document VQA model (M3D) is an end-to-end encoder-decoder system built upon the Fusion-in-Decoder (FiD) framework initialized with pre-trained T5 weights. M3D unifies evidence selection and question answering into a sequence-to-sequence multi-task setup with intermediate arithmetic expression generation.

    1. Multimodal Independent Encoding: For each slide image IkI_k (k∈{1,…,K}k \in \{1, \dots, K\}), an input sequence xkx_k is constructed incorporating task prefix t∈{"Evidence Selection","Question Answering"}t \in \{\text{"Evidence Selection"}, \text{"Question Answering"}\}, question qq, slide page number eke_k, and multimodal context ckc_k: xk=(task:t, question:q, page:ek, context:ck)x_k = (\text{task:} t, \text{ question:} q, \text{ page:} e_k, \text{ context:} c_k) where ck=([Rkr1],Wkr1,[Rkr2],Wkr2,…,[RkrN],WkrN)c_k = ([R_k^{r_1}], \mathbf{W}_k^{r_1}, [R_k^{r_2}], \mathbf{W}_k^{r_2}, \dots, [R_k^{r_N}], \mathbf{W}_k^{r_N}), with [Rkri][R_k^{r_i}] denoting the detected region category token of region rir_i and Wkri\mathbf{W}_k^{r_i} its OCR sub-word tokens. Each of the KK input sequences xkx_k is independently encoded by a multi-modal Transformer encoder into a sequence of hidden vectors xk∈RL×d\mathbf{x}_k \in \mathbb{R}^{L \times d}, and these are concatenated along the sequence length dimension into a joint representation X∈RK×L×d\mathbf{X} \in \mathbb{R}^{K \times L \times d}.

    2. Unified Autoregressive Decoding: A shared Transformer decoder performs cross-attention over X\mathbf{X} to output target sequences depending on the task prefix:

    • For Evidence Selection, the decoder generates page identifiers: I^pages="Evidence pages: e^1,…,e^K′"\hat{I}_{\text{pages}} = \text{"Evidence pages: } \hat{e}_1, \dots, \hat{e}_{K'}\text{"}.
    • For Question Answering, the decoder generates either a direct answer string prefixed by "Answer: "\text{"Answer: "} or an arithmetic expression string prefixed by "Expression: "\text{"Expression: "} (e.g., "Expression: 30 - 28"\text{"Expression: 30 - 28"}). When an arithmetic expression is produced, a deterministic calculator computes the final numerical answer value.
    1. Training Objective: The model is trained end-to-end to minimize the combined loss: L=Ldec+Lsel\mathcal{L} = \mathcal{L}_{\text{dec}} + \mathcal{L}_{\text{sel}} where Ldec\mathcal{L}_{\text{dec}} is the negative log-likelihood of generating the target answer string (or arithmetic expression) given the question answering task prefix, and Lsel\mathcal{L}_{\text{sel}} is the negative log-likelihood of generating the ground-truth evidence page string given the evidence selection task prefix.
  4. Knowl 4 — M3D Multimodal Input Embedding Formulation

    equation

    For a sequence of LL input tokens extracted from a slide image IkI_k, the input representation z∈RL×d\mathbf{z} \in \mathbb{R}^{L \times d} fed into the M3D multi-modal Transformer encoder is computed via layer normalization over the element-wise sum of four distinct embedding vectors:

    z=LN(ztoken+zseg+zlay+zvis)∈RL×d\mathbf{z} = \text{LN}(\mathbf{z}^{\text{token}} + \mathbf{z}^{\text{seg}} + \mathbf{z}^{\text{lay}} + \mathbf{z}^{\text{vis}}) \in \mathbb{R}^{L \times d}

    where:

    • LN(⋅)\text{LN}(\cdot) denotes Layer Normalization.
    • LL is the maximum input sequence length per slide (set to 200 tokens), and dd is the hidden vector dimension.
    • ztoken∈RL×d\mathbf{z}^{\text{token}} \in \mathbb{R}^{L \times d} is the standard sub-word token embedding.
    • zseg∈RL×d\mathbf{z}^{\text{seg}} \in \mathbb{R}^{L \times d} is the segment embedding indicating which detected semantic region contains the token.
    • zlay∈RL×d\mathbf{z}^{\text{lay}} \in \mathbb{R}^{L \times d} is the 2D layout embedding representing the normalized bounding box coordinates (xmin⁡,ymin⁡,xmax⁡,ymax⁡)(x_{\min}, y_{\min}, x_{\max}, y_{\max}) of the token within the slide image, embedded using learned x-axis and y-axis positional tables.
    • zvis∈RL×d\mathbf{z}^{\text{vis}} \in \mathbb{R}^{L \times d} is the visual appearance feature embedding corresponding to the visual region and OCR bounding box extracted by a Faster-RCNN model with a ResNet-101 backbone.
    • For tokens corresponding to the task prefix tt, question qq, and page number eke_k, the layout embedding zlay\mathbf{z}^{\text{lay}} and visual embedding zvis\mathbf{z}^{\text{vis}} are set to zero vectors.
  5. Knowl 5 — Benchmark Performance of M3D versus Baselines on SlideVQA

    empirical result

    Experimental evaluation on the SlideVQA test set compares M3D against language-only models (BERT, T5, PreasM), multimodal document models (LayoutLM, LayoutLMv2, LayoutT5), video QA models (UniVL), and human performance across three evaluation settings: Joint End-to-End Task (Exact Match JEM and F1 JF1), Evidence Selection alone (EM and F1), and Question Answering alone (EM and F1). For pipeline models, the QA module receives the top-3 slides retrieved by a Hierarchical LayoutLMv2 (H-LayoutLMv2) evidence selector (which achieves 96.0% Recall@3).

    Model Modality Main Task Evidence Select Question Answering
    JEM JF1 EM F1 EM F1
    Q-only – – – – – 10.7 13.5
    UniVL V – – – – 10.6 14.1
    BM25 T – – 35.9 47.5 – –
    CLIP (zero-shot) V – – 30.6 34.4 – –
    CLIP (fine-tuned) V – – 39.3 43.5 – –
    BERT T – – 50.3 69.2 – –
    BERT + zlay\mathbf{z}^{\text{lay}} TL – – 52.7 71.0 – –
    LayoutLM TL – – 42.0 59.9 – –
    LayoutLMv2 TLV – – 51.7 71.5 – –
    H-LayoutLMv2 TLV – – 69.8 85.6 – –
    PreasM (pipeline) T 23.4 34.7 – – 30.7 38.2
    T5 (pipeline) T 22.6 34.2 – – 29.3 37.9
    T5 + zlay\mathbf{z}^{\text{lay}} (pipeline) TL 23.6 35.7 – – 31.0 39.7
    LayoutT5 (pipeline) TLV 24.3 36.1 – – 31.7 39.9
    LayoutLMv2 (extractive pipeline) TLV 16.5 26.5 – – 21.4 29.3
    FiD T – – – – 30.4 38.9
    FiD + zlay\mathbf{z}^{\text{lay}} TL – – – – 30.6 38.9
    M3D (Ours) TLV 28.0 37.3 75.0 83.8 33.5 41.7
    M3DGT_{\text{GT}} (oracle evidence) TLV 35.4 44.7 – – – –
    Human Performance – 88.6 91.9 97.7 98.0 89.8 93.0

    Note: Modality denotes Text (T), Layout (L), and Visual (V).

    M3D achieves the highest main task scores (28.0 JEM / 37.3 JF1) and QA scores (33.5 EM / 41.7 F1) among all automated systems. Generative models (LayoutT5) strongly outperform extractive models (LayoutLMv2, 16.5 JEM) due to multi-span and non-span answers. However, a large performance gap remains between M3D and human performance (88.6 JEM / 91.9 JF1).

  6. Knowl 6 — Ablation Analysis of M3D Architecture Components

    empirical result

    An ablation study on the SlideVQA development set demonstrates the contribution of multimodal input features, arithmetic expression (AE) prediction, and joint multi-task learning in M3D:

    M3D Configuration Main Task Evidence Select Question Answering
    JEM JF1 EM F1 EM F1
    Full M3D 36.2 42.8 83.1 87.7 41.3 47.1
    w/o AE prediction 35.7 42.3 82.9 87.7 40.5 46.3
    w/o Evidence selection – – – – 40.6 46.4
    w/o Layout features (zlay\mathbf{z}^{\text{lay}}) 35.1 42.0 82.4 87.1 40.3 46.3
    w/o Visual features (zvis\mathbf{z}^{\text{vis}}) 34.2 40.9 81.5 86.3 39.0 44.9
    w/o Text features (ztoken\mathbf{z}^{\text{token}}) 1.0 1.5 8.4 9.8 9.8 12.0

    Key takeaways:

    1. Textual context: Removing text features leads to complete model failure (JEM drops to 1.0).
    2. Visual and Layout features: Removing visual features causes a 2.0-point drop in JEM (36.2 to 34.2) and a 2.2-point drop in QA F1; removing layout features reduces JEM by 1.1 points.
    3. Arithmetic Expression (AE) Prediction: Disabling AE prediction reduces QA EM from 41.3 to 40.5. On the subset of questions requiring arithmetic operations specifically, AE prediction increases F1 by +10.4% over direct numerical output generation.
    4. Joint Evidence Selection: Removing joint evidence selection during training reduces QA EM from 41.3 to 40.6.
  7. Knowl 7 — Comparison of Evidence Selection Formulations in FiD

    empirical result

    Different evidence selection mechanisms integrated into the Fusion-in-Decoder (FiD) framework were compared on the SlideVQA development set:

    Evidence Selection Method Main Task Evidence Select Question Answering
    JEM JF1 EM F1 EM F1
    M3D backbone (no selection module) – – – – 39.0 44.8
    + BinaryClass (MLP classifier on encoder) 24.7 34.8 54.5 68.5 38.8 44.8
    + ChainGen (joint page-answer sequence) 34.0 40.8 81.1 86.1 39.8 45.4
    + MultiGen (M3D task-prefixed decoding) 35.7 42.3 82.9 87.7 40.5 46.3
    • BinaryClass: A classification baseline using a two-layer MLP with sigmoid activation on top of each slide's start-of-sequence representation. It underperforms generative approaches substantially (54.5 EM vs. 82.9 EM on Evidence Selection), showing that generative modeling benefits from pre-trained language model representations.
    • ChainGen: An autoregressive approach that generates the selected page numbers and the final answer sequentially in a single output stream.
    • MultiGen (M3D's approach): Trains the decoder to generate evidence page lists and answers separately under distinct task prefixes (Evidence Selection vs. Question Answering). MultiGen outperforms ChainGen (35.7 vs. 34.0 JEM; 82.9 vs. 81.1 Selection EM) because optimizing separate decoding trajectories for selection and QA is easier to learn than chaining them into a single long target sequence.
  8. Knowl 8 — Impact of OCR Engine Quality and Semantic Region Detection Accuracy

    empirical result

    The performance of M3D depends significantly on upstream OCR quality and semantic region extraction accuracy:

    1. OCR Engine Sensitivity (SlideVQA Dev Set):
    OCR Engine Main Task Evidence Select Question Answering
    JEM JF1 EM F1 EM F1
    Google Cloud Vision API 36.2 42.8 83.1 87.7 41.3 47.1
    Tesseract OCR 22.5 28.3 69.6 74.7 28.3 34.0

    Using Google Cloud Vision API provides large performance margins over open-source Tesseract (+13.7 JEM, +13.0 QA EM), showing high sensitivity to OCR accuracy.

    1. Faster-RCNN Detection AP by Category: Faster-RCNN (ResNet-101 backbone) trained on SlideVQA achieves the following Average Precision (AP @ IoU=0.5\text{IoU} = 0.5) across semantic object classes on the test set:
    • Title: 87.5% AP
    • Page-text: 76.9% AP
    • Figure: 68.8% AP
    • Table: 65.6% AP
    • Diagram: 64.0% AP
    • Image: 62.2% AP
    • Other-text: 39.4% AP
    • Obj-text: 33.4% AP
    • Caption: 24.9% AP

    Fixed and larger elements (Title, Page-text) achieve significantly higher detection accuracy than small, randomly located textual components (Obj-text, Caption).

  9. Knowl 9 — Limitations of SlideVQA and M3D

    limitation

    The SlideVQA dataset and the M3D architecture have two main limitations:

    1. Synthesized Multi-hop Question Formulations: Multi-hop QA instances were constructed by identifying bridge entities in single-hop questions and replacing them with referring expressions from linked slides. While this editing strategy reliably generates multi-hop constraints at scale, it can yield questions with phrasing that is less natural than spontaneous human queries.
    2. Cross-Attention Computational Scaling: M3D concatenates representations of all K=20K=20 candidate slides into X∈RK×L×d\mathbf{X} \in \mathbb{R}^{K \times L \times d} and computes full cross-attention across all tokens during decoding. This mechanism does not scale computationally to open-domain settings or document collections with hundreds or thousands of pages without an initial coarse document filtering phase.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Ba, L. J.; Kiros, R.; and Hinton, G. E. 2016. Layer Normalization. arXiv:1607.06450.
  2. 2.Bansal, A.; Zhang, Y.; and Chellappa, R. 2020. Visual question answering on image sets. In ECCV, 51–67.
  3. 3.Chen, X.; Zhao, Z.; Chen, L.; Ji, J.; Zhang, D.; Luo, A.; Xiong, Y.; and Yu, K. 2021. WebSRC: A Dataset for Web-Based Structural Reading Comprehension. In EMNLP, 4173–4185.
  4. 4.Cui, L.; Huang, S.; Wei, F.; Tan, C.; Duan, C.; and Zhou, M. 2017. SuperAgent: A Customer Service Chatbot for E-commerce Websites. In ACL, 97–102.
  5. 5.Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT, 4171–4186.
  6. 6.Dua, D.; Wang, Y.; Dasigi, P.; Stanovsky, G.; Singh, S.; and Gardner, M. 2019. DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs. In ACL, 2368–2378.
  7. 7.Fu, T.; Wang, W. Y.; McDuff, D.; and Song, Y. 2022. DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents. In AAAI, 634–642.
  8. 8.Geva, M.; Gupta, A.; and Berant, J. 2020. Injecting Numerical Reasoning Skills into Language Models. In ACL, 946–958.
  9. 9.Haurilet, M.; Roitberg, A.; Martinez, M.; and Stiefelhagen, R. 2019. Wise—slide segmentation in the wild. In ICDAR, 343–348.
  10. 10.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In CVPR, 770–778.
  11. 11.Hong, Y.; Wang, J.; Jia, Y.; Zhang, W.; and Wang, X. 2019. Academic Reader: An Interactive Question Answering System on Academic Literatures. In AAAI, 9855–9856.
  12. 12.Izacard, G.; and Grave, E. 2021. Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In EACL, 874–880.
  13. 13.Kembhavi, A.; Seo, M. J.; Schwenk, D.; Choi, J.; Farhadi, A.; and Hajishirzi, H. 2017. Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension. In CVPR, 5376–5384.
  14. 14.Le, H.; and Hoi, S. C. H. 2020. Video-Grounded Dialogues with Pretrained Generation Language Models. In ACL, 5842–5848.
  15. 15.Lei, J.; Yu, L.; Bansal, M.; and Berg, T. 2018. TVQA: Localized, Compositional Video Question Answering. In EMNLP, 1369–1379.
  16. 16.Lei, J.; Yu, L.; Berg, T.; and Bansal, M. 2020. TVQA+: Spatio-Temporal Grounding for Video Question Answering. In ACL, 8211–8225.
  17. 17.Loshchilov, I.; and Hutter, F. 2017. Decoupled weight decay regularization. arXiv:1711.05101.
  18. 18.Luo, H.; Ji, L.; Shi, B.; Huang, H.; Duan, N.; Li, T.; Li, J.; Bharti, T.; and Zhou, M. 2020. UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation. arXiv:2002.06353.
  19. 19.Mathew, M.; Bagal, V.; Tito, R.; Karatzas, D.; Valveny, E.; and Jawahar, C. 2022. InfographicVQA. In WACV, 1697–1706.
  20. 20.Mathew, M.; Karatzas, D.; and Jawahar, C. V. 2021. DocVQA: A Dataset for VQA on Document Images. In WACV, 2200–2209.
  21. 21.Monica Haurilet, Z. A.-H.; and Stiefelhagen, R. 2019. SPaSe - Multi-Label Page Segmentation for Presentation Slides. In WACV, 726–734.
  22. 22.Powalski, R.; Borchmann, Ł.; Jurkiewicz, D.; Dwojak, T.; Pietruszka, M.; and Pałka, G. 2021. Going full-tilt boogie on document understanding with text-image-layout transformer. In ICDAR, 732–747.
  23. 23.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In ICML, 8748–8763.
  24. 24.Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. JMLR, 21(140): 1–67.
  25. 25.Rajpurkar, P.; Jia, R.; and Liang, P. 2018. Know What You Don't Know: Unanswerable Questions for SQuAD. In ACL, 784–789.
  26. 26.Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. In EMNLP, 2383–2392.
  27. 27.Ren; Shaoqing; He; Kaiming; Girshicka; Ross; Sun; and Jian. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS, 91–99.
  28. 28.Robertson, S.; Zaragoza, H.; et al. 2009. The probabilistic relevance framework: BM25 and beyond. Foundations and Trends® in Information Retrieval, 3(4): 333–389.
  29. 29.Ruder, S. 2016. An overview of gradient descent optimization algorithms. arXiv:1609.04747.
  30. 30.Sachan, D.; Patwary, M.; Shoeybi, M.; Kant, N.; Ping, W.; Hamilton, W. L.; and Catanzaro, B. 2021a. End-to-End Training of Neural Retrievers for Open-Domain Question Answering. In ACL, 6648–6662.
  31. 31.Sachan, D. S.; Reddy, S.; Hamilton, W. L.; Dyer, C.; and Yogatama, D. 2021b. End-to-End Training of Multi-Document Reader and Retriever for Open-Domain Question Answering. In NeurIPS, 25968–25981.
  32. 32.Sun, E.; Hou, Y.; Wang, D.; Zhang, Y.; and Wang, N. X. R. 2021. D2S: Document-to-Slide Generation Via Query-Based Text Summarization. In NAACL-HLT, 1405–1418.
  33. 33.Talmor, A.; and Berant, J. 2018. The Web as a Knowledge-Base for Answering Complex Questions. In NAACL-HLT, 641–651.
  34. 34.Talmor, A.; Yoran, O.; Catav, A.; Lahav, D.; Wang, Y.; Asai, A.; Ilharco, G.; Hajishirzi, H.; and Berant, J. 2021. MultiModalQA: complex question answering over text, tables and images. In ICLR.
  35. 35.Tanaka, R.; Nishida, K.; and Yoshida, S. 2021. VisualMRC: Machine Reading Comprehension on Document Images. In AAAI, 13878–13888.
  36. 36.Tapaswi, M.; Zhu, Y.; Stiefelhagen, R.; Torralba, A.; Urtasun, R.; and Fidler, S. 2016. Movieqa: Understanding stories in movies through question-answering. In CVPR, 4631–4640.
  37. 37.Tito, R.; Karatzas, D.; and Valveny, E. 2021. Document Collection Visual Question Answering. In ICADR, 778–792.
  38. 38.Tu, M.; Huang, K.; Wang, G.; Huang, J.; He, X.; and Zhou, B. 2020. Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents. In AAAI, 9073–9080.
  39. 39.Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017. Attention is All you Need. In NIPS, 6000–6010.
  40. 40.Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Chi, E.; Le, Q.; and Zhou, D. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv:2201.11903.
  41. 41.Xu, Y.; Li, M.; Cui, L.; Huang, S.; Wei, F.; and Zhou, M. 2020. LayoutLM: Pre-training of Text and Layout for Document Image Understanding. In KDD, 1192–1200.
  42. 42.Xu, Y.; Xu, Y.; Lv, T.; Cui, L.; Wei, F.; Wang, G.; Lu, Y.; Florêncio, D. A. F.; Zhang, C.; Che, W.; Zhang, M.; and ˆZhou, L. 2021. LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding. In ACL/IJCNLP, 2579–2591.
  43. 43.Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W. W.; Salakhutdinov, R.; and Manning, C. D. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In EMNLP, 2369–2380.
  44. 44.Yoran, O.; Talmor, A.; and Berant, J. 2022. Turning Tables: Generating Examples from Semi-structured Tables for Endowing Language Models with Reasoning Skills. In ACL, 6016–6031.
  45. 45.Zhang, Q.; Wang, L.; Yu, S.; Wang, S.; Wang, Y.; Jiang, J.; and Lim, E.-P. 2021. NOAHQA: Numerical Reasoning with Interpretable Graph Question Answering Dataset. In Findings of EMNLP, 4147–4161.
  46. 46.Zhang, X.; Ramachandran, D.; Tenney, I.; Elazar, Y.; and Roth, D. 2020. Do Language Embeddings capture Scales? In BlackboxNLP Workshop.

Citation

MLA
Tanaka, R., et al. “SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images”. arXiv, 2023, http://arxiv.org/abs/2301.04883v1.
APA
Tanaka, R., Nishida, K., Nishida, K., Hasegawa, T., Saito, I., & Saito, K. (2023). SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images. arXiv. http://arxiv.org/abs/2301.04883v1
Chicago
Tanaka, R., K. Nishida, K. Nishida, T. Hasegawa, I. Saito, and K. Saito. 2023. “SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images”. arXiv. http://arxiv.org/abs/2301.04883v1.
Harvard
Tanaka, R. et al. (2023) “SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2301.04883v1.
Vancouver
1. Tanaka R, Nishida K, Nishida K, Hasegawa T, Saito I, Saito K (2023) SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images. arXiv

BibTeX

@article{tanaka2023slidevqa,
  title = {SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images},
  author = {Tanaka, Ryota and Nishida, Kyosuke and Nishida, Kosuke and Hasegawa, Taku and Saito, Itsumi and Saito, Kuniko},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2301.04883v1},
  eprint = {2301.04883}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF