SlideVQA: A Dataset for Document Visual Question Answering on Multiple Images
Ryota TanakaKyosuke NishidaKosuke NishidaTaku HasegawaItsumi SaitoKuniko Saito
Introduces SlideVQA, a large-scale multi-image benchmark paired with a unified sequence-to-sequence model that challenges document visual question answering systems to perform multi-hop and numerical reasoning across multi-page slide decks.
Organizations increasingly rely on automated systems to process and interpret visually rich business materials, such as presentation slides, which integrate text, graphic layout, and data charts. However, existing automated document reading tools primarily focus on single images, leaving artificial intelligence systems unable to effectively search, synthesize, and reason across multi-image documents. Developing AI capable of cross-page information retrieval and numerical calculation is crucial for building reliable virtual assistants and enterprise search agents.
The article introduces SlideVQA, a benchmark dataset designed to evaluate multi-image visual question answering on presentation decks, and proposes M3D, an end-to-end multi-modal artificial intelligence model. The study evaluates how well automated systems can identify relevant evidence slides across multi-page presentations and generate accurate, explainable answers requiring multi-step and arithmetic reasoning.
To establish the benchmark, the researchers collected and curated 2,619 slide decks containing over 52,000 slide images and 14,500 questions covering 39 topics. The dataset features nearly 891,000 bounding-box annotations across nine visual categories and provides arithmetic expressions for questions requiring numerical calculation. The authors designed the M3D model using a sequence-to-sequence framework that jointly selects evidence pages and generates either direct answers or mathematical expressions, using an external calculator to compute final numerical values.
The evaluation revealed several key findings. First, the proposed M3D model outperformed existing state-of-the-art baselines across joint evidence selection and question answering metrics, achieving a Joint Exact Match of 28.0% compared to 24.3% for the best alternative model. Second, predicting intermediate arithmetic expressions directly enhanced performance, yielding an approximate 10% improvement in accuracy on arithmetic questions. Third, incorporating layout and visual data alongside text significantly boosted accuracy across all baseline models, confirming that slide comprehension requires multi-modal understanding. Finally, a substantial performance gap remains between automated systems and people: while human evaluators achieved an 88.6% Joint Exact Match score, the best model reached only 28.0%.
These findings indicate that current automated tools are not yet reliable enough for fully autonomous document analysis in high-stakes operational, compliance, or financial settings. Relying on current models poses risks of inaccurate data extraction and faulty multi-step reasoning. However, the study demonstrates that joint multi-modal training and explicit intermediate reasoning steps offer a viable pathway toward more explainable and dependable document processing.
The article recommends adopting unified, generative architectures rather than disjointed pipelines for multi-document comprehension tasks. For future technical development, the authors suggest exploring two-stage retrieval systems to handle larger document sets efficiently and designing architectures that are robust to variations in optical character recognition quality. Readers should note that multi-hop questions were constructed by synthetically editing single-hop questions, which may not fully match the natural phrasing of human inquiries, and computational demands currently limit the model's scalability to very large slide collections.
- Paper: DocVQA: A Dataset for VQA on Document Images, Minesh Mathew et al. (2020). This benchmark established the core task of Document Visual Question Answering on single document images, providing the foundational paradigm that SlideVQA expands to multi-image slide presentations.
- Paper: ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning, Ahmed Masry et al. (2022). ChartQA introduces visual and arithmetic reasoning tasks over complex data figures, which directly informs SlideVQA's numerical and multi-hop reasoning formulations.
- Paper: Towards VQA Models That Can Read, Amanpreet Singh et al. (2019). TextVQA formulated the challenge of integrating OCR tokens and visual text reading into visual question answering, a core requirement for document and slide comprehension.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This seminal work introduced the foundational visual question answering task, establishing the standard problem framing that subsequent document and multi-image VQA datasets build upon.
- Paper: GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering, Drew A. Hudson et al. (2019). GQA pioneered compositional and multi-step reasoning evaluation in visual question answering, influencing how structured multi-hop reasoning is formulated in multimodal settings.
- Paper: HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering, Zhilin Yang et al. (2018). HotpotQA formalizes multi-hop reasoning and explicit supporting-fact selection, providing the multi-hop evidence-selection blueprint adapted by multi-image document VQA models.
- Paper: ConvFinQA: Exploring the Chain of Numerical Reasoning in Conversational Finance Question Answering, Zhiyu Chen et al. (2022). ConvFinQA provides methodology for multi-step numerical and arithmetic reasoning over complex documents containing text and tables, directly relevant to SlideVQA's arithmetic annotations.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). LLaVA-OneVision develops unified multimodal foundation models capable of scaling across single-image, multi-image, and video contexts, directly addressing the multi-image visual reasoning challenges highlighted by SlideVQA.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). InternVL 2.5 advances open-source multimodal scaling specifically for complex multi-image, document, and chart reasoning pipelines.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). Visual CoT expands upon evidence selection in document VQA by implementing localized bounding-box zooming and multi-modal chain-of-thought reasoning over text-dense images.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). M3CoT establishes a dedicated benchmark to evaluate multi-domain, multi-step multimodal chain-of-thought reasoning that extends complex multi-hop visual reasoning.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). PreFLMR tackles scaled multi-modal information retrieval and late-interaction evidence selection across multi-document knowledge bases, generalizing SlideVQA's evidence-retrieval setting.
- Paper: ScreenQA: Large-Scale Question-Answer Pairs Over Mobile App Screenshots, Yu-Chung Hsiao et al. (2025). ScreenQA applies document and interface question answering to complex mobile layouts, extending visual reading comprehension across practical digital screens.
