keyword
multi-image Document VQA
Multi-image document visual question answering is a multimodal artificial intelligence task in which a system answers natural language questions by analyzing, locating, and reasoning across information distributed throughout a collection of multiple document images, such as presentation slide decks or multi-page reports. Unlike standard document question answering approaches that operate on a single page, this task requires models to jointly interpret text, visual elements, and spatial layouts across several interrelated images. Effectively answering questions in this setting typically requires identifying relevant evidence pages from the broader set, performing multi-hop reasoning to connect dispersed details, and executing comparative or arithmetic operations across the document collection.
1 item

