Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?
Yang ChenHexiang HuYi LuanHaitian SunSoravit ChangpinyoAlan RitterMing-Wei Chang
Existing multimodal artificial intelligence systems have shown impressive performance on visual question answering tasks, yet they frequently falter when answering visual information-seeking questions. These real-world inquiries demand fine-grained factual knowledge about an image that goes beyond common sense, such as identifying the specific construction year or architect of a photographed building. Prior benchmarks largely evaluated either visual common sense or allowed language models to guess answers from the text alone without examining the image. Consequently, it remained unclear whether state-of-the-art vision and language models could reliably recall and reason over specialized knowledge tied to visual entities.
The article introduces INFOSEEK, a large-scale visual question answering benchmark specifically designed to assess whether models can answer fine-grained, knowledge-intensive questions grounded in visual content. The primary objective is to evaluate both end-to-end vision-language models and modular, external knowledge-retrieval pipelines to determine how effectively they store, recall, and retrieve specialized factual information about visual entities.
To construct this benchmark, the authors developed a dataset comprising two core parts: a human-annotated test set of 8.9 thousand natural information-seeking questions and a semi-automated set of 1.35 million image-question-answer triplets generated from Wikidata across 11 thousand distinct visual entities and 2.7 thousand entity types. The evaluation is split to test generalization on both unseen questions and entirely unseen entities. The authors evaluated state-of-the-art end-to-end models, including PaLI and BLIP2, alongside two-stage pipeline systems that first recognize visual entities using CLIP and then extract answers from external knowledge sources such as Wikipedia using language models and passage readers like Fusion-in-Decoder.
The investigation produced several key findings. First, advanced end-to-end models exhibit weak zero-shot capabilities on visual information-seeking queries, but fine-tuning on targeted data successfully awakens knowledge learned during initial pre-training, enabling PaLI-X to achieve a leading end-to-end score of 22.1% on Wikidata and 10.8% on human questions. Second, retrieval-based pipeline systems with access to external knowledge consistently outperform end-to-end models on human-written queries, achieving an 18.2% accuracy score using Fusion-in-Decoder. Third, the primary performance bottleneck in pipeline systems is visual entity recognition; simulating perfect entity recognition elevated accuracy from around 18% to 45.6%, representing an improvement of roughly 150%. Finally, while pipeline systems dominate on popular entities, end-to-end models demonstrate a distinct advantage on long-tail, less common entities for coarse visual and geographic attributes.
These findings indicate that relying solely on internal model memory or simple instruction tuning is insufficient for fine-grained factual visual queries. For organizations building real-world visual search and assistant technologies, deploying modular architectures that ground images in external knowledge bases provides higher factual precision and better error interpretability. However, because end-to-end models retain advantages on tail entities and broader geographic reasoning, hybrid approaches may be required to balance precision and broad coverage.
Organizations developing visual question answering systems should prioritize investments in high-accuracy visual entity recognition systems, as improvements in entity linking offer the most immediate performance gains. When building such systems, teams should avoid instruction-tuning solely on coarse visual datasets, which was shown to degrade fine-grained factual precision. Development should focus on hybrid frameworks that combine internal parametric reasoning for long-tail visual recognition with retrieval pipelines for fine-grained factual extraction.
The benchmark currently focuses on English-language text and knowledge primarily curated from Wikipedia, which limits direct applicability to specialized domains such as medical imaging or multilingual use cases. Nevertheless, the rigorous multi-stage annotation and verified 95% human accuracy level provide high confidence in the evaluation findings, highlighting visual entity recognition and knowledge retrieval as critical frontiers for multimodal artificial intelligence.
- Paper: OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge, Kenneth Marino et al. (2019). This paper establishes the foundational benchmark for knowledge-based visual question answering requiring external knowledge, providing the key paradigm that INFOSEEK builds upon and evaluates against.
- Paper: KAT: A Knowledge Augmented Transformer for Vision-and-Language, Liangke Gui et al. (2022). This work introduces techniques for augmenting vision-language models with explicit external knowledge bases and implicit knowledge, directly preceding the knowledge-seeking VQA methodologies examined in the source paper.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). This work introduces the BLIP vision-language pre-training framework, which forms the direct foundation of the BLIP-2 architectures analyzed and fine-tuned in the source paper.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This study analyzes how parametric versus non-parametric knowledge retrieval behaves when answering fact-intensive entity queries, motivating the source's retrieval-augmented strategies for information-seeking questions.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This foundational paper defines the visual question answering task that the source extends toward fine-grained, information-seeking domains.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). This work scales up multi-modal knowledge retrieval models across diverse formats, directly extending the document and entity retrieval solutions highlighted for INFOSEEK.
- Paper: Multimodal Reasoning with Multimodal Knowledge Graph, Junlin Lee et al. (2024). This paper advances knowledge-augmented multimodal question answering by incorporating structured multimodal knowledge graphs into language models.
- Paper: Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning, Hao Shao et al. (2024). This study introduces chain-of-thought visual reasoning to extract localized, fine-grained visual details necessary for answering complex information queries.
- Paper: MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI, Kaining Ying et al. (2024). This benchmark broadens the evaluation of large vision-language models across extensive expert-level knowledge domains and diverse task types.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). This paper demonstrates how progressive scaling of vision encoders and language backbones improves multi-step and knowledge-intensive multimodal reasoning.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). This model family scales synthetic knowledge injection and multimodal instruction tuning across single-image, multi-image, and video contexts.
