Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Yang ChenHexiang HuYi LuanHaitian SunSoravit ChangpinyoAlan RitterMing-Wei Chang

article2023EMNLP262 citationsBest Social Impact Paper Award
Listen

Existing multimodal artificial intelligence systems have shown impressive performance on visual question answering tasks, yet they frequently falter when answering visual information-seeking questions. These real-world inquiries demand fine-grained factual knowledge about an image that goes beyond common sense, such as identifying the specific construction year or architect of a photographed building. Prior benchmarks largely evaluated either visual common sense or allowed language models to guess answers from the text alone without examining the image. Consequently, it remained unclear whether state-of-the-art vision and language models could reliably recall and reason over specialized knowledge tied to visual entities.

The article introduces INFOSEEK, a large-scale visual question answering benchmark specifically designed to assess whether models can answer fine-grained, knowledge-intensive questions grounded in visual content. The primary objective is to evaluate both end-to-end vision-language models and modular, external knowledge-retrieval pipelines to determine how effectively they store, recall, and retrieve specialized factual information about visual entities.

To construct this benchmark, the authors developed a dataset comprising two core parts: a human-annotated test set of 8.9 thousand natural information-seeking questions and a semi-automated set of 1.35 million image-question-answer triplets generated from Wikidata across 11 thousand distinct visual entities and 2.7 thousand entity types. The evaluation is split to test generalization on both unseen questions and entirely unseen entities. The authors evaluated state-of-the-art end-to-end models, including PaLI and BLIP2, alongside two-stage pipeline systems that first recognize visual entities using CLIP and then extract answers from external knowledge sources such as Wikipedia using language models and passage readers like Fusion-in-Decoder.

The investigation produced several key findings. First, advanced end-to-end models exhibit weak zero-shot capabilities on visual information-seeking queries, but fine-tuning on targeted data successfully awakens knowledge learned during initial pre-training, enabling PaLI-X to achieve a leading end-to-end score of 22.1% on Wikidata and 10.8% on human questions. Second, retrieval-based pipeline systems with access to external knowledge consistently outperform end-to-end models on human-written queries, achieving an 18.2% accuracy score using Fusion-in-Decoder. Third, the primary performance bottleneck in pipeline systems is visual entity recognition; simulating perfect entity recognition elevated accuracy from around 18% to 45.6%, representing an improvement of roughly 150%. Finally, while pipeline systems dominate on popular entities, end-to-end models demonstrate a distinct advantage on long-tail, less common entities for coarse visual and geographic attributes.

These findings indicate that relying solely on internal model memory or simple instruction tuning is insufficient for fine-grained factual visual queries. For organizations building real-world visual search and assistant technologies, deploying modular architectures that ground images in external knowledge bases provides higher factual precision and better error interpretability. However, because end-to-end models retain advantages on tail entities and broader geographic reasoning, hybrid approaches may be required to balance precision and broad coverage.

Organizations developing visual question answering systems should prioritize investments in high-accuracy visual entity recognition systems, as improvements in entity linking offer the most immediate performance gains. When building such systems, teams should avoid instruction-tuning solely on coarse visual datasets, which was shown to degrade fine-grained factual precision. Development should focus on hybrid frameworks that combine internal parametric reasoning for long-tail visual recognition with retrieval pipelines for fine-grained factual extraction.

The benchmark currently focuses on English-language text and knowledge primarily curated from Wikipedia, which limits direct applicability to specialized domains such as medical imaging or multilingual use cases. Nevertheless, the rigorous multi-stage annotation and verified 95% human accuracy level provide high confidence in the evaluation findings, highlighting visual entity recognition and knowledge retrieval as critical frontiers for multimodal artificial intelligence.

arXiv: 2302.11713
Cover for Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?

Abstract

Pre-trained vision and language models (Chen et al., 2023b,a; Dai et al., 2023; Li et al., 2023b) have demonstrated state-of-the-art capabilities over existing tasks involving images and texts, including visual question answering. However, it remains unclear whether these models possess the capability to answer questions that are not only querying visual content but knowledge-intensive and information-seeking. In this study, we introduce INFOSEEK^1, a visual question answering dataset tailored for information-seeking questions that cannot be answered with only common sense knowledge. Using INFOSEEK, we analyze various pre-trained visual question answering models and gain insights into their characteristics. Our findings reveal that state-of-the-art pre-trained multi-modal models (e.g., PaLI-X, BLIP2, etc.) face challenges in answering visual information-seeking questions, but fine-tuning on the INFOSEEK dataset elicits models to use fine-grained knowledge that was learned during their pre-training. Furthermore, we show that accurate visual entity recognition can be used to improve performance on INFOSEEK by retrieving relevant documents, showing a significant space for improvement.

Table of Contents

  • 1 Introduction
  • 2 The Need for a New Visual Information-seeking Benchmark
  • 3 INFOSEEK: A VQA Benchmark of Visual Information-seeking Questions
  • 3.1 INFOSEEKHuman: Natural Info-Seeking VQA Data Annotated by Humans
  • 3.2 INFOSEEKWikidata : 1 Million Automated VQA Data from Wikipedia
  • 3.3 Evaluation of INFOSEEK
  • 4 Protocols and Models for INFOSEEK
  • 4.1 Models without KB Information
  • 4.2 Models with KB Information
  • 5 Experiments
  • 5.1 Results for No-KB Models
  • 5.2 Results for With-KB Models
  • 6 Related Work
  • 7 Conclusion
  • 8 Limitation
  • Acknowledgement
  • References
  • A Details of the Dataset.
  • A.1 Human Annotation Quality Control
  • A.2 Filtering and Subsampling
  • A.3 Evaluation Metric.
  • A.4 Image Sources.
  • B Implementation details of the baseline systems
  • B.1 without-KB Models
  • B.2 With-KB Models
  • C Additional Experiment Results
  • C.1 OK-VQA Annotation Guidelines

Knowls

  1. Knowl 1 — INFOSEEK Visual Question Answering Benchmark

    definition

    INFOSEEK is a multimodal visual question answering (VQA) benchmark designed to evaluate models on knowledge-intensive, visual information-seeking questions where answers cannot be deduced from visual common sense or visual attributes alone, but require fine-grained world knowledge tied to specific visual entities. INFOSEEK consists of two sub-datasets:

    1. INFOSEEK_Human: A collection of 8,931 natural, human-authored {image, question, answer} triplets covering 806 visual entities and 527 fine-grained entity types. Questions were written by human annotators without access to answers, reflecting genuine information-seeking curiosity, and answers were extracted as concise spans from corresponding Wikipedia articles. The question distribution comprises 26.8% TIME, 26.4% NUMERICAL, and 46.8% STRING questions.
    2. INFOSEEK_Wikidata: A large-scale semi-automated dataset comprising 1,355,648 triplets across 11,481 visual entities and 2,739 entity types. It was generated by translating structured Wikidata knowledge triples across 300 relations into natural language queries via human-written templates, followed by entity linking to Wikipedia, text-grounding filtering, and relation subsampling. Its test split question distribution consists of 4.8% TIME, 22.9% NUMERICAL, and 72.3% STRING questions.

    Images are sourced across nine visual recognition and instance retrieval domains, including landmarks (17.4%), landforms (15.0%), animals (13.1%), items (10.3%), plants (7.5%), birds (7.1%), food (5.3%), vehicles (4.6%), and aircraft (2.8%).

  2. Knowl 2 — INFOSEEK Evaluation Protocols (No-KB and With-KB)

    experimental setup

    INFOSEEK defines two distinct evaluation protocols to evaluate how models access and utilize world knowledge:

    1. No-KB Protocol (Closed-Book / End-to-End): Evaluates the parametric memory of multimodal models. During training and validation, models receive standard {Image,Question,Answer}\{\text{Image}, \text{Question}, \text{Answer}\} triplets. At test time, models must predict the answer directly given only {Image,Question}\{\text{Image}, \text{Question}\} without external retrieval.
    2. With-KB Protocol (Open-Book / Decoupled Pipeline): Evaluates multimodal retrieval-augmented systems with access to an external knowledge base consisting of 100,000 Wikipedia articles and associated infobox images. The task is decomposed into two sub-tasks:
      • Visual Entity Recognition: Grounding the visual query {Image,Question}\{\text{Image}, \text{Question}\} to a candidate entity in the 100,000 Wikipedia knowledge base using a multimodal retriever (such as CLIP fine-tuned on entity image-text representations).
      • Textual Question Answering: Using the identified entity name and/or reading the corresponding Wikipedia text passages (via large language models such as PaLM with in-context prompting, or retrieval-augmented generative readers such as Fusion-in-Decoder) to answer the query.

    At test time under both protocols, models receive only the input image and question {Image,Question}\{\text{Image}, \text{Question}\}.

  3. Knowl 3 — Benchmark Performance of Multimodal Models on INFOSEEK

    data/table

    The following table summarizes the performance (in VQA accuracy score %) of various models evaluated on the INFOSEEK_Wikidata and INFOSEEK_Human test splits, along with comparator benchmarks (OK-VQA and VQAv2).

    Model LLM Backbone # Params INFOSEEK Wikidata INFOSEEK Human
    Unseen Q Unseen E Overall Unseen Q Unseen E Overall
    Baselines
    Random - - 0.1 0.1 0.1 0.2 0.1 0.1
    Question Prior (4-gram) - - 3.9 2.7 3.2 0.3 0.3 0.3
    PaLM (Q-only, 5-shot) PaLM 540B 5.1 3.7 4.3 4.8 6.6 5.6
    No-KB (End-to-End)
    BLIP2 Flan-T5-XXL 12B 14.5 13.3 13.9 10.0 8.9 9.4
    InstructBLIP Flan-T5-XXL 12B 14.3 13.2 13.7 10.6 9.3 9.9
    PaLI-17B mT5-XXL 17B 20.7 16.0 18.1 13.3 5.9 8.2
    PaLI-X UL2-32B 55B 23.5 20.8 22.1 12.9 9.3 10.8
    With-KB (Pipeline)
    CLIP →\rightarrow PaLM (5-shot) PaLM 540B 21.9 18.6 20.1 15.6 14.9 15.2
    CLIP →\rightarrow FiD T5-Large 1B 20.7 18.1 19.3 18.9 17.6 18.2
    Oracle Entity →\rightarrow FiD T5-Large 1B 52.1 51.9 52.0 44.8 46.4 45.6

    The data demonstrates that:

    1. Question-only language models achieve minimal accuracy on INFOSEEK (5.6% on Human test set), confirming that image understanding is required.
    2. In the No-KB setting, fine-tuned PaLI-X achieves the highest overall performance (22.1% on Wikidata, 10.8% on Human), but a persistent generalization gap remains between seen and unseen visual entities.
    3. In the With-KB setting, reading retrieved Wikipedia passages (CLIP →\rightarrow FiD) outperforms end-to-end parametric models on natural human-written questions (18.2% vs. 10.8% on INFOSEEK_Human).
    4. An oracle entity linking upper bound (Oracle →\rightarrow FiD) reaches 52.0% on Wikidata and 45.6% on Human, demonstrating that entity recognition is the primary performance bottleneck for knowledge-grounded VQA.
  4. Knowl 4 — INFOSEEK Evaluation Metrics and Split Formulation

    equation

    INFOSEEK evaluates model generalization across two distinct held-out evaluation splits:

    • UNSEEN QUESTION: QA pairs regarding visual entities present in the training set, but querying relations not paired with those entities during training.
    • UNSEEN ENTITY: QA pairs for visual entities held out completely from the training set.

    The overall benchmark score is computed as the harmonic mean of the accuracies achieved on these two splits: Overall Accuracy=2⋅Accunseen_q⋅Accunseen_eAccunseen_q+Accunseen_e\text{Overall Accuracy} = \frac{2 \cdot \text{Acc}_{\text{unseen\_q}} \cdot \text{Acc}_{\text{unseen\_e}}}{\text{Acc}_{\text{unseen\_q}} + \text{Acc}_{\text{unseen\_e}}} where Accunseen_q\text{Acc}_{\text{unseen\_q}} is the accuracy on the UNSEEN QUESTION split, and Accunseen_e\text{Acc}_{\text{unseen\_e}} is the accuracy on the UNSEEN ENTITY split.

    Question-level accuracy is computed according to question type:

    • STRING Questions: Standard VQA accuracy evaluated against multiple reference aliases from Wikidata (average 4.5 references) or multiple human annotations (average 2.4 references).
    • TIME Questions: Exact match against reference date formats, incorporating a relaxed ±1\pm 1 year error tolerance to accommodate historical date variations.
    • NUMERICAL Questions: Relaxed Accuracy evaluated against a reference range [refmin⁡,refmax⁡][\text{ref}_{\min}, \text{ref}_{\max}] (or a target value with a ±10%\pm 10\% tolerance window). A single predicted value pred\text{pred} is correct if: refmin⁡≤pred≤refmax⁡\text{ref}_{\min} \le \text{pred} \le \text{ref}_{\max} A predicted interval [predmin⁡,predmax⁡][\text{pred}_{\min}, \text{pred}_{\max}] is correct if the Intersection over Union (IoU\text{IoU}) with the reference interval satisfies: IoU([predmin⁡,predmax⁡],[refmin⁡,refmax⁡])≥50%\text{IoU}([\text{pred}_{\min}, \text{pred}_{\max}], [\text{ref}_{\min}, \text{ref}_{\max}]) \ge 50\%
  5. Knowl 5 — Elicitation of Latent Pre-Trained Knowledge via INFOSEEK Fine-Tuning

    empirical result

    End-to-end vision-language models pre-trained on massive web image-text datasets without instruction tuning (such as PaLI-17B and PaLI-X) exhibit near-zero accuracy on INFOSEEK in zero-shot evaluations. However, fine-tuning on the INFOSEEK_Wikidata training corpus elicits this latent factual knowledge:

    • PaLI-17B overall accuracy increases from near 0% zero-shot to 18.1% on INFOSEEK_Wikidata and 8.2% on INFOSEEK_Human.
    • PaLI-X overall accuracy increases from near 0% zero-shot to 22.1% on INFOSEEK_Wikidata and 10.8% on INFOSEEK_Human.
    • Fine-tuned PaLI models generalize to out-of-distribution real-world images and unseen entity domains (such as fashion items and artwork captured by smartphone cameras), where PaLI-17B and PaLI-X achieve 22.2% and 38.9% accuracy respectively on test queries.
    • For unseen entities (such as Amberd fortress), BLIP2 accuracy on country-location queries improves from 18% (2/11) in zero-shot to 92% (10/11) after fine-tuning on INFOSEEK, despite the entity never appearing in the fine-tuning set.

    This demonstrates that INFOSEEK fine-tuning serves to awaken pre-existing parametric knowledge rather than merely inducing memorization of the fine-tuning data.

  6. Knowl 6 — Answer Coarsening Degradation in Multi-Task Instruction-Tuned Vision-Language Models

    empirical result

    When evaluated zero-shot on fine-grained visual info-seeking questions (INFOSEEK_Wikidata), InstructBLIP performs substantially worse than its base model BLIP2 (achieving 7.4% versus 11.3% zero-shot accuracy).

    This degradation occurs because standard visual question answering datasets used during instruction tuning (such as VQAv2 and OK-VQA) favor coarse-grained answers and general category terms over specific named entities. For instance, given a query asking for a bridge designer ("Who designed this bridge?"), zero-shot BLIP2 outputs a specific name ("john nash", target: "Thomas Telford"), whereas zero-shot InstructBLIP predicts the generic category ("architect"). Similarly, for a geographical question ("Which body of water is this mountain located in or next to?"), BLIP2 predicts "lake como" (target: "Lake Como") while InstructBLIP predicts "lake". Fine-tuning on INFOSEEK data eliminates this coarsening bias and restores fine-grained answer generation.

  7. Knowl 7 — Head vs. Tail Entity Performance Dynamics in End-to-End and Pipeline VQA Models

    empirical result

    A comparative analysis across Wikipedia entity popularity (measured by average monthly Wikipedia pageviews) reveals diverging strengths between end-to-end (No-KB) models and retrieval-pipeline (With-KB) models:

    1. Tail Entities (Low Pageviews ≤1K\le 1\text{K}): End-to-end models like PaLI-17B outperform two-stage retrieval pipeline systems (such as CLIP →\rightarrow PaLM and CLIP →\rightarrow FiD) by a substantial margin. End-to-end models can exploit peripheral visual clues (e.g., architectural style, landscape, or typography) to answer coarse-grained questions (such as country or continent) without requiring explicit named entity recognition, thereby bypassing entity-linking failures on rare entities.
    2. Head Entities (High Pageviews ≥50K\ge 50\text{K}): Pipeline systems with knowledge base retrieval substantially outperform end-to-end models. On popular entities, visual entity linking accuracy is high, enabling text retrieval and reader modules to extract exact fine-grained facts that parametric memory fails to retain accurately.
  8. Knowl 8 — Model Performance Discrepancy Across Question Types in Info-Seeking VQA

    empirical result

    Evaluation across distinct question types on the INFOSEEK_Wikidata validation set (unseen question split) reveals systematic performance gaps across model paradigms:

    Model TIME (Acc. %) NUMERICAL (Relaxed Acc. %) STRING (Acc. %)
    No-KB Setting
    Prior 0.0 4.4 5.0
    PaLM (Q-only) 0.0 11.4 4.0
    InstructBLIP 7.9 7.5 17.8
    BLIP2 6.9 5.8 18.5
    PaLI-17B 3.8 18.4 27.4
    PaLI-X 7.7 16.1 30.0
    With-KB Setting
    CLIP →\rightarrow PaLM 12.5 27.7 21.7
    CLIP →\rightarrow FiD 12.3 23.4 23.9

    Key observations:

    1. End-to-end vision-language models (e.g., PaLI-X) achieve reasonable performance on STRING questions (30.0%), which involve entity relations or categorizations, but struggle severely on TIME questions (7.7%) and NUMERICAL measurements (16.1%).
    2. With-KB models (such as CLIP →\rightarrow PaLM and CLIP →\rightarrow FiD) significantly outperform end-to-end models on TIME (12.3--12.5% vs. 3.8--7.9%) and NUMERICAL questions (23.4--27.7% vs. 5.8--18.4%).
    3. On human-written info-seeking questions (INFOSEEK_Human), reading retrieved Wikipedia text (FiD) yields superior generalization over prompted LLMs (PaLM) for fine-grained factual values: TIME (21.5% for FiD vs. 14.6% for PaLM) and NUMERICAL (25.6% for FiD vs. 21.3% for PaLM).
  9. Knowl 9 — Two-Stage Natural Question-Answer Annotation Protocol for Visual Info-Seeking

    algorithm

    To collect natural visual info-seeking questions without visual attribute shortcuts or answer-leakage bias, INFOSEEK employs an information-seeking annotation workflow inspired by TyDiQA:

    Input: Visual entity EE, Entity images IEI_E, 15-word entity summary SES_E, Wikipedia section titles TET_E, Wikipedia article text WEW_E
    Output: Validated {I, Q, A} triples
    Stage 1: Question Writing (Annotator Group 1)
    Prompt Annotator 1 with (IEI_E, SES_E, TET_E) without showing full article or answers
    Annotator 1 writes 3 to 5 questions QQ querying sections TET_E driven by curiosity
    Reject QQ if: binary, asks for visual attributes (color/shape), copies entity name or section title, or is ambiguous
    Stage 2: Answer Labeling (Annotator Group 2)
    Present Annotator 2 with QQ and Wikipedia article WEW_E
    if answer to QQ cannot be derived from WEW_E:
        Discard QQ
    else:
        Extract shortest continuous answer text span AA from WEW_E (∣A∣≤10|A| \le 10 words)
        Assign question category C∈{TIME,NUMERICAL,STRING}C \in \{\text{TIME}, \text{NUMERICAL}, \text{STRING}\}
        Extract supporting evidence sentence from WEW_E
    Stage 3: Expert Verification and Image Assignment
    Pair (QQ, AA) with image I∈IEI \in I_E
    Expert annotators verify correctness of AA against WEW_E, correct type CC, format numerical units, and resolve multi-object image ambiguity
    return {I, Q, A}

    Expert evaluation on a sample of 200 examples confirms a 95% annotation accuracy rate under this procedure.

  10. Knowl 10 — Subsampling Formulation for Mitigating Entity and Relation Skew in INFOSEEK

    algorithm

    To avoid question prior shortcuts caused by skewed distributions of crowd-sourced Wikidata triples, INFOSEEK applies a probabilistic subsampling algorithm over relation-category pairs.

    Input: Triple set T={(s,r,o)}T = \{(s, r, o)\}, Entity categories {cs}\{c_s\} for subjects ss, Relations {r}\{r\}
    Output: Filtered triple subset TsubT_{sub}
    Compute P(r,c)←∣{(s,r,o)∈T∣category(s)=c}∣∣T∣P(r, c) \leftarrow \frac{|\{(s, r, o) \in T \mid \text{category}(s) = c\}|}{|T|}
    for each relation-category pair (r,c)(r, c) present in TT:
        P′(r,c)←1∣{(r′,c′) present in T}∣P'(r, c) \leftarrow \frac{1}{|\{(r', c') \text{ present in } T\}|}
    Initialize Tsub←∅T_{sub} \leftarrow \emptyset
    for each triple (s,r,o)∈T(s, r, o) \in T with category(s)=c(s) = c:
        rdrop←1−min⁡(1,P′(r,c)P(r,c))1/2r_{drop} \leftarrow 1 - \min\left(1, \frac{P'(r, c)}{P(r, c)}\right)^{1/2}
        Draw random variable u∼Uniform(0,1)u \sim \text{Uniform}(0, 1)
        if u≥rdropu \ge r_{drop}:
            Add (s,r,o)(s, r, o) to TsubT_{sub}
    return TsubT_{sub}

    The subsampling reduces model exploitation of question frequency priors, depressing a 4-gram question-prior baseline's VQA score on INFOSEEK_Wikidata to 3.2%.

Coverage note — Specific fine-tuning hyperparameter grids (e.g., exact learning rates and batch sizes per baseline model in Table 9) and minor baseline runs (OFA, mPLUG-owl) were omitted in favor of the core benchmark methodology, algorithmic data synthesis procedures, and primary empirical findings.

References

  1. 1.Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018. Don't just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4971–4980.
  2. 2.Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning. Conference on Neural Information Processing Systems.
  3. 3.Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2021. Beit: Bert pre-training of image transformers. International Conference on Learning Representations.
  4. 4.Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. 2014. Food-101–mining discriminative components with random forests. In European Conference on Computer Vision. Springer.
  5. 5.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. Conference on Neural Information Processing Systems.
  6. 6.Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, et al. 2023a. Pali-x: On scaling up a multilingual vision and language model. arXiv preprint arXiv:2305.18565.
  7. 7.Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, et al. 2023b. PaLi: A jointly-scaled multilingual language-image model. International Conference on Learning Representations.
  8. 8.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. Palm: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  10. 10.Jonathan H Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. 2020. TyDi QA: A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Linguistics.
  11. 11.Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. 2023. Instructblip: Towards general-purpose vision-language models with instruction tuning. arXiv preprint arXiv:2305.06500.
  12. 12.Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos, Ibrahim Alabdulmohsin, et al. 2023. Scaling vision transformers to 22 billion parameters. arXiv preprint arXiv:2302.05442.
  13. 13.Gerry. 2021. 100 sports image classification.
  14. 14.Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In Computer Vision and Pattern Recognition.
  15. 15.Liangke Gui, Borui Wang, Qiuyuan Huang, Alex Hauptmann, Yonatan Bisk, and Jianfeng Gao. 2022. KAT: A knowledge augmented transformer for vision-and-language. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  16. 16.Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham. 2018. Vizwiz Grand Challenge: Answering visual questions from blind people. In Computer Vision and Pattern Recognition.
  17. 17.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.
  18. 18.Hexiang Hu, Yi Luan, Yang Chen, Urvashi Khandelwal, Mandar Joshi, Kenton Lee, Kristina Toutanova, and Ming-Wei Chang. 2023. Open-domain visual entity recognition: Towards recognizing millions of wikipedia entities. arXiv preprint arXiv:2302.11154.
  19. 19.Yushi Hu, Hang Hua, Zhengyuan Yang, Weijia Shi, Noah A Smith, and Jiebo Luo. 2022a. Promptcap: Prompt-guided task-aware image captioning. arXiv preprint arXiv:2211.09699.
  20. 20.Ziniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang, Kai-Wei Chang, Yizhou Sun, Cordelia Schmid, David A Ross, and Alireza Fathi. 2022b. Reveal: Retrieval-augmented visual-language pre-training with multi-source multimodal knowledge memory. arXiv preprint arXiv:2212.05221.
  21. 21.Robert Iv, Alexandre Passos, Sameer Singh, and Ming-Wei Chang. 2022. FRUIT: Faithfully reflecting updated information in text. In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Seattle, United States. Association for Computational Linguistics.
  22. 22.Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. Proceedings of the European Chapter of the Association for Computational Linguistics.
  23. 23.Aman Jain, Mayank Kothyari, Vishwajeet Kumar, Preethi Jyothi, Ganesh Ramakrishnan, and Soumen Chakrabarti. 2021. Select, substitute, search: A new benchmark for knowledge-augmented visual question answering. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2491–2498.
  24. 24.Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Annual Meeting of the Association for Computational Linguistics.
  25. 25.Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2022. Large language models struggle to learn long-tail knowledge. arXiv preprint arXiv:2211.08411.
  26. 26.Jing Yu Koh, Ruslan Salakhutdinov, and Daniel Fried. 2023. Grounding language models to images for multimodal generation. International Conference on Machine Learning.
  27. 27.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 2013. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops.
  28. 28.Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics.
  29. 29.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. In Annual Meeting of the Association for Computational Linguistics.
  30. 30.Paul Lerner, Olivier Ferret, Camille Guinaudeau, Hervé Le Borgne, Romaric Besançon, Jose G. Moreno, and Jesús Lovón Melgarejo. 2022. Viquae, a dataset for knowledge-based visual question answering about named entities. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '22.
  31. 31.Dongxu Li, Junnan Li, Hung Le, Guangsen Wang, Silvio Savarese, and Steven C.H. Hoi. 2023a. LAVIS: A one-stop library for language-vision intelligence. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). Association for Computational Linguistics.
  32. 32.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023b. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. International Conference on Machine Learning.
  33. 33.Yuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu, Chenguang Zhu, and Lu Yuan. 2022. REVIVE: Regional visual representation matters in knowledge-based visual question answering. Conference on Neural Information Processing Systems.
  34. 34.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. arXiv preprint arXiv:2304.08485.
  35. 35.Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. 2022. Unified-io: A unified model for vision, language, and multimodal tasks. arXiv preprint arXiv:2206.08916.
  36. 36.Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew Blaschko, and Andrea Vedaldi. 2013. Fine-grained visual classification of aircraft. Technical report.
  37. 37.Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi. 2022. When not to trust language models: Investigating effectiveness and limitations of parametric and non-parametric memories. arXiv preprint arXiv:2212.10511.
  38. 38.Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta, and Marcus Rohrbach. 2021. Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa. In Computer Vision and Pattern Recognition.
  39. 39.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. OK-VQA: A visual question answering benchmark requiring external knowledge. In Computer Vision and Pattern Recognition.
  40. 40.Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. 2022. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. Annual Meeting of the Association for Computational Linguistics (Findings).
  41. 41.Thomas Mensink, Jasper Uijlings, Lluis Castrejon, Arushi Goel, Felipe Cadar, Howard Zhou, Fei Sha, André Araujo, and Vittorio Ferrari. 2023. Encyclopedic vqa: Visual questions about detailed properties of fine-grained categories. arXiv preprint arXiv:2306.09224.
  42. 42.Nitesh Methani, Pritha Ganguly, Mitesh M Khapra, and Pratyush Kumar. 2020. PlotQA: Reasoning over scientific plots. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1527–1536.
  43. 43.Maria-Elena Nilsback and Andrew Zisserman. 2008. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing.
  44. 44.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning.
  45. 45.Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, W. Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR.
  46. 46.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. Proceedings of Emperical Methods in Natural Language Processing.
  47. 47.Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. 2021. Imagenet-21k pretraining for the masses. arXiv preprint arXiv:2104.10972.
  48. 48.Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? CoRR, abs/2002.08910.
  49. 49.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. 2015. Imagenet large scale visual recognition challenge. International journal of computer vision, 115:211–252.
  50. 50.Naganand Yadati Sanket Shah, Anand Mishra and Partha Pratim Talukdar. 2019. KVQA: Knowledge-aware visual question answering.
  51. 51.Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-okvqa: A benchmark for visual question answering using world knowledge. arXiv preprint arXiv:2206.01718.
  52. 52.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Neil Houlsby, and Donald Metzler. 2022. Unifying language learning paradigms. arXiv preprint arXiv:2205.05131.
  53. 53.Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge Belongie. 2018. The inaturalist species classification and detection dataset. In Computer Vision and Pattern Recognition.
  54. 54.Denny Vrandečić and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM.
  55. 55.Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. 2017. FVQA: Fact-based visual question answering. IEEE transactions on pattern analysis and machine intelligence.
  56. 56.Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony Dick. 2015. Explicit knowledge-based reasoning for visual question answering. arXiv preprint arXiv:1511.02570.
  57. 57.Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang. 2022. Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework. International Conference on Machine Learning.
  58. 58.Tobias Weyand, Andre Araujo, Bingyi Cao, and Jack Sim. 2020. Google landmarks dataset v2-a large-scale benchmark for instance-level recognition and retrieval. In Computer Vision and Pattern Recognition.
  59. 59.Jialin Wu and Raymond J Mooney. 2022. Entity-focused dense passage retrieval for outside-knowledge visual question answering. Proceedings of Emperical Methods in Natural Language Processing.
  60. 60.Jianxiong Xiao, James Hays, Krista A Ehinger, Aude Oliva, and Antonio Torralba. 2010. Sun database: Large-scale scene recognition from abbey to zoo. In Computer Vision and Pattern Recognition.
  61. 61.Zhiyang Xu, Ying Shen, and Lifu Huang. 2023. Multi-instruct: Improving multi-modal zero-shot learning via instruction tuning. Annual Meeting of the Association for Computational Linguistics.
  62. 62.Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020. mT5: A massively multilingual pre-trained text-to-text transformer. Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  63. 63.Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang. 2022. An empirical study of GPT-3 for few-shot knowledge-based VQA. In Proceedings of the AAAI Conference on Artificial Intelligence.
  64. 64.Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al. 2023. mplug-owl: Modularization empowers large language models with multimodality. arXiv preprint arXiv:2304.14178.
  65. 65.Victor Zhong, Weijia Shi, Wen-tau Yih, and Luke Zettlemoyer. 2022. RoMQA: A benchmark for robust, multi-evidence, multi-answer question answering. arXiv preprint arXiv:2210.14353.
  66. 66.Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. 2020. Unified vision-language pre-training for image captioning and VQA. In Proceedings of the AAAI Conference on Artificial Intelligence.
  67. 67.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Chen, Y., et al. “Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 14948–68, https://doi.org/10.18653/v1/2023.emnlp-main.925.
APA
Chen, Y., Hu, H., Luan, Y., Sun, H., Changpinyo, S., Ritter, A., & Chang, M.-W. (2023). Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 14948–14968. https://doi.org/10.18653/v1/2023.emnlp-main.925
Chicago
Chen, Y., H. Hu, Y. Luan, et al. 2023. “Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 14948–68. https://doi.org/10.18653/v1/2023.emnlp-main.925.
Harvard
Chen, Y. et al. (2023) “Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 14948–14968. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.925.
Vancouver
1. Chen Y, Hu H, Luan Y, Sun H, Changpinyo S, Ritter A, Chang M-W (2023) Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 14948–14968

BibTeX

@inproceedings{chen-etal-2023-pre-trained,
    title = "Can Pre-trained Vision and Language Models Answer Visual Information-Seeking Questions?",
    author = "Chen, Yang  and
      Hu, Hexiang  and
      Luan, Yi  and
      Sun, Haitian  and
      Changpinyo, Soravit  and
      Ritter, Alan  and
      Chang, Ming-Wei",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.925/",
    doi = "10.18653/v1/2023.emnlp-main.925",
    pages = "14948--14968"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/