KAT: A Knowledge Augmented Transformer for Vision-and-Language

Liangke GuiBorui WangQiuyuan HuangAlexander HauptmannYonatan BiskJianfeng Gao

article2022NAACL229 citations

Proposes an encoder-decoder framework that jointly reasons over explicit knowledge retrieved from Wikidata and implicit knowledge extracted from GPT-3, setting a new state of the art on open-domain visual question answering while improving prediction interpretability.

Listen

Modern artificial intelligence systems frequently struggle with multimodal reasoning tasks when answering open-ended questions that require external information beyond the image itself. Existing approaches typically rely on either internal commonsense knowledge stored implicitly in large language models or factual data retrieved explicitly from structured knowledge bases. However, relying on either source alone creates limitations: language models often lack grounded factual precision, while structured knowledge bases often introduce irrelevant noise and lack commonsense reasoning.

The article introduces and evaluates the Knowledge Augmented Transformer (KAT), a unified sequence-to-sequence model designed to integrate both explicit structured facts and implicit commonsense knowledge for knowledge-intensive vision-and-language tasks.

To evaluate the system, the authors conducted experiments on the Outside Knowledge Visual Question Answering (OK-VQA) benchmark, which contains over 14,000 images and open-domain questions. The approach extracts explicit knowledge by using a contrastive visual model (CLIP) to match image patches against a curated knowledge base of over 423,000 Wikidata entities. In parallel, it queries a large language model (GPT-3) to retrieve implicit commonsense knowledge along with supporting rationales. Both knowledge streams are processed through an encoder-decoder transformer architecture (T5) with a dedicated cross-attention reasoning module that jointly evaluates all evidence to generate open-ended textual answers.

The findings demonstrate significant performance gains. KAT established a new state-of-the-art accuracy of 54.41% on the OK-VQA benchmark, outperforming the previous top system (PICa-Full at 48.0%) by more than 6 absolute percentage points. Ablation experiments showed that combining explicit and implicit knowledge yields an approximate 4 percentage point improvement over using either knowledge source alone. Furthermore, the specialized joint reasoning module outperformed simple text concatenation of knowledge sources by 2.43 percentage points, demonstrating its effectiveness at filtering irrelevant data. Performance scaled consistently as the number of retrieved explicit knowledge entries increased.

These results demonstrate that future multimodal AI applications, such as autonomous agents and visual search engines, should not rely solely on larger model parameters. Integrating structured retrieval with generative language models provides a more accurate, cost-effective, and interpretable alternative to purely scaling model size. It also shifts model evaluation from restrictive multiple-choice classification to flexible, open-vocabulary generation.

Organizations developing knowledge-based vision systems should adopt hybrid architectures that combine retrieval mechanisms with generative reasoning rather than depending on single-source paradigms. Future development should focus on enhancing visual-semantic entity alignment, expanding explicit knowledge sources beyond curated subsets, and optimizing multi-source retrieval efficiency.

While the findings are supported by strong empirical results, the authors note clear boundary conditions: the explicit knowledge base was limited to 423,520 English Wikidata entities across eight categories, and the implicit knowledge retriever depended on an external, frozen language model. Consequently, teams implementing these architectures in broader domains should validate retrieval quality and entity coverage before production deployment.

Cover for KAT: A Knowledge Augmented Transformer for Vision-and-Language

Abstract

The primary focus of recent work with large-scale transformers has been on optimizing the amount of information packed into the model's parameters. In this work, we ask a complementary question: Can multimodal transformers leverage explicit knowledge in their reasoning? Existing, primarily unimodal, methods have explored approaches under the paradigm of knowledge retrieval followed by answer prediction, but leave open questions about the quality and relevance of the retrieved knowledge used, and how the reasoning processes over implicit and explicit knowledge should be integrated. To address these challenges, we propose a - Knowledge Augmented Transformer (KAT) - which achieves a strong state-of-the-art result (+6% absolute) on the open-domain multimodal task of OK-VQA. Our approach integrates implicit and explicit knowledge in an encoder-decoder architecture, while still jointly reasoning over both knowledge sources during answer generation. Additionally, explicit knowledge integration improves interpretability of model predictions in our analysis. Code and pre-trained models are released at https://github.com/guilk/KAT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Overview
  • 3.2 Explicit Knowledge Retrieval
  • 3.2.1 Explicit Knowledge Extraction
  • 3.2.2 Knowledge Base Construction
  • 3.3 Implicit Knowledge Retrieval
  • 3.4 KAT Model
  • 4 Experiment
  • 4.1 Dataset
  • 4.2 Implementation Details
  • 4.3 Comparison with Existing Approaches
  • 5 Ablation Study
  • 5.1 Effectiveness of Knowledge Reasoning
  • 5.2 Extracting Explicit Knowledge
  • 5.3 Category Results on OK-VQA
  • 5.4 Qualitative Analysis
  • 6 Conclusion
  • Acknowledgement
  • References
  • Appendix
  • A Figure of Explicit Knowledge
  • B Examples of Prompts of Implicit Knowledge
  • C Analysis on More Examples

Knowls

  1. Knowl 1 — Knowledge Augmented Transformer Architecture for Knowledge-Based VQA

    model/method

    The Knowledge Augmented Transformer (KAT) is an encoder-decoder framework designed for open-domain, knowledge-based Visual Question Answering (such as the OK-VQA task). KAT formulates question answering as an open-ended autoregressive sequence generation problem rather than a classification task over a fixed candidate set.

    The overall architecture operates in three stages:

    1. Explicit Knowledge Retrieval: Fine-grained regional visual patches extracted from the input image are matched against a structured knowledge base (derived from Wikidata) via a contrastive vision-language model (CLIP) to retrieve top-mm explicit entity descriptions xexpx^{\text{exp}}.
    2. Implicit Knowledge Retrieval: The image is captioned, and few-shot in-context prompts are supplied to a frozen large language model (GPT-3 175B) to generate tentative answer candidates along with natural language supporting rationales/evidence, forming implicit knowledge ximpx^{\text{imp}}.
    3. Knowledge Reasoning and Generation: A sequence-to-sequence model (initialized from pre-trained T5) separately encodes question-knowledge pairs for each explicit and implicit entry to prevent noisy cross-contamination, concatenates the resulting representations, and uses cross-attention inside the decoder to jointly reason over all knowledge sources during autoregressive answer decoding.
  2. Knowl 2 — Contrastive Vision-Language Retrieval for Explicit Knowledge Extraction

    model/method

    To ground image content to structured factual entities without relying on fixed-vocabulary object detector tags, KAT retrieves knowledge from an external database using visual-semantic contrastive representations.

    Given an input image viv_i, a sliding window with a predefined stride extracts NN local image regions {vi1,vi2,…,viN}\{v_i^1, v_i^2, \dots, v_i^N\}. A dense multimodal image encoder Eimg(⋅)E_{\text{img}}(\cdot) maps each region patch into a drd_r-dimensional embedding space. An entity encoder Eent(⋅)E_{\text{ent}}(\cdot) encodes knowledge entries e∈Ke \in \mathcal{K} (where each entry ee is a text string concatenating an entity's name and its textual description). The similarity between image patch vijv_i^j and knowledge entry ee is computed via the inner product of their normalized representations:

    sim(vij,e)=Eent(e)TEimg(vij)\text{sim}(v_i^j, e) = E_{\text{ent}}(e)^T E_{\text{img}}(v_i^j)

    For each patch vijv_i^j, the top-kk closest knowledge entries in K\mathcal{K} are retrieved using dense similarity search (indexed with FAISS). Across all NN patches, the top-mm unique retrieved knowledge entries ranked by similarity score are selected as the explicit knowledge set xexp={e1,e2,…,em}x^{\text{exp}} = \{e_1, e_2, \dots, e_m\}. In the default implementation, CLIP (ViT-B/16 variant) provides EimgE_{\text{img}} and EentE_{\text{ent}} using the representation of the [CLS] token.

  3. Knowl 3 — Implicit Knowledge and Rationale Generation via Prompted Large Language Models

    model/method

    To incorporate commonsense knowledge and multi-step inference not captured by entity descriptions, KAT extracts implicit knowledge from a frozen large language model (GPT-3, 175B parameters) using a two-stage prompting mechanism:

    1. Tentative Answer Generation: The input image viv_i is converted into a textual caption CC using an image captioning model. A text prompt is constructed consisting of an instruction, the caption CC, the question qiq_i, and context-question-answer triplets selected from training examples that are semantically closest to (vi,qi)(v_i, q_i). Querying frozen GPT-3 with this prompt yields candidate answers.
    2. Supporting Evidence Generation: For each tentative answer candidate aa, GPT-3 is queried with a secondary prompt of the format: "(question qi)? (answer a). This is because"\text{"(question } q_i \text{)? (answer } a \text{). This is because"} GPT-3 completes the prompt by generating a explanatory rationale explaining why aa is the answer.

    The combined set of generated answer candidates and their supporting rationales constitutes the implicit knowledge source ximp={(a1,evidence1),…,(ap,evidencep)}x^{\text{imp}} = \{ (a_1, \text{evidence}_1), \dots, (a_p, \text{evidence}_p) \}, containing pp candidate-evidence pairs.

  4. Knowl 4 — Knowledge Reasoning Module with Decoupled Encoding and Joint Cross-Attention

    model/method

    To prevent noisy or irrelevant retrieved knowledge from contaminating the representations of valid knowledge, KAT processes each question-knowledge pair independently in the encoder before jointly attending to them in the decoder.

    1. Sentinel Token Formatting: Each explicit knowledge entry ej=(entityj,descriptionj)e_j = (\text{entity}_j, \text{description}_j) is combined with question qiq_i using sentinel tokens: question: qi entity: entityj description: descriptionj\text{question: } q_i \text{ entity: } \text{entity}_j \text{ description: } \text{description}_j Similarly, each implicit knowledge entry (ak,evidencek)(a_k, \text{evidence}_k) is formatted as: question: qi candidate: ak evidence: evidencek\text{question: } q_i \text{ candidate: } a_k \text{ evidence: } \text{evidence}_k

    2. Decoupled Encoding: Each formatted string is independently passed through the transformer encoder. The token representations from the final encoder layer are mean-pooled across token positions, yielding explicit embedding matrix Xexp∈Rm×dX^{\text{exp}} \in \mathbb{R}^{m \times d} and implicit embedding matrix Ximp∈Rp×dX^{\text{imp}} \in \mathbb{R}^{p \times d}, where dd is the hidden embedding dimension, mm is the number of explicit entries, and pp is the number of implicit entries.

    3. Joint Cross-Attention: The explicit and implicit matrices are concatenated into a global knowledge representation X=[Xexp;Ximp]∈R(m+p)×dX = [X^{\text{exp}}; X^{\text{imp}}] \in \mathbb{R}^{(m+p) \times d}. During decoder generation, let H∈RdH \in \mathbb{R}^{d} be the output of the decoder's prior self-attention layer. The decoder applies scaled dot-product cross-attention over XX:

      Qv=softmax(QKTd)VQ_v = \text{softmax}\left( \frac{Q K^T}{\sqrt{d}} \right) V

      where Q=WQHQ = W_Q H, K=WKXK = W_K X, and V=WVXV = W_V X, with learnable linear projection matrices WQ,WK,WV∈Rd×dW_Q, W_K, W_V \in \mathbb{R}^{d \times d}. This enables the model to dynamically weight both knowledge sources simultaneously during output token prediction.

  5. Knowl 5 — Autoregressive Answer Generation Training Objective

    equation

    KAT generates answers token by token across the full vocabulary rather than selecting an answer from a fixed predefined candidate set. Given retrieved explicit knowledge xexpx^{\text{exp}} and implicit knowledge ximpx^{\text{imp}}, model parameters θ\theta are trained end-to-end to maximize the log-likelihood of target answer sequence y=(y1,y2,…,yn)y = (y_1, y_2, \dots, y_n) via cross-entropy loss:

    LCE=−∑t=1nlog⁡pθ(yt∣y<t,xexp,ximp)L_{\text{CE}} = - \sum_{t=1}^n \log p_\theta(y_t \mid y_{<t}, x^{\text{exp}}, x^{\text{imp}})

    where yty_t is the target token at step tt, y<ty_{<t} denotes the previously generated prefix tokens (y1,…,yt−1)(y_1, \dots, y_{t-1}), and nn is the sequence length of the target answer.

  6. Knowl 6 — Curated Wikidata Explicit Knowledge Base

    data/table

    To support explicit knowledge retrieval for visual reasoning, an external knowledge base K\mathcal{K} was extracted from the English Wikidata dump (dated September 20, 2021, containing 95,870,584 raw entities). The dump was filtered to retain 8 core subclasses covering everyday entities, removing any entries with empty or non-English labels and descriptions. Each item is stored as an entity triplet ⟨Wikidata ID,Label,Description⟩\langle \text{Wikidata ID}, \text{Label}, \text{Description} \rangle.

    Subclass Wikidata ID Number of Entries
    Role Q214339 162,027
    Point of interest Q960648 85,900
    Tool Q39546 78,621
    Vehicle Q42889 44,274
    Animal Q729 18,581
    Clothing Q11460 17,711
    Company Q891723 12,173
    Sport Q349 4,233
    Total 423,520

    The resulting knowledge base contains 423,520 triplets providing fine-grained entity definitions (e.g., ⟨Q2813,Coca-Cola,carbonated brown colored soft drink⟩\langle \text{Q2813}, \text{Coca-Cola}, \text{carbonated brown colored soft drink} \rangle).

  7. Knowl 7 — OK-VQA Benchmark Performance Comparison

    data/table

    Performance of KAT (initialized with T5-Large, 770M parameters) evaluated on the full testing set of the OK-VQA benchmark against standard baselines, multimodal knowledge methods, and GPT-3-based models.

    Method Knowledge Resources Accuracy (%)
    No knowledge
    Q only - 14.93
    Vanilla T5 - 18.56
    MLP - 20.67
    BAN - 25.10
    MUTAN - 26.41
    With explicit knowledge
    BAN+AN Wikipedia 25.61
    BAN+KG-AUG Wikipedia+ConceptNet 26.71
    MUTAN+AN Wikipedia 27.84
    ConceptBERT ConceptNet 33.66
    KRISP Wikipedia+ConceptNet 38.35
    Vis-DPR Google Search 39.20
    MAVEx Wikipedia+ConceptNet+Google Images 39.40
    With implicit knowledge (GPT-3)
    PICa-Base Frozen GPT-3 (175B) 43.30
    PICa-Full Frozen GPT-3 (175B) 48.00
    KAT Variants
    KAT-explicit (w/ reasoning) Wikidata 44.25
    KAT-implicit (w/ reasoning) Frozen GPT-3 (175B) 49.72
    KAT (w/o reasoning) Wikidata+Frozen GPT-3 (175B) 51.97
    KAT (single model) Wikidata+Frozen GPT-3 (175B) 53.09
    KAT (ensemble, 3 seeds) Wikidata+Frozen GPT-3 (175B) 54.41

    KAT achieves 53.09% accuracy with a single model and 54.41% when ensembled across 3 random seeds, outperforming prior state of the art (PICa-Full at 48.00%) by +6.41% absolute, and outperforming explicit-retrieval methods (such as MAVEx at 39.40%) by +15.01% absolute.

  8. Knowl 8 — Ablation Study on Model Scale, Knowledge Types, and Reasoning Mechanism

    data/table

    Ablation experiments evaluate the impact of transformer capacity (T5-Base with 220M parameters vs. T5-Large with 770M parameters), the inclusion of explicit vs. implicit knowledge, and the knowledge reasoning module architecture on OK-VQA accuracy.

    Model Architecture Knowledge Used OK-VQA Accuracy (%)
    Base (220M) Large (770M) Explicit Implicit
    ✓ 18.56
    ✓ ✓ 40.93
    ✓ ✓ 44.25
    ✓ ✓ 47.60
    ✓ ✓ 49.72
    ✓ ✓ ✓ 50.58
    ✓ ✓ ✓ 54.41

    Key findings from these ablations include:

    1. Complementarity of Knowledge: Combining both explicit and implicit knowledge provides a consistent gain of ~4% accuracy over using either source alone (e.g., 54.41% vs. 44.25% explicit-only and 49.72% implicit-only on T5-Large).
    2. Model Capacity: Scaling the backbone from T5-Base to T5-Large improves overall accuracy from 50.58% to 54.41% when both knowledge sources are used.
    3. Reasoning Module vs. Flat Concatenation: Removing the decoupled knowledge reasoning module and instead concatenating all explicit and implicit knowledge into a single text sequence (capped at 256 tokens) yields 51.97% accuracy on T5-Large, which is 2.44% lower than KAT's decoupled reasoning module (54.41%). This indicates that separate encoding avoids introducing noise across unrelated knowledge entries.
  9. Knowl 9 — Effect of Retrieved Explicit Knowledge Quantity and Visual Backbone

    empirical result

    Experiments on KAT-base measuring the effect of varying the number of retrieved explicit Wikidata entities mm from 0 to 60, comparing two visual retrieval backbones (CLIP ViT-B/16 vs. CLIP ResNet-50):

    1. Zero-Entity Baseline: With 0 retrieved entities (implicit knowledge only), the model achieves 47.60% accuracy.
    2. Monotonic Gain with Entity Count: Accuracy increases as more explicit entities are integrated, reaching 48.62% at m=5m=5, 49.16% at m=10m=10, 49.50% at m=20m=20, 49.97% at m=30m=30, and peaking around 50.58% at m=40m=40 (using ViT-B/16).
    3. Backbone Comparison: CLIP ViT-B/16 consistently outperforms CLIP ResNet-50 across all entity counts (e.g., at m=40m=40, ViT-B/16 achieves 50.58% vs. 49.23% for ResNet-50). However, the performance gap between visual backbones narrows as more entities are retrieved, because the knowledge reasoning module adaptively attends to relevant knowledge entries even when retrieved candidates contain noise.
  10. Knowl 10 — Category-Specific Accuracy on OK-VQA

    data/table

    Performance breakdown across the 11 question categories of the OK-VQA testing set comparing explicit-only (Exp), implicit-only (Imp), and full KAT models (Ours):

    Question Category Exp (%) Imp (%) Ours (%) Δ\Delta (Ours −- best single)
    Plants and Animals 42.2 51.5 54.7 +3.2
    Science and Technology 44.4 43.3 52.8 +8.3
    Sports and Recreation 49.7 53.8 60.4 +6.7
    Geo, History, Lang, and Culture 45.6 45.4 55.8 +10.2
    Brands, Companies, and Products 41.7 38.2 48.5 +6.8
    Vehicles and Transportation 41.5 42.9 51.3 +8.4
    Cooking and Food 47.9 47.7 52.7 +4.8
    Weather and Climate 51.7 46.3 54.8 +3.1
    People and Everyday 43.1 44.4 51.5 +7.1
    Objects, Material and Clothing 42.9 45.4 49.3 +3.9
    Other 41.5 50.2 51.2 +1.0

    Explicit knowledge alone outperforms implicit knowledge alone in categories that rely on recognizing specific visual entities and products (e.g., 'Brands, Companies, and Products' at 41.7% vs. 38.2%; 'Weather and Climate' at 51.7% vs. 46.3%). Conversely, implicit knowledge dominates in categories requiring commonsense biological and behavioural facts (e.g., 'Plants and Animals' at 51.5% vs. 42.2%). Joint reasoning over both sources yields substantial improvements across every category, with the largest absolute gains occurring in 'Geo, History, Lang, and Culture' (+10.2%), 'Vehicles and Transportation' (+8.4%), and 'Science and Technology' (+8.3%).

Coverage note — None was omitted; all key architectural components, algorithms, loss functions, curated dataset statistics, main leaderboard results, ablation studies, and category-level analyses are covered.

References

  1. 1.Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In CVPR.
  2. 2.Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007. Dbpedia: A nucleus for a web of open data. In The semantic web, pages 722–735. Springer.
  3. 3.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165.
  4. 4.Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017. Reading wikipedia to answer open-domain questions. arXiv preprint arXiv:1704.00051.
  5. 5.Kezhen Chen, Qiuyuan Huang, Yonatan Bisk, Daniel McDuff, and Jianfeng Gao. 2021a. Kb-vlp: Knowledge based vision and language pretraining. In ICML, workshop.
  6. 6.Kezhen Chen, Qiuyuan Huang, Daniel McDuff, Xiang Gao, Hamid Palangi, Jianfeng Wang, Kenneth Forbus, and Jianfeng Gao. 2021b. Nice: Neural image commenting with empathy. In EMNLP.
  7. 7.Kezhen Chen, Qiuyuan Huang, Paul Smolensky, Kenneth Forbus, and Jianfeng Gao. 2020. Learning inference rules with neural tp-reasoner. In NeurIPS, workshop.
  8. 8.François Garderes, Maryam Ziaeefard, Baptiste Abeloos, and Freddy Lecue. 2020. Conceptbert: Concept-aware representation for visual question answering. In EMNLP.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR.
  10. 10.Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. In EACL.
  11. 11.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML.
  12. 12.Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with gpus. IEEE Transactions on Big Data.
  13. 13.Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open-domain question answering. In EMNLP.
  14. 14.Kenton Lee, Ming-Wei Chang, and Kristina Toutanova. 2019. Latent retrieval for weakly supervised open domain question answering. arXiv preprint arXiv:1906.00300.
  15. 15.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2020a. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In ACL.
  16. 16.Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020b. Retrieval-augmented generation for knowledge-intensive nlp tasks. In NeurIPS.
  17. 17.Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou. 2020a. Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training. Proceedings of AAAI.
  18. 18.Guohao Li, Xin Wang, and Wenwu Zhu. 2020b. Boosting visual question answering with context-aware knowledge aggregation. In ACM MM.
  19. 19.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557.
  20. 20.Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020c. Oscar: Object-semantics aligned pre-training for vision-language tasks. In ECCV.
  21. 21.Hugo Liu and Push Singh. 2004. Conceptnet—a practical commonsense reasoning tool-kit. BT technology journal, 22(4):211–226.
  22. 22.Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In ICLR.
  23. 23.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In NeurIPS.
  24. 24.Man Luo, Yankai Zeng, Pratyay Banerjee, and Chitta Baral. 2021. Weakly-supervised visual-retriever-reader for knowledge-based question answering. In EMNLP.
  25. 25.Kenneth Marino, Xinlei Chen, Devi Parikh, Abhinav Gupta, and Marcus Rohrbach. 2021. Krisp: Integrating implicit and symbolic knowledge for open-domain knowledge-based vqa. In CVPR.
  26. 26.Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. Ok-vqa: A visual question answering benchmark requiring external knowledge. In CVPR.
  27. 27.Medhini Narasimhan, Svetlana Lazebnik, and Alexander Schwing. 2018. Out of the box: Reasoning with graph convolution nets for factual visual question answering. NeurIPS.
  28. 28.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In ICML.
  29. 29.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. JMLR, 21(140):1–67.
  30. 30.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS.
  31. 31.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020. Vl-bert: Pre-training of generic visual-linguistic representations. In ICLR.
  32. 32.Hao Tan and Mohit Bansal. 2019. Lxmert: Learning cross-modality encoder representations from transformers. In EMNLP.
  33. 33.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In NeurIPS.
  34. 34.Denny Vrandecic and Markus Krotzsch. 2014. Wikidata: a free collaborative knowledgebase. Communications of the ACM.
  35. 35.Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel. 2017a. Fvqa: Fact-based visual question answering. TPAMI, 40(10):2413–2427.
  36. 36.Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony Dick. 2017b. Explicit knowledge-based reasoning for visual question answering. In IJCAI.
  37. 37.Peter West, Chandra Bhagavatula, Jack Hessel, Jena D. Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2021. Symbolic knowledge distillation: from general language models to commonsense models. In ArXiv.
  38. 38.Jialin Wu, Jiasen Lu, Ashish Sabharwal, and Roozbeh Mottaghi. 2022. Multi-modal answer validation for knowledge-based vqa. In AAAI.
  39. 39.Zhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu, Yumao Lu, Zicheng Liu, and Lijuan Wang. 2022. An empirical study of gpt-3 for few-shot knowledge-based vqa. In AAAI.

Citation

MLA
Gui, L., et al. “KAT: A Knowledge Augmented Transformer for Vision-and-Language”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 956–68, https://doi.org/10.18653/v1/2022.naacl-main.70.
APA
Gui, L., Wang, B., Huang, Q., Hauptmann, A. G., Bisk, Y., & Gao, J. (2022). KAT: A Knowledge Augmented Transformer for Vision-and-Language. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 956–968. https://doi.org/10.18653/v1/2022.naacl-main.70
Chicago
Gui, L., B. Wang, Q. Huang, A. G. Hauptmann, Y. Bisk, and J. Gao. 2022. “KAT: A Knowledge Augmented Transformer for Vision-and-Language”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 956–68. https://doi.org/10.18653/v1/2022.naacl-main.70.
Harvard
Gui, L. et al. (2022) “KAT: A Knowledge Augmented Transformer for Vision-and-Language”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 956–968. Available at: https://doi.org/10.18653/v1/2022.naacl-main.70.
Vancouver
1. Gui L, Wang B, Huang Q, Hauptmann AG, Bisk Y, Gao J (2022) KAT: A Knowledge Augmented Transformer for Vision-and-Language. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 956–968

BibTeX

@inproceedings{gui-etal-2022-kat,
    title = "{KAT}: A Knowledge Augmented Transformer for Vision-and-Language",
    author = "Gui, Liangke  and
      Wang, Borui  and
      Huang, Qiuyuan  and
      Hauptmann, Alexander  and
      Bisk, Yonatan  and
      Gao, Jianfeng",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.70/",
    doi = "10.18653/v1/2022.naacl-main.70",
    pages = "956--968"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/