KAT: A Knowledge Augmented Transformer for Vision-and-Language
Liangke GuiBorui WangQiuyuan HuangAlexander HauptmannYonatan BiskJianfeng Gao
Proposes an encoder-decoder framework that jointly reasons over explicit knowledge retrieved from Wikidata and implicit knowledge extracted from GPT-3, setting a new state of the art on open-domain visual question answering while improving prediction interpretability.
Modern artificial intelligence systems frequently struggle with multimodal reasoning tasks when answering open-ended questions that require external information beyond the image itself. Existing approaches typically rely on either internal commonsense knowledge stored implicitly in large language models or factual data retrieved explicitly from structured knowledge bases. However, relying on either source alone creates limitations: language models often lack grounded factual precision, while structured knowledge bases often introduce irrelevant noise and lack commonsense reasoning.
The article introduces and evaluates the Knowledge Augmented Transformer (KAT), a unified sequence-to-sequence model designed to integrate both explicit structured facts and implicit commonsense knowledge for knowledge-intensive vision-and-language tasks.
To evaluate the system, the authors conducted experiments on the Outside Knowledge Visual Question Answering (OK-VQA) benchmark, which contains over 14,000 images and open-domain questions. The approach extracts explicit knowledge by using a contrastive visual model (CLIP) to match image patches against a curated knowledge base of over 423,000 Wikidata entities. In parallel, it queries a large language model (GPT-3) to retrieve implicit commonsense knowledge along with supporting rationales. Both knowledge streams are processed through an encoder-decoder transformer architecture (T5) with a dedicated cross-attention reasoning module that jointly evaluates all evidence to generate open-ended textual answers.
The findings demonstrate significant performance gains. KAT established a new state-of-the-art accuracy of 54.41% on the OK-VQA benchmark, outperforming the previous top system (PICa-Full at 48.0%) by more than 6 absolute percentage points. Ablation experiments showed that combining explicit and implicit knowledge yields an approximate 4 percentage point improvement over using either knowledge source alone. Furthermore, the specialized joint reasoning module outperformed simple text concatenation of knowledge sources by 2.43 percentage points, demonstrating its effectiveness at filtering irrelevant data. Performance scaled consistently as the number of retrieved explicit knowledge entries increased.
These results demonstrate that future multimodal AI applications, such as autonomous agents and visual search engines, should not rely solely on larger model parameters. Integrating structured retrieval with generative language models provides a more accurate, cost-effective, and interpretable alternative to purely scaling model size. It also shifts model evaluation from restrictive multiple-choice classification to flexible, open-vocabulary generation.
Organizations developing knowledge-based vision systems should adopt hybrid architectures that combine retrieval mechanisms with generative reasoning rather than depending on single-source paradigms. Future development should focus on enhancing visual-semantic entity alignment, expanding explicit knowledge sources beyond curated subsets, and optimizing multi-source retrieval efficiency.
While the findings are supported by strong empirical results, the authors note clear boundary conditions: the explicit knowledge base was limited to 423,520 English Wikidata entities across eight categories, and the implicit knowledge retriever depended on an external, frozen language model. Consequently, teams implementing these architectures in broader domains should validate retrieval quality and entity coverage before production deployment.
- Paper: OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge, Kenneth Marino et al. (2019). This paper establishes the OK-VQA benchmark requiring external knowledge that KAT is explicitly designed to tackle and evaluate on.
- Paper: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Patrick Lewis et al. (2020). This foundational work introduces retrieval-augmented generation (RAG) for knowledge-intensive generation tasks, providing the architectural paradigm KAT extends to multimodal vision-and-language reasoning.
- Paper: Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering, Gautier Izacard et al. (2021). This work establishes the Fusion-in-Decoder architecture that utilizes T5 cross-attention over retrieved evidence, forming a core technical mechanism adapted by KAT's encoder-decoder transformer.
- Paper: What Makes Good In-Context Examples for GPT-3?, Jiachang Liu et al. (2021). This study introduces retrieval-guided in-context prompting for GPT-3 (KATE), which directly informs KAT's strategy for querying implicit commonsense knowledge and rationales from large language models.
- Paper: ERNIE: Enhanced Language Representation with Informative Entities, Zhengyan Zhang et al. (2019). This paper demonstrates how to integrate structured knowledge base entities into transformer models, laying key groundwork for KAT's explicit knowledge integration.
- Paper: VisualBERT: A Simple and Performant Baseline for Vision and Language, Liunian Harold Li et al. (2019). This paper introduces a unified self-attention transformer baseline for multimodal vision-and-language tasks that KAT builds upon.
- Paper: VQA: Visual Question Answering, Stanislaw Antol et al. (2015). This landmark paper defines the visual question answering task, establishing the foundational problem setting addressed by knowledge-augmented systems like KAT.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). PreFLMR directly advances the multimodal external knowledge retrieval stage essential to systems like KAT by scaling multi-task fine-grained late-interaction retrievers across vision-language benchmarks.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). This comprehensive roadmap formalizes the paradigms of unifying large language models with structured knowledge graphs, generalizing KAT's hybrid retrieval-and-reasoning approach.
- Paper: StructGPT: A General Framework for Large Language Model to Reason over Structured Data, Jinhao Jiang et al. (2023). StructGPT expands external structured knowledge utilization by enabling large language models to iteratively read and reason over heterogeneous structured schemas such as knowledge graphs and tables without task-specific fine-tuning.
- Paper: Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge Graphs, Liyi Chen et al. (2024). Plan-on-Graph advances beyond fixed-retrieval graph augmentation by implementing adaptive exploration and self-correcting path planning on knowledge graphs for language models.
- Paper: Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-Based Retrofitting, Xinyan Guan et al. (2024). This paper presents an autonomous retrofitting framework that leverages structured knowledge graphs to iteratively correct multi-step reasoning errors and hallucinations in generated language model outputs.
- Paper: When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories, Alex Troy Mallen et al. (2022). This study analyzes the boundary between parametric model memory and non-parametric external retrieval, offering critical empirical insights into when hybrid systems like KAT should trigger retrieval.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This survey provides a comprehensive synthesis of multimodal large language model architectures, connector interfaces, and training strategies that succeeded transformer-based pipelines like KAT.
