Multimodal Reasoning with Multimodal Knowledge Graph
Junlin LeeYequan WangJing LiMin Zhang
Presents a parameter-efficient framework that integrates multimodal knowledge graphs into large language models via relation graph attention networks and cross-modal alignment to curb hallucinations and improve multimodal reasoning.
Large language models often struggle with multimodal tasks that require joint reasoning over images and text, frequently generating inaccurate information due to outdated or deficient internal knowledge. While external text-only knowledge graphs have been used to supplement language models, their lack of visual context restricts cross-modal understanding and leaves systems vulnerable to hallucinations.
The article aims to demonstrate that integrating structured multimodal knowledge graphs into language models significantly improves multimodal reasoning while maintaining high computational efficiency.
To evaluate this framework, the authors developed a method called MR-MKG and tested it across two distinct benchmarks: multimodal question answering on the ScienceQA dataset (comprising over 21,000 instances across diverse academic subjects) and multimodal analogical reasoning on the MARS dataset. The architecture encodes retrieved multimodal graph subgraphs using a relation graph attention network, projects both knowledge and visual tokens into the language model's representation space via lightweight adapters, and refines image-text alignment using a contrastive matching task. Credibility is further established by pretraining on an 18,448-instance multimodal knowledge dataset derived from Visual Genome scene graphs, all while keeping the core language model and visual encoder completely frozen.
The evaluation yielded several key findings. First, the method achieved state-of-the-art average accuracy of 93.63% on ScienceQA, surpassing the previous leading benchmark by 1.95% and exceeding estimated human performance of 88.40%. Second, on the MARS analogical reasoning benchmark, the approach delivered a 10.4% absolute gain in top-one accuracy (Hits@1) over baseline visual language models, reaching 40.5%. Third, the framework demonstrated remarkable parameter efficiency by updating only approximately 2.25% of the total parameter count (77M to 248M parameters), outperforming fully fine-tuned models and 13-billion-parameter systems like LLaVA across nearly all evaluation categories. Fourth, ablation experiments confirmed that structured multimodal graph knowledge contributed the single largest performance increase, adding over 5.6 to 6.7 percentage points compared to models lacking graph grounding.
These findings indicate that multimodal knowledge graphs offer a cost-effective alternative to full-parameter retraining. By updating only a tiny fraction of adapter weights, organizations can deploy high-performing multimodal AI systems with lower compute budgets, reduced latency, and shorter training schedules. Furthermore, providing explicit graph-based factual grounding directly addresses the enterprise risk of model hallucinations in complex visual-textual tasks.
Organizations developing multimodal decision systems should prioritize adapter-based integration with domain-specific multimodal knowledge graphs rather than costly end-to-end model retraining. Before broad operational deployment, technical teams should optimize knowledge retrieval pipelines, as the system's accuracy remains bounded by the quality and relevance of retrieved graph data. Future work should validate the framework on larger foundational backbones and across higher-risk operational environments where entity ambiguity or incomplete knowledge graphs could lead to retrieval errors.
- Paper: Unifying Large Language Models and Knowledge Graphs: A Roadmap, Shirui Pan et al. (2023). This survey outlines fundamental strategies and paradigms for unifying large language models with knowledge graphs to combat hallucinations, establishing the conceptual foundation for multimodal graph extensions.
- Paper: KAT: A Knowledge Augmented Transformer for Vision-and-Language, Liangke Gui et al. (2022). This paper presents the integration of explicit external structured knowledge into transformer architectures for visual question answering, providing key architectural groundwork for knowledge-augmented vision-language models.
- Paper: JointLK: Joint Reasoning with Language Models and Knowledge Graphs for Commonsense Question Answering, Yueqing Sun et al. (2022). This work develops joint reasoning mechanisms between language models and graph neural networks over structured knowledge, which directly precedes graph-based knowledge reasoning in multimodal settings.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This survey provides essential background on modern multimodal large language model architectures, alignment interfaces, and training pipelines.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). This work introduces contrastive cross-modal alignment prior to multimodal fusion, providing fundamental principles for the alignment modules utilized in multimodal reasoning.
- Paper: MMCoQA: Conversational Question Answering over Text, Tables, and Images, Yongqi Li et al. (2022). This paper establishes the task and formulation of conversational question answering over diverse multimodal knowledge sources including text, tables, and images.
- Paper: MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI, Xiang Yue et al. (2023). This benchmark defines standards and challenges in expert-level multi-discipline multimodal reasoning that require external factual knowledge.
- Paper: InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Jinguo Zhu et al. (2025). This work advances native pretraining recipes and scaling for open-source multimodal models, extending parameter-efficient alignment and reasoning techniques.
- Paper: Qwen3-VL Technical Report, Shuai Bai et al. (2025). This report explores scaled multi-stage pretraining and cross-layer visual token injection to enhance complex multimodal reasoning without degrading pure-text capabilities.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). This paper investigates progressive model and test-time reasoning scaling to expand the multimodal reasoning frontiers of vision-language models.
- Paper: PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers, Weizhe Lin et al. (2024). This research scales fine-grained cross-modal retrieval over external knowledge bases, providing a complementary retrieval-augmented generation framework for knowledge-intensive multimodal QA.
- Paper: Multi-Modal Hallucination Control by Visual Information Grounding, Alessandro Favero et al. (2024). This work proposes decoding and preference optimization methods to control multimodal hallucinations, tackling the visual grounding and hallucination issues addressed by MR-MKG.
- Paper: Hallucination Augmented Contrastive Learning for Multimodal Large Language Model, Chaoya Jiang et al. (2024). This study develops counterfactual cross-modal contrastive learning to suppress multimodal hallucinations and refine feature alignment.
- Paper: T-SciQ: Teaching Multimodal Chain-of-Thought Reasoning via Large Language Model Signals for Science Question Answering, Lei Wang et al. (2024). This paper advances chain-of-thought multimodal reasoning by integrating external signals and rationales for complex scientific question answering.
- Paper: M³CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought, Qiguang Chen et al. (2024). This benchmark evaluates multi-domain, multi-step visual reasoning and chain-of-thought capabilities, providing a rigorous testbed for multimodal reasoning models.
