Multimodal Reasoning with Multimodal Knowledge Graph

Junlin LeeYequan WangJing LiMin Zhang

article2024ACL73 citations

Presents a parameter-efficient framework that integrates multimodal knowledge graphs into large language models via relation graph attention networks and cross-modal alignment to curb hallucinations and improve multimodal reasoning.

Listen

Large language models often struggle with multimodal tasks that require joint reasoning over images and text, frequently generating inaccurate information due to outdated or deficient internal knowledge. While external text-only knowledge graphs have been used to supplement language models, their lack of visual context restricts cross-modal understanding and leaves systems vulnerable to hallucinations.

The article aims to demonstrate that integrating structured multimodal knowledge graphs into language models significantly improves multimodal reasoning while maintaining high computational efficiency.

To evaluate this framework, the authors developed a method called MR-MKG and tested it across two distinct benchmarks: multimodal question answering on the ScienceQA dataset (comprising over 21,000 instances across diverse academic subjects) and multimodal analogical reasoning on the MARS dataset. The architecture encodes retrieved multimodal graph subgraphs using a relation graph attention network, projects both knowledge and visual tokens into the language model's representation space via lightweight adapters, and refines image-text alignment using a contrastive matching task. Credibility is further established by pretraining on an 18,448-instance multimodal knowledge dataset derived from Visual Genome scene graphs, all while keeping the core language model and visual encoder completely frozen.

The evaluation yielded several key findings. First, the method achieved state-of-the-art average accuracy of 93.63% on ScienceQA, surpassing the previous leading benchmark by 1.95% and exceeding estimated human performance of 88.40%. Second, on the MARS analogical reasoning benchmark, the approach delivered a 10.4% absolute gain in top-one accuracy (Hits@1) over baseline visual language models, reaching 40.5%. Third, the framework demonstrated remarkable parameter efficiency by updating only approximately 2.25% of the total parameter count (77M to 248M parameters), outperforming fully fine-tuned models and 13-billion-parameter systems like LLaVA across nearly all evaluation categories. Fourth, ablation experiments confirmed that structured multimodal graph knowledge contributed the single largest performance increase, adding over 5.6 to 6.7 percentage points compared to models lacking graph grounding.

These findings indicate that multimodal knowledge graphs offer a cost-effective alternative to full-parameter retraining. By updating only a tiny fraction of adapter weights, organizations can deploy high-performing multimodal AI systems with lower compute budgets, reduced latency, and shorter training schedules. Furthermore, providing explicit graph-based factual grounding directly addresses the enterprise risk of model hallucinations in complex visual-textual tasks.

Organizations developing multimodal decision systems should prioritize adapter-based integration with domain-specific multimodal knowledge graphs rather than costly end-to-end model retraining. Before broad operational deployment, technical teams should optimize knowledge retrieval pipelines, as the system's accuracy remains bounded by the quality and relevance of retrieved graph data. Future work should validate the framework on larger foundational backbones and across higher-risk operational environments where entity ambiguity or incomplete knowledge graphs could lead to retrieval errors.

arXiv: 2406.02030
Cover for Multimodal Reasoning with Multimodal Knowledge Graph

Abstract

Multimodal reasoning with large language models (LLMs) often suffers from hallucinations and the presence of deficient or outdated knowledge within LLMs. Some approaches have sought to mitigate these issues by employing textual knowledge graphs, but their singular modality of knowledge limits comprehensive cross-modal understanding. In this paper, we propose the Multimodal Reasoning with Multimodal Knowledge Graph (MR-MKG) method, which leverages multimodal knowledge graphs (MMKGs) to learn rich and semantic knowledge across modalities, significantly enhancing the multimodal reasoning capabilities of LLMs. In particular, a relation graph attention network is utilized for encoding MMKGs and a cross-modal alignment module is designed for optimizing image-text alignment. A MMKG-grounded dataset is constructed to equip LLMs with initial expertise in multimodal reasoning through pretraining. Remarkably, MR-MKG achieves superior performance while training on only a small fraction of parameters, approximately 2.25% of the LLM's parameter size. Experimental results on multimodal question answering and multimodal analogy reasoning tasks demonstrate that our MR-MKG method outperforms previous state-of-the-art models.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multimodal Knowledge Graph
  • 2.2 Knowledge-Augmented LLMs
  • 2.3 Multimodal Large Language Models
  • 3 Method
  • 3.1 MR-MKG Overview
  • 3.2 The MR-MKG Architecture
  • 3.3 Training Objectives
  • 4 Experiments
  • 4.1 Setups
  • 4.2 Main Results
  • 4.3 Ablation Study
  • 4.4 Further Analysis
  • 4.5 Qualitative Analysis
  • 5 Conclusion
  • Acknowledgements
  • Limitations
  • Ethical Considerations
  • References
  • A Additional Experimental Setups
  • A.1 Datasets
  • A.2 Multimodal Knowledge Graphs
  • A.3 MMKG-Grounded Dataset Construction
  • A.4 Large Language Models
  • A.5 Detailed Evaluation Metrics
  • A.6 Knowledge Retrieve Schemes
  • A.7 Implementation Details
  • B Additional Examples of Case Studies

Knowls

  1. Knowl 1 — MR-MKG integrates retrieved multimodal graph knowledge into frozen LLMs

    model/method

    MR-MKG is a parameter-efficient method for multimodal reasoning that supplies a large language model with three synchronized inputs: the task text, an input image, and a retrieved subgraph from a multimodal knowledge graph (MMKG). A language encoder embeds the text, a pretrained visual encoder embeds the image, and a relation graph attention network (RGAT) encodes the retrieved graph while preserving its entities and relations. Trainable visual and knowledge adapters map the visual and graph representations into the LLM’s word-embedding space. The resulting visual embeddings, knowledge-node embeddings, and text embeddings are concatenated as prompt embedding tokens and passed to the frozen LLM. A cross-modal alignment module additionally trains image-linked and text-linked graph entities to match. The page-3 architecture diagram depicts these three embedding streams, their adapter layers, prompt concatenation, and the auxiliary alignment objective. The LLM and visual encoder remain frozen; only the adapters, RGAT-based knowledge encoder, and alignment-related parameters are updated, amounting to approximately 2.25% of the LLM parameter count.

  2. Knowl 2 — Visual and graph embeddings are projected and attention-aligned with language embeddings

    model/method

    For an input text embedding HTH_T, image feature XIX_I, and retrieved MMKG subgraph GG, MR-MKG uses a linear visual adapter and a relation-aware graph encoder. The visual-language representation is HI=WIXI+bIH_I = W_I X_I + b_I, where WIW_I and bIb_I are trainable parameters and HIH_I has the same embedding dimension as the LLM word vectors. Text-conditioned visual features are computed as HI′=Softmax⁡(HTHI⊤/dk)HIH'_I = \operatorname{Softmax}(H_T H_I^{\top}/\sqrt{d_k})H_I, where dkd_k is the embedding dimension. The RGAT produces graph-node features XK=fRGAT(G)X_K=f_{\mathrm{RGAT}}(G) after node and relation initialization with CLIP representations. A knowledge adapter maps them to HK=WKXK+bKH_K=W_KX_K+b_K, where WKW_K and bKb_K are trainable. Knowledge features are then conditioned on either the text or visual representation, Q∈{HT,HI}Q\in\{H_T,H_I\}, using HK′=Softmax⁡(QHK⊤/dk)HKH'_K=\operatorname{Softmax}(QH_K^{\top}/\sqrt{d_k})H_K. The final prompt is the concatenation HK′⊕HI′⊕HTH'_K\oplus H'_I\oplus H_T, which is consumed by the LLM.

  3. Knowl 3 — Sub-MMKG retrieval expands relevant triples with one-hop graph context

    algorithm

    MR-MKG retrieves a task-specific sub-MMKG rather than passing the entire knowledge graph to the LLM.

    Input: Query text or image, MMKG triples, initial retrieval size n, final triple count N
    Output: Retrieved sub-MMKG G
    Embed the query and every MMKG triple in a shared representation space.
    Compute cosine similarity between the query representation and each triple representation.
    Select the entities E' occurring in the top-n most similar triples.
    Construct a candidate graph from E', their one-hop neighboring entities, and the relations connecting them.
    Rank candidate triples by cosine similarity to the query.
    Return the top-N ranked triples as G.

    For ScienceQA, retrieval uses text formed from the question, context, and answer options; for MARS, it is based on the question entity. The experiments use one-hop expansion and retrieve either 10 or 20 triples. This procedure is intended to reduce irrelevant graph noise while retaining structural relations that would be lost by simply verbalizing triples as a sequence.

  4. Knowl 4 — Cross-modal triplet alignment is combined with autoregressive answer generation

    equation

    The cross-modal alignment module randomly selects image entities from a retrieved MMKG and uses their corresponding text entities as positive matches. For image anchor xax_a, matching text entity xpx_p, nonmatching text entity xnx_n, Euclidean distance dd, number of selected image entities MM, and margin α\alpha, the alignment loss is La=∑i=1Mmax⁡(d(xa,xp)−d(xa,xn)+α,0)L_a=\sum_{i=1}^{M}\max\big(d(x_a,x_p)-d(x_a,x_n)+\alpha,0\big). The generative objective predicts the target answer tokens autoregressively: Lg=∑i=1Llog⁡p(Ai∣prompt,A0:i−1;θa)L_g=\sum_{i=1}^{L}\log p(A_i\mid\mathrm{prompt},A_{0:i-1};\theta_a), where AA is the target answer of length LL, AiA_i is its iith token, A0:i−1A_{0:i-1} is the preceding token sequence, and θa\theta_a denotes trainable adaptation parameters. MR-MKG combines the two objectives as L=Lg+λLaL=L_g+\lambda L_a, with trade-off weight λ\lambda. Training has two stages: pretraining on an MMKG-grounded visual question-answering dataset, followed by task-specific multimodal reasoning training. The LLM and visual encoder remain unchanged in both stages.

  5. Knowl 5 — The MMKG-grounded pretraining dataset is built from region-level Visual Genome graphs

    data/table

    The authors construct 18,448 MMKG-grounded pretraining instances from Visual Genome. Each instance contains an image, a region-based visual question-answer pair, and a modified scene graph representing knowledge needed to answer the question. Object entities from the original scene graph are linked to cropped images of their bounding-box regions through an image-of relation. Object attributes are linked to their entities through an attribute-of relation. Only region-based questions are retained because their associated scene graphs contain localized visual knowledge relevant to the answer; freeform whole-image questions are excluded. The resulting graph combines textual entities, object attributes, relations, and object images, allowing MR-MKG to pretrain both multimodal graph understanding and visual question answering.

  6. Knowl 6 — Evaluation uses domain-matched MMKGs for ScienceQA and MARS

    experimental setup

    MR-MKG is evaluated on ScienceQA multimodal multiple-choice question answering and MARS multimodal analogical reasoning. ScienceQA contains 12,726 training, 4,241 validation, and 4,241 test instances across 26 topics, 127 categories, and 379 skills; only 48.7% of its instances contain images. MARS contains 10,685 training, 1,228 validation, and 1,415 test instances and evaluates entity prediction with Hits@1, Hits@3, Hits@5, Hits@10, and mean reciprocal rank (MRR). ScienceQA uses the MMKG constructed from FreeBase, DBpedia, and YAGO, containing 812,899 triples, 14,951 entities, and 1,345 relations, with most entities associated with images. MARS uses MarKG, which contains 11,292 entities, 192 relations, and 76,424 images and shares entities and relations with MARS. The implementation uses ViT-L/32 for visual encoding and RGAT for graph encoding; FLAN-T5-3B, FLAN-T5-11B, FLAN-UL2-19B, and LLaMA-2 7B serve as language backbones across the experiments. ScienceQA uses Multimodal-CoT prompting, text-based retrieval, up to three task-training epochs, learning rate 4×10−54\times10^{-5}, batch size 1, and maximum input/output length 512. MMKG-grounded pretraining uses two epochs, learning rate 5×10−55\times10^{-5}, batch size 2, maximum input length 512, maximum output length 128, and AdamW with weight decay 0.01.

  7. Knowl 7 — MR-MKG substantially improves ScienceQA accuracy with few trainable parameters

    data/table

    On the ScienceQA test set, the metric is average multiple-choice accuracy in percent; the number in parentheses is the number of trainable parameters. The comparison demonstrates that MR-MKG reaches or exceeds stronger fully trained and larger multimodal baselines while updating only a small parameter subset. Reported average accuracies are: Human 88.40; GPT-3.5 with chain-of-thought 75.17; GPT-4 82.69; UnifiedQABase 74.11 with 223M trainable parameters; UnifiedQABase with MM-CoT 84.91 with 223M; UnifiedQALarge with MM-CoT 91.68 with 738M; LLaVA 90.92 with 13B; LLaMA-Adapter 85.19 with 1.8M; LaVIN-7B 89.41 with 3.8M; LaVIN-13B 90.83 with 5.4M; MR-MKG with FLAN-T5-3B 88.47 with 77M; MR-MKG with FLAN-T5-11B 92.78 with 248M; and MR-MKG with FLAN-UL2-19B 93.63 with 248M. MR-MKG with FLAN-T5-11B exceeds the previous UnifiedQABase MM-CoT result by 7.87 percentage points and LLaVA by 1.86 points while training 248M rather than 13B parameters. Increasing the FLAN-T5 backbone from 3B to 11B improves MR-MKG’s average accuracy by 4.31 points, while the FLAN-UL2-19B configuration gives the highest reported accuracy.

  8. Knowl 8 — MR-MKG achieves the strongest MARS analogical-reasoning results

    data/table

    On the MARS test set, MR-MKG with a visual LLaMA-2 7B backbone obtains Hits@1 0.4050.405, Hits@3 0.4650.465, Hits@5 0.4970.497, Hits@10 0.5310.531, and MRR 0.4490.449. The strongest non-MR-MKG baseline, Visual_LLaMA-2 7B, obtains 0.2860.286, 0.3730.373, 0.4090.409, 0.4570.457, and 0.3470.347 on these metrics, respectively. The strongest earlier multimodal knowledge-graph model, MarT_MKGformer, obtains 0.3010.301, 0.3670.367, 0.3800.380, 0.4080.408, and 0.3410.341. Other reported baselines include IKRL (0.266/0.294/0.301/0.310/0.2830.266/0.294/0.301/0.310/0.283), TransAE (0.261/0.285/0.289/0.293/0.2760.261/0.285/0.289/0.293/0.276), RSME (0.266/0.298/0.307/0.311/0.2850.266/0.298/0.307/0.311/0.285), MarT_VisualBERT (0.261/0.292/0.308/0.321/0.2840.261/0.292/0.308/0.321/0.284), MarT_ViLT (0.245/0.275/0.287/0.303/0.2660.245/0.275/0.287/0.303/0.266), MarT_ViLBERT (0.256/0.312/0.327/0.347/0.2920.256/0.312/0.327/0.347/0.292), and MarT_FLAVA (0.264/0.303/0.309/0.319/0.2880.264/0.303/0.309/0.319/0.288), where each sequence is ordered as Hits@1, Hits@3, Hits@5, Hits@10, and MRR. MR-MKG improves Hits@1 by 0.119 over Visual_LLaMA-2 7B and by 0.104 over MarT_MKGformer. Replacing the task-matched MarKG with the generic MMKG reduces Hits@1 from 0.4050.405 to 0.3840.384, while removing graph knowledge entirely gives 0.2860.286, showing that both the presence and task compatibility of multimodal graph knowledge matter.

  9. Knowl 9 — Ablations validate graph knowledge, multimodal content, alignment, and pretraining

    empirical result

    The ScienceQA ablation starts from Visual_FLAN-T5-11B and adds MR-MKG components cumulatively. Average accuracy rises from 86.08% without graph knowledge, to 91.74% with a textual KG, 92.21% when the KG is replaced by an MMKG, 92.36% after adding cross-modal alignment, and 92.78% after MMKG-grounded pretraining. On 1,973 manually selected ScienceQA examples containing images and belonging to natural or social science, the corresponding accuracies are 86.59%, 90.37%, 91.78%, and 92.32% for the visual baseline, KG, MMKG, and alignment configurations. On MARS, Hits@1 increases from 0.286 for Visual_LLaMA-2 7B to 0.352 with KG, 0.381 with MMKG, and 0.394 with cross-modal alignment. The larger gains on image-containing ScienceQA examples and on MARS support the claim that MMKG content and image-text alignment are most useful when visual knowledge is central to the task. Additional architectural comparisons show that RGAT is strongest: GNN, GAT, and RGAT obtain ScienceQA accuracies of 92.23%, 91.94%, and 92.78%, and MARS Hits@1 values of 0.391, 0.396, and 0.405, respectively. For ScienceQA retrieval, text-only, text-plus-image, and image-only retrieval obtain 92.78%, 92.03%, and 91.58%, respectively.

  10. Knowl 10 — MR-MKG depends on retrieval quality and has limited demonstrated scope

    limitation

    MR-MKG’s effectiveness depends directly on whether retrieval supplies relevant and unambiguous multimodal graph knowledge. If the retrieved sub-MMKG lacks information needed for the question, or retrieves semantically ambiguous entities, the LLM can receive misleading evidence; the paper’s error analysis includes a fish question for which a movie titled Big Fish is retrieved instead of useful fish knowledge. Adding too many triples can also introduce irrelevant information: performance improves as retrieval grows from zero toward roughly 10 triples but declines when the graph is expanded toward 30 triples. The evaluation covers only four LLM configurations and two tasks, ScienceQA and MARS, because of computational constraints, so performance on larger models, other datasets, and other domains remains uncertain. The authors therefore caution that the method should be verified before use on privacy-sensitive or high-risk applications.

Coverage note — Detailed qualitative case studies from the MARS and ScienceQA examples, the full category-wise ScienceQA breakdown, and plotted sensitivity curves were omitted because they illustrate or refine the quantified method and ablation findings rather than add independent load-bearing contributions.

References

  1. 1.Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  2. 2.Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007. Dbpedia: A nucleus for a web of open data. In Proceedings of the International Semantic Web Conference (ISWC), pages 722–735. Springer.
  3. 3.Jinheon Baek, Alham Fikri Aji, and Amir Saffari. 2023. Knowledge-augmented language model prompting for zero-shot knowledge graph question answering. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).
  4. 4.Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. Translating embeddings for modeling multi-relational data. Advances in neural information processing systems (NeurIPS), 26.
  5. 5.Yapei Chang, Kyle Lo, Tanya Goyal, and Mohit Iyyer. 2024. Booookscore: A systematic exploration of book-length summarization in the era of llms. In Proceedings of the International Conference on Learning Representations (ICLR).
  6. 6.Jiangjie Chen, Rui Xu, Ziquan Fu, Wei Shi, Zhongqiao Li, Xinbo Zhang, Changzhi Sun, Lei Li, Yanghua Xiao, and Hao Zhou. 2022a. E-kar: A benchmark for rationalizing natural language analogical reasoning. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL).
  7. 7.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton. 2020. Big self-supervised models are strong semi-supervised learners. Advances in neural information processing systems (NeurIPS), 33:22243–22255.
  8. 8.Xiang Chen, Ningyu Zhang, Lei Li, Shumin Deng, Chuanqi Tan, Changliang Xu, Fei Huang, Luo Si, and Huajun Chen. 2022b. Hybrid transformer with multi-level fusion for multimodal knowledge graph completion. In Proceedings of the International Conference on Research and Development in Information Retrieva (SIGIR), pages 904–915.
  9. 9.Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416.
  10. 10.Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra. 2023. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15180–15190.
  11. 11.Anna Gladkova, Aleksandr Drozd, and Satoshi Matsuoka. 2016. Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn’t. In Proceedings of the Student Research Workshop, SRW@HLT-NAACL 2016, pages 8–15.
  12. 12.Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. 2023. Language is not all you need: Aligning perception with language models. arXiv preprint arXiv:2302.14045.
  13. 13.Taichi Ishiwatari, Yuki Yasuda, Taro Miyazaki, and Jun Goto. 2020. Relation-aware graph attention networks with relational position encodings for emotion recognition in conversations. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7360–7370.
  14. 14.Erik Jones, Hamid Palangi, Clarisse Simões, Varun Chandrasekaran, Subhabrata Mukherjee, Arindam Mitra, Ahmed Awadallah, and Ece Kamar. 2024. Teaching language models to hallucinate less with synthetic tasks. In Proceedings of the International Conference on Learning Representations (ICLR).
  15. 15.Jiho Kim, Yeonsu Kwon, Yohan Jo, and Edward Choi. 2023. Kg-gpt: A general framework for reasoning on knowledge graphs using large language models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).
  16. 16.Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. In Proceedings of the International Conference on Machine Learning (ICML), volume 139, pages 5583–5594.
  17. 17.Jing Yu Koh, Daniel Fried, and Ruslan Salakhutdinov. 2023. Generating images with multimodal language models. Advances in Neural Information Processing Systems (NeurIPS).
  18. 18.Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al. 2017. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73.
  19. 19.Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. 2023a. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning (ICML), pages 19730–19742.
  20. 20.Kunchang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao. 2023b. Videochat: Chat-centric video understanding. arXiv preprint arXiv:2305.06355.
  21. 21.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. Visualbert: A simple and performant baseline for vision and language. arXiv preprint arXiv:1908.03557.
  22. 22.Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in Neural Information Processing Systems (NeurIPS).
  23. 23.Ye Liu, Hui Li, Alberto Garcia-Duran, Mathias Niepert, Daniel Onoro-Rubio, and David S Rosenblum. 2019. Mmkg: multi-modal knowledge graphs. In Proceedings of the Extended Semantic Web Conference (ESWC), pages 459–474. Springer.
  24. 24.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems (NeurIPS), pages 13–23.
  25. 25.Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to explain: Multimodal reasoning via thought chains for science question answering. Advances in Neural Information Processing Systems (NeurIPS), 35:2507–2521.
  26. 26.Gen Luo, Yiyi Zhou, Tianhe Ren, Shengxin Chen, Xiaoshuai Sun, and Rongrong Ji. 2023. Cheap and quick: Efficient vision-language instruction tuning for large language models. Advances in Neural Information Processing Systems (NeurIPS).
  27. 27.Debjyoti Mondal, Suraj Modi, Subhadarshi Panda, Rituraj Singh, and Godawari Sudhakar Rao. 2024. Kamcot: Knowledge augmented multimodal chain-of-thoughts reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), pages 18798–18806.
  28. 28.Hatem Mousselly-Sergieh, Teresa Botschen, Iryna Gurevych, and Stefan Roth. 2018. A multimodal translation-based approach for knowledge graph representation learning. In Proceedings of the Joint Conference on Lexical and Computational Semantics, pages 225–234.
  29. 29.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning (ICML), pages 8748–8763.
  30. 30.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551.
  31. 31.Joshua Robinson, Christopher Michael Rytting, and David Wingate. 2023. Leveraging large language models for multiple choice question answering. In Proceedings of the International Conference on Learning Representations (ICLR).
  32. 32.Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018. Object hallucination in image captioning. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).
  33. 33.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. 2008. The graph neural network model. IEEE transactions on neural networks, 20(1):61–80.
  34. 34.Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 815–823.
  35. 35.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems (NeurIPS), 35:25278–25294.
  36. 36.Priyanka Sen, Sandeep Mavadia, and Amir Saffari. 2023. Knowledge graph-augmented language models for complex question answering. In The Workshop on Natural Language Reasoning and Structured Explanations (NLRSE), pages 1–8.
  37. 37.Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022. FLAVA: A foundational language and vision alignment model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15617–15629.
  38. 38.Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. 2023. Pandagpt: One model to instruction-follow them all. arXiv preprint arXiv:2305.16355.
  39. 39.Yixuan Su, Tian Lan, Yahui Liu, Fangyu Liu, Dani Yogatama, Yan Wang, Lingpeng Kong, and Nigel Collier. 2022. Language models can see: Plugging visual controls in text generation. arXiv preprint arXiv:2205.02655.
  40. 40.Fabian M Suchanek, Gjergji Kasneci, and Gerhard Weikum. 2007. Yago: a core of semantic knowledge. In Proceedings of the International Conference on World Wide Web (WWW), pages 697–706.
  41. 41.Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Heung-Yeung Shum, and Jian Guo. 2024. Think-on-graph: Deep and responsible reasoning of large language model with knowledge graph. In Proceedings of the International Conference on Learning Representations (ICLR).
  42. 42.Rui Sun, Xuezhi Cao, Yan Zhao, Junchen Wan, Kun Zhou, Fuzheng Zhang, Zhongyuan Wang, and Kai Zheng. 2020. Multi-modal knowledge graphs for recommender systems. In Proceedings of the ACM International Conference on Information & Knowledge Management (CIKM), pages 1405–1414.
  43. 43.Yi Tay, Mostafa Dehghani, Vinh Q Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Steven Zheng, et al. 2022. Ul2: Unifying language learning paradigms. In Proceedings of the International Conference on Learning Representations (ICLR).
  44. 44.Yijun Tian, Huan Song, Zichen Wang, Haozhu Wang, Ziqing Hu, Fang Wang, Nitesh V Chawla, and Panpan Xu. 2023. Graph neural prompting with large language models. Advances in Neural Information Processing Systems (NeurIPS).
  45. 45.Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.
  46. 46.Petar Velickovi ˇ c, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903.
  47. 47.Meng Wang, Sen Wang, Han Yang, Zheng Zhang, Xi Chen, and Guilin Qi. 2021. Is visual context really helpful for knowledge graph? A representation learning perspective. In Proceedings of the ACM International Conference on Multimedia (MM), pages 2735–2743.
  48. 48.Zikang Wang, Linjing Li, Qiudan Li, and Daniel Zeng. 2019. Multimodal data enhanced representation learning for knowledge graphs. In Proceedings of the International Joint Conference on Neural Networks (IJCNN), pages 1–8.
  49. 49.Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan. 2023a. Visual chatgpt: Talking, drawing and editing with visual foundation models. arXiv preprint arXiv:2303.04671.
  50. 50.Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. 2023b. Next-gpt: Any-to-any multimodal llm. arXiv preprint arXiv:2309.05519.
  51. 51.Yike Wu, Nan Hu, Guilin Qi, Sheng Bi, Jie Ren, Anhuan Xie, and Wei Song. 2023c. Retrieve-rewrite-answer: A kg-to-text enhanced llms framework for knowledge graph question answering. arXiv preprint arXiv:2309.11206.
  52. 52.Ruobing Xie, Zhiyuan Liu, Huanbo Luan, and Maosong Sun. 2017. Image-embodied knowledge representation learning. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), pages 3140–3146.
  53. 53.Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher D Manning, Percy S Liang, and Jure Leskovec. 2022. Deep bidirectional language-knowledge graph pretraining. In Advances in neural information processing systems (NeurIPS).
  54. 54.Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023a. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. arXiv preprint arXiv:2305.11000.
  55. 55.Hang Zhang, Xin Li, and Lidong Bing. 2023b. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858.
  56. 56.Ningyu Zhang, Lei Li, Xiang Chen, Xiaozhuan Liang, Shumin Deng, and Huajun Chen. 2022. Multimodal analogical reasoning over knowledge graphs. Proceedings of the International Conference on Learning Representations (ICLR).
  57. 57.Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao. 2024a. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. Proceedings of the International Conference on Learning Representations (ICLR).
  58. 58.Tianyi Zhang, Faisal Ladhak, Esin Durmus, Percy Liang, Kathleen McKeown, and Tatsunori B Hashimoto. 2024b. Benchmarking large language models for news summarization. Transactions of the Association for Computational Linguistics, 12:39–57.
  59. 59.Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023c. Multimodal chain-of-thought reasoning in language models. arXiv preprint arXiv:2302.00923.
  60. 60.Wentian Zhao and Xinxiao Wu. 2023. Boosting entity-aware image captioning with multi-modal knowledge graph. IEEE Transactions on Multimedia.
  61. 61.Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592.

Citation

MLA
Lee, J., et al. “Multimodal Reasoning with Multimodal Knowledge Graph”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 10767–82, https://doi.org/10.18653/v1/2024.acl-long.579.
APA
Lee, J., Wang, Y., Li, J., & Zhang, M. (2024). Multimodal Reasoning with Multimodal Knowledge Graph. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10767–10782. https://doi.org/10.18653/v1/2024.acl-long.579
Chicago
Lee, J., Y. Wang, J. Li, and M. Zhang. 2024. “Multimodal Reasoning with Multimodal Knowledge Graph”. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 10767–82. https://doi.org/10.18653/v1/2024.acl-long.579.
Harvard
Lee, J. et al. (2024) “Multimodal Reasoning with Multimodal Knowledge Graph”, Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 10767–10782. Available at: https://doi.org/10.18653/v1/2024.acl-long.579.
Vancouver
1. Lee J, Wang Y, Li J, Zhang M (2024) Multimodal Reasoning with Multimodal Knowledge Graph. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 10767–10782

BibTeX

@inproceedings{lee-etal-2024-multimodal,
    title = "Multimodal Reasoning with Multimodal Knowledge Graph",
    author = "Lee, Junlin  and
      Wang, Yequan  and
      Li, Jing  and
      Zhang, Min",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.579/",
    doi = "10.18653/v1/2024.acl-long.579",
    pages = "10767--10782"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/