Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-modal knowledge transfer

Multi-modal knowledge transfer is a machine learning process in which representations, semantic alignments, and contextual relationships acquired across multiple data modalities, such as vision and text, are transferred to improve the performance of a model on target tasks. Unlike unimodal transfer methods that rely on information from a single source, such as isolated textual embeddings or visual features, multi-modal knowledge transfer leverages joint embeddings and cross-modal correspondences typically learned by vision-language pre-training models. By utilizing techniques such as knowledge distillation and feature alignment to convey these rich multi-modal associations, the process enables systems to better bridge semantic gaps across different modalities, recognize novel or unseen categories in open-vocabulary scenarios, and enhance generalizability across diverse recognition tasks.

1 item

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Sunan He, Taian Guo, Tao Dai, Ruizhi Qiao, Xiujun Shu, Bo Ren, Shu-Tao Xia

OrganizationsPeng Cheng LaboratoryShenzhen UniversityTencentTsinghua University

Why you should read this

Proposes a multi-modal knowledge transfer framework that adapts vision-language pre-trained models via knowledge distillation, prompt tuning, and a two-stream feature extractor to recognize unseen object labels in multi-label classification.

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge from a language model, while ignoring the rich semantic information inherent in image-text pairs. Instead, recently developed open-vocabulary (OV) based methods succeed in exploiting such information of image-text pairs in object detection, and achieve impressive performance. Inspired by the success of OV-based methods, we propose a novel open-vocabulary framework, named multi-modal knowledge transfer (MKT), for multi-label classification. Specifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model. To facilitate transferring the image-text matching ability of VLP model, knowledge distillation is employed to guarantee the consistency of image and label embeddings, along with prompt tuning to further update the label embeddings. To further enable the recognition of multiple objects, a simple but effective two-stream module is developed to capture both local and global features. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on public benchmark datasets.

Added

2026-09-26