Built independently by an author, for readers. Read the story and support ChapterPal

keyword

relevant label embeddings

Relevant label embeddings are high-dimensional vector representations that correspond specifically to the target or ground-truth categories associated with a given data instance within a shared semantic feature space. Typically derived from pre-trained language models, vision-language encoders, or textual descriptions of class names, these vectors capture the semantic meanings of labels as well as the relationships between different categories. In machine learning frameworks such as multi-label classification and zero-shot learning, algorithms align the feature representations of input data with their corresponding relevant label embeddings while maximizing distance from non-relevant ones. This geometric alignment facilitates cross-modal knowledge transfer, enabling systems to accurately identify multiple co-occurring concepts and generalize predictions to unseen classes.

1 item

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Sunan He, Taian Guo, Tao Dai, Ruizhi Qiao, Xiujun Shu, Bo Ren, Shu-Tao Xia

OrganizationsPeng Cheng LaboratoryShenzhen UniversityTencentTsinghua University

Why you should read this

Proposes a multi-modal knowledge transfer framework that adapts vision-language pre-trained models via knowledge distillation, prompt tuning, and a two-stream feature extractor to recognize unseen object labels in multi-label classification.

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge from a language model, while ignoring the rich semantic information inherent in image-text pairs. Instead, recently developed open-vocabulary (OV) based methods succeed in exploiting such information of image-text pairs in object detection, and achieve impressive performance. Inspired by the success of OV-based methods, we propose a novel open-vocabulary framework, named multi-modal knowledge transfer (MKT), for multi-label classification. Specifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model. To facilitate transferring the image-text matching ability of VLP model, knowledge distillation is employed to guarantee the consistency of image and label embeddings, along with prompt tuning to further update the label embeddings. To further enable the recognition of multiple objects, a simple but effective two-stream module is developed to capture both local and global features. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on public benchmark datasets.

Added

2026-09-26