Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-label zero-shot learning

Multi-label zero-shot learning is a machine learning paradigm where a model learns to identify and assign multiple relevant category labels to a single data instance, such as an image or text document, even when some or all of those labels were never present in the training data. Unlike conventional zero-shot learning, which typically assumes each instance belongs to only one unseen class, this task addresses complex scenarios containing multiple co-occurring objects or concepts. To recognize novel categories without direct supervision, models map input features and auxiliary semantic information, such as word embeddings, class attributes, or vision-language representations, into a shared embedding space. This alignment enables the system to transfer knowledge from seen to unseen classes while modeling relationships, dependencies, and feature alignments across multiple labels simultaneously.

1 item

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Sunan He, Taian Guo, Tao Dai, Ruizhi Qiao, Xiujun Shu, Bo Ren, Shu-Tao Xia

OrganizationsPeng Cheng LaboratoryShenzhen UniversityTencentTsinghua University

Why you should read this

Proposes a multi-modal knowledge transfer framework that adapts vision-language pre-trained models via knowledge distillation, prompt tuning, and a two-stream feature extractor to recognize unseen object labels in multi-label classification.

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge from a language model, while ignoring the rich semantic information inherent in image-text pairs. Instead, recently developed open-vocabulary (OV) based methods succeed in exploiting such information of image-text pairs in object detection, and achieve impressive performance. Inspired by the success of OV-based methods, we propose a novel open-vocabulary framework, named multi-modal knowledge transfer (MKT), for multi-label classification. Specifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model. To facilitate transferring the image-text matching ability of VLP model, knowledge distillation is employed to guarantee the consistency of image and label embeddings, along with prompt tuning to further update the label embeddings. To further enable the recognition of multiple objects, a simple but effective two-stream module is developed to capture both local and global features. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on public benchmark datasets.

Added

2026-09-26