keyword
partial labels
Partial labels refer to an incomplete annotation setting in multi-label classification where each training instance is labeled for only a subset of all possible target categories, leaving the presence or absence of the remaining categories unannotated or unknown. Unlike fully supervised learning, which requires exhaustive ground-truth verification across every candidate class for every sample, learning with partial labels reduces the cost and effort of dataset curation. The objective under this setting is to train a model capable of accurately predicting the complete set of relevant labels for unseen data, commonly achieved by leveraging label correlations, generating pseudo-labels for unobserved categories, or transferring semantic representations from known categories to unknown ones.
3 items

Texts as Images in Prompt Tuning for Multi-Label Image Recognition
Zixian Guo, Bowen Dong, Zhilong Ji, Jinfeng Bai, Yiwen Guo, Wangmeng Zuo
Why you should read this
Proposes a prompt tuning framework that trains vision-language models for multi-label image recognition using easily accessible text descriptions instead of labeled images, achieving strong classification performance without visual training data.
Prompt tuning has been employed as an efficient way to adapt large vision-language pre-trained models (e.g. CLIP) to various downstream tasks in data-limited or label-limited settings. Nonetheless, visual data (e.g., images) is by default prerequisite for learning prompts in existing methods. In this work, we advocate that the effectiveness of image-text contrastive learning in aligning the two modalities (for training CLIP) further makes it feasible to treat texts as images for prompt tuning and introduce TaI prompting. In contrast to the visual data, text descriptions are easy to collect, and their class labels can be directly derived. Particularly, we apply TaI prompting to multi-label image recognition, where sentences in the wild serve as alternatives to images for prompt tuning. Moreover, with TaI, dual-grained prompt tuning (TaI-DPT) is further presented to extract both coarse-grained and fine-grained embeddings for enhancing the multi-label recognition performance. Experimental results show that our proposed TaI-DPT outperforms zero-shot CLIP by a large margin on multiple benchmarks, e.g., MS-COCO, VOC2007, and NUS-WIDE, while it can be combined with existing methods of prompting from images to improve recognition performance further. The code is released at https://github.com/guozix/TaI-DPT.
Added
2026-09-26

Semantic-Aware Representation Blending for Multi-Label Image Recognition with Partial Labels
Tao Pu, Tianshui Chen, Hefeng Wu, Liang Lin
Why you should read this
Proposes a semantic-aware representation blending framework that transfers category-specific features across images at both instance and prototype levels to effectively complement missing annotations in multi-label image recognition without requiring pre-trained pseudo-labeling models.
Training the multi-label image recognition models with partial labels, in which merely some labels are known while others are unknown for each image, is a considerably challenging and practical task. To address this task, current algorithms mainly depend on pre-training classification or similarity models to generate pseudo labels for the unknown labels. However, these algorithms depend on sufficient multi-label annotations to train the models, leading to poor performance especially with low known label proportion. In this work, we propose to blend category-specific representation across different images to transfer information of known labels to complement unknown labels, which can get rid of pre-training models and thus does not depend on sufficient annotations. To this end, we design a unified semantic-aware representation blending (SARB) framework that exploits instance-level and prototype-level semantic representation to complement unknown labels by two complementary modules: 1) an instance-level representation blending (ILRB) module blends the representations of the known labels in an image to the representations of the unknown labels in another image to complement these unknown labels. 2) a prototype-level representation blending (PLRB) module learns more stable representation prototypes for each category and blends the representation of unknown labels with the prototypes of corresponding labels to complement these labels. Extensive experiments on the MS-COCO, Visual Genome, Pascal VOC 2007 datasets show that the proposed SARB framework obtains superior performance over current leading competitors on all known label proportion settings, i.e., with the mAP improvement of 4.6%, 4.6%, 2.2% on these three datasets when the known label proportion is 10%. Codes are available at https://github.com/HCPLab-SYSU/HCP-MLR-PL.
Added
2026-09-26

DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations
Ximeng Sun, Ping Hu, Kate Saenko
Why you should read this
Introduces DualCoOp, a lightweight prompt-learning framework for pretrained vision-language models that optimizes pairs of positive and negative context prompts to efficiently adapt CLIP for both partial-label and zero-shot multi-label image recognition.
Solving multi-label recognition (MLR) for images in the low-label regime is a challenging task that has many real-world applications. Recent work learns an alignment between textual and visual spaces to compensate for insufficient image labels, but loses accuracy because of the limited amount of available MLR annotations. In this work, we utilize the strong alignment of textual and visual features pretrained with millions of auxiliary image-text pairs and propose Dual Context Optimization (DualCoOp) as a unified framework for partial-label MLR and zero-shot MLR. DualCoOp encodes positive and negative contexts with class names as part of the linguistic input (i.e. prompts). Since DualCoOp only introduces a very light learnable overhead upon the pretrained vision-language framework, it can quickly adapt to multi-label recognition tasks that have limited annotations and even unseen classes. Experiments on standard multi-label recognition benchmarks across two challenging low-label settings demonstrate the advantages of our approach over state-of-the-art methods. Project page: https://cs-people.bu.edu/sunxm/DualCoOp/project.html
Added
2026-09-26
