Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Open Images dataset

The Open Images dataset is a large-scale, publicly accessible collection of approximately nine million images designed to train and evaluate computer vision and machine learning models. Developed by Google in collaboration with academic institutions, the dataset features diverse, real-world scenes annotated across thousands of semantic categories. It provides multiple levels of annotations, including image-level classification labels, object bounding boxes, instance segmentation masks, visual relationships, localized narratives, and point labels. Due to its extensive scale and granular annotations, the dataset is widely used as a benchmark for core visual recognition tasks such as multi-label classification, object detection, instance segmentation, and vision-language representation learning.

2 items

Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling

Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling

Dat Huynh, Jason Kuen, Zhe Lin, Jiuxiang Gu, Ehsan Elhamifar

OrganizationsAdobeNortheastern University

Why you should read this

Proposes a cross-modal pseudo-labeling framework that aligns caption words with visual mask features and filters label noise to segment novel object classes without mask annotations.

Open-vocabulary instance segmentation aims at segmenting novel classes without mask annotations. It is an important step toward reducing laborious human supervision. Most existing works first pretrain a model on captioned images covering many novel classes and then finetune it on limited base classes with mask annotations. However, the high-level textual information learned from caption pre-training alone cannot effectively encode the details required for pixel-wise segmentation. To address this, we propose a cross-modal pseudo-labeling framework, which generates training pseudo masks by aligning word semantics in captions with visual features of object masks in images. Thus, our framework is capable of labeling novel classes in captions via their word semantics to self-train a student model. To account for noises in pseudo masks, we design a robust student model that selectively distills mask knowledge by estimating the mask noise levels, hence mitigating the adverse impact of noisy pseudo masks. By extensive experiments, we show the effectiveness of our framework, where we significantly improve mAP score by 4.5% on MS-COCO and 5.1% on the large-scale Open Images & Conceptual Captions datasets compared to the state-of-the-art.

Added

2026-09-26

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

Sunan He, Taian Guo, Tao Dai, Ruizhi Qiao, Xiujun Shu, Bo Ren, Shu-Tao Xia

OrganizationsPeng Cheng LaboratoryShenzhen UniversityTencentTsinghua University

Why you should read this

Proposes a multi-modal knowledge transfer framework that adapts vision-language pre-trained models via knowledge distillation, prompt tuning, and a two-stream feature extractor to recognize unseen object labels in multi-label classification.

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge from a language model, while ignoring the rich semantic information inherent in image-text pairs. Instead, recently developed open-vocabulary (OV) based methods succeed in exploiting such information of image-text pairs in object detection, and achieve impressive performance. Inspired by the success of OV-based methods, we propose a novel open-vocabulary framework, named multi-modal knowledge transfer (MKT), for multi-label classification. Specifically, our method exploits multi-modal knowledge of image-text pairs based on a vision and language pre-training (VLP) model. To facilitate transferring the image-text matching ability of VLP model, knowledge distillation is employed to guarantee the consistency of image and label embeddings, along with prompt tuning to further update the label embeddings. To further enable the recognition of multiple objects, a simple but effective two-stream module is developed to capture both local and global features. Extensive experimental results show that our method significantly outperforms state-of-the-art methods on public benchmark datasets.

Added

2026-09-26