Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling
Dat HuynhJason KuenZhe LinJiuxiang GuEhsan Elhamifar
Proposes a cross-modal pseudo-labeling framework that aligns caption words with visual mask features and filters label noise to segment novel object classes without mask annotations.
Modern computer vision applications such as autonomous navigation, surveillance, and medical imaging depend heavily on instance segmentation, which requires identifying and outlining individual objects at the pixel level. However, conventional systems require massive quantities of human-annotated pixel masks for every target category, making expansion to thousands of novel concepts prohibitively costly and labor-intensive. While low-cost captioned images provide a broader vocabulary, existing vision-language pretraining approaches only capture high-level visual features and fail to teach models how to generate precise, pixel-wise object boundaries for novel classes.
The article develops and evaluates a cross-modal pseudo-labeling framework called XPM. The objective is to enable open-vocabulary instance segmentation—identifying and segmenting previously unseen object categories without requiring human-drawn mask annotations—by leveraging paired image and caption data alongside an automated noise-filtering mechanism.
The evaluated approach uses a teacher-student model architecture. First, a teacher model trained on base annotated classes identifies image regions that match word semantics in captions and segments them into candidate "pseudo masks." Second, a student model is trained on these generated masks while simultaneously estimating their visual noise levels. This noise estimation downweights low-quality, erroneous teacher predictions during training to prevent error propagation. The methodology was validated across standard benchmarks (MS-COCO) and scaled using extensive datasets (Open Images and Conceptual Captions, spanning millions of annotated instances and captioned images).
The findings demonstrate substantial performance gains over existing state-of-the-art methods. In the generalized instance segmentation setting, where the system must identify both base and novel classes simultaneously, the proposed framework improved segmentation accuracy (mean Average Precision) by 4.5 percentage points on MS-COCO and 5.1 percentage points on the large-scale Open Images benchmark. For object detection tasks with mask supervision, novel class performance improved by 9.2 percentage points over previous baselines. The ablation experiments confirmed that combining text-guided region selection with explicit pseudo-mask noise estimation was critical to preventing performance degradation from noisy web captions.
These results indicate that automated cross-modal pseudo-labeling can substitute for expensive manual pixel-level annotation when scaling visual recognition systems. This provides a direct path to lowering data labeling costs, accelerating deployment timelines, and expanding coverage to rare or long-tail categories that are traditionally absent from training sets. The framework also demonstrates that student models can systematically outperform their teacher models when supplied with reliable cross-modal filtering.
Organizations developing large-scale visual recognition pipelines should adopt cross-modal alignment and noise-aware self-training strategies to expand category coverage without incurring proportional annotation costs. When scaling to noisy web captions, practitioners must implement adaptive loss reweighting rather than naive self-training to avoid compounding prediction errors. The primary boundary condition is that the system still requires an initial set of diverse base classes with ground-truth masks; performance may degrade if base categories are overly restricted or homogeneous. Confidence in the reported results is high across medium and large public benchmarks, though teams should conduct preliminary domain-specific pilots when applying the framework to specialized imagery with limited initial vocabulary coverage.
- Paper: Learning Transferable Visual Models From Natural Language Supervision, Alec Radford et al.. This foundational work introduces contrastive image-text pre-training (CLIP) that establishes the open-vocabulary visual-language semantic representations leveraged by cross-modal segmentation frameworks.
- Paper: Open-vocabulary Object Detection via Vision and Language Knowledge Distillation, Xiuye Gu et al. (2021). ViLD establishes the foundational paradigm of distilling pre-trained vision-language representations into region-level detectors for open-vocabulary visual recognition.
- Paper: Mask R-CNN, Kaiming He et al. (2017). Mask R-CNN provides the standard architecture and benchmarking framework for instance mask prediction upon which modern open-vocabulary instance segmentation systems build.
- Paper: Self-Training With Noisy Student Improves ImageNet Classification, Qizhe Xie et al. (2019). This paper establishes the teacher-student self-training paradigm with noisy student updates that underlies the pseudo-label generation and noise-filtering mechanisms in XPM.
- Paper: Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image Captioning, Piyush Sharma et al. (2018). Conceptual Captions provides the foundational dataset and methodology for web-scale image-alt-text pairing utilized in cross-modal pseudo-labeling.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). MaskFormer formulates segmentation as a universal mask classification problem, simplifying the generation and assignment of instance pseudo-masks.
- Paper: Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation, Golnaz Ghiasi et al. (2020). Copy-Paste data augmentation supplies essential instance-level data-scaling techniques for training robust instance segmentation models under limited supervision.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). PACL extends open-vocabulary segmentation by directly aligning local image patches with text representations from paired data without requiring explicit pseudo-mask generation.
- Paper: CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor, Shuyang Sun et al. (2024). CLIP as RNN advances open-vocabulary segmentation by developing a training-free recurrent alignment strategy that bypasses the need for pseudo-labeling pipelines entirely.
- Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything scales promptable segmentation foundation models to unprecedented data scales, offering an alternative prompt-driven paradigm for open-world mask generation.
- Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). This survey provides a comprehensive synthesis of vision-language foundation models across downstream visual recognition tasks, contextualizing cross-modal pseudo-labeling within broader transfer strategies.
- Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). Generative Semantic Segmentation explores a generative reformulation of pixel-level prediction to overcome domain transfer limitations faced by discriminative segmentation frameworks.
- Paper: Token Contrast for Weakly-Supervised Semantic Segmentation, Lixiang Ru et al. (2023). Token Contrast tackles the internal representation collapse and over-smoothing issues of vision transformers when generating segmentation masks from weak supervision.
