Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling

Dat HuynhJason KuenZhe LinJiuxiang GuEhsan Elhamifar

article2022CVPR117 citations

Proposes a cross-modal pseudo-labeling framework that aligns caption words with visual mask features and filters label noise to segment novel object classes without mask annotations.

Listen

Modern computer vision applications such as autonomous navigation, surveillance, and medical imaging depend heavily on instance segmentation, which requires identifying and outlining individual objects at the pixel level. However, conventional systems require massive quantities of human-annotated pixel masks for every target category, making expansion to thousands of novel concepts prohibitively costly and labor-intensive. While low-cost captioned images provide a broader vocabulary, existing vision-language pretraining approaches only capture high-level visual features and fail to teach models how to generate precise, pixel-wise object boundaries for novel classes.

The article develops and evaluates a cross-modal pseudo-labeling framework called XPM. The objective is to enable open-vocabulary instance segmentation—identifying and segmenting previously unseen object categories without requiring human-drawn mask annotations—by leveraging paired image and caption data alongside an automated noise-filtering mechanism.

The evaluated approach uses a teacher-student model architecture. First, a teacher model trained on base annotated classes identifies image regions that match word semantics in captions and segments them into candidate "pseudo masks." Second, a student model is trained on these generated masks while simultaneously estimating their visual noise levels. This noise estimation downweights low-quality, erroneous teacher predictions during training to prevent error propagation. The methodology was validated across standard benchmarks (MS-COCO) and scaled using extensive datasets (Open Images and Conceptual Captions, spanning millions of annotated instances and captioned images).

The findings demonstrate substantial performance gains over existing state-of-the-art methods. In the generalized instance segmentation setting, where the system must identify both base and novel classes simultaneously, the proposed framework improved segmentation accuracy (mean Average Precision) by 4.5 percentage points on MS-COCO and 5.1 percentage points on the large-scale Open Images benchmark. For object detection tasks with mask supervision, novel class performance improved by 9.2 percentage points over previous baselines. The ablation experiments confirmed that combining text-guided region selection with explicit pseudo-mask noise estimation was critical to preventing performance degradation from noisy web captions.

These results indicate that automated cross-modal pseudo-labeling can substitute for expensive manual pixel-level annotation when scaling visual recognition systems. This provides a direct path to lowering data labeling costs, accelerating deployment timelines, and expanding coverage to rare or long-tail categories that are traditionally absent from training sets. The framework also demonstrates that student models can systematically outperform their teacher models when supplied with reliable cross-modal filtering.

Organizations developing large-scale visual recognition pipelines should adopt cross-modal alignment and noise-aware self-training strategies to expand category coverage without incurring proportional annotation costs. When scaling to noisy web captions, practitioners must implement adaptive loss reweighting rather than naive self-training to avoid compounding prediction errors. The primary boundary condition is that the system still requires an initial set of diverse base classes with ground-truth masks; performance may degrade if base categories are overly restricted or homogeneous. Confidence in the reported results is high across medium and large public benchmarks, though teams should conduct preliminary domain-specific pilots when applying the framework to specialized imagery with limited initial vocabulary coverage.

  • Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). PACL extends open-vocabulary segmentation by directly aligning local image patches with text representations from paired data without requiring explicit pseudo-mask generation.
  • Paper: CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor, Shuyang Sun et al. (2024). CLIP as RNN advances open-vocabulary segmentation by developing a training-free recurrent alignment strategy that bypasses the need for pseudo-labeling pipelines entirely.
  • Paper: Segment Anything, Alexander M. Kirillov et al. (2023). Segment Anything scales promptable segmentation foundation models to unprecedented data scales, offering an alternative prompt-driven paradigm for open-world mask generation.
  • Paper: Vision-Language Models for Vision Tasks: A Survey, Jingyi Zhang et al. (2023). This survey provides a comprehensive synthesis of vision-language foundation models across downstream visual recognition tasks, contextualizing cross-modal pseudo-labeling within broader transfer strategies.
  • Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). Generative Semantic Segmentation explores a generative reformulation of pixel-level prediction to overcome domain transfer limitations faced by discriminative segmentation frameworks.
  • Paper: Token Contrast for Weakly-Supervised Semantic Segmentation, Lixiang Ru et al. (2023). Token Contrast tackles the internal representation collapse and over-smoothing issues of vision transformers when generating segmentation masks from weak supervision.
Cover for Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling

Abstract

Open-vocabulary instance segmentation aims at segmenting novel classes without mask annotations. It is an important step toward reducing laborious human supervision. Most existing works first pretrain a model on captioned images covering many novel classes and then finetune it on limited base classes with mask annotations. However, the high-level textual information learned from caption pre-training alone cannot effectively encode the details required for pixel-wise segmentation. To address this, we propose a cross-modal pseudo-labeling framework, which generates training pseudo masks by aligning word semantics in captions with visual features of object masks in images. Thus, our framework is capable of labeling novel classes in captions via their word semantics to self-train a student model. To account for noises in pseudo masks, we design a robust student model that selectively distills mask knowledge by estimating the mask noise levels, hence mitigating the adverse impact of noisy pseudo masks. By extensive experiments, we show the effectiveness of our framework, where we significantly improve mAP score by 4.5% on MS-COCO and 5.1% on the large-scale Open Images & Conceptual Captions datasets compared to the state-of-the-art.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Robust Cross-Modal Pseudo-Labeling on Captioned Images
  • 3.1. Problem Setting
  • 3.2. Proposed Method
  • 3.2.1 Designing Teacher Model
  • 3.2.2 Cross-Modal Pseudo-Labeling
  • 3.2.3 Estimating Pseudo-Mask Noises
  • 3.2.4 Training Robust Student Model
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Experimental Results
  • 5. Conclusions
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Problem Formulation for Open-Vocabulary Instance Segmentation with Captions

    definition

    In open-vocabulary instance segmentation, a model is trained to detect and segment novel object classes that lack instance mask annotations during training, using ground-truth mask annotations on a restricted set of base classes together with captioned images covering a broader vocabulary.

    The training data consists of two sets:

    1. A base dataset DB={(Im,Ym)}m=1NB\mathcal{D}_B = \{(I_m, Y_m)\}_{m=1}^{N_B}, where each image ImI_m is provided with ground-truth instance annotations YmY_m consisting of bounding boxes, instance masks, and class labels from a set of base classes VB\mathcal{V}_B.
    2. A captioned dataset DC={(Ic,Yc)}c=1NC\mathcal{D}_C = \{(I_c, Y_c)\}_{c=1}^{N_C}, where each image IcI_c is paired only with an image-level caption YcY_c. From each caption YcY_c, a set of object nouns Oc⊂YcO_c \subset Y_c is extracted belonging to a large caption vocabulary VC\mathcal{V}_C, where ∣VC∣≫∣VB∣|\mathcal{V}_C| \gg |\mathcal{V}_B|.

    Evaluation is performed on target classes VT\mathcal{V}_T that have no instance annotations during training (VT∩VB=∅\mathcal{V}_T \cap \mathcal{V}_B = \emptyset). Class recognition across all sets o∈VB∪VC∪VTo \in \mathcal{V}_B \cup \mathcal{V}_C \cup \mathcal{V}_T is enabled by projecting text labels into semantic word embedding vectors vo∈Rdv_o \in \mathbb{R}^d using a pretrained language model (e.g., BERT).

  2. Knowl 2 — Teacher Model with Joint Visual-Semantic Embedding and Class-Agnostic Mask Head

    model/method

    The teacher model hh builds upon a two-stage Mask R-CNN architecture to detect and segment novel classes based on text embeddings. It consists of three main modules:

    1. Class-Agnostic Region Proposal Network (RPN): Proposes candidate object bounding regions p(I)={ri}i=1NRp(I) = \{r_i\}_{i=1}^{N_R} for an input image II.
    2. Visual-Semantic Embedding Head (hEmbh_{\text{Emb}}): Replaces the standard discrete classification layer. For each region r∈p(I)r \in p(I) with visual RoIAlign feature frf_r, hEmbh_{\text{Emb}} maps frf_r into the word embedding space. The compatibility score for class oo with embedding vector vov_o is given by the inner product:
    s(o,r)=vo⊤hEmb(fr)s(o, r) = v_o^\top h_{\text{Emb}}(f_r)

    To reject non-object proposals, the background embedding is fixed to the zero vector 0\mathbf{0}, meaning a proposal is classified as background if s(o,r)<0s(o, r) < 0 for all candidate classes. 3. Class-Agnostic Mask Head (hMaskh_{\text{Mask}}): Predicts pixel-wise binary segmentation logit scores hMask(fr)h_{\text{Mask}}(f_r) for any region proposal rr.

    The teacher is trained on the base dataset DB\mathcal{D}_B using standard supervised detection and instance segmentation losses LGT\mathcal{L}_{GT}.

  3. Knowl 3 — Cross-Modal Bounding Box Grounding and Classification Alignment

    model/method

    To generate pseudo-supervision from captioned images DC\mathcal{D}_C, object nouns Oc⊂YcO_c \subset Y_c are extracted from caption YcY_c (identified as synset descendants of the 'Object' node in the WordNet hierarchy). For each object noun o∈Oco \in O_c with semantic embedding vov_o, the teacher model selects the most compatible region proposal bob_o from p(Ic)p(I_c) via cross-modal visual-semantic alignment:

    bo=arg⁡max⁡r∈p(Ic)vo⊤hEmb(fr)b_o = \arg\max_{r \in p(I_c)} v_o^\top h_{\text{Emb}}(f_r)

    Selecting only the highest-confidence bounding region per caption noun minimizes false-positive region selections.

    A student network gg with embedding head gEmbg_{\text{Emb}} is trained to identify each aligned region bob_o with its matching caption noun oo via the cross-modal classification loss:

    LX(Yc∣Ic;g)=−∑o∈Oclog⁡exp⁡(vo⊤gEmb(fbo))∑w∈VCexp⁡(vw⊤gEmb(fbo))\mathcal{L}_X(Y_c \mid I_c; g) = - \sum_{o \in O_c} \log \frac{\exp(v_o^\top g_{\text{Emb}}(f_{b_o}))}{\sum_{w \in \mathcal{V}_C} \exp(v_w^\top g_{\text{Emb}}(f_{b_o}))}

    where VC\mathcal{V}_C is the entire caption vocabulary.

  4. Knowl 4 — Heteroscedastic Noise-Aware Mask Loss for Pseudo-Mask Distillation

    model/method

    Given aligned region proposals bob_o for each caption object o∈Oco \in O_c, the teacher mask head hMaskh_{\text{Mask}} generates a binary pseudo mask MoM_o:

    Mo=I[hMask(fbo)≥0]M_o = \mathbb{I}[h_{\text{Mask}}(f_{b_o}) \ge 0]

    where I[⋅]\mathbb{I}[\cdot] is the indicator function.

    Because teacher pseudo-masks on novel objects can be noisy and contain severe segmentation errors, directly enforcing binary cross-entropy loss causes error propagation. To prevent this, the student model incorporates a mask prediction head gMaskg_{\text{Mask}} alongside an auxiliary pixel-wise noise prediction network gNoiseg_{\text{Noise}}. The pseudo-mask loss LM\mathcal{L}_M models pixel-level uncertainty by corrupting student mask logits with zero-mean Gaussian noise whose variance is estimated by gNoiseg_{\text{Noise}}:

    LM(Yc∣Ic,g)=∑o∈Oc∑x,yLBCE(Moxy  |  gMaskxy(fbo)+ϵoxy),ϵoxy∼N(0,gNoisexy(fbo))\mathcal{L}_M(Y_c \mid I_c, g) = \sum_{o \in O_c} \sum_{x,y} \mathcal{L}_{BCE}\left(M_o^{xy} \;\middle|\; g_{\text{Mask}}^{xy}(f_{b_o}) + \epsilon_o^{xy}\right), \quad \epsilon_o^{xy} \sim \mathcal{N}\left(0, g_{\text{Noise}}^{xy}(f_{b_o})\right)

    where LBCE\mathcal{L}_{BCE} is the binary cross-entropy loss at pixel location (x,y)(x,y). The reparameterization trick is used to backpropagate gradients through the sampled noise ϵoxy\epsilon_o^{xy} into gNoiseg_{\text{Noise}}.

  5. Knowl 5 — Robust Pseudo-Mask Reliability Reweighting and Joint Student Objective

    equation

    The overall reliability score α(o∣Ic)\alpha(o \mid I_c) of the pseudo mask for object o∈Oco \in O_c in image IcI_c is defined as the inverse of the average predicted pixel noise over region bob_o:

    α(o∣Ic)=η1∣bo∣∑x,ygNoisexy(fbo)\alpha(o \mid I_c) = \frac{\eta}{\frac{1}{|b_o|} \sum_{x,y} g_{\text{Noise}}^{xy}(f_{b_o})}

    where ∣bo∣|b_o| is the number of pixels in region bob_o, and η\eta is a fixed constant set to the minimum average noise level across captioned training images (set empirically to η=0.01\eta = 0.01).

    The student model g={gEmb,gMask,gNoise}g = \{g_{\text{Emb}}, g_{\text{Mask}}, g_{\text{Noise}}\} is trained by minimizing the combined loss over captioned images DC\mathcal{D}_C and base ground-truth images DB\mathcal{D}_B:

    min⁡g∑c∈DC[LM(Yc∣Ic;g)+LXα(Yc∣Ic;g)]+∑m∈DBLGT(Ym∣Im;g)\min_{g} \sum_{c \in \mathcal{D}_C} \left[ \mathcal{L}_M(Y_c \mid I_c; g) + \mathcal{L}_X^\alpha(Y_c \mid I_c; g) \right] + \sum_{m \in \mathcal{D}_B} \mathcal{L}_{GT}(Y_m \mid I_m; g)

    where LXα\mathcal{L}_X^\alpha is the noise-reweighted cross-modal classification loss:

    LXα(Yc∣Ic;g)=−∑o∈Ocα(o∣Ic)log⁡exp⁡(vo⊤gEmb(fbo))∑w∈VCexp⁡(vw⊤gEmb(fbo))\mathcal{L}_X^\alpha(Y_c \mid I_c; g) = - \sum_{o \in O_c} \alpha(o \mid I_c) \log \frac{\exp(v_o^\top g_{\text{Emb}}(f_{b_o}))}{\sum_{w \in \mathcal{V}_C} \exp(v_w^\top g_{\text{Emb}}(f_{b_o}))}

    To prevent the trivial solution where gNoiseg_{\text{Noise}} predicts arbitrarily high noise to zero out LXα\mathcal{L}_X^\alpha, gradients from LXα\mathcal{L}_X^\alpha are not propagated into gNoiseg_{\text{Noise}}.

  6. Knowl 6 — Cross-Modal Pseudo-Mask (XPM) Training Algorithm

    algorithm

    The complete training pipeline of the Cross-Modal Pseudo-Mask (XPM) framework is structured as follows:

    Input: Base dataset DB\mathcal{D}_B, captioned dataset DC\mathcal{D}_C, pretrained word embeddings {vo}\{v_o\}, reference noise constant η=0.01\eta = 0.01
    Output: Trained robust student segmentation model g={gEmb,gMask,gNoise}g = \{g_{\text{Emb}}, g_{\text{Mask}}, g_{\text{Noise}}\}
    1. Train Teacher Model h={hEmb,hMask}h = \{h_{\text{Emb}}, h_{\text{Mask}}\} on base annotations DB\mathcal{D}_B using standard loss LGT\mathcal{L}_{GT}.
    2. Initialize Student Model gg weights with Teacher Model hh weights.
    3. for each training iteration do
    4. Sample base batch from DB\mathcal{D}_B and captioned batch from DC\mathcal{D}_C.
    5. for each captioned image (Ic,Yc)(I_c, Y_c) in captioned batch do
    6. Extract object nouns Oc⊂YcO_c \subset Y_c using WordNet.
    7. Generate candidate region proposals p(Ic)p(I_c) using the RPN.
    8. for each object noun o∈Oco \in O_c do
    9. Identify best region: bo←arg⁡max⁡r∈p(Ic)vo⊤hEmb(fr)b_o \leftarrow \arg\max_{r \in p(I_c)} v_o^\top h_{\text{Emb}}(f_r)
    10. Generate binary pseudo mask: Mo←I[hMask(fbo)≥0]M_o \leftarrow \mathbb{I}[h_{\text{Mask}}(f_{b_o}) \ge 0]
    11. Compute mask noise variance gNoise(fbo)g_{\text{Noise}}(f_{b_o}) and sample ϵo∼N(0,gNoise(fbo))\epsilon_o \sim \mathcal{N}(0, g_{\text{Noise}}(f_{b_o}))
    12. Compute mask reliability weight: α(o∣Ic)←η1∣bo∣∑x,ygNoisexy(fbo)\alpha(o \mid I_c) \leftarrow \frac{\eta}{\frac{1}{|b_o|} \sum_{x,y} g_{\text{Noise}}^{xy}(f_{b_o})}
    13. end for
    14. Compute mask loss LM(Yc∣Ic;g)\mathcal{L}_M(Y_c \mid I_c; g) and reweighted classification loss LXα(Yc∣Ic;g)\mathcal{L}_X^\alpha(Y_c \mid I_c; g).
    15. end for
    16. Compute ground-truth loss LGT(Ym∣Im;g)\mathcal{L}_{GT}(Y_m \mid I_m; g) on base batch.
    17. Update student parameters gg via gradient descent on ∑c[LM+LXα]+∑mLGT\sum_{c} [\mathcal{L}_M + \mathcal{L}_X^\alpha] + \sum_{m} \mathcal{L}_{GT}, with LXα\mathcal{L}_X^\alpha gradients detached from gNoiseg_{\text{Noise}}.
    18. end for
    19. return gg
  7. Knowl 7 — Benchmark Results on Open-Vocabulary Instance Segmentation

    data/table

    Instance segmentation performance is evaluated using mean Average Precision at IoU=0.5\text{IoU} = 0.5 (mAP50\text{mAP}_{50}) across base classes, target (novel unseen) classes, and all classes combined. Results are reported in two evaluation modes: Constrained (evaluated on test images containing only base or only target classes) and Generalized (jointly evaluated across all test images containing base and target classes).

    Method MS-COCO Open Images Conceptual Captions
    Constrained Generalized Constrained Generalized
    Base Target Base Target All Base Target Base Target All
    OVR 42.0 20.9 41.6 17.1 35.2 52.6 23.8 45.6 17.5 36.2
    SB 41.6 20.8 41.0 16.0 34.5 52.8 24.8 46.4 17.3 36.6
    BA-RPN 41.8 20.1 41.3 15.4 34.5 52.9 25.3 47.3 16.9 37.1
    OVR+OMP 31.3 14.1 30.5 8.3 24.7 52.5 24.9 47.1 16.8 36.9
    Soft-Teacher 41.8 14.8 41.5 9.6 33.2 52.0 25.9 46.6 17.6 36.8
    Unbiased-Teacher 41.8 15.1 41.4 9.8 33.1 51.7 22.2 45.3 14.5 34.9
    XPM (Ours) 42.4 24.0 41.5 21.6 36.3 55.1 31.6 49.8 22.7 40.7

    On MS-COCO (48 base, 17 target classes), XPM achieves 21.6%21.6\% generalized target mAP (+4.5%+4.5\% over OVR). On Open Images (200 base, 100 target classes) combined with 3M Conceptual Captions, XPM achieves 22.7%22.7\% generalized target mAP (+5.1%+5.1\% over Soft-Teacher and +5.2%+5.2\% over OVR) and 40.7%40.7\% overall mAP (+3.6%+3.6\% gain across all classes).

  8. Knowl 8 — Benchmark Results on Open-Vocabulary Object Detection

    data/table

    Object detection performance (mAP at IoU=0.5\text{IoU} = 0.5) on MS-COCO evaluated on 48 base and 17 target classes under base supervision with bounding boxes or instance masks:

    Method Bounding Box Supervision Instance Mask Supervision
    Constrained Generalized Constrained Generalized
    Base Target Base Target All Base Target Base Target All
    Zero-Shot Training
    SB 29.7 0.7 29.2 0.3 24.9 - - - - -
    BA-RPN - 11.4 46.5 4.8 35.6 - - - - -
    Caption Pretraining
    OVR 46.8 27.5 46.0 22.8 39.9 47.2 25.9 46.7 20.7 39.9
    SB 46.9 26.9 46.3 21.2 39.7 45.9 25.7 45.3 19.6 38.6
    BA-RPN 46.8 26.0 46.2 20.7 39.5 46.0 25.0 45.5 19.3 38.7
    OVR+OMP - - - - - 34.1 16.9 33.2 10.0 27.1
    Pseudo-Labeling
    Soft-Teacher 47.4 18.8 47.1 12.4 38.0 46.6 16.0 46.2 10.4 36.8
    Unbiased-Teacher 47.5 20.5 47.2 13.8 38.4 46.6 16.8 46.1 10.8 36.9
    Cap2Det - - 20.1 20.3 20.1 - - - - -
    XPM (Ours) 46.8 29.9 46.3 27.0 41.2 47.3 33.2 46.3 29.9 42.0

    Under bounding box supervision, XPM achieves 27.0%27.0\% target mAP in the generalized setting (+4.2%+4.2\% improvement over OVR). Under mask supervision, XPM achieves 29.9%29.9\% target mAP (+9.2%+9.2\% improvement over OVR), demonstrating the strong complementary value of cross-modal pseudo-mask training.

  9. Knowl 9 — Ablation of Pseudo-Mask Noise Estimation and Reweighting Strategies

    data/table

    Comparison of different noise estimation and loss-weighting strategies on the Open Images dataset (evaluating generalized segmentation mAP at IoU=0.5\text{IoU} = 0.5 on 200 base and 100 target classes):

    Method Loss Applied On Base mAP Target mAP All mAP
    No Noise Estimation - 53.3 30.2 39.1
    Stochastic BCE LM\mathcal{L}_M 53.8 29.8 39.2
    Class Score LXα\mathcal{L}_X^\alpha 54.0 28.4 38.8
    Pixel Score LXα\mathcal{L}_X^\alpha 53.2 30.1 38.5
    DropOut Entropy LXα\mathcal{L}_X^\alpha 53.6 29.7 38.5
    Robust Student (Ours) LXα+LM\mathcal{L}_X^\alpha + \mathcal{L}_M 55.1 31.6 40.7

    Methods relying on classification confidence (Class Score), mask prediction confidence (Pixel Score), or Monte Carlo dropout uncertainty (DropOut Entropy) fail to improve novel class performance because they are trained on base classes and do not reflect pseudo-mask noise. Stochastic BCE regulates only LM\mathcal{L}_M and lacks cross-modal classification reweighting. Regulating both LXα\mathcal{L}_X^\alpha and LM\mathcal{L}_M via predicted pseudo-mask noise yields the highest performance across base (55.1%55.1\%), target (31.6%31.6\%), and all classes (40.7%40.7\%).

  10. Knowl 10 — Limitation: Dependence on Base Class Diversity

    limitation

    The cross-modal pseudo-labeling framework assumes that the base training classes are sufficiently diverse to allow the teacher model to learn generalizable visual-semantic embeddings and class-agnostic mask representations. If trained on a very small or restricted set of base classes (such as limited single-domain datasets), the teacher's ability to localize and segment novel objects in captioned images degrades, reducing the quality and quantity of distilled pseudo masks.

Coverage note — None was omitted; all key architectural components, equations, algorithmic details, primary benchmarks, noise estimation ablations, and stated limitations are included.

References

  1. 1.M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” IEEE Conference on Computer Vision and Pattern Recognition, 2016.
  2. 2.A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, T. Duerig, and V. Ferrari., “The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale,” International Journal of Computer Vision, 2016.
  3. 3.A. Gupta, P. Dollár, and R. B. Girshick, “Lvis: A dataset for large vocabulary instance segmentation,” IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  4. 4.S. W. Zamir, A. Arora, A. Gupta, S. H. Khan, G. Sun, F. S. Khan, F. Zhu, L. Shao, G. Xia, and X. Bai, “isaid: A largescale dataset for instance segmentation in aerial images,” CVPR Workshops, 2019.
  5. 5.S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, pp. 1137–1149, 2015.
  6. 6.K. He, G. Gkioxari, P. Dollár, and R. B. Girshick, “Mask rcnn,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  7. 7.S. Liu, L. Qi, H. Qin, J. Shi, and J. Jia, “Path aggregation network for instance segmentation,” IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  8. 8.Z. M. Chen, X. S. Wei, P. Wang, and Y. Guo, “Multi-label image recognition with graph convolutional networks,” IEEE Conference on Computer Vision and Pattern Recognition, vol. abs/1904.03582, 2019.
  9. 9.H. Huang, C. Wang, P. S. Yu, and C. D. Wang, “Generative dual adversarial network for generalized zero-shot learning,” IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  10. 10.A. Arnab and P. H. S. Torr, “Pixelwise instance segmentation with a dynamically instantiated network,” IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  11. 11.Z. Tian, C. Shen, and H. Chen, “Conditional convolutions for instance segmentation,” European Conference on Computer Vision, 2020.
  12. 12.A. Kirillov, Y. Wu, K. He, and R. B. Girshick, “Pointrend: Image segmentation as rendering,” IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  13. 13.G. Zhang, X. Lu, J. Tan, J. Li, Z. Zhang, Q. Li, and X. Hu, “Refinemask: Towards high-quality instance segmentation with fine-grained features,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  14. 14.C. Tang, H. Chen, X. Li, J. Li, Z. Zhang, and X. Hu, “Look closer to segment better: Boundary patch refinement for instance segmentation,” 2021.
  15. 15.C.-C. Hsu, K.-J. Hsu, C.-C. Tsai, Y.-Y. Lin, and Y.-Y. Chuang, “Weakly supervised instance segmentation using the bounding box tightness prior,” Neural Information Processing Systems, 2019.
  16. 16.J. Ahn, S. Cho, and S. Kwak, “Weakly supervised learning of instance segmentation with inter-pixel relations,” IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  17. 17.A. Arun, C. V. Jawahar, and M. P. Kumar, “Weakly supervised instance segmentation by learning annotation consistent instances,” European Conference on Computer Vision, 2020.
  18. 18.R. Hu, P. Dollár, K. He, T. Darrell, and R. B. Girshick, “Learning to segment every thing,” IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  19. 19.D. Biertimpel, S. Shkodrani, A. S. Baslamisli, and N. Baka, “Prior to segment: Foreground cues for novel objects in partially supervised instance segmentation,” IEEE International Conference on Computer Vision, 2021.
  20. 20.T. Zhou, W. Wang, S. Qi, H. Ling, and J. Shen, “Cascaded human-object interaction recognition,” IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  21. 21.Z. Tian, C. Shen, X. Wang, and H. Chen, “Boxinst: Highperformance instance segmentation with box annotations,” IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  22. 22.J. Lee, J. Yi, C. Shin, and S. Yoon, “Bbam: Bounding box attribution map for weakly supervised semantic and instance segmentation,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  23. 23.X. Wang, J. Feng, B. Hu, Q. Ding, L. Ran, X. Chen, and W. Liu, “Weakly-supervised instance segmentation via class-agnostic learning with salient images,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  24. 24.A. Bansal, K. Sikka, G. Sharma, R. Chellappa, and A. Divakaran, “Zero-shot object detection,” European Conference on Computer Vision, 2018.
  25. 25.S. Rahman, S. Khan, and N. Barnes, “Improved visualsemantic alignment for zero-shot object detection,” AAAI Conference on Artificial Intelligence, 2020.
  26. 26.P. Zhu, H. Wang, and V. Saligrama, “Don’t even look once: Synthesizing features for zero-shot detection,” IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  27. 27.Y. Zheng, J. Wu, Y. Qin, F. Zhang, and L. Cui, “Zero-shot instance segmentation,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  28. 28.A. Zareian, K. D. Rosa, D. H. Hu, and S. F. Chang, “Openvocabulary object detection using captions,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  29. 29.A. L. Bearman, O. Russakovsky, V. Ferrari, and L. Fei-Fei, “What’s the point: Semantic segmentation with point supervision,” European Conference on Computer Vision, 2016.
  30. 30.A. Khoreva, R. Benenson, J. H. Hosang, M. Hein, and B. Schiele, “Simple does it: Weakly supervised instance and semantic segmentation,” IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  31. 31.S. Lan, Z. Yu, C. B. Choy, S. Radhakrishnan, G. Liu, Y. Zhu, L. Davis, and A. Anandkumar, “Discobox: Weakly supervised instance segmentation and semantic correspondence from box supervision,” IEEE International Conference on Computer Vision, 2021.
  32. 32.W. Kuo, A. Angelova, J. Malik, and T.-Y. Lin, “Shapemask: Learning to segment novel objects by refining shape priors,” IEEE International Conference on Computer Vision, 2019.
  33. 33.Q. Fan, L. Ke, W. Pei, C.-K. Tang, and Y.-W. Tai, “Commonality-parsing network across shape and appearance for partially supervised instance segmentation,” European Conference on Computer Vision, 2020.
  34. 34.Y. Zhou, Y. Zhu, Q. Ye, Q. Qiu, and J. Jiao, “Weakly supervised instance segmentation using class peak response,” IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  35. 35.W. Ge, S. Guo, W. Huang, and M. R. Scott, “Label-penet: Sequential label propagation and enhancement networks for weakly supervised instance segmentation,” IEEE International Conference on Computer Vision, 2019.
  36. 36.P. Zhu, H. Wang, and V. Saligrama, “Generalized zero-shot recognition based on visually semantic embedding,” IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  37. 37.H. Cholakkal, G. Sun, F. S. Khan, and L. Shao, “Object counting and instance segmentation with image-level supervision,” IEEE Conference on Computer Vision and Pattern Recognition, pp. 12 389–12 397, 2019.
  38. 38.Y. Shen, L. Cao, Z. Chen, B. Zhang, C. Su, Y. Wu, F. Huang, and R. Ji, “Parallel detection-and-segmentation learning for weakly supervised instance segmentation,” IEEE International Conference on Computer Vision, 2021.
  39. 39.I. H. Laradji, N. Rostamzadeh, P. H. O. Pinheiro, D. Vazquez, and M. W. Schmidt, “Proposal-based instance segmentation with point supervision,” IEEE International Conference on Image Processing, 2020.
  40. 40.B. Cheng, O. Parkhi, and A. Kirillov, “Pointly-supervised instance segmentation,” ArXiv, 2021.
  41. 41.Y. Li, H. Zhao, X. Qi, Y. Chen, L. Qi, L. Wang, Z. Li, J. Sun, and J. Jia, “Fully convolutional networks for panoptic segmentation with point-based supervision,” ArXiv, 2021.
  42. 42.I. Radosavovic, P. Dollár, R. B. Girshick, G. Gkioxari, and K. He, “Data distillation: Towards omni-supervised learning,” IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  43. 43.K. Wang, X. Yan, D. Zhang, L. Zhang, and L. Lin, “Towards human-machine cooperation: Self-supervised sample mining for object detection,” European Conference on Computer Vision, 2018.
  44. 44.K. Sohn, Z. Zhang, C.-L. Li, H. Zhang, C.-Y. Lee, and T. Pfister, “A simple semi-supervised learning framework for object detection,” ArXiv, 2020.
  45. 45.J. Li, C. Zhang, P. Zhu, B. Wu, L. Chen, and Q. Hu, “Splmll: Selecting predictable landmarks for multi-label learning,” European Conference on Computer Vision, 2020.
  46. 46.B. Zoph, G. Ghiasi, T. Y. Lin, Y. Cui, H. Liu, E. D. Cubuk, and Q. V. Le, “Rethinking pre-training and self-training,” Neural Information Processing Systems, 2020.
  47. 47.M. Xu, Z. Zhang, H. Hu, J. Wang, L. Wang, F. Wei, X. Bai, and Z. Liu, “End-to-end semi-supervised object detection with soft teacher,” IEEE International Conference on Computer Vision, 2021.
  48. 48.Y. C. Liu, C. Y. Ma, Z. He, C. W. Kuo, K. Chen, P. Zhang, B. Wu, Z. Kira, and P. Vajda, “Unbiased teacher for semisupervised object detection,” International Conference on Learning Representations, 2021.
  49. 49.D. Huynh and E. Elhamifar, “Fine-grained generalized zero-shot learning via dense attribute-based attention,” IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  50. 50.Y. Xian, B. Schiele, and Z. Akata, “Zero-shot learning — the good, the bad and the ugly,” IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  51. 51.E. Schönfeld, S. Ebrahimi, S. Sinha, T. Darrell, and Z. Akata, “Generalized zero- and few-shot learning via aligned variational autoencoders,” IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  52. 52.R. Felix, B. G. V. Kumar, I. D. Reid, and G. Carneiro, “Multi-modal cycle-consistent generalized zero-shot learning,” European Conference on Computer Vision, 2018.
  53. 53.H. Jiang, R. Wang, S. Shan, and X. Chen, “Transferable contrastive network for generalized zero-shot learning,” IEEE International Conference on Computer Vision, 2019.
  54. 54.Y. Atzmon and G. Chechik, “Adaptive confidence smoothing for generalized zero-shot learning,” IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  55. 55.Y. Gong, S. Karanam, Z. Wu, K. Peng, J. Ernst, and P. Doerschuk, “Learning compositional visual concepts with mutual consistency,” IEEE Conference on Computer Vision and Pattern Recognition, 2018.
  56. 56.D. Huynh and E. Elhamifar, “A shared multi-attention framework for multi-label zero-shot learning,” IEEE Conference on Computer Vision and Pattern Recognition, 2020.
  57. 57.— — —, “Compositional zero-shot learning via fine-grained dense feature composition,” Neural Information Processing Systems, 2020.
  58. 58.— — —, “Interaction compass: Multi-label zero-shot learning of human-object interactions via spatial relations,” International Conference on Computer Vision, 2021.
  59. 59.Z. Li, L. Yao, X. Zhang, X. Wang, S. S. Kanhere, and H. Zhang, “Zero-shot object detection with textual descriptions,” AAAI Conference on Artificial Intelligence, 2019.
  60. 60.Y. Xian, S. Sharma, B. Schiele, and Z. Akata, “f-vaegand2: A feature generating framework for any-shot learning,” IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  61. 61.N. Kato, T. Yamasaki, and K. Aizawa, “Zero-shot semantic segmentation via variational mapping,” Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2019.
  62. 62.P. Hu, S. Sclaroff, and K. Saenko, “Uncertainty-aware learning for zero-shot semantic segmentation,” Neural Information Processing Systems, 2020.
  63. 63.G. Tian, S. Wang, J. Feng, L. Zhou, and Y. Mu, “Cap2seg: Inferring semantic and spatial context from captions for zero-shot image segmentation,” Proceedings of the 28th ACM International Conference on Multimedia, 2020.
  64. 64.P. Li, Y. Wei, and Y. Yang, “Consistent structural relation learning for zero-shot segmentation,” Neural Information Processing Systems, 2020.
  65. 65.H. Zhang and H. Ding, “Prototypical matching and open set rejection for zero-shot semantic segmentation,” IEEE International Conference on Computer Vision, 2021.
  66. 66.D. Baek, Y. Oh, and B. Ham, “Exploiting a joint embedding space for generalized zero-shot semantic segmentation,” IEEE International Conference on Computer Vision, 2021.
  67. 67.J. Cheng, S. Nandi, P. Natarajan, and W. Abd-Almageed, “Sign: Spatial-information incorporated generative network for generalized zero-shot semantic segmentation,” IEEE International Conference on Computer Vision, 2021.
  68. 68.M. Bucher, T. H. Vu, M. Cord, and P. Pérez, “Zero-shot semantic segmentation,” Neural Information Processing Systems, 2019.
  69. 69.S. Rahman, S. Khan, and N. Barnes, “Transductive learning for zero-shot object detection,” IEEE International Conference on Computer Vision, 2019.
  70. 70.G. Pastore, F. Cermelli, Y. Xian, M. Mancini, Z. Akata, and B. Caputo, “A closer look at self-training for zero-label semantic segmentation,” Conference on Computer Vision and Pattern Recognition Workshops, 2021.
  71. 71.H. H. Tan and M. Bansal, “Lxmert: Learning crossmodality encoder representations from transformers,” Empirical Methods in Natural Language Processing, 2019.
  72. 72.J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-andlanguage tasks,” Neural Information Processing Systems, 2019.
  73. 73.Y. Li, J. He, X. Zhou, Y. Zhang, and J. Baldridge, “Mapping natural language instructions to mobile ui action sequences,” Annual Meeting of the Association for Computational Linguistics, 2020.
  74. 74.G. Li, N. Duan, Y. Fang, D. Jiang, and M. Zhou, “Unicodervl: A universal encoder for vision and language by crossmodal pre-training,” AAAI Conference on Artificial Intelligence, 2020.
  75. 75.Y.-C. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “Uniter: Universal image-text representation learning,” European Conference on Computer Vision, 2020.
  76. 76.X. Yuan, Z. L. Lin, J. Kuen, J. Zhang, Y. Wang, M. Maire, A. Kale, and B. Faieta, “Multimodal contrastive training for visual representation learning,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  77. 77.K. Desai and J. Johnson, “Virtex: Learning visual representations from textual annotations,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  78. 78.A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” International Conference on Machine learning, 2021.
  79. 79.C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, “Understanding deep learning requires rethinking generalization,” International Conference on Learning Representations, 2017.
  80. 80.N. Natarajan, I. S. Dhillon, P. Ravikumar, and A. Tewari, “Learning with noisy labels,” Neural Information processing Systems, 2013.
  81. 81.E. A. Sanchez, D. Ortego, P. Albert, N. E. O’Connor, and K. McGuinness, “Unsupervised label noise modeling and loss correction,” International Conference on Machine learning, 2019.
  82. 82.T. Wang, R. Anwer, M. H. Khan, F. Khan, Y. Pang, L. Shao, and J. Laaksonen, “Deep contextual attention for humanobject interaction detection,” IEEE International Conference on Computer Vision, 2019.
  83. 83.X. Zhou, X. Liu, C. Wang, D. Zhai, J. Jiang, and X. Ji, “Learning with noisy labels via sparse regularization,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  84. 84.H. Zhang, X. Xing, and L. Liu, “Dualgraph: A graph-based method for reasoning about label noise,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  85. 85.Y. Xu, L. Zhu, L. Jiang, and Y. Yang, “Faster meta update strategy for noise-robust deep learning,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  86. 86.D. Ortego, E. Arazo, P. Albert, N. E. O’Connor, and K. McGuinness, “Multi-objective interpolation training for robustness to label noise,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  87. 87.M. Collier, B. Mustafa, E. Kokiopoulou, R. Jenatton, and J. Berent, “Correlated input-dependent label noise in largescale image classification,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  88. 88.Z. Zhu, T. Liu, and Y. Liu, “A second-order approach to learning with instance-dependent label noise,” IEEE Conference on Computer Vision and Pattern Recognition, 2021.
  89. 89.A. Veit, N. G. Alldrin, G. Chechik, I. Krasin, A. Gupta, and S. J. Belongie, “Learning from noisy large-scale datasets with minimal supervision,” IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  90. 90.K. Yi and J. Wu, “Probabilistic end-to-end noise correction for learning with noisy labels,” IEEE Conference on Computer Vision and Pattern Recognition, 2019.
  91. 91.J. Li, C. Xiong, and S. C. Hoi, “Learning from noisy data with robust representation learning,” IEEE International Conference on Computer Vision, 2021.
  92. 92.Y. Ding, L. Wang, D. Fan, and B. Gong, “A semi-supervised two-stage approach to learning from noisy labels,” IEEE Winter Conference on Applications of Computer Vision, 2018.
  93. 93.J. Li, R. Socher, and S. C. H. Hoi, “Dividemix: Learning with noisy labels as semi-supervised learning,” International Conference on Learning Representations, 2020.
  94. 94.A. Kendall and Y. Gal, “What uncertainties do we need in bayesian deep learning for computer vision?” 2017.
  95. 95.J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2019.
  96. 96.P. Tang, X. Wang, X. Bai, and W. Liu, “Multiple instance detection network with online instance classifier refinement,” IEEE Conference on Computer Vision and Pattern Recognition, 2017.
  97. 97.K. Ye, M. Zhang, A. Kovashka, W. Li, D. Qin, and J. Berent, “Cap2det: Learning to amplify weak caption supervision for object detection,” IEEE International Conference on Computer Vision, 2019.
  98. 98.T.-Y. Lin, M. Maire, S. J. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” European Conference on Computer Vision, 2014.
  99. 99.P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” Association for Computational Linguistics, 2018.
  100. 100.D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” International Conference on Learning Representations.
  101. 101.L. Yang, Q. Song, Z. Wang, Z. Liu, S. Xu, and Z. Li, “Quality-aware network for human parsing,” ArXiv, 2021.
  102. 102.Y. Gal and Z. Ghahramani, “Dropout as a bayesian approximation: Representing model uncertainty in deep learning,” International Conference on Machine learning, 2016.
  103. 103.S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun, “Objects365: A large-scale, high-quality dataset for object detection,” IEEE International Conference on Computer Vision, 2019.

Citation

MLA
Huynh, D., et al. “Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling”. arXiv, 2021, http://arxiv.org/abs/2111.12698v2.
APA
Huynh, D., Kuen, J., Lin, Z., Gu, J., & Elhamifar, E. (2021). Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling. arXiv. http://arxiv.org/abs/2111.12698v2
Chicago
Huynh, D., J. Kuen, Z. Lin, J. Gu, and E. Elhamifar. 2021. “Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling”. arXiv. http://arxiv.org/abs/2111.12698v2.
Harvard
Huynh, D. et al. (2021) “Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2111.12698v2.
Vancouver
1. Huynh D, Kuen J, Lin Z, Gu J, Elhamifar E (2021) Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling. arXiv

BibTeX

@article{huynh2021open,
  title = {Open-Vocabulary Instance Segmentation via Robust Cross-Modal Pseudo-Labeling},
  author = {Huynh, Dat and Kuen, Jason and Lin, Zhe and Gu, Jiuxiang and Elhamifar, Ehsan},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2111.12698v2},
  eprint = {2111.12698}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE