Learning Mask-aware CLIP Representations for Zero-Shot Segmentation

Siyu JiaoYunchao WeiYaowei WangYao ZhaoHumphrey Shi

article2023NeurIPS82 citations

Proposes a mask-aware fine-tuning strategy that makes pre-trained CLIP sensitive to region-level proposals without adding new parameters or sacrificing zero-shot transferability, substantially boosting unseen-class segmentation accuracy across standard benchmarks.

Listen

Modern computer vision systems are increasingly tasked with identifying and segmenting novel objects that were never seen during initial training. A common industry strategy pairs a region proposal generator with a frozen, pre-trained vision-language model, specifically CLIP, to classify image segments. However, because these pre-trained vision models are originally trained on full images rather than distinct regions, they struggle to differentiate between clean object boundaries and low-quality proposals that contain background noise or partial shapes. This limitation leads to frequent false positives, high computational redundancy, and degraded segmentation accuracy.

The article evaluates a fine-tuning strategy called Mask-Aware Fine-Tuning (MAFT), designed to make vision-language models responsive to specific mask regions without losing their general transfer capabilities. To achieve this, the authors introduce an Image-Proposals CLIP Encoder that uses masked attention mechanisms to process an entire image alongside multiple candidate regions simultaneously. They optimize this framework using two complementary loss functions: a mask-aware loss that ties classification confidence directly to mask overlap quality, and a self-distillation loss that uses the original frozen model as a teacher network to prevent catastrophic forgetting and overfitting.

Evaluating the technique across established benchmarks—including COCO-Stuff, Pascal-VOC, ADE20K, and Pascal-Context—demonstrated substantial performance gains. When integrated into leading baseline frameworks such as FreeSeg, MAFT improved segmentation accuracy for unseen object classes by +8.2% on COCO-Stuff (from 42.2% to 50.4%), +3.2% on Pascal-VOC (from 78.6% to 81.8%), and nearly doubled accuracy on ADE20K from 4.4% to 8.7%. In broader open-vocabulary scenarios, performance jumped significantly, including a +19.1% boost on Pascal-Context and +11.2% on ADE20K. Additionally, the unified image-proposal processing architecture drastically cut redundant computation, reducing image encoder floating-point operations from 1,127.0 GFLOPs to 53.4 GFLOPs—over a 95% reduction.

These findings indicate that the standard practice of keeping vision-language models strictly frozen is an unnecessary constraint that limits precision. By incorporating region-aware fine-tuning, organizations can achieve substantially higher visual recognition accuracy while drastically lowering inference and computing costs. Because the method introduces no extra parameters and operates as a lightweight, plug-and-play module requiring less than a single epoch of training, it presents minimal implementation risk and engineering overhead.

Engineering and research teams deploying open-vocabulary or zero-shot image segmentation should integrate region-aware fine-tuning into their existing pipelines. The method successfully generalizes across various architectures, including ViT, ResNet backbones, and modern proposal generators like the Segment Anything Model. While these results show high reliability across diverse benchmarks, the ultimate zero-shot classification ceiling remains constrained by the baseline knowledge of the underlying pre-trained vision-language model, representing an important area for future improvements.

Cover for Learning Mask-aware CLIP Representations for Zero-Shot Segmentation

Abstract

Recently, pre-trained vision-language models have been increasingly used to tackle the challenging zero-shot segmentation task. Typical solutions follow the paradigm of first generating mask proposals and then adopting CLIP to classify them. To maintain the CLIP's zero-shot transferability, previous practices favour to freeze CLIP during training. However, in the paper, we reveal that CLIP is insensitive to different mask proposals and tends to produce similar predictions for various mask proposals of the same image. This insensitivity results in numerous false positives when classifying mask proposals. This issue mainly relates to the fact that CLIP is trained with image-level supervision. To alleviate this issue, we propose a simple yet effective method, named Mask-aware Fine-tuning (MAFT). Specifically, Image-Proposals CLIP Encoder (IP-CLIP Encoder) is proposed to handle arbitrary numbers of image and mask proposals simultaneously. Then, mask-aware loss and self-distillation loss are designed to fine-tune IP-CLIP Encoder, ensuring CLIP is responsive to different mask proposals while not sacrificing transferability. In this way, mask-aware representations can be easily learned to make the true positives stand out. Notably, our solution can seamlessly plug into most existing methods without introducing any new parameters during the fine-tuning process. We conduct extensive experiments on the popular zero-shot benchmarks. With MAFT, the performance of the state-of-the-art methods is promoted by a large margin: 50.4% (+ 8.2%) on COCO, 81.8% (+ 3.2%) on Pascal-VOC, and 8.7% (+4.3%) on ADE20K in terms of mIoU for unseen classes. Code is available at github.com/jiaosiyu1999/MAFT.git.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 4 Methodology
  • 4.1 Image-Proposal CLIP Encoder (IP-CLIP Encoder)
  • 4.2 Objective
  • 5 Experiments
  • 5.1 Setting
  • 5.2 Comparisons with State-of-the-art Methods
  • 5.3 Ablation Study
  • 5.4 Extending MAFT with SAM
  • 5.5 Extending MAFT with more Vision-Language Models
  • 5.6 Qualitative Study
  • 6 Conclusion
  • A Technical details of the "frozen CLIP" approaches
  • B Dataset
  • C Additional experiments
  • C.1 Analysis of the Upper Bound of MAFT
  • C.2 Analysis of the Self-Training (ST) strategy
  • D Visualization
  • References

Knowls

  1. Knowl 1 — Mask-Aware Fine-Tuning (MAFT) Framework for Zero-Shot Segmentation

    model/method

    Mask-Aware Fine-Tuning (MAFT) is a framework designed to adapt pre-trained vision-language models (such as CLIP) for zero-shot and open-vocabulary semantic segmentation. Prior methods following the "frozen CLIP" paradigm generate class-agnostic mask proposals and crop or mask sub-images before passing them to a frozen CLIP classifier. Because CLIP is trained with image-level contrastive objectives, it is insensitive to regional mask boundaries and produces high-confidence false positive predictions for partial or background proposals.

    MAFT resolves this limitation by fine-tuning the vision encoder directly on mask proposals without modifying its architecture or introducing new model parameters. The framework operates in two main components:

    1. Image-Proposal CLIP Encoder (IP-CLIP Encoder): Modifies the attention mechanism of the pre-trained vision encoder to accept the whole image alongside an arbitrary number of binary mask proposals simultaneously, utilizing proposal masks as attention biases in Transformer layers.

    2. Joint Training Objectives: Optimizes the IP-CLIP Encoder using a mask-aware loss Lma\mathcal{L}_{ma} (which aligns classification confidence scores with proposal-ground-truth Intersection over Union) and a self-distillation loss Ldis\mathcal{L}_{dis} (which aligns whole-image representations with a frozen CLIP teacher model to retain zero-shot generalizability).

    During fine-tuning, the proposal generator (e.g., MaskFormer, Mask2Former, or SAM) and the CLIP Text Encoder remain completely frozen. Fine-tuning is computationally lightweight, requiring only a few iterations (less than 1 epoch).

  2. Knowl 2 — Image-Proposal CLIP Encoder Architecture and Masked Attention

    model/method

    The Image-Proposal CLIP Encoder (IP-CLIP Encoder) adapts a standard Vision Transformer (ViT) image encoder to process an input image together with NN class-agnostic binary mask proposals M∈{0,1}N×h×wM \in \{0, 1\}^{N \times h \times w} in a single forward pass.

    For a 12-layer ViT, token propagation is divided at layer LL:

    1. Early Layers (i=1,…,Li = 1, \dots, L): Tokens propagate via standard full self-attention. The feature representation is Fi=[Fclsi;Ffeati]∈R(1+hw)×dF^i = [F_{cls}^i; F_{feat}^i] \in \mathbb{R}^{(1 + hw) \times d}, where Fclsi∈R1×dF_{cls}^i \in \mathbb{R}^{1 \times d} is the class token, Ffeati∈Rhw×dF_{feat}^i \in \mathbb{R}^{hw \times d} represents flattened spatial image patch tokens, and dd is the hidden channel dimension. This stage preserves global context across the entire image.

    2. Proposal Expansion (at layer LL): The class token FclsLF_{cls}^L is duplicated NN times to produce proposal-specific classification tokens FclsL∗∈RN×dF_{cls}^{L*} \in \mathbb{R}^{N \times d}, yielding expanded features FL∗=[FclsL∗;FfeatL]∈R(N+hw)×dF^{L*} = [F_{cls}^{L*}; F_{feat}^L] \in \mathbb{R}^{(N + hw) \times d}.

    3. Masked Transformer Layers (i=L+1,…,12i = L+1, \dots, 12):

    • The class tokens Fclsi∗F_{cls}^{i*} are updated via masked Multihead Attention using an attention bias B∈RN×(N+hw)B \in \mathbb{R}^{N \times (N + hw)}:

    B(j,k)={0if M^(j,k)=1−∞if M^(j,k)=0B(j, k) = \begin{cases} 0 & \text{if } \hat{M}(j, k) = 1 \\ -\infty & \text{if } \hat{M}(j, k) = 0 \end{cases}

    where M^=[I(N,N);Flat(M)]\hat{M} = [I(N, N); \text{Flat}(M)], I(N,N)I(N, N) is the N×NN \times N identity matrix, and Flat(M)∈{0,1}N×hw\text{Flat}(M) \in \{0, 1\}^{N \times hw} is the flattened mask tensor. The class token update is:

    Fcls(i+1)∗=Softmax(Que(Fclsi∗) Key(Fi∗)Td+B)Val(Fi∗)F_{cls}^{(i+1)*} = \text{Softmax}\left(\frac{\text{Que}(F_{cls}^{i*}) \, \text{Key}(F^{i*})^T}{\sqrt{d}} + B\right) \text{Val}(F^{i*})

    • The spatial patch features FfeatiF_{feat}^i continue to propagate via standard self-attention without attention masks:

    Ffeati+1=Softmax(Que(Ffeati) Key(Ffeati)Td)Val(Ffeati)F_{feat}^{i+1} = \text{Softmax}\left(\frac{\text{Que}(F_{feat}^i) \, \text{Key}(F_{feat}^i)^T}{\sqrt{d}}\right) \text{Val}(F_{feat}^i)

    This structure ensures proposal class tokens attend strictly to their masked foreground regions and full image context, while patch features retain global interaction. Computation drops from 1127.0 GFLOPs1127.0\text{ GFLOPs} (in standard crop/merge pipelines) to 53.4 GFLOPs53.4\text{ GFLOPs}.

  3. Knowl 3 — Mask-Aware Loss with Min-Max IoU Normalization

    equation

    To supervise the IP-CLIP Encoder so that proposal classification scores reflect proposal localization quality, the mask-aware loss Lma\mathcal{L}_{ma} aligns the predicted classification confidence with the Intersection-over-Union (IoU) scores between the proposals and ground-truth masks.

    Let SIoU∈[0,1]K×NS_{IoU} \in [0, 1]^{K \times N} denote the matrix of IoU values between NN mask proposals and KK ground-truth binary class masks present in an image. Because raw maximum IoU scores typically range between 0.750.75 and 0.990.99 while softmax-derived classification scores Ac∈[0,1]N×CA^c \in [0, 1]^{N \times C} approach 1.01.0, a min-max normalization across proposals is applied to SIoUS_{IoU}:

    SIoUnorm=SIoU−min⁡(SIoU)max⁡(SIoU)−min⁡(SIoU),SIoU∈RK×NS_{IoU}^{norm} = \frac{S_{IoU} - \min(S_{IoU})}{\max(S_{IoU}) - \min(S_{IoU})}, \quad S_{IoU} \in \mathbb{R}^{K \times N}

    Let Aselectc∈RK×NA_{select}^c \in \mathbb{R}^{K \times N} denote the predicted proposal classification scores corresponding to the KK present ground-truth classes. The mask-aware loss is computed using the Smooth L1 loss:

    Lma(Aselectc,SIoUnorm)=SmoothL1(Aselectc,SIoUnorm)\mathcal{L}_{ma}(A_{select}^c, S_{IoU}^{norm}) = \text{SmoothL1}(A_{select}^c, S_{IoU}^{norm})

    where the elementwise Smooth L1 loss function is defined as:

    SmoothL1(x,y)={0.5⋅(x−y)2if ∣x−y∣<1∣x−y∣−0.5otherwise\text{SmoothL1}(x, y) = \begin{cases} 0.5 \cdot (x - y)^2 & \text{if } |x - y| < 1 \\ |x - y| - 0.5 & \text{otherwise} \end{cases}

  4. Knowl 4 — Self-Distillation Objective for Zero-Shot Knowledge Preservation

    equation

    To prevent catastrophic forgetting of CLIP's broad open-vocabulary knowledge and prevent overfitting on seen training classes CseenC_{seen}, fine-tuning incorporates a self-distillation objective Ldis\mathcal{L}_{dis}.

    A frozen pre-trained CLIP model serves as the teacher network, and the trainable IP-CLIP Encoder serves as the student network. When processing an input image without mask proposals (denoted w.o. M\text{w.o. } M), the IP-CLIP Encoder produces an unmasked global classification prediction vector AS∈RC×1A_S \in \mathbb{R}^{C \times 1}. The teacher network produces prediction vector AT∈RC×1A_T \in \mathbb{R}^{C \times 1}.

    The self-distillation loss minimizes the discrepancy between student and teacher whole-image predictions:

    Ldis(AS,AT)=SmoothL1(AS,AT)\mathcal{L}_{dis}(A_S, A_T) = \text{SmoothL1}(A_S, A_T)

    The total training loss for fine-tuning the IP-CLIP Encoder is:

    L=Lma+λLdis\mathcal{L} = \mathcal{L}_{ma} + \lambda \mathcal{L}_{dis}

    where λ\lambda is a balancing hyperparameter set to λ=1.0\lambda = 1.0.

  5. Knowl 5 — Zero-Shot Semantic Segmentation Performance on Popular Benchmarks

    empirical result

    The effectiveness of Mask-Aware Fine-Tuning (MAFT) was evaluated across three zero-shot semantic segmentation benchmarks: COCO-Stuff (156 seen / 15 unseen classes), Pascal-VOC (15 seen / 5 unseen classes), and ADE20K (572 seen / 275 unseen classes), using ResNet-101 backbones for proposal generators and ViT-B/16 for CLIP.

    Method COCO-Stuff Pascal-VOC ADE20K
    mIoUs\text{mIoU}^s mIoUu\text{mIoU}^u hIoU\text{hIoU} mIoUs\text{mIoU}^s mIoUu\text{mIoU}^u hIoU\text{hIoU} mIoUs\text{mIoU}^s mIoUu\text{mIoU}^u hIoU\text{hIoU}
    SPNet 34.6 26.9 30.3 77.8 25.8 38.8 - - -
    ZS5 34.9 10.6 16.2 78.0 21.2 33.3 - - -
    CaGNet 35.6 13.4 19.5 78.6 30.3 43.7 - - -
    STRICT 35.3 30.3 32.6 82.7 35.6 73.3 - - -
    ZegFormer 36.7 36.2 36.4 90.1 70.6 79.2 17.4 5.1 7.9
    ZegFormer + MAFT 36.4 40.1 38.1 91.5 80.7 85.7 16.6 7.0 9.8
    ZSSeg 40.4 36.5 38.3 86.6 59.7 69.4 18.0 4.5 7.2
    ZSSeg + MAFT 40.6 40.1 40.3 88.4 66.2 75.7 18.9 6.7 9.9
    FreeSeg 42.4 42.2 42.3 91.9 78.6 84.7 22.3 4.4 7.3
    FreeSeg + MAFT 43.3 50.4 46.5 91.4 81.8 86.3 21.4 8.7 12.4

    When isolating the proposal classifier by removing score ensembling between the proposal generator and CLIP, MAFT yields substantial standalone gains on FreeSeg: unseen mIoU increases from 29.3%29.3\% to 49.7%49.7\% (+20.4%+20.4\%) on COCO-Stuff, from 74.7%74.7\% to 84.7%84.7\% (+10.0%+10.0\%) on Pascal-VOC, and from 2.8%2.8\% to 8.7%8.7\% (+5.9%+5.9\%) on ADE20K.

  6. Knowl 6 — Open-Vocabulary Semantic Segmentation Evaluation Across Five Datasets

    empirical result

    Models trained exclusively on COCO-Stuff (171 classes) with ViT-B/16 CLIP were evaluated in an open-vocabulary setting across five target benchmarks: ADE20K-847 (A-847), ADE20K-150 (A-150), Pascal-Context-459 (PC-459), Pascal-Context-59 (PC-59), and Pascal-VOC 20 (PAS-20).

    Method A-847 A-150 PC-459 PC-59 PAS-20
    SPNet - - - 24.3 18.3
    ZSSeg - - - 19.4 38.3
    LSeg+ 2.5 13.0 5.2 36.0 59.0
    OVSeg 7.1 24.8 11.0 53.3 92.6
    OpenSeg* 8.8 28.6 12.2 48.2 72.2
    FreeSeg 7.1 17.9 6.4 34.4 85.6
    FreeSeg + MAFT 10.1 29.1 12.8 53.5 90.0

    (Note: OpenSeg utilizes additional external image-caption supervision data.)*

    Plugging MAFT into FreeSeg improves mIoU by +3.0%+3.0\% on A-847, +11.2%+11.2\% on A-150, +6.4%+6.4\% on PC-459, +19.1%+19.1\% on PC-59, and +4.4%+4.4\% on PAS-20, establishing superior performance compared to prior methods without requiring additional training data.

  7. Knowl 7 — Ablation Analysis of MAFT Components, Losses, and Architectural Choices

    empirical result

    Ablation experiments conducted on the COCO-Stuff dataset (without proposal generator score ensembling) reveal the contribution of each design choice:

    1. Component Contributions:
    • FreeSeg baseline (frozen CLIP): 22.3% mIoUs22.3\% \text{ mIoU}^s, 29.3% mIoUu29.3\% \text{ mIoU}^u, 25.3% hIoU25.3\% \text{ hIoU}, 1127.0 GFLOPs1127.0\text{ GFLOPs}.
    • Adding IP-CLIP Encoder: 29.4% mIoUs29.4\% \text{ mIoU}^s, 36.2% mIoUu36.2\% \text{ mIoU}^u, 32.4% hIoU32.4\% \text{ hIoU}, 53.4 GFLOPs53.4\text{ GFLOPs}.
    • Adding Lma\mathcal{L}_{ma} fine-tuning: 39.9% mIoUs39.9\% \text{ mIoU}^s, 47.1% mIoUu47.1\% \text{ mIoU}^u, 43.1% hIoU43.1\% \text{ hIoU}.
    • Adding Ldis\mathcal{L}_{dis} self-distillation: 40.1% mIoUs40.1\% \text{ mIoU}^s, 49.7% mIoUu49.7\% \text{ mIoU}^u, 44.4% hIoU44.4\% \text{ hIoU}.
    1. Loss Function Formulation for Lma\mathcal{L}_{ma}: Smooth L1 loss achieves 47.1% mIoUu47.1\% \text{ mIoU}^u, outperforming L1 (45.8%45.8\%), L2 (45.8%45.8\%), and Kullback-Leibler (KL) divergence (41.8%41.8\%). KL loss overfits seen classes (40.9% mIoUs40.9\% \text{ mIoU}^s) at the expense of transferability.

    2. Training Iterations: Zero-shot performance on unseen classes peaks at 1,000 iterations (49.7% mIoUu49.7\% \text{ mIoU}^u). Training for 5,000 iterations improves seen mIoU to 42.0%42.0\% but degrades unseen mIoU to 45.7%45.7\% due to overfitting.

    3. Selective Layer Freezing: Freezing the convolution patch projection, class embedding, positional embedding, and MLP blocks while tuning only attention parameters improves unseen mIoU from 44.7%44.7\% (full tuning) to 49.7%49.7\%.

    4. Start Layer LL for Masked Attention: Setting L=8L = 8 yields optimal performance (49.7% mIoUu49.7\% \text{ mIoU}^u), compared to L=0L = 0 (46.4%46.4\%) or L=10L = 10 (45.7%45.7\%), balancing context aggregation in early layers with mask specialization in later layers.

  8. Knowl 8 — Generalization of MAFT to Segment Anything Model Proposals and Diverse CLIP Backbones

    empirical result

    MAFT generalizes effectively across alternative proposal generators and vision-language backbone architectures:

    1. Integration with Segment Anything Model (SAM): Using SAM-H as the proposal generator, applying MAFT improves zero-shot performance over SAM with frozen CLIP from 86.7%86.7\% to 88.6% mIoUu88.6\% \text{ mIoU}^u on Pascal-VOC, and from 43.3%43.3\% to 51.5% mIoUu51.5\% \text{ mIoU}^u on COCO-Stuff. In open-vocabulary evaluations, SAM + MAFT achieves 12.7%12.7\% on A-847, 33.0%33.0\% on A-150, 16.2%16.2\% on PC-459, 59.0%59.0\% on PC-59, and 92.7%92.7\% on PAS-20.

    2. Scaling to Larger and Convolutional Backbones:

    Method Backbone A-847 A-150 PC-459 PC-59 PAS-20
    OVSeg CLIP-ViT-L 9.0 29.6 12.4 55.7 94.5
    FreeSeg CLIP-ViT-L 8.5 21.0 7.6 33.8 86.4
    FreeSeg + MAFT CLIP-ViT-L 12.1 32.0 15.7 58.5 92.1
    FreeSeg CLIP-ResNet-50 5.3 15.5 5.4 28.2 87.1
    FreeSeg + MAFT CLIP-ResNet-50 8.4 27.0 9.9 50.8 89.0

    For ResNet-50 CLIP, MAFT modifies the AttentionPool2d layer by duplicating FclsF_{cls} and inserting proposal mask attention biases, raising PC-59 performance by +22.6%+22.6\%.

  9. Knowl 9 — Upper Bound and Self-Training Analysis of Proposal-Based Zero-Shot Segmentation

    empirical result
    1. Oracle Proposal Upper Bound: Replacing predicted CLIP proposal classification scores AcA^c with oracle ground-truth IoU scores SIoUS_{IoU} during inference (with Mask2Former proposals) yields 77.2% mIoUs77.2\% \text{ mIoU}^s, 82.1% mIoUu82.1\% \text{ mIoU}^u, and 77.6% mIoU77.6\% \text{ mIoU} on COCO-Stuff. This demonstrates that modern proposal generators provide high-quality mask candidates, but highlights an approximate 30% mIoU30\% \text{ mIoU} headroom between current state-of-the-art classifier performance (50.4% mIoUu50.4\% \text{ mIoU}^u) and the theoretical upper limit.

    2. Self-Training (ST) Strategy: Using FreeSeg + MAFT to generate pseudo-labels for unseen classes on the training set and subsequently retraining the model further boosts unseen class mIoU from 81.8%81.8\% to 86.3%86.3\% on Pascal-VOC (+4.5%+4.5\%) and from 50.4%50.4\% to 55.2%55.2\% on COCO-Stuff (+4.8%+4.8\%). However, self-training requires pre-specifying target unseen class vocabularies during training, limiting its practical deployment in fully open-vocabulary settings.

  10. Knowl 10 — Fundamental Novel-Class Recognition Bound of Pre-Trained Vision-Language Encoders

    limitation

    Although Mask-Aware Fine-Tuning (MAFT) effectively suppresses background false positives and makes CLIP sensitive to proposal mask geometry, the ultimate zero-shot recognition capability on novel, unseen classes remains strictly bounded by the semantic representation capacity and vocabulary coverage of the underlying pre-trained vision-language model.

Coverage note — None was omitted; all key contributed methods, architectural specifications, loss formulations, zero-shot and open-vocabulary experiments, ablations, SAM and backbone extensions, upper-bound analyses, and limitations are fully covered.

References

  1. 1.Maxime Bucher, Tuan-Hung Vu, Matthieu Cord, and Patrick Pérez. Zero-shot semantic segmentation. Advances in Neural Information Processing Systems, 32, 2019.
  2. 2.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1209–1218, 2018.
  3. 3.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017.
  4. 4.Bowen Cheng, Anwesa Choudhuri, Ishan Misra, Alexander Kirillov, Rohit Girdhar, and Alexander G Schwing. Mask2former for video instance segmentation. arXiv preprint arXiv:2112.10764, 2021.
  5. 5.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. Advances in Neural Information Processing Systems, 34, 2021.
  6. 6.Jian Ding, Nan Xue, Gui-Song Xia, and Dengxin Dai. Decoupling zero-shot semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11583–11592, 2022.
  7. 7.Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International journal of computer vision, 111:98–136, 2015.
  8. 8.Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Open-vocabulary image segmentation. arXiv preprint arXiv:2112.12143, 2021.
  9. 9.Golnaz Ghiasi, Xiuye Gu, Yin Cui, and Tsung-Yi Lin. Scaling open-vocabulary image segmentation with image-level labels. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVI, pages 540–557. Springer, 2022.
  10. 10.Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang. Context-aware feature generation for zero-shot semantic segmentation. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1921–1929, 2020.
  11. 11.Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao, and Liqing Zhang. Context-aware feature generation for zero-shot semantic segmentation. In Proceedings of the 28th ACM International Conference on Multimedia, pages 1921–1929, 2020.
  12. 12.Zixian Guo, Bowen Dong, Zhilong Ji, Jinfeng Bai, Yiwen Guo, and Wangmeng Zuo. Texts as images in prompt tuning for multi-label image recognition. arXiv preprint arXiv:2211.12739, 2022.
  13. 13.Kunyang Han, Yong Liu, Jun Hao Liew, Henghui Ding, Yunchao Wei, Jiajun Liu, Yitong Wang, Yansong Tang, Yujiu Yang, Jiashi Feng, et al. Global knowledge calibration for fast open-vocabulary segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023.
  14. 14.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  15. 15.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 603–612, 2019.
  16. 16.Zilong Huang, Yunchao Wei, Xinggang Wang, Wenyu Liu, Thomas S Huang, and Humphrey Shi. Alignseg: Feature-aligned segmentation networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(1):550–557, 2021.
  17. 17.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR, 2021.
  18. 18.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643, 2023.
  19. 19.Harold W Kuhn. The hungarian method for the assignment problem. Naval research logistics quarterly, 2(1-2):83–97, 1955.
  20. 20.Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. Open-vocabulary semantic segmentation with mask-adapted clip. arXiv preprint arXiv:2210.04150, 2022.
  21. 21.Jiaxu Miao, Xiaohan Wang, Yu Wu, Wei Li, Xu Zhang, Yunchao Wei, and Yi Yang. Large-scale video panoptic segmentation in the wild: A benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21033–21043, 2022.
  22. 22.Jiaxu Miao, Yunchao Wei, Yu Wu, Chen Liang, Guangrui Li, and Yi Yang. Vspw: A large-scale dataset for video scene parsing in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4133–4143, 2021.
  23. 23.Juhong Min, Dahyun Kang, and Minsu Cho. Hypercorrelation squeeze for few-shot segmentation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6941–6952, 2021.
  24. 24.Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 891–898, 2014.
  25. 25.Giuseppe Pastore, Fabio Cermelli, Yongqin Xian, Massimiliano Mancini, Zeynep Akata, and Barbara Caputo. A closer look at self-training for zero-label semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2693–2702, 2021.
  26. 26.Giuseppe Pastore, Fabio Cermelli, Yongqin Xian, Massimiliano Mancini, Zeynep Akata, and Barbara Caputo. A closer look at self-training for zero-label semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2693–2702, 2021.
  27. 27.Jie Qin, Jie Wu, Pengxiang Yan, Ming Li, Ren Yuxi, Xuefeng Xiao, Yitong Wang, Rui Wang, Shilei Wen, Xin Pan, et al. Freeseg: Unified, universal and open-vocabulary image segmentation. arXiv preprint arXiv:2303.17225, 2023.
  28. 28.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021.
  29. 29.Amirreza Shaban, Shray Bansal, Zhen Liu, Irfan Essa, and Byron Boots. One-shot learning for semantic segmentation. arXiv preprint arXiv:1709.03410, 2017.
  30. 30.Yanpeng Sun, Qiang Chen, Xiangyu He, Jian Wang, Haocheng Feng, Junyu Han, Errui Ding, Jian Cheng, Zechao Li, and Jingdong Wang. Singular value fine-tuning: Few-shot segmentation requires few-parameters fine-tuning. arXiv preprint arXiv:2206.06122, 2022.
  31. 31.Yongqin Xian, Subhabrata Choudhury, Yang He, Bernt Schiele, and Zeynep Akata. Semantic projection network for zero-and few-label semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8256–8265, 2019.
  32. 32.Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang. Groupvit: Semantic segmentation emerges from text supervision. CVPR, 2022.
  33. 33.Mengde Xu, Zheng Zhang, Fangyun Wei, Yutong Lin, Yue Cao, Han Hu, and Xiang Bai. A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIX, pages 736–753. Springer, 2022.
  34. 34.Nir Zabari and Yedid Hoshen. Semantic segmentation in-the-wild without seeing any segmentation examples. arXiv preprint arXiv:2112.03185, 2021.
  35. 35.Gengwei Zhang, Guoliang Kang, Yi Yang, and Yunchao Wei. Few-shot segmentation via cycle-consistent transformer. Advances in Neural Information Processing Systems, 34:21984–21996, 2021.
  36. 36.Gengwei Zhang, Shant Navasardyan, Ling Chen, Yao Zhao, Yunchao Wei, Honghui Shi, et al. Mask matching transformer for few-shot segmentation. Advances in Neural Information Processing Systems, 35:823–836, 2022.
  37. 37.Zekang Zhang, Guangyu Gao, Zhiyuan Fang, Jianbo Jiao, and Yunchao Wei. Mining unseen classes via regional objectness: A simple baseline for incremental segmentation. Advances in Neural Information Processing Systems, 35:24340–24353, 2022.
  38. 38.Zekang Zhang, Guangyu Gao, Jianbo Jiao, Chi Harold Liu, and Yunchao Wei. Coinseg: Contrast inter-and intra-class representations for incremental segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023.
  39. 39.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2881–2890, 2017.
  40. 40.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 633–641, 2017.
  41. 41.Chong Zhou, Chen Change Loy, and Bo Dai. Extract free dense labels from clip. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXVIII, pages 696–712. Springer, 2022.
  42. 42.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Conditional prompt learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16816–16825, 2022.
  43. 43.Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu. Learning to prompt for vision-language models. International Journal of Computer Vision, 130(9):2337–2348, 2022.

Citation

MLA
Jiao, S., et al. “Learning Mask-aware CLIP Representations for Zero-Shot Segmentation”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 35631–53, https://proceedings.neurips.cc/paper_files/paper/2023/file/6ffe484a646db13891bb6435ca39d667-Paper-Conference.pdf.
APA
Jiao, S., Wei, Y., Wang, Y., Zhao, Y., & Shi, H. (2023). Learning Mask-aware CLIP Representations for Zero-Shot Segmentation. Advances in Neural Information Processing Systems, 36, 35631–35653. https://proceedings.neurips.cc/paper_files/paper/2023/file/6ffe484a646db13891bb6435ca39d667-Paper-Conference.pdf
Chicago
Jiao, S., Y. Wei, Y. Wang, Y. Zhao, and H. Shi. 2023. “Learning Mask-aware CLIP Representations for Zero-Shot Segmentation”. Advances in Neural Information Processing Systems 36: 35631–53. https://proceedings.neurips.cc/paper_files/paper/2023/file/6ffe484a646db13891bb6435ca39d667-Paper-Conference.pdf.
Harvard
Jiao, S. et al. (2023) “Learning Mask-aware CLIP Representations for Zero-Shot Segmentation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 35631–35653. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/6ffe484a646db13891bb6435ca39d667-Paper-Conference.pdf.
Vancouver
1. Jiao S, Wei Y, Wang Y, Zhao Y, Shi H (2023) Learning Mask-aware CLIP Representations for Zero-Shot Segmentation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 35631–35653

BibTeX

@inproceedings{jiao2023learning,
  title = {Learning Mask-aware CLIP Representations for Zero-Shot Segmentation},
  author = {Jiao, Siyu and Wei, Yunchao and Wang, Yaowei and Zhao, Yao and Shi, Humphrey},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {35631-35653},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/6ffe484a646db13891bb6435ca39d667-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors