SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

Huaishao LuoJunwei BaoYouzheng WuXiaodong HeTianrui Li

article2023ICML216 citations

Proposes SegCLIP, an annotation-free open-vocabulary semantic segmentation framework that dynamically aggregates Vision Transformer patches into irregular semantic regions via learnable centers while training solely on image-text pairs with auxiliary reconstruction and superpixel-guided losses.

Listen

Traditional computer vision models for image segmentation—the process of identifying and outlining specific objects within an image—depend on labor-intensive, pixel-level human annotations and remain restricted to a fixed set of predefined categories. While recent vision-language foundation models capture rich visual concepts across vast vocabularies using web-scale text, adapting this broad knowledge to fine-grained pixel-level segmentation without expensive retraining or manual annotations remains a core technical challenge.

The article introduces and evaluates SegCLIP, an open-vocabulary semantic segmentation framework designed to segment images using arbitrary text labels in an annotation-free manner. The objective is to demonstrate that integrating a dynamic patch-aggregation mechanism into an existing pre-trained vision-language model enables accurate zero-shot semantic segmentation without requiring dense pixel labels or specialized segmentation decoders.

The authors develop a dual-encoder architecture that inserts a plug-in semantic grouping module into the middle layers of a Vision Transformer image encoder. This module uses learnable centers and cross-attention to dynamically group regular image patches into irregular semantic regions. The system is trained on approximately 3.4 million paired image-caption examples using an end-to-end objective combining standard contrastive learning with two novel self-supervised enhancements: a visual reconstruction loss on masked patches and a superpixel-based consistency loss derived from unsupervised graph segmentation. Evaluated benchmarks include standard validation sets from PASCAL VOC 2012, PASCAL Context, and COCO.

The evaluation yields several key findings. First, SegCLIP sets new state-of-the-art results for zero-shot text-supervised segmentation, achieving mean Intersection over Union (mIoU) scores of 52.6% on PASCAL VOC, 24.7% on PASCAL Context, and 26.5% on COCO, outperforming prior text-supervised baselines like GroupViT. Second, initializing SegCLIP with pre-trained vision-language weights dramatically boosts accuracy over training from scratch, increasing performance by up to 19.3 percentage points on PASCAL VOC. Third, the two auxiliary training objectives provide substantial gains, with the reconstruction and consistency losses jointly driving multi-point accuracy increases across all benchmarks. Finally, the framework operates with high computational efficiency, requiring roughly six hours of training on eight standard graphics processing units.

These results demonstrate that organizations can achieve highly flexible, open-vocabulary image understanding without incurring the massive labor costs and turnaround times associated with pixel-level labeling. By reusing pre-trained vision-language foundation models, enterprises can significantly cut training compute budgets while retaining the ability to recognize novel, arbitrary object classes during deployment, thereby reducing risks associated with rigid category definitions.

Decision-makers should consider adopting patch-aggregation and foundation model transfer strategies for cost-effective visual segmentation workflows. For production pipelines, technical teams should prioritize smaller patch inputs, which the article demonstrates improve boundary detail, and should evaluate whether off-the-shelf superpixel modules can be integrated directly into end-to-end training pipelines.

Confidence in these findings is high for standard object-recognition benchmarks; however, stakeholders should exercise caution when deploying the system in highly complex environments. In tests on dense scene understanding benchmarks, accuracy dropped significantly (8.7% on ADE20K and 11.0% on Cityscapes), indicating that complex visual scenes and boundary precision require further development and larger-scale dataset pretraining before deployment in high-stakes operational environments.

Cover for SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation

Abstract

Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual concepts for images by learning from a large scale of text-image data. However, transferring the learned visual knowledge to open-vocabulary semantic segmentation is still under-explored. In this paper, we propose a CLIP-based model named SegCLIP for the topic of open-vocabulary segmentation in an annotation-free manner. The SegCLIP achieves segmentation based on ViT and the main idea is to gather patches with learnable centers to semantic regions through training on text-image pairs. The gathering operation can dynamically capture the semantic groups, which can be used to generate the final segmentation results. We further propose a reconstruction loss on masked patches and a superpixel-based KL loss with pseudo-labels to enhance the visual representation. Experimental results show that our model achieves comparable or superior segmentation accuracy on the PASCAL VOC 2012 (+0.3% mIoU), PASCAL Context (+2.3% mIoU), and COCO (+2.2% mIoU) compared with baselines. We release the code at https://github.com/ArrowLuo/SegCLIP.

Table of Contents

  • 1. Introduction
  • 2. Model
  • 2.1. Main Architecture
  • 2.2. Semantic Group Module
  • 2.3. Reconstruction Loss
  • 2.4. Superpixel based KL Loss
  • 2.5. Training and Inference
  • 3. Experiments
  • 3.1. Datasets
  • 3.2. Experimental Details
  • 3.3. Ablation Studies
  • 3.4. Comparisons with State-of-the-Art Methods
  • 3.5. Qualitative Results
  • 4. Related Work
  • 4.1. Vision-Language pre-training
  • 4.2. Open-Vocabulary Semantic Segmentation
  • 5. Conclusion and Future Work
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — SegCLIP Semantic Group Module

    model/method

    The SegCLIP image encoder incorporates a pluggable semantic group module inserted into the intermediate layers of a Vision Transformer (ViT) to aggregate regular grid patches into irregular, arbitrarily shaped semantic region representations.

    Let Hp={hpjs}j=1N∈RN×DH_p = \{h_{p_j}^s\}_{j=1}^N \in \mathbb{R}^{N \times D} denote the hidden representations of NN image patches output by the ss-th (last layer of the first stage) Transformer layer. The semantic group module maintains a set of LL learnable center embeddings Hc={ck}k=1L∈RL×DH_c = \{c_k\}_{k=1}^L \in \mathbb{R}^{L \times D}. It first updates these centers through tt layers of cross-attention to obtain contextual centers H^c={c^k}k=1L∈RL×D\hat{H}_c = \{\hat{c}_k\}_{k=1}^L \in \mathbb{R}^{L \times D}: H^ct=CrossAttention(Hct,Hp,Hp)\hat{H}_c^t = \text{CrossAttention}(H_c^t, H_p, H_p) where the queries are the center representations HctH_c^t (with initial state Hc1=HcH_c^1 = H_c), and the keys and values are the patch representations HpH_p.

    A patch-to-center assignment mapping matrix M∈RN×LM \in \mathbb{R}^{N \times L} is computed using the Gumbel-Softmax operation: M=Gumbel-Softmax(HpH^c⊤)M = \text{Gumbel-Softmax}(H_p \hat{H}_c^\top) Each row of MM forms a one-hot assignment vector, ensuring that patch jj is uniquely assigned to center kk when Mjk=1M_{jk} = 1.

    The aggregated semantic region representations H^p∈RL×D\hat{H}_p \in \mathbb{R}^{L \times D} are then computed by averaging the patch embeddings assigned to each center, adding the contextual centers, and applying a multilayer perceptron block: H^p=MLP(MEAN(M⊤Hp)+H^c)\hat{H}_p = \text{MLP}(\text{MEAN}(M^\top H_p) + \hat{H}_c) where MEAN(⋅)\text{MEAN}(\cdot) denotes patch-averaging per center, and MLP(⋅)\text{MLP}(\cdot) consists of two fully connected linear layers with a GELU non-linearity between them. The resulting LL semantic region representations H^p\hat{H}_p are passed into the subsequent Transformer layers (the third stage of the image encoder) to yield interactive region features Zp∈RL×DZ_p \in \mathbb{R}^{L \times D}.

  2. Knowl 2 — Irregular Masked Patch Reconstruction Loss

    model/method

    To enhance the visual representation learned by the image encoder without supervised masks, SegCLIP uses a self-supervised masked image reconstruction objective defined over irregular semantic segments.

    Given an image II, a subset of patches is randomly masked at a masking ratio of 0.75. The unmasked patches are passed through the first stage of the image encoder and the semantic group module to generate masked region representations H^p(m)\hat{H}_p^{(m)} and a corresponding mapping matrix M(m)M^{(m)}.

    To restore patch-level tokens from the grouped region tokens, a reconstruction layer applies a transposed linear transformation followed by a GELU activation: H~p(m)=GELU(Linear(M(m))⊤H^p(m))\tilde{H}_p^{(m)} = \text{GELU}\left(\text{Linear}(M^{(m)})^\top \hat{H}_p^{(m)}\right) where Linear(⋅)\text{Linear}(\cdot) is a fully connected layer.

    The restored patch features H~p(m)\tilde{H}_p^{(m)} are processed by additional Transformer layers matching the third stage of the image encoder to obtain representations Zp(m)Z_p^{(m)}. A Masked Autoencoder (MAE) decoder then reconstructs the original pixel image I(m)I^{(m)} from Zp(m)Z_p^{(m)}. The reconstruction loss is the Mean Squared Error (MSE) computed against the original input image II: Lrec=MSE(I(m),I)\mathcal{L}_{\text{rec}} = \text{MSE}(I^{(m)}, I)

  3. Knowl 3 — Superpixel-Guided KL Consistency Loss

    model/method

    To preserve fine-grained spatial and pixel-level consistency during patch grouping, SegCLIP employs a superpixel-guided Kullback-Leibler (KL) divergence loss.

    Superpixel boundaries are extracted offline using an unsupervised graph-based segmentation algorithm (such as Felzenszwalb and Huttenlocher's method). Each patch j∈{1,…,N}j \in \{1, \dots, N\} is assigned a super-patch label determined by the majority superpixel ID of its constituent pixels. Let GjG_j denote the index set of patches belonging to the same super-patch as patch jj.

    Let Pj∈RLP_j \in \mathbb{R}^L denote the region probability distribution for patch jj, obtained by taking the softmax across the jj-th row of the patch-to-center affinity matrix before one-hot discretization. The average super-patch region distribution P^j\hat{P}_j is calculated as: P^j=softmax(1∣Gj∣∑j^∈GjPj^)\hat{P}_j = \text{softmax}\left(\frac{1}{|G_j|} \sum_{\hat{j} \in G_j} P_{\hat{j}}\right)

    A symmetric KL divergence loss Lsup\mathcal{L}_{\text{sup}} penalizes deviations between the assignment distribution of patch jj and its corresponding neighborhood average P^j\hat{P}_j: Lsup=12N∑j=1N(KL(Pj,P^j)+KL(P^j,Pj))\mathcal{L}_{\text{sup}} = \frac{1}{2N} \sum_{j=1}^N \left( \text{KL}(P_j, \hat{P}_j) + \text{KL}(\hat{P}_j, P_j) \right) This objective forces patches belonging to the same homogeneous superpixel to be grouped into the same semantic center.

  4. Knowl 4 — Open-Vocabulary Zero-Shot Inference Pipeline

    algorithm

    SegCLIP performs zero-shot open-vocabulary semantic segmentation by computing cross-modal similarity between text embeddings of candidate category prompts and region features, projecting similarities back to patches via the assignment matrix MM, and interpolating to full image resolution.

    Input: Image II, candidate class names C={c1,c2,…,cT}\mathcal{C} = \{c_1, c_2, \dots, c_T\}, foreground similarity threshold θ\theta
    Output: Pixel-level semantic segmentation mask Y∈{0,1,…,T}Himg×WimgY \in \{0, 1, \dots, T\}^{H_{\text{img}} \times W_{\text{img}}}
    for τ=1\tau = 1 to TT do
        text_input ←\leftarrow "a photo of a " + cτc_\tau
        z_w^{(\tau)} \leftarrow E_T(\text{text_input}) // Extract text feature from [SEP] token
    end for
    Zp,M←EI(I)Z_p, M \leftarrow E_I(I) // Zp∈RL×HZ_p \in \mathbb{R}^{L \times H} region features, M∈RN×LM \in \mathbb{R}^{N \times L} mapping matrix
    for k=1k = 1 to LL do
        for τ=1\tau = 1 to TT do
            S^k,τ←Zp,k⋅zw(τ)∥Zp,k∥∥zw(τ)∥\hat{S}_{k, \tau} \leftarrow \frac{Z_{p, k} \cdot z_w^{(\tau)}}{\|Z_{p, k}\| \|z_w^{(\tau)}\|} // Region-to-class cosine similarity
        end for
    end for
    S←MS^S \leftarrow M \hat{S} // Patch-to-class similarity matrix S∈RN×TS \in \mathbb{R}^{N \times T}
    Sdense←Interpolate(S,target_size=(Himg,Wimg))S_{\text{dense}} \leftarrow \text{Interpolate}(S, \text{target\_size}=(H_{\text{img}}, W_{\text{img}}))
    for each pixel (u,v)(u, v) do
        τ∗←arg⁡max⁡τSdense(u,v,τ)\tau^* \leftarrow \arg\max_\tau S_{\text{dense}}(u, v, \tau)
        if Sdense(u,v,τ∗)≥θS_{\text{dense}}(u, v, \tau^*) \ge \theta then
            Y(u,v)←τ∗Y(u, v) \leftarrow \tau^*
        else
            Y(u,v)←0Y(u, v) \leftarrow 0 // Assign background
        end if
    end for
    return YY
  5. Knowl 5 — SegCLIP Training Objective and Optimization Setup

    experimental setup

    SegCLIP is trained end-to-end using a joint objective combining image-text contrastive learning, masked autoencoding reconstruction, and superpixel KL consistency: Ltotal=Lcon+Lrec+Lsup\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{con}} + \mathcal{L}_{\text{rec}} + \mathcal{L}_{\text{sup}}

    The contrastive loss Lcon=12(Lt2i+Li2t)\mathcal{L}_{\text{con}} = \frac{1}{2}(\mathcal{L}_{t2i} + \mathcal{L}_{i2t}) is computed across a batch of size BB using the text representation zwz_w (from the [SEP] token) and the image representation zpz_p (obtained via max-pooling over the output region tokens of the final image encoder layer): Lt2i=−1B∑i=1Blog⁡exp⁡(s(zw(i),zp(i)))∑j=1Bexp⁡(s(zw(j),zp(i))),Li2t=−1B∑i=1Blog⁡exp⁡(s(zw(i),zp(i)))∑j=1Bexp⁡(s(zw(i),zp(j)))\mathcal{L}_{t2i} = -\frac{1}{B} \sum_{i=1}^B \log \frac{\exp(s(z_w^{(i)}, z_p^{(i)}))}{\sum_{j=1}^B \exp(s(z_w^{(j)}, z_p^{(i)}))}, \quad \mathcal{L}_{i2t} = -\frac{1}{B} \sum_{i=1}^B \log \frac{\exp(s(z_w^{(i)}, z_p^{(i)}))}{\sum_{j=1}^B \exp(s(z_w^{(i)}, z_p^{(j)}))} where s(zi,zj)=zizj⊤∥zi∥∥zj∥s(z_i, z_j) = \frac{z_i z_j^\top}{\|z_i\| \|z_j\|}.

    Architecture & Initialization:

    • Both the text encoder and the image encoder consist of 12 Transformer layers initialized from pretrained ViT-B/16 CLIP weights.
    • Image resolution is 224×224224 \times 224 with 16×1616 \times 16 patch resolution (N=196N = 196). Maximum text token length is 32.
    • The semantic group module is plugged after the 10th Transformer layer of the image encoder, configured with L=8L = 8 learnable centers and 2 cross-attention layers.
    • The MAE decoder consists of 3 Transformer layers with a patch masking ratio of 0.75.
    • Newly introduced components (group module, MAE decoder, linear projection heads) are randomly initialized.

    Training Details:

    • Pretraining datasets: Conceptual Captions (CC, 3M pairs) and COCO (400K pairs).
    • Optimizer: Adam with cosine learning rate schedule.
    • Learning rates: 4×10−64 \times 10^{-6} for pretrained CLIP weights (embeddings, text encoder, image Transformer layers 1–10); 4×10−34 \times 10^{-3} for newly initialized parameters.
    • Hardware: 8 NVIDIA A100 GPUs, total batch size 768, trained for 10 epochs (~6 hours).
  6. Knowl 6 — Open-Vocabulary Semantic Segmentation Performance

    data/table

    Zero-shot semantic segmentation performance of SegCLIP was evaluated on the validation splits of PASCAL VOC 2012 (20 classes, threshold 0.75), PASCAL Context (59 classes, threshold 0.25), and COCO (80 classes, threshold 0.65) using mean Intersection over Union (mIoU, %).

    Model Architecture Initialization Training Data VOC Context COCO
    DeiT ViT – ImageNet (Class Sup.) 53.0 35.9 –
    DINO ViT – CC12M+YFCC (Self-Sup.) 37.6 22.8 –
    MoCo ViT – CC12M+YFCC (Self-Sup.) 36.1 23.0 –
    GroupViT ViT – CC12M+YFCC (Text Sup.) 52.3 22.4 24.3
    GroupViT1-s_{\text{1-s}} ViT – CC+COCO (Text Sup.) 28.1 14.8 12.9
    GroupViT2-s_{\text{2-s}} ViT – CC+COCO (Text Sup.) 19.7 10.4 8.0
    SegCLIP (ours) ViT Scratch CC+COCO (Text Sup.) 33.3 19.1 15.2
    SegCLIP (ours) ViT CLIP Init CC+COCO (Text Sup.) 52.6 24.7 26.5

    When trained on the CC+COCO dataset (3.4M pairs), CLIP-initialized SegCLIP outperforms GroupViT trained on the larger CC12M+YFCC dataset (15M pairs) by +0.3% mIoU on VOC, +2.3% on Context, and +2.2% on COCO. Compared to GroupViT1-s_{\text{1-s}} trained on the same CC+COCO data from scratch, SegCLIP from scratch achieves +5.2%, +4.3%, and +2.3% mIoU gains.

  7. Knowl 7 — Ablation on Reconstruction and Superpixel KL Losses

    data/table

    Ablation results on the impact of the masked patch reconstruction loss (R-Loss, Lrec\mathcal{L}_{\text{rec}}) and the superpixel-based KL divergence loss (S-KL, Lsup\mathcal{L}_{\text{sup}}) on semantic segmentation mIoU (%):

    Model R-Loss S-KL VOC Context COCO
    SegCLIP 47.95 23.43 24.86
    SegCLIP ✓ 49.14 24.35 25.52
    SegCLIP ✓ 48.49 24.15 25.93
    SegCLIP ✓ ✓ 52.60 24.71 26.45

    Individually, R-Loss yields gains of +1.19% (VOC), +0.92% (Context), and +0.66% (COCO) over baseline contrastive training. S-KL alone achieves improvements of +0.54% (VOC), +0.72% (Context), and +1.07% (COCO). Combining both auxiliary losses produces an aggregate improvement of +4.65% on VOC, +1.28% on Context, and +1.59% on COCO over pure contrastive learning.

  8. Knowl 8 — Influence of Group Module Plugged Layer, Center Count, and Cross-Attention Depth

    empirical result

    Ablation experiments conducted on SegCLIP using only contrastive loss demonstrate the following effects of architectural hyperparameters:

    1. Plugged Layer Position: Inserting the semantic group module after the 10th Transformer layer yields the highest performance across all benchmarks (47.95% on VOC, 23.43% on Context, 24.86% on COCO). Placing the module too early (layer 6: 35.28% VOC, 19.28% Context, 16.73% COCO; layer 8: 43.75% VOC) damages pretrained representations, while placing it too late (layer 11: 22.07% VOC, 10.76% Context, 12.08% COCO) leaves insufficient subsequent layers to learn region-level representations.

    2. Number of Learnable Centers (LL): Performance remains stable between 6 and 10 centers, with L=8L = 8 performing optimally (47.95% VOC, 23.43% Context, 24.86% COCO) compared to L=6L = 6 (47.03% VOC, 23.36% Context, 24.85% COCO) and L=10L = 10 (44.89% VOC, 23.46% Context, 24.74% COCO).

    3. Cross-Attention Depth: Using 2 cross-attention layers in the semantic group module provides the best results (47.95% VOC, 23.43% Context, 24.86% COCO) compared to 0 layers (44.44% VOC), 1 layer (47.63% VOC), 3 layers (47.83% VOC), and 4 layers (45.39% VOC). Excess cross-attention layers degrade performance due to data scarcity during training.

  9. Knowl 9 — Limitations on Resolution, Scene Complexity, and Offline Superpixels

    limitation

    SegCLIP exhibits three notable limitations identified through empirical evaluation:

    1. Patch Resolution Dependency: Because the model groups coarse ViT patches (16×1616 \times 16 pixels) and applies linear interpolation to recover full-resolution masks, boundaries can be coarse. Increasing patch size to 32×3232 \times 32 (yielding 49 patches per image) causes severe performance drops compared to 16×1616 \times 16 (196 patches): VOC drops from 52.5% to 44.2% mIoU, Context drops from 24.7% to 22.0%, and COCO drops from 26.5% to 21.4%.

    2. Complex Scene Degradation: Evaluation on cluttered, complex scene parsing datasets reveals low absolute zero-shot performance: SegCLIP obtains 8.7% mIoU on ADE20K and 11.0% mIoU on Cityscapes (though exceeding GroupViT1-s_{\text{1-s}}'s 4.9% and 4.2%).

    3. Non-Differentiable Offline Superpixel Extraction: The superpixel generation is an offline, non-end-to-end preprocessing step using graph-based segmentation. Overly fine superpixels can bias patch grouping, requiring exploration of differentiable class-agnostic proposal modules.

Coverage note — None was omitted; all key architectural components, equations, training setups, ablation studies, main benchmark results, and stated limitations are fully represented.

References

  1. 1.Bain, M., Nagrani, A., Varol, G., and Zisserman, A. Frozen in time: A joint video and image encoder for end-to-end retrieval. In ICCV, pp. 1708–1718, 2021.
  2. 2.Bucher, M., Vu, T., Cord, M., and Pérez, P. Zero-shot semantic segmentation. In NeurIPS, 2019.
  3. 3.Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In ICCV, pp. 9630–9640, 2021.
  4. 4.Changpinyo, S., Sharma, P., Ding, N., and Soricut, R. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In CVPR, pp. 3558–3568, 2021.
  5. 5.Chefer, H., Gur, S., and Wolf, L. Transformer interpretability beyond attention visualization. In CVPR, pp. 782–791, 2021.
  6. 6.Chen, L., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L. Semantic image segmentation with deep convolutional nets and fully connected crfs. In ICLR, 2015.
  7. 7.Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., and Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018.
  8. 8.Chen, X., Xie, S., and He, K. An empirical study of training self-supervised vision transformers. In ICCV, pp. 9620–9629, 2021.
  9. 9.Chen, Y., Li, L., Yu, L., Kholy, A. E., Ahmed, F., Gan, Z., Cheng, Y., and Liu, J. UNITER: universal image-text representation learning. In ECCV, volume 12375, pp. 104–120, 2020.
  10. 10.Cheng, B., Schwing, A. G., and Kirillov, A. Per-pixel classification is not all you need for semantic segmentation. In NeurIPS, pp. 17864–17875, 2021.
  11. 11.Cheng, B., Misra, I., Schwing, A. G., Kirillov, A., and Girdhar, R. Masked-attention mask transformer for universal image segmentation. In CVPR, pp. 1290–1299, 2022.
  12. 12.Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. The cityscapes dataset for semantic urban scene understanding. In CVPR, pp. 3213–3223, 2016.
  13. 13.Devlin, J., Chang, M., Lee, K., and Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT, pp. 4171–4186, 2019.
  14. 14.Ding, J., Xue, N., Xia, G., and Dai, D. Decoupling zeroshot semantic segmentation. In CVPR, pp. 11573–11582, 2022a.
  15. 15.Ding, Z., Wang, J., and Tu, Z. Open-vocabulary panoptic segmentation with maskclip. arXiv preprint arXiv:2208.08984, 2022b.
  16. 16.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021.
  17. 17.Everingham, M., Gool, L. V., Williams, C. K. I., Winn, J. M., and Zisserman, A. The pascal visual object classes (VOC) challenge. Int. J. Comput. Vis., 88(2):303–338, 2010.
  18. 18.Felzenszwalb, P. F. and Huttenlocher, D. P. Efficient graphbased image segmentation. International journal of computer vision, 59(2):167–181, 2004.
  19. 19.Gan, Z., Li, L., Li, C., Wang, L., Liu, Z., and Gao, J. Visionlanguage pre-training: Basics, recent advances, and future trends. arXiv preprint arXiv:2210.09263, 2022.
  20. 20.Ghiasi, G., Gu, X., Cui, Y., and Lin, T.-Y. Scaling openvocabulary image segmentation with image-level labels. arXiv:2112.12143, 2021.
  21. 21.Gu, X., Lin, T.-Y., Kuo, W., and Cui, Y. Open-vocabulary object detection via vision and language knowledge distillation. In ICLR, 2022.
  22. 22.He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. B. Masked autoencoders are scalable vision learners. In CVPR, pp. 15979–15988, 2022.
  23. 23.Hendrycks, D. and Gimpel, K. Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415, 2016.
  24. 24.Huang, Z., Zeng, Z., Liu, B., Fu, D., and Fu, J. Pixel-bert: Aligning image pixels with text by deep multi-modal transformers. arXiv preprint arXiv:2004.00849, 2020.
  25. 25.Jain, J., Li, J., Chiu, M., Hassani, A., Orlov, N., and Shi, H. Oneformer: One transformer to rule universal image segmentation. arXiv preprint arXiv:abs/2211.06220, 2022.
  26. 26.Jang, E., Gu, S., and Poole, B. Categorical reparameterization with gumbel-softmax. In ICLR, 2017.
  27. 27.Jia, C., Yang, Y., Xia, Y., Chen, Y., Parekh, Z., Pham, H., Le, Q. V., Sung, Y., Li, Z., and Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In ICML, volume 139, pp. 4904–4916, 2021.
  28. 28.Kim, W., Son, B., and Kim, I. Vilt: Vision-and-language transformer without convolution or region supervision. In ICML, volume 139, pp. 5583–5594, 2021.
  29. 29.Li, B., Weinberger, K. Q., Belongie, S., Koltun, V., and Ranftl, R. Language-driven semantic segmentation. In ICLR, 2022a.
  30. 30.Li, J., Li, D., Xiong, C., and Hoi, S. C. H. BLIP: bootstrapping language-image pre-training for unified visionlanguage understanding and generation. In ICML, volume 162, pp. 12888–12900, 2022b.
  31. 31.Li, L., Gan, Z., Lin, K., Lin, C., Liu, Z., Liu, C., and Wang, L. LAVENDER: unifying video-language understanding as masked language modeling. arXiv preprint arXiv:2206.07160, 2022c.
  32. 32.Li, L. H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J., Chang, K., and Gao, J. Grounded language-image pre-training. In CVPR, pp. 10955–10965, 2022d.
  33. 33.Li, W., Gao, C., Niu, G., Xiao, X., Liu, H., Liu, J., Wu, H., and Wang, H. UNIMO: towards unified-modal understanding and generation via cross-modal contrastive learning. In ACL/IJCNLP, pp. 2592–2607, 2021.
  34. 34.Li, Y., Liang, F., Zhao, L., Cui, Y., Ouyang, W., Shao, J., Yu, F., and Yan, J. Supervision exists everywhere: A data efficient contrastive language-image pre-training paradigm. In ICLR, 2022e.
  35. 35.Liang, F., Wu, B., Dai, X., Li, K., Zhao, Y., Zhang, H., Zhang, P., Vajda, P., and Marculescu, D. Open-vocabulary semantic segmentation with mask-adapted CLIP. arXiv preprint arXiv:abs/2210.04150, 2022.
  36. 36.Lin, T., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft COCO: common objects in context. In ECCV, volume 8693, pp. 740–755, 2014.
  37. 37.Long, J., Shelhamer, E., and Darrell, T. Fully convolutional networks for semantic segmentation. In CVPR, pp. 3431–3440, 2015.
  38. 38.Lüddecke, T. and Ecker, A. S. Image segmentation using text and image prompts. In CVPR, pp. 7076–7086, 2022.
  39. 39.Luo, H., Ji, L., Shi, B., Huang, H., Duan, N., Li, T., Chen, X., and Zhou, M. Univl: A unified video and language pre-training model for multimodal understanding and generation. arXiv preprint arXiv:2002.06353, 2020.
  40. 40.Ma, C., Yang, Y., Wang, Y., Zhang, Y., and Xie, W. Openvocabulary semantic segmentation with frozen visionlanguage models. arXiv preprint arXiv:abs/2210.15138, 2022.
  41. 41.Miech, A., Zhukov, D., Alayrac, J., Tapaswi, M., Laptev, I., and Sivic, J. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In ICCV, pp. 2630–2640, 2019.
  42. 42.Mottaghi, R., Chen, X., Liu, X., Cho, N., Lee, S., Fidler, S., Urtasun, R., and Yuille, A. L. The role of context for object detection and semantic segmentation in the wild. In CVPR, pp. 891–898, 2014.
  43. 43.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In ICML, volume 139, pp. 8748–8763, 2021.
  44. 44.Rao, Y., Zhao, W., Chen, G., Tang, Y., Zhu, Z., Huang, G., Zhou, J., and Lu, J. Denseclip: Language-guided dense prediction with context-aware prompting. In CVPR, pp. 18061–18070, 2022.
  45. 45.Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, pp. 234–241, 2015.
  46. 46.Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A. LAION-400M: open dataset of clipfiltered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021.
  47. 47.Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV, pp. 618–626, 2017.
  48. 48.Sharma, P., Ding, N., Goodman, S., and Soricut, R. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In ACL, pp. 2556–2565, 2018.
  49. 49.Sun, C., Myers, A., Vondrick, C., Murphy, K., and Schmid, C. Videobert: A joint model for video and language representation learning. In ICCV, pp. 7463–7472, 2019.
  50. 50.Tan, H. and Bansal, M. LXMERT: learning cross-modality encoder representations from transformers. In EMNLP, 2019.
  51. 51.Thomee, B., Shamma, D. A., Friedland, G., Elizalde, B., Ni, K., Poland, D., Borth, D., and Li, L. YFCC100M: the new data in multimedia research. Commun. ACM, 59(2): 64–73, 2016.
  52. 52.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H. Training data-efficient image transformers & distillation through attention. In ICML, volume 139, pp. 10347–10357, 2021.
  53. 53.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is all you need. In NeurIPS, pp. 5998–6008, 2017.
  54. 54.Wang, R., Chen, D., Wu, Z., Chen, Y., Dai, X., Liu, M., Jiang, Y.-G., Zhou, L., and Yuan, L. Bevt: Bert pretraining of video transformers. In CVPR, pp. 14733–14743, 2022a.
  55. 55.Wang, Z., Yu, J., Yu, A. W., Dai, Z., Tsvetkov, Y., and Cao, Y. Simvlm: Simple visual language model pretraining with weak supervision. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022, 2022b.
  56. 56.Wen, X., Zhao, B., Zheng, A., Zhang, X., and Qi, X. Selfsupervised visual representation learning with semantic grouping. In NeurIPS, 2022.
  57. 57.Wu, C., Lin, Z., Cohen, S., Bui, T., and Maji, S. Phrasecut: Language-based image segmentation in the wild. In CVPR, pp. 10213–10222, 2020.
  58. 58.Xian, Y., Choudhury, S., He, Y., Schiele, B., and Akata, Z. Semantic projection network for zero- and few-label semantic segmentation. In CVPR, pp. 8256–8265, 2019.
  59. 59.Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., and Luo, P. Segformer: Simple and efficient design for semantic segmentation with transformers. In NeurIPS, pp. 12077–12090, 2021.
  60. 60.Xu, J., De Mello, S., Liu, S., Byeon, W., Breuel, T., Kautz, J., and Wang, X. Groupvit: Semantic segmentation emerges from text supervision. In CVPR, pp. 18134–18144, 2022a.
  61. 61.Xu, M., Zhang, Z., Wei, F., Lin, Y., Cao, Y., Hu, H., and Bai, X. A simple baseline for open vocabulary semantic segmentation with pre-trained vision-language model. ECCV, 2022b.
  62. 62.Yao, L., Huang, R., Hou, L., Lu, G., Niu, M., Xu, H., Liang, X., Li, Z., Jiang, X., and Xu, C. FILIP: fine-grained interactive language-image pre-training. In ICLR, 2022.
  63. 63.Zabari, N. and Hoshen, Y. Semantic segmentation inthe-wild without seeing any segmentation examples. arXiv:2112.03185, 2021.
  64. 64.Zeng, Y., Zhang, X., and Li, H. Multi-grained vision language pre-training: Aligning texts with visual concepts. In ICML, volume 162, pp. 25994–26009, 2022.
  65. 65.Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. Pyramid scene parsing network. In CVPR, pp. 6230–6239, 2017.
  66. 66.Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P. H., and Zhang, L. Rethinking semantic segmentation from a sequence-tosequence perspective with transformers. In CVPR, 2021.
  67. 67.Zhong, Y., Yang, J., Zhang, P., Li, C., Codella, N., Li, L. H., Zhou, L., Dai, X., Yuan, L., Li, Y., and Gao, J. Regionclip: Region-based language-image pretraining. In CVPR, pp. 16772–16782, 2022.
  68. 68.Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A. Scene parsing through ade20k dataset. In CVPR, pp. 633–641, 2017.
  69. 69.Zhou, C., Loy, C. C., and Dai, B. Extract free dense labels from clip. In ECCV, 2022a.
  70. 70.Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A. L., and Kong, T. iBOT: Image BERT pre-training with online tokenizer. In ICLR, 2022b.

Citation

MLA
Luo, H., et al. “SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation”. International Conference on Machine Learning, vol. 202, 2023, pp. 23033–44, https://proceedings.mlr.press/v202/luo23a.html.
APA
Luo, H., Bao, J., Wu, Y., He, X., & Li, T. (2023). SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation. International Conference on Machine Learning, 202, 23033–23044. https://proceedings.mlr.press/v202/luo23a.html
Chicago
Luo, H., J. Bao, Y. Wu, X. He, and T. Li. 2023. “SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation”. International Conference on Machine Learning 202: 23033–44. https://proceedings.mlr.press/v202/luo23a.html.
Harvard
Luo, H. et al. (2023) “SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation”, International Conference on Machine Learning. PMLR, pp. 23033–23044. Available at: https://proceedings.mlr.press/v202/luo23a.html.
Vancouver
1. Luo H, Bao J, Wu Y, He X, Li T (2023) SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation. In: International Conference on Machine Learning. PMLR, pp 23033–23044

BibTeX

@InProceedings{pmlr-v202-luo23a,
  title = 	 {{S}eg{CLIP}: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic Segmentation},
  author =       {Luo, Huaishao and Bao, Junwei and Wu, Youzheng and He, Xiaodong and Li, Tianrui},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {23033--23044},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/luo23a/luo23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/luo23a.html},
  abstract = 	 {Recently, the contrastive language-image pre-training, e.g., CLIP, has demonstrated promising results on various downstream tasks. The pre-trained model can capture enriched visual concepts for images by learning from a large scale of text-image data. However, transferring the learned visual knowledge to open-vocabulary semantic segmentation is still under-explored. In this paper, we propose a CLIP-based model named SegCLIP for the topic of open-vocabulary segmentation in an annotation-free manner. The SegCLIP achieves segmentation based on ViT and the main idea is to gather patches with learnable centers to semantic regions through training on text-image pairs. The gathering operation can dynamically capture the semantic groups, which can be used to generate the final segmentation results. We further propose a reconstruction loss on masked patches and a superpixel-based KL loss with pseudo-labels to enhance the visual representation. Experimental results show that our model achieves comparable or superior segmentation accuracy on the PASCAL VOC 2012 (+0.3% mIoU), PASCAL Context (+2.3% mIoU), and COCO (+2.2% mIoU) compared with baselines. We release the code at https://github.com/ArrowLuo/SegCLIP.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/