CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

Yuqi LinMinghao ChenWenxiao WangBoxi WuKe LiBinbin LinHaifeng LiuXiaofei He

article2023CVPR204 citations

Presents CLIP-ES, a training-free framework that adapts frozen CLIP models with softmax-modified GradCAM, text-driven prompt strategies, and real-time attention refinement to generate high-quality pseudo segmentation masks tenfold faster than traditional multi-stage weakly supervised methods.

Listen

Semantic segmentation—identifying and outlining specific objects within digital images at the pixel level—is critical for computer vision applications but traditionally demands labor-intensive, expensive pixel-by-pixel manual annotations. Weakly supervised semantic segmentation addresses this bottleneck by training models using only image-level tags. However, conventional weak supervision pipelines require complex, multi-stage workflows that train separate classification models and refinement networks, resulting in high computational costs, prolonged training cycles, and noisy object boundaries.

The article evaluates whether a frozen, pre-trained vision-language model (Contrastive Language-Image Pre-training, or CLIP) can directly generate high-quality segmentation masks without task-specific training, and it demonstrates a streamlined framework called CLIP-ES that improves accuracy and efficiency across the entire segmentation pipeline.

The evaluation tests the proposed approach on standard benchmark datasets, including PASCAL VOC 2012 (over 10,500 training images) and MS COCO 2014 (over 82,000 training images). Instead of fine-tuning the base model, the framework leverages the zero-shot capabilities of a Vision Transformer-based CLIP model, applying targeted modifications to gradient-based localization, text prompt engineering, attention-based refinement, and downstream loss calculations.

The analysis yields four primary findings. First, modifying gradient activation maps with a softmax function and a defined background category set resolves class confusion, boosting initial activation quality from 49.4% to 58.6% mean Intersection over Union (mIoU) on PASCAL VOC. Second, task-specific text engineering—specifically selecting low-dispersion prompts and fusing synonyms—significantly sharpens localization, improving person category segmentation from 43.6% to 51.6% mIoU. Third, the real-time Class-Aware Attention-based Affinity module refines activation maps to 75.0% mIoU when paired with standard post-processing, eliminating the need to train a dedicated affinity network. Finally, the overall framework generates pseudo-segmentation masks in 0.6 hours using 2 GB of memory—representing a more than tenfold reduction in time and memory compared to leading multi-stage baselines requiring 6 to 77 hours and 18 GB—while establishing state-of-the-art segmentation accuracy of 73.8% mIoU on PASCAL VOC and 45.4% on MS COCO.

These findings indicate that organizations can drastically reduce compute costs, hardware requirements, and development timelines for dense visual recognition tasks by repurposing foundation models rather than training multi-stage architectures from scratch. The results challenge standard practices by showing that single-prompt selection and single-scale inference outperform traditional prompt ensembling and multi-scale aggregation in multi-label segmentation contexts.

Decision-makers should consider adopting training-free vision-language feature extraction for weak supervision pipelines to accelerate model deployment and lower labeling expenses. When implementing this approach, teams should utilize text-driven background suppression and confidence-guided loss functions to filter out boundary noise automatically. Future technical efforts should focus on refining the framework for crowded scenes, occlusions, and very small objects, where performance bottlenecks persist.

Confidence in the reported efficiency and accuracy gains is high within the tested benchmark environments. However, leaders should note that the system's localization relies on semantic representations learned during large-scale pre-training, which may exhibit lower fidelity when applied to highly specialized domain vocabularies or severe object occlusions.

Cover for CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation

Abstract

Weakly supervised semantic segmentation (WSSS) with image-level labels is a challenging task. Mainstream approaches follow a multi-stage framework and suffer from high training costs. In this paper, we explore the potential of Contrastive Language-Image Pre-training models (CLIP) to localize different categories with only image-level labels and without further training. To efficiently generate high-quality segmentation masks from CLIP, we propose a novel WSSS framework called CLIP-ES. Our framework improves all three stages of WSSS with special designs for CLIP: 1) We introduce the softmax function into GradCAM and exploit the zero-shot ability of CLIP to suppress the confusion caused by non-target classes and backgrounds. Meanwhile, to take full advantage of CLIP, we re-explore text inputs under the WSSS setting and customize two text-driven strategies: sharpness-based prompt selection and synonym fusion. 2) To simplify the stage of CAM refinement, we propose a real-time class-aware attention-based affinity (CAA) module based on the inherent multi-head self-attention (MHSA) in CLIP-ViTs. 3) When training the final segmentation model with the masks generated by CLIP, we introduced a confidence-guided loss (CGL) focus on confident regions. Our CLIP-ES achieves SOTA performance on Pascal VOC 2012 and MS COCO 2014 while only taking 10% time of previous methods for the pseudo mask generation. Code is available at https://github.com/linyq2117/CLIP-ES.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Weakly Supervised Semantic Segmentation
  • 2.2. Contrastive Language-Image Pretraining
  • 3. Method
  • 3.1. Softmax-GradCAM
  • 3.2. Text-driven Strategies
  • 3.2.1 Sharpness-based Prompt Selection
  • 3.2.2 Synonym Fusion
  • 3.3. Class-aware Attention-based Affinity (CAA)
  • 3.4. Confidence-guided Loss (CGL)
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Experimental Results
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Softmax-GradCAM and Class-Related Background Suppression

    model/method

    In weakly supervised semantic segmentation (WSSS) with multi-label images, vanilla GradCAM computes gradients directly from unnormalized class logits YcY^c, lacking mutual competition between classes and leading to activation confusion across foreground categories as well as between foreground and co-occurring background elements.

    Softmax-GradCAM introduces a softmax normalization across C=K+MC = K + M categories (comprising KK target categories present in the image and MM predefined common class-related background categories):

    sc=exp⁡(Yc)∑c′=1Cexp⁡(Yc′)s^c = \frac{\exp(Y^c)}{\sum_{c'=1}^C \exp(Y^{c'})}

    where scs^c is the softmax probability of category cc. The gradient-based class feature weight wkcw_k^c for the kk-th feature map Ak∈Rh×wA^k \in \mathbb{R}^{h \times w} with Z=h×wZ = h \times w pixels is formulated as:

    wkc=1Z∑i∑j∂Yc∂Aijksc(1−sc)+1Z∑i∑j∑c′≠c∂Yc′∂Aijksc(−sc′)w_k^c = \frac{1}{Z}\sum_i \sum_j \frac{\partial Y^c}{\partial A_{ij}^k} s^c(1 - s^c) + \frac{1}{Z}\sum_i \sum_j \sum_{c' \neq c} \frac{\partial Y^{c'}}{\partial A_{ij}^k} s^c(-s^{c'})

    The resulting Class Activation Map (CAM) for class cc at location (i,j)(i, j) is computed via:

    CAMijc=ReLU(∑kwkcAijk)CAM_{ij}^c = \text{ReLU}\left(\sum_k w_k^c A_{ij}^k\right)

    By including the background category set MM, non-target background pixels (e.g., water for boat, railway for train) are actively suppressed through negative gradients without requiring model retraining.

  2. Knowl 2 — Class-Aware Attention-based Affinity (CAA) for CAM Refinement

    model/method

    Multi-Head Self-Attention (MHSA) in Vision Transformers (ViT) captures pairwise patch affinity but is class-agnostic, which risks propagating noisy activations to unrelated semantically similar areas if applied directly. Class-Aware Attention-based Affinity (CAA) converts MHSA into a symmetric class-aware spatial affinity operator.

    Given the asymmetric self-attention weight matrix Wattn∈Rhw×hwW^{\text{attn}} \in \mathbb{R}^{hw \times hw} from ViT, Sinkhorn normalization is applied to produce a doubly stochastic matrix D=Sinkhorn(Wattn)D = \text{Sinkhorn}(W^{\text{attn}}), from which a symmetric affinity matrix AA is formed:

    A=D+DT2A = \frac{D + D^T}{2}

    For each target class CAM map Mc∈Rh×wM_c \in \mathbb{R}^{h \times w}, the activation map is thresholded at λ\lambda. Connected components are identified, and their minimal bounding boxes are combined into a binary box mask Bc∈R1×hwB_c \in \mathbb{R}^{1 \times hw}. The refined activation map McaffM_c^{\text{aff}} after tt propagation iterations is given by:

    Mcaff=Bc⊙At⋅vec(Mc)M_c^{\text{aff}} = B_c \odot A^t \cdot \text{vec}(M_c)

    where ⊙\odot is the Hadamard product and vec(⋅)\text{vec}(\cdot) denotes vectorization. Using bounding boxes instead of tight pixel masks allows the propagation to recover missing object parts while bounding the spread of class-agnostic attention noise.

  3. Knowl 3 — Sharpness-Guided Text Prompt Selection for Multi-Label WSSS

    model/method

    While prompt ensembling improves single-label zero-shot classification in CLIP, it degrades multi-label localization in WSSS because boosting the most prominent category suppresses the softmax scores and gradients of co-occurring target categories.

    To evaluate prompts for multi-label localization without pixel annotations, the sharpness metric measures the score dispersion across target classes over a dataset of nn images:

    sharpness(prompt)=∑i=1nvar(si1,…,sik)∑i=1nmean(si1,…,sik)\text{sharpness}(\text{prompt}) = \frac{\sum_{i=1}^n \text{var}(s_{i1}, \dots, s_{ik})}{\sum_{i=1}^n \text{mean}(s_{i1}, \dots, s_{ik})}

    where sijs_{ij} is the post-softmax score of the jj-th target category in image ii (k≥1k \ge 1 classes per image). Empirical evaluations show a negative correlation between prompt sharpness and CAM mIoU. Prompts incorporating abstract descriptions and adjectives yield the lowest sharpness and highest localization quality; the prompt template selected based on this metric is "a clean origami {}.".

  4. Knowl 4 — Sentence-Level Synonym Fusion and Semantic Customization

    model/method

    To disambiguate polysemous category names and enrich semantic coverage without incurring the computational cost of multiple forward passes, synonym fusion concatenates multiple synonymous terms into a single text prompt at the sentence level (e.g., "A clean origami of person, people, human").

    Additionally, category-specific visual biases can be corrected by customizing descriptive phrases. For example, CLIP standard embeddings for "person" bias CAM activations towards human faces rather than whole bodies; replacing "person" with "person with clothes" expands the activation to cover the full body area matching semantic segmentation ground truth.

  5. Knowl 5 — Confidence-Guided Loss (CGL) for Segmentation Training

    model/method

    Standard pseudo-mask generation applies hard thresholds to CAMs, introducing label noise at semantically ambiguous regions such as object boundaries. Confidence-Guided Loss (CGL) filters out low-confidence pixel locations during semantic segmentation model training.

    Given class activation maps X∈Rh×w×cX \in \mathbb{R}^{h \times w \times c} normalized for an image with cc target classes, the confidence map Conf(i,j)\text{Conf}(i, j) is computed as:

    Conf(i,j)=max⁡(1−max⁡cX(i,j,c),  max⁡cX(i,j,c))\text{Conf}(i, j) = \max\left(1 - \max_c X(i, j, c), \; \max_c X(i, j, c)\right)

    The training loss L^(i,j)\hat{L}(i, j) at pixel location (i,j)(i, j) modulates the standard cross-entropy loss L(i,j)L(i, j) using a confidence threshold parameter μ\mu:

    L^(i,j)={L(i,j),if Conf(i,j)≥μ0,if Conf(i,j)<μ\hat{L}(i, j) = \begin{cases} L(i, j), & \text{if } \text{Conf}(i, j) \ge \mu \\ 0, & \text{if } \text{Conf}(i, j) < \mu \end{cases}

  6. Knowl 6 — CLIP-ES Architecture and CAM Extraction Pipeline

    model/method

    The CLIP-ES framework uses a frozen pre-trained CLIP model with a ViT-B/16 vision backbone.

    1. Feature Extraction: CAMs are extracted from the feature map immediately preceding the final self-attention layer in the vision transformer.
    2. Logit Computation: Rather than using the class token ([CLS][\text{CLS}]), the final classification logits are calculated by taking the average of the remaining spatial patch tokens, which improves localization quality.
    3. Inference & Refinement: CAM generation and CAA refinement are executed in a single forward pass without multi-scale input aggregation. Threshold λ\lambda for CAA box masking is set to 0.40.4 for PASCAL VOC and 0.70.7 for MS COCO.
    4. Supervision: Pseudo masks generated from refined CAMs (further post-processed by dense CRF) supervise a ResNet-101 DeepLabV2 segmentation network trained with Confidence-Guided Loss.
  7. Knowl 7 — Quality of Pseudo Masks on PASCAL VOC 2012

    data/table

    On the PASCAL VOC 2012 training set, Class Activation Maps generated by CLIP-ES outperform previous weakly supervised methods on both initial seeds and refined pseudo masks without requiring auxiliary classification training or separate affinity networks.

    Method Seed dCRF RW
    IRN 48.8% 54.3% 66.3%
    SC-CAM 50.9% 55.3% 63.4%
    SEAM 55.4% 56.8% 63.6%
    AdvCAM 55.6% 62.1% 68.0%
    CLIMS 56.6% 62.4% 70.5%
    RIB 56.5% 62.9% 70.6%
    OoD 59.1% 65.5% 72.1%
    MCTformer 61.7% 64.5% 69.1%
    CLIP-ES (Ours) 70.8% 75.0% -

    Metric is mIoU (%). Seed is the initial CAM; dCRF denotes dense CRF post-processing; RW denotes random walk / affinity network refinement. CAA refinement combined with dCRF achieves 75.0% mIoU, exceeding prior methods that train dedicated affinity networks.

  8. Knowl 8 — Semantic Segmentation Benchmarks on PASCAL VOC 2012 and MS COCO 2014

    data/table

    When training a DeepLabV2 (ResNet-101) segmentation network using the generated pseudo masks, CLIP-ES achieves state-of-the-art weakly supervised semantic segmentation performance on PASCAL VOC 2012 and MS COCO 2014 under image-level plus language supervision.

    Dataset Supervision Backbone Val mIoU Test mIoU
    PASCAL VOC 2012 Image + Lang ResNet-101 (V2) 71.1% 71.4%
    PASCAL VOC 2012 Image + Lang ResNet-101 (V2, COCO pretrain) 73.8% 73.9%
    MS COCO 2014 Image + Lang ResNet-101 (V2) 45.4% -

    On PASCAL VOC 2012, CLIP-ES outperforms previous image-level supervised methods (e.g., MCTformer at 71.9% Val / 71.6% Test) as well as methods utilizing saliency maps (e.g., PPC+EPS at 72.6% Val / 73.6% Test). On MS COCO 2014, CLIP-ES reaches 45.4% Val mIoU, outperforming AMN (44.7%) and L2G (44.2%).

  9. Knowl 9 — Pseudo-Mask Generation Time and Memory Efficiency

    data/table

    CLIP-ES eliminates separate classification model training and affinity network training stages, significantly reducing the computational and memory footprint for generating pseudo masks on the PASCAL VOC 2012 augmented training set (10,582 images).

    Method Train Time (h) Inference Time (h) dCRF Time (h) Affinity Time (h) Total Time (h) Peak Memory
    AdvCAM - 70.5 0.2 6.5 77.2 18 GB
    CLIMS 2.1 0.3 0.2 6.5 9.1 18 GB
    MCTformer 0.5 2.5 - 3.0 6.0 18 GB
    CLIP-ES (Ours) - 0.4 0.2 - 0.6 2 GB

    CLIP-ES runs in 0.6 hours total with a peak GPU memory consumption of 2 GB, achieving over a 10x speedup and 9x memory reduction compared to prior methods.

  10. Knowl 10 — Ablation of Softmax-GradCAM, Background Sets, and Component Contributions

    empirical result

    Ablation experiments on PASCAL VOC 2012 confirm the contributions of individual framework components:

    • Softmax-GradCAM: Introducing softmax across the 20 VOC target classes improves initial CAM mIoU from 49.4% to 53.3%.
    • Background Category Suppression: Adding the class-related background set MM to Softmax-GradCAM further increases initial CAM mIoU from 53.3% to 58.6%. Classes heavily prone to background confusion show large gains: "boat" (confused with water) increases from 24.1% to 46.9% mIoU (+22.8%), and "train" (confused with railway) increases from 43.8% to 57.5% mIoU (+13.7%).
    • CAA Refinement vs. Vanilla MHSA: On initial CAMs (58.6% mIoU), applying vanilla MHSA achieves 68.2% (72.1% with dCRF), whereas CAA achieves 70.8% (75.0% with dCRF).
    • Synonym Fusion: Evaluating initial CAMs shows improvements across categories (e.g., "person" increases from 43.6% to 51.6% mIoU; "chair" from 40.7% to 44.1% mIoU).
    • Confidence-Guided Loss (CGL): In final segmentation training, CGL improves VOC validation performance from 70.6% to 71.1% (from 73.3% to 73.8% with COCO pretraining) and COCO validation performance from 45.1% to 45.4%.

Coverage note — None was omitted; all key contributions including Softmax-GradCAM, CAA refinement, prompt selection, synonym fusion, CGL, and empirical benchmarks were captured.

References

  1. 1.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In CVPR, 2019. 1, 3, 6, 7
  2. 2.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In CVPR, 2018. 1, 3, 6, 7
  3. 3.Nikita Araslanov and Stefan Roth. Single-stage semantic segmentation from image labels. In CVPR, June 2020. 2
  4. 4.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What's the point: Semantic segmentation with point supervision. In ECCV, 2016. 1
  5. 5.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weaklysupervised semantic segmentation via sub-category exploration. In CVPR, 2020. 2, 6, 7
  6. 6.Liyin Chen, Weiwei Wu, Chenchen Fu, Xiao Han, and Yuntao Zhang. Weakly supervised semantic segmentation with boundary exploration. In ECCV, 2020. 3, 7
  7. 7.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017. 1
  8. 8.Qi Chen, Lingxiao Yang, Jian-Huang Lai, and Xiaohua Xie. Self-supervised image-specific prototype exploration for weakly supervised semantic segmentation. In CVPR, June 2022. 2, 7
  9. 9.Zhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua, Hanwang Zhang, and Qianru Sun. Class re-activation maps for weakly-supervised semantic segmentation. In CVPR, 2022. 2, 7
  10. 10.Jifeng Dai, Kaiming He, and Jian Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In ICCV, 2015. 1
  11. 11.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR. Ieee, 2009. 4
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2021. 2
  13. 13.Ye Du, Zehua Fu, Qingjie Liu, and Yunhong Wang. Weakly supervised semantic segmentation by pixel-to-prototype contrast. In CVPR, 2022. 7
  14. 14.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010. 4, 6
  15. 15.Junsong Fan, Zhaoxiang Zhang, Chunfeng Song, and Tieniu Tan. Learning integral objects with intra-class discriminator for weakly-supervised semantic segmentation. In CVPR, 2020. 3, 7
  16. 16.Junsong Fan, Zhaoxiang Zhang, and Tieniu Tan. Cian: Cross-image affinity net for weakly supervised semantic segmentation. In AAAI, 2020. 2
  17. 17.Qibin Hou, Peng-Tao Jiang, Yunchao Wei, and Ming-Ming Cheng. Self-erasing network for integral object attention. In NeurIPS, 2018. 2
  18. 18.Peng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng, Yunchao Wei, and Hongkai Xiong. Integral object mining via online attention accumulation. In ICCV, 2019. 1, 2, 7
  19. 19.Peng-Tao Jiang, Yuqi Yang, Qibin Hou, and Yunchao Wei. L2g: A simple local-to-global knowledge transfer framework for weakly supervised semantic segmentation. In CVPR, 2022. 3, 7
  20. 20.Beomyoung Kim, Sangeun Han, and Junmo Kim. Discriminative region suppression for weakly-supervised semantic segmentation. In AAAI, 2021. 2, 7
  21. 21.Philipp Krahenbuhl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. In NeurIPS, 2011. 6
  22. 22.Hyeok Ryool Kweon, Sung-Hoon Yoon, Hyeonseong Kim, Dae-Soon Park, and Kuk-Jin Yoon. Unlocking the potential of ordinary classifier: Class-specific adversarial erasing framework for weakly supervised semantic segmentation. In ICCV, 2021. 2
  23. 23.Jungbeom Lee, Jooyoung Choi, Ji-Yoon Choi Ji-Hyeok Moon Young-Ilc Mok, and Sungroh Yoon. Reducing information bottleneck for weakly supervised semantic segmentation. In NeurIPS, 2021. 6, 7
  24. 24.Jungbeom Lee, Eunji Kim, and Sungroh Yoon. Anti-adversarially manipulated attributions for weakly and semisupervised semantic segmentation. In CVPR, 2021. 1, 2, 6, 7
  25. 25.Jungbeom Lee, Seong Joon Oh, Sangdoo Yun, Junsuk Choe, Eunji Kim, and Sungroh Yoon. Weakly supervised semantic segmentation using out-of-distribution data. In CVPR, 2022. 2, 6
  26. 26.Minhyun Lee, Dongseob Kim, and Hyunjung Shim. Threshold matters in wsss: Manipulating the activation for the robust and accurate segmentation model against thresholds. In CVPR, 2022. 7
  27. 27.Seungho Lee, Minhyun Lee, Jongwuk Lee, and Hyunjung Shim. Railroad is not a train: Saliency as pseudo-pixel supervision for weakly supervised semantic segmentation. In CVPR, 2021. 3, 6, 7
  28. 28.Xueyi Li, Tianfei Zhou, Jianwu Li, Yi Zhou, and Zhaoxiang Zhang. Group-wise semantic mining for weakly supervised semantic segmentation. In AAAI, 2021. 2
  29. 29.Yi Li, Yiqun Duan, Zhanghui Kuang, Yimin Chen, Wayne Zhang, and Xiaomeng Li. Uncertainty estimation via response scaling for pseudo-mask noise mitigation in weaklysupervised semantic segmentation. In AAAI, 2022. 3, 7
  30. 30.Yi Li, Zhanghui Kuang, Liyang Liu, Yimin Chen, and Wayne Zhang. Pseudo-mask matters in weakly-supervised semantic segmentation. In ICCV, 2021. 3
  31. 31.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In CVPR, 2016. 1
  32. 32.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 6
  33. 33.George Papandreou, Liang-Chieh Chen, Kevin P. Murphy, and Alan Loddon Yuille. Weakly-and semi-supervised learning of a deep convolutional network for semantic image segmentation. In ICCV, 2015. 1
  34. 34.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021. 2, 3, 6
  35. 35.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 2
  36. 36.Lixiang Ru, Bo Du, and Chen Wu. Learning visual words for weakly-supervised semantic segmentation. In IJCAI, 2021. 2
  37. 37.Lixiang Ru, Yibing Zhan, Baosheng Yu, and Bo Du. Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers. In CVPR, 2022. 2, 5
  38. 38.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV, 2017. 3, 4
  39. 39.Richard Sinkhorn. A relationship between arbitrary positive matrices and doubly stochastic matrices. Annals of Mathematical Statistics, 35:876–879, 1964. 5
  40. 40.Robin Strudel, Ricardo Garcia Pinel, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. In ICCV, 2021. 1
  41. 41.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In ECCV, 2020. 1, 2, 7
  42. 42.Paul Vernaza and Manmohan Chandraker. Learning random-walk label propagation for weakly-supervised semantic segmentation. In CVPR, 2017. 1
  43. 43.Yude Wang, Jie Zhang, Meina Kan, S. Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In CVPR, 2020. 1, 2, 4, 6, 7
  44. 44.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In CVPR, 2017. 2
  45. 45.Tong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei, Xiaolin Wei, Xuan Luo, and Chi Harold Liu. Embedded discriminative attention mechanism for weakly supervised semantic segmentation. In CVPR, 2021. 2, 4, 7
  46. 46.Jinheng Xie, Xianxu Hou, Kai Ye, and Linlin Shen. CLIMS: Cross language image matching for weakly supervised semantic segmentation. In CVPR, June 2022. 1, 3, 6, 7
  47. 47.Lian Xu, Wanli Ouyang, Bennamoun, Farid Boussaid, Ferdous Sohel, and Dan Xu. Leveraging auxiliary tasks with affinity learning for weakly supervised semantic segmentation. In ICCV, 2021. 2
  48. 48.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, and Dan Xu. Multi-class token transformer for weakly supervised semantic segmentation. In CVPR, 2022. 1, 2, 5, 6, 7
  49. 49.Yazhou Yao, Tao Chen, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu, Zhen min Tang, and Jian Zhang. Non-salient region object mining for weakly supervised semantic segmentation. In CVPR, 2021. 2, 7
  50. 50.Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang. Reliability does matter: An end-toend weakly supervised semantic segmentation approach. In AAAI, 2020. 2
  51. 51.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In CVPR, 2016. 2, 3
  52. 52.Tianfei Zhou, Meijie Zhang, Fang Zhao, and Jianwu Li. Regional semantic contrast and aggregation for weakly supervised semantic segmentation. In CVPR, 2022. 7

Citation

MLA
Lin, Y., et al. “CLIP Is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation”. arXiv, 2022, http://arxiv.org/abs/2212.09506v3.
APA
Lin, Y., Chen, M., Wang, W., Wu, B., Li, K., Lin, B., Liu, H., & He, X. (2022). CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation. arXiv. http://arxiv.org/abs/2212.09506v3
Chicago
Lin, Y., M. Chen, W. Wang, et al. 2022. “CLIP Is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation”. arXiv. http://arxiv.org/abs/2212.09506v3.
Harvard
Lin, Y. et al. (2022) “CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.09506v3.
Vancouver
1. Lin Y, Chen M, Wang W, Wu B, Li K, Lin B, Liu H, He X (2022) CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation. arXiv

BibTeX

@article{lin2022clip,
  title = {CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic Segmentation},
  author = {Lin, Yuqi and Chen, Minghao and Wang, Wenxiao and Wu, Boxi and Li, Ke and Lin, Binbin and Liu, Haifeng and He, Xiaofei},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.09506v3},
  eprint = {2212.09506}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE