Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation

Qi ChenLingxiao YangJianhuang LaiXiaohua Xie

article2022CVPR186 citations

Proposes a self-supervised framework that tailors image-specific prototypes and enforces general-specific consistency to overcome incomplete class activation maps, achieving state-of-the-art weakly supervised semantic segmentation using only image-level labels.

Listen

Training computer vision systems to precisely identify and outline objects at the pixel level typically requires large volumes of detailed manual annotations. Producing these annotations is expensive and time-consuming, creating a strong practical need for methods that rely only on simple image-level tags. However, traditional weakly supervised methods struggle because standard classification models focus only on the most distinctive visual cues, leaving object boundaries incomplete and missing key areas.

The article evaluates a framework designed to overcome this limitation by generating complete object and background localization maps using only image-level labels. The objective is to demonstrate that tailoring feature representations to individual images and enforcing consistency can significantly narrow the performance gap between weakly supervised and fully supervised segmentation systems.

The researchers conducted an experimental evaluation on two established computer vision benchmarks: PASCAL VOC 2012, encompassing over ten thousand augmented training images across 20 object classes, and MS COCO 2014, covering 80 complex object categories. The proposed method first discovers robust seed regions by matching structural pixel relationships to class activation patterns, then combines multi-level visual features to construct tailored prototypes for foreground objects and background scenes. A self-supervised training signal enforces consistency between general category weights and image-specific representations without requiring extra pixel annotations or auxiliary visual cues.

The experimental findings show substantial improvements over previous approaches. Initial localization maps achieved an intersection-over-union score of 58.6 percent on the primary benchmark, rising to 64.7 percent after standard boundary refinement, outperforming the previous top baseline of 56.6 percent. When these generated labels were used to train a standard segmentation network, the system achieved a 69.7 percent test score on PASCAL VOC 2012 and 43.6 percent on MS COCO 2014, establishing new state-of-the-art benchmarks for weakly supervised methods. Notably, the framework surpassed competing techniques that rely on additional external data or saliency inputs.

These results indicate that organizations deploying visual recognition models can achieve competitive segmentation accuracy while avoiding the high costs and operational delays associated with manual pixel-level data labeling. By successfully capturing full object areas and filtering background noise without human intervention, the approach reduces annotation budgets, shortens project timelines, and simplifies training pipelines across domains like automated inspection and visual analytics.

Organizations developing segmentation capabilities should consider adopting image-specific prototype generation and consistency training as standard practices in their weakly supervised data pipelines. Future efforts should evaluate deploying this architecture across specialized operational domains, such as medical diagnostics or aerial remote sensing, and explore adapting the framework to real-time inference constraints.

Confidence in these findings is strong given the rigorous testing across standard multi-object datasets and clear ablation analyses. However, readers should consider that evaluation was limited to standard benchmark datasets, and performance in specialized real-world settings with heavy clutter or poor lighting may require domain-specific parameter calibration.

Cover for Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation

Abstract

Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has attracted much attention due to low annotation costs. Existing methods often rely on Class Activation Mapping (CAM) that measures the correlation between image pixels and classifier weight. However, the classifier focuses only on the discriminative regions while ignoring other useful information in each image, resulting in incomplete localization maps. To address this issue, we propose a Self-supervised Image-specific Prototype Exploration (SIPE) that consists of an Image-specific Prototype Exploration (IPE) and a General-Specific Consistency (GSC) loss. Specifically, IPE tailors prototypes for every image to capture complete regions, formed our Image-Specific CAM (IS-CAM), which is realized by two sequential steps. In addition, GSC is proposed to construct the consistency of general CAM and our specific IS-CAM, which further optimizes the feature representation and empowers a self-correction ability of prototype exploration. Extensive experiments are conducted on PASCAL VOC 2012 and MS COCO 2014 segmentation benchmark and results show our SIPE achieves new state-of-the-art performance using only image-level labels. The code is available at https://github.com/chenqi1126/SIPE.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Approach
  • 3.1. Class Activation Mapping
  • 3.2. Image-specific Prototype Exploration
  • 3.3. Self-supervised Learning with GSC
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Comparison with State-of-the-arts
  • 4.3. Ablation Studies
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Self-Supervised Image-Specific Prototype Exploration (SIPE) Framework

    model/method

    The Self-supervised Image-specific Prototype Exploration (SIPE) framework is designed for weakly supervised semantic segmentation (WSSS) using only image-level labels. Standard Class Activation Mapping (CAM) computes activations by projecting pixel features onto fixed classifier weight vectors (class centers), which tends to highlight only the most discriminative regions of an object. SIPE overcomes this by constructing image-specific prototypes to capture complete object boundaries.

    SIPE consists of two core components:

    1. Image-Specific Prototype Exploration (IPE): Generates an Image-Specific CAM (IS-CAM) through a two-step procedure: (a) Structure-aware seed locating, which determines foreground and background seed regions by evaluating inter-pixel semantic correlation and matching spatial structures against initial CAM templates via Intersection-over-Union (IoU); and (b) Background-aware prototype modeling, which extracts hierarchical multi-scale feature representations to construct foreground and background prototypes as the spatial centroids of the located seeds, followed by prototypical cosine correlation to produce IS-CAM.

    2. General-Specific Consistency (GSC): Imposes a self-supervised L1L_1 consistency regularization between the general CAMs (activated by classifier weights) and the IS-CAMs (activated by image-specific prototypes). This mutual regularization forces the feature representations to incorporate missing object regions and suppresses noisy background activations.

  2. Knowl 2 — Structure-Aware Seed Locating

    model/method

    Structure-aware seed locating identifies reliable seed regions for each class without relying on fixed heuristic thresholds.

    Given the semantic feature map Fs∈RCs×H×WF_s \in \mathbb{R}^{C_s \times H \times W} from the backbone's final layer and the feature vector fi=Fs(i)f^i = F_s(i) at pixel location ii:

    1. The spatial structure map Si∈RH×WS^i \in \mathbb{R}^{H \times W} of pixel ii is obtained via cosine correlation with all spatial locations jj, suppressing negative values:

    Si(j)=ReLU(fi⋅Fs(j)∥fi∥2∥Fs(j)∥2)S^i(j) = \text{ReLU}\left( \frac{f^i \cdot F_s(j)}{\|f^i\|_2 \|F_s(j)\|_2} \right)

    1. The class-wise spatial structure similarity CkiC^i_k between the structure map SiS^i and the class activation map MkM_k of class kk is computed using IoU:

    Cki=∑jMk(j)Si(j)∑j[Mk(j)+Si(j)−Mk(j)Si(j)]C^i_k = \frac{\sum_j M_k(j) S^i(j)}{\sum_j \left[ M_k(j) + S^i(j) - M_k(j)S^i(j) \right]}

    where k∈{1,…,K,background}k \in \{1, \dots, K, \text{background}\} covers KK foreground classes and one background class.

    1. Pixel ii is assigned to the class that maximizes the structure similarity:

    Rki={1,if k=arg⁡max⁡k′Ck′i0,otherwiseR^i_k = \begin{cases} 1, & \text{if } k = \arg\max_{k'} C^i_{k'} \\ 0, & \text{otherwise} \end{cases}

    where Rk∈{0,1}H×WR_k \in \{0, 1\}^{H \times W} denotes the binary seed mask for class kk.

  3. Knowl 3 — Background-Aware Prototype Modeling and Image-Specific CAM

    model/method

    Because background regions lack distinct high-level category semantics, modeling background requires low-level visual features (such as color and texture) alongside semantic features.

    1. Hierarchical Feature Representation: Four convolutional layers process the feature outputs from four backbone stages (F1,F2,F3,F4F_1, F_2, F_3, F_4). These multi-scale representations are downsampled or upsampled to a common spatial resolution and concatenated to form a hierarchical feature map Fh∈RCh×H×WF_h \in \mathbb{R}^{C_h \times H \times W}.

    2. Prototype Modeling: For each class k∈{1,…,K}∪{background}k \in \{1, \dots, K\} \cup \{\text{background}\}, the image-specific prototype Pk∈RChP_k \in \mathbb{R}^{C_h} is calculated as the centroid of the seed region RkR_k in the hierarchical feature space:

    Pk=∑iFh(i)⋅1(Rki=1)∑i1(Rki=1)P_k = \frac{\sum_i F_h(i) \cdot \mathbb{1}(R^i_k = 1)}{\sum_i \mathbb{1}(R^i_k = 1)}

    where 1(⋅)\mathbb{1}(\cdot) is the indicator function.

    1. Image-Specific CAM (IS-CAM): The IS-CAM activation M~k(j)\tilde{M}_k(j) for class kk at pixel jj is calculated via cosine similarity between the hierarchical feature vector Fh(j)F_h(j) and prototype PkP_k:

    M~k(j)=ReLU(Fh(j)⋅Pk∥Fh(j)∥2∥Pk∥2)\tilde{M}_k(j) = \text{ReLU}\left( \frac{F_h(j) \cdot P_k}{\|F_h(j)\|_2 \|P_k\|_2} \right)

  4. Knowl 4 — SIPE Training Objective and General-Specific Consistency Loss

    equation

    The SIPE model is trained end-to-end under image-level supervision using the total loss:

    Ltotal=Lcls+Lgsc\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{cls}} + \mathcal{L}_{\text{gsc}}

    1. Multi-Label Classification Loss (Lcls\mathcal{L}_{\text{cls}}):

    Lcls=1K∑k=1K[yklog⁡σ(y^k)+(1−yk)log⁡(1−σ(y^k))]\mathcal{L}_{\text{cls}} = \frac{1}{K} \sum_{k=1}^K \left[ y_k \log \sigma(\hat{y}_k) + (1 - y_k) \log(1 - \sigma(\hat{y}_k)) \right]

    where yk∈{0,1}y_k \in \{0, 1\} is the image-level label for class kk, σ(⋅)\sigma(\cdot) is the sigmoid function, and y^k\hat{y}_k is the class score obtained by global average pooling over the foreground CAM Mk=ReLU(θkTFs)M_k = \text{ReLU}(\theta_k^T F_s), with classifier weights θk\theta_k and semantic feature FsF_s.

    1. Background CAM Estimation:

    The background map MbM_b is estimated from the foreground CAMs Mf={Mk}k=1KM_f = \{M_k\}_{k=1}^K using an attenuation coefficient α=0.5\alpha = 0.5:

    Mb=α(1−max⁡1≤k≤KMk)M_b = \alpha \left( 1 - \max_{1 \le k \le K} M_k \right)

    The complete general CAM set is M=Mf∪{Mb}M = M_f \cup \{M_b\}.

    1. General-Specific Consistency Loss (Lgsc\mathcal{L}_{\text{gsc}}):

    Lgsc=1K+1∥M−M~∥1=1(K+1)HW∑k=1K+1∑j∣Mk(j)−M~k(j)∣\mathcal{L}_{\text{gsc}} = \frac{1}{K + 1} \|M - \tilde{M}\|_1 = \frac{1}{(K + 1)HW} \sum_{k=1}^{K+1} \sum_{j} \left| M_k(j) - \tilde{M}_k(j) \right|

    where M~\tilde{M} represents the IS-CAMs generated from image-specific prototypes.

  5. Knowl 5 — Image-Specific CAM Generation and Pseudo-Label Extraction

    algorithm

    The procedure for generating Image-Specific CAMs (IS-CAM) and extracting pseudo segmentation labels from an input image is formalized as follows:

    Input: Input image II, trained backbone with classifier weights {θk}k=1K\{\theta_k\}_{k=1}^K, background attenuation factor α=0.5\alpha = 0.5
    Output: IS-CAM activation maps {M~k}k=1K+1\{\tilde{M}_k\}_{k=1}^{K+1} and pseudo-label map Y∈{1,…,K+1}H×WY \in \{1, \dots, K+1\}^{H \times W}
    1: Extract semantic feature FsF_s and hierarchical fused feature FhF_h from II
    2: for k←1k \leftarrow 1 to KK do
    3: Mk←ReLU(θkTFs)M_k \leftarrow \text{ReLU}(\theta_k^T F_s)
    4: Mk←Mk/max⁡u,v(Mk(u,v))M_k \leftarrow M_k / \max_{u, v}(M_k(u, v))
    5: end for
    6: MK+1←α(1−max⁡1≤k≤KMk)M_{K+1} \leftarrow \alpha (1 - \max_{1 \le k \le K} M_k)
    7: for each spatial location ii do
    8: Compute pixel correlation structure: Si(j)←ReLU(Fs(i)⋅Fs(j)∥Fs(i)∥2∥Fs(j)∥2)S^i(j) \leftarrow \text{ReLU}\left(\frac{F_s(i) \cdot F_s(j)}{\|F_s(i)\|_2 \|F_s(j)\|_2}\right) for all jj
    9: for k←1k \leftarrow 1 to K+1K+1 do
    10: Cki←∑jMk(j)Si(j)∑j[Mk(j)+Si(j)−Mk(j)Si(j)]C^i_k \leftarrow \frac{\sum_j M_k(j) S^i(j)}{\sum_j [M_k(j) + S^i(j) - M_k(j) S^i(j)]}
    11: end for
    12: Ri←arg⁡max⁡kCkiR^i \leftarrow \arg\max_k C^i_k
    13: end for
    14: for k←1k \leftarrow 1 to K+1K+1 do
    15: Pk←∑iFh(i)⋅1(Ri=k)∑i1(Ri=k)P_k \leftarrow \frac{\sum_i F_h(i) \cdot \mathbb{1}(R^i = k)}{\sum_i \mathbb{1}(R^i = k)}
    16: for each spatial location jj do
    17: M~k(j)←ReLU(Fh(j)⋅Pk∥Fh(j)∥2∥Pk∥2)\tilde{M}_k(j) \leftarrow \text{ReLU}\left(\frac{F_h(j) \cdot P_k}{\|F_h(j)\|_2 \|P_k\|_2}\right)
    18: end for
    19: end for
    20: for each spatial location jj do
    21: Y(j)←arg⁡max⁡kM~k(j)Y(j) \leftarrow \arg\max_k \tilde{M}_k(j)
    22: end for
    23: return {M~k}k=1K+1,Y\{\tilde{M}_k\}_{k=1}^{K+1}, Y
  6. Knowl 6 — Localization Map Quality on PASCAL VOC 2012

    data/table

    The quality of localization maps generated by SIPE on the PASCAL VOC 2012 train set was evaluated in terms of mean Intersection-over-Union (mIoU, %), both directly on initial maps and after DenseCRF refinement:

    Could not parse LaTeX table

    SIPE achieves 58.6% mIoU without post-processing (+2.0% over ECS) and 64.7% mIoU with DenseCRF (+1.9% over CSE). The more complete object coverage and cleaner background boundaries in SIPE provide higher-quality inputs for DenseCRF post-processing.

  7. Knowl 7 — Weakly Supervised Semantic Segmentation on PASCAL VOC 2012 Benchmark

    data/table

    DeepLabV2 was trained using pseudo-labels refined with Inter-pixel Relation Network (IRN) derived from SIPE localization maps on PASCAL VOC 2012. Performance (mIoU, %) is compared against weakly supervised methods using image-level labels (I) and methods using auxiliary saliency supervision (I + S):

    Could not parse LaTeX table

    Using only image-level supervision (I), SIPE achieves 68.8% val mIoU and 69.7% test mIoU with ResNet101, outperforming all prior image-level supervised methods and approaching methods that require additional saliency maps (I + S).

  8. Knowl 8 — Semantic Segmentation Performance on MS COCO 2014

    data/table

    Performance of SIPE on the MS COCO 2014 validation set (81 classes, 40k images) without IRN refinement:

    Could not parse LaTeX table

    SIPE achieves 43.6% mIoU with ResNet38 and 40.6% with ResNet101, surpassing the previous state-of-the-art image-level method CSE (36.4%) by 7.2% mIoU and exceeding methods that rely on external saliency supervision (such as EPS at 35.7%).

  9. Knowl 9 — Ablation Study on SIPE Components

    empirical result

    An ablation study evaluating the individual contributions of Image-specific Prototype Exploration (IPE) and General-Specific Consistency (GSC) on PASCAL VOC 2012 localization map mIoU (%):

    • Baseline CAM: 50.1% mIoU.
    • CAM + IPE: 53.2% mIoU (+3.1% over baseline). Generating IS-CAM directly from image-specific prototypes activates broader and more complete object regions than standard classifier weights.
    • CAM + IPE + GSC (Full SIPE): 58.6% mIoU (+5.4% over IPE alone, +8.5% over baseline). The GSC loss forces mutual alignment between CAM and IS-CAM, optimizing the underlying feature representation and enhancing foreground/background separation.
  10. Knowl 10 — Ablation on Seed Locating: Structure-Aware Locating vs. Heuristics

    empirical result

    Comparing different seed generation strategies for prototype extraction on the PASCAL VOC 2012 train set:

    1. Fixed Thresholding on CAM: Thresholds tested from 0.10 to 0.50 (in increments of 0.05) produce mIoU scores between 52.1% (at 0.25) and 53.3% (at 0.45). Fixed thresholds cannot adapt to varying object scales and scene complexities across images.
    2. Argmax on CAM: Assigning seed class by pixel-wise argmax over CAM channels yields 53.7% mIoU (+0.4% over the best fixed threshold).
    3. Structure-Aware Seed Locating: Evaluating inter-pixel correlation structure maps and matching them to CAM templates via IoU achieves 58.6% mIoU (+4.9% over argmax, +5.3% over optimal thresholding). Exploiting spatial contextual correlations captures coherent semantic regions that single-pixel probability methods overlook.
  11. Knowl 11 — Ablation on Feature Selection and Background Prototype Modeling

    data/table

    Ablation on feature levels (semantic vs. multi-scale hierarchical) and Background Prototype Modeling (BPM) trained with GSC on PASCAL VOC 2012, comparing threshold-searched pseudo-labels against estimated background map pseudo-labels:

    Could not parse LaTeX table

    Key findings:

    1. With semantic features alone, adding BPM does not improve performance (56.0% vs. 55.9% with threshold) because high-level features lack generic background descriptors.
    2. Using hierarchical features without BPM degrades performance (53.9% vs. 56.0%) because shallow features introduce noise to foreground localization.
    3. Combining hierarchical features with BPM achieves the best result (58.6%), allowing direct background map estimation to outperform empirical threshold searching (58.6% vs. 57.6%).

Coverage note — Qualitative segmentation visualizations and standard DeepLabV2 retraining protocols adopted directly from prior work were omitted as they do not constitute novel contributed findings.

References

  1. 1.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2209–2218, 2019.
  2. 2.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4981–4990, 2018.
  3. 3.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What’s the point: Semantic segmentation with point supervision. In European Conference on Computer Vision, pages 549–565, 2016.
  4. 4.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation via sub-category exploration. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 8991–9000, 2020.
  5. 5.Hongjun Chen, Jinbao Wang, Hong Cai Chen, Xiantong Zhen, Feng Zheng, Rongrong Ji, and Ling Shao. Seminar learning for click-level weakly supervised semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 6920–6929, October 2021.
  6. 6.Liyi Chen, Weiwei Wu, Chenchen Fu, Xiao Han, and Yuntao Zhang. Weakly supervised semantic segmentation with boundary exploration. In European Conference on Computer Vision, pages 347–362, 2020.
  7. 7.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4):834–848, 2017.
  8. 8.Junsuk Choe, Seungho Lee, and Hyunjung Shim. Attention-based dropout layer for weakly supervised single object localization and semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  10. 10.Zhang Dong, Zhang Hanwang, Tang Jinhui, Hua Xiansheng, and Sun Qianru. Causal intervention for weakly supervised semantic segmentation. In Advances in Neural Information Processing Systems, 2020.
  11. 11.Mark Everingham, SM Ali Eslami, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes challenge: A retrospective. International Journal of Computer Vision, 111(1):98–136, 2015.
  12. 12.Junsong Fan, Zhaoxiang Zhang, Chunfeng Song, and Tieniu Tan. Learning integral objects with intra-class discriminator for weakly-supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4283–4292, 2020.
  13. 13.Junsong Fan, Zhaoxiang Zhang, Tieniu Tan, Chunfeng Song, and Jun Xiao. Cian: Cross-image affinity net for weakly supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 10762–10769, 2020.
  14. 14.Di Feng, Christian Haase-Schuetz, Lars Rosenbaum, Heinz Hertlein, Claudius Glaeser, Fabian Timm, Werner Wiesbeck, and Klaus Dietmayer. Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges. IEEE Transactions on Intelligent Transportation Systems, 2020.
  15. 15.Bharath Hariharan, Pablo Arbelaez, Lubomir Bourdev, Subhransu Maji, and Jitendra Malik. Semantic contours from inverse detectors. In Proceedings of the IEEE International Conference on Computer Vision, pages 991–998, 2011.
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016.
  17. 17.Mohammad D Hossain and Dongmei Chen. Segmentation for object-based image analysis (obia): A review of algorithms and challenges from remote sensing perspective. ISPRS Journal of Photogrammetry and Remote Sensing, 150:115–134, 2019.
  18. 18.Qibin Hou, Peng-Tao Jiang, Yunchao Wei, and Ming-Ming Cheng. Self-erasing network for integral object attention. In Advances in Neural Information Processing Systems, pages 547–557, 2018.
  19. 19.Zilong Huang, Xinggang Wang, Jiasi Wang, Wenyu Liu, and Jingdong Wang. Weakly-supervised semantic segmentation network with deep seeded region growing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7014–7023, 2018.
  20. 20.Peng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng, Yunchao Wei, and Hong-Kai Xiong. Integral object mining via online attention accumulation. In Proceedings of the IEEE International Conference on Computer Vision, pages 2070–2079, 2019.
  21. 21.Beomyoung Kim, Sangeun Han, and Junmo Kim. Discriminative region suppression for weakly-supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1754–1761, 2021.
  22. 22.Alexander Kolesnikov and Christoph H. Lampert. Seed, expand and constrain: Three principles for weakly-supervised image segmentation. In European Conference on Computer Vision, 2016.
  23. 23.Hyeokjun Kweon, Sung-Hoon Yoon, Hyeonseong Kim, Daehee Park, and Kuk-Jin Yoon. Unlocking the potential of ordinary classifier: Class-specific adversarial erasing framework for weakly supervised semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, pages 6994–7003, October 2021.
  24. 24.Jungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee, and Sungroh Yoon. Ficklenet: Weakly and semi-supervised semantic image segmentation using stochastic inference. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5267–5276, 2019.
  25. 25.Jungbeom Lee, Eunji Kim, and Sungroh Yoon. Anti-adversarially manipulated attributions for weakly and semi-supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4071–4080, June 2021.
  26. 26.Jungbeom Lee, Jihun Yi, Chaehun Shin, and Sungroh Yoon. Bbam: Bounding box attribution map for weakly supervised semantic and instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2643–2652, 2021.
  27. 27.Seungho Lee, Minhyun Lee, Jongwuk Lee, and Hyunjung Shim. Railroad is not a train: Saliency as pseudo-pixel supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5495–5505, June 2021.
  28. 28.Xueyi Li, Tianfei Zhou, Jianwu Li, Yi Zhou, and Zhaoxiang Zhang. Group-wise semantic mining for weakly supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 1984–1992, May 2021.
  29. 29.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3159–3167, 2016.
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755, 2014.
  31. 31.Weide Liu, Chi Zhang, Guosheng Lin, Tzu-Yi Hung, and Chunyan Miao. Weakly supervised segmentation with maximum bipartite graph matching. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2085–2094, 2020.
  32. 32.Yun Liu, Yu-Huan Wu, Pei-Song Wen, Yu-Jun Shi, Yu Qiu, and Ming-Ming Cheng. Leveraging instance-, image- and dataset-level information for weakly supervised instance segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  33. 33.Yongfei Liu, Xiangyi Zhang, Songyang Zhang, and Xuming He. Part-aware prototype network for few-shot semantic segmentation. In European Conference on Computer Vision, pages 142–158. Springer, 2020.
  34. 34.Zhiyi Pan, Peng Jiang, Yunhai Wang, Changhe Tu, and Anthony G. Cohn. Scribble-supervised semantic segmentation by uncertainty reduction on neural representation and self-supervision on neural eigenspace. In Proceedings of the IEEE International Conference on Computer Vision, pages 7416–7425, October 2021.
  35. 35.Yukun Su, Ruizhou Sun, Guosheng Lin, and Qingyao Wu. Context decoupling augmentation for weakly supervised semantic segmentation. 2021.
  36. 36.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In European Conference on Computer Vision, pages 347–365, 2020.
  37. 37.Kunyang Sun, Haoqing Shi, Zhengming Zhang, and Yongming Huang. Ecs-net: Improving weakly supervised semantic segmentation by using connections between class activation maps. In Proceedings of the IEEE International Conference on Computer Vision, pages 7283–7292, October 2021.
  38. 38.Nima Tajbakhsh, Laura Jeyaseelan, Qian Li, Jeffrey N Chiang, Zhihao Wu, and Xiaowei Ding. Embracing imperfect datasets: A review of deep learning solutions for medical image segmentation. Medical Image Analysis, 63:101693, 2020.
  39. 39.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  40. 40.Kaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou, and Jiashi Feng. Panet: Few-shot image semantic segmentation with prototype alignment. In Proceedings of the IEEE International Conference on Computer Vision, October 2019.
  41. 41.Xiang Wang, Sifei Liu, Huimin Ma, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation by iterative affinity learning. International Journal of Computer Vision, 128(6):1736–1749, 2020.
  42. 42.Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 12275–12284, 2020.
  43. 43.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1568–1576, 2017.
  44. 44.Yunchao Wei, Huaxin Xiao, Honghui Shi, Zequn Jie, Jiashi Feng, and Thomas S Huang. Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7268–7277, 2018.
  45. 45.Tong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei, Xiaolin Wei, Xuan Luo, and Chi Harold Liu. Embedded discriminative attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 16765–16774, June 2021.
  46. 46.Jingshan Xu, Chuanwei Zhou, Zhen Cui, Chunyan Xu, Yuge Huang, Pengcheng Shen, Shaoxin Li, and Jian Yang. Scribble-supervised semantic segmentation inference. In Proceedings of the IEEE International Conference on Computer Vision, pages 15354–15363, October 2021.
  47. 47.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, Ferdous Sohel, and Dan Xu. Leveraging auxiliary tasks with affinity learning for weakly supervised semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, 2021.
  48. 48.Yazhou Yao, Tao Chen, Guo-Sen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu, Zhenmin Tang, and Jian Zhang. Non-salient region object mining for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2623–2632, June 2021.
  49. 49.Bingfeng Zhang, Jimin Xiao, Jianbo Jiao, Yunchao Wei, and Yao Zhao. Affinity attention graph neural network for weakly supervised semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  50. 50.Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang. Reliability does matter: An end-to-end weakly supervised semantic segmentation approach. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 12765–12772, 2020.
  51. 51.Fei Zhang, Chaochen Gu, Chenyue Zhang, and Yuchao Dai. Complementary patch for weakly supervised semantic segmentation. In Proceedings of the IEEE International Conference on Computer Vision, 2021.
  52. 52.Xiaolin Zhang, Yunchao Wei, Yi Yang, and Thomas S Huang. Sg-one: Similarity guidance network for one-shot semantic segmentation. IEEE transactions on cybernetics, 50(9):3855–3865, 2020.
  53. 53.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2921–2929, June 2016.

Citation

MLA
Chen, Q., et al. “Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 4278–88, https://doi.org/10.1109/CVPR52688.2022.00425.
APA
Chen, Q., Yang, L., Lai, J., & Xie, X. (2022). Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4278–4288. https://doi.org/10.1109/CVPR52688.2022.00425
Chicago
Chen, Q., L. Yang, J. Lai, and X. Xie. 2022. “Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation”. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4278–88. https://doi.org/10.1109/CVPR52688.2022.00425.
Harvard
Chen, Q. et al. (2022) “Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 4278–4288. Available at: https://doi.org/10.1109/CVPR52688.2022.00425.
Vancouver
1. Chen Q, Yang L, Lai J, Xie X (2022) Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation. In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 4278–4288

BibTeX

@inproceedings{Chen_2022, title={Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation}, url={http://dx.doi.org/10.1109/CVPR52688.2022.00425}, DOI={10.1109/cvpr52688.2022.00425}, booktitle={2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Chen, Qi and Yang, Lingxiao and Lai, Jianhuang and Xie, Xiaohua}, year={2022}, month=June, pages={4278–4288} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE