Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor

Hyeokjun KweonSung-Hoon YoonKuk-Jin Yoon

article2023CVPR81 citations

Proposes an adversarial framework pitting a CAM-generating classifier against an image reconstructor to prevent over-erasing and under-activation by minimizing cross-segment inferability, achieving state-of-the-art weakly supervised semantic segmentation performance on PASCAL VOC and MS COCO.

Listen

Training computer vision systems to identify and outline objects at the pixel level typically requires expensive and time-consuming manual annotations. Weakly supervised semantic segmentation addresses this bottleneck by using simple, low-cost image-level tags that only indicate whether an object category exists in an image. However, standard methods relying on class activation maps consistently fail to capture entire objects and frequently bleed into irrelevant backgrounds, creating noisy training data that degrades segmentation quality.

The article demonstrates a novel framework called Adversarial learning of the Classifier and the Reconstructor to resolve these localization errors. The main objective is to significantly enhance the precision and completeness of object localization maps without relying on pixel-level annotations or auxiliary datasets.

The authors approach the problem through the concept of mutual segment independence: if an image is cleanly divided into target and non-target segments, neither piece should contain enough residual information to infer the appearance of the other. The method pairs a classifier model, which generates segmentation maps, against an image reconstructor model in a two-player competitive learning setup. The reconstructor attempts to recreate the missing target segment using only the non-target background, while the classifier learns to produce clean boundaries that starve the reconstructor of helpful leftover visual clues. To prevent the reconstructor from simply memorizing training images, the authors introduce a synthetic noise strategy that injects controlled remnants during training. The framework was evaluated across standard benchmarks including PASCAL VOC 2012 (21 categories) and MS COCO 2014 (81 categories).

The evaluation yielded several critical findings. First, the proposed framework improved initial localization quality from a baseline of 48.4% mean Intersection over Union to 60.3% on the PASCAL VOC training set, representing a nearly 12-percentage-point jump. Second, incorporating both target and non-target adversarial loss terms proved essential; omitting either term caused severe under-segmentation or over-expansion. Third, the resulting segmentation models established new benchmark records, achieving 71.9% validation accuracy on PASCAL VOC and 45.3% on MS COCO when using convolutional backbones, and reaching 72.4% accuracy when paired with vision transformer architectures.

These findings demonstrate that image reconstruction can serve as an effective, self-supervised regularizer for weakly supervised visual learning. By outperforming existing adversarial erasing methods, the approach avoids the common pitfall of over-expanding object boundaries while completely removing the need for external saliency models or costly manual pixel annotations. This translates to substantial labor and cost savings when deploying high-accuracy segmentation tools in real-world computer vision pipelines.

Organizations developing segmentation systems should adopt this adversarial reconstruction framework when only category-level image labels are available. Because the methodology is compatible with both convolutional and transformer backbones, engineering teams can integrate it directly into existing training workflows. When deploying this training scheme, teams should include the synthetic remnant feeding mechanism to guard against model memorization.

The study notes certain operational boundaries. Reconstructor training is sensitive to specific data augmentations, requiring color jittering to be omitted to ensure stable model convergence. Furthermore, while the method operates efficiently on standard hardware—training in approximately 12 hours on a single commercial graphics processor—its performance has been primarily validated on natural object benchmarks, meaning performance on specialized domains such as medical imaging or industrial inspection warrants further verification.

No sufficiently relevant recommendations were found.

Cover for Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor

Abstract

In Weakly Supervised Semantic Segmentation (WSSS), Class Activation Maps (CAMs) usually 1) do not cover the whole object and 2) be activated on irrelevant regions. To address the issues, we propose a novel WSSS framework via adversarial learning of a classifier and an image reconstructor. When an image is perfectly decomposed into class-wise segments, information (i.e., color or texture) of a single segment could not be inferred from the other segments. Therefore, inferability between the segments can represent the preciseness of segmentation. We quantify the inferability as a reconstruction quality of one segment from the other segments. If one segment could be reconstructed from the others, then the segment would be imprecise. To bring this idea into WSSS, we simultaneously train two models: a classifier generating CAMs that decompose an image into segments and a reconstructor that measures the inferability between the segments. As in GANs, while being alternatively trained in an adversarial manner, two networks provide positive feedback to each other. We verify the superiority of the proposed framework with extensive ablation studies. Our method achieves new state-of-the-art performances on both PASCAL VOC 2012 and MS COCO 2014. The code is available at https://github.com/sangrockEG/ACR.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. CAMs Improvements
  • 2.2. Mask Refinements
  • 3. Motivation
  • 4. Methods
  • 4.1. Overall Framework
  • 4.2. Reconstructor-Updating Phase
  • 4.2.1. Basic Formulation
  • 4.2.2. Stochastic Remnant Feeding
  • 4.3. Classifier-Updating Phase
  • 5. Experimental Results
  • 5.1. Dataset and Evaluation Metric
  • 5.2. Implementation Details
  • 5.3. Ablation Studies
  • 5.4. Comparisons to State-of-The-Arts
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Adversarial Classifier-Reconstructor Framework for Weakly Supervised Semantic Segmentation

    model/method

    The Adversarial Classifier-Reconstructor (ACR) framework addresses Weakly Supervised Semantic Segmentation (WSSS) under image-level supervision by exploiting the concept of segment inferability. If an image is accurately segmented into distinct class regions, the color and texture information of one segment cannot be inferred from the remaining segments (low inferability). Conversely, if segmentation is imprecise, leftover regions (remnants) leak visual clues that enable reconstructing one segment from another.

    ACR operationalizes this insight by formulating a two-player min-max game between an image classifier FF (which generates Class Activation Maps, CAMs) and an image reconstructor G=(GE,GD)G = (G_E, G_D) composed of a feature encoder GEG_E and a decoder GDG_D:

    1. For an input image I∈[0,1]3×H×WI \in [0, 1]^{3 \times H \times W}, the classifier produces CAMs A∈[0,1]C×h×wA \in [0, 1]^{C \times h \times w} and a multi-label class prediction vector p∈[0,1]Cp \in [0, 1]^C, while the encoder extracts a dense feature map X=GE(I)∈Rd×h×wX = G_E(I) \in \mathbb{R}^{d \times h \times w}: A,p=F(I),X=GE(I)A, p = F(I), \quad X = G_E(I)

    2. A target class tt present in the image is sampled, and its corresponding CAM At∈[0,1]h×wA_t \in [0, 1]^{h \times w} serves to decompose the feature XX into a target segment XtX_t and a non-target segment XntX_{nt} via element-wise multiplication ⊙\odot: Xt=At⊙X,Xnt=(1−At)⊙XX_t = A_t \odot X, \quad X_{nt} = (1 - A_t) \odot X

    3. The reconstructor GG is optimized to accurately reconstruct the target image region from XntX_{nt} and the non-target region from XtX_t by exploiting segmentation remnants. Concurrently, the classifier FF is optimized to adjust AtA_t so as to eliminate remnants and maximize the reconstructor's error. This adversarial game drives the CAMs toward precise object boundaries without suffering from the unbounded over-erasing common in adversarial erasing methods.

  2. Knowl 2 — Reconstructor-Updating Phase with Stochastic Remnant Feeding

    model/method

    In the Reconstructor-Updating (RU) phase of the ACR framework, the reconstructor parameters (GE,GDG_E, G_D) are optimized while the classifier FF is kept frozen.

    To prevent the reconstructor from trivial memorization (i.e., generating the original training image from global context rather than relying on segmentation remnants), ACR introduces Stochastic Remnant Feeding (SRF). A random binary grid g∈{0,1}h×wg \in \{0, 1\}^{h \times w} composed of s×ss \times s spatial patches is generated by sampling each patch independently from a Bernoulli distribution with probability qq. The target and non-target features are then augmented with synthetic remnants: XtRU=Xt+g⊙XntX_t^{RU} = X_t + g \odot X_{nt} XntRU=Xnt+g⊙XtX_{nt}^{RU} = X_{nt} + g \odot X_t

    Reconstructions are produced by the decoder GDG_D: I^tRU=GD(XtRU),I^ntRU=GD(XntRU)\hat{I}_t^{RU} = G_D(X_t^{RU}), \quad \hat{I}_{nt}^{RU} = G_D(X_{nt}^{RU})

    To validate cross-reconstruction outside the injected synthetic remnants, the soft validation masks are defined as: VtRU=(1−At)⊙(1−g)V_t^{RU} = (1 - A_t) \odot (1 - g) VntRU=At⊙(1−g)V_{nt}^{RU} = A_t \odot (1 - g)

    The reconstructor is trained to minimize the masked L1L_1 reconstruction loss: LtRU=∣VtRU⊙(I−I^tRU)∣1\mathcal{L}_t^{RU} = | V_t^{RU} \odot (I - \hat{I}_t^{RU}) |_1 LntRU=∣VntRU⊙(I−I^ntRU)∣1\mathcal{L}_{nt}^{RU} = | V_{nt}^{RU} \odot (I - \hat{I}_{nt}^{RU}) |_1 LRU=λtRULtRU+λntRULntRU\mathcal{L}^{RU} = \lambda_t^{RU} \mathcal{L}_t^{RU} + \lambda_{nt}^{RU} \mathcal{L}_{nt}^{RU}

    where λtRU\lambda_t^{RU} and λntRU\lambda_{nt}^{RU} are balancing hyperparameters set to 0.50.5.

  3. Knowl 3 — Classifier-Updating Phase with Reconstruction Adversarial Loss

    model/method

    In the Classifier-Updating (CU) phase of the ACR framework, only the classifier FF is updated while the reconstructor (GE,GDG_E, G_D) remains frozen. Stochastic Remnant Feeding is disabled in this phase.

    Reconstructions are obtained directly from the decomposed segment features Xt=At⊙XX_t = A_t \odot X and Xnt=(1−At)⊙XX_{nt} = (1 - A_t) \odot X: I^tCU=GD(Xt),I^ntCU=GD(Xnt)\hat{I}_t^{CU} = G_D(X_t), \quad \hat{I}_{nt}^{CU} = G_D(X_{nt})

    The classifier is trained to adjust the target CAM AtA_t to spoil the reconstructor's cross-reconstruction capability, maximizing the difference between the reconstructed output and the input image II on the opposing segment: LtCU=−∣VtCU⊙(I−I^tCU)∣1,where VtCU=1−At\mathcal{L}_t^{CU} = - | V_t^{CU} \odot (I - \hat{I}_t^{CU}) |_1, \quad \text{where } V_t^{CU} = 1 - A_t LntCU=−∣VntCU⊙(I−I^ntCU)∣1,where VntCU=At\mathcal{L}_{nt}^{CU} = - | V_{nt}^{CU} \odot (I - \hat{I}_{nt}^{CU}) |_1, \quad \text{where } V_{nt}^{CU} = A_t

    Combined with the multi-label binary cross-entropy classification loss LclsCU\mathcal{L}_{cls}^{CU} computed between the predicted class logits pp and the ground-truth image-level labels yy, the total classifier loss is: LCU=LclsCU+λtCULtCU+λntCULntCU\mathcal{L}^{CU} = \mathcal{L}_{cls}^{CU} + \lambda_t^{CU} \mathcal{L}_t^{CU} + \lambda_{nt}^{CU} \mathcal{L}_{nt}^{CU}

    where λtCU=0.8\lambda_t^{CU} = 0.8 and λntCU=0.3\lambda_{nt}^{CU} = 0.3. The negative sign ensures that minimizing LCU\mathcal{L}^{CU} maximizes the reconstruction discrepancy. Minimizing LtCU\mathcal{L}_t^{CU} penalizes false-positive activations (increasing precision), while minimizing LntCU\mathcal{L}_{nt}^{CU} penalizes missed object regions (increasing recall).

  4. Knowl 4 — Alternating Adversarial Optimization of Classifier and Reconstructor

    algorithm

    The ACR framework alternates between updating the reconstructor parameters θG=(θGE,θGD)\theta_G = (\theta_{G_E}, \theta_{G_D}) and the classifier parameters θF\theta_F in each iteration over training mini-batches.

    Input: Training images and image-level labels {I, y}, learning rates eta_F, eta_G, patch size s, probability q, loss weights lambda_t^RU, lambda_nt^RU, lambda_t^CU, lambda_nt^CU
    Output: Trained classifier F with parameters theta_F
    for each epoch do
        for each mini-batch (I, y) do
            Forward pass: A, p = F(I) and X = G_E(I)
            Sample target class t from classes present in y
            Extract target CAM A_t
            Compute features X_t = A_t * X and X_nt = (1 - A_t) * X
            // Reconstructor-Updating (RU) Phase
            Sample binary patch grid g from Bernoulli(q) with s x s patches
            Compute X_t^RU = X_t + g * X_nt and X_nt^RU = X_nt + g * X_t
            Compute reconstructions I_hat_t^RU = G_D(X_t^RU) and I_hat_nt^RU = G_D(X_nt^RU)
            Compute validation masks V_t^RU = (1 - A_t) * (1 - g) and V_nt^RU = A_t * (1 - g)
            Compute L^RU = lambda_t^RU * |V_t^RU * (I - I_hat_t^RU)|_1 + lambda_nt^RU * |V_nt^RU * (I - I_hat_nt^RU)|_1
            Update reconstructor: theta_G = theta_G - eta_G * grad_{theta_G}(L^RU)
            // Classifier-Updating (CU) Phase
            Compute reconstructions I_hat_t^CU = G_D(X_t) and I_hat_nt^CU = G_D(X_nt)
            Compute L_t^CU = - |(1 - A_t) * (I - I_hat_t^CU)|_1
            Compute L_nt^CU = - |A_t * (I - I_hat_nt^CU)|_1
            Compute classification loss L_cls^CU = BCE(p, y)
            Compute L^CU = L_cls^CU + lambda_t^CU * L_t^CU + lambda_nt^CU * L_nt^CU
            Update classifier: theta_F = theta_F - eta_F * grad_{theta_F}(L^CU)
        end for
    end for
    return F
  5. Knowl 5 — Experimental Implementation Configuration of ACR

    experimental setup

    The ACR framework is implemented and evaluated with the following specifications:

    • Classifier Architecture: Wide ResNet38 (WRN38) backbone with an attached 1×11 \times 1 convolution classification head to generate CAMs. For Vision Transformer experiments, a ViT backbone with patch attention refinement (as in MCTformer-V2) is used.
    • Reconstructor Architecture: UNet-based network. The encoder GEG_E aggregates multi-scale feature representations from multiple layers to supply low-level primitive details directly to the decoder GDG_D.
    • Segmentation Network: DeepLab architecture with ResNet38 backbone, trained on pseudo-labels generated by refining ACR CAMs using Inter-pixel Relation Network (IRNet).
    • Data Augmentations: Random cropping with a crop size of 256, random resizing in the scale range [0.5,1.3][0.5, 1.3], and random horizontal flipping. Color jittering is excluded due to observed reconstructor training instability.
    • Optimization Policy: Polynomial learning rate decay schedule multiplying the initial learning rate by (1−itermax_iter)0.9(1 - \frac{\text{iter}}{\text{max\_iter}})^{0.9}. Initial learning rate is 0.010.01 for the ACR classifier and reconstructor, and 0.0010.001 for DeepLab. The framework is trained for 40 epochs (~12 hours on a single NVIDIA RTX 3090 Ti GPU).
    • Loss Weight Parameters: λtRU=0.5\lambda_t^{RU} = 0.5, λntRU=0.5\lambda_{nt}^{RU} = 0.5, λtCU=0.8\lambda_t^{CU} = 0.8, and λntCU=0.3\lambda_{nt}^{CU} = 0.3.
  6. Knowl 6 — Ablation Study on Adversarial Learning and Stochastic Remnant Feeding

    data/table

    An ablation study evaluated the impact of adversarial co-training versus pre-training the reconstructor, as well as the necessity of Stochastic Remnant Feeding (SRF), measured by CAM mean Intersection over Union (mIoU) on the PASCAL VOC 2012 training set:

    Learning strategy for reconstructor SRF mIoU (%)
    Baseline (Classification only) 48.4
    Pre-trained 52.9
    Pre-trained ✓ 54.6
    Adversarial 55.8
    Adversarial (ACR full) ✓ 60.3

    Fixing a pre-trained reconstructor confirms the core hypothesis by improving CAM mIoU from 48.4% to 54.6% (with SRF). However, joint adversarial training allows both networks to provide continuous positive feedback to each other, achieving 60.3% mIoU. SRF provides consistent gains (+1.7% in the pre-trained setting, +4.5% in the adversarial setting) by regularizing against reconstructor memorization.

  7. Knowl 7 — Ablation Study on Classifier Loss Components

    data/table

    An ablation study evaluated the individual contributions of the target reconstruction loss LtCU\mathcal{L}_t^{CU} and non-target reconstruction loss LntCU\mathcal{L}_{nt}^{CU} in the classifier objective on the PASCAL VOC 2012 training set:

    Setting LclsCU\mathcal{L}_{cls}^{CU} LtCU\mathcal{L}_t^{CU} LntCU\mathcal{L}_{nt}^{CU} Precision Recall mIoU (%)
    Baseline ✓ 0.61 0.72 48.4
    (a) Target only ✓ ✓ 0.73 (+0.12) 0.68 (-0.04) 53.4 (+5.0)
    (b) Non-target only ✓ ✓ 0.58 (-0.03) 0.77 (+0.05) 51.5 (+3.1)
    (c) Both (ACR) ✓ ✓ ✓ 0.75 (+0.14) 0.76 (+0.04) 60.3 (+13.6)

    When optimizing only with the target loss LtCU\mathcal{L}_t^{CU} (setting a), the classifier maximizes error in non-target regions by shrinking the CAM, resulting in high precision (0.73) but reduced recall (0.68). Conversely, optimizing only with the non-target loss LntCU\mathcal{L}_{nt}^{CU} (setting b) causes the CAM to over-expand to spoil target reconstruction, yielding high recall (0.77) but poor precision (0.58). Combining both terms (setting c) balances precision and recall simultaneously, yielding the highest mIoU of 60.3%.

  8. Knowl 8 — CAM and Pseudo-Mask Quality Benchmarks on PASCAL VOC 2012

    data/table

    Evaluation of initial CAM seeds, DenseCRF-refined CAMs, and generated pseudo-segmentation masks (via IRNet) on the PASCAL VOC 2012 train set across Weakly Supervised Semantic Segmentation methods:

    Methods Backbone Seed mIoU (%) w/ CRF mIoU (%) Mask mIoU (%)
    CONTA WRN38 56.2 65.4 66.1
    EDAM WRN38 52.8 58.2 68.1
    AdvCAM RN50 55.6 62.1 68.0
    ECS WRN38 56.6 58.6 -
    OC-CSE WRN38 56.0 62.8 66.9
    CDA WRN38 58.4 - 66.4
    PMM WRN38 58.2 61.5 61.0
    RIB RN50 56.5 62.9 70.6
    AMR RN50 56.8 - 69.7
    ReCAM RN50 54.8 - 70.5
    SIPE RN50 58.6 64.7 -
    CLIMS WRN38 56.6 - 70.5
    W-OoD RN50 53.3 58.4 -
    PPC WRN38 61.5 64.0 70.1
    AEFT WRN38 56.0 63.5 71.0
    Ours (ACR) WRN38 60.3 65.9 72.3
    MCT ViT 61.7 - 69.1
    Ours (ACR + ViT) ViT 65.5 - 70.9

    ACR with Wide ResNet38 achieves a mask mIoU of 72.3%, outperforming all prior ResNet-based methods. When integrated with a Vision Transformer backbone, ACR achieves 65.5% seed mIoU and 70.9% mask mIoU, demonstrating compatibility across network backbones.

  9. Knowl 9 — Weakly Supervised Semantic Segmentation Benchmark Results on PASCAL VOC 2012 and MS COCO 2014

    data/table

    Semantic segmentation performance (mIoU %) of DeepLab segmentation models trained using pseudo-labels generated by methods relying solely on image-level classification labels on PASCAL VOC 2012 (val and test sets) and MS COCO 2014 (val set):

    Methods Backbone VOC val VOC test COCO val
    AffinityNet WRN38 61.7 63.7 -
    IRNet RN50 63.5 64.8 41.4
    SEAM WRN38 64.5 65.7 31.9
    OC-CSE WRN38 68.4 68.2 36.4
    CPN WRN38 67.8 68.5 -
    RIB RN101 68.3 68.6 43.8
    PMM WRN38 68.5 69.0 36.7
    ReCAM RN101 68.5 68.4 42.9
    SIPE RN101 68.8 69.7 -
    SIPE WRN38 - - 43.6
    CLIMS RN50 69.3 68.7 -
    W-OoD WRN38 70.7 70.1 -
    Spatial-BCE RN101 70.0 71.3 -
    Spatial-BCE VGG16 - - 35.2
    AEFT WRN38 70.9 71.7 44.8
    Ours (ACR) WRN38 71.9 71.9 45.3
    MCT WRN38 71.9 71.6 42.0
    Ours (ACR + ViT) WRN38 72.4 72.4 -

    ACR with Wide ResNet38 achieves 71.9% mIoU on both VOC val and test sets, and 45.3% mIoU on the MS COCO 2014 val set, setting a new state-of-the-art among image-level weakly supervised segmentation methods. Incorporating ViT attention refinement during pseudo-label generation further pushes performance to 72.4% on VOC val and test.

Coverage note — None was omitted; all key theoretical formulations, loss functions, algorithms, implementation details, ablation experiments, and benchmark results are fully covered.

References

  1. 1.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2209–2218, 2019. 1, 3, 7, 8
  2. 2.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4981–4990, 2018. 1, 3, 6, 8
  3. 3.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation via sub-category exploration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8991–9000, 2020. 1, 2
  4. 4.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In Yoshua Bengio and Yann LeCun, editors, 3rd International Conference on Learning Representations, ICLR, 2015. 6
  5. 5.Liyi Chen, Weiwei Wu, Chenchen Fu, Xiao Han, and Yuntao Zhang. Weakly supervised semantic segmentation with boundary exploration. In European Conference on Computer Vision, pages 347–362. Springer, 2020. 3
  6. 6.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE transactions on pattern analysis and machine intelligence, 40(4):834–848, 2017. 6
  7. 7.Qi Chen, Lingxiao Yang, Jian-Huang Lai, and Xiaohua Xie. Self-supervised image-specific prototype exploration for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4288–4298, 2022. 1, 2, 7, 8
  8. 8.Zhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua, Hanwang Zhang, and Qianru Sun. Class re-activation maps for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 969–978, 2022. 7, 8
  9. 9.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 7, 8
  10. 10.Ye Du, Zehua Fu, Qingjie Liu, and Yunhong Wang. Weakly supervised semantic segmentation by pixel-to-prototype contrast. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4320–4329, 2022. 7
  11. 11.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International journal of computer vision, 88(2):303–338, 2010. 2, 6
  12. 12.Junsong Fan, Zhaoxiang Zhang, Chunfeng Song, and Tieniu Tan. Learning integral objects with intra-class discriminator for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4283–4292, 2020. 1, 2, 3
  13. 13.Junsong Fan, Zhaoxiang Zhang, Tieniu Tan, Chunfeng Song, and Jun Xiao. Cian: Cross-image affinity net for weakly supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 10762–10769, 2020. 2
  14. 14.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020. 2, 4, 6
  15. 15.Peng-Tao Jiang, Yuqi Yang, Qibin Hou, and Yunchao Wei. L2g: A simple local-to-global knowledge transfer framework for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16886–16896, 2022. 3
  16. 16.Anna Khoreva, Rodrigo Benenson, Jan Hosang, Matthias Hein, and Bernt Schiele. Simple does it: Weakly supervised instance and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 876–885, 2017. 1
  17. 17.Philipp Krahenb uhl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. In Advances in neural information processing systems, pages 109–117, 2011. 7
  18. 18.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM, 60(6):84–90, 2017. 6
  19. 19.Hyeokjun Kweon, Sung-Hoon Yoon, Hyeonseong Kim, Daehee Park, and Kuk-Jin Yoon. Unlocking the potential of ordinary classifier: Class-specific adversarial erasing framework for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6994–7003, 2021. 1, 2, 3, 4, 6, 7, 8
  20. 20.Jungbeom Lee, Jooyoung Choi, Jisoo Mok, and Sungroh Yoon. Reducing information bottleneck for weakly supervised semantic segmentation. Advances in Neural Information Processing Systems, 34:27408–27421, 2021. 2, 7, 8
  21. 21.Jungbeom Lee, Eunji Kim, and Sungroh Yoon. Anti-adversarially manipulated attributions for weakly and semi-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4071–4080, 2021. 1, 3, 7
  22. 22.Jungbeom Lee, Seong Joon Oh, Sangdoo Yun, Junsuk Choe, Eunji Kim, and Sungroh Yoon. Weakly supervised semantic segmentation using out-of-distribution data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16897–16906, 2022. 2, 7, 8
  23. 23.Jungbeom Lee, Jihun Yi, Chaehun Shin, and Sungroh Yoon. Bbam: Bounding box attribution map for weakly supervised semantic and instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2643–2652, 2021. 1
  24. 24.Seungho Lee, Minhyun Lee, Jongwuk Lee, and Hyunjung Shim. Railroad is not a train: Saliency as pseudo-pixel supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5495–5505, 2021. 3
  25. 25.Kunpeng Li, Ziyan Wu, Kuan-Chuan Peng, Jan Ernst, and Yun Fu. Tell me where to look: Guided attention inference network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9215–9223, 2018. 2
  26. 26.Xueyi Li, Tianfei Zhou, Jianwu Li, Yi Zhou, and Zhaoxiang Zhang. Group-wise semantic mining for weakly supervised semantic segmentation. arXiv preprint arXiv:2012.05007, 2020. 2, 3
  27. 27.Yi Li, Yiqun Duan, Zhanghui Kuang, Yimin Chen, Wayne Zhang, and Xiaomeng Li. Uncertainty estimation via response scaling for pseudo-mask noise mitigation in weakly-supervised semantic segmentation. arXiv preprint arXiv:2112.07431, 2021. 3
  28. 28.Yi Li, Zhanghui Kuang, Liyang Liu, Yimin Chen, and Wayne Zhang. Pseudo-mask matters in weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6964–6973, 2021. 1, 3, 6, 7, 8
  29. 29.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3159–3167, 2016. 1
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. 2, 6
  31. 31.George Papandreou, Liang-Chieh Chen, Kevin P Murphy, and Alan L Yuille. Weakly-and semi-supervised learning of a deep convolutional network for semantic image segmentation. In Proceedings of the IEEE international conference on computer vision, pages 1742–1750, 2015. 1
  32. 32.Jie Qin, Jie Wu, Xuefeng Xiao, Lujun Li, and Xingang Wang. Activation modulation and recalibration scheme for weakly supervised semantic segmentation. arXiv preprint arXiv:2112.08996, 2021. 2, 7
  33. 33.Yukun Su, Ruizhou Sun, Guosheng Lin, and Qingyao Wu. Context decoupling augmentation for weakly supervised semantic segmentation. arXiv preprint arXiv:2103.01795, 2021. 2, 7
  34. 34.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. arXiv preprint arXiv:2007.01947, 2020. 2, 3
  35. 35.Kunyang Sun, Haoqing Shi, Zhengming Zhang, and Yongming Huang. Ecs-net: Improving weakly supervised semantic segmentation by using connections between class activation maps. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7283–7292, 2021. 2, 6, 7
  36. 36.Paul Vernaza and Manmohan Chandraker. Learning randomwalk label propagation for weakly-supervised semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7158–7166, 2017. 1
  37. 37.Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12275–12284, 2020. 1, 2, 8
  38. 38.Tong Wu, Guangyu Gao, Junshi Huang, Xiaolin Wei, Xiaoming Wei, and Chi Harold Liu. Adaptive spatial-bce loss for weakly supervised semantic segmentation. In European Conference on Computer Vision, pages 199–216. Springer, 2022. 8
  39. 39.Tong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei, Xiaolin Wei, Xuan Luo, and Chi Harold Liu. Embedded discriminative attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16765–16774, 2021. 2, 7
  40. 40.Zifeng Wu, Chunhua Shen, and Anton Van Den Hengel. Wider or deeper: Revisiting the resnet model for visual recognition. Pattern Recognition, 90:119–133, 2019. 6
  41. 41.Jinheng Xie, Xianxu Hou, Kai Ye, and Linlin Shen. Cross language image matching for weakly supervised semantic segmentation. arXiv preprint arXiv:2203.02668, 2022. 7, 8
  42. 42.Jinheng Xie, Jianfeng Xiang, Junliang Chen, Xianxu Hou, Xiaodong Zhao, and Linlin Shen. C2am: Contrastive learning of class-agnostic activation map for weakly supervised object localization and semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 989–998, 2022. 1, 2
  43. 43.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, Ferdous Sohel, and Dan Xu. Leveraging auxiliary tasks with affinity learning for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6984–6993, 2021. 3
  44. 44.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, and Dan Xu. Multi-class token transformer for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4310–4319, 2022. 6, 7, 8
  45. 45.Yazhou Yao, Tao Chen, Guo-Sen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu, Zhenmin Tang, and Jian Zhang. Nonsalient region object mining for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2623–2632, 2021. 3
  46. 46.Sung-Hoon Yoon, Hyeokjun Kweon, Jegyeong Cho, Shinjeong Kim, and Kuk-Jin Yoon. Adversarial erasing framework via triplet with gated pyramid pooling layer for weakly supervised semantic segmentation. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXIX, pages 326–344. Springer Nature Switzerland Cham, 2022. 1, 2, 6, 7, 8
  47. 47.Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang. Reliability does matter: An end-to-end weakly supervised semantic segmentation approach. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 12765–12772, 2020. 1
  48. 48.Dong Zhang, Hanwang Zhang, Jinhui Tang, Xiansheng Hua, and Qianru Sun. Causal intervention for weakly-supervised semantic segmentation. Advances in Neural Information Processing Systems, 2020. 7
  49. 49.Fei Zhang, Chaochen Gu, Chenyue Zhang, and Yuchao Dai. Complementary patch for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7242–7251, 2021. 1, 2, 6, 8
  50. 50.Xiaolin Zhang, Yunchao Wei, Jiashi Feng, Yi Yang, and Thomas S Huang. Adversarial complementary learning for weakly supervised object localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1325–1334, 2018. 2, 6
  51. 51.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2921–2929, 2016. 1
  52. 52.Tianfei Zhou, Meijie Zhang, Fang Zhao, and Jianwu Li. Regional semantic contrast and aggregation for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4299–4309, 2022. 1, 2, 3

Citation

MLA
Kweon, H., et al. “Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 11329–39, https://doi.org/10.1109/CVPR52729.2023.01090.
APA
Kweon, H., Yoon, S.-H., & Yoon, K.-J. (2023). Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11329–11339. https://doi.org/10.1109/CVPR52729.2023.01090
Chicago
Kweon, H., S.-H. Yoon, and K.-J. Yoon. 2023. “Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11329–39. https://doi.org/10.1109/CVPR52729.2023.01090.
Harvard
Kweon, H., Yoon, S.-H. and Yoon, K.-J. (2023) “Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 11329–11339. Available at: https://doi.org/10.1109/CVPR52729.2023.01090.
Vancouver
1. Kweon H, Yoon S-H, Yoon K-J (2023) Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 11329–11339

BibTeX

@inproceedings{Kweon_2023, title={Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01090}, DOI={10.1109/cvpr52729.2023.01090}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Kweon, Hyeokjun and Yoon, Sung-Hoon and Yoon, Kuk-Jin}, year={2023}, month=June, pages={11329–11339} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE