Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation

Zesen ChengPengchong QiaoKehan LiSiheng LiPengxu WeiXiangyang JiLi YuanChang LiuJie Chen

article2023CVPR61 citations

Proposes a plug-and-play rectification mechanism that detects and corrects out-of-candidate pixel errors contradicting image-level labels via a differentiable group ranking loss, consistently boosting existing weakly supervised semantic segmentation baselines with minimal overhead.

Listen

Training computer vision models to accurately outline and identify objects in images typically demands dense, pixel-by-pixel manual annotations, which are expensive and time-consuming to produce. Weakly supervised semantic segmentation addresses this challenge by training models using only image-level tags that indicate which objects are present. However, because standard intermediate cues capture only the most prominent parts of an object, generated training masks contain substantial noise. This noise causes segmentation models to routinely assign pixels to categories that are entirely absent from the image tags—a flaw defined in the article as out-of-candidate errors.

The article demonstrates an out-of-candidate rectification framework designed to systematically detect and correct these out-of-candidate pixel errors during model training. The approach evaluates whether enforcing consistency between pixel predictions and known image-level tags can improve final segmentation performance across standard computer vision benchmarks.

The proposed framework operates in three steps during network training. First, it automatically flags out-of-candidate pixels whenever predicted categories contradict known image tags. Second, it adaptively divides categories into in-candidate and out-of-candidate groups by combining historical co-occurrence patterns across the dataset with current model prediction probabilities. Third, it applies a smooth, differentiable ranking loss that forces model activations for valid in-candidate classes to exceed those for out-of-candidate classes. The authors integrated this plug-and-play module into three established baseline systems and tested them on standard benchmarks, including the PASCAL VOC 2012 and MS COCO 2014 datasets.

The analysis produced several key findings. First, applying the rectification method reduced the rate of out-of-candidate pixel errors substantially across evaluated baselines, cutting error rates on the PASCAL VOC validation set from baseline levels of 21.5%–32.4% down to 7.9%–10.1%. Second, this error suppression produced consistent accuracy gains across all evaluated models, boosting mean intersection-over-union scores by 0.8% to 3.3% on PASCAL VOC and by 0.5% to 1.3% on the more complex MS COCO dataset. Third, combining the rectification module with a state-of-the-art transformer baseline established new top-tier benchmark performance on both datasets. Finally, the framework achieved these gains with negligible overhead, adding only 0.56 to 1.18 minutes per training epoch and requiring zero additional computation during final inference.

These findings indicate that directly penalizing logical contradictions against image-level labels is an effective, low-risk way to enhance weakly supervised vision pipelines. Organizations developing vision systems can reduce manual labeling costs without sacrificing accuracy by incorporating this loss function into existing training workflows. Because the rectification step is removed during inference, the performance improvements come with no operational latency or deployment cost trade-offs.

Teams training semantic segmentation models under weak supervision should consider integrating out-of-candidate rectification as a standard regularizer. When deploying the method, teams should use adaptive group splitting rather than simple rule-based assignment, as ablations show adaptive filtering provides the strongest accuracy improvements. While confidence in the reported experimental improvements is high across the evaluated datasets, validation on specialized, non-standard visual domains or with alternative network architectures would further confirm the module's broader generalizability.

Cover for Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation

Abstract

Weakly supervised semantic segmentation is typically inspired by class activation maps, which serve as pseudo masks with class-discriminative regions highlighted. Although tremendous efforts have been made to recall precise and complete locations for each class, existing methods still commonly suffer from the unsolicited Out-of-Candidate (OC) error predictions that do not belong to the label candidates, which could be avoidable since the contradiction with image-level class tags is easy to be detected. In this paper, we develop a group ranking-based Out-of-Candidate Rectification (OCR) mechanism in a plug-and-play fashion. Firstly, we adaptively split the semantic categories into In-Candidate (IC) and OC groups for each OC pixel according to their prior annotation correlation and posterior prediction correlation. Then, we derive a differentiable rectification loss to force OC pixels to shift to the IC group. Incorporating OCR with seminal baselines (e.g., AffinityNet, SEAM, MCTformer), we can achieve remarkable performance gains on both Pascal VOC (+3.2%, +3.3%, +0.8% mIoU) and MS COCO (+1.0%, +1.3%, +0.5% mIoU) datasets with negligible extra training overhead, which justifies the effectiveness and generality of OCR.

Table of Contents

  • 1. Introduction
  • 2. Related works
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Overall Pipeline
  • 3.3. Out-of-Candidate Rectification
  • Posterior Prediction
  • 4. Experiments
  • 4.1. Experimental Settings
  • 4.2. Main Results
  • 4.3. Quantitative Analysis
  • 4.4. Qualitative Analysis
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Out-of-Candidate (OC) and In-Candidate (IC) Errors in Weakly Supervised Semantic Segmentation

    definition

    In weakly supervised semantic segmentation (WSSS) supervised only by image-level tags, an image xix_i has an associated candidate label set Si⊂{1,2,…,C}S_i \subset \{1, 2, \dots, C\}, where CC is the number of semantic foreground classes.

    An Out-of-Candidate (OC) error occurs when a pixel in xix_i is predicted by the segmentation network as a semantic class ll that is not present in the candidate set and is not background, i.e., l∉Si∪{bg}l \notin S_i \cup \{bg\}, where bgbg represents the background class.

    • OC pixels: Pixels whose predicted class label belongs to the complement set Si∪{bg}‾\overline{S_i \cup \{bg\}}.
    • OC categories (GocG_{oc}): Categories that do not belong to the candidate set of image xix_i, i.e., Goc={l∣l∉Si∪{bg}}G_{oc} = \{l \mid l \notin S_i \cup \{bg\}\}.
    • In-Candidate (IC) categories (GicG_{ic}): Categories that belong to the candidate label set including background (Si∪{bg}S_i \cup \{bg\}) and are potentially correct semantic labels for the given pixel.
  2. Knowl 2 — Out-of-Candidate Rectification (OCR) Mechanism

    model/method

    Out-of-Candidate Rectification (OCR) is a plug-and-play training framework designed to eliminate Out-of-Candidate (OC) prediction errors during the segmentation network training stage in weakly supervised semantic segmentation (WSSS). Rather than relying solely on noisy pseudo-labels generated from Class Activation Maps (CAMs), OCR utilizes the ground-truth image-level label candidates SiS_i to enforce group-level prediction ranking constraints on pixel logits.

    The OCR workflow operates in three sequential steps:

    1. OC Pixel Selection: Identifies pixels whose predicted class contradicts the image-level candidate label set Si∪{bg}S_i \cup \{bg\}.
    2. IC and OC Group Split: Adaptively partitions the class label space into an In-Candidate group GicG_{ic} and an Out-of-Candidate group GocG_{oc} for each OC pixel by filtering out irrelevant candidate classes using prior label co-occurrence correlation and posterior network predictions.
    3. Group Ranking Rectification: Applies a differentiable rectification loss Lrec\mathcal{L}_{rec} to penalize instances where the maximum logit in GocG_{oc} exceeds the minimum logit in GicG_{ic} by a margin Δ\Delta, pulling OC pixels toward valid IC classes and pushing them away from invalid OC classes.

    OCR adds negligible computation during training and zero computation during inference, as it modifies only the training loss function.

  3. Knowl 3 — Adaptive IC and OC Group Splitting Algorithm

    algorithm

    The adaptive splitting procedure determines the subset of valid candidate classes GicG_{ic} and invalid candidate classes GocG_{oc} for a given pixel identified as an Out-of-Candidate (OC) error.

    Input: Segmentation logits vector z∈RC+1z \in \mathbb{R}^{C+1}, posterior probability vector P=softmax(z)P = \mathrm{softmax}(z), image-level candidate label set SiS_i, dataset label co-occurrence matrix M∈R(C+1)×(C+1)\mathcal{M} \in \mathbb{R}^{(C+1) \times (C+1)}, threshold tt
    Output: OC indicator mocm_{oc}, In-Candidate class set GicG_{ic}, Out-of-Candidate class set GocG_{oc}
    predicted_class = argmaxk(zk)\mathrm{argmax}_{k} (z^k)
    if predicted_class ∈Si∪{bg}‾\in \overline{S_i \cup \{bg\}}:
        moc=1m_{oc} = 1
    else:
        moc=0m_{oc} = 0
        return moc,∅,∅m_{oc}, \emptyset, \emptyset
    A=argmaxj∈Si∪{bg}(zj)\mathcal{A} = \mathrm{argmax}_{j \in S_i \cup \{bg\}} (z^j)
    Goc={l∣l∉Si∪{bg}}G_{oc} = \{l \mid l \notin S_i \cup \{bg\}\}
    Gic=∅G_{ic} = \emptyset
    for each class k∈Si∪{bg}k \in S_i \cup \{bg\}:
        if PA−Pk×MA,k<tP^{\mathcal{A}} - P^k \times \mathcal{M}_{\mathcal{A}, k} < t:
            Gic=Gic∪{k}G_{ic} = G_{ic} \cup \{k\}
    return moc,Gic,Gocm_{oc}, G_{ic}, G_{oc}

    The co-occurrence prior correlation matrix M\mathcal{M} is computed across the training dataset of LL images as: Mk,l=1L∑i=1LI(k∈Si∧l∈Si)\mathcal{M}_{k,l} = \frac{1}{L} \sum_{i=1}^L \mathbb{I}(k \in S_i \land l \in S_i) where I(⋅)\mathbb{I}(\cdot) is the indicator function. The anchor class A\mathcal{A} acts as the most confident valid candidate and is used to prune unpromising candidate tags from GicG_{ic} using threshold t=0.2t = 0.2.

  4. Knowl 4 — Smooth Differentiable Group-Ranking Rectification Loss

    equation

    The goal of Out-of-Candidate Rectification is to ensure that for every Out-of-Candidate pixel, the network's logits for all classes in the In-Candidate group GicG_{ic} are strictly higher than those in the Out-of-Candidate group GocG_{oc} by a margin Δ>0\Delta > 0: max⁡l∈Goczl<min⁡k∈Giczk\max_{l \in G_{oc}} z^l < \min_{k \in G_{ic}} z^k

    The non-differentiable margin ranking loss is formulated as: Lrecraw=max⁡l∈Goczl−min⁡k∈Giczk+Δ=max⁡l∈Goczl+max⁡k∈Gic(−zk)+Δ\mathcal{L}_{rec}^{\text{raw}} = \max_{l \in G_{oc}} z^l - \min_{k \in G_{ic}} z^k + \Delta = \max_{l \in G_{oc}} z^l + \max_{k \in G_{ic}} (-z^k) + \Delta

    Applying the smooth approximation max⁡(u1,…,un)≈log⁡∑i=1neui\max(u_1, \dots, u_n) \approx \log \sum_{i=1}^n e^{u_i} and softplus approximation for ReLU(u)≈log⁡(1+eu)\mathrm{ReLU}(u) \approx \log(1 + e^u), the pixel-level differentiable rectification loss is: Lrec=moclog⁡[1+(∑k∈Gice−zk)×(∑l∈Gocezl+Δ)]\mathcal{L}_{rec} = m_{oc} \log \left[ 1 + \left( \sum_{k \in G_{ic}} e^{-z^k} \right) \times \left( \sum_{l \in G_{oc}} e^{z^l + \Delta} \right) \right] where moc∈{0,1}m_{oc} \in \{0, 1\} is the OC pixel selection mask indicating whether the predicted class is in Si∪{bg}‾\overline{S_i \cup \{bg\}}.

    The total training objective for the segmentation network is: L=Lseg+αLrec\mathcal{L} = \mathcal{L}_{seg} + \alpha \mathcal{L}_{rec} where Lseg\mathcal{L}_{seg} is the standard cross-entropy loss against CAM-derived pseudo-labels y^\hat{y}, and α\alpha is the loss modulation coefficient (set to α=1.0,Δ=2.0\alpha = 1.0, \Delta = 2.0).

  5. Knowl 5 — Experimental Setup for Weakly Supervised Semantic Segmentation with OCR

    experimental setup

    The Out-of-Candidate Rectification (OCR) method is evaluated on two standard semantic segmentation benchmarks under image-level weak supervision:

    • PASCAL VOC 2012: 20 foreground classes plus background. Augmented training set contains 10,582 images; validation set contains 1,449 images; test set contains 1,456 images.
    • MS COCO 2014: 80 foreground classes plus background, using COCO stuff ground-truth annotations. Training set contains 82,081 images; validation set contains 40,137 images.

    Network Architecture & Training:

    • Segmentation Network: DeepLab-LargeFOV (DeepLabv1) with a ResNet38 backbone pre-trained on ImageNet, using an output stride of 8.
    • Baselines: AffinityNet (ResNet38 classifier), SEAM (ResNet38 classifier), and MCTformer (DeiT-S classifier).
    • Optimizer: SGD with momentum 0.90.9, weight decay 5×10−45 \times 10^{-4}, initial learning rate 1×10−31 \times 10^{-3} with exponential decay.
    • Training Details: Batch size 16, trained for 30 epochs. Training images are randomly rescaled between scales 0.70.7 and 1.31.3, then cropped to 321×321321 \times 321.
    • Hyperparameters: Filtering threshold t=0.2t = 0.2, margin Δ=2.0\Delta = 2.0, loss weight α=1.0\alpha = 1.0.
    • Evaluation & Post-processing: Evaluated using mean Intersection-over-Union (mIoU, %). Test-time augmentation and DenseCRF post-processing are applied.
  6. Knowl 6 — Performance Improvement of OCR on PASCAL VOC 2012 and MS COCO 2014

    data/table

    Integrating Out-of-Candidate Rectification (OCR) into existing WSSS baselines consistently improves validation and test mIoU across both PASCAL VOC 2012 and MS COCO 2014 datasets.

    Method Classification Backbone Segmentation Network VOC val mIoU (%) VOC test mIoU (%)
    AffinityNet ResNet38 V1-Res38 61.7 63.7
    OCR + AffinityNet ResNet38 V1-Res38 64.9 (+3.2) 65.2 (+1.5)
    SEAM ResNet38 V1-Res38 64.5 65.7
    OCR + SEAM ResNet38 V1-Res38 67.8 (+3.3) 68.4 (+2.7)
    MCTformer DeiT-S V1-Res38 71.9 71.6
    OCR + MCTformer DeiT-S V1-Res38 72.7 (+0.8) 72.0 (+0.4)
    Method Classification Backbone Segmentation Network COCO val mIoU (%)
    AffinityNet ResNet38 V1-Res38 29.5
    OCR + AffinityNet ResNet38 V1-Res38 30.5 (+1.0)
    SEAM ResNet38 V1-Res38 31.9
    OCR + SEAM ResNet38 V1-Res38 33.2 (+1.3)
    MCTformer DeiT-S V1-Res38 42.0
    OCR + MCTformer DeiT-S V1-Res38 42.5 (+0.5)

    On PASCAL VOC 2012, OCR provides gains of +3.2% mIoU on AffinityNet, +3.3% mIoU on SEAM, and +0.8% mIoU on MCTformer. On MS COCO 2014 (which contains 80 classes where OC confusion is more frequent), OCR yields improvements of +1.0%, +1.3%, and +0.5% mIoU over AffinityNet, SEAM, and MCTformer, respectively.

  7. Knowl 7 — Ablation on IC/OC Group Splitting and Target Pixel Selection Strategies

    data/table

    Ablation experiments evaluated on the PASCAL VOC 2012 validation set using SEAM as the baseline (64.5% baseline mIoU) demonstrate the necessity of adaptive candidate splitting and restricting rectification strictly to OC pixels.

    Strategy Variant GicG_{ic} Specification GocG_{oc} Specification mIoU (%)
    Baseline (No OCR) None None 64.5
    All Candidates: all(⋅\cdot) S∪{bg}S \cup \{bg\} S∪{bg}‾\overline{S \cup \{bg\}} 63.4 (-1.1)
    Max Only: max(⋅\cdot) {A}\{\mathcal{A}\} S∪{bg}‾\overline{S \cup \{bg\}} 67.2 (+2.7)
    Adaptive: ada(⋅\cdot) {k∈S∪{bg}∣PA−PkMA,k<t}\{k \in S \cup \{bg\} \mid P^{\mathcal{A}} - P^k \mathcal{M}_{\mathcal{A},k} < t\} S∪{bg}‾\overline{S \cup \{bg\}} 67.8 (+3.3)
    Target Rectified Pixels Applied Pixel Subset mIoU (%)
    None None 64.5
    IC Pixels Only {p∣argmaxk(zpk)∈S∪{bg}}\{p \mid \mathrm{argmax}_k(z_p^k) \in S \cup \{bg\}\} 64.7 (+0.2)
    All Pixels All image pixels 66.9 (+2.4)
    OC Pixels Only {p∣argmaxk(zpk)∉S∪{bg}}\{p \mid \mathrm{argmax}_k(z_p^k) \notin S \cup \{bg\}\} 67.8 (+3.3)

    Key takeaways:

    1. Including all candidate tags in GicG_{ic} degrades performance below the baseline (−1.1%-1.1\%) because irrelevant candidate tags mislead pixel rectification.
    2. Pruning via the adaptive strategy ada(·) outperforms the single top-prediction strategy max(·) by +0.6%+0.6\%, demonstrating that the top anchor class A\mathcal{A} is not always the sole ground-truth class for the pixel.
    3. Applying rectification exclusively to OC pixels yields the optimal +3.3%+3.3\% gain; applying it to all pixels dilutes precision on already correct pixels.
  8. Knowl 8 — Out-of-Candidate Error Rate Suppression Across Baseline Models

    empirical result

    On the PASCAL VOC 2012 validation set, incorporating Out-of-Candidate Rectification (OCR) substantially reduces the fraction of images exhibiting Out-of-Candidate (OC) prediction errors:

    • AffinityNet: OC error image rate drops from 32.4% (baseline) to 9.7% with OCR (relative reduction of 70.1%). Corresponding segmentation mIoU increases from 61.7% to 64.9%.
    • SEAM: OC error image rate drops from 21.5% (baseline) to 10.1% with OCR (relative reduction of 53.0%). Corresponding segmentation mIoU increases from 64.5% to 67.8%.
    • MCTformer: OC error image rate drops from 22.4% (baseline) to 7.9% with OCR (relative reduction of 64.7%). Corresponding segmentation mIoU increases from 71.9% to 72.7%.

    This confirms that suppressing OC errors directly correlates with improved final semantic segmentation accuracy.

  9. Knowl 9 — Hyperparameter Sensitivity and Computational Overhead of OCR

    empirical result

    Empirical sensitivity analysis and computational overhead measurements on PASCAL VOC 2012 validation using SEAM show:

    • Margin Δ\Delta: Swept across values in [0,8][0, 8]. The peak validation mIoU occurs at Δ=2.0\Delta = 2.0. Smaller values provide insufficient separation between IC and OC classes, while larger values impose over-regularization that impairs performance.
    • Loss Coefficient α\alpha: Swept across {0.1,1.0,10.0}\{0.1, 1.0, 10.0\}. The optimal coefficient is α=1.0\alpha = 1.0; setting α=10.0\alpha = 10.0 leads to degraded performance due to over-penalization.
    • Training Speed & Efficiency: Benchmarked with SEAM across various crop sizes:
      • At 256×256256 \times 256: Training time is 14.3814.38 min/epoch (+0.56+0.56 min/epoch overhead vs. baseline), yielding 63.6%63.6\% mIoU (+3.2%+3.2\% gain).
      • At 321×321321 \times 321: Training time is 23.3223.32 min/epoch (+0.83+0.83 min/epoch overhead vs. baseline), yielding 67.8%67.8\% mIoU (+3.3%+3.3\% gain).
      • At 448×448448 \times 448: Training time is 42.2942.29 min/epoch (+1.18+1.18 min/epoch overhead vs. baseline), yielding 68.4%68.4\% mIoU (+3.7%+3.7\% gain).
    • Inference Overhead: Zero additional parameters or compute are required at test time because OCR is strictly an auxiliary loss during training.

Coverage note — None was omitted; all key contributions including definitions, algorithms, loss formulations, benchmark results, ablations, and analysis were included.

References

  1. 1.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2209–2218, 2019.
  2. 2.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.
  3. 3.Nikita Araslanov and Stefan Roth. Single-stage semantic segmentation from image labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  4. 4.Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al. A closer look at memorization in deep networks. In International Conference on Machine Learning, pages 233–242. PMLR, 2017.
  5. 5.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What’s the point: Semantic segmentation with point supervision. In European Conference on Computer Vision, 2016.
  6. 6.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1209–1218, 2018.
  7. 7.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation via sub-category exploration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  8. 8.Liyi Chen, Weiwei Wu, Chenchen Fu, Xiao Han, and Yuntao Zhang. Weakly supervised semantic segmentation with boundary exploration. In European Conference on Computer Vision, 2020.
  9. 9.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. In International Conference on Learning Representations, 2015.
  10. 10.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4):834–848, 2017.
  11. 11.Qi Chen, Lingxiao Yang, Jian-Huang Lai, and Xiaohua Xie. Self-supervised image-specific prototype exploration for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4288–4298, 2022.
  12. 12.Zhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua, Hanwang Zhang, and Qianru Sun. Class re-activation maps for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 969–978, 2022.
  13. 13.Junsuk Choe, Seungho Lee, and Hyunjung Shim. Attention-based dropout layer for weakly supervised single object localization and semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 43(12):4256–4271, 2020.
  14. 14.Jifeng Dai, Kaiming He, and Jian Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2015.
  15. 15.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2009.
  16. 16.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021.
  17. 17.Ye Du, Zehua Fu, Qingjie Liu, and Yunhong Wang. Weakly supervised semantic segmentation by pixel-to-prototype contrast. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4320–4329, 2022.
  18. 18.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. International Journal of Computer Vision, 88(2):303–338, 2010.
  19. 19.Junsong Fan, Zhaoxiang Zhang, Chunfeng Song, and Tieniu Tan. Learning integral objects with intra-class discriminator for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  20. 20.Junsong Fan, Zhaoxiang Zhang, and Tieniu Tan. Employing multi-estimations for weakly-supervised semantic segmentation. In European Conference on Computer Vision, pages 332–348. Springer, 2020.
  21. 21.Junsong Fan, Zhaoxiang Zhang, Tieniu Tan, Chunfeng Song, and Jun Xiao. Cian: Cross-image affinity net for weakly supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020.
  22. 22.Wei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng, Qi Tian, Zhenjun Han, Bolei Zhou, and Qixiang Ye. Ts-cam: Token semantic coupled attention map for weakly supervised object localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  23. 23.Xavier Glorot, Antoine Bordes, and Yoshua Bengio. Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics, pages 315–323. JMLR Workshop and Conference Proceedings, 2011.
  24. 24.Bharath Hariharan, Pablo Arbeláez, Lubomir Bourdev, Subhransu Maji, and Jitendra Malik. Semantic contours from inverse detectors. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2011.
  25. 25.Alexander Hermans, Lucas Beyer, and Bastian Leibe. In defense of the triplet loss for person re-identification. arXiv preprint arXiv:1703.07737, 2017.
  26. 26.Seunghoon Hong, Donghun Yeo, Suha Kwak, Honglak Lee, and Bohyung Han. Weakly supervised semantic segmentation using web-crawled videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017.
  27. 27.Qibin Hou, PengTao Jiang, Yunchao Wei, and Ming-Ming Cheng. Self-erasing network for integral object attention. In Advances in Neural Information Processing Systems, 2018.
  28. 28.Peng-Tao Jiang, Yuqi Yang, Qibin Hou, and Yunchao Wei. L2g: A simple local-to-global knowledge transfer framework for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16886–16896, 2022.
  29. 29.Tsung-Wei Ke, Jyh-Jing Hwang, and Stella X Yu. Universal weakly supervised segmentation by pixel-to-segment contrastive learning. In International Conference on Learning Representations, 2021.
  30. 30.Anna Khoreva, Rodrigo Benenson, Jan Hendrik Hosang, Matthias Hein, and Bernt Schiele. Simple does it: Weakly supervised instance and semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017.
  31. 31.Hyeokjun Kweon, Sung-Hoon Yoon, Hyeonseong Kim, Daehee Park, and Kuk-Jin Yoon. Unlocking the potential of ordinary classifier: Class-specific adversarial erasing framework for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  32. 32.Jungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee, and Sungroh Yoon. Ficklenet: Weakly and semi-supervised segmentation using stochastic inference. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019.
  33. 33.Jungbeom Lee, Eunji Kim, and Sungroh Yoon. Anti-adversarially manipulated attributions for weakly and semi-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  34. 34.Jungbeom Lee, Jihun Yi, Chaehun Shin, and Sungroh Yoon. Bbam: Bounding box attribution map for weakly supervised semantic and instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2643–2652, 2021.
  35. 35.Seungho Lee, Minhyun Lee, Jongwuk Lee, and Hyunjung Shim. Railroad is not a train: Saliency as pseudo-pixel supervision for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  36. 36.Kunpeng Li, Ziyan Wu, Kuan-Chuan Peng, Jan Ernst, and Yun Fu. Tell me where to look: Guided attention inference network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.
  37. 37.Xueyi Li, Tianfei Zhou, Jianwu Li, Yi Zhou, and Zhaoxiang Zhang. Group-wise semantic mining for weakly supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021.
  38. 38.Yi Li, Yiqun Duan, Zhanghui Kuang, Yimin Chen, Wayne Zhang, and Xiaomeng Li. Uncertainty estimation via response scaling for pseudo-mask noise mitigation in weakly-supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 1447–1455, 2022.
  39. 39.Yi Li, Zhanghui Kuang, Liyang Liu, Yimin Chen, and Wayne Zhang. Pseudo-mask matters in weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6964–6973, 2021.
  40. 40.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016.
  41. 41.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014.
  42. 42.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3431–3440, 2015.
  43. 43.Wenfeng Luo and Meng Yang. Learning saliency-free model with generic features for weakly-supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020.
  44. 44.Frank Nielsen and Ke Sun. Guaranteed bounds on information-theoretic measures of univariate mixtures using piecewise log-sum-exp inequalities. Entropy, 18(12):442, 2016.
  45. 45.Pengchong Qiao, Zhidan Wei, Yu Wang, Chang Liu, Zhennan Wang, Guoli Song, and Jie Chen. Pcr: Pessimistic consistency regularization for semi-supervised segmentation. arXiv preprint arXiv:2210.08519, 2022.
  46. 46.Jie Qin, Jie Wu, Xuefeng Xiao, Lujun Li, and Xingang Wang. Activation modulation and recalibration scheme for weakly supervised semantic segmentation. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 2117–2125, 2022.
  47. 47.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  48. 48.Simone Rossetti, Damiano Zappia, Marta Sanzari, Marco Schaerf, and Fiora Pirri. Max pooling with vision transformers reconciles class and shape in weakly supervised semantic segmentation. In European Conference on Computer Vision, pages 446–463. Springer, 2022.
  49. 49.Lixiang Ru, Yibing Zhan, Baosheng Yu, and Bo Du. Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16846–16855, 2022.
  50. 50.Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 815–823, 2015.
  51. 51.Wataru Shimoda and Keiji Yanai. Self-supervised difference detection for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019.
  52. 52.Krishna Kumar Singh and Yong Jae Lee. Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2017.
  53. 53.Yukun Su, Ruizhou Sun, Guosheng Lin, and Qingyao Wu. Context decoupling augmentation for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  54. 54.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In European Conference on Computer Vision, 2020.
  55. 55.Yifan Sun, Changmao Cheng, Yuhan Zhang, Chi Zhang, Liang Zheng, Zhongdao Wang, and Yichen Wei. Circle loss: A unified perspective of pair similarity optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6398–6407, 2020.
  56. 56.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning, 2021.
  57. 57.Paul Vernaza and Manmohan Chandraker. Learning random-walk label propagation for weakly-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7158–7166, 2017.
  58. 58.Feng Wang, Xiang Xiang, Jian Cheng, and Alan Loddon Yuille. Normface: L2 hypersphere embedding for face verification. In Proceedings of the 25th ACM international conference on Multimedia, pages 1041–1049, 2017.
  59. 59.Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu. Cosface: Large margin cosine loss for deep face recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5265–5274, 2018.
  60. 60.Xiang Wang, Sifei Liu, Huimin Ma, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation by iterative affinity learning. International Journal of Computer Vision, 128(6):1736–1749, 2020.
  61. 61.Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020.
  62. 62.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2017.
  63. 63.Yunchao Wei, Huaxin Xiao, Honghui Shi, Zequn Jie, Jiashi Feng, and Thomas S Huang. Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018.
  64. 64.Tong Wu, Guangyu Gao, Junshi Huang, Xiaolin Wei, Xiaoming Wei, and Chi Harold Liu. Adaptive spatial-bce loss for weakly supervised semantic segmentation. In European Conference on Computer Vision, pages 199–216. Springer, 2022.
  65. 65.Tong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei, Xiaolin Wei, Xuan Luo, and Chi Harold Liu. Embedded discriminative attention mechanism for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  66. 66.Zifeng Wu, Chunhua Shen, and Anton Van Den Hengel. Wider or deeper: Revisiting the resnet model for visual recognition. Pattern Recognition, 90:119–133, 2019.
  67. 67.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems, 34:12077–12090, 2021.
  68. 68.Jinheng Xie, Jianfeng Xiang, Junliang Chen, Xianxu Hou, Xiaodong Zhao, and Linlin Shen. Contrastive learning of class-agnostic activation map for weakly supervised object localization and semantic segmentation. arXiv preprint arXiv:2203.13505, 2022.
  69. 69.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, Ferdous Sohel, and Dan Xu. Leveraging auxiliary tasks with affinity learning for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  70. 70.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, and Dan Xu. Multi-class token transformer for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4310–4319, 2022.
  71. 71.Lian Xu, Hao Xue, Mohammed Bennamoun, Farid Boussaid, and Ferdous Sohel. Atrous convolutional feature network for weakly supervised semantic segmentation. Neurocomputing, 421:115–126, 2021.
  72. 72.Yazhou Yao, Tao Chen, Guo-Sen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu, Zhenmin Tang, and Jian Zhang. Non-salient region object mining for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  73. 73.Sung-Hoon Yoon, Hyeokjun Kweon, Jegyeong Cho, Shinjeong Kim, and Kuk-Jin Yoon. Adversarial erasing framework via triplet with gated pyramid pooling layer for weakly supervised semantic segmentation. In European Conference on Computer Vision, pages 326–344. Springer, 2022.
  74. 74.Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang. Reliability does matter: An end-to-end weakly supervised semantic segmentation approach. In Proceedings of the AAAI Conference on Artificial Intelligence, 2020.
  75. 75.Dong Zhang, Hanwang Zhang, Jinhui Tang, Xiansheng Hua, and Qianru Sun. Causal intervention for weakly-supervised semantic segmentation. In Advances in Neural Information Processing Systems, 2020.
  76. 76.Fei Zhang, Chaochen Gu, Chenyue Zhang, and Yuchao Dai. Complementary patch for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021.
  77. 77.Tianyi Zhang, Guosheng Lin, Weide Liu, Jianfei Cai, and Alex Kot. Splitting vs. merging: Mining object regions with discrepancy and intersection loss for weakly supervised semantic segmentation. In European Conference on Computer Vision, 2020.
  78. 78.Xiangrong Zhang, Zelin Peng, Peng Zhu, Tianyang Zhang, Chen Li, Huiyu Zhou, and Licheng Jiao. Adaptive affinity loss and erroneous pseudo-label refinement for weakly supervised semantic segmentation. In Proceedings of the 29th ACM International Conference on Multimedia, pages 5463–5472, 2021.
  79. 79.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016.
  80. 80.Tianfei Zhou, Meijie Zhang, Fang Zhao, and Jianwu Li. Regional semantic contrast and aggregation for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4299–4309, 2022.

Citation

MLA
Cheng, Z., et al. “Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation”. arXiv, 2022, http://arxiv.org/abs/2211.12268v3.
APA
Cheng, Z., Qiao, P., Li, K., Li, S., Wei, P., Ji, X., Yuan, L., Liu, C., & Chen, J. (2022). Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation. arXiv. http://arxiv.org/abs/2211.12268v3
Chicago
Cheng, Z., P. Qiao, K. Li, et al. 2022. “Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation”. arXiv. http://arxiv.org/abs/2211.12268v3.
Harvard
Cheng, Z. et al. (2022) “Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.12268v3.
Vancouver
1. Cheng Z, Qiao P, Li K, Li S, Wei P, Ji X, Yuan L, Liu C, Chen J (2022) Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation. arXiv

BibTeX

@article{cheng2022out,
  title = {Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation},
  author = {Cheng, Zesen and Qiao, Pengchong and Li, Kehan and Li, Siheng and Wei, Pengxu and Ji, Xiangyang and Yuan, Li and Liu, Chang and Chen, Jie},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.12268v3},
  eprint = {2211.12268}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE