Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation

Tianfei ZhouMeijie ZhangFang ZhaoJianwu Li

article2022CVPR165 citations

Introduces Regional Semantic Contrast and Aggregation (RCA), a framework that utilizes a dataset-wide regional memory bank to contrast and aggregate categorical object patterns across training images, achieving state-of-the-art weakly supervised semantic segmentation on PASCAL VOC and COCO.

Listen

Semantic segmentation allows computer vision systems to identify and classify visual objects at the exact pixel level, supporting crucial applications such as autonomous driving and medical imaging. However, standard systems require expensive and labor-intensive manual pixel annotations. To lower these costs, weakly supervised methods use simple image-level category tags instead. The central challenge with this approach is that image classifiers usually focus only on the most distinct parts of an object rather than its full shape. Existing techniques struggle because they analyze images individually or in small groups, missing broader contextual patterns present across entire datasets.

The article evaluates whether extracting dataset-wide visual relationships can overcome this limitation and improve segmentation accuracy. To demonstrate this, the authors introduce a framework called Regional Semantic Contrast and Aggregation. The approach constructs a dynamic memory bank that stores and continuously refines regional visual representations from training images. Using this stored knowledge, the system contrasts regions of the same and different categories to sharpen object boundaries and applies a non-parametric attention mechanism to aggregate broad context across images, with regional blending used to enhance robustness against noisy label estimates.

The experimental findings show significant performance gains on standard industry benchmarks, specifically PASCAL VOC 2012 and COCO 2014. Generating training masks with the new framework improved base baseline accuracy by 2.7 to 3.8 percentage points. On standard segmentation tests, the approach consistently outperformed established baselines by 1.1 to 3.6 percentage points, establishing a new state-of-the-art result. Ablation studies confirmed that combining contrastive comparison with context aggregation yields much stronger results than using either technique alone. Furthermore, diagnostic tests revealed that maintaining compact prototype representations per class preserved strong performance without requiring vast memory storage.

These results demonstrate that organizations can achieve high-grade semantic segmentation while relying primarily on inexpensive image-level tags. By reducing the reliance on costly manual pixel annotations, teams can lower data labeling budgets and shorten development timelines for computer vision deployment. The framework is flexible and can integrate into existing pipelines to improve accuracy across complex scenes, scale variations, and overlapping objects.

Organizations developing vision systems should consider piloting this regional contrast and aggregation approach within their existing weakly supervised workflows. In deployment, teams should retain only the compressed class prototypes during model inference to maximize performance while minimizing computational overhead. Although the experimental evidence provides high confidence across benchmark datasets, decision-makers should note that the approach still relies on initial classifier activations and supplementary background cues; validating performance on specialized operational data is advised before full production rollout.

Cover for Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation

Abstract

Learning semantic segmentation from weakly-labeled (e.g., image tags only) data is challenging since it is hard to infer dense object regions from sparse semantic tags. Despite being broadly studied, most current efforts directly learn from limited semantic annotations carried by individual image or image pairs, and struggle to obtain integral localization maps. Our work alleviates this from a novel perspective, by exploring rich semantic contexts synergistically among abundant weakly-labeled training data for network learning and inference. In particular, we propose regional semantic contrast and aggregation (RCA). RCA is equipped with a regional memory bank to store massive, diverse object patterns appearing in training data, which acts as strong support for exploration of dataset-level semantic structure. Particularly, we propose i) semantic contrast to drive network learning by contrasting massive categorical object regions, leading to a more holistic object pattern understanding, and ii) semantic aggregation to gather diverse relational contexts in the memory to enrich semantic representations. In this manner, RCA earns a strong capability of fine-grained semantic understanding, and eventually establishes new state-of-the-art results on two popular benchmarks, i.e., PASCAL VOC 2012 and COCO 2014.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our Approach
  • 3.1 Problem Statement
  • 3.2 Regional Semantic Contrast and Aggregation
  • 3.2.1 Pseudo-Region Representation
  • 3.2.2 Pseudo-Region Memory Bank
  • 3.2.3 Regional Semantic Contrast (RSC)
  • 3.2.4 Regional Semantic Aggregation (RSA)
  • 3.2.5 Class Activation Map Prediction
  • 3.3 Detailed Network Architecture
  • 4 Experiment
  • 4.1 Experimental Setting
  • 4.2 Diagnostic Experiment
  • 4.3 Comparison with Prior Art
  • 4.4 Visualization Result
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Regional Semantic Contrast and Aggregation (RCA) Framework for WSSS

    model/method

    Weakly supervised semantic segmentation (WSSS) with image-level labels typically relies on Class Activation Maps (CAM) generated from single-image classification models, which tend to identify only the most discriminative object parts. Regional Semantic Contrast and Aggregation (RCA) overcomes this limitation by mining dataset-level contextual relationships across images through a non-parametric categorical memory bank.

    The RCA framework operates via the following workflow:

    1. Feature Extraction and Intermediate CAM: An input image II is processed by a fully convolutional network (FCN) backbone to produce visual feature maps, from which an initial class-aware attention map is generated via a 1×11 \times 1 convolutional layer.
    2. Pseudo-Region Representation: Categorical region embeddings are extracted using masked average pooling guided by the strongly activated regions of the intermediate CAM.
    3. Dataset-Level Memory Bank: A dynamic memory bank stores historical region embeddings for each category, updated using a momentum mechanism.
    4. Regional Semantic Contrast (RSC): The network is trained using a contrastive loss with region mixup regularization to pull pseudo-region embeddings toward same-class memory entries and push them away from different-class memory entries.
    5. Regional Semantic Aggregation (RSA): Memory embeddings are compressed via kk-means clustering into categorical prototypes. A non-parametric cross-attention mechanism aggregates these global prototypes into the local feature representation to capture inter-image context.
    6. Final CAM Generation: The augmented feature maps are fed into a second class-aware convolutional layer to produce more integral and accurate localization maps.
  2. Knowl 2 — Pseudo-Region Feature Extraction via Masked Average Pooling

    model/method

    In the RCA framework for weakly supervised semantic segmentation, dense feature representations of an image are summarized into compact categorical region embeddings to enable scalable cross-image relational learning.

    Given an input image I∈Rw×h×3I \in \mathbb{R}^{w \times h \times 3} and its multi-label ground truth y=[y1,y2,…,yL]∈{0,1}Ly = [y_1, y_2, \dots, y_L] \in \{0, 1\}^L for LL categories:

    1. A convolutional backbone FFCN\mathcal{F}_{\text{FCN}} extracts a dense feature map F=FFCN(I)∈RW×H×DF = \mathcal{F}_{\text{FCN}}(I) \in \mathbb{R}^{W \times H \times D}, where W×HW \times H is the spatial resolution and DD is the channel dimension.
    2. A class-aware 1×11 \times 1 convolutional layer FCAM\mathcal{F}_{\text{CAM}} generates intermediate class activation maps P=FCAM(F)∈RW×H×LP = \mathcal{F}_{\text{CAM}}(F) \in \mathbb{R}^{W \times H \times L}, where Pl∈RW×HP_l \in \mathbb{R}^{W \times H} corresponds to the activation map for class ll.
    3. For each category present in the image (yl=1y_l = 1), a binary mask Ml∈{0,1}W×HM_l \in \{0, 1\}^{W \times H} is constructed to isolate strongly activated pixels:

    Ml(x,y)=1(Pl(x,y)>μ),M_l(x, y) = \mathbb{1}(P_l(x, y) > \mu),

    where 1(⋅)\mathbb{1}(\cdot) is the indicator function and μ\mu is the spatial mean value of PlP_l. 4. The categorical pseudo-region representation fl∈RDf_l \in \mathbb{R}^D is computed via masked average pooling (MAP):

    fl=∑x=1W∑y=1HMl(x,y)F(x,y)∑x=1W∑y=1HMl(x,y).f_l = \frac{\sum_{x=1}^W \sum_{y=1}^H M_l(x, y) F(x, y)}{\sum_{x=1}^W \sum_{y=1}^H M_l(x, y)}.

    This compact DD-dimensional vector flf_l summarizes the appearance of class ll within image II and facilitates computationally efficient metric learning across large datasets.

  3. Knowl 3 — Pseudo-Region Memory Bank and Momentum Update Rule

    model/method

    To capture dataset-level semantic patterns across different training images without prohibitive memory overhead, RCA maintains a dynamic, non-parametric memory bank M\mathcal{M}. The memory bank is structured as LL category-specific dictionaries:

    M={M1,M2,…,ML},\mathcal{M} = \{\mathcal{M}_1, \mathcal{M}_2, \dots, \mathcal{M}_L\},

    where each dictionary Ml\mathcal{M}_l stores holistic region-aware feature embeddings ml∈RDm_l \in \mathbb{R}^D corresponding to instances of class ll across the dataset.

    During backward propagation at each training step, when an image II containing category ll (yl=1y_l = 1) is processed, its extracted pseudo-region feature fl∈RDf_l \in \mathbb{R}^D is used to update the corresponding memory entry mlm_l via a momentum update:

    ml←γml+(1−γ)fl,m_l \leftarrow \gamma m_l + (1 - \gamma) f_l,

    where γ∈[0,1)\gamma \in [0, 1) is the momentum coefficient (set by default to γ=0.99\gamma = 0.99).

    The update is conditionally executed only if the classification prediction score pl=GAP(Pl)p_l = \text{GAP}(P_l) (where GAP\text{GAP} denotes global average pooling over the intermediate CAM PlP_l) exceeds a confidence threshold ν\nu (set by default to ν=0.7\nu = 0.7). If pl≤νp_l \le \nu, the entry mlm_l remains unchanged.

    This momentum mechanism smoothly aggregates intermediate region features produced across different training epochs. Because the classifier activates different parts of an object over time, accumulating these complementary states enables mlm_l to capture progressively more complete object semantics.

  4. Knowl 4 — Regional Semantic Contrast with Region Mixup Regularization

    model/method

    Regional Semantic Contrast (RSC) optimizes dense object representations by contrasting pseudo-region embeddings against categorical embeddings stored in a global memory bank M={M1,…,ML}\mathcal{M} = \{\mathcal{M}_1, \dots, \mathcal{M}_L\}.

    For a pseudo-region feature fl∈RDf_l \in \mathbb{R}^D of category ll in an image II, the standard region-aware contrastive InfoNCE loss pulls flf_l toward positive memory embeddings of the same class {ml+∈Ml}\{m_l^+ \in \mathcal{M}_l\} and pushes it away from negative memory embeddings of other classes {ml−∈M∖Ml}\{m_l^- \in \mathcal{M} \setminus \mathcal{M}_l\}:

    LlNCE(fl,yl)=1∣Ml∣∑ml+∈Ml−log⁡exp⁡(sim(fl,ml+)/τ)exp⁡(sim(fl,ml+)/τ)+∑ml−∈M∖Mlexp⁡(sim(fl,ml−)/τ),\mathcal{L}_l^{\text{NCE}}(f_l, y_l) = \frac{1}{|\mathcal{M}_l|} \sum_{m_l^+ \in \mathcal{M}_l} -\log \frac{\exp(\text{sim}(f_l, m_l^+) / \tau)}{\exp(\text{sim}(f_l, m_l^+) / \tau) + \sum_{m_l^- \in \mathcal{M} \setminus \mathcal{M}_l} \exp(\text{sim}(f_l, m_l^-) / \tau)},

    where τ\tau is a temperature hyperparameter and sim(u,v)=u⋅v∥u∥2∥v∥2\text{sim}(u, v) = \frac{u \cdot v}{\|u\|_2 \|v\|_2} represents cosine similarity.

    Because pseudo-region features derived from weak image-level labels can be noisy or inaccurate, RSC introduces region mixup as a regularizer. A synthetic region embedding f^l\hat{f}_l is generated by linearly interpolating flf_l with a feature fl−f_{l^-} of a different class l−≠ll^- \neq l from another image in the mini-batch:

    f^l=ωfl+(1−ω)fl−,\hat{f}_l = \omega f_l + (1 - \omega) f_{l^-},

    where ω∼Beta(β,β)\omega \sim \text{Beta}(\beta, \beta) is sampled from a symmetric Beta distribution (with shape parameter β=8\beta = 8).

    The regularized region mixup contrastive loss is defined as:

    LlRM-NCE=ωLlNCE(f^l,yl)+(1−ω)LlNCE(f^l,yl−).\mathcal{L}_l^{\text{RM-NCE}} = \omega \mathcal{L}_l^{\text{NCE}}(\hat{f}_l, y_l) + (1 - \omega) \mathcal{L}_l^{\text{NCE}}(\hat{f}_l, y_{l^-}).

    The overall contrastive loss for image II is the average of LlRM-NCE\mathcal{L}_l^{\text{RM-NCE}} across all categories present in II.

  5. Knowl 5 — Regional Semantic Aggregation via Prototypical Attention

    model/method

    Regional Semantic Aggregation (RSA) enriches local image feature maps by gathering inter-image contextual knowledge from the entire dataset using prototype-based cross-attention.

    Directly attending to all entries in the large-scale memory bank M\mathcal{M} is computationally expensive and susceptible to noisy representations. RSA compresses the memory bank representations via the following steps:

    1. Prototype Extraction: For each class l∈{1,…,L}l \in \{1, \dots, L\}, kk-means clustering is performed over all memory embeddings in dictionary Ml\mathcal{M}_l at the start of each epoch to yield KK class centroids organized into matrix Ql∈RK×DQ_l \in \mathbb{R}^{K \times D} (with default K=10K = 10).
    2. Global Prototype Matrix: Prototypes from all LL classes are concatenated into a holistic prototypical representation Q=[Q1,Q2,…,QL]∈R(LK)×DQ = [Q_1, Q_2, \dots, Q_L] \in \mathbb{R}^{(LK) \times D}.
    3. Affinity Matrix Computation: For an image with backbone feature map F∈R(WH)×DF \in \mathbb{R}^{(WH) \times D}, the normalized affinity matrix S∈R(WH)×(LK)S \in \mathbb{R}^{(WH) \times (LK)} against QQ is computed as:

    S=softmax(F⊗Q⊤),S = \text{softmax}(F \otimes Q^\top),

    where ⊗\otimes denotes matrix multiplication and softmax(⋅)\text{softmax}(\cdot) normalizes along each row. 4. Context Summarization: Inter-image contextual summaries F′∈R(WH)×DF' \in \mathbb{R}^{(WH) \times D} are computed by aggregating prototype representations:

    F′=S⊗Q.F' = S \otimes Q.

    1. Feature Augmentation: F′F' is reshaped to RW×H×D\mathbb{R}^{W \times H \times D} and concatenated with the original local feature map F∈RW×H×DF \in \mathbb{R}^{W \times H \times D}:

    F^=[F,F′]∈RW×H×2D.\hat{F} = [F, F'] \in \mathbb{R}^{W \times H \times 2D}.

    1. Final CAM Generation: F^\hat{F} is fed into a second 1×11 \times 1 convolutional layer FCAM\mathcal{F}_{\text{CAM}} to generate the refined class activation map O∈RW×H×LO \in \mathbb{R}^{W \times H \times L}:

    O=FCAM(F^).O = \mathcal{F}_{\text{CAM}}(\hat{F}).

    During model deployment/inference, the full memory bank M\mathcal{M} is discarded and only the compact prototype matrix QQ is retained.

  6. Knowl 6 — Training Loss for the RCA Classification Network

    equation

    The classification network in the Regional Semantic Contrast and Aggregation (RCA) framework is trained end-to-end using a composite multi-task objective function defined over each training image II:

    L=∑I(α1LRM-NCE+α2LCE(GAP(P),y)+LCE(GAP(O),y)),\mathcal{L} = \sum_I \left( \alpha_1 \mathcal{L}^{\text{RM-NCE}} + \alpha_2 \mathcal{L}_{\text{CE}}(\text{GAP}(P), y) + \mathcal{L}_{\text{CE}}(\text{GAP}(O), y) \right),

    where:

    • LRM-NCE\mathcal{L}^{\text{RM-NCE}} is the region mixup contrastive loss averaged over all categories present in image II, enforcing metric learning on pseudo-region embeddings against a categorical memory bank.
    • P∈RW×H×LP \in \mathbb{R}^{W \times H \times L} is the intermediate class activation map generated from the raw backbone features FF.
    • O∈RW×H×LO \in \mathbb{R}^{W \times H \times L} is the final class activation map produced after regional semantic aggregation.
    • GAP(⋅)\text{GAP}(\cdot) denotes global average pooling over spatial dimensions, mapping activation maps to un-normalized class score vectors in RL\mathbb{R}^L.
    • y∈{0,1}Ly \in \{0, 1\}^L is the multi-hot ground-truth image-level label vector.
    • LCE\mathcal{L}_{\text{CE}} is the multi-label binary cross-entropy classification loss.
    • α1\alpha_1 and α2\alpha_2 are balancing hyperparameters set to α1=0.01\alpha_1 = 0.01 and α2=0.4\alpha_2 = 0.4.

    During the first training epoch (warm-up phase), contrastive supervision is disabled by setting α1=0\alpha_1 = 0.

  7. Knowl 7 — Pseudo-Label Generation and Segmentation Pipeline in RCA

    model/method

    In the RCA framework for weakly supervised semantic segmentation, full pixel-level segmentation models are trained in a two-stage pipeline:

    1. Foreground Seed Generation: Using the trained classification network with regional semantic contrast and aggregation, the final class activation map O∈RW×H×LO \in \mathbb{R}^{W \times H \times L} is computed for each training image to serve as foreground localization cues.
    2. Background Estimation: An off-the-shelf salient object detection model is applied to each training image to generate a class-agnostic saliency map providing background cues.
    3. Pseudo-Label Synthesis: Foreground activation maps and background saliency maps are fused to create initial pixel-level pseudo-labels.
    4. Boundary Refinement: Dense Conditional Random Fields (Dense CRF) are applied to the initial pseudo-masks to align object boundaries with image edges.
    5. Segmentation Network Training: The refined pseudo-labels serve as supervision to train a fully supervised semantic segmentation model (specifically DeepLabV2 with ResNet backbones).
  8. Knowl 8 — Pseudo-Segmentation Label Quality on PASCAL VOC 2012

    data/table

    The table below evaluates the quality of pseudo-segmentation labels generated on the PASCAL VOC 2012 training set measured by mean Intersection-over-Union (mIoU, %):

    Method Backbone mIoU (%)
    SS-WSSS (CVPR 2020) ResNet38 62.2
    ICD (CVPR 2020) VGG16 62.2
    SubCat (CVPR 2020) ResNet38 63.4
    CONTA (NeurIPS 2020) ResNet38 65.4
    GroupWSSS (TIP 2021) VGG16 65.7
    IRNet (CVPR 2019) ResNet50 66.5
    BES (ECCV 2020) ResNet50 67.2
    EDAM (CVPR 2021) ResNet38 68.1
    OAA++ VGG16 68.2
    RCA + OAA++ VGG16 71.4 (+3.2)
    OAA++ ResNet38 69.4
    RCA + OAA++ ResNet38 73.2 (+3.8)
    EPS (CVPR 2021) ResNet38 71.4
    RCA + EPS ResNet38 74.1 (+2.7)

    Incorporating RCA yields substantial improvements across different baseline frameworks and backbones: +3.2% mIoU over OAA++ with VGG16, +3.8% mIoU over OAA++ with ResNet38, and +2.7% mIoU over EPS with ResNet38, demonstrating that cross-image semantic contrast and aggregation improve pseudo-mask quality over single-image and pairwise localization methods.

  9. Knowl 9 — Semantic Segmentation Performance on PASCAL VOC 2012 and MS COCO 2014

    data/table

    The tables below compare the semantic segmentation performance (mIoU, %) of RCA integrated with baseline WSSS methods (OAA++ and EPS) against prior weakly supervised semantic segmentation approaches on PASCAL VOC 2012 (val and test sets) and MS COCO 2014 (val set). DeepLabV2 with a ResNet backbone is used as the segmentation model across all methods on PASCAL VOC 2012.

    PASCAL VOC 2012 val and test performance:

    Method Classification Backbone Val mIoU (%) Test mIoU (%)
    SSNet (ICCV 2019) ResNet50 63.3 64.3
    RNet (CVPR 2019) ResNet50 63.5 64.8
    CIAN (AAAI 2020) VGG16 64.3 65.3
    FickleNet (CVPR 2019) VGG16 64.9 65.3
    SSDD (ICCV 2019) ResNet38 64.9 65.5
    SEAM (CVPR 2020) ResNet38 64.5 65.7
    SubCat (CVPR 2020) ResNet38 66.1 65.9
    OAA+ (ICCV 2019) VGG16 65.2 66.4
    BES (ECCV 2020) ResNet50 65.7 66.6
    CONTA (NeurIPS 2020) ResNet38 66.1 66.7
    MCIS (ECCV 2020) VGG16 66.2 66.9
    ICD (CVPR 2020) VGG16 67.8 68.0
    CPN (ICCV 2021) ResNet38 67.8 68.5
    NSROM (CVPR 2021) VGG16 68.3 68.5
    AuxSegNet (ICCV 2021) ResNet38 69.0 68.6
    PMM (ICCV 2021) ResNet50 68.5 69.0
    GroupWSSS (TIP 2021) VGG16 68.7 69.0
    EDAM (CVPR 2021) ResNet38 70.9 70.6
    SPML (ICLR 2021) ResNet38 69.5 71.6
    OAA++ VGG16 67.7 67.4
    RCA + OAA++ VGG16 70.6 (+2.9) 71.0 (+3.6)
    OAA++ ResNet38 68.1 68.2
    RCA + OAA++ ResNet38 71.1 (+3.0) 71.6 (+3.4)
    EPS (CVPR 2021) ResNet38 70.9 70.8
    RCA + EPS ResNet38 72.2 (+1.3) 72.8 (+2.0)

    MS COCO 2014 val performance:

    Method Classification Backbone Val mIoU (%)
    BFBP (ECCV 2016) VGG16 20.4
    SEC (ECCV 2016) VGG16 22.4
    DSRG (CVPR 2018) VGG16 26.0
    IAL (IJCV 2020) VGG16 27.7
    GroupWSSS (TIP 2021) VGG16 28.7
    ADL (PAMI 2020) VGG16 30.8
    SEAM (CVPR 2020) ResNet38 32.8
    CONTA (NeurIPS 2020) ResNet38 32.8
    AuxSegNet (ICCV 2021) ResNet38 33.9
    OAA+ (ICCV 2019) VGG16 24.6
    RCA + OAA+ VGG16 26.7 (+2.1)
    EPS (CVPR 2021) VGG16 35.7
    RCA + EPS VGG16 36.8 (+1.1)
  10. Knowl 10 — Ablation Analysis of RCA Components and Hyperparameters

    empirical result

    Ablation experiments conducted on PASCAL VOC 2012 with a VGG16 classification backbone evaluate the individual components and hyperparameters of RCA:

    1. Interaction between RSC and RSA:

      • Baseline (OAA++): 68.2% pseudo-label mIoU, 67.7% val segmentation mIoU.
      • Regional Semantic Contrast only (w/ RSC): 69.5% pseudo-label mIoU (+1.3%), 69.3% val segmentation mIoU (+1.6%).
      • Regional Semantic Aggregation only (w/ RSA): 68.5% pseudo-label mIoU (+0.3%), 68.6% val segmentation mIoU (+0.9%).
      • Full model (w/ RSC and RSA): 71.4% pseudo-label mIoU (+3.2%), 70.6% val segmentation mIoU (+2.9%). While RSA alone provides limited gain, combining it with RSC yields substantial improvements, indicating that contrastive learning is essential to establish discriminative and well-structured memory embeddings before prototype aggregation can be effective.
    2. Region Mixup Regularization: Removing region mixup from the contrastive loss reduces pseudo-label mIoU from 71.4% to 70.6% (-0.8%), demonstrating its role in mitigating label noise inherent in weak annotations.

    3. Memory Momentum Coefficient γ\gamma: Pseudo-label mIoU across updating rates:

      • γ=0\gamma = 0 (no momentum): 69.9%
      • γ=0.5\gamma = 0.5: 70.9%
      • γ=0.8\gamma = 0.8: 71.2%
      • γ=0.9\gamma = 0.9: 71.2%
      • γ=0.99\gamma = 0.99 (optimal/default): 71.4%
      • γ=0.999\gamma = 0.999: 70.9% Momentum updating is crucial, with stable performance when γ∈[0.8,0.99]\gamma \in [0.8, 0.99].
    4. Number of Prototypes KK per Class:

      • K=1K = 1 (single average prototype): 70.4%
      • K=10K = 10: 71.4%
      • K=20K = 20: 71.1%
      • K=50K = 50: 71.1%
      • K=100K = 100: 71.3%
      • "all" (no clustering, using all memory entries): 70.0% K=10K = 10 provides the best trade-off between clustering out noise and preserving multi-modal intra-class diversity.
    5. Memory Capacity: Restricting the memory size to 100, 500, or all embeddings per class yields 70.8%, 71.2%, and 71.4% mIoU respectively, confirming scalability to large datasets like COCO 2014 where caching all instances is infeasible.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In CVPR, 2019.
  2. 2.Nikita Araslanov and Stefan Roth. Single-stage semantic segmentation from image labels. In CVPR, 2020.
  3. 3.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What’s the point: Semantic segmentation with point supervision. In ECCV, 2016.
  4. 4.Mathilde Caron, Piotr Bojanowski, Armand Joulin, and Matthijs Douze. Deep clustering for unsupervised learning of visual features. In ECCV, 2018.
  5. 5.Krishna Chaitanya, Ertunc Erdil, Neerav Karani, and Ender Konukoglu. Contrastive learning of global and local features for medical image segmentation with limited annotations. In NeurIPS, 2020.
  6. 6.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation via sub-category exploration. In CVPR, 2020.
  7. 7.Arslan Chaudhry, Puneet K Dokania, and Philip HS Torr. Discovering class-specific pixels for weakly-supervised semantic segmentation. arXiv preprint arXiv:1707.05821, 2017.
  8. 8.Liyi Chen, Weiwei Wu, Chenchen Fu, Xiao Han, and Yuntao Zhang. Weakly supervised semantic segmentation with boundary exploration. In ECCV, 2020.
  9. 9.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE TPAMI, 40(4):834–848, 2017.
  10. 10.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. 2020.
  11. 11.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In CVPR, 2021.
  12. 12.Yunpeng Chen, Yannis Kalantidis, Jianshu Li, Shuicheng Yan, and Jiashi Feng. A2-nets: Double attention networks. In NeurIPS, 2018.
  13. 13.Junsuk Choe, Seungho Lee, and Hyunjung Shim. Attention-based dropout layer for weakly supervised single object localization and semantic segmentation. IEEE TPAMI, 2020.
  14. 14.Jifeng Dai, Kaiming He, and Jian Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In ICCV, 2015.
  15. 15.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 88(2):303–338, 2010.
  16. 16.Junsong Fan, Zhaoxiang Zhang, Chunfeng Song, and Tieniu Tan. Learning integral objects with intra-class discriminator for weakly-supervised semantic segmentation. In CVPR, 2020.
  17. 17.Junsong Fan, Zhaoxiang Zhang, Tienniu Tan, Chunfeng Song, and Jun Xiao. Cian: Cross-image affinity net for weakly supervised semantic segmentation. In AAAI, 2020.
  18. 18.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene segmentation. In CVPR, 2019.
  19. 19.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. In NeurIPS, 2020.
  20. 20.Bharath Hariharan, Pablo Arbelaez, Lubomir Bourdev, Subhransu Maji, and Jitendra Malik. Semantic contours from inverse detectors. In ICCV, 2011.
  21. 21.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020.
  22. 22.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  23. 23.Qibin Hou, PengTao Jiang, Yunchao Wei, and Ming-Ming Cheng. Self-erasing network for integral object attention. In NeurIPS, 2018.
  24. 24.Zilong Huang, Xinggang Wang, Jiasi Wang, Wenyu Liu, and Jingdong Wang. Weakly-supervised semantic segmentation network with deep seeded region growing. In CVPR, 2018.
  25. 25.Peng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng, Yunchao Wei, and Hong-Kai Xiong. Integral object mining via online attention accumulation. In ICCV, 2019.
  26. 26.Zhenchao Jin, Tao Gong, Dongdong Yu, Qi Chu, Jian Wang, Changhu Wang, and Jie Shao. Mining contextual information beyond image for semantic segmentation. In ICCV, 2021.
  27. 27.Tsung-Wei Ke, Jyh-Jing Hwang, and Stella X Yu. Universal weakly supervised segmentation by pixel-to-segment contrastive learning. In ICLR, 2021.
  28. 28.Anna Khoreva, Rodrigo Benenson, Jan Hosang, Matthias Hein, and Bernt Schiele. Simple does it: Weakly supervised instance and semantic segmentation. In CVPR, 2017.
  29. 29.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In NeurIPS, 2020.
  30. 30.Alexander Kolesnikov and Christoph H Lampert. Seed, expand and constrain: Three principles for weakly-supervised image segmentation. In ECCV, 2016.
  31. 31.Philipp Krahenbuhl and Vladlen Koltun. Efficient inference in fully connected crfs with gaussian edge potentials. NeurIPS, 2011.
  32. 32.Krishna Kumar Singh and Yong Jae Lee. Hide-and-seek: Forcing a network to be meticulous for weakly-supervised object and action localization. In ICCV, 2017.
  33. 33.Jungbeom Lee, Eunji Kim, Sungmin Lee, Jangho Lee, and Sungroh Yoon. Ficklenet: Weakly and semi-supervised semantic image segmentation using stochastic inference. In CVPR, 2019.
  34. 34.Jungbeom Lee, Jihun Yi, Chaehun Shin, and Sungroh Yoon. Bbam: Bounding box attribution map for weakly supervised semantic and instance segmentation. In CVPR, 2021.
  35. 35.Seungho Lee, Minhyun Lee, Jongwuk Lee, and Hyunjung Shim. Railroad is not a train: Saliency as pseudo-pixel supervision for weakly supervised semantic segmentation. In CVPR, 2021.
  36. 36.Kunpeng Li, Ziyan Wu, Kuan-Chuan Peng, Jan Ernst, and Yun Fu. Tell me where to look: Guided attention inference network. In CVPR, 2018.
  37. 37.Xueyi Li, Tianfei Zhou, Jianwu Li, Yi Zhou, and Zhaoxiang Zhang. Group-wise semantic mining for weakly supervised semantic segmentation. In AAAI, 2021.
  38. 38.Yi Li, Zhanghui Kuang, Liyang Liu, Yimin Chen, and Wayne Zhang. Pseudo-mask matters inweakly-supervised semantic segmentation. In ICCV, 2021.
  39. 39.Zhiyuan Liang, Tiancai Wang, Xiangyu Zhang, Jian Sun, and Jianbing Shen. Tree energy loss: Towards sparsely annotated semantic segmentation. In CVPR, 2022.
  40. 40.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In CVPR, 2016.
  41. 41.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014.
  42. 42.Jiang-Jiang Liu, Qibin Hou, Ming-Ming Cheng, Jiashi Feng, and Jianmin Jiang. A simple pooling-based design for realtime salient object detection. In CVPR, 2019.
  43. 43.Ishan Misra and Laurens van der Maaten. Self-supervised learning of pretext-invariant representations. In CVPR, 2020.
  44. 44.Youngmin Oh, Beomjun Kim, and Bumsub Ham. Background-aware pooling and noise-aware loss for weakly-supervised semantic segmentation. In CVPR, 2021.
  45. 45.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  46. 46.Fatemehsadat Saleh, Mohammad Sadegh Aliakbarian, Mathieu Salzmann, Lars Petersson, Stephen Gould, and Jose M Alvarez. Built-in foreground/background prior for weakly-supervised semantic segmentation. In ECCV, 2016.
  47. 47.Wataru Shimoda and Keiji Yanai. Self-supervised difference detection for weakly-supervised semantic segmentation. In ICCV, 2019.
  48. 48.Mennatullah Siam, Boris N Oreshkin, and Martin Jagersand. Amp: Adaptive masked proxies for few-shot segmentation. In ICCV, 2019.
  49. 49.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  50. 50.Kihyuk Sohn. Improved deep metric learning with multiclass n-pair loss objective. In NeurIPS, 2016.
  51. 51.Chunfeng Song, Yan Huang, Wanli Ouyang, and Liang Wang. Box-driven class-wise region masking and filling rate guided loss for weakly supervised semantic segmentation. In CVPR, 2019.
  52. 52.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In ECCV, 2020.
  53. 53.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. In ECCV, 2020.
  54. 54.Paul Vernaza and Manmohan Chandraker. Learning random-walk label propagation for weakly-supervised semantic segmentation. In CVPR, 2017.
  55. 55.Binglu Wang, Yongqiang Zhao, and Xuelong Li. Multiple instance graph learning for weakly supervised remote sensing object detection. IEEE TGRS, 60:1–12, 2021.
  56. 56.Wenguan Wang, Tianfei Zhou, Fatih Porikli, David Crandall, and Luc Van Gool. A survey on deep learning technique for video segmentation. arXiv preprint arXiv:2107.01153, 2021.
  57. 57.Wenguan Wang, Tianfei Zhou, Siyuan Qi, Jianbing Shen, and Song-Chun Zhu. Hierarchical human semantic parsing with comprehensive part-relation modeling. IEEE TPAMI, 2021.
  58. 58.Wenguan Wang, Tianfei Zhou, Fisher Yu, Jifeng Dai, Ender Konukoglu, and Luc Van Gool. Exploring cross-image pixel contrast for semantic segmentation. In ICCV, 2021.
  59. 59.Xiang Wang, Sifei Liu, Huimin Ma, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation by iterative affinity learning. IJCV, pages 1–14, 2020.
  60. 60.Xiang Wang, Shaodi You, Xi Li, and Huimin Ma. Weakly-supervised semantic segmentation by iteratively mining common object features. In CVPR, 2018.
  61. 61.Xun Wang, Haozhi Zhang, Weilin Huang, and Matthew R Scott. Cross-batch memory for embedding learning. In CVPR, 2020.
  62. 62.Xinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong, and Lei Li. Dense contrastive learning for self-supervised visual pre-training. In CVPR, 2021.
  63. 63.Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In CVPR, 2020.
  64. 64.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In CVPR, 2017.
  65. 65.Yunchao Wei, Xiaodan Liang, Yunpeng Chen, Xiaohui Shen, Ming-Ming Cheng, Jiashi Feng, Yao Zhao, and Shuicheng Yan. Stc: A simple to complex framework for weakly-supervised semantic segmentation. IEEE TPAMI, 39(11):2314–2320, 2016.
  66. 66.Yunchao Wei, Huaxin Xiao, Honghui Shi, Zequn Jie, Jiashi Feng, and Thomas S Huang. Revisiting dilated convolution: A simple approach for weakly-and semi-supervised semantic segmentation. In CVPR, 2018.
  67. 67.Tong Wu, Junshi Huang, Guangyu Gao, Xiaoming Wei, Xiaolin Wei, Xuan Luo, and Chi Harold Liu. Embedded discriminative attention mechanism for weakly supervised semantic segmentation. In CVPR, 2021.
  68. 68.Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. Unsupervised feature learning via non-parametric instance discrimination. In CVPR, 2018.
  69. 69.Tong Xiao, Shuang Li, Bochao Wang, Liang Lin, and Xiaogang Wang. Joint detection and identification feature learning for person search. In CVPR, 2017.
  70. 70.Zhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao, Stephen Lin, and Han Hu. Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning. In CVPR, 2021.
  71. 71.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, Ferdous Sohel, and Dan Xu. Leveraging auxiliary tasks with affinity learning for weakly supervised semantic segmentation. In ICCV, 2021.
  72. 72.Yazhou Yao, Tao Chen, Guo-Sen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu, Zhenmin Tang, and Jian Zhang. Non-salient region object mining for weakly supervised semantic segmentation. In CVPR, 2021.
  73. 73.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In ECCV, 2020.
  74. 74.Yu Zeng, Yunzhi Zhuge, Huchuan Lu, and Lihe Zhang. Joint learning of saliency detection and weakly supervised semantic segmentation. In ICCV, 2019.
  75. 75.Dong Zhang, Hanwang Zhang, Jinhui Tang, Xiansheng Hua, and Qianru Sun. Causal intervention for weakly-supervised semantic segmentation. In NeurIPS, 2020.
  76. 76.Fei Zhang, Chaochen Gu, Chenyue Zhang, and Yuchao Dai. Complementary patch for weakly supervised semantic segmentation. In ICCV, 2021.
  77. 77.Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018.
  78. 78.Hang Zhang, Kristin Dana, Jianping Shi, Zhongyue Zhang, Xiaogang Wang, Ambrish Tyagi, and Amit Agrawal. Context encoding for semantic segmentation. In CVPR, 2018.
  79. 79.Xiaolin Zhang, Yunchao Wei, Jiashi Feng, Yi Yang, and Thomas S. Huang. Adversarial complementary learning for weakly supervised object localization. In CVPR, 2018.
  80. 80.Xiaolin Zhang, Yunchao Wei, and Yi Yang. Inter-image communication for weakly supervised localization. In ECCV, 2020.
  81. 81.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In CVPR, 2016.
  82. 82.Tianfei Zhou, Jianwu Li, Xueyi Li, and Ling Shao. Target-aware object discovery and association for unsupervised video multi-object segmentation. In CVPR, 2021.
  83. 83.Tianfei Zhou, Liulei Li, Xueyi Li, Chun-Mei Feng, Jianwu Li, and Ling Shao. Group-wise learning for weakly supervised semantic segmentation. IEEE TIP, 31:799–811, 2021.
  84. 84.Tianfei Zhou, Shunzhou Wang, Yi Zhou, Yazhou Yao, Jianwu Li, and Ling Shao. Motion-attentive transition for zero-shot video object segmentation. In AAAI, 2020.
  85. 85.Tianfei Zhou, Wenguan Wang, Si Liu, Yi Yang, and Luc Van Gool. Differentiable multi-granularity human representation learning for instance-aware human semantic parsing. In CVPR, 2021.

Citation

MLA
Zhou, T., et al. “Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation”. arXiv, 2022, http://arxiv.org/abs/2203.09653v2.
APA
Zhou, T., Zhang, M., Zhao, F., & Li, J. (2022). Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation. arXiv. http://arxiv.org/abs/2203.09653v2
Chicago
Zhou, T., M. Zhang, F. Zhao, and J. Li. 2022. “Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation”. arXiv. http://arxiv.org/abs/2203.09653v2.
Harvard
Zhou, T. et al. (2022) “Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.09653v2.
Vancouver
1. Zhou T, Zhang M, Zhao F, Li J (2022) Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation. arXiv

BibTeX

@article{zhou2022regional,
  title = {Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation},
  author = {Zhou, Tianfei and Zhang, Meijie and Zhao, Fang and Li, Jianwu},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.09653v2},
  eprint = {2203.09653}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE