Token Contrast for Weakly-Supervised Semantic Segmentation

Lixiang RuHeliang ZhengYibing ZhanBo Du

article2023CVPR164 citations

Proposes Token Contrast, a framework that counters feature over-smoothing in Vision Transformers by using intermediate layer representations and class token contrasts to produce accurate pseudo-labels for weakly-supervised semantic segmentation.

Listen

Training computer vision models for semantic segmentation—the task of classifying every pixel in an image—normally demands massive volumes of expensive, manually annotated data. Weakly-supervised semantic segmentation addresses this challenge by relying only on simple image-level tags (such as indicating whether a cat is present anywhere in an image). However, conventional approaches typically highlight only small, highly distinctive fragments of an object rather than its full shape. While modern Vision Transformer architectures capture broader context, their internal mechanics cause feature representations in later network layers to become excessively uniform—a failure mode known as over-smoothing that degrades segmentation accuracy.

The article demonstrates a novel framework called Token Contrast (ToCo) designed to resolve over-smoothing in Vision Transformers and produce highly accurate pixel-level segmentations using only image-level labels. The approach leverages two key components: a patch contrast module that uses diverse intermediate-layer representations to prevent later layers from collapsing into uniformity, and a class contrast module that compares local uncertain regions against global object features to activate less prominent object parts. The authors evaluated ToCo against standard benchmark datasets, specifically PASCAL VOC and MS COCO, under a streamlined, single-stage training setup.

The findings show that ToCo substantially outperforms existing single-stage methods and matches the performance of complex multi-stage pipelines. On the PASCAL VOC benchmark, the baseline approach generated class activation maps with an accuracy of only 27.9% mean Intersection over Union (mIoU), whereas incorporating ToCo increased this baseline to 70.5% mIoU. In final segmentation evaluations, ToCo achieved 71.1% on the PASCAL VOC validation set and 42.3% on MS COCO. This performance reached approximately 86.4% to 86.7% of the performance achieved by fully supervised models trained with complete, manual pixel-level annotations.

These results demonstrate that organizations can deploy high-performing image segmentation models while cutting the substantial labor and time expenses associated with detailed manual annotation. By eliminating the need for multi-stage pipelines, complex architectural modifications, or costly test-time gradient calculations, ToCo streamlines the model development workflow and reduces training overhead.

Organizations developing automated visual inspection or segmentation systems should consider adopting intermediate feature supervision and token contrast strategies to simplify annotation requirements. Further development is recommended to test the methodology across specialized industrial domains, evaluate real-time performance trade-offs, and refine the selection of intermediate layers when using alternative network depths. Confidence in these results is high across benchmark conditions, though practitioners should note that tuning thresholds and local crop sizes remains important for achieving optimal results across varying image domains.

Cover for Token Contrast for Weakly-Supervised Semantic Segmentation

Abstract

Weakly-Supervised Semantic Segmentation (WSSS) using image-level labels typically utilizes Class Activation Map (CAM) to generate the pseudo labels. Limited by the local structure perception of CNN, CAM usually cannot identify the integral object regions. Though the recent Vision Transformer (ViT) can remedy this flaw, we observe it also brings the over-smoothing issue, i.e., the final patch tokens incline to be uniform. In this work, we propose Token Contrast (ToCo) to address this issue and further explore the virtue of ViT for WSSS. Firstly, motivated by the observation that intermediate layers in ViT can still retain semantic diversity, we designed a Patch Token Contrast module (PTC). PTC supervises the final patch tokens with the pseudo token relations derived from intermediate layers, allowing them to align the semantic regions and thus yield more accurate CAM. Secondly, to further differentiate the low-confidence regions in CAM, we devised a Class Token Contrast module (CTC) inspired by the fact that class tokens in ViT can capture high-level semantics. CTC facilitates the representation consistency between uncertain local regions and global objects by contrasting their class tokens. Experiments on the PASCAL VOC and MS COCO datasets show the proposed ToCo can remarkably surpass other single-stage competitors and achieve comparable performance with state-of-the-art multi-stage methods. Code is available at https://github.com/rulixiang/ToCo.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 3.1. Class Activation Map
  • 3.2. Vision Transformer & Over-smoothing
  • 4. Methodology
  • 4.1. Overview
  • 4.2. Patch Token Contrast
  • 4.3. Class Token Contrast
  • 4.4. ToCo for WSSS
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Experimental Results
  • 5.3. Ablation and Analysis
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Token Contrast (ToCo) Framework for Single-Stage Weakly-Supervised Semantic Segmentation

    model/method

    Token Contrast (ToCo) is a single-stage weakly-supervised semantic segmentation (WSSS) framework based on the Vision Transformer (ViT) architecture that operates using only image-level class labels. Standard ViT models suffer from an over-smoothing problem where repeated multi-head self-attention acts as a low-pass filter, causing patch token representations in deep layers to become uniform and producing class activation maps (CAMs) that assign monotonous labels across entire images.

    ToCo overcomes over-smoothing and enhances activation completeness through two contrastive mechanisms and an end-to-end training pipeline:

    1. Auxiliary Feature Extraction: An auxiliary classification head is attached to an intermediate Transformer block (block mm, where m=10m=10 for a 12-block ViT-Base). Intermediate patch tokens Fm∈Rn×d\mathbf{F}^m \in \mathbb{R}^{n \times d} are aggregated via global max-pooling (GMP) and projected via a linear classifier parameterized by θm\theta^m to yield an auxiliary CAM Mm=CAM(Fm,θm)\mathbf{M}^m = \text{CAM}(\mathbf{F}^m, \theta^m). Dual background thresholds βl\beta_l and βh\beta_h (0<βl<βh<10 < \beta_l < \beta_h < 1) segment Mm\mathbf{M}^m into pseudo token labels Ym\mathbf{Y}^m designating reliable foreground, reliable background, and uncertain regions.

    2. Patch Token Contrast (PTC): Reliable pairwise relations derived from Ym\mathbf{Y}^m supervise the final layer patch tokens F∈Rn×d\mathbf{F} \in \mathbb{R}^{n \times d}, pulling same-label patch tokens together and pushing different-label tokens apart in absolute cosine similarity space.

    3. Class Token Contrast (CTC): Local views are randomly cropped from uncertain regions (positive proposals) and background regions (negative proposals) guided by Mm\mathbf{M}^m. The class token of the global image and those of the local crops are projected through linear projection heads and contrasted via an InfoNCE loss with exponential moving average (EMA) parameter updates, enforcing local-to-global semantic consistency.

    4. End-to-End Single-Stage Segmentation: The final CAM M=CAM(F,θ)\mathbf{M} = \text{CAM}(\mathbf{F}, \theta) produces pseudo labels that are refined online via Pixel-Adaptive Refinement (PAR) to align with low-level boundaries and directly supervise an integrated convolutional segmentation decoder.

  2. Knowl 2 — Patch Token Contrast (PTC) Module

    model/method

    The Patch Token Contrast (PTC) module counters the patch token uniformity (over-smoothing) in deep layers of a Vision Transformer (ViT) by transferring the semantic diversity preserved in intermediate layers to the final layer.

    For an input image tokenized into nn patches and passed through a ViT encoder, let Fm∈Rn×d\mathbf{F}^m \in \mathbb{R}^{n \times d} denote the output patch tokens of an intermediate layer mm, and let F∈Rn×d\mathbf{F} \in \mathbb{R}^{n \times d} denote the final layer patch tokens. An auxiliary linear classification head with weights θm\theta^m is applied to the global max-pooled intermediate tokens to compute the auxiliary Class Activation Map (CAM): Mm=CAM(Fm,θm)\mathbf{M}^m = \text{CAM}(\mathbf{F}^m, \theta^m)

    Two background thresholds βl\beta_l and βh\beta_h (0<βl<βh<10 < \beta_l < \beta_h < 1) partition Mm\mathbf{M}^m into a pseudo token label map Ym∈{0,1,…,c,uncertain}\mathbf{Y}^m \in \{0, 1, \dots, c, \text{uncertain}\}:

    • A token ii is labeled as class cc (reliable foreground) if its auxiliary activation exceeds βh\beta_h.
    • A token ii is labeled as background if its activation is below βl\beta_l.
    • Tokens with intermediate activations [βl,βh][\beta_l, \beta_h] are labeled as uncertain.

    Pairwise relations are constructed strictly over reliable tokens (ignoring uncertain tokens): pairs sharing the same semantic class label (Yim=YjmY_i^m = Y_j^m) form the positive pair set P\mathcal{P}, while pairs with distinct labels (Yim≠YjmY_i^m \neq Y_j^m) form the negative pair set N\mathcal{N}. PTC applies contrastive supervision on the final patch tokens F\mathbf{F} to maximize cosine similarity for pairs in P\mathcal{P} and minimize absolute cosine similarity for pairs in N\mathcal{N}.

  3. Knowl 3 — Patch Token Contrast Loss Function

    equation

    Given final-layer patch token representations F=[F1,F2,…,Fn]⊤∈Rn×d\mathbf{F} = [\mathbf{F}_1, \mathbf{F}_2, \dots, \mathbf{F}_n]^\top \in \mathbb{R}^{n \times d} and intermediate-layer pseudo token labels Ym=[Y1m,…,Ynm]⊤\mathbf{Y}^m = [Y_1^m, \dots, Y_n^m]^\top, the Patch Token Contrast (PTC) loss is formulated as:

    Lptc=1N+∑Yim=Yjm(1−∣CosSim(Fi,Fj)∣)+1N−∑Yim≠Yjm∣CosSim(Fi,Fj)∣\mathcal{L}_{ptc} = \frac{1}{N^+} \sum_{Y_i^m = Y_j^m} \left(1 - |\text{CosSim}(\mathbf{F}_i, \mathbf{F}_j)|\right) + \frac{1}{N^-} \sum_{Y_i^m \neq Y_j^m} |\text{CosSim}(\mathbf{F}_i, \mathbf{F}_j)|

    where:

    • CosSim(Fi,Fj)=Fi⊤Fj∥Fi∥2∥Fj∥2\text{CosSim}(\mathbf{F}_i, \mathbf{F}_j) = \frac{\mathbf{F}_i^\top \mathbf{F}_j}{\|\mathbf{F}_i\|_2 \|\mathbf{F}_j\|_2} denotes the cosine similarity between the ii-th and jj-th final patch tokens.
    • ∣⋅∣|\cdot| denotes the absolute value operation, used because minimizing signed cosine similarity allows tokens with negative cosine similarity (e.g., −1-1) to remain collinear and correlated, whereas absolute cosine similarity enforces true linear independence and feature diversity for negative pairs.
    • N+N^+ is the total number of positive token pairs (Yim=YjmY_i^m = Y_j^m) belonging to reliable foreground or background classes.
    • N−N^- is the total number of negative token pairs (Yim≠YjmY_i^m \neq Y_j^m) belonging to reliable foreground or background classes.
    • Uncertain tokens (where YimY_i^m falls between background threshold βl\beta_l and foreground threshold βh\beta_h) are excluded from both summations.
  4. Knowl 4 — Class Token Contrast (CTC) Module

    model/method

    The Class Token Contrast (CTC) module activates non-salient and low-confidence object regions by enforcing representation consistency between global objects and local ambiguous regions using the ViT class token ([CLS][\text{CLS}]).

    1. Proposal Generation and Cropping: Guided by the auxiliary CAM Mm\mathbf{M}^m, local image patches are randomly cropped from the global input image XX. Crops falling inside uncertain regions (where activation is in [βl,βh][\beta_l, \beta_h]) are designated as positive local proposals X+X^+, while crops falling strictly inside background regions (activation <βl< \beta_l) are designated as negative local proposals X−X^-.

    2. Projection and Class Token Extraction: The global image XX and cropped local images X+/X−X^+ / X^- are fed into the ViT encoder to extract their respective class tokens. The global class token is projected via global head PgP^g, yielding normalized vector p∈Rk\mathbf{p} \in \mathbb{R}^k. The local class tokens are projected via local head PlP^l, yielding normalized sets Q+={q+}\mathcal{Q}^+ = \{\mathbf{q}^+\} and Q−={q−}\mathcal{Q}^- = \{\mathbf{q}^-\}. Both projection heads comprise 3 linear layers followed by an L2L_2-normalization layer.

    3. Asymmetric Optimization with EMA: To guide the local representations toward the global view without representation collapse, gradients through PgP^g are stopped. The parameters θg\theta^g of PgP^g are updated using an Exponential Moving Average (EMA) of the parameters θl\theta^l of PlP^l: θg←ρθg+(1−ρ)θl\theta^g \leftarrow \rho \theta^g + (1 - \rho)\theta^l where ρ∈[0,1)\rho \in [0, 1) is the momentum factor (default ρ=0.9\rho = 0.9).

  5. Knowl 5 — Class Token Contrast InfoNCE Loss

    equation

    The Class Token Contrast (CTC) module is optimized using the InfoNCE contrastive loss over the projected L2L_2-normalized class tokens:

    Lctc=1N+∑q+∈Q+−log⁡exp⁡(p⊤q+/τ)exp⁡(p⊤q+/τ)+∑q−∈Q−exp⁡(p⊤q−/τ)+ϵ\mathcal{L}_{ctc} = \frac{1}{N^+} \sum_{\mathbf{q}^+ \in \mathcal{Q}^+} -\log \frac{\exp(\mathbf{p}^\top \mathbf{q}^+ / \tau)}{\exp(\mathbf{p}^\top \mathbf{q}^+ / \tau) + \sum_{\mathbf{q}^- \in \mathcal{Q}^-} \exp(\mathbf{p}^\top \mathbf{q}^- / \tau) + \epsilon}

    where:

    • p∈Rk\mathbf{p} \in \mathbb{R}^k is the L2L_2-normalized global class token projected by PgP^g, with gradients detached.
    • Q+\mathcal{Q}^+ is the set of N+N^+ projected L2L_2-normalized local class tokens cropped from uncertain regions of the auxiliary CAM.
    • Q−\mathcal{Q}^- is the set of projected L2L_2-normalized local class tokens cropped from background regions.
    • τ>0\tau > 0 is a temperature scaling parameter (set to 0.50.5).
    • ϵ>0\epsilon > 0 is a small positive constant added for numerical stability.
  6. Knowl 6 — Overall Optimization Objective for Single-Stage Token Contrast

    equation

    The overall end-to-end training loss L\mathcal{L} of the single-stage Token Contrast (ToCo) framework integrates image classification, token contrast, online segmentation, and spatial regularization:

    L=Lcls+Lclsm+λ1Lptc+λ2Lctc+λ3Lseg+Lreg\mathcal{L} = \mathcal{L}_{cls} + \mathcal{L}_{cls}^m + \lambda_1 \mathcal{L}_{ptc} + \lambda_2 \mathcal{L}_{ctc} + \lambda_3 \mathcal{L}_{seg} + \mathcal{L}_{reg}

    where:

    • Lcls\mathcal{L}_{cls} is the multi-label soft margin classification loss computed on the final classification head.
    • Lclsm\mathcal{L}_{cls}^m is the auxiliary multi-label soft margin classification loss computed on the intermediate classification head at block mm.
    • Lptc\mathcal{L}_{ptc} is the Patch Token Contrast loss regularizing final patch token similarities against intermediate token relations.
    • Lctc\mathcal{L}_{ctc} is the Class Token Contrast loss aligning local uncertain class tokens with the global class token.
    • Lseg\mathcal{L}_{seg} is the cross-entropy segmentation loss supervising the convolutional decoder using pseudo masks generated from the final CAM after Pixel-Adaptive Refinement (PAR).
    • Lreg\mathcal{L}_{reg} is an auxiliary spatial regularization loss enforcing boundary and spatial consistency in the predicted segmentation masks.
    • The balancing hyperparameters are set to λ1=0.2\lambda_1 = 0.2, λ2=0.5\lambda_2 = 0.5, and λ3=0.1\lambda_3 = 0.1.
  7. Knowl 7 — Component Ablation Study of ToCo

    empirical result

    An ablation study evaluated on the PASCAL VOC 2012 validation set demonstrates the incremental performance contributions of each module in the ToCo framework, tracking the quality (mIoU) of the final CAM M\mathbf{M}, the auxiliary intermediate CAM Mm\mathbf{M}^m, and the final semantic segmentation mask Seg.\text{Seg.}

    Method PTC CTC Lseg\mathcal{L}_{seg} Lreg\mathcal{L}_{reg} M\mathbf{M} Mm\mathbf{M}^m Seg.\text{Seg.}
    Lcls\mathcal{L}_{cls} 27.8 – –
    Lcls+Lclsm\mathcal{L}_{cls} + \mathcal{L}_{cls}^m 27.9 53.8 –
    Lcls+Lclsm\mathcal{L}_{cls} + \mathcal{L}_{cls}^m ✓ 62.5 57.8 –
    Lcls+Lclsm\mathcal{L}_{cls} + \mathcal{L}_{cls}^m ✓ ✓ 67.2 60.7 –
    Lcls+Lclsm\mathcal{L}_{cls} + \mathcal{L}_{cls}^m ✓ ✓ ✓ 69.9 61.2 66.6
    Lcls+Lclsm\mathcal{L}_{cls} + \mathcal{L}_{cls}^m ✓ ✓ ✓ ✓ 70.5 62.5 68.1

    Key takeaways from the ablation:

    1. Training ViT-Base with standard classification loss Lcls\mathcal{L}_{cls} achieves only 27.8%27.8\% CAM mIoU due to severe patch over-smoothing.
    2. Adding an intermediate classifier (Lclsm\mathcal{L}_{cls}^m) yields an intermediate CAM Mm\mathbf{M}^m of 53.8%53.8\% mIoU, but the final CAM M\mathbf{M} remains degraded at 27.9%27.9\%.
    3. Introducing PTC bridges this gap, increasing final CAM mIoU from 27.9%27.9\% to 62.5%62.5\% (+34.6%+34.6\% mIoU).
    4. Adding CTC further improves final CAM mIoU by +4.7%+4.7\% to 67.2%67.2\%.
    5. Online segmentation supervision with PAR (Lseg\mathcal{L}_{seg}) and regularization (Lreg\mathcal{L}_{reg}) brings final CAM mIoU to 70.5%70.5\% and semantic segmentation performance to 68.1%68.1\% mIoU (without dense CRF or multi-scale evaluation).
  8. Knowl 8 — Pseudo Label Quality Comparison on PASCAL VOC 2012

    empirical result

    The quality of pseudo labels generated by ToCo was evaluated in terms of mean Intersection-over-Union (mIoU, %) on the PASCAL VOC 2012 train and val sets against state-of-the-art weakly-supervised methods:

    Method Backbone train val
    RRM (AAAI'2020) WR38 – 65.4
    1Stage (CVPR'2020) WR38 66.9 65.3
    AALR (ACM MM'2021) WR38 68.2 65.8
    SLRNet (IJCV'2022) WR38 67.1 66.2
    AFA (CVPR'2022) MiT-B1 68.7 66.5
    ViT-PCM (ECCV'2022) ViT-B†^\dagger 67.7 66.0
    ViT-PCM + CRF (ECCV'2022) ViT-B†^\dagger 71.4 69.3
    ToCo ViT-B 72.2 70.5
    ToCo†^\dagger ViT-B†^\dagger 73.6 72.3

    Here, †\dagger denotes models initialized with ImageNet-21k pretrained weights, while untagged models use ImageNet-1k weights (DeiT for ViT-B).

    ToCo with ImageNet-1k pretraining achieves 72.2%72.2\% mIoU on train and 70.5%70.5\% on val, outperforming all prior single-stage methods as well as ViT-PCM with ImageNet-21k and CRF post-processing. With ImageNet-21k pretraining (ToCo†^\dagger), pseudo label mIoU reaches 73.6%73.6\% on train and 72.3%72.3\% on val.

  9. Knowl 9 — Weakly-Supervised Semantic Segmentation Benchmarks on PASCAL VOC 2012 and MS COCO 2014

    empirical result

    Semantic segmentation performance was evaluated using mIoU (%) on the PASCAL VOC 2012 (val and test) and MS COCO 2014 (val) benchmarks under image-level supervision (II) and saliency-assisted supervision (I+SI+S):

    Method Sup. Net. VOC COCO
    val test val
    Multi-stage WSSS methods
    RIB (NeurIPS'2021) I+SI+S DL-V2 70.2 70.0 –
    EPS (CVPR'2021) I+SI+S DL-V2 71.0 71.8 –
    L2G (CVPR'2022) I+SI+S DL-V2 72.1 71.7 44.2
    RCA (CVPR'2022) I+SI+S DL-V2 72.2 72.8 36.8
    Du et al. (CVPR'2022) I+SI+S DL-V2 72.6 73.6 –
    RIB (NeurIPS'2021) II DL-V2 68.3 68.6 43.8
    ReCAM (CVPR'2022) II DL-V2 68.4 68.2 45.0
    VWL (IJCV'2022) II DL-V2 69.2 69.2 36.2
    W-OoD (CVPR'2022) II WR38 70.7 70.1 –
    MCTformer (CVPR'2022) II WR38 71.9 71.6 42.0
    ESOL (NeurIPS'2022) II DL-V2 69.9 69.3 42.6
    Single-stage WSSS methods
    RRM (AAAI'2020) II WR38 62.6 62.9 –
    1Stage (CVPR'2020) II WR38 62.7 64.3 –
    AFA (CVPR'2022) II MiT-B1 66.0 66.3 38.9
    SLRNet (IJCV'2022) II WR38 67.2 67.6 35.0
    ToCo II ViT-B 69.8 70.5 41.3
    ToCo†^\dagger II ViT-B†^\dagger 71.1 72.2 42.3

    Here, †\dagger denotes ImageNet-21k pretraining. Evaluation includes multi-scale testing and dense CRF post-processing.

    ToCo establishes new state-of-the-art results for single-stage WSSS, scoring 69.8%69.8\% / 70.5%70.5\% (ImageNet-1k) and 71.1%71.1\% / 72.2%72.2\% (ImageNet-21k) on VOC 2012 val/test, and 42.3%42.3\% on COCO val. This matches or surpasses multi-stage pipelines using image-level labels only and approaches multi-stage methods requiring external saliency maps.

  10. Knowl 10 — Hyperparameter Sensitivity for Auxiliary Block Depth, Local Crop Size, and EMA Momentum

    empirical result

    Ablation experiments on hyperparameter configurations evaluated on the PASCAL VOC 2012 validation set reveal optimal choices for the auxiliary classifier block index, local crop size in CTC, and EMA momentum factor ρ\rho:

    Block M\mathbf{M} Mm\mathbf{M}^m Seg.\text{Seg.} Size M\mathbf{M} Seg.\text{Seg.} Momentum ρ\rho M\mathbf{M} Seg.\text{Seg.}
    #8 63.5 48.1 60.5 64264^2 68.1 65.5 0 69.2 66.9
    #9 67.6 55.1 64.8 80280^2 69.3 67.3 0.1 69.0 67.1
    #10 70.5 62.5 68.1 96296^2 70.5 68.1 0.9 70.5 68.1
    #11 43.1 45.2 40.3 1122112^2 69.4 67.0 0.99 68.7 66.6
    1. Auxiliary Block Index: For a 12-block ViT-Base, placing the auxiliary head at block #10 achieves peak performance (70.5%70.5\% CAM mIoU, 68.1%68.1\% Seg mIoU). Earlier blocks (#8, #9) lack sufficient high-level semantic abstraction, while later blocks (#11) already suffer from over-smoothing.
    2. Crop Size: A crop size of 96×9696 \times 96 yields optimal performance. Smaller crops (64264^2) often miss foreground content, while larger crops (1122112^2) contain excessive discriminative foreground, reducing the contrastive benefit for uncertain boundary regions.
    3. EMA Momentum: ρ=0.9\rho = 0.9 achieves the highest CAM and segmentation accuracy, while the model demonstrates stable performance across ρ∈[0,0.99]\rho \in [0, 0.99].
  11. Knowl 11 — Supervised Performance Retention in Single-Stage WSSS

    empirical result

    The efficacy of single-stage weakly-supervised semantic segmentation methods was evaluated by measuring their performance ratio relative to their fully-supervised counterparts on the PASCAL VOC 2012 validation set:

    Method Backbone val (Fully Supervised) val (Image-level) Ratio
    1Stage WR38 80.8 62.7 77.59%
    AFA MiT-B1 78.7 66.0 83.86%
    SLRNet WR38 80.8 67.2 83.17%
    ToCo ViT-B 80.5 69.8 86.71%
    ToCo†^\dagger ViT-B†^\dagger 82.3 71.1 86.39%

    Here, †\dagger denotes ImageNet-21k pretraining.

    While WideResNet38 (WR38) achieves fully-supervised accuracy (80.8%80.8\%) comparable to ViT-Base (80.5%80.5\%), ToCo achieves 69.8%69.8\% mIoU under weak supervision, retaining 86.71%86.71\% of its upper-bound performance. In contrast, SLRNet and AFA retain only 83.17%83.17\% and 83.86%83.86\%, demonstrating that ToCo bridges the gap between weak and full supervision more effectively.

Coverage note — None. All primary contributions, modules (PTC, CTC, single-stage framework), theoretical justifications, loss formulations, and empirical findings were extracted.

References

  1. 1.Jiwoon Ahn, Sunghyun Cho, and Suha Kwak. Weakly supervised learning of instance segmentation with inter-pixel relations. In CVPR, pages 2209–2218, 2019.
  2. 2.Jiwoon Ahn and Suha Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In CVPR, pages 4981–4990, 2018.
  3. 3.Nikita Araslanov and Stefan Roth. Single-stage semantic segmentation from image labels. In CVPR, pages 4253–4262, 2020.
  4. 4.Amy Bearman, Olga Russakovsky, Vittorio Ferrari, and Li Fei-Fei. What’s the point: Semantic segmentation with point supervision. In ECCV, pages 549–565. Springer, 2016.
  5. 5.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In ECCV, pages 213–229. Springer, 2020.
  6. 6.Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In ICCV, pages 9650–9660, 2021.
  7. 7.Yu-Ting Chang, Qiaosong Wang, Wei-Chih Hung, Robinson Piramuthu, Yi-Hsuan Tsai, and Ming-Hsuan Yang. Weakly-supervised semantic segmentation via sub-category exploration. In CVPR, pages 8991–9000, 2020.
  8. 8.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. TPAMI, 40(4):834–848, 2017.
  9. 9.Qi Chen, Lingxiao Yang, Jianhuang Lai, and Xiaohua Xie. Self-supervised image-specific prototype exploration for weakly supervised semantic segmentation. CVPR, 2022.
  10. 10.Tianlong Chen, Zhenyu Zhang, Yu Cheng, Ahmed Awadallah, and Zhangyang Wang. The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy. In CVPR, pages 12020–12030, 2022.
  11. 11.Zhaozheng Chen, Tan Wang, Xiongwei Wu, Xian-Sheng Hua, Hanwang Zhang, and Qianru Sun. Class re-activation maps for weakly-supervised semantic segmentation. CVPR, 2022.
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  13. 13.Ye Du, Zehua Fu, Qingjie Liu, and Yunhong Wang. Weakly supervised semantic segmentation by pixel-to-prototype contrast. In CVPR, pages 4320–4329, 2022.
  14. 14.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 88(2):303–338, 2010.
  15. 15.Wei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng, Qi Tian, Zhenjun Han, Bolei Zhou, and Qixiang Ye. Ts-cam: Token semantic coupled attention map for weakly supervised object localization. In ICCV, pages 2886–2895, October 2021.
  16. 16.Chengyue Gong, Dilin Wang, Meng Li, Vikas Chandra, and Qiang Liu. Vision transformers with patch diversification. arXiv preprint arXiv:2104.12753, 2021.
  17. 17.Bharath Hariharan, Pablo Arbeláez, Lubomir Bourdev, Subhransu Maji, and Jitendra Malik. Semantic contours from inverse detectors. In ICCV, pages 991–998, 2011.
  18. 18.Peng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng, Yunchao Wei, and Hong-Kai Xiong. Integral object mining via online attention accumulation. In ICCV, pages 2070–2079, 2019.
  19. 19.Peng-Tao Jiang, Yuqi Yang, Qibin Hou, and Yunchao Wei. L2g: A simple local-to-global knowledge transfer framework for weakly supervised semantic segmentation. In CVPR, pages 16886–16896, 2022.
  20. 20.Meng Lan, Jing Zhang, Lefei Zhang, and Dacheng Tao. Learning to learn better for video object segmentation. arXiv preprint arXiv:2212.02112, 2022.
  21. 21.Jungbeom Lee, Jooyoung Choi, Jisoo Mok, and Sungroh Yoon. Reducing information bottleneck for weakly supervised semantic segmentation. NeurIPS, 34, 2021.
  22. 22.Jungbeom Lee, Seong Joon Oh, Sangdoo Yun, Junsuk Choe, Eunji Kim, and Sungroh Yoon. Weakly supervised semantic segmentation using out-of-distribution data. CVPR, 2022.
  23. 23.Jungbeom Lee, Jihun Yi, Chaehun Shin, and Sungroh Yoon. Bbam: Bounding box attribution map for weakly supervised semantic and instance segmentation. In CVPR, pages 2643–2652, 2021.
  24. 24.Seungho Lee, Minhyun Lee, Jongwuk Lee, and Hyunjung Shim. Railroad is not a train: Saliency as pseudo-pixel supervision for weakly supervised semantic segmentation. In CVPR, pages 5495–5505, 2021.
  25. 25.Jinlong Li, Zequn Jie, Xu Wang, Xiaolin Wei, and Lin Ma. Expansion and shrinkage of localization for weakly-supervised semantic segmentation. NeurIPS, 2022.
  26. 26.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, pages 740–755. Springer, 2014.
  27. 27.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  28. 28.Junwen Pan, Pengfei Zhu, Kaihua Zhang, Bing Cao, Yu Wang, Dingwen Zhang, Junwei Han, and Qinghua Hu. Learning self-supervised low-rank network for single-stage weakly and semi-supervised semantic segmentation. IJCV, 130(5):1181–1195, 2022.
  29. 29.Namuk Park and Songkuk Kim. How do vision transformers work? In ICLR, 2021.
  30. 30.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, pages 8748–8763. PMLR, 2021.
  31. 31.Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. Imagenet-21k pretraining for the masses. arXiv preprint arXiv:2104.10972, 2021.
  32. 32.Simone Rossetti, Damiano Zappia, Marta Sanzari, Marco Schaerf, and Fiora Pirri. Max pooling with vision transformers reconciles class and shape in weakly supervised semantic segmentation. In ECCV, 2022.
  33. 33.Lixiang Ru, Bo Du, Yibing Zhan, and Chen Wu. Weakly-supervised semantic segmentation with visual words learning and hybrid pooling. IJCV, 130(4):1127–1144, 2022.
  34. 34.Lixiang Ru, Yibing Zhan, Baosheng Yu, and Bo Du. Learning affinity from attention: End-to-end weakly-supervised semantic segmentation with transformers. In CVPR, 2022.
  35. 35.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV, pages 618–626, 2017.
  36. 36.Han Shi, Jiahui Gao, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen MS Lee, and James Kwok. Revisiting over-smoothing in bert from the perspective of graph. In ICLR, 2021.
  37. 37.Yukun Su, Ruizhou Sun, Guosheng Lin, and Qingyao Wu. Context decoupling augmentation for weakly supervised semantic segmentation. In ICCV, pages 7004–7014, 2021.
  38. 38.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In ECCV, pages 347–365. Springer, 2020.
  39. 39.Weixuan Sun, Jing Zhang, Zheyuan Liu, Yiran Zhong, and Nick Barnes. Getam: Gradient-weighted element-wise transformer attention map for weakly-supervised semantic segmentation. arXiv preprint arXiv:2112.02841, 2021.
  40. 40.Meng Tang, Federico Perazzi, Abdelaziz Djelouah, Ismail Ben Ayed, Christopher Schroers, and Yuri Boykov. On regularized losses for weakly-supervised cnn segmentation. In ECCV, pages 507–522, 2018.
  41. 41.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. Training data-efficient image transformers & distillation through attention. In ICML, pages 10347–10357. PMLR, 2021.
  42. 42.Peihao Wang, Wenqing Zheng, Tianlong Chen, and Zhangyang Wang. Anti-oversmoothing in deep vision transformers via the fourier domain analysis: From theory to practice. In ICLR, 2021.
  43. 43.Yude Wang, Jie Zhang, Meina Kan, Shiguang Shan, and Xilin Chen. Self-supervised equivariant attention mechanism for weakly supervised semantic segmentation. In CVPR, pages 12275–12284, 2020.
  44. 44.Yunchao Wei, Jiashi Feng, Xiaodan Liang, Ming-Ming Cheng, Yao Zhao, and Shuicheng Yan. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In CVPR, pages 1568–1576, 2017.
  45. 45.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. NeurIPS, 2021.
  46. 46.Jinheng Xie, Xianxu Hou, Kai Ye, and Linlin Shen. Cross language image matching for weakly supervised semantic segmentation. CVPR, 2022.
  47. 47.Lian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaid, and Dan Xu. Multi-class token transformer for weakly supervised semantic segmentation. CVPR, 2022.
  48. 48.Yangyang Xu, Yibo Yang, and Lefei Zhang. Demt: Deformable mixer transformer for multi-task learning of dense prediction. arXiv e-prints, pages arXiv–2301, 2023.
  49. 49.Yazhou Yao, Tao Chen, Guo-Sen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu, Zhenmin Tang, and Jian Zhang. Non-salient region object mining for weakly supervised semantic segmentation. In CVPR, pages 2623–2632, 2021.
  50. 50.Bingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun, and Kaizhu Huang. Reliability does matter: An end-to-end weakly supervised semantic segmentation approach. In AAAI, pages 12765–12772, 2020.
  51. 51.Bingfeng Zhang, Jimin Xiao, and Yao Zhao. Dynamic feature regularized loss for weakly supervised semantic segmentation. arXiv preprint arXiv:2108.01296, 2021.
  52. 52.Xiangrong Zhang, Zelin Peng, Peng Zhu, Tianyang Zhang, Chen Li, Huiyu Zhou, and Licheng Jiao. Adaptive affinity loss and erroneous pseudo-label refinement for weakly supervised semantic segmentation. In ACM MM, 2021.
  53. 53.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In CVPR, pages 2921–2929, 2016.
  54. 54.Tianfei Zhou, Meijie Zhang, Fang Zhao, and Jianwu Li. Regional semantic contrast and aggregation for weakly supervised semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4299–4309, 2022.

Citation

MLA
Ru, L., et al. “Token Contrast for Weakly-Supervised Semantic Segmentation”. arXiv, 2023, http://arxiv.org/abs/2303.01267v1.
APA
Ru, L., Zheng, H., Zhan, Y., & Du, B. (2023). Token Contrast for Weakly-Supervised Semantic Segmentation. arXiv. http://arxiv.org/abs/2303.01267v1
Chicago
Ru, L., H. Zheng, Y. Zhan, and B. Du. 2023. “Token Contrast for Weakly-Supervised Semantic Segmentation”. arXiv. http://arxiv.org/abs/2303.01267v1.
Harvard
Ru, L. et al. (2023) “Token Contrast for Weakly-Supervised Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.01267v1.
Vancouver
1. Ru L, Zheng H, Zhan Y, Du B (2023) Token Contrast for Weakly-Supervised Semantic Segmentation. arXiv

BibTeX

@article{ru2023token,
  title = {Token Contrast for Weakly-Supervised Semantic Segmentation},
  author = {Ru, Lixiang and Zheng, Heliang and Zhan, Yibing and Du, Bo},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.01267v1},
  eprint = {2303.01267}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE