Token Contrast for Weakly-Supervised Semantic Segmentation
Lixiang RuHeliang ZhengYibing ZhanBo Du
Proposes Token Contrast, a framework that counters feature over-smoothing in Vision Transformers by using intermediate layer representations and class token contrasts to produce accurate pseudo-labels for weakly-supervised semantic segmentation.
Training computer vision models for semantic segmentation—the task of classifying every pixel in an image—normally demands massive volumes of expensive, manually annotated data. Weakly-supervised semantic segmentation addresses this challenge by relying only on simple image-level tags (such as indicating whether a cat is present anywhere in an image). However, conventional approaches typically highlight only small, highly distinctive fragments of an object rather than its full shape. While modern Vision Transformer architectures capture broader context, their internal mechanics cause feature representations in later network layers to become excessively uniform—a failure mode known as over-smoothing that degrades segmentation accuracy.
The article demonstrates a novel framework called Token Contrast (ToCo) designed to resolve over-smoothing in Vision Transformers and produce highly accurate pixel-level segmentations using only image-level labels. The approach leverages two key components: a patch contrast module that uses diverse intermediate-layer representations to prevent later layers from collapsing into uniformity, and a class contrast module that compares local uncertain regions against global object features to activate less prominent object parts. The authors evaluated ToCo against standard benchmark datasets, specifically PASCAL VOC and MS COCO, under a streamlined, single-stage training setup.
The findings show that ToCo substantially outperforms existing single-stage methods and matches the performance of complex multi-stage pipelines. On the PASCAL VOC benchmark, the baseline approach generated class activation maps with an accuracy of only 27.9% mean Intersection over Union (mIoU), whereas incorporating ToCo increased this baseline to 70.5% mIoU. In final segmentation evaluations, ToCo achieved 71.1% on the PASCAL VOC validation set and 42.3% on MS COCO. This performance reached approximately 86.4% to 86.7% of the performance achieved by fully supervised models trained with complete, manual pixel-level annotations.
These results demonstrate that organizations can deploy high-performing image segmentation models while cutting the substantial labor and time expenses associated with detailed manual annotation. By eliminating the need for multi-stage pipelines, complex architectural modifications, or costly test-time gradient calculations, ToCo streamlines the model development workflow and reduces training overhead.
Organizations developing automated visual inspection or segmentation systems should consider adopting intermediate feature supervision and token contrast strategies to simplify annotation requirements. Further development is recommended to test the methodology across specialized industrial domains, evaluate real-time performance trade-offs, and refine the selection of intermediate layers when using alternative network depths. Confidence in these results is high across benchmark conditions, though practitioners should note that tuning thresholds and local crop sizes remains important for achieving optimal results across varying image domains.
- Paper: Bridging the Gap between Classification and Localization for Weakly Supervised Object Localization, Eunji Kim et al. (2022). This paper analyzes the intrinsic limitations of Class Activation Maps in weakly supervised localization, establishing the foundational problem that Token Contrast aims to resolve using Vision Transformers.
- Paper: Do Vision Transformers See Like Convolutional Neural Networks?, Maithra Raghu et al. (2021). This work analyzes the internal representation structure and token behavior of Vision Transformers versus CNNs, providing the empirical basis for understanding token over-smoothing in dense prediction tasks.
- Paper: Segmenter: Transformer for Semantic Segmentation, Robin Strudel et al. (2021). This paper introduces pure Vision Transformer architectures for semantic segmentation, providing the core framework for dense patch and class token representation.
- Paper: Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers, Sixiao Zheng et al. (2020). This work formulates semantic segmentation as a sequence-to-sequence prediction task using Transformers, establishing how patch tokens model global context across an entire image.
- Paper: SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers, Enze Xie et al. (2021). This paper presents a Transformer-based semantic segmentation architecture that demonstrates how hierarchical multi-layer token representations capture varying spatial and semantic scales.
- Paper: Out-of-Candidate Rectification for Weakly Supervised Semantic Segmentation, Zesen Cheng et al. (2023). This work tackles out-of-candidate localization errors in weakly supervised semantic segmentation, providing a complementary post-generation rectification mechanism for CAM-derived pseudo labels.
- Paper: Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor, Hyeokjun Kweon et al. (2023). This paper presents an adversarial learning framework to refine class activation maps and delineate precise boundaries in weakly supervised semantic segmentation.
- Paper: Vision Transformers Need Registers, Timothée Darcet et al. (2024). This paper investigates why Vision Transformers develop feature artifacts and over-smoothing across deep layers, offering an architectural register solution to preserve local token quality.
- Paper: Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning, Jishnu Mukhoti et al. (2023). This work extends token-level contrastive alignment to open-vocabulary semantic segmentation by aligning patch tokens with text class representations.
