Built independently by an author, for readers. Read the story and support ChapterPal

keyword

self-supervised ViT

A self-supervised Vision Transformer (ViT) is a deep learning model based on the transformer architecture that learns visual representations from unlabeled image data without relying on human annotations. Instead of using predefined class labels, it is pre-trained through self-supervised objectives such as contrastive learning, self-distillation, or masked patch reconstruction, which encourage the network to discover visual patterns and underlying structures autonomously. By dividing images into discrete patches and processing them with multi-head self-attention mechanisms, a self-supervised ViT captures both global context and fine-grained spatial relationships, often naturally producing attention maps that delineate object boundaries and group semantically consistent regions. Consequently, these models generate versatile feature representations that serve as robust foundations for various downstream computer vision tasks, including image classification, object detection, and dense unsupervised semantic segmentation.

1 item

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

Kehan Li, Zhennan Wang, Zesen Cheng, Runyi Yu, Yian Zhao, Guoli Song, Chang Liu, Li Yuan, Jie Chen

OrganizationsDalian University of TechnologyPeking UniversityPeng Cheng LaboratoryTsinghua University

Why you should read this

Proposes an unsupervised semantic segmentation framework that dynamically maps learnable prototypes into image-specific semantic concepts using attention mechanisms and a modularity loss, overcoming over- and under-clustering issues without requiring manual annotations.

Recently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg.

Added

2026-09-26