Built independently by an author, for readers. Read the story and support ChapterPal

keyword

region-level representations

Region-level representations are feature embeddings in computer vision that capture the visual, spatial, and semantic attributes of specific coherent sub-areas or segments within an image. Unlike global representations that summarize an entire image or pixel-level representations that describe individual coordinate points, region-level representations aggregate localized visual information corresponding to distinct objects, parts, or semantic concepts. They are typically generated by pooling, clustering, or projecting features across bounded areas, masks, or contiguous groups of pixels that share visual characteristics. These representations serve as a crucial intermediate level of abstraction in dense visual perception tasks, facilitating localized reasoning, object detection, instance segmentation, and semantic segmentation by bridging granular pixel data with higher-level conceptual understanding.

1 item

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

Kehan Li, Zhennan Wang, Zesen Cheng, Runyi Yu, Yian Zhao, Guoli Song, Chang Liu, Li Yuan, Jie Chen

OrganizationsDalian University of TechnologyPeking UniversityPeng Cheng LaboratoryTsinghua University

Why you should read this

Proposes an unsupervised semantic segmentation framework that dynamically maps learnable prototypes into image-specific semantic concepts using attention mechanisms and a modularity loss, overcoming over- and under-clustering issues without requiring manual annotations.

Recently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg.

Added

2026-09-26