Built independently by an author, for readers. Read the story and support ChapterPal

keyword

pixel-level annotations

Pixel-level annotations are detailed ground-truth labels in computer vision where every individual pixel in a digital image is assigned a specific semantic category, object class, or instance identifier. Unlike coarser labeling methods such as whole-image tags, bounding boxes, or scribbles, pixel-level annotations provide exact spatial boundaries and dense segmentations for all visual elements across a scene. These per-pixel masks are fundamental for training and evaluating dense prediction models in tasks such as semantic segmentation, instance segmentation, and panoptic segmentation, enabling algorithms to accurately delineate complex contours, separate overlapping objects, and categorize background regions. Because manually labeling each pixel requires significant time and human labor, datasets featuring these dense annotations serve as high-precision benchmarks in supervised learning and motivate the development of weakly supervised and unsupervised segmentation alternatives.

5 items

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

ACSeg: Adaptive Conceptualization for Unsupervised Semantic Segmentation

Kehan Li, Zhennan Wang, Zesen Cheng, Runyi Yu, Yian Zhao, Guoli Song, Chang Liu, Li Yuan, Jie Chen

OrganizationsDalian University of TechnologyPeking UniversityPeng Cheng LaboratoryTsinghua University

Why you should read this

Proposes an unsupervised semantic segmentation framework that dynamically maps learnable prototypes into image-specific semantic concepts using attention mechanisms and a modularity loss, overcoming over- and under-clustering issues without requiring manual annotations.

Recently, self-supervised large-scale visual pre-training models have shown great promise in representing pixel-level semantic relationships, significantly promoting the development of unsupervised dense prediction tasks, e.g., unsupervised semantic segmentation (USS). The extracted relationship among pixel-level representations typically contains rich class-aware information that semantically identical pixel embeddings in the representation space gather together to form sophisticated concepts. However, leveraging the learned models to ascertain semantically consistent pixel groups or regions in the image is non-trivial since over/ under-clustering overwhelms the conceptualization procedure under various semantic distributions of different images. In this work, we investigate the pixel-level semantic aggregation in self-supervised ViT pre-trained models as image Segmentation and propose the Adaptive Conceptualization approach for USS, termed ACSeg. Concretely, we explicitly encode concepts into learnable prototypes and design the Adaptive Concept Generator (ACG), which adaptively maps these prototypes to informative concepts for each image. Meanwhile, considering the scene complexity of different images, we propose the modularity loss to optimize ACG independent of the concept number based on estimating the intensity of pixel pairs belonging to the same concept. Finally, we turn the USS task into classifying the discovered concepts in an unsupervised manner. Extensive experiments with state-of-the-art results demonstrate the effectiveness of the proposed ACSeg.

Added

2026-09-26

Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation

Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation

Tianfei Zhou, Meijie Zhang, Fang Zhao, Jianwu Li

OrganizationsBeijing Institute of TechnologyETH ZurichInception Institute of AI

Why you should read this

Introduces Regional Semantic Contrast and Aggregation (RCA), a framework that utilizes a dataset-wide regional memory bank to contrast and aggregate categorical object patterns across training images, achieving state-of-the-art weakly supervised semantic segmentation on PASCAL VOC and COCO.

Learning semantic segmentation from weakly-labeled (e.g., image tags only) data is challenging since it is hard to infer dense object regions from sparse semantic tags. Despite being broadly studied, most current efforts directly learn from limited semantic annotations carried by individual image or image pairs, and struggle to obtain integral localization maps. Our work alleviates this from a novel perspective, by exploring rich semantic contexts synergistically among abundant weakly-labeled training data for network learning and inference. In particular, we propose regional semantic contrast and aggregation (RCA). RCA is equipped with a regional memory bank to store massive, diverse object patterns appearing in training data, which acts as strong support for exploration of dataset-level semantic structure. Particularly, we propose i) semantic contrast to drive network learning by contrasting massive categorical object regions, leading to a more holistic object pattern understanding, and ii) semantic aggregation to gather diverse relational contexts in the memory to enrich semantic representations. In this manner, RCA earns a strong capability of fine-grained semantic understanding, and eventually establishes new state-of-the-art results on two popular benchmarks, i.e., PASCAL VOC 2012 and COCO 2014.

Added

2026-09-26

Generative Semantic Segmentation

Generative Semantic Segmentation

Jiaqi Chen, Jiachen Lu, Xiatian Zhu, Li Zhang

OrganizationsFudan UniversityUniversity of Surrey

Why you should read this

Proposes Generative Semantic Segmentation to cast semantic segmentation as an image-conditioned mask generation problem using discrete latent priors and mask-as-image representations, delivering state-of-the-art generalization on challenging cross-domain benchmarks.

We present Generative Semantic Segmentation (GSS), a generative learning approach for semantic segmentation. Uniquely, we cast semantic segmentation as an image-conditioned mask generation problem. This is achieved by replacing the conventional per-pixel discriminative learning with a latent prior learning process. Specifically, we model the variational posterior distribution of latent variables given the segmentation mask. To that end, the segmentation mask is expressed with a special type of image (dubbed as maskige). This posterior distribution allows to generate segmentation masks unconditionally. To achieve semantic segmentation on a given image, we further introduce a conditioning network. It is optimized by minimizing the divergence between the posterior distribution of maskige (i.e. segmentation masks) and the latent prior distribution of input training images. Extensive experiments on standard benchmarks show that our GSS can perform competitively to prior art alternatives in the standard semantic segmentation setting, whilst achieving a new state of the art in the more challenging cross-domain setting.

Added

2026-09-26

PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment

PANet: Few-Shot Image Semantic Segmentation With Prototype Alignment

Kaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou, Jiashi Feng

OrganizationsNational University of Singapore

Why you should read this

Proposes a metric learning framework for few-shot image semantic segmentation that learns representative class prototypes and enforces bidirectional alignment between support and query images to improve generalization to unseen categories.

Despite the great progress made by deep CNNs in image semantic segmentation, they typically require a large number of densely-annotated images for training and are difficult to generalize to unseen object categories. Few-shot segmentation has thus been developed to learn to perform segmentation from only a few annotated examples. In this paper, we tackle the challenging few-shot segmentation problem from a metric learning perspective and present PANet, a novel prototype alignment network to better utilize the information of the support set. Our PANet learns class-specific prototype representations from a few support images within an embedding space and then performs segmentation over the query images through matching each pixel to the learned prototypes. With non-parametric metric learning, PANet offers high-quality prototypes that are representative for each semantic class and meanwhile discriminative for different classes. Moreover, PANet introduces a prototype alignment regularization between support and query. With this, PANet fully exploits knowledge from the support and provides better generalization on few-shot segmentation. Significantly, our model achieves the mIoU score of 48.1% and 55.7% on PASCAL-5i for 1-shot and 5-shot settings respectively, surpassing the state-of-the-art method by 1.8% and 8.6%.

Added

2026-09-25