Built independently by an author, for readers. Read the story and support ChapterPal

keyword

foreground localization

Foreground localization is a computer vision process that involves identifying and determining the spatial extent or pixel-level regions of target objects and salient visual elements within an image while distinguishing them from the surrounding background. In weakly supervised learning paradigms, such as semantic segmentation and object detection where detailed pixel-level annotations are unavailable, foreground localization commonly relies on activation maps and feature representations derived from image-level category labels to generate initial spatial cues. A key objective in foreground localization is expanding beyond the most discriminative parts of an object to capture its entire semantic structure, providing reliable pseudo-ground-truth masks for downstream dense prediction models.

1 item

Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation

Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic Segmentation

Qi Chen, Lingxiao Yang, Jianhuang Lai, Xiaohua Xie

OrganizationsSun Yat-sen University

Why you should read this

Proposes a self-supervised framework that tailors image-specific prototypes and enforces general-specific consistency to overcome incomplete class activation maps, achieving state-of-the-art weakly supervised semantic segmentation using only image-level labels.

Weakly Supervised Semantic Segmentation (WSSS) based on image-level labels has attracted much attention due to low annotation costs. Existing methods often rely on Class Activation Mapping (CAM) that measures the correlation between image pixels and classifier weight. However, the classifier focuses only on the discriminative regions while ignoring other useful information in each image, resulting in incomplete localization maps. To address this issue, we propose a Self-supervised Image-specific Prototype Exploration (SIPE) that consists of an Image-specific Prototype Exploration (IPE) and a General-Specific Consistency (GSC) loss. Specifically, IPE tailors prototypes for every image to capture complete regions, formed our Image-Specific CAM (IS-CAM), which is realized by two sequential steps. In addition, GSC is proposed to construct the consistency of general CAM and our specific IS-CAM, which further optimizes the feature representation and empowers a self-correction ability of prototype exploration. Extensive experiments are conducted on PASCAL VOC 2012 and MS COCO 2014 segmentation benchmark and results show our SIPE achieves new state-of-the-art performance using only image-level labels. The code is available at https://github.com/chenqi1126/SIPE.

Added

2026-09-26