keyword
pixel-wise supervision
Pixel-wise supervision is a training approach in which a model receives learning signals at individual image-pixel locations, typically using labels that specify the class or target associated with each pixel. Unlike supervision provided only at the image or object level, it guides the model toward spatially detailed predictions, such as segmentation masks or dense per-pixel labels.
2 items

Dense Learning based Semi-Supervised Object Detection
Binghui Chen, Pengyu Li, Xiang Chen, Biao Wang, Lei Zhang, Xian-Sheng Hua
Why you should read this
Proposes an anchor-free semi-supervised object detection framework that assigns dense pixel-level pseudo-labels via adaptive filtering and scale-consistent regularization to substantially outperform anchor-based methods on limited labeled data.
Semi-supervised object detection (SSOD) aims to facilitate the training and deployment of object detectors with the help of a large amount of unlabeled data. Though various self-training based and consistency-regularization based SSOD methods have been proposed, most of them are anchor-based detectors, ignoring the fact that in many real-world applications anchor-free detectors are more demanded. In this paper, we intend to bridge this gap and propose a DenSe Learning (DSL) based anchor-free SSOD algorithm. Specifically, we achieve this goal by introducing several novel techniques, including an Adaptive Filtering strategy for assigning multi-level and accurate dense pixel-wise pseudo-labels, an Aggregated Teacher for producing stable and precise pseudo-labels, and an uncertainty-consistency-regularization term among scales and shuffled patches for improving the generalization capability of the detector. Extensive experiments are conducted on MS-COCO and PASCAL-VOC, and the results show that our proposed DSL method records new state-of-the-art SSOD performance, surpassing existing methods by a large margin. Codes can be found at https://github.com/chenbinghui1/DSL.
Added
2026-09-26

Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor
Hyeokjun Kweon, Sung-Hoon Yoon, Kuk-Jin Yoon
Why you should read this
Proposes an adversarial framework pitting a CAM-generating classifier against an image reconstructor to prevent over-erasing and under-activation by minimizing cross-segment inferability, achieving state-of-the-art weakly supervised semantic segmentation performance on PASCAL VOC and MS COCO.
In Weakly Supervised Semantic Segmentation (WSSS), Class Activation Maps (CAMs) usually 1) do not cover the whole object and 2) be activated on irrelevant regions. To address the issues, we propose a novel WSSS framework via adversarial learning of a classifier and an image reconstructor. When an image is perfectly decomposed into class-wise segments, information (i.e., color or texture) of a single segment could not be inferred from the other segments. Therefore, inferability between the segments can represent the preciseness of segmentation. We quantify the inferability as a reconstruction quality of one segment from the other segments. If one segment could be reconstructed from the others, then the segment would be imprecise. To bring this idea into WSSS, we simultaneously train two models: a classifier generating CAMs that decompose an image into segments and a reconstructor that measures the inferability between the segments. As in GANs, while being alternatively trained in an adversarial manner, two networks provide positive feedback to each other. We verify the superiority of the proposed framework with extensive ablation studies. Our method achieves new state-of-the-art performances on both PASCAL VOC 2012 and MS COCO 2014. The code is available at https://github.com/sangrockEG/ACR.
Added
2026-09-26
