Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures
Linshan WuZhun ZhongLeyuan FangXingxin HeQiang LiuJiayi MaHao Chen
Proposes an adaptive Gaussian mixture model framework that enables reliable end-to-end self-supervision from point or scribble annotations, eliminating the need for complex multi-stage pseudo-label generation in weakly supervised semantic segmentation.
Training computer vision models to accurately segment objects in images typically demands dense, pixel-by-pixel manual annotations. This labeling process is labor-intensive, expensive, and difficult to scale across industrial applications. While using sparse annotations like individual points or scribbles substantially cuts manual costs, conventional methods struggle with the severe lack of supervisory data. Prior solutions rely on brittle low-level image cues or multi-stage self-training that produces inaccurate pseudo-labels, ultimately compromising model quality and throughput.
The article introduces and evaluates an Adaptive Gaussian Mixture Model framework designed to improve sparsely annotated semantic segmentation. The main objective is to establish an end-to-end framework that effectively transfers reliable supervisory signals from sparse labeled points to unlabeled regions by dynamically modeling feature similarities without requiring complex multi-stage pipelines or supplemental edge annotations.
The authors develop a dual-branch neural network architecture that pairs a standard segmentation head with an adaptive probabilistic Gaussian mixture branch. The model establishes labeled pixels as class centroids in high-dimensional feature space, calculates dynamic variances across unlabeled pixels, and generates soft probabilistic predictions to guide mutual self-supervision. The system was evaluated on benchmark computer vision datasets—PASCAL VOC 2012 and Cityscapes—under point-level and scribble-level supervision settings using ResNet architectures.
Evaluation shows that the proposed approach consistently outperforms existing state-of-the-art sparse segmentation techniques. On the PASCAL VOC benchmark, the method achieved 69.6% mean intersection-over-union with point supervision and 76.4% with scribble supervision, outperforming baseline models by 10.4% and 9.1% while surpassing leading alternatives by up to 4.7%. On the complex Cityscapes urban dataset, it exceeded the prior state-of-the-art by 4.0% to 5.8% across varying point densities (20, 50, and 100 clicks per image). Furthermore, ablation studies confirmed that using soft probabilistic modeling and continuous gradient updates across classes delivered substantial gains over static, hard-threshold pseudo-labeling.
These findings indicate that sparse labeling can achieve performance levels close to fully supervised systems without inflating computational overhead during deployment. Because the auxiliary Gaussian mixture branch is active only during training and discarded at inference time, operational runtime latency and memory costs remain unaffected. Organizations can achieve significant cost savings and faster training cycles by transitioning from dense manual segmentation annotations to low-cost point or scribble inputs.
Decision-makers and engineering teams should consider piloting this adaptive framework within active computer vision pipelines that face annotation bottlenecks, such as automated visual inspection or autonomous driving perception. Where maximum segmentation accuracy is essential, teams can optionally combine the method with multi-stage training, which the article showed can provide an additional 2% to 5% accuracy boost at the cost of longer overall training times.
Confidence in these findings is supported by extensive empirical validation across multiple public benchmarks and supervision regimes. However, the evaluation focused on standard 2D natural and urban image datasets using established convolutional backbones. Practitioners should conduct targeted validation when extending the framework to specialized domains such as aerial imaging, medical scans, or transformer-based network architectures to verify generalizability.
- Paper: Pointly-Supervised Instance Segmentation, Bowen Cheng et al. (2022). Introduces foundational techniques and benchmarks for training segmentation models using sparse point annotations, establishing the sparse-supervision setup that AGMM builds upon.
- Paper: Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic Segmentation, Tianfei Zhou et al. (2022). Pioneers category-level semantic contrast and regional feature aggregation for weakly supervised segmentation, providing the foundation for contrastive boundary learning in AGMM.
- Paper: Learning with Local and Global Consistency, Dengyong Zhou et al. (2003). Establishes classic graph-based label propagation and feature-space affinity principles connecting labeled and unlabeled data points utilized in adaptive mixture modeling.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). Provides the foundational end-to-end fully convolutional network architecture upon which modern dense semantic segmentation frameworks are constructed.
- Paper: Weakly Supervised Semantic Segmentation via Adversarial Learning of Classifier and Reconstructor, Hyeokjun Kweon et al. (2023). Explores an alternative competitive reconstruction paradigm to resolve boundary errors and feature leakage in weakly annotated semantic segmentation.
- Paper: Boosting Semi-Supervised Learning by Exploiting All Unlabeled Data, Yuhao Chen et al. (2023). Extends semi-supervised training concepts by utilizing all ambiguous unlabeled samples through non-target prediction distribution modeling and adaptive negative learning.
