Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures

Linshan WuZhun ZhongLeyuan FangXingxin HeQiang LiuJiayi MaHao Chen

article2023CVPR70 citations

Proposes an adaptive Gaussian mixture model framework that enables reliable end-to-end self-supervision from point or scribble annotations, eliminating the need for complex multi-stage pseudo-label generation in weakly supervised semantic segmentation.

Listen

Training computer vision models to accurately segment objects in images typically demands dense, pixel-by-pixel manual annotations. This labeling process is labor-intensive, expensive, and difficult to scale across industrial applications. While using sparse annotations like individual points or scribbles substantially cuts manual costs, conventional methods struggle with the severe lack of supervisory data. Prior solutions rely on brittle low-level image cues or multi-stage self-training that produces inaccurate pseudo-labels, ultimately compromising model quality and throughput.

The article introduces and evaluates an Adaptive Gaussian Mixture Model framework designed to improve sparsely annotated semantic segmentation. The main objective is to establish an end-to-end framework that effectively transfers reliable supervisory signals from sparse labeled points to unlabeled regions by dynamically modeling feature similarities without requiring complex multi-stage pipelines or supplemental edge annotations.

The authors develop a dual-branch neural network architecture that pairs a standard segmentation head with an adaptive probabilistic Gaussian mixture branch. The model establishes labeled pixels as class centroids in high-dimensional feature space, calculates dynamic variances across unlabeled pixels, and generates soft probabilistic predictions to guide mutual self-supervision. The system was evaluated on benchmark computer vision datasets—PASCAL VOC 2012 and Cityscapes—under point-level and scribble-level supervision settings using ResNet architectures.

Evaluation shows that the proposed approach consistently outperforms existing state-of-the-art sparse segmentation techniques. On the PASCAL VOC benchmark, the method achieved 69.6% mean intersection-over-union with point supervision and 76.4% with scribble supervision, outperforming baseline models by 10.4% and 9.1% while surpassing leading alternatives by up to 4.7%. On the complex Cityscapes urban dataset, it exceeded the prior state-of-the-art by 4.0% to 5.8% across varying point densities (20, 50, and 100 clicks per image). Furthermore, ablation studies confirmed that using soft probabilistic modeling and continuous gradient updates across classes delivered substantial gains over static, hard-threshold pseudo-labeling.

These findings indicate that sparse labeling can achieve performance levels close to fully supervised systems without inflating computational overhead during deployment. Because the auxiliary Gaussian mixture branch is active only during training and discarded at inference time, operational runtime latency and memory costs remain unaffected. Organizations can achieve significant cost savings and faster training cycles by transitioning from dense manual segmentation annotations to low-cost point or scribble inputs.

Decision-makers and engineering teams should consider piloting this adaptive framework within active computer vision pipelines that face annotation bottlenecks, such as automated visual inspection or autonomous driving perception. Where maximum segmentation accuracy is essential, teams can optionally combine the method with multi-stage training, which the article showed can provide an additional 2% to 5% accuracy boost at the cost of longer overall training times.

Confidence in these findings is supported by extensive empirical validation across multiple public benchmarks and supervision regimes. However, the evaluation focused on standard 2D natural and urban image datasets using established convolutional backbones. Practitioners should conduct targeted validation when extending the framework to specialized domains such as aerial imaging, medical scans, or transformer-based network architectures to verify generalizability.

Cover for Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures

Abstract

Sparsely annotated semantic segmentation (SASS) aims to learn a segmentation model by images with sparse labels (i.e., points or scribbles). Existing methods mainly focus on introducing low-level affinity or generating pseudo labels to strengthen supervision, while largely ignoring the inherent relation between labeled and unlabeled pixels. In this paper, we observe that pixels that are close to each other in the feature space are more likely to share the same class. Inspired by this, we propose a novel SASS framework, which is equipped with an Adaptive Gaussian Mixture Model (AGMM). Our AGMM can effectively endow reliable supervision for unlabeled pixels based on the distributions of labeled and unlabeled pixels. Specifically, we first build Gaussian mixtures using labeled pixels and their relatively similar unlabeled pixels, where the labeled pixels act as centroids, for modeling the feature distribution of each class. Then, we leverage the reliable information from labeled pixels and adaptively generated GMM predictions to supervise the training of unlabeled pixels, achieving online, dynamic, and robust self-supervision. In addition, by capturing category-wise Gaussian mixtures, AGMM encourages the model to learn discriminative class decision boundaries in an end-to-end contrastive learning manner. Experimental results conducted on the PASCAL VOC 2012 and Cityscapes datasets demonstrate that our AGMM can establish new state-of-the-art SASS performance. Code is available at https://github.com/Luffy03/AGMM-SASS.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methodology
  • 3.1. Motivation
  • 3.2. Overall Framework
  • 3.3. GMM Formulation
  • 3.4. Training with AGMM
  • 3.5. Discussion
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Implementation Details
  • 4.3. Ablation Study
  • 4.4. Comparison with State-of-the-art Methods
  • 5. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Adaptive Gaussian mixture framework for sparse segmentation

    model/method

    AGMM is a sparsely annotated semantic segmentation framework for training a pixelwise segmentation network when only point or scribble labels are available. For an input image, the main segmentation branch produces per-pixel class scores PP, while a second GMM branch receives the pixel features f(x)f(x) from the layer immediately before the segmentation classifier and produces class-wise soft scores GG. Sparse labels supervise both branches, and GG provides online probabilistic supervision for pixels without labels.

    The labeled pixels of each class are treated as reliable centroids of class-specific feature distributions. The resulting Gaussian mixtures model the relation between labeled and unlabeled pixels directly in the learned feature space, rather than relying on image-level affinity, manually supplied edges, or externally generated pseudo-labels. The GMM branch is used only during training; it is discarded at inference, when sparse labels are unavailable.

  2. Knowl 2 — Adaptive category-wise GMM construction

    equation

    Consider one training image with KK annotated semantic classes. Let Xℓi\mathcal{X}_{\ell i} be the set of labeled pixels assigned to class ii, Xu\mathcal{X}_u the unlabeled pixels, and f(x)∈RDf(x)\in\mathbb{R}^D the learned DD-dimensional feature vector at pixel xx. AGMM computes the class-ii centroid from labeled features as

    μi=1∣Xℓi∣∑x∈Xℓif(x).\mu_i=\frac{1}{|\mathcal{X}_{\ell i}|}\sum_{x\in\mathcal{X}_{\ell i}}f(x).

    For a pixel xx, its displacement from class ii is di(x)=f(x)−μid_i(x)=f(x)-\mu_i. Let Pi(x)P_i(x) be the segmentation branch's score for class ii at xx, and let ∣Pi∣|P_i| denote the normalization used for the class-ii scores over unlabeled pixels. AGMM estimates the adaptive class variance by

    σi=1∣Pi∣∑x∈XuPi(x)di(x)2.\sigma_i=\sqrt{\frac{1}{|P_i|}\sum_{x\in\mathcal{X}_u}P_i(x)d_i(x)^2}.

    The unnormalized Gaussian score for class ii is

    gi(x)=exp⁡(−di(x)22σi2),G(x)=(g1(x),…,gK(x)).g_i(x)=\exp\left(-\frac{d_i(x)^2}{2\sigma_i^2}\right),\qquad G(x)=\big(g_1(x),\ldots,g_K(x)\big).

    Here G(x)G(x) is the vector of soft category-wise GMM predictions; the paper also denotes the collection of component scores compactly as a GMM. The usual factor 1/2πσi21/\sqrt{2\pi\sigma_i^2} is omitted, so every component score lies in [0,1][0,1]. The number of components KK, centroids μi\mu_i, and variances σi\sigma_i are recomputed from each input's features and sparse labels, avoiding expectation-maximization and fixed, globally estimated mixture parameters.

  3. Knowl 3 — Mutual self-supervision and contrastive training objectives

    equation

    Let NℓN_\ell be the number of labeled pixels, NN the number of pixels used for training, yi(x)∈{0,1}y_i(x)\in\{0,1\} the sparse one-hot label for class ii, Pi(x)P_i(x) the segmentation score, and Gi(x)G_i(x) the AGMM score. The segmentation branch is trained on labeled pixels with partial cross-entropy:

    Lseg=−1Nℓ∑x∈Xℓ∑i=1Kyi(x)log⁡Pi(x).L_{\mathrm{seg}}=-\frac{1}{N_\ell}\sum_{x\in\mathcal{X}_\ell}\sum_{i=1}^{K}y_i(x)\log P_i(x).

    AGMM supplies soft supervision to the segmentation branch over labeled and unlabeled pixels through a per-class binary cross-entropy:

    Lself=−1N∑x∑i=1K[Gi(x)log⁡Pi(x)+(1−Gi(x))log⁡(1−Pi(x))].L_{\mathrm{self}}=-\frac{1}{N}\sum_x\sum_{i=1}^{K}\left[G_i(x)\log P_i(x)+(1-G_i(x))\log\big(1-P_i(x)\big)\right].

    The GMM branch is itself anchored to the sparse labels by

    Lspar=−1Nℓ∑x∈Xℓ∑i=1Kyi(x)log⁡Gi(x).L_{\mathrm{spar}}=-\frac{1}{N_\ell}\sum_{x\in\mathcal{X}_\ell}\sum_{i=1}^{K}y_i(x)\log G_i(x).

    To separate the centroids of different classes, AGMM uses

    Lcon=2K(K+1)∑i,j=1i≠jKexp⁡(−∥μi−μj∥22).L_{\mathrm{con}}=\frac{2}{K(K+1)}\sum_{\substack{i,j=1\\i\ne j}}^{K}\exp\left(-\lVert\mu_i-\mu_j\rVert_2^2\right).

    The GMM-related objective and the complete training objective are

    LGMM=Lself+Lspar+Lcon,L=Lseg+LGMM.L_{\mathrm{GMM}}=L_{\mathrm{self}}+L_{\mathrm{spar}}+L_{\mathrm{con}},\qquad L=L_{\mathrm{seg}}+L_{\mathrm{GMM}}.

    Thus, the segmentation predictions learn from the adaptive GMM, the GMM remains tied to reliable sparse labels, and the centroid-separation term encourages discriminative class distributions.

  4. Knowl 4 — End-to-end online optimization of the GMM branch

    model/method

    AGMM does not stop gradients through the soft GMM predictions GG. Because the centroids, variances, and Gaussian scores are differentiable functions of the feature extractor and segmentation predictions, back-propagation through LGMML_{\mathrm{GMM}} progressively updates the feature representations and the input-dependent Gaussian mixtures. The self-supervision term makes each unlabeled pixel move toward or away from class-specific mixtures according to its soft assignment, while the centroid-separation term enlarges the separation between different class distributions.

    This optimization is online: mixture parameters and supervision scores are generated anew during training for each input rather than produced in a separate pseudo-label-generation stage. The trained segmentation branch alone is retained for inference, so prediction requires neither sparse labels nor GMM computation.

  5. Knowl 5 — Feature-space similarity predicts unlabeled-pixel correctness

    empirical result

    The paper's analysis finds that an unlabeled pixel's feature-space similarity to a labeled pixel is strongly associated with whether its predicted class agrees with the labeled class. For class ii, pixels with distance di(x)<σid_i(x)<\sigma_i are treated as similar, whereas pixels with di(x)>σid_i(x)>\sigma_i are treated as dissimilar. Across semantic categories on PASCAL VOC 2012, the similar unlabeled pixels have substantially higher intersection-over-union than the dissimilar unlabeled pixels, and their performance is closer to the IoU over all unlabeled pixels.

    This observation motivates using learned high-level feature distributions as soft supervision. AGMM avoids converting the distance comparison into a rigid one-hot decision, because a fixed similarity threshold can incorrectly assign ambiguous pixels or discard useful supervision.

  6. Knowl 6 — Datasets and training configuration

    experimental setup

    Experiments evaluate point- and scribble-supervised semantic segmentation on PASCAL VOC 2012 and point-supervised segmentation on Cityscapes. PASCAL VOC 2012 is augmented with SBD data to form 10,582 training images, with 20 foreground classes plus background and 1,449 validation images. Point and scribble annotations are used. Cityscapes contains 2,975 training images and 500 validation images, with 19 classes and image resolution 2048×10242048\times1024; sparse labels are randomly sampled as 20, 50, or 100 clicks per image.

    The default architecture uses an ImageNet-pretrained ResNet backbone and DeepLabV3+. Random resizing, random cropping, and horizontal flipping are applied; crop sizes are 321×321321\times321 for PASCAL VOC and 769×769769\times769 for Cityscapes. SGD uses momentum 0.90.9, weight decay 0.00010.0001, and a polynomial learning-rate schedule with power 0.90.9. Backbone base learning rates are 0.0010.001 for PASCAL VOC and 0.0040.004 for Cityscapes, while the randomly initialized segmentation head uses ten times the backbone rate. Training lasts 80 epochs on PASCAL VOC and 240 epochs on Cityscapes, using four NVIDIA 3090 GPUs.

  7. Knowl 7 — Ablation of AGMM losses and multi-stage training

    data/table

    On PASCAL VOC 2012, the ablation progressively adds AGMM objectives to a baseline trained with only partial cross-entropy. The reported metric is validation mIoU, with point-supervised and scribble-supervised results listed respectively:

    • LsegL_{\mathrm{seg}} only: 59.2 / 67.3.
    • LsegL_{\mathrm{seg}} plus multi-stage training (MT): 66.3 / 72.4.
    • Lseg+LselfL_{\mathrm{seg}}+L_{\mathrm{self}}: 68.5 / 75.2.
    • Lseg+Lself+LsparL_{\mathrm{seg}}+L_{\mathrm{self}}+L_{\mathrm{spar}}: 69.3 / 76.1.
    • Lseg+Lself+Lspar+LconL_{\mathrm{seg}}+L_{\mathrm{self}}+L_{\mathrm{spar}}+L_{\mathrm{con}}: 69.6 / 76.4.
    • The full AGMM objectives with additional MT: 74.7 / 77.2.

    Thus, online mutual supervision from the GMM produces the largest improvement over the sparse-label baseline before the remaining AGMM terms are added; sparse-label supervision of the GMM and centroid contrast further improve both annotation settings. Multi-stage training can increase performance, especially for point supervision, but it is not required by the core AGMM method.

  8. Knowl 8 — Soft adaptive supervision outperforms hard or frozen alternatives

    data/table

    Ablations on PASCAL VOC 2012 compare hard similarity-based pseudo-label assignment, soft AGMM supervision, and whether the GMM branch is optimized. Point- and scribble-supervised mIoU values are reported respectively:

    • Baseline with no pseudo-label mechanism: 59.2 / 67.3.
    • Baseline with multi-stage training: 66.3 / 72.4.
    • Hard similarity-based label assignment with online use: 66.5 / 73.4.
    • AGMM with soft scores and online use but stopped GMM gradients (AGMM-SG): 67.4 / 74.6.
    • Full AGMM with soft scores and trainable GMM gradients: 69.6 / 76.4.

    The results support both design choices: soft probabilistic scores are more effective than assigning only pixels that satisfy a fixed hard distance rule, and retaining gradients through the GMM branch is important. The variance experiment also shows low sensitivity to the variance value. For point supervision, mIoU is 69.2, 69.5, 69.5, 69.6, 69.4, and 69.6 for σ=0.1,0.5,0.8,1.0,1.5\sigma=0.1,0.5,0.8,1.0,1.5, and the adaptive variance, respectively. For scribble supervision, the corresponding values are 76.0, 76.3, 76.2, 76.3, 76.1, and 76.4.

  9. Knowl 9 — PASCAL VOC state-of-the-art performance

    empirical result

    On the PASCAL VOC 2012 validation set, the DeepLabV3+ and ResNet-101 baseline obtains 59.2% mIoU with point supervision and 67.3% mIoU with scribble supervision. AGMM reaches 69.6% and 76.4% mIoU, improving over the corresponding baseline by 10.4 and 9.1 percentage points without extra data, multi-stage training, or DenseCRF post-processing.

    In the reported comparison, AGMM is above the strongest listed competing point-supervised result, TEL at 63.3% mIoU, and above the strongest listed competing scribble-supervised result, TEL at 75.8% mIoU. Other listed point-supervised results include What’s the Point 43.4%, KernelCut Loss 57.0%, A2GNN 66.8%, and DBFNet 66.8%. Other listed scribble-supervised results range from ScribbleSup 63.1%, RAWKS 61.4%, and GraphNet 70.3% to PSI 74.9%, CCL 74.4%, and SPML 74.2%. AGMM therefore establishes the best reported sparse-label result under the paper's no-extra-data, no-multi-stage, and no-DenseCRF setting.

  10. Knowl 10 — Cityscapes performance under varying click budgets

    empirical result

    On Cityscapes with randomly sampled point labels, AGMM improves as the annotation budget increases and outperforms the compared low-level regularization methods. Validation mIoU for 20, 50, and 100 clicks per 2048×10242048\times1024 image is:

    • Baseline: 53.5, 60.3, 64.2.
    • DenseCRF Loss: 54.2, 61.6, 65.5.
    • TEL: 56.3, 62.8, 67.6.
    • AGMM: 62.1, 68.3, 71.6.
    • AGMM with multi-stage training: 66.5, 71.7, 73.4.

    Without multi-stage training, AGMM exceeds TEL by 5.8, 5.5, and 4.0 percentage points for 20, 50, and 100 clicks, respectively. The results indicate that modeling learned feature distributions is effective in the complex, cluttered Cityscapes scenes where low-level image affinity provides limited gains.

Coverage note — No substantial contributed material was omitted; background, related work, acknowledgments, and references were excluded.

References

  1. 1.J. Ahn and S. Kwak. Learning pixel-level semantic affinity with image-level supervision for weakly supervised semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2018. 3
  2. 2.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. SegNet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 39(12):2481–2495, Dec. 2017. 1
  3. 3.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 39(12):2481–2495, Dec. 2017. 1
  4. 4.A. Bearman, O. Russakovsky, V. Ferrari, and F. F. Li. What’s the Point: Semantic segmentation with point supervision. In Eur. Conf. Comput. Vis., 2016. 2, 3, 6, 8
  5. 5.Yuri Boykov and Gareth Funka-Lea. Graph cuts and efficient nd image segmentation. Int. J. Comput. Vis., 70(2):109–131, 2006. 3, 4
  6. 6.Yuri Boykov, Olga Veksler, and Ramin Zabih. Fast approximate energy minimization via graph cuts. IEEE Trans. Pattern Anal. Mach. Intell., 23(11):1222–1239, 2001. 3, 4
  7. 7.Hongjun Chen, Jinbao Wang, Hong Cai Chen, Xiantong Zhen, Feng Zheng, Rongrong Ji, and Ling Shao. Seminar learning for click-level weakly supervised semantic segmentation. In Proc. IEEE/CVF Int. Conf. Comput. Vis., pages 6920–6929, 2021. 2, 3, 4, 6
  8. 8.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell., 40(4):834–848, Apr. 2017. 1, 7, 8
  9. 9.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell., 40(4):834–848, Apr. 2017. 1
  10. 10.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Euro. Conf. Comput. Vis. (ECCV), pages 801–818, 2018. 6, 7, 8
  11. 11.Yung-Yu Chuang, Brian Curless, David H Salesin, and Richard Szeliski. A bayesian approach to digital matting. In IEEE Conf. Comput. Vis. Pattern Recog. CVPR 2001, volume 2, pages II–II. IEEE, 2001. 3, 4
  12. 12.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In IEEE Conf. Comput. Vis. Pattern Recog., pages 3213–3223, 2016. 3, 6
  13. 13.Jifeng Dai, Kaiming He, and Jian Sun. Boxsup: Exploiting bounding boxes to supervise convolutional networks for semantic segmentation. In Int. Conf. Comput. Vis., pages 1635–1643, 2015. 3
  14. 14.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. Int. Jour. Comput. Vision, 88(2):303–338, Jun. 2010. 3, 6
  15. 15.Bharath Hariharan, Pablo Arbeláez, Lubomir Bourdev, Subhransu Maji, and Jitendra Malik. Semantic contours from inverse detectors. In Int. Conf. Comput. Vis., pages 991–998, 2011. 6
  16. 16.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc. IEEE Conf. Comput Vis. Pattern Recognit., pages 770–778, 2016. 6, 7, 8
  17. 17.Xingxin He, Leyuan Fang, Mingkui Tan, and Xiangdong Chen. Intra-and inter-slice contrastive learning for point supervised oct fluid segmentation. IEEE Trans. Image Process., 31:1870–1881, 2022. 2
  18. 18.Z. Huang, X. Wang, J. Wang, W. Liu, and J. Wang. Weakly-supervised semantic segmentation network with deep seeded region growing. In IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), 2018. 3
  19. 19.Tsung-Wei Ke, Jyh-Jing Hwang, and Stella X Yu. Universal weakly supervised segmentation by pixel-to-segment contrastive learning. In ICLR, 2021. 2, 3, 6, 7, 8
  20. 20.Alexander Kolesnikov and Christoph H. Lampert. Seed, expand and constrain: Three principles for weakly-supervised image segmentation. In Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling, editors, Eur. Conf. Comput. Vis., pages 695–711, 2016. 3
  21. 21.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. Adv. Neural Inf. Process. Syst., 25, 2012. 6
  22. 22.Mingchun Li, Dali Chen, and Shixin Liu. Weakly supervised segmentation loss based on graph cuts and superpixel algorithm. Neural Process. Letters, pages 1–24, 2022. 8
  23. 23.Stan Z Li. Markov random field models in computer vision. In Eur. Conf. Comput. Vis., pages 361–370. Springer, 1994. 7
  24. 24.Chen Liang, Wenguan Wang, Jiaxu Miao, and Yi Yang. Gmmseg: Gaussian mixture based generative semantic segmentation models. Adv. Neural Inform. Process. Syst., 2022. 3
  25. 25.Chen Liang, Wenguan Wang, Jiaxu Miao, and Yi Yang. Gmmseg: Gaussian mixture based generative semantic segmentation models. Adv. Neural Inform. Process. Syst., 2022. 4
  26. 26.Zhiyuan Liang, Tiancai Wang, Xiangyu Zhang, Jian Sun, and Jianbing Shen. Tree energy loss: Towards sparsely annotated semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 16907–16916, 2022. 2, 3, 4, 6, 7, 8
  27. 27.Di Lin, Jifeng Dai, Jiaya Jia, Kaiming He, and Jian Sun. Scribblesup: Scribble-supervised convolutional networks for semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 3159–3167, 2016. 2, 3, 4, 6, 8
  28. 28.Wei Liu, Andrew Rabinovich, and Alexander C Berg. Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579, 2015. 6
  29. 29.K.-K. Maninis, S. Caelles, J. Pont-Tuset, and L. Van Gool. Deep extreme cut: From extreme points to object segmentation. In Proc. IEEE Conf. Comput. Vis. Pattern Recog. (CVPR), pages 616–625, 2018. 2, 3
  30. 30.Dmitrii Marin, Meng Tang, Ismail Ben Ayed, and Yuri Boykov. Beyond gradient descent for regularized segmentation losses. In IEEE Conf. Comput. Vis. Pattern Recog., pages 10187–10196, 2019. 2, 3, 6
  31. 31.Dmitrii Marin, Meng Tang, Ismail Ben Ayed, and Yuri Boykov. Beyond gradient descent for regularized segmentation losses. In IEEE Conf. Comput. Vis. Pattern Recog., pages 10187–10196, 2019. 3, 8
  32. 32.Christoph Mayer, Radu Timofte, and Grégory Paul. Towards closing the gap in weakly supervised semantic segmentation with dcnns: Combining local and global models. Comput. Vis. Image Under., 208:103209, 2021. 3
  33. 33.Todd K Moon. The expectation-maximization algorithm. IEEE Signal Process. magazine, 13(6):47–60, 1996. 3
  34. 34.Anton Obukhov, Stamatios Georgoulis, Dengxin Dai, and Luc Van Gool. Gated crf loss for weakly supervised semantic image segmentation. arXiv preprint arXiv:1906.04651, 2019. 2, 3, 6
  35. 35.Zhiyi Pan, Peng Jiang, Yunhai Wang, Changhe Tu, and Anthony G Cohn. Scribble-supervised semantic segmentation by uncertainty reduction on neural representation and self-supervision on neural eigenspace. In Int. Conf. Comput. Vis., pages 7416–7425, 2021. 2, 3, 6, 8
  36. 36.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. Adv. Neural Inform. Process. Syst., 32, 2019. 6
  37. 37.Mengyang Pu, Yaping Huang, Qingji Guan, and Qi Zou. GraphNet: Learning image pseudo annotations for weakly-supervised semantic segmentation. In ACM Int. Conf. Multimedia, pages 483–491, 2018. 3, 6, 8
  38. 38.Rui Qian, Yunchao Wei, Honghui Shi, Jiachen Li, Jiaying Liu, and Thomas Huang. Weakly supervised scene parsing with point-based distance metric learning. In AAAI, volume 33, pages 8843–8850, 2019. 2, 6
  39. 39.Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. " grabcut" interactive foreground extraction using iterated graph cuts. ACM Trans. Graphics (TOG), 23(3):309–314, 2004. 3
  40. 40.Carsten Rother, Vladimir Kolmogorov, and Andrew Blake. Interactive foreground extraction using iterated graph cuts. ACM Trans. Graphics, 23:3, 2012. 3, 4
  41. 41.Mark A Ruzon and Carlo Tomasi. Alpha estimation in natural images. In IEEE Conf. Comput. Vis. Pattern Recog. CVPR 2000, volume 1, pages 18–25. IEEE, 2000. 3, 4
  42. 42.Evan Shelhamer, Jonathan Long, and Trevor Darrell. Fully convolutional networks for semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 39(4):640–651, Apr. 2016. 1
  43. 43.Evan Shelhamer, Jonathan Long, and Trevor Darrell. Fully convolutional networks for semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 39(4):640–651, Apr. 2016. 1
  44. 44.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 8
  45. 45.Lin Song, Yanwei Li, Zhengkai Jiang, Zeming Li, Xiangyu Zhang, Hongbin Sun, Jian Sun, and Nanning Zheng. Rethinking learnable tree filter for generic feature transform. Adv. Neural Inform. Process. Syst., 33:3991–4002, 2020. 7
  46. 46.Lin Song, Yanwei Li, Zeming Li, Gang Yu, Hongbin Sun, Jian Sun, and Nanning Zheng. Learnable tree filter for structure-preserving feature transform. Adv. Neural Inform. Process. Syst., 32, 2019. 7
  47. 47.Meng Tang, Abdelaziz Djelouah, Federico Perazzi, Yuri Boykov, and Christopher Schroers. Normalized cut loss for weakly-supervised cnn segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 1818–1827, 2018. 2, 3, 6, 8
  48. 48.Meng Tang, Federico Perazzi, Abdelaziz Djelouah, Ismail Ben Ayed, Christopher Schroers, and Yuri Boykov. On regularized losses for weakly-supervised cnn segmentation. In Proc. Euro. Conf. Comput. Vis. (ECCV), pages 507–522, 2018. 2, 3, 6, 8
  49. 49.Ozan Unal, Dengxin Dai, and Luc Van Gool. Scribble-supervised lidar semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 2697–2707, 2022. 3
  50. 50.Paul Vernaza and Manmohan Chandraker. Learning random-walk label propagation for weakly-supervised semantic segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., pages 7158–7166, 2017. 3, 6, 7, 8
  51. 51.Bin Wang, Guojun Qi, Sheng Tang, Tianzhu Zhang, Yunchao Wei, Linghui Li, and Yongdong Zhang. Boundary perception guidance: A scribble-supervised semantic segmentation approach. In IJCAI Int. Joint Conf. Artifi. Intell., 2019. 3, 6, 7, 8
  52. 52.Bin Wang, Yu Qiao, Dahua Lin, Stephen DH Yang, and Weijia Li. Cycle-consistent learning for weakly supervised semantic segmentation. In 3rd Inter. Workshop Human-Centric Multi. Anal., pages 7–13, 2022. 8
  53. 53.Zhihui Wang, Shijie Wang, Shuhui Yang, Haojie Li, Jianjun Li, and Zezhou Li. Weakly supervised fine-grained image classification via gaussian mixture model oriented discriminative learning. In IEEE Conf. Comput. Vis. Pattern Recog., pages 9749–9758, 2020. 3
  54. 54.Y. Wei, H. Xiao, H. Shi, Z. Jie, J. Feng, and T.S. Huang. Revisiting Dilated Convolution: A Simple Approach for Weakly- and Semi- Supervised Semantic Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 3
  55. 55.Linshan Wu, Leyuan Fang, Xingxin He, Min He, Jiayi Ma, and Zhun Zhong. Querying labeled for unlabeled: Cross-image semantic consistency guided semi-supervised semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell., pages 1–18, 2022. 2, 3
  56. 56.Linshan Wu, Leyuan Fang, Jun Yue, Bob Zhang, Pedram Ghamisi, and Min He. Deep bilateral filtering network for point-supervised semantic segmentation in remote sensing images. IEEE Trans. Image Process., 31:7419–7434, 2022. 8
  57. 57.Linshan Wu, Ming Lu, and Leyuan Fang. Deep covariance alignment for domain adaptive remote sensing image segmentation. IEEE Trans. Geosci. Remote Sens., 60:1–11, 2022. 3
  58. 58.Saining Xie and Zhuowen Tu. Holistically-nested edge detection. In Int. Conf. Comput. Vis., pages 1395–1403, 2015. 7
  59. 59.Jingshan Xu, Chuanwei Zhou, Zhen Cui, Chunyan Xu, Yuge Huang, Pengcheng Shen, Shaoxin Li, and Jian Yang. Scribble-supervised semantic segmentation inference. In Int. Conf. Comput. Vis., pages 15354–15363, 2021. 2, 3, 4, 6, 8
  60. 60.Bingfeng Zhang, Jimin Xiao, Jianbo Jiao, Yunchao Wei, and Yao Zhao. Affinity attention graph neural network for weakly supervised semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell., 2021. 2, 3, 4, 6, 8
  61. 61.Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. Learning deep features for discriminative localization. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), pages 2921–2929, 2016. 3

Citation

MLA
Wu, L., et al. “Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 15454–64, https://doi.org/10.1109/CVPR52729.2023.01483.
APA
Wu, L., Zhong, Z., Fang, L., He, X., Liu, Q., Ma, J., & Chen, H. (2023). Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15454–15464. https://doi.org/10.1109/CVPR52729.2023.01483
Chicago
Wu, L., Z. Zhong, L. Fang, et al. 2023. “Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 15454–64. https://doi.org/10.1109/CVPR52729.2023.01483.
Harvard
Wu, L. et al. (2023) “Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 15454–15464. Available at: https://doi.org/10.1109/CVPR52729.2023.01483.
Vancouver
1. Wu L, Zhong Z, Fang L, He X, Liu Q, Ma J, Chen H (2023) Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 15454–15464

BibTeX

@inproceedings{Wu_2023, title={Sparsely Annotated Semantic Segmentation with Adaptive Gaussian Mixtures}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01483}, DOI={10.1109/cvpr52729.2023.01483}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wu, Linshan and Zhong, Zhun and Fang, Leyuan and He, Xingxin and Liu, Qiang and Ma, Jiayi and Chen, Hao}, year={2023}, month=June, pages={15454–15464} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE