GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models

Chen LiangWenguan WangJiaxu MiaoYi Yang

article2022NeurIPS185 citations

Proposes a hybrid semantic segmentation framework that models pixel feature densities via Gaussian Mixture Models using online Expectation-Maximization alongside end-to-end discriminative representation learning, delivering superior closed-set accuracy while naturally detecting out-of-distribution anomalies without architectural changes or post-processing.

Listen

Modern computer vision models for semantic segmentation—the task of classifying every pixel in an image—almost universally rely on discriminative classifiers using softmax layers. While effective for standard benchmarks, these systems learn only simple decision boundaries between known categories rather than the actual distribution of visual data. Consequently, they assume that all samples within a class look similar and frequently produce overconfident, incorrect predictions when encountering unfamiliar objects. This poses substantial safety and reliability risks in high-stakes environments such as autonomous driving, where identifying unexpected road hazards and out-of-distribution anomalies is vital.

The article introduces and evaluates GMMSeg, a new generative neural framework designed to resolve these shortcomings. The main objective is to demonstrate that replacing standard discriminative classifiers with a generative Gaussian Mixture Model (GMM) improves both standard segmentation accuracy and out-of-distribution anomaly detection within a single, end-to-end trained architecture.

To achieve this, the authors designed a hybrid training strategy that pairs generative statistical optimization with deep feature learning. The system models the feature distribution of each semantic class using five Gaussian components, which are updated iteratively using an online momentum version of Sinkhorn Expectation-Maximization alongside an external feature memory. Simultaneously, the underlying deep neural network learns visual representations using standard discriminative cross-entropy loss. The approach was evaluated across multiple network backbones on three standard segmentation benchmarks (ADE20K, Cityscapes, and COCO-Stuff) and two specialized road anomaly detection benchmarks (Fishyscapes Lost&Found and Road Anomaly).

Key findings show consistent performance advantages over conventional approaches. First, GMMSeg systematically outperformed standard discriminative baselines on closed-set benchmarks, achieving gains of 0.6% to 1.5% mean Intersection over Union on ADE20K, 0.5% to 0.8% on Cityscapes, and 0.7% to 1.7% on COCO-Stuff across various convolutional and Transformer architectures. Second, without any retraining, specialized anomaly exposure, or image reconstruction modules, the Cityscapes-trained GMMSeg model established state-of-the-art results on anomaly segmentation, improving the Average Precision on Fishyscapes Lost&Found from the baseline range of 6.02%–36.55% up to 43.47%–50.03%. Third, confidence calibration improved markedly, reducing the Expected Calibration Error from 0.1065 down to 0.0766. Finally, diagnostic ablations confirmed that optimal transport-based Sinkhorn optimization and multi-component modeling were critical, whereas post-hoc fitting of Gaussian mixtures to pre-trained networks degraded segmentation performance by over 14 percentage points.

These results demonstrate that generative density modeling can be seamlessly combined with deep neural networks without sacrificing predictive accuracy. The framework improves operational safety and reliability by naturally rejecting unknown objects through class-conditional likelihoods rather than relying on brittle post-processing calibration. Crucially, GMMSeg delivers these benefits with virtually no computational runtime penalty, processing images at 13.37 frames per second compared to 14.16 frames per second for standard softmax models.

Organizations developing safety-critical perception systems should consider transitioning from standard softmax classification layers to hybrid generative frameworks like GMMSeg to enhance anomaly detection without needing complex auxiliary pipelines. For future deployment, engineering teams should evaluate this methodology across broader computer vision tasks, such as general image classification and open-world scene understanding, while conducting pilot validations on application-specific edge hardware.

Confidence in these findings is supported by consistent empirical gains across multiple independent datasets and diverse model architectures. However, decision-makers should note that the computational cost of running multiple training iterations across large GPU clusters prevented the reporting of multi-seed error bars in the study. Additionally, setting the number of Gaussian components per class requires balanced tuning, as increasing beyond five to ten components yielded diminishing returns due to overparameterization.

arXiv: 2210.02025
Cover for GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models

Abstract

Prevalent semantic segmentation solutions are, in essence, a dense discriminative classifier of p(class | pixel feature). Though straightforward, this de facto paradigm neglects the underlying data distribution p(pixel feature | class), and struggles to identify out-of-distribution data. Going beyond this, we propose GMMSeg, a new family of segmentation models that rely on a dense generative classifier for the joint distribution p(pixel feature, class). For each class, GMMSeg builds Gaussian Mixture Models (GMMs) via Expectation-Maximization (EM), so as to capture class-conditional densities. Meanwhile, the deep dense representation is end-to-end trained in a discriminative manner, i.e., maximizing p(class | pixel feature). This endows GMMSeg with the strengths of both generative and discriminative models. With a variety of segmentation architectures and backbones, GMMSeg outperforms the discriminative counterparts on three closed-set datasets. More impressively, without any modification, GMMSeg even performs well on open-world datasets. We believe this work brings fundamental insights into the related fields.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Existing Segmentation Solutions: Dense Discriminative Classifier
  • 3.2 GMMSeg: Dense GMM Generative Classification
  • 3.3 Implementation Details
  • 4 Experiments
  • 4.1 Experiments on Semantic Segmentation
  • 4.2 Experiments on Anomaly Segmentation
  • 4.3 Diagnostic Experiments
  • 5 Conclusion
  • References
  • Checklist

Knowls

  1. Knowl 1 — GMMSeg Generative-Discriminative Framework for Semantic Segmentation

    model/method

    GMMSeg is a semantic segmentation framework that combines a dense generative classifier for modeling class-conditional data densities with discriminatively trained dense feature representations. Given an input pixel image xx, a dense feature extractor f_{oldsymbol{ heta}}: \mathbb{R}^3 \to \mathbb{R}^D parameterized by heta\boldsymbol{ heta} maps each pixel to a DD-dimensional feature embedding \mathbf{x} = f_{oldsymbol{ heta}}(x) \in \mathbb{R}^D (with D=64D=64 via a 1×11 \times 1 convolution). For each class c∈{1,…,C}c \in \{1, \dots, C\}, GMMSeg models the class-conditional feature distribution p(x∣c;ϕc)p(\mathbf{x}|c; \boldsymbol{\phi}_c) as a Gaussian Mixture Model (GMM) with MM multivariate Gaussian components:

    p(x∣c;ϕc)=∑m=1MπcmN(x;μcm,Σcm)p(\mathbf{x}|c; \boldsymbol{\phi}_c) = \sum_{m=1}^{M} \pi_{cm} \mathcal{N}(\mathbf{x}; \boldsymbol{\mu}_{cm}, \boldsymbol{\Sigma}_{cm})

    where ϕc={πc,μc,Σc}\boldsymbol{\phi}_c = \{\boldsymbol{\pi}_c, \boldsymbol{\mu}_c, \boldsymbol{\Sigma}_c\}, πc=[πc1,…,πcM]⊤\boldsymbol{\pi}_c = [\pi_{c1}, \dots, \pi_{cM}]^\top are mixing coefficients satisfying ∑m=1Mπcm=1\sum_{m=1}^M \pi_{cm} = 1, μcm∈RD\boldsymbol{\mu}_{cm} \in \mathbb{R}^D is the mean vector, and Σcm∈RD×D\boldsymbol{\Sigma}_{cm} \in \mathbb{R}^{D \times D} is constrained to be diagonal for computational efficiency. Using Bayes' rule and assuming a uniform class prior p(c)=1/Cp(c) = 1/C, the class posterior is:

    p(c∣x;{ϕc}c=1C,θ)=p(x∣c;ϕc)∑c′=1Cp(x∣c′;ϕc′)=∑m=1MπcmN(fheta(x);μcm,Σcm)∑c′=1C∑m=1Mπc′mN(fheta(x);μc′m,Σc′m)p(c|\mathbf{x}; \{\boldsymbol{\phi}_c\}_{c=1}^C, \boldsymbol{\theta}) = \frac{p(\mathbf{x}|c; \boldsymbol{\phi}_c)}{\sum_{c'=1}^C p(\mathbf{x}|c'; \boldsymbol{\phi}_{c'})} = \frac{\sum_{m=1}^M \pi_{cm} \mathcal{N}(f_{\boldsymbol{ heta}}(x); \boldsymbol{\mu}_{cm}, \boldsymbol{\Sigma}_{cm})}{\sum_{c'=1}^C \sum_{m=1}^M \pi_{c'm} \mathcal{N}(f_{\boldsymbol{ heta}}(x); \boldsymbol{\mu}_{c'm}, \boldsymbol{\Sigma}_{c'm})}

    Training proceeds through a hybrid objective: the dense feature extractor parameters θ\boldsymbol{\theta} are optimized via stochastic gradient descent by minimizing the cross-entropy loss over the posteriors:

    θ∗=arg⁡min⁡θ−∑(x,c)∈Dlog⁡p(c∣x;{ϕc}c=1C,θ)\boldsymbol{\theta}^* = \arg\min_{\boldsymbol{\theta}} - \sum_{(x, c) \in \mathcal{D}} \log p(c|\mathbf{x}; \{\boldsymbol{\phi}_c\}_{c=1}^C, \boldsymbol{\theta})

    while the GMM parameters {ϕc}c=1C\{\boldsymbol{\phi}_c\}_{c=1}^C are updated purely through generative maximum-likelihood estimation via Expectation-Maximization (EM).

  2. Knowl 2 — Momentum Sinkhorn EM Optimization for Class-Conditional GMMs

    algorithm

    To train the generative GMM classifier alongside the evolving feature extractor fhetaf_{\boldsymbol{ heta}}, GMMSeg employs an online Sinkhorn Expectation-Maximization (EM) algorithm with momentum updates and an external memory queue.

    For class cc, let NcN_c denote the number of training samples labeled as cc across the current mini-batch and an external first-in-first-out (FIFO) feature memory queue. The memory queue holds 32K32\text{K} pixel representations per Gaussian component per class, populated by sampling a sparse set of 100 pixels per class from each training image. Under a uniform mixture prior constraint (equipartition constraint πcm=1/M\pi_{cm} = 1/M), the E-step is formulated as an entropy-regularized optimal transport problem:

    min⁡Qc∈Qc′∑n=1Nc∑m=1MQc(n,m)Oc(n,m)+ϵH(Qc)\min_{\mathbf{Q}_c \in \mathcal{Q}'_c} \sum_{n=1}^{N_c} \sum_{m=1}^M \mathbf{Q}_c(n, m) \mathbf{O}_c(n, m) + \epsilon H(\mathbf{Q}_c)

    where Oc(n,m)=−log⁡N(xn;μcm,Σcm)\mathbf{O}_c(n, m) = -\log \mathcal{N}(\mathbf{x}_n; \boldsymbol{\mu}_{cm}, \boldsymbol{\Sigma}_{cm}) is the cost matrix, H(Qc)=−∑n,mQc(n,m)log⁡Qc(n,m)H(\mathbf{Q}_c) = -\sum_{n, m} \mathbf{Q}_c(n, m) \log \mathbf{Q}_c(n, m) is the entropy regularizer with regularization weight ϵ=0.05\epsilon = 0.05, and Qc′={Qc∈R+Nc×M:Qc1M=1Nc,Qc⊤1Nc=NcM1M}\mathcal{Q}'_c = \left\{\mathbf{Q}_c \in \mathbb{R}_+^{N_c \times M} : \mathbf{Q}_c \mathbf{1}^M = \mathbf{1}^{N_c}, \mathbf{Q}_c^\top \mathbf{1}^{N_c} = \frac{N_c}{M} \mathbf{1}^M\right\}. This optimal transport plan is solved via Sinkhorn-Knopp iterations.

    The procedure executed at each training iteration is:

    Input: Current batch features and labels B={(xn,cn)}\mathcal{B} = \{(\mathbf{x}_n, c_n)\}, external memory queues M\mathcal{M}, current GMM parameters ϕc(t−1)={μcm,Σcm}\boldsymbol{\phi}^{(t-1)}_c = \{\boldsymbol{\mu}_{cm}, \boldsymbol{\Sigma}_{cm}\}, momentum coefficient τ=0.999\tau = 0.999, entropy weight ϵ=0.05\epsilon = 0.05.
    Output: Updated GMM parameters ϕc(t)\boldsymbol{\phi}^{(t)}_c and gradients for feature extractor θ\boldsymbol{\theta}.
    for each class c∈{1,…,C}c \in \{1, \dots, C\} do
        Gather all features {xn:cn=c}\{\mathbf{x}_n : c_n = c\} from B∪M\mathcal{B} \cup \mathcal{M}
        Compute cost matrix Oc(n,m)=−log⁡N(xn;μcm(t−1),Σcm(t−1))\mathbf{O}_c(n, m) = -\log \mathcal{N}(\mathbf{x}_n; \boldsymbol{\mu}^{(t-1)}_{cm}, \boldsymbol{\Sigma}^{(t-1)}_{cm}) for m∈{1,…,M}m \in \{1, \dots, M\}
        Solve for optimal transport assignment matrix Qc\mathbf{Q}_c using Sinkhorn-Knopp iterations with entropy scale ϵ\epsilon
        Let qcn[m]=Qc(n,m)q_{cn}[m] = \mathbf{Q}_c(n, m) and Ncm=∑n:cn=cqcn[m]N_{cm} = \sum_{n: c_n=c} q_{cn}[m]
        Compute candidate mean: μ^cm=1Ncm∑n:cn=cqcn[m]xn\hat{\boldsymbol{\mu}}_{cm} = \frac{1}{N_{cm}} \sum_{n: c_n=c} q_{cn}[m] \mathbf{x}_n
        Compute candidate diagonal covariance: Σ^cm=diag⁡(1Ncm∑n:cn=cqcn[m](xn−μ^cm)(xn−μ^cm)⊤)\hat{\boldsymbol{\Sigma}}_{cm} = \operatorname{diag}\left(\frac{1}{N_{cm}} \sum_{n: c_n=c} q_{cn}[m] (\mathbf{x}_n - \hat{\boldsymbol{\mu}}_{cm})(\mathbf{x}_n - \hat{\boldsymbol{\mu}}_{cm})^\top\right)
        Apply momentum update: μcm(t)=(1−τ)μcm(t−1)+τμ^cm\boldsymbol{\mu}^{(t)}_{cm} = (1 - \tau) \boldsymbol{\mu}^{(t-1)}_{cm} + \tau \hat{\boldsymbol{\mu}}_{cm}
        Apply momentum update: Σcm(t)=(1−τ)Σcm(t−1)+τΣ^cm\boldsymbol{\Sigma}^{(t)}_{cm} = (1 - \tau) \boldsymbol{\Sigma}^{(t-1)}_{cm} + \tau \hat{\boldsymbol{\Sigma}}_{cm}
    end for
    Compute cross-entropy loss Lce\mathcal{L}_{\text{ce}} on B\mathcal{B} using updated GMM parameters and backpropagate gradients to θ\boldsymbol{\theta}
    Update external memory M\mathcal{M} with up to 100 sampled pixel features per class from B\mathcal{B}
  3. Knowl 3 — Closed-Set and Out-of-Distribution Inference in GMMSeg

    model/method

    GMMSeg supports both standard closed-set semantic segmentation and out-of-distribution (anomaly) segmentation at inference time using the same learned model instance without architectural changes, retraining, or post-hoc calibration.

    1. Closed-Set Semantic Segmentation: Given an input feature vector x=fheta(x)\mathbf{x} = f_{\boldsymbol{ heta}}(x) extracted from a pixel xx, the predicted class label c^\hat{c} is obtained via Bayes' rule under a uniform prior p(c)=1/Cp(c) = 1/C (simplified under the winner-take-all assumption):

    c^=arg⁡max⁡c∈{1,…,C}p(c∣x)=arg⁡max⁡c∈{1,…,C}p(x∣c;ϕc)\hat{c} = \arg\max_{c \in \{1, \dots, C\}} p(c|\mathbf{x}) = \arg\max_{c \in \{1, \dots, C\}} p(\mathbf{x}|c; \boldsymbol{\phi}_c)

    1. Open-World Anomaly / Out-of-Distribution Segmentation: Because the GMM explicitly measures class-conditional feature likelihoods p(x∣c)p(\mathbf{x}|c), anomalous or out-of-distribution pixels naturally fall into low-probability density regions across all known classes. The pixel-level anomaly score s(x)s(\mathbf{x}) is directly computed as:

    s(x)=−max⁡c∈{1,…,C}p(x∣c;ϕc)s(\mathbf{x}) = -\max_{c \in \{1, \dots, C\}} p(\mathbf{x}|c; \boldsymbol{\phi}_c)

    Pixels with an anomaly score exceeding a chosen decision threshold are identified as out-of-distribution objects.

  4. Knowl 4 — Closed-Set Semantic Segmentation Benchmarking

    data/table

    GMMSeg was evaluated on three closed-set semantic segmentation benchmarks: ADE20K validation set (150 classes), Cityscapes validation set (19 classes), and COCO-Stuff test set (171 classes). Evaluated segmentation architectures include DeepLabV3+, OCRNet, UPerNet, and SegFormer using ResNet-101, HRNetV2-W48, Swin-Base, and MiT-B5 backbones. Performance is reported in mean Intersection over Union (mIoU, %).

    Method Backbone ADE20K Cityscapes COCO-Stuff
    FCN ResNet101 39.9 75.5 32.6
    PSPNet ResNet101 44.4 79.8 37.8
    SETR ViT-Large 48.2 79.2 -
    Segmenter ViT-Large 51.8 79.1 -
    MaskFormer Swin-Base 52.7 - -
    DeepLabV3+ ResNet101 45.5 80.6 33.8
    GMMSeg ResNet101 46.7 (+1.2) 81.1 (+0.5) 35.5 (+1.7)
    OCRNet HRNetV2-W48 43.3 80.4 37.6
    GMMSeg HRNetV2-W48 44.8 (+1.5) 81.2 (+0.8) 39.2 (+1.6)
    UPerNet Swin-Base 48.0 81.1 43.4
    GMMSeg Swin-Base 49.0 (+1.0) 81.8 (+0.7) 44.3 (+0.9)
    SegFormer MiT-B5 50.0 82.0 44.0
    GMMSeg MiT-B5 50.6 (+0.6) 82.6 (+0.6) 44.7 (+0.7)

    Across all four segmentation backbones/architectures, replacing the standard discriminative softmax classifier with GMMSeg improves mIoU by 0.6%–1.5% on ADE20K, 0.5%–0.8% on Cityscapes, and 0.7%–1.7% on COCO-Stuff.

  5. Knowl 5 — Open-World Anomaly Segmentation Benchmarking

    data/table

    Models trained exclusively on Cityscapes (19 classes) were evaluated on out-of-distribution anomaly segmentation benchmarks without retraining or post-hoc calibrators: Fishyscapes Lost&Found validation set and Road Anomaly dataset. Metrics include Area Under the Receiver Operating Characteristic curve (AUROC, %), Average Precision (AP, %), and False Positive Rate at 95% True Positive Rate (FPR95_{95}, %, lower is better).

    Extra OOD Fishyscapes LostFound Road Anomaly
    Method Resynthesis Data mIoU AUROC ↑\uparrow AP ↑\uparrow FPR95↓_{95} \downarrow AUROC ↑\uparrow AP ↑\uparrow FPR95↓_{95} \downarrow
    SynthCP ✓ ✓ 80.3 88.34 6.54 45.95 76.08 24.86 64.69
    SynBoost ✓ ✓ - 96.21 60.58 31.02 81.91 38.21 64.75
    MSP × × 80.3 86.99 6.02 45.63 73.76 20.59 68.44
    Entropy × × 80.3 88.32 13.91 44.85 75.12 22.38 68.15
    SML × × 80.3 96.88 36.55 14.53 81.96 25.82 49.74
    Mahalanobis × × 80.3 92.51 27.83 30.17 76.73 22.85 59.20
    GMMSeg-FCN × × 76.7 96.28 32.94 16.07 78.99 24.51 56.95
    GMMSeg-DeepLabV3+ × × 81.1 97.34 43.47 13.11 84.71 34.42 47.90
    GMMSeg-SegFormer × × 82.6 97.83 50.03 12.55 89.37 57.65 44.34

    Under the setting without extra resynthesis models or auxiliary out-of-distribution training data, GMMSeg-DeepLabV3+ achieves 97.34% AUROC, 43.47% AP, and 13.11% FPR95_{95} on Fishyscapes Lost&Found, outperforming discriminative post-processing methods (MSP, SML, Mahalanobis). GMMSeg-SegFormer further achieves 97.83% AUROC / 50.03% AP on Fishyscapes Lost&Found and 89.37% AUROC / 57.65% AP on Road Anomaly.

  6. Knowl 6 — Efficacy of Online Hybrid Training vs. Post-Hoc and Fully Discriminative GMM Optimization

    empirical result

    The hybrid training strategy of GMMSeg was analyzed against two alternative optimization paradigms using DeepLabV3+-ResNet101:

    1. Post-Hoc GMM Fitting vs. Online Hybrid Training: Fitting a GMM classifier directly onto a pre-trained frozen feature space produced by a standard softmax classifier (DeepLabV3+ + GMM) results in 31.6% mIoU on ADE20K validation. In contrast, online iterative co-adaptation of the feature extractor and GMM classifier (GMMSeg-DeepLabV3+) achieves 46.0% mIoU (+14.4% mIoU).

    2. Generative vs. Discriminative GMM Updates: When all GMM parameters {ϕ,θ}\{\boldsymbol{\phi}, \boldsymbol{\theta}\} are optimized end-to-end using purely discriminative cross-entropy loss (max⁡ϕ,θp(c∣x;ϕ,θ)\max_{\boldsymbol{\phi}, \boldsymbol{\theta}} p(c|\mathbf{x}; \boldsymbol{\phi}, \boldsymbol{\theta})), closed-set performance on Cityscapes remains comparable (81.0% vs. 81.1% mIoU for generative Sinkhorn EM). However, out-of-distribution anomaly detection performance on Fishyscapes Lost&Found degrades significantly under discriminative GMM training: AUROC drops from 97.34% to 89.77%, AP drops from 43.47% to 17.68%, and FPR95_{95} increases from 13.11% to 51.81%.

  7. Knowl 7 — Sinkhorn EM vs. Vanilla EM and EM Loop Count in GMMSeg

    empirical result

    Ablation on ADE20K validation using DeepLabV3+-ResNet101 compares vanilla EM with Sinkhorn EM and evaluates the number of EM loops executed per training iteration:

    1. Vanilla EM vs. Sinkhorn EM: Running 1 loop of vanilla EM yields 42.7% mIoU, increasing to 44.8% mIoU with 10 loops. In contrast, Sinkhorn EM reaches 46.0% mIoU with a single loop (t=1t=1), matching the performance of 5 or 10 loops (both 46.0% mIoU). The equipartition constraint in Sinkhorn EM provides stronger regularization against suboptimal local minima.

    2. Number of EM Loops: In the online momentum formulation, executing t=1t=1 Sinkhorn EM loop per training iteration is sufficient to track the slowly shifting feature space representations.

  8. Knowl 8 — Effect of Gaussian Component Count per Class in GMMSeg

    empirical result

    Ablation of the number of Gaussian components per class MM in GMMSeg using DeepLabV3+-ResNet101 on ADE20K validation demonstrates:

    • M=1M = 1 (single Gaussian per class estimated via Gaussian Discriminant Analysis): 44.2% mIoU.
    • M=3M = 3: 45.3% mIoU.
    • M=5M = 5: 46.0% mIoU.
    • M=10M = 10: 46.0% mIoU.
    • M=15M = 15: 45.7% mIoU.

    Increasing MM from 1 to 5 confirms that modeling within-class multimodality improves segmentation accuracy. Beyond M=5M=5, performance plateaus and slightly declines due to overparameterization.

  9. Knowl 9 — Prediction Confidence Calibration and Computational Efficiency of GMMSeg

    empirical result

    GMMSeg improves probability calibration over standard softmax-based discriminative networks while adding minimal runtime overhead:

    1. Confidence Calibration: On Cityscapes validation evaluated with DeepLabV3+, GMMSeg achieves an Expected Calibration Error (ECE) of 0.0766 compared to 0.1065 for the baseline softmax classifier. Reliability diagrams confirm smaller discrepancies between predicted confidence and empirical accuracy across confidence bins.

    2. Inference Speed: Measured on a single NVIDIA GeForce RTX 3090 GPU with batch size 1, GMMSeg-DeepLabV3+ operates at 13.37 frames per second (fps), representing negligible inference overhead compared to the softmax baseline at 14.16 fps.

Coverage note — No substantial contributed material was omitted from the main paper.

References

  1. 1.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 1, 3, 8
  2. 2.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE TPAMI, 2017. 1, 3
  3. 3.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In CVPR, 2017. 1, 3, 8
  4. 4.Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In CVPR, 2018. 1, 3
  5. 5.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene segmentation. In CVPR, 2019. 1, 3
  6. 6.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 1
  7. 7.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. In NeurIPS, 2021. 1, 2, 3, 7, 8, 16
  8. 8.Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. Segmenter: Transformer for semantic segmentation. In ICCV, 2021. 1, 3, 8, 16
  9. 9.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, 2021. 1, 3, 8
  10. 10.Yuhui Yuan, Rao Fu, Lang Huang, Weihong Lin, Chao Zhang, Xilin Chen, and Jingdong Wang. Hrformer: High-resolution transformer for dense prediction. In NeurIPS, 2021. 1, 3, 16
  11. 11.JM Bernardo, MJ Bayarri, JO Berger, AP Dawid, D Heckerman, AFM Smith, and M West. Generative or discriminative? getting the best of both worlds. Bayesian statistics, 2007. 1, 4, 5
  12. 12.Hideaki Hayashi and Seiichi Uchida. A discriminative gaussian mixture model with sparsity. In ICLR, 2021. 1, 2, 3, 4
  13. 13.Tianfei Zhou, Wenguan Wang, Ender Konukoglu, and Luc Van Gool. Rethinking semantic segmentation: A prototype view. In CVPR, 2022. 1, 3, 4, 16
  14. 14.Lynton Ardizzone, Radek Mackowiak, Carsten Rother, and Ullrich Köthe. Training normalizing flows with the information bottleneck for competitive generative classification. In NeurIPS, 2020. 2, 3
  15. 15.Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In ICML, 2017. 2, 10
  16. 16.Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. NeurIPS, 2019. 2
  17. 17.Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In ICLR, 2017. 2, 3, 9
  18. 18.Sanghun Jung, Jungsoo Lee, Daehoon Gwak, Sungha Choi, and Jaegul Choo. Standardized max logits: A simple yet effective approach for identifying unexpected road obstacles in urban-scene segmentation. In ICCV, 2021. 2, 3, 8, 9
  19. 19.Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. NeurIPS, 2018. 2, 9
  20. 20.Bradley Efron. The efficiency of logistic regression compared to normal discriminant analysis. Journal of the American Statistical Association, 1975. 2
  21. 21.Andrew Ng and Michael Jordan. On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. In NeurIPS, 2001. 2, 3
  22. 22.Eric Nalisnick, Akihiro Matsukawa, Yee Whye Teh, Dilan Gorur, and Balaji Lakshminarayanan. Hybrid models with deep and invertible features. In ICML, 2019. 2, 3
  23. 23.Pavel Izmailov, Polina Kirichenko, Marc Finzi, and Andrew Gordon Wilson. Semi-supervised learning with normalizing flows. In ICML, 2020. 2, 3
  24. 24.Radek Mackowiak, Lynton Ardizzone, Ullrich Kothe, and Carsten Rother. Generative classifiers as a basis for trustworthy image classification. In CVPR, 2021. 2, 3
  25. 25.Lukas Schott, Jonas Rauber, Matthias Bethge, and Wieland Brendel. Towards the first adversarially robust neural network model on mnist. In NeurIPS, 2018. 2, 3
  26. 26.Yang Song, Taesup Kim, Sebastian Nowozin, Stefano Ermon, and Nate Kushman. Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. In ICLR, 2018. 2, 3
  27. 27.Joan Serrà, David Álvarez, Vicenç Gómez, Olga Slizovskaia, José F Núñez, and Jordi Luque. Input complexity and out-of-distribution detection with likelihood-based generative models. In ICLR, 2019. 2, 3
  28. 28.Gonzalo Mena, Amin Nejatbakhsh, Erdem Varol, and Jonathan Niles-Weed. Sinkhorn em: an expectation-maximization algorithm based on entropic optimal transport. arXiv preprint arXiv:2006.16548, 2020. 2, 6, 10
  29. 29.Ehsan Variani, Erik McDermott, and Georg Heigold. A gaussian mixture model layer jointly optimized with discriminative features within a deep neural network architecture. In ICASSP, 2015. 2, 3
  30. 30.Zoltán Tüske, Muhammad Ali Tahir, Ralf Schlüter, and Hermann Ney. Integrating gaussian mixtures into deep neural networks: Softmax layer with hidden variables. In ICASSP, 2015. 2, 3
  31. 31.Aldebaro Klautau, Nikola Jevtic, and Alon Orlitsky. Discriminative gaussian mixture models: A comparison with kernel classifiers. In ICML, 2003. 2
  32. 32.Zhihao Zheng and Pengyu Hong. Robust detection of adversarial attacks by modeling the intrinsic properties of deep neural networks. In NeurIPS, 2018. 2
  33. 33.Kimin Lee, Sukmin Yun, Kibok Lee, Honglak Lee, Bo Li, and Jinwoo Shin. Robust inference via generative classifiers for handling noisy labels. In ICML, 2019. 2
  34. 34.Giancarlo Di Biase, Hermann Blum, Roland Siegwart, and Cesar Cadena. Pixel-wise anomaly detection in complex driving scenes. In CVPR, 2021. 2, 3, 9
  35. 35.Yingda Xia, Yi Zhang, Fengze Liu, Wei Shen, and Alan L Yuille. Synthesize then compare: Detecting failures and anomalies for semantic segmentation. In ECCV, 2020. 2, 3, 8, 9
  36. 36.Krzysztof Lis, Krishna Nakka, Pascal Fua, and Mathieu Salzmann. Detecting the unexpected via image resynthesis. In ICCV, 2019. 2, 3, 8, 9
  37. 37.Tomas Vojir, Tomáš Šipka, Rahaf Aljundi, Nikolay Chumerin, Daniel Olmeda Reino, and Jiri Matas. Road anomaly detection by partial image reconstruction with segmentation coupling. In ICCV, 2021. 2, 3
  38. 38.Petra Bevandic, Ivan Krešo, Marin Orši ´ c, and Siniša Šegvi ´ c. Simultaneous semantic segmentation and ´ outlier detection in presence of domain shift. In GCPR, 2019. 2, 3
  39. 39.Robin Chan, Matthias Rottmann, and Hanno Gottschalk. Entropy maximization and meta classification for out-of-distribution detection in semantic segmentation. In ICCV, 2021. 2, 3
  40. 40.Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich. Deep anomaly detection with outlier exposure. arXiv preprint arXiv:1812.04606, 2018. 2, 3
  41. 41.Shiyu Liang, Yixuan Li, and Rayadurgam Srikant. Enhancing the reliability of out-of-distribution image detection in neural networks. arXiv preprint arXiv:1706.02690, 2017. 2, 3
  42. 42.KIMIN LEE, Kibok Lee, Honglak Lee, and Jinwoo Shin. Training confidence-calibrated classifiers for detecting out-of-distribution samples. In ICLR, 2018. 2, 3
  43. 43.Matthias Rottmann, Pascal Colling, Thomas Paul Hack, Robin Chan, Fabian Hüger, Peter Schlicht, and Hanno Gottschalk. Prediction error meta classification in semantic segmentation: Detection via aggregated dispersion measures of softmax probabilities. In IJCNN, 2020. 2, 3
  44. 44.Alex Kendall and Yarin Gal. What uncertainties do we need in bayesian deep learning for computer vision? In NeurIPS, 2017. 2, 3
  45. 45.Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles. In NeurIPS, 2017. 2, 3
  46. 46.Jishnu Mukhoti and Yarin Gal. Evaluating bayesian deep learning methods for semantic segmentation. arXiv preprint arXiv:1811.12709, 2018. 2, 3
  47. 47.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018. 2, 3, 7, 8, 9
  48. 48.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In ECCV, 2020. 2, 3, 7, 8
  49. 49.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, 2018. 2, 7, 8, 10
  50. 50.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 2, 7, 9
  51. 51.Jingdong Wang, Ke Sun, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao, Dong Liu, Yadong Mu, Mingkui Tan, Xinggang Wang, et al. Deep high-resolution representation learning for visual recognition. IEEE TPAMI, 2020. 2, 7
  52. 52.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 2, 7
  53. 53.Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ade20k dataset. In CVPR, 2017. 2, 7, 8, 9, 10
  54. 54.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 2, 7, 8
  55. 55.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In CVPR, 2018. 2, 7, 8
  56. 56.Hermann Blum, Paul-Edouard Sarlin, Juan Nieto, Roland Siegwart, and Cesar Cadena. The fishyscapes benchmark: Measuring blind spots in semantic segmentation. IJCV, 2021. 2, 3, 4, 8, 9
  57. 57.Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. Deformable convolutional networks. In CVPR, 2017. 3
  58. 58.Maoke Yang, Kun Yu, Chi Zhang, Zhiwei Li, and Kuiyuan Yang. Denseaspp for semantic segmentation in street scenes. In CVPR, 2018. 3
  59. 59.Fisher Yu and Vladlen Koltun. Multi-scale context aggregation by dilated convolutions. In ICLR, 2016. 3
  60. 60.Guolei Sun, Wenguan Wang, Jifeng Dai, and Luc Van Gool. Mining cross-image semantics for weakly supervised semantic segmentation. In ECCV, 2020. 3
  61. 61.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In MICCAI, 2015. 3
  62. 62.Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip HS Torr. Conditional random fields as recurrent neural networks. In ICCV, 2015. 3
  63. 63.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE TPAMI, 2017. 3
  64. 64.Ziwei Liu, Xiaoxiao Li, Ping Luo, Chen Change Loy, and Xiaoou Tang. Deep learning markov random field for semantic segmentation. IEEE TPAMI, 2017. 3
  65. 65.Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In CVPR, 2017. 3
  66. 66.Sachin Mehta, Mohammad Rastegari, Anat Caspi, Linda Shapiro, and Hannaneh Hajishirzi. Espnet: Efficient spatial pyramid of dilated convolutions for semantic segmentation. In ECCV, 2018. 3
  67. 67.Hang Zhang, Kristin Dana, Jianping Shi, Zhongyue Zhang, Xiaogang Wang, Ambrish Tyagi, and Amit Agrawal. Context encoding for semantic segmentation. In CVPR, 2018. 3
  68. 68.Junjun He, Zhongying Deng, Lei Zhou, Yali Wang, and Yu Qiao. Adaptive pyramid context network for semantic segmentation. In CVPR, 2019. 3
  69. 69.Hanzhe Hu, Deyi Ji, Weihao Gan, Shuai Bai, Wei Wu, and Junjie Yan. Class-wise dynamic graph convolution for semantic segmentation. In ECCV, 2020. 3
  70. 70.Changqian Yu, Jingbo Wang, Changxin Gao, Gang Yu, Chunhua Shen, and Nong Sang. Context prior for scene segmentation. In CVPR, 2020. 3
  71. 71.Mingyuan Liu, Dan Schonfeld, and Wei Tang. Exploit visual dependency relations for semantic segmentation. In CVPR, 2021. 3
  72. 72.Chi-Wei Hsiao, Cheng Sun, Hwann-Tzong Chen, and Min Sun. Specialize and fuse: Pyramidal output representation for semantic segmentation. In ICCV, 2021. 3
  73. 73.Zhenchao Jin, Bin Liu, Qi Chu, and Nenghai Yu. Isnet: Integrate image-level and semantic-level context for semantic segmentation. In ICCV, 2021. 3
  74. 74.Zhenchao Jin, Tao Gong, Dongdong Yu, Qi Chu, Jian Wang, Changhu Wang, and Jie Shao. Mining contextual information beyond image for semantic segmentation. In ICCV, 2021. 3
  75. 75.Wenguan Wang, Tianfei Zhou, Fisher Yu, Jifeng Dai, Ender Konukoglu, and Luc Van Gool. Exploring cross-image pixel contrast for semantic segmentation. In ICCV, 2021. 3
  76. 76.Jiaxu Miao, Yunchao Wei, Yu Wu, Chen Liang, Guangrui Li, and Yi Yang. Vspw: A large-scale dataset for video scene parsing in the wild. In CVPR, 2021. 3
  77. 77.Zongxin Yang, Yunchao Wei, and Yi Yang. Collaborative video object segmentation by multi-scale foreground-background integration. IEEE TPAMI, 2021. 3
  78. 78.Adam W Harley, Konstantinos G Derpanis, and Iasonas Kokkinos. Segmentation-aware convolutional networks using local attention masks. In ICCV, 2017. 3
  79. 79.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018. 3
  80. 80.Hanchao Li, Pengfei Xiong, Jie An, and Lingxue Wang. Pyramid attention network for semantic segmentation. arXiv preprint arXiv:1805.10180, 2018. 3
  81. 81.Hengshuang Zhao, Yi Zhang, Shu Liu, Jianping Shi, Chen Change Loy, Dahua Lin, and Jiaya Jia. Psanet: Point-wise spatial attention network for scene parsing. In ECCV, 2018. 3
  82. 82.Junjun He, Zhongying Deng, and Yu Qiao. Dynamic multi-scale filters for semantic segmentation. In ICCV, 2019. 3
  83. 83.Xia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang, Zhouchen Lin, and Hong Liu. Expectation-maximization attention networks for semantic segmentation. In ICCV, 2019. 3
  84. 84.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In ICCV, 2019. 3
  85. 85.Liulei Li, Tianfei Zhou, Wenguan Wang, Jianwu Li, and Yi Yang. Deep hierarchical semantic segmentation. In CVPR, 2022. 3
  86. 86.Wenguan Wang, Hailong Zhu, Jifeng Dai, Yanwei Pang, Jianbing Shen, and Ling Shao. Hierarchical human parsing with typed part-relation reasoning. In CVPR, 2020. 3
  87. 87.Wenguan Wang, Zhijie Zhang, Siyuan Qi, Jianbing Shen, Yanwei Pang, and Ling Shao. Learning compositional neural information fusion for human parsing. In ICCV, 2019. 3
  88. 88.Zongxin Yang, Yunchao Wei, and Yi Yang. Associating objects with transformers for video object segmentation. NeurIPS, 2021. 3
  89. 89.Zongxin Yang, Jiaxu Miao, Xiaohan Wang, Yunchao Wei, and Yi Yang. Associating objects with scalable transformers for video object segmentation. arXiv preprint arXiv:2203.11442, 2022. 3
  90. 90.Bowen Cheng, Alexander G. Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. In NeurIPS, 2021. 3, 8
  91. 91.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. CVPR, 2022. 3
  92. 92.Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang. Large-margin softmax loss for convolutional neural networks. In ICML, 2016. 3
  93. 93.Yi Yang, Yueting Zhuang, and Yunhe Pan. Multiple knowledge representation for big data artificial intelligence: framework, applications, and case studies. Frontiers of Information Technology & Electronic Engineering, 2021. 3
  94. 94.Wenguan Wang and Yi Yang. Towards data-and knowledge-driven artificial intelligence: A survey on neuro-symbolic computing. arXiv preprint arXiv:2210.15889, 2022. 3
  95. 95.Wenguan Wang, Cheng Han, Tianfei Zhou, and Dongfang Liu. Visual recognition with deep nearest centroids. arXiv preprint arXiv:2209.07383, 2022. 3
  96. 96.Jyh-Jing Hwang, Stella X Yu, Jianbo Shi, Maxwell D Collins, Tien-Ju Yang, Xiao Zhang, and Liang-Chieh Chen. Segsort: Segmentation by discriminative sorting of segments. In ICCV, 2019. 3
  97. 97.Arindam Banerjee, Inderjit S Dhillon, Joydeep Ghosh, Suvrit Sra, and Greg Ridgeway. Clustering on the unit hypersphere using von mises-fisher distributions. Journal of Machine Learning Research, 2005. 3
  98. 98.Rajat Raina, Yirong Shen, Andrew Mccallum, and Andrew Ng. Classification with hybrid generative/discriminative models. NeurIPS, 2003. 3
  99. 99.Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. Improved techniques for training gans. In NeurIPS, 2016. 3
  100. 100.Ricky TQ Chen, Jens Behrmann, David K Duvenaud, and Jörn-Henrik Jacobsen. Residual flows for invertible generative modeling. In NeurIPS, 2019. 3
  101. 101.Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky. Your classifier is secretly an energy based model and you should treat it like one. In ICLR, 2019. 3
  102. 102.Yingzhen Li, John Bradshaw, and Yash Sharma. Are generative classifiers more robust to adversarial attacks? In ICML, 2019. 3
  103. 103.Xinshuai Dong, Hong Liu, Rongrong Ji, Liujuan Cao, Qixiang Ye, Jianzhuang Liu, and Qi Tian. Api-net: Robust generative classifier via a single discriminator. In ECCV, 2020. 3
  104. 104.Ethan Fetaya, Jörn-Henrik Jacobsen, Will Grathwohl, and Richard Zemel. Understanding the limitations of conditional generative models. In ICLR, 2020. 3
  105. 105.Eric Nalisnick, Akihiro Matsukawa, Yee Whye Teh, Dilan Gorur, and Balaji Lakshminarayanan. Do deep generative models know what they don’t know? In ICLR, 2019. 3
  106. 106.Florian Wenzel, Théo Galy-Fajou, Christan Donner, Marius Kloft, and Manfred Opper. Efficient gaussian process classification using pòlya-gamma data augmentation. In AAAI, 2019. 3
  107. 107.Tommi Jaakkola and David Haussler. Exploiting generative models in discriminative classifiers. In NeurIPS, 1998. 3
  108. 108.Julia A Lasserre, Christopher M Bishop, and Thomas P Minka. Principled hybrids of generative and discriminative models. In CVPR, 2006. 3
  109. 109.Hyunsun Choi, Eric Jang, and Alexander A Alemi. Waic, but why? generative ensembles for robust anomaly detection. arXiv preprint arXiv:1810.01392, 2018. 3, 4
  110. 110.Jie Ren, Peter J Liu, Emily Fertig, Jasper Snoek, Ryan Poplin, Mark Depristo, Joshua Dillon, and Balaji Lakshminarayanan. Likelihood ratios for out-of-distribution detection. In NeurIPS, 2019. 3, 4
  111. 111.Durk P Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling. Semi-supervised learning with deep generative models. In NeurIPS, 2014. 4
  112. 112.Guillaume Bouchard and Bill Triggs. The tradeoff between generative and discriminative classifiers. In IASC International Symposium on Computational Statistics, 2004. 4
  113. 113.Thomas Mensink, Jakob Verbeek, Florent Perronnin, and Gabriela Csurka. Distance-based image classification: Generalizing to new classes at near-zero cost. IEEE TPAMI, 2013. 4
  114. 114.Kateryna Chumachenko, Alexandros Iosifidis, and Moncef Gabbouj. Feedforward neural networks initialization based on discriminant learning. Neural Networks, 146, 2022. 4
  115. 115.Murat Sensoy, Lance Kaplan, and Melih Kandemir. Evidential deep learning to quantify classification uncertainty. In NeurIPS, 2020. 4
  116. 116.Arthur P Dempster, Nan M Laird, and Donald B Rubin. Maximum likelihood from incomplete data via the em algorithm. Journal of the Royal Statistical Society: Series B (Methodological), 1977. 5
  117. 117.Radford M Neal and Geoffrey E Hinton. A view of the em algorithm that justifies incremental, sparse, and other variants. In Learning in graphical models. 1998. 5
  118. 118.Naonori Ueda and Ryohei Nakano. Deterministic annealing variant of the em algorithm. NeurIPS, 1994. 6
  119. 119.Yuki Markus Asano, Christian Rupprecht, and Andrea Vedaldi. Self-labelling via simultaneous clustering and representation learning. In ICLR, 2020. 6
  120. 120.Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In NeurIPS, 2013. 6
  121. 121.Steven Nowlan. Maximum likelihood competitive learning. NeurIPS, 1989. 7
  122. 122.Nanda Kambhatla and Todd Leen. Classifying with gaussian mixtures and clusters. NeurIPS, 1994. 7
  123. 123.Xuefeng Du, Xin Wang, Gabriel Gozum, and Yixuan Li. Unknown-aware object detection: Learning what you don’t know from videos in the wild. In CVPR, 2022. 7
  124. 124.MMSegmentation Contributors. MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020. 7
  125. 125.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li. Imagenet large scale visual recognition challenge. IJCV, 2015. 7
  126. 126.Peter Pinggera, Sebastian Ramos, Stefan Gehrig, Uwe Franke, Carsten Rother, and Rudolf Mester. Lost and found: detecting small road hazards for self-driving vehicles. In IROS, 2016. 8
  127. 127.Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li. Energy-based out-of-distribution detection. NeurIPS, 2020. 9

Citation

MLA
Liang, C., et al. “GMMSeg: Gaussian Mixture Based Generative Semantic Segmentation Models”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 31360–75, https://proceedings.neurips.cc/paper_files/paper/2022/file/cb1c4782f159b55380b4584671c4fd88-Paper-Conference.pdf.
APA
Liang, C., Wang, W., Miao, J., & Yang, Y. (2022). GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models. Advances in Neural Information Processing Systems, 35, 31360–31375. https://proceedings.neurips.cc/paper_files/paper/2022/file/cb1c4782f159b55380b4584671c4fd88-Paper-Conference.pdf
Chicago
Liang, C., W. Wang, J. Miao, and Y. Yang. 2022. “GMMSeg: Gaussian Mixture Based Generative Semantic Segmentation Models”. Advances in Neural Information Processing Systems 35: 31360–75. https://proceedings.neurips.cc/paper_files/paper/2022/file/cb1c4782f159b55380b4584671c4fd88-Paper-Conference.pdf.
Harvard
Liang, C. et al. (2022) “GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 31360–31375. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/cb1c4782f159b55380b4584671c4fd88-Paper-Conference.pdf.
Vancouver
1. Liang C, Wang W, Miao J, Yang Y (2022) GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 31360–31375

BibTeX

@inproceedings{liang2022gmmseg,
  title = {GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models},
  author = {Liang, Chen and Wang, Wenguan and Miao, Jiaxu and Yang, Yi},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {31360-31375},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/cb1c4782f159b55380b4584671c4fd88-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors