Built independently by an author, for readers. Read the story and support ChapterPal

keyword

generative semantic segmentation

Generative semantic segmentation is an approach in computer vision that assigns semantic category labels to individual image pixels by modeling underlying data distributions or formulating mask prediction as a generative process. Unlike conventional discriminative methods that rely strictly on direct per-pixel classification, generative semantic segmentation incorporates generative modeling principles, such as estimating joint or class-conditional feature densities and learning latent priors for segmentation masks. By capturing the statistical structures and representations of visual scenes and their corresponding label layouts, this paradigm aims to improve model robustness, out-of-distribution detection, and cross-domain generalization while maintaining dense pixel-level labeling accuracy.

2 items

GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models

GMMSeg: Gaussian Mixture based Generative Semantic Segmentation Models

Chen Liang, Wenguan Wang, Jiaxu Miao, Yi Yang

OrganizationsBaiduUniversity of Technology SydneyZhejiang University

Why you should read this

Proposes a hybrid semantic segmentation framework that models pixel feature densities via Gaussian Mixture Models using online Expectation-Maximization alongside end-to-end discriminative representation learning, delivering superior closed-set accuracy while naturally detecting out-of-distribution anomalies without architectural changes or post-processing.

Prevalent semantic segmentation solutions are, in essence, a dense discriminative classifier of p(class | pixel feature). Though straightforward, this de facto paradigm neglects the underlying data distribution p(pixel feature | class), and struggles to identify out-of-distribution data. Going beyond this, we propose GMMSeg, a new family of segmentation models that rely on a dense generative classifier for the joint distribution p(pixel feature, class). For each class, GMMSeg builds Gaussian Mixture Models (GMMs) via Expectation-Maximization (EM), so as to capture class-conditional densities. Meanwhile, the deep dense representation is end-to-end trained in a discriminative manner, i.e., maximizing p(class | pixel feature). This endows GMMSeg with the strengths of both generative and discriminative models. With a variety of segmentation architectures and backbones, GMMSeg outperforms the discriminative counterparts on three closed-set datasets. More impressively, without any modification, GMMSeg even performs well on open-world datasets. We believe this work brings fundamental insights into the related fields.

Added

2026-09-26

Generative Semantic Segmentation

Generative Semantic Segmentation

Jiaqi Chen, Jiachen Lu, Xiatian Zhu, Li Zhang

OrganizationsFudan UniversityUniversity of Surrey

Why you should read this

Proposes Generative Semantic Segmentation to cast semantic segmentation as an image-conditioned mask generation problem using discrete latent priors and mask-as-image representations, delivering state-of-the-art generalization on challenging cross-domain benchmarks.

We present Generative Semantic Segmentation (GSS), a generative learning approach for semantic segmentation. Uniquely, we cast semantic segmentation as an image-conditioned mask generation problem. This is achieved by replacing the conventional per-pixel discriminative learning with a latent prior learning process. Specifically, we model the variational posterior distribution of latent variables given the segmentation mask. To that end, the segmentation mask is expressed with a special type of image (dubbed as maskige). This posterior distribution allows to generate segmentation masks unconditionally. To achieve semantic segmentation on a given image, we further introduce a conditioning network. It is optimized by minimizing the divergence between the posterior distribution of maskige (i.e. segmentation masks) and the latent prior distribution of input training images. Extensive experiments on standard benchmarks show that our GSS can perform competitively to prior art alternatives in the standard semantic segmentation setting, whilst achieving a new state of the art in the more challenging cross-domain setting.

Added

2026-09-26