Generative Semantic Segmentation

Jiaqi ChenJiachen LuXiatian ZhuLi Zhang

article2023CVPR75 citations

Proposes Generative Semantic Segmentation to cast semantic segmentation as an image-conditioned mask generation problem using discrete latent priors and mask-as-image representations, delivering state-of-the-art generalization on challenging cross-domain benchmarks.

Listen

Semantic segmentation, which involves classifying every individual pixel within an image, is a foundational task for automated visual systems in robotics, autonomous driving, and scene parsing. Conventional approaches treat this strictly as a per-pixel discriminative classification problem, which often struggles when transferring across diverse visual environments and fails to exploit powerful, pre-trained image generation systems.

The article demonstrates that semantic segmentation can be effectively reformulated as an image-conditioned mask generation task using a framework named Generative Semantic Segmentation. The central objective is to evaluate whether casting segmentation as generative modeling can achieve competitive accuracy in standard environments while delivering superior generalization in challenging cross-domain applications.

To overcome the computational bottleneck typically associated with training large generative architectures from scratch, the authors introduce a concept called "maskige," which expresses multi-class segmentation masks in a standard three-channel color image format. This formulation allows the framework to directly utilize off-the-shelf, pre-trained discrete generative autoencoders (specifically VQ-VAE) alongside lightweight transformation modules and standard vision backbones. The approach was evaluated through extensive benchmark experiments on established datasets, including Cityscapes for urban driving, ADE20K for complex scene parsing, and MSeg for cross-domain evaluation.

The empirical findings demonstrate three primary outcomes. First, the generative approach establishes a new state of the art in cross-domain segmentation on the composite MSeg benchmark, outperforming dedicated domain generalization and multi-task baselines with a 61.9 harmonic mean score. Second, in standard single-domain benchmarks, the model achieves competitive accuracy relative to leading discriminative models, recording mean Intersection over Union scores of up to 80.05% on Cityscapes and 48.54% on ADE20K. Third, the maskige design drastically reduces training overhead; the training-free linear variant requires zero first-stage reconstruction training, whereas prior generative approaches required thousands of computing hours.

These results imply that generative modeling offers a robust, domain-agnostic alternative to traditional per-pixel discriminative frameworks without demanding prohibitive computational budgets. By leveraging existing large generative representations, organizations can achieve better out-of-domain reliability and mitigate the performance drops typically encountered when deploying visual models in novel visual domains.

For practical adoption, organizations requiring rapid deployment with minimal compute should implement the training-free linear variant, while applications prioritizing maximum segmentation precision should adopt the Transformer-enhanced non-linear transformation. When handling real-world datasets with missing or incomplete annotations, teams should integrate the proposed auxiliary pseudo-labeling strategy to prevent generative bias toward unlabeled regions.

Confidence in these findings is strong across the evaluated urban, indoor, and composite benchmarks. However, stakeholders should note that the framework's efficiency relies heavily on pre-trained discrete generative codebooks, and applications involving fine-grained non-linear transformations still require moderate training overhead.

Cover for Generative Semantic Segmentation

Abstract

We present Generative Semantic Segmentation (GSS), a generative learning approach for semantic segmentation. Uniquely, we cast semantic segmentation as an image-conditioned mask generation problem. This is achieved by replacing the conventional per-pixel discriminative learning with a latent prior learning process. Specifically, we model the variational posterior distribution of latent variables given the segmentation mask. To that end, the segmentation mask is expressed with a special type of image (dubbed as maskige). This posterior distribution allows to generate segmentation masks unconditionally. To achieve semantic segmentation on a given image, we further introduce a conditioning network. It is optimized by minimizing the divergence between the posterior distribution of maskige (i.e. segmentation masks) and the latent prior distribution of input training images. Extensive experiments on standard benchmarks show that our GSS can perform competitively to prior art alternatives in the standard semantic segmentation setting, whilst achieving a new state of the art in the more challenging cross-domain setting.

Table of Contents

  • 2. Related work
  • Latent prior learning
  • 3. Methodology
  • 3.1. GSS formulation
  • 3.2. ELBO optimization for semantic segmentation
  • 3.3. Stage I: Efficient latent posterior learning
  • 3.4. Stage II: Latent prior learning
  • 3.5. Generative inference
  • 4. Experiment
  • 4.1. Experimental setup
  • 4.2. Ablation studies
  • 4.3. Single-domain semantic segmentation
  • 4.4. Cross-domain semantic segmentation
  • 4.5. Qualitative evaluation
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Generative Semantic Segmentation Formulation and Evidence Lower Bound

    model/method

    Generative Semantic Segmentation (GSS) reformulates semantic segmentation from per-pixel discriminative classification into image-conditioned discrete mask generation. Given an input image x∈RH×W×3x \in \mathbb{R}^{H \times W \times 3} and a target segmentation mask c∈{0,1}H×W×Kc \in \{0, 1\}^{H \times W \times K} across KK semantic categories, GSS introduces a discrete latent variable z∈ZLz \in \mathcal{Z}^L, where Z={0,1,…,V−1}\mathcal{Z} = \{0, 1, \dots, V-1\} represents a discrete codebook of size VV (e.g., V=8192V = 8192) and L=Hd×WdL = \frac{H}{d} \times \frac{W}{d} with downsample ratio dd.

    The log-likelihood of the segmentation mask conditioned on the image, log⁡p(c∣x)\log p(c|x), is optimized via its Evidence Lower Bound (ELBO): log⁡p(c∣x)≥Eqϕ(z∣c)[log⁡pθ(c∣z)]−DKL(qϕ(z∣c)∥pψ(z∣x))\log p(c|x) \ge \mathbb{E}_{q_\phi(z|c)} [\log p_\theta(c|z)] - D_{\text{KL}}\big(q_\phi(z|c) \parallel p_\psi(z|x)\big) where:

    • qϕ(z∣c)q_\phi(z|c) is the variational posterior distribution parameterized by a mask encoder EϕE_\phi and transformation X\mathcal{X}, encoding the mask cc into discrete latent tokens zz.
    • pθ(c∣z)p_\theta(c|z) is the mask reconstruction distribution parameterized by a mask decoder DθD_\theta and inverse transformation X−1\mathcal{X}^{-1}.
    • pψ(z∣x)p_\psi(z|x) is the image-conditioned prior distribution over latent tokens modeled by an image encoder IψI_\psi.
    • DKL(⋅∥⋅)D_{\text{KL}}(\cdot \parallel \cdot) denotes the Kullback-Leibler divergence.

    Optimization proceeds in two decoupled stages: Stage I optimizes the reconstruction term Eqϕ(z∣c)[log⁡pθ(c∣z)]\mathbb{E}_{q_\phi(z|c)} [\log p_\theta(c|z)] (latent posterior learning), and Stage II optimizes the prior alignment DKL(qϕ(z∣c)∥pψ(z∣x))D_{\text{KL}}\big(q_\phi(z|c) \parallel p_\psi(z|x)\big) (latent prior learning).

  2. Knowl 2 — Maskige Concept and Stage I Latent Posterior Learning

    model/method

    To exploit off-the-shelf discrete generative models (e.g., VQVAE pretrained on large-scale RGB image datasets) without retraining from scratch for varying numbers of semantic categories KK, GSS introduces maskige, an RGB representation x(c)∈RH×W×3x^{(c)} \in \mathbb{R}^{H \times W \times 3} of the categorical segmentation mask c∈{0,1}H×W×Kc \in \{0, 1\}^{H \times W \times K}. A forward transformation X:RK→R3\mathcal{X}: \mathbb{R}^K \to \mathbb{R}^3 maps each category to a specific color, and a pseudo-inverse transformation X−1:R3→RK\mathcal{X}^{-1}: \mathbb{R}^3 \to \mathbb{R}^K maps RGB values back to category predictions.

    The Stage I posterior reconstruction objective is decomposed as: min⁡ϕ^,θ^Eqϕ^(z^∣x(c))∥Dθ^(z^)−x(c)∥+min⁡X−1Eqϕ^(z^∣X(c))∥X−1(x^(c))−c∥\min_{\hat{\phi}, \hat{\theta}} \mathbb{E}_{q_{\hat{\phi}}(\hat{z}|x^{(c)})} \| D_{\hat{\theta}}(\hat{z}) - x^{(c)} \| + \min_{\mathcal{X}^{-1}} \mathbb{E}_{q_{\hat{\phi}}(\hat{z}|\mathcal{X}(c))} \| \mathcal{X}^{-1}(\hat{x}^{(c)}) - c \| where x^(c)=Dθ^(z^)\hat{x}^{(c)} = D_{\hat{\theta}}(\hat{z}). The first term corresponds to standard RGB image reconstruction, which is fulfilled by freezing an off-the-shelf pretrained VQVAE (with encoder EϕE_\phi and decoder DθD_\theta, total 29.1M parameters). Consequently, latent posterior learning reduces to optimizing only the lightweight transformations X\mathcal{X} and X−1\mathcal{X}^{-1} (between 0.9K and 466.7K parameters) using cross-entropy loss against the ground truth cc.

  3. Knowl 3 — Maskige Transformation Variants and Maximal Distance Initialization

    model/method

    GSS realizes the forward mapping X:RK→R3\mathcal{X}: \mathbb{R}^K \to \mathbb{R}^3 and inverse mapping X−1:R3→RK\mathcal{X}^{-1}: \mathbb{R}^3 \to \mathbb{R}^K under linear and non-linear designs:

    1. Linear Design with Maximal Distance Assumption: Under the linear assumption, the maskige is formed as x(c)=cβx^{(c)} = c\beta with matrix β∈RK×3\beta \in \mathbb{R}^{K \times 3}, and inverted as c^=x^(c)β†\hat{c} = \hat{x}^{(c)}\beta^\dagger with β†∈R3×K\beta^\dagger \in \mathbb{R}^{3 \times K}. Given β\beta, the optimal pseudo-inverse is determined analytically via least squares: β†=β⊤(ββ⊤)−1\beta^\dagger = \beta^\top (\beta \beta^\top)^{-1} To eliminate the training cost of β\beta, the maximal distance assumption initializes β\beta so that the color coordinates of the KK classes are maximally dispersed in 3D Euclidean space R3\mathbb{R}^3.

    2. Concrete Maskige Architectures:

    • GSS-FF (Free-Free): X\mathcal{X} initialized with maximal distance assumption and X−1=β†\mathcal{X}^{-1} = \beta^\dagger computed by least squares; requires 0 training GPU hours.
    • GSS-FF-R: Linear mappings with β\beta randomly initialized.
    • GSS-FT: Linear X\mathcal{X} initialized with maximal distance assumption; X−1\mathcal{X}^{-1} is a non-linear 3-layer Convolutional Neural Network (CNN) trained with gradient descent.
    • GSS-FT-W: Linear X\mathcal{X} initialized with maximal distance assumption; X−1\mathcal{X}^{-1} is a single-layer Shifted Window Transformer block trained with gradient descent.
    • GSS-TF: Linear X\mathcal{X} (parameterized by β\beta) trained with gradient descent using hard Gumbel-Softmax relaxation, and β†\beta^\dagger computed by least squares.
    • GSS-TT: Linear X\mathcal{X} and 3-layer CNN X−1\mathcal{X}^{-1} jointly trained end-to-end via hard Gumbel-Softmax relaxation.
  4. Knowl 4 — Stage II Latent Prior Learning and Image Encoder Architecture

    model/method

    During Stage II, the maskige encoder EϕE_\phi, decoder DθD_\theta, and transformations X,X−1\mathcal{X}, \mathcal{X}^{-1} are frozen. The conditioning network IψI_\psi is trained to predict the latent token distribution pψ(z∣x)p_\psi(z|x) conditioned on the input RGB image x∈RH×W×3x \in \mathbb{R}^{H \times W \times 3}.

    Network Architecture: IψI_\psi consists of:

    1. A visual backbone (e.g., ResNet-101 or Swin-Large) extracting multi-scale feature maps.
    2. A Multi-Level Aggregation (MLA) module constructed from DD shifted window Transformer layers and a linear projection layer, outputting discrete code logits z∈ZH/d×W/dz \in \mathcal{Z}^{H/d \times W/d} where dd is the downsampling ratio (e.g., d=4d = 4 or d=8d = 8).

    Training Loss: Because the ground truth posterior token sequence z∗=Eϕ(X(c))z^* = E_\phi(\mathcal{X}(c)) has fixed entropy, minimizing DKL(qϕ(z∣c)∥pψ(z∣x))D_{\text{KL}}\big(q_\phi(z|c) \parallel p_\psi(z|x)\big) is equivalent to minimizing the spatial multi-class cross-entropy loss: min⁡ψE(x,c)[−∑i=1Llog⁡pψ(zi=zi∗∣x)]\min_\psi \mathbb{E}_{(x, c)} \left[ - \sum_{i=1}^{L} \log p_\psi(z_i = z_i^* \mid x) \right] where L=Hd×WdL = \frac{H}{d} \times \frac{W}{d}.

  5. Knowl 5 — Auxiliary Pseudo-Labeling for Unlabeled Image Regions

    model/method

    In discriminative semantic segmentation, unlabeled or missing-label pixels are ignored during loss computation. In generative semantic segmentation, training operates on discrete latent tokens z∈ZH/d×W/dz \in \mathcal{Z}^{H/d \times W/d}, where each token aggregates a spatial patch. Unlabeled pixels within patches cause heterogeneous token representations, which causes generative models to misclassify difficult labeled pixels as unlabeled.

    To overcome this, GSS introduces an auxiliary prediction head pξ(cˉ∣z)p_\xi(\bar{c}|z) during Stage II prior learning to generate pseudo-labels cˉ\bar{c} for unlabeled pixels. An enhanced ground-truth mask c~\tilde{c} is composed as: c~=Mu⊙cˉ+(1−Mu)⊙c\tilde{c} = M_u \odot \bar{c} + (1 - M_u) \odot c where Mu∈{0,1}H×WM_u \in \{0, 1\}^{H \times W} is a binary mask equal to 1 at unlabeled pixel locations and 0 at labeled pixel locations, and ⊙\odot denotes element-wise multiplication.

    The Stage II optimization objective with the auxiliary pseudo-labeling head is formulated as: min⁡ψDKL(qϕ(z∣c~)∥pψ(z∣x))+pξ(cˉ∣z)\min_{\psi} D_{\text{KL}}\big(q_\phi(z|\tilde{c}) \parallel p_\psi(z|x)\big) + p_\xi(\bar{c}|z)

  6. Knowl 6 — Generative Semantic Segmentation Inference Pipeline

    algorithm

    Inference in GSS generates the full segmentation mask c^\hat{c} for an input image xx via sequential latent prior sampling, maskige decoding, and maskige inversion.

    Input: Input image x∈RH×W×3x \in \mathbb{R}^{H \times W \times 3}, image conditioning network IψI_\psi, pretrained maskige decoder DθD_\theta, inverse transformation X−1\mathcal{X}^{-1}
    Output: Predicted semantic segmentation mask c^∈{0,1}H×W×K\hat{c} \in \{0, 1\}^{H \times W \times K}
    Step 1: Predict latent tokens from the input image
    z←arg⁡max⁡zpψ(z∣x)z \leftarrow \arg\max_z p_\psi(z \mid x) where z∈ZH/d×W/dz \in \mathcal{Z}^{H/d \times W/d}
    Step 2: Generate maskige in RGB space
    x^(c)←Dθ(z)\hat{x}^{(c)} \leftarrow D_\theta(z) where x^(c)∈RH×W×3\hat{x}^{(c)} \in \mathbb{R}^{H \times W \times 3}
    Step 3: Invert maskige to class categorical mask
    c^←X−1(x^(c))\hat{c} \leftarrow \mathcal{X}^{-1}(\hat{x}^{(c)}) where c^∈{0,1}H×W×K\hat{c} \in \{0, 1\}^{H \times W \times K}
    return c^\hat{c}

    For linear variants (e.g., GSS-FF), Step 3 computes c^=arg⁡max⁡k(x^(c)β†)\hat{c} = \arg\max_k (\hat{x}^{(c)} \beta^\dagger) using the pseudo-inverse matrix β†∈R3×K\beta^\dagger \in \mathbb{R}^{3 \times K}. For non-linear variants (e.g., GSS-FT-W), Step 3 passes x^(c)\hat{x}^{(c)} through a Shifted Window Transformer block.

  7. Knowl 7 — Single-Domain Semantic Segmentation Benchmark Performance

    data/table

    Comparison of GSS against discriminative and generative segmentation models on the validation sets of Cityscapes (19 classes) and ADE20K (150 classes), reported in mean Intersection over Union (mIoU, %).

    Method Pretrain Backbone Cityscapes mIoU ADE20K mIoU
    Discriminative modeling:
    FCN 1K ResNet-101 77.02 41.40
    PSPNet 1K ResNet-101 79.77 -
    DeepLabV3+ 1K ResNet-101 80.65 45.47
    NonLocal 1K ResNet-101 79.40 -
    CCNet 1K ResNet-101 79.45 43.71
    MaskFormer 1K ResNet-101 78.50 45.50
    Mask2Former 1K ResNet-101 80.10 47.80
    UperNet 22K Swin-Large 82.89 43.82
    MaskFormer 22K Swin-Large 78.50 -
    Mask2Former 22K Swin-Large 83.30 -
    SegFormer 1K MiT-B5 82.25 50.08
    SETR 22K ViT-Large 78.10 48.28
    Generative modeling:
    UViM 22K Swin-Large 70.77 43.71
    GSS-FF (Ours) 1K ResNet-101 77.76 -
    GSS-FT-W (Ours) 1K ResNet-101 78.46 -
    GSS-FF (Ours) 22K Swin-Large 78.90 46.29
    GSS-FT-W (Ours) 22K Swin-Large 80.05 48.54

    GSS-FT-W with Swin-Large achieves 80.05% mIoU on Cityscapes and 48.54% mIoU on ADE20K, outperforming the previous generative baseline UViM by +9.28% on Cityscapes and +4.83% on ADE20K while delivering performance competitive with discriminative models such as SETR and MaskFormer.

  8. Knowl 8 — Cross-Domain Semantic Segmentation on MSeg Benchmark

    data/table

    Zero-shot cross-domain semantic segmentation performance on the MSeg test split across 6 unseen datasets: Pascal VOC, Pascal Context, CamVid, WildDash, KITTI, and ScanNet. Models are trained on the unified MSeg training split. Evaluation metric is mIoU (%) and harmonic mean (h. mean, %).

    Method Backbone VOC Context CamVid WildDash KITTI ScanNet h. mean
    Discriminative modeling:
    CCSA HRNet-W48 48.9 - 52.4 36.0 - 27.0 39.7
    MGDA HRNet-W48 69.4 - 57.5 39.9 - 33.5 46.1
    MSeg HRNet-W48 (500k) 70.7 42.7 83.3 62.0 67.0 48.2 59.2
    MSeg HRNet-W48 (160k) 63.8 39.6 73.9 60.9 65.1 43.5 54.9
    MSeg Swin-Large (160k) 78.7 47.5 75.1 66.1 68.1 49.0 61.7
    Generative modeling:
    GSS-FF (Ours) HRNet-W48 (160k) 64.1 37.1 72.3 59.3 62.0 40.6 52.6
    GSS-FT-W (Ours) HRNet-W48 (160k) 65.2 38.8 75.2 62.5 66.2 43.1 55.2
    GSS-FF (Ours) Swin-Large (160k) 78.7 45.8 74.2 61.8 65.4 46.9 59.5
    GSS-FT-W (Ours) Swin-Large (160k) 79.5 47.7 75.9 65.3 68.0 49.7 61.9

    GSS-FT-W achieves 55.2% harmonic mean mIoU with HRNet-W48 and 61.9% with Swin-Large under 160k training iterations, surpassing the discriminative MSeg model (54.9% and 61.7% respectively) and establishing a new state of the art in zero-shot cross-domain generalization.

  9. Knowl 9 — Ablation of Stage I Latent Posterior Designs and VQVAE Architectures

    data/table

    Ablation of Stage I reconstruction performance and training efficiency across maskige transformation settings on ADE20K validation split, along with comparison across different VQVAE codebook designs on Cityscapes and ADE20K.

    Variant ADE20K Recon. mIoU (%) Training Time (GPU hours)
    GSS-FF-R 62.83 0
    GSS-FF 84.31 0
    GSS-FT 86.10 ≤20\le 20
    GSS-TF 84.37 ≤5\le 5
    GSS-TT 36.11 ≤5\le 5
    GSS-FT-W 87.73 ≤350\le 350
    Design Maskige? Cityscapes mIoU (%) ADE20K mIoU (%) Train Time (GPU hours)
    VQGAN No 82.16 81.89 ≤500\le 500
    VQGAN Yes 75.09 42.70 ≤100\le 100
    UViM No 89.14 78.98 ≤2,000\le 2,000
    DALL-E Yes 95.17 87.73 ≤350\le 350

    Key observations:

    • Maximal distance initialization (GSS-FF) improves training-free reconstruction mIoU by +21.48% over random initialization (GSS-FF-R).
    • GSS-FT-W achieves the best reconstruction accuracy (87.73% mIoU).
    • Using DALL-E pretrained VQVAE with maskige achieves 95.17% mIoU on Cityscapes and 87.73% on ADE20K, outperforming training from scratch without maskige (UViM at 89.14% and 78.98%) while cutting training time from ≤2000\le 2000 to ≤350\le 350 GPU hours.
  10. Knowl 10 — Ablation of Stage II Latent Prior Components

    data/table

    Ablation study on Stage II latent prior learning evaluated on the ADE20K validation set using a Swin-Large backbone.

    Downsample Ratio dd Unlabeled Area Auxiliary MLA mIoU (%) mAcc (%)
    1/81/8 40.64 52.55
    1/81/8 ✓ 43.72 56.08
    1/41/4 ✓ 43.98 56.11
    1/41/4 ✓ ✓ 46.29 57.84
    • The unlabeled area auxiliary head provides a +3.08% mIoU improvement (40.64%→43.72%40.64\% \to 43.72\%), demonstrating that resolving label ambiguity is critical for generative latent token training.
    • Multi-Level Aggregation (MLA) provides a +2.31% mIoU improvement (43.98%→46.29%43.98\% \to 46.29\%).
    • Increasing the discrete token resolution from downsampling ratio d=1/8d = 1/8 to d=1/4d = 1/4 improves mIoU by +0.26% (43.72%→43.98%43.72\% \to 43.98\%).
  11. Knowl 11 — Domain Generality and Transferability of Maskige

    empirical result

    Because maskige maps class indices to RGB colors independently of input image appearance, it functions as an intrinsically domain-generic representation. When evaluating cross-dataset transfer by transferring models trained on MSeg to the Cityscapes validation split:

    • Transferring only the maskige transformation (while training the image conditioning network IψI_\psi on Cityscapes) achieves 79.5% mIoU with GSS-FT-W, representing only a 1.0% drop compared to the 80.5% mIoU achieved when training maskige directly on Cityscapes.
    • Transferring both the maskige transformation and the image encoder IψI_\psi achieves 78.4% mIoU with GSS-FT-W (a 2.1% drop).

    In contrast, directly transferring image visual representations results in more than double the performance degradation, demonstrating the domain generality of the maskige representation.

Coverage note — None was omitted; all key theoretical formulations, maskige variants, stage I and II training pipelines, inference procedures, and experimental evaluation results were extracted as knowls.

References

  1. 1.Gabriel J Brostow, Julien Fauqueur, and Roberto Cipolla. Semantic object classes in video: A high-definition ground truth database. Pattern Recognition Letters, 2009. 5, 7
  2. 2.Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In CVPR, 2018. 5
  3. 3.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv preprint, 2014. 2, 5
  4. 4.Liang-Chieh Chen, George Papandreou, Florian Schroff, and Hartwig Adam. Rethinking atrous convolution for semantic image segmentation. arXiv preprint, 2017. 1
  5. 5.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018. 6
  6. 6.Ting Chen, Lala Li, Saurabh Saxena, Geoffrey Hinton, and David J Fleet. A generalist framework for panoptic segmentation of images and videos. arXiv preprint, 2022. 2, 4
  7. 7.Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton. Pix2seq: A language modeling framework for object detection. In ICLR, 2021. 2
  8. 8.Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. In CVPR, 2022. 1, 2, 6
  9. 9.Bowen Cheng, Alex Schwing, and Alexander Kirillov. Per-pixel classification is not all you need for semantic segmentation. In NeurIPS, 2021. 1, 2, 6, 7
  10. 10.MMSegmentation Contributors. Openmmlab semantic segmentation toolbox and benchmark. https://github.com/open-mmlab/mmsegmentation, 2020. 5, 7
  11. 11.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016. 5, 6, 7, 8
  12. 12.Angela Dai, Angel X Chang, Manolis Savva, Maciej Halber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 5, 7
  13. 13.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 6
  14. 14.Patrick Esser, Robin Rombach, and Bjorn Ommer. A disentangling invertible interpretation network for explaining latent representations. In CVPR, 2020. 2
  15. 15.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, 2021. 2, 4, 6, 7, 8
  16. 16.Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. The pascal visual object classes (voc) challenge. IJCV, 2010. 5, 7
  17. 17.Jun Fu, Jing Liu, Haijie Tian, Yong Li, Yongjun Bao, Zhiwei Fang, and Hanqing Lu. Dual attention network for scene segmentation. In CVPR, 2019. 6
  18. 18.Andreas Geiger, Philip Lenz, Christoph Stiller, and Raquel Urtasun. Vision meets robotics: The kitti dataset. The International Journal of Robotics Research, 2013. 5, 7
  19. 19.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 2020. 2
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016. 3
  21. 21.Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. Ccnet: Criss-cross attention for semantic segmentation. In ICCV, 2019. 2, 6
  22. 22.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In CVPR, 2017. 2
  23. 23.Ge-Peng Ji, Keren Fu, Zhe Wu, Deng-Ping Fan, Jianbing Shen, and Ling Shao. Full-duplex strategy for video object segmentation. In ICCV, 2021. 1
  24. 24.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint, 2013. 1, 2, 3
  25. 25.Alexander Kolesnikov, André Susano Pinto, Lucas Beyer, Xiaohua Zhai, Jeremiah Harmsen, and Neil Houlsby. Uvim: A unified modeling approach for vision with learned guiding codes. arXiv preprint, 2022. 2, 4, 5, 6, 7, 8
  26. 26.John Lambert, Zhuang Liu, Ozan Sener, James Hays, and Vladlen Koltun. Mseg: A composite dataset for multi-domain semantic segmentation. In CVPR, 2020. 2, 5, 7, 8
  27. 27.Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther. Autoencoding beyond pixels using a learned similarity metric. In ICML, 2016. 2
  28. 28.Daiqing Li, Junlin Yang, Karsten Kreis, Antonio Torralba, and Sanja Fidler. Semantic segmentation with generative models: Semi-supervised learning and strong out-of-domain generalization. In CVPR, 2021. 2
  29. 29.Chen Liang, Wenguan Wang, Jiaxu Miao, and Yi Yang. Gmmseg: Gaussian mixture based generative semantic segmentation models. In NeurIPS, 2022. 2
  30. 30.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In ECCV, 2014. 5
  31. 31.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In ICCV, 2021. 3, 4, 6, 7, 8
  32. 32.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015. 1, 2, 6
  33. 33.Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi, and Aniruddha Kembhavi. Unified-io: A unified model for vision, language, and multi-modal tasks. arXiv preprint, 2022. 2
  34. 34.Chris J Maddison, Andriy Mnih, and Yee Whye Teh. The concrete distribution: A continuous relaxation of discrete random variables. arXiv preprint, 2016. 3, 5
  35. 35.Saeid Motiian, Marco Piccirilli, Donald A Adjeroh, and Gianfranco Doretto. Unified deep supervised domain adaptation and generalization. In ICCV, 2017. 7, 8
  36. 36.Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille. The role of context for object detection and semantic segmentation in the wild. In CVPR, 2014. 5, 7
  37. 37.Kevin P Murphy. Machine learning: a probabilistic perspective. MIT press, 2012. 3, 4
  38. 38.Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo, and Peter Kontschieder. The mapillary vistas dataset for semantic understanding of street scenes. In ICCV, 2017. 5
  39. 39.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021. 1, 2, 4, 5, 6, 7
  40. 40.Robin Rombach, Patrick Esser, and Björn Ommer. Making sense of cnns: Interpreting deep representations and their invariances with inns. In ECCV, 2020. 2
  41. 41.Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. NeurIPS, 2018. 7, 8
  42. 42.Shuran Song, Samuel P Lichtenberg, and Jianxiong Xiao. Sun rgb-d: A rgb-d scene understanding benchmark suite. In CVPR, 2015. 5
  43. 43.Ke Sun, Yang Zhao, Borui Jiang, Tianheng Cheng, Bin Xiao, Dong Liu, Yadong Mu, Xinggang Wang, Wenyu Liu, and Jingdong Wang. High-resolution representations for labeling pixels and regions. arXiv preprint, 2019. 8
  44. 44.Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. NeurIPS, 2017. 1, 2, 3, 4
  45. 45.Girish Varma, Anbumani Subramanian, Anoop Namboodiri, Manmohan Chandraker, and CV Jawahar. Idd: A dataset for exploring problems of autonomous navigation in unconstrained environments. In WACV, 2019. 5
  46. 46.Qiang Wan, Zilong Huang, Jiachen Lu, YU Gang, and Li Zhang. Seaformer: Squeeze-enhanced axial transformer for mobile semantic segmentation. In ICLR. 1
  47. 47.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, 2018. 2, 6
  48. 48.Tete Xiao, Yingcheng Liu, Bolei Zhou, Yuning Jiang, and Jian Sun. Unified perceptual parsing for scene understanding. In ECCV, 2018. 6
  49. 49.Zhisheng Xiao, Qing Yan, and Yali Amit. Generative latent flow. arXiv preprint, 2019. 2
  50. 50.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. In NeurIPS, 2021. 1, 2, 6
  51. 51.Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In CVPR, 2020. 5
  52. 52.Yuhui Yuan, Xilin Chen, and Jingdong Wang. Object-contextual representations for semantic segmentation. In ECCV, 2020. 6
  53. 53.Oliver Zendel, Katrin Honauer, Markus Murschitz, Daniel Steininger, and Gustavo Fernandez Dominguez. Wilddash-creating hazard-aware benchmarks. In ECCV, 2018. 5, 7
  54. 54.Li Zhang, Dan Xu, Anurag Arnab, and Philip HS Torr. Dynamic graph message passing networks. In CVPR, 2020. 2
  55. 55.Hengshuang Zhao, Jianping Shi, Xiaojuan Qi, Xiaogang Wang, and Jiaya Jia. Pyramid scene parsing network. In CVPR, 2017. 2, 6
  56. 56.Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In CVPR, 2021. 1, 2, 6, 7
  57. 57.Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Semantic understanding of scenes through the ade20k dataset. IJCV, 2019. 5, 7, 8

Citation

MLA
Chen, J., et al. “Generative Semantic Segmentation”. arXiv, 2023, http://arxiv.org/abs/2303.11316v2.
APA
Chen, J., Lu, J., Zhu, X., & Zhang, L. (2023). Generative Semantic Segmentation. arXiv. http://arxiv.org/abs/2303.11316v2
Chicago
Chen, J., J. Lu, X. Zhu, and L. Zhang. 2023. “Generative Semantic Segmentation”. arXiv. http://arxiv.org/abs/2303.11316v2.
Harvard
Chen, J. et al. (2023) “Generative Semantic Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.11316v2.
Vancouver
1. Chen J, Lu J, Zhu X, Zhang L (2023) Generative Semantic Segmentation. arXiv

BibTeX

@article{chen2023generative,
  title = {Generative Semantic Segmentation},
  author = {Chen, Jiaqi and Lu, Jiachen and Zhu, Xiatian and Zhang, Li},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.11316v2},
  eprint = {2303.11316}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE