Generative Semantic Segmentation
Jiaqi ChenJiachen LuXiatian ZhuLi Zhang
Proposes Generative Semantic Segmentation to cast semantic segmentation as an image-conditioned mask generation problem using discrete latent priors and mask-as-image representations, delivering state-of-the-art generalization on challenging cross-domain benchmarks.
Semantic segmentation, which involves classifying every individual pixel within an image, is a foundational task for automated visual systems in robotics, autonomous driving, and scene parsing. Conventional approaches treat this strictly as a per-pixel discriminative classification problem, which often struggles when transferring across diverse visual environments and fails to exploit powerful, pre-trained image generation systems.
The article demonstrates that semantic segmentation can be effectively reformulated as an image-conditioned mask generation task using a framework named Generative Semantic Segmentation. The central objective is to evaluate whether casting segmentation as generative modeling can achieve competitive accuracy in standard environments while delivering superior generalization in challenging cross-domain applications.
To overcome the computational bottleneck typically associated with training large generative architectures from scratch, the authors introduce a concept called "maskige," which expresses multi-class segmentation masks in a standard three-channel color image format. This formulation allows the framework to directly utilize off-the-shelf, pre-trained discrete generative autoencoders (specifically VQ-VAE) alongside lightweight transformation modules and standard vision backbones. The approach was evaluated through extensive benchmark experiments on established datasets, including Cityscapes for urban driving, ADE20K for complex scene parsing, and MSeg for cross-domain evaluation.
The empirical findings demonstrate three primary outcomes. First, the generative approach establishes a new state of the art in cross-domain segmentation on the composite MSeg benchmark, outperforming dedicated domain generalization and multi-task baselines with a 61.9 harmonic mean score. Second, in standard single-domain benchmarks, the model achieves competitive accuracy relative to leading discriminative models, recording mean Intersection over Union scores of up to 80.05% on Cityscapes and 48.54% on ADE20K. Third, the maskige design drastically reduces training overhead; the training-free linear variant requires zero first-stage reconstruction training, whereas prior generative approaches required thousands of computing hours.
These results imply that generative modeling offers a robust, domain-agnostic alternative to traditional per-pixel discriminative frameworks without demanding prohibitive computational budgets. By leveraging existing large generative representations, organizations can achieve better out-of-domain reliability and mitigate the performance drops typically encountered when deploying visual models in novel visual domains.
For practical adoption, organizations requiring rapid deployment with minimal compute should implement the training-free linear variant, while applications prioritizing maximum segmentation precision should adopt the Transformer-enhanced non-linear transformation. When handling real-world datasets with missing or incomplete annotations, teams should integrate the proposed auxiliary pseudo-labeling strategy to prevent generative bias toward unlabeled regions.
Confidence in these findings is strong across the evaluated urban, indoor, and composite benchmarks. However, stakeholders should note that the framework's efficiency relies heavily on pre-trained discrete generative codebooks, and applications involving fine-grained non-linear transformations still require moderate training overhead.
- Paper: Learning Structured Output Representation using Deep Conditional Generative Models, Kihyuk Sohn et al. (2015). It establishes the foundational conditional variational autoencoder (CVAE) framework for structured output prediction, which directly underpins formulating semantic segmentation as conditional mask generation.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). It introduces the paradigm shift from traditional per-pixel classification to mask-level modeling that motivates generative approaches to mask prediction.
- Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). It provides the essential principles of conditional generative adversarial networks for image-to-image translation between images and semantic masks.
- Paper: Learning to Adapt Structured Output Space for Semantic Segmentation, Yi-Hsuan Tsai et al. (2018). It demonstrates how aligning structured output spaces across domains enables robust cross-domain semantic segmentation performance.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It establishes the standard discriminative fully convolutional network baseline that Generative Semantic Segmentation seeks to replace.
- Paper: GSVA: Generalized Segmentation via Multimodal Large Language Models, Zhuofan Xia et al. (2024). It extends generalized mask generation and segmentation paradigms by combining multimodal large language models with foundation mask decoders.
