Unsupervised Domain Adaptation for Semantic Segmentation via Class-Balanced Self-training
Yang ZouZhiding YuB. V. K. Vijaya KumarJinsong Wang
Proposes a class-balanced self-training framework for unsupervised domain adaptation in semantic segmentation that prevents dominant categories from biasing pseudo-label generation while using spatial priors to refine target predictions.
Deep neural networks have achieved remarkable accuracy in semantic segmentation, which assigns a specific category label to every pixel in an image. However, deploying these models in real-world systems such as autonomous driving remains difficult because models trained in one environment often fail when exposed to new conditions, such as different cities, lighting, or weather. While training models on synthetic, computer-generated data offers an inexpensive alternative to manual image annotation, a large visual domain gap persists between synthetic simulations and real-world environments. The article addresses this challenge by studying unsupervised domain adaptation, which aims to adapt models trained on labeled source data to entirely unlabeled target environments.
The main objective of the article is to demonstrate an effective, self-training framework that aligns synthetic and real data representations without relying on complex adversarial training techniques, while resolving class-imbalance issues in generated pseudo-labels.
To achieve this, the authors develop an iterative self-training approach formulated as a unified loss minimization problem. The algorithm alternately generates high-confidence pseudo-labels for unlabeled target images and updates the neural network using both the original source data and these pseudo-labeled target samples. Because standard self-training naturally favors dominant and easily transferred classes (such as roads or sky) over rarer or harder categories (such as riders, traffic signs, and bikes), the authors introduce Class-Balanced Self-Training (CBST), which normalizes prediction confidence scores independently for each class. Additionally, they incorporate spatial scene priors derived from geometric layout consistencies in driving scenes to further guide label selection. The framework was evaluated across major synthetic-to-real benchmarks (GTA5 to Cityscapes and SYNTHIA to Cityscapes) and cross-city transfers (Cityscapes to NTHU).
The evaluation yielded several key findings. First, self-training combined with modern deep architectures matches or surpasses prevailing adversarial distribution-matching methods without requiring separate discriminator networks. Second, class-balanced selection significantly boosts accuracy on underrepresented and difficult categories; on the SYNTHIA-to-Cityscapes benchmark, CBST improved overall mean Intersection-over-Union (mIoU) to 42.5% (48.4% on a standard 13-class subset), outperforming competing approaches. Third, on the GTA5-to-Cityscapes adaptation task, combining class balancing with spatial priors achieved an mIoU of 46.2%, and 47.0% with multi-scale testing, establishing new state-of-the-art benchmark results. Fourth, the approach demonstrated strong transferability in real-world cross-city testing across Rome, Rio, Tokyo, and Taipei.
These findings indicate that organizations developing vision systems for robotics and autonomous driving can substantially lower the cost and turnaround time of manual data labeling by relying more heavily on synthetic data. By unifying feature alignment and task training into a single optimization process, engineering teams can also simplify model training pipelines and eliminate the training instabilities typical of adversarial architectures.
Decision-makers should consider adopting class-balanced self-training workflows when expanding perception models into new operating domains or geographic markets. For immediate implementation, teams should leverage spatial priors when camera viewpoints and scene geometry remain relatively consistent across environments. Prior to broad deployment, further validation is recommended to evaluate performance under severe environmental variations, such as nighttime conditions or extreme weather, where structural and visual assumptions may diverge significantly.
- Paper: Learning to Adapt Structured Output Space for Semantic Segmentation, Yi-Hsuan Tsai et al. (2018). It introduces output-space adversarial learning for unsupervised domain adaptation in semantic segmentation, establishing the foundational problem setup and benchmarks adapted by CBST.
- Paper: CyCADA: Cycle-Consistent Adversarial Domain Adaptation, Judy Hoffman et al. (2018). It establishes key synthetic-to-real semantic segmentation adaptation pipelines (GTA5 and SYNTHIA to Cityscapes) that motivate the need for alternative self-training paradigms.
- Paper: DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs, Liang-Chieh Chen et al. (2016). It provides the underlying DeepLab semantic segmentation architecture with atrous convolution and spatial priors that serve as the backbone for dense target-domain prediction.
- Book: Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al. (2016). It introduces the core principles of domain-adversarial training that framed standard unsupervised domain adaptation before class-balanced self-training methods emerged.
- Paper: The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes, Germán Ros et al. (2016). It provides the SYNTHIA synthetic urban scene dataset, one of the primary source benchmarks used to evaluate domain adaptation to real-world target environments.
- Paper: Maximum Classifier Discrepancy for Unsupervised Domain Adaptation, Kuniaki Saito et al. (2017). It presents a task-specific boundary alignment method for unsupervised domain adaptation that serves as an essential baseline and conceptual contrast to pseudo-labeling.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). It introduces end-to-end fully convolutional networks, defining the core pixel-level classification paradigm used in dense urban scene segmentation.
- Paper: ADVENT: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation, Tuan-Hung Vu et al. (2018). It builds directly on unsupervised domain adaptation for segmentation by minimizing predictive entropy and uncertainty on target data, offering a complementary alternative to iterative pseudo-labeling.
- Paper: Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain Adaptation, Jian Liang et al. (2020). It advances target self-training and pseudo-labeling mechanisms by eliminating the requirement of accessing source data entirely during domain adaptation.
- Paper: Self-Training With Noisy Student Improves ImageNet Classification, Qizhe Xie et al. (2019). It scales up iterative self-training and pseudo-labeling paradigms using teacher-student noise injection, extending the foundational principles of self-training beyond domain-specific shifts.
- Paper: Unsupervised Data Augmentation for Consistency Training, Qizhe Xie et al. (2020). It extends self-supervised consistency and pseudo-labeling frameworks by integrating advanced data augmentations for robust semi-supervised learning.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). It surveys methods for generalizing models across domain shifts to completely unseen environments without requiring access to unlabeled target data.
- Paper: Per-Pixel Classification is Not All You Need for Semantic Segmentation, Bowen Cheng et al. (2021). It rethinks the per-pixel classification paradigm underlying segmentation adaptation by reformulating semantic segmentation as a unified mask classification problem.
