Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation
Zhun ZhongYuyang ZhaoGim Hee LeeNicu Sebe
Proposes a dynamic data augmentation framework that optimizes channel-wise image statistics via adversarial training to synthesize challenging styles from synthetic data alone, substantially boosting cross-domain semantic segmentation performance on unseen real-world driving scenes.
Autonomous systems like self-driving vehicles rely heavily on computer vision models to accurately identify street elements, such as roads, pedestrians, and vehicles. While training these models using synthetic data from video games reduces expensive manual labeling costs, the resulting models often fail in real-world deployment due to domain shifts caused by variations in lighting, weather, and geographic environments. Existing approaches often require real-world target data during training or introduce complex architectures that do not reliably maintain pixel-level semantic consistency.
The article evaluates whether dynamically generating challenging visual styles during synthetic training can enhance a model's robustness and accuracy on unseen real-world domains. The researchers introduce and evaluate Adversarial Style Augmentation (AdvStyle), a model-agnostic technique designed to prevent models from overfitting to synthetic source datasets.
The evaluation used standard synthetic-to-real urban scene segmentation benchmarks, training models on synthetic environments (GTAV and SYNTHIA) and testing them directly across three real-world datasets (CityScapes, BDD-100K, and Mapillary). The approach isolates an image's style into basic statistical features—channel-wise mean and variance—and applies an adversarial update step during training to dynamically construct the most challenging visual style for the model before optimizing segmentation performance. Experiments evaluated multiple standard neural network backbones across both single-source and multi-source generalization configurations, as well as single-domain image classification tasks.
The analysis demonstrates several critical findings. First, AdvStyle substantially improved generalization accuracy across all evaluated neural architectures. For baseline models trained on synthetic GTAV data, the average mean Intersection-over-Union (mIoU) increased from 27.42% to 37.39% for a ResNet-50 backbone, and from 31.47% to 37.34% for a ResNet-101 backbone. Second, when integrated with existing domain generalization architectures, AdvStyle set new state-of-the-art benchmarks, raising ResNet-101 performance on unseen real domains to 42.22% average mIoU without requiring extra real-world training images. Third, AdvStyle outperformed traditional data augmentations, such as color jittering and pixel-level adversarial perturbations, by about 8% in average mIoU. Fourth, the method proved adaptable beyond semantic segmentation, increasing classification accuracy across unseen target domains by up to 8.0% on standard digit and object benchmarks.
These findings suggest that addressing style distribution gaps at the image level provides a practical, low-overhead solution for deploying vision systems into unpredictable real-world environments. By adjusting only a six-dimensional style representation per image during training, organizations can significantly improve safety and operational reliability in autonomous perception systems without collecting expensive target-domain annotations or altering deployment architectures.
Based on these results, machine learning teams should adopt style-level adversarial training in synthetic-data pipelines, especially when deploying models to unseen visual environments. When implementing AdvStyle, teams should maintain the adversarial learning rate within moderate bounds (between 1.0 and 10.0), as excessively large parameters generate unrealistic visual styles that hinder training. Further testing on diverse edge cases and physical hardware environments is recommended before final production deployment.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). This survey provides an essential foundational taxonomy of domain generalization paradigms and benchmarks, contextualizing the need for data-manipulation strategies over unseen environments.
- Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). Reading this overview clarifies the distinctions between representation learning and data-augmentation strategies for domain generalization that motivate AdvStyle's lightweight training technique.
- Paper: Playing for Data: Ground Truth from Computer Games, Stephan R. Richter et al. (2016). This paper introduces the GTA V synthetic urban driving dataset, which forms one of the core source training benchmarks evaluated in the target paper.
- Paper: The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes, Germán Ros et al. (2016). This work establishes the SYNTHIA dataset, providing the second primary synthetic urban scene benchmark used to train and test synthetic-to-real domain generalization.
- Paper: Explaining and Harnessing Adversarial Examples, Ian J. Goodfellow et al. (2015). This work introduces foundational principles of gradient-based adversarial perturbations that AdvStyle adapts from pixel space to channel-wise style representations.
- Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). This work provides critical empirical baselines and standardized evaluation protocols for domain generalization methods, highlighting the importance of robust data-augmentation techniques.
- Paper: AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty, Dan Hendrycks et al. (2019). This study establishes key concepts in data augmentation mixing and consistency loss to improve model robustness across unseen corruptions and domain shifts.
- Paper: Learning to Generalize: Meta-Learning for Domain Generalization, Da Li et al. (2017). This paper presents the meta-learning formulation for domain generalization that AdvStyle contrasts against and integrates with to boost out-of-distribution performance.
- Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). This work extends beyond standard discriminative cross-domain segmentation by reformulating the segmentation task entirely into a generative modeling framework.
