Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation

Zhun ZhongYuyang ZhaoGim Hee LeeNicu Sebe

article2022NeurIPS116 citations

Proposes a dynamic data augmentation framework that optimizes channel-wise image statistics via adversarial training to synthesize challenging styles from synthetic data alone, substantially boosting cross-domain semantic segmentation performance on unseen real-world driving scenes.

Listen

Autonomous systems like self-driving vehicles rely heavily on computer vision models to accurately identify street elements, such as roads, pedestrians, and vehicles. While training these models using synthetic data from video games reduces expensive manual labeling costs, the resulting models often fail in real-world deployment due to domain shifts caused by variations in lighting, weather, and geographic environments. Existing approaches often require real-world target data during training or introduce complex architectures that do not reliably maintain pixel-level semantic consistency.

The article evaluates whether dynamically generating challenging visual styles during synthetic training can enhance a model's robustness and accuracy on unseen real-world domains. The researchers introduce and evaluate Adversarial Style Augmentation (AdvStyle), a model-agnostic technique designed to prevent models from overfitting to synthetic source datasets.

The evaluation used standard synthetic-to-real urban scene segmentation benchmarks, training models on synthetic environments (GTAV and SYNTHIA) and testing them directly across three real-world datasets (CityScapes, BDD-100K, and Mapillary). The approach isolates an image's style into basic statistical features—channel-wise mean and variance—and applies an adversarial update step during training to dynamically construct the most challenging visual style for the model before optimizing segmentation performance. Experiments evaluated multiple standard neural network backbones across both single-source and multi-source generalization configurations, as well as single-domain image classification tasks.

The analysis demonstrates several critical findings. First, AdvStyle substantially improved generalization accuracy across all evaluated neural architectures. For baseline models trained on synthetic GTAV data, the average mean Intersection-over-Union (mIoU) increased from 27.42% to 37.39% for a ResNet-50 backbone, and from 31.47% to 37.34% for a ResNet-101 backbone. Second, when integrated with existing domain generalization architectures, AdvStyle set new state-of-the-art benchmarks, raising ResNet-101 performance on unseen real domains to 42.22% average mIoU without requiring extra real-world training images. Third, AdvStyle outperformed traditional data augmentations, such as color jittering and pixel-level adversarial perturbations, by about 8% in average mIoU. Fourth, the method proved adaptable beyond semantic segmentation, increasing classification accuracy across unseen target domains by up to 8.0% on standard digit and object benchmarks.

These findings suggest that addressing style distribution gaps at the image level provides a practical, low-overhead solution for deploying vision systems into unpredictable real-world environments. By adjusting only a six-dimensional style representation per image during training, organizations can significantly improve safety and operational reliability in autonomous perception systems without collecting expensive target-domain annotations or altering deployment architectures.

Based on these results, machine learning teams should adopt style-level adversarial training in synthetic-data pipelines, especially when deploying models to unseen visual environments. When implementing AdvStyle, teams should maintain the adversarial learning rate within moderate bounds (between 1.0 and 10.0), as excessively large parameters generate unrealistic visual styles that hinder training. Further testing on diverse edge cases and physical hardware environments is recommended before final production deployment.

arXiv: 2207.04892
  • Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). This work extends beyond standard discriminative cross-domain segmentation by reformulating the segmentation task entirely into a generative modeling framework.
Cover for Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation

Abstract

In this paper, we consider the problem of domain generalization in semantic segmentation, which aims to learn a robust model using only labeled synthetic (source) data. The model is expected to perform well on unseen real (target) domains. Our study finds that the image style variation can largely influence the model's performance and the style features can be well represented by the channel-wise mean and standard deviation of images. Inspired by this, we propose a novel adversarial style augmentation (AdvStyle) approach, which can dynamically generate hard stylized images during training and thus can effectively prevent the model from overfitting on the source domain. Specifically, AdvStyle regards the style feature as a learnable parameter and updates it by adversarial training. The learned adversarial style feature is used to construct an adversarial image for robust model training. AdvStyle is easy to implement and can be readily applied to different models. Experiments on two synthetic-to-real semantic segmentation benchmarks demonstrate that AdvStyle can significantly improve the model performance on unseen real domains and show that we can achieve the state of the art. Moreover, AdvStyle can be employed to domain generalized image classification and produces a clear improvement on the considered datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Overview
  • 3.2 Adversarial Style Learning
  • 3.3 Robust Model Training
  • 3.4 Discussion
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Evaluation
  • 4.3 Comparison with State-of-The-Art Methods
  • 4.4 Performance under Multi-Source Setting
  • 4.5 Visualization
  • 4.6 Evaluation on Image Classification Task
  • 5 Conclusion
  • Acknowledgement
  • References
  • Checklist

Knowls

  1. Knowl 1 — Adversarial Style Augmentation Framework

    model/method

    Adversarial Style Augmentation (AdvStyle) is a domain generalization approach for semantic segmentation that dynamically synthesizes hard stylized training samples during model optimization. The approach decomposes each input image into a normalized content representation and an image-level style representation parameterized by channel-wise mean and standard deviation.

    Given an input source image x∈RC×H×Wx \in \mathbb{R}^{C \times H \times W} (where C=3C=3 for RGB channels and H×WH \times W denotes spatial resolution), AdvStyle executes two steps in each training iteration:

    1. Adversarial Style Learning:

      • Compute the channel-wise mean μ∈RC\mu \in \mathbb{R}^C and standard deviation σ∈RC\sigma \in \mathbb{R}^C of xx: μ=1HW∑h=1H∑w=1Wxc,h,w,σ=1HW∑h=1H∑w=1W(xc,h,w−μc)2\mu = \frac{1}{HW} \sum_{h=1}^H \sum_{w=1}^W x_{c,h,w}, \quad \sigma = \sqrt{\frac{1}{HW} \sum_{h=1}^H \sum_{w=1}^W (x_{c,h,w} - \mu_c)^2}
      • Standardize xx to obtain the normalized content image xˉ∈RC×H×W\bar{x} \in \mathbb{R}^{C \times H \times W}: xˉ=x−μσ\bar{x} = \frac{x - \mu}{\sigma}
      • Initialize learnable style parameters μ+←μ\mu^+ \leftarrow \mu and σ+←σ\sigma^+ \leftarrow \sigma, and reconstruct the stylized image x^=xˉ⋅σ++μ+\hat{x} = \bar{x} \cdot \sigma^+ + \mu^+.
      • With model parameters θ\theta fixed, update the 6-dimensional style parameters (μ+,σ+)(\mu^+, \sigma^+) via one step of gradient ascent on the cross-entropy segmentation loss Lseg\mathcal{L}_{\text{seg}}: μ+←μ++γ∇μ+Lseg(θ;x^)\mu^+ \leftarrow \mu^+ + \gamma \nabla_{\mu^+} \mathcal{L}_{\text{seg}}(\theta; \hat{x}) σ+←σ++γ∇σ+Lseg(θ;x^)\sigma^+ \leftarrow \sigma^+ + \gamma \nabla_{\sigma^+} \mathcal{L}_{\text{seg}}(\theta; \hat{x}) where γ>0\gamma > 0 is the adversarial style learning rate.
    2. Robust Model Training:

      • Compose the adversarial sample x+=xˉ⋅σ++μ+x^+ = \bar{x} \cdot \sigma^+ + \mu^+ using the optimized style statistics (μ+,σ+)(\mu^+, \sigma^+) and the normalized image xˉ\bar{x}.
      • Update the segmentation model parameters θ\theta by minimizing the joint loss on the clean original image and the adversarial image: min⁡θ(Lseg(θ;x)+Lseg(θ;x+))\min_\theta \left( \mathcal{L}_{\text{seg}}(\theta; x) + \mathcal{L}_{\text{seg}}(\theta; x^+) \right)

    During inference on unseen target domains, AdvStyle is not applied and testing images are processed directly by the trained network.

  2. Knowl 2 — Properties of Style-Level Versus Pixel-Level Adversarial Augmentation

    model/method

    Traditional adversarial data augmentation for domain generalization (e.g., AdvPixel) generates pixel-wise perturbation tensors δ∈RC×H×W\delta \in \mathbb{R}^{C \times H \times W} added to an image xx. In dense prediction tasks such as semantic segmentation, this approach exhibits two key drawbacks compared to Adversarial Style Augmentation (AdvStyle):

    • Pixel-Wise Semantic Consistency: Perturbing every pixel independently can distort high-frequency semantic cues, edges, and object structures, corrupting the correspondence between input pixels and their ground-truth segmentation masks unless regularized by additional auxiliary constraints. In contrast, AdvStyle optimizes only global channel statistics (μ+,σ+)∈R2C(\mu^+, \sigma^+) \in \mathbb{R}^{2C} to scale and shift the normalized image xˉ\bar{x}. This modifies global contrast, illumination, and color while preserving edge layouts, spatial configurations, and per-pixel ground-truth labels.
    • Optimization Efficiency: Pixel-wise adversarial methods optimize C×H×WC \times H \times W parameters per image, whereas AdvStyle optimizes a 6-dimensional parameter vector per RGB image, reducing gradient computation overhead.
  3. Knowl 3 — Experimental Setup for Synthetic-to-Real Domain Generalized Semantic Segmentation

    experimental setup

    The synthetic-to-real domain generalization benchmark assesses models trained exclusively on synthetic urban-scene datasets and evaluated on real-world datasets without access to target data during training:

    • Source Datasets:
      • GTAV: 24,966 synthetic street-view images (or a standard training split of 12,403 images) with 19 shared semantic categories.
      • SYNTHIA (SYNTHIA-RAND-CITYSCAPES split): 9,400 synthetic images evaluated on 16 shared semantic categories (omitting train, truck, and terrain).
    • Target (Evaluation) Datasets: Validation splits of CityScapes (500 images), BDD-100K (1,000 images), and Mapillary Vistas (2,000 images).
    • Architecture: DeepLabV3+ with backbones MobileNetV2, ResNet-50, or ResNet-101.
    • Optimization Details:
      • Optimizer: Stochastic Gradient Descent (SGD) with initial learning rate 0.010.01, momentum 0.90.9, and weight decay 5×10−45 \times 10^{-4}.
      • Schedule: Polynomial learning rate decay (1−iter/max_iter)0.9(1 - \text{iter}/\text{max\_iter})^{0.9}.
      • Training duration: 40,000 iterations with batch size 16.
      • Default augmentations: Random cropping to 768×768768 \times 768, random horizontal flip, color jittering, and Gaussian blur.
      • AdvStyle hyperparameter: γ=3\gamma = 3.
    • Evaluation Metric: Mean Intersection-over-Union (mIoU, in %) computed on full-resolution target validation sets at the final training iteration (averaged over 3 runs).
  4. Knowl 4 — AdvStyle Performance Across Backbones and Normalization Modules

    data/table

    AdvStyle is model-agnostic and can be applied to standard batch normalization baselines as well as domain-generalization normalization architectures including Instance-Batch Normalization (IBN-Net) and Instance Selective Whitening (ISW).

    The table below reports mIoU (%) on the validation sets of CityScapes (C), BDD-100K (B), and Mapillary (M), along with their mean, for models trained on GTAV (19 classes) or SYNTHIA (16 classes):

    GTAV Source MobileNetV2 ResNet-50 ResNet-101
    Methods C B M Mean C B M Mean C B M Mean
    Baseline 25.92 25.73 26.45 26.03 28.95 25.14 28.18 27.42 32.97 30.77 30.68 31.47
    +AdvStyle 31.81 33.01 31.50 32.11 39.62 35.54 37.00 37.39 39.52 36.39 36.10 37.34
    IBN-Net 30.14 27.66 27.07 28.29 33.85 32.30 37.75 34.63 37.37 34.21 36.81 36.13
    +AdvStyle 32.45 31.55 33.09 32.36 39.32 36.42 40.82 38.85 44.04 39.96 42.67 42.22
    ISW 30.86 30.05 30.67 30.53 36.58 35.20 40.33 37.37 37.20 33.36 35.57 35.38
    +AdvStyle 33.23 31.84 32.00 32.36 39.60 38.59 41.89 40.03 43.44 40.32 41.96 41.91
    SYNTHIA Source ResNet-101 Backbone
    Methods CityScapes BDD-100K Mapillary Mean
    Baseline 34.94 21.96 27.94 28.28
    +AdvStyle 37.59 27.45 31.76 32.27
    IBN-Net 35.83 23.62 28.88 29.44
    +AdvStyle 38.72 28.55 33.59 33.62
    ISW 35.27 23.54 26.72 28.51
    +AdvStyle 39.74 28.33 32.87 33.65

    AdvStyle increases the mean mIoU of the baseline by +6.08%+6.08\%, +9.97%+9.97\%, and +5.87%+5.87\% on MobileNetV2, ResNet-50, and ResNet-101 respectively when trained on GTAV, and improves the performance of IBN-Net and ISW across all configurations.

  5. Knowl 5 — Comparison of AdvStyle with Data Augmentation and Style Manipulation Techniques

    data/table

    AdvStyle was evaluated against standard data augmentation methods (Color Jittering, Gaussian Blur), pixel-level adversarial perturbation (AdvPixel), and alternative image-level style randomization techniques (RandStyle, MixStyle, CrossStyle).

    All experiments use DeepLabV3+ with a ResNet-50 backbone trained on GTAV (19 classes) and evaluated on CityScapes, BDD-100K, and Mapillary validation sets (mIoU in %):

    Comparison with Augmentation Techniques
    Color Jittering Gaussian Blur AdvPixel AdvStyle CityScapes BDD-100K Mapillary Mean
    - - - - 21.64 22.85 24.22 22.91
    ✓ - - - 26.36 23.82 26.33 25.50
    - ✓ - - 25.77 24.05 26.71 25.51
    - - ✓ - 23.34 28.42 30.64 27.46
    - - - ✓ 37.51 33.74 34.73 35.32
    ✓ ✓ - - 28.95 25.14 28.18 27.42
    ✓ ✓ ✓ - 35.42 33.28 33.23 33.97
    ✓ ✓ - ✓ 39.62 35.54 37.00 37.39
    Comparison with Style-Aware Methods
    Method CityScapes BDD-100K Mapillary Mean
    Baseline (with Color Jittering + Gaussian Blur) 28.95 25.14 28.18 27.42
    RandStyle (Random Gaussian noise added to style) 33.40 34.14 31.67 33.07
    MixStyle (Convex combination of sample styles) 35.53 32.41 35.87 34.60
    CrossStyle (Swapping styles between sample pairs) 37.26 32.40 34.09 34.58
    AdvStyle (Ours) 39.62 35.54 37.00 37.39

    AdvStyle alone achieves 35.32%35.32\% mean mIoU without color jittering or Gaussian blur (an improvement of +12.41%+12.41\% over the clean baseline), outperforming AdvPixel (27.46%27.46\%). When combined with standard augmentations, AdvStyle achieves 37.39%37.39\% mean mIoU, surpassing image-level RandStyle (33.07%33.07\%), MixStyle (34.60%34.60\%), and CrossStyle (34.58%34.58\%).

  6. Knowl 6 — Comparison with State-of-the-Art Synthetic-to-Real Domain Generalization Methods

    data/table

    Comparison of AdvStyle against state-of-the-art domain generalization methods for semantic segmentation trained on GTAV and evaluated on CityScapes, BDD-100K, and Mapillary validation sets (mIoU in %).

    In the table below, §\S indicates methods that incorporate extra real-world ImageNet images during training, and ∗* indicates selecting the best checkpoint individually for each target domain (as opposed to evaluating the final training iteration checkpoint):

    Backbone Method CityScapes BDD Mapillary Mean Gain
    ResNet-50 Baseline 28.95 25.14 28.18 27.42 -
    Baseline + AdvStyle 39.62 35.54 37.00 37.39 +9.97
    SW 29.91 27.48 29.71 29.03 +1.61
    IterNorm 31.81 32.70 33.88 32.79 +5.37
    IBN-Net 33.85 32.30 37.75 34.63 +7.21
    IBN-Net + AdvStyle 39.32 36.42 40.82 38.85 +11.43
    ISW 36.58 35.20 40.33 37.37 +9.95
    ISW + AdvStyle 39.60 38.59 41.89 40.03 +12.61
    ResNet-101 Baseline 32.97 30.77 30.68 31.47 -
    Baseline + AdvStyle 39.52 36.39 36.10 37.34 +5.87
    IBN-Net 37.37 34.21 36.81 36.13 +4.66
    IBN-Net + AdvStyle 44.04 39.96 42.67 42.22 +10.75
    ISW 37.20 33.36 35.57 35.38 +3.91
    ISW + AdvStyle 43.44 40.32 41.96 41.91 +10.44
    DRPC§∗^{\S*} 42.53 38.72 38.05 39.76 +9.88
    FSDR§∗^{\S*} 44.80 41.20 43.40 43.13 +13.60
    ISW + AdvStyle∗^* 45.62 41.71 46.69 44.67 +13.20

    Under the standard protocol (final checkpoint evaluation without ImageNet data), IBN-Net + AdvStyle with ResNet-101 achieves 42.22%42.22\% mean mIoU. Under the best-checkpoint evaluation protocol (∗^*), ISW + AdvStyle achieves 44.67%44.67\% mean mIoU, outperforming FSDR (43.13%43.13\%).

  7. Knowl 7 — Multi-Source Domain Generalization for Semantic Segmentation

    empirical result

    When trained simultaneously on two synthetic source datasets (GTAV + SYNTHIA) using DeepLabV3+ with a ResNet-50 backbone and evaluated on real-world validation sets (CityScapes, BDD-100K, Mapillary), the methods achieve the following mIoU (%):

    • Baseline: CityScapes: 35.46%35.46\%, BDD-100K: 25.09%25.09\%, Mapillary: 31.94%31.94\%, Mean: 30.83%30.83\%.
    • IBN-Net: CityScapes: 35.55%35.55\%, BDD-100K: 32.18%32.18\%, Mapillary: 38.09%38.09\%, Mean: 35.27%35.27\%.
    • ISW: CityScapes: 37.69%37.69\%, BDD-100K: 34.09%34.09\%, Mapillary: 38.49%38.49\%, Mean: 36.75%36.75\%.
    • ISW + AdvStyle: CityScapes: 39.29%39.29\%, BDD-100K: 39.26%39.26\%, Mapillary: 41.14%41.14\%, Mean: 39.90%39.90\%.

    Adding AdvStyle to ISW yields an absolute gain of +3.15%+3.15\% in mean mIoU (+1.60%+1.60\% on CityScapes, +5.17%+5.17\% on BDD-100K, and +2.65%+2.65\% on Mapillary).

  8. Knowl 8 — Performance on Single Domain Generalization Image Classification Benchmarks

    data/table

    AdvStyle was evaluated on single-source domain generalization for image classification across the Digits and PACS benchmarks.

    Digits Benchmark (Source: MNIST; Targets: SVHN, MNIST-M, SYN, USPS; classification accuracy in %):

    • ERM: SVHN 27.8%27.8\%, MNIST-M 52.7%52.7\%, SYN 39.7%39.7\%, USPS 76.9%76.9\%, Average: 49.3%49.3\%.
    • CCSA: Average 49.1%49.1\%.
    • d-SNE: Average 52.1%52.1\%.
    • JiGen: Average 53.1%53.1\%.
    • ADA: SVHN 35.5%35.5\%, MNIST-M 60.4%60.4\%, SYN 45.3%45.3\%, USPS 77.3%77.3\%, Average: 54.6%54.6\%.
    • M-ADA: Average 59.5%59.5\%.
    • ME-ADA: SVHN 42.6%42.6\%, MNIST-M 63.3%63.3\%, SYN 50.4%50.4\%, USPS 81.0%81.0\%, Average: 59.3%59.3\%.
    • ERM + AdvStyle: SVHN 50.4%50.4\%, MNIST-M 73.4%73.4\%, SYN 58.7%58.7\%, USPS 81.6%81.6\%, Average: 66.0%66.0\%.
    • ME-ADA + AdvStyle: SVHN 55.5%55.5\%, MNIST-M 74.1%74.1\%, SYN 59.3%59.3\%, USPS 80.1%80.1\%, Average: 67.3%67.3\%.

    PACS Benchmark (Leave-one-domain-in setting: trained on one domain listed in the column and tested on the remaining three; accuracy in %):

    Method Art Painting Cartoon Sketch Photo Average
    ERM 67.4 74.4 51.4 42.6 58.9
    JiGen 69.1 74.6 52.4 41.5 59.4
    RSC 68.8 74.5 53.6 41.9 59.7
    L2D 74.3 77.5 54.4 45.9 63.0
    ERM + AdvStyle 75.8 76.6 58.1 51.1 65.4
    RSC + AdvStyle 75.1 78.0 58.9 55.5 66.8
    L2D + AdvStyle 80.6 78.4 58.3 59.7 69.3

    AdvStyle improves the ERM baseline average accuracy by +16.7%+16.7\% on Digits and +6.5%+6.5\% on PACS, and provides gains when combined with state-of-the-art DG classification methods (ME-ADA, RSC, L2D).

  9. Knowl 9 — Hyperparameter Sensitivity to Adversarial Style Learning Rate

    empirical result

    The adversarial style learning rate γ\gamma determines the magnitude of the gradient ascent update applied to the style statistics (μ+,σ+)(\mu^+, \sigma^+) in Adversarial Style Augmentation. An empirical evaluation varying γ∈[0.1,40]\gamma \in [0.1, 40] with a ResNet-50 backbone trained on GTAV demonstrates the following behavior across CityScapes, BDD-100K, and Mapillary:

    • The model achieves improvements over the baseline (27.42%27.42\% mean mIoU) across the full range γ∈[0.1,40]\gamma \in [0.1, 40].
    • Small learning rates in the range γ∈[0.1,0.5]\gamma \in [0.1, 0.5] increase mean mIoU from 27.42%27.42\% to above 33%33\%.
    • The optimal performance band is observed between γ=1.0\gamma = 1.0 and γ=10.0\gamma = 10.0, peaking around γ=3\gamma = 3 (37.39%37.39\% mean mIoU).
    • At very large values such as γ=40\gamma = 40, the adversarial perturbation introduces extreme style distortions that reduce performance relative to the peak, though accuracy remains higher than the unaugmented baseline.

Coverage note — All core contributions from the paper—the AdvStyle formulation, comparative analyses against alternative augmentations, segmentation benchmarks under single- and multi-source setups, hyperparameter sensitivity, and image classification evaluations—are included; qualitative visual figures (t-SNE distributions and segmentation mask visualizations) were omitted as they illustrate the quantitative results.

References

  1. 1.Vijay Badrinarayanan, Alex Kendall, and Roberto Cipolla. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. TPAMI, 2017.
  2. 2.Anand Bhattad, Min Jin Chong, Kaizhao Liang, Bo Li, and DA Forsyth. Unrestricted adversarial examples via semantic manipulation. In ICLR, 2020.
  3. 3.Fabio M Carlucci, Antonio D’Innocente, Silvia Bucci, Barbara Caputo, and Tatiana Tommasi. Domain generalization by solving jigsaw puzzles. In CVPR, 2019.
  4. 4.Chen Chen, Zeju Li, Cheng Ouyang, Matt Sinclair, Wenjia Bai, and Daniel Rueckert. Maxstyle: Adversarial style composition for robust medical image segmentation. In MICCAI, 2022.
  5. 5.Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. TPAMI, 2018.
  6. 6.Liang-Chieh Chen, Yukun Zhu, George Papandreou, Florian Schroff, and Hartwig Adam. Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, 2018.
  7. 7.Sungha Choi, Sanghun Jung, Huiwon Yun, Joanne T Kim, Seungryong Kim, and Jaegul Choo. Robustnet: Improving domain generalization in urban-scene segmentation via instance selective whitening. In CVPR, 2021.
  8. 8.Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. The cityscapes dataset for semantic urban scene understanding. In CVPR, 2016.
  9. 9.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009.
  10. 10.Vincent Dumoulin, Jonathon Shlens, and Manjunath Kudlur. A learned representation for artistic style. In ICLR, 2017.
  11. 11.Xinjie Fan, Qifei Wang, Junjie Ke, Feng Yang, Boqing Gong, and Mingyuan Zhou. Adversarially adaptive normalization for single domain generalization. In CVPR, 2021.
  12. 12.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. ICLR, 2015.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  14. 14.Jiaxing Huang, Dayan Guan, Aoran Xiao, and Shijian Lu. Fsdr: Frequency space domain randomization for domain generalization. In CVPR, 2021.
  15. 15.Lei Huang, Yi Zhou, Fan Zhu, Li Liu, and Ling Shao. Iterative normalization: Beyond standardization towards efficient whitening. In CVPR, 2019.
  16. 16.Xun Huang and Serge Belongie. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV, 2017.
  17. 17.Zeyi Huang, Haohan Wang, Eric P. Xing, and Dong Huang. Self-challenging improves cross-domain generalization. In ECCV, 2020.
  18. 18.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning (ICML), 2015.
  19. 19.Wei Liu, Andrew Rabinovich, and Alexander C Berg. Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579, 2015.
  20. 20.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, 2015.
  21. 21.Yawei Luo, Liang Zheng, Tao Guan, Junqing Yu, and Yi Yang. Taking a closer look at domain shift: Category-level adversaries for semantics consistent domain adaptation. In CVPR, 2019.
  22. 22.Saeid Motiian, Marco Piccirilli, Donald A Adjeroh, and Gianfranco Doretto. Unified deep supervised domain adaptation and generalization. In ICCV, 2017.
  23. 23.Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo, and Peter Kontschieder. The mapillary vistas dataset for semantic understanding of street scenes. In ICCV, 2017.
  24. 24.Xingang Pan, Ping Luo, Jianping Shi, and Xiaoou Tang. Two at once: Enhancing learning and generalization capacities via ibn-net. In ECCV, 2018.
  25. 25.Xingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang, and Ping Luo. Switchable whitening for deep representation learning. In ICCV, 2019.
  26. 26.Fengchun Qiao and Xi Peng. Uncertainty-guided model generalization to unseen domains. In CVPR, 2021.
  27. 27.Fengchun Qiao, Long Zhao, and Xi Peng. Learning to learn single domain generalization. In CVPR, 2020.
  28. 28.Stephan R. Richter, Vibhav Vineet, Stefan Roth, and Vladlen Koltun. Playing for data: Ground truth from computer games. In ECCV, 2016.
  29. 29.German Ros, Laura Sellart, Joanna Materzynska, David Vazquez, and Antonio M Lopez. The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In CVPR, 2016.
  30. 30.Nataniel Ruiz, Samuel Schulter, and Manmohan Chandraker. Learning to simulate. In ICLR, 2019.
  31. 31.Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In CVPR, 2018.
  32. 32.Manli Shu, Zuxuan Wu, Micah Goldblum, and Tom Goldstein. Encoding robustness to image style via adversarial feature perturbations. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, NeurIPS, 2021.
  33. 33.Zhiqiang Tang, Yunhe Gao, Yi Zhu, Zhi Zhang, Mu Li, and Dimitris Metaxas. Selfnorm and crossnorm for out-of-distribution robustness. ICCV, 2021.
  34. 34.Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022, 2016.
  35. 35.Vladimir Vapnik. The nature of statistical learning theory. Springer science & business media, 2013.
  36. 36.Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese. Generalizing to unseen domains via adversarial data augmentation. In NeurIPS, 2018.
  37. 37.Zijian Wang, Yadan Luo, Ruihong Qiu, Zi Huang, and Mahsa Baktashmotlagh. Learning to diversify for single domain generalization. In ICCV, 2021.
  38. 38.Xiang Xu, Xiong Zhou, Ragav Venkatesan, Gurumurthy Swaminathan, and Orchid Majumder. d-sne: Domain adaptation using stochastic neighborhood embedding. In CVPR, 2019.
  39. 39.Fisher Yu, Haofeng Chen, Xin Wang, Wenqi Xian, Yingying Chen, Fangchen Liu, Vashisht Madhavan, and Trevor Darrell. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In CVPR, 2020.
  40. 40.Xiangyu Yue, Yang Zhang, Sicheng Zhao, Alberto Sangiovanni-Vincentelli, Kurt Keutzer, and Boqing Gong. Domain randomization and pyramid consistency: Simulation-to-real generalization without accessing target domain data. In ICCV, 2019.
  41. 41.Pan Zhang, Bo Zhang, Ting Zhang, Dong Chen, Yong Wang, and Fang Wen. Prototypical pseudo label denoising and target structure learning for domain adaptive semantic segmentation. In CVPR, 2021.
  42. 42.Long Zhao, Ting Liu, Xi Peng, and Dimitris Metaxas. Maximum-entropy adversarial data augmentation for improved generalization and robustness. NeurIPS, 2020.
  43. 43.Yuyang Zhao, Zhun Zhong, Fengxiang Yang, Zhiming Luo, Yaojin Lin, Shaozi Li, and Sebe Nicu. Learning to generalize unseen domains via memory-based multi-source meta-learning for person re-identification. In CVPR, 2021.
  44. 44.Yuyang Zhao, Zhun Zhong, Zhiming Luo, Gim Hee Lee, and Nicu Sebe. Source-free open compound domain adaptation in semantic segmentation. TCSVT, 2022.
  45. 45.Yuyang Zhao, Zhun Zhong, Nicu Sebe, and Gim Hee Lee. Novel class discovery in semantic segmentation. In CVPR, 2022.
  46. 46.Yuyang Zhao, Zhun Zhong, Na Zhao, Nicu Sebe, and Gim Hee Lee. Style-hallucinated dual consistency learning for domain generalized semantic segmentation. In ECCV, 2022.
  47. 47.Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang. Domain generalization with mixstyle. In ICLR, 2021.

Citation

MLA
Zhong, Z., et al. “Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 338–50, https://proceedings.neurips.cc/paper_files/paper/2022/file/023d94f44110b9a3c62329beec739772-Paper-Conference.pdf.
APA
Zhong, Z., Zhao, Y., Lee, G. H., & Sebe, N. (2022). Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation. Advances in Neural Information Processing Systems, 35, 338–350. https://proceedings.neurips.cc/paper_files/paper/2022/file/023d94f44110b9a3c62329beec739772-Paper-Conference.pdf
Chicago
Zhong, Z., Y. Zhao, G. H. Lee, and N. Sebe. 2022. “Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation”. Advances in Neural Information Processing Systems 35: 338–50. https://proceedings.neurips.cc/paper_files/paper/2022/file/023d94f44110b9a3c62329beec739772-Paper-Conference.pdf.
Harvard
Zhong, Z. et al. (2022) “Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 338–350. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/023d94f44110b9a3c62329beec739772-Paper-Conference.pdf.
Vancouver
1. Zhong Z, Zhao Y, Lee GH, Sebe N (2022) Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 338–350

BibTeX

@inproceedings{zhong2022adversarial,
  title = {Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation},
  author = {Zhong, Zhun and Zhao, Yuyang and Lee, Gim Hee and Sebe, Nicu},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {338-350},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/023d94f44110b9a3c62329beec739772-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors