DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation

Qi XuLiang YaoZhengkai JiangGuannan JiangWenqing ChuWenhui HanWei ZhangChengjie WangYing Tai

article2022AAAI97 citations

Proposes a domain generalization framework for semantic segmentation that quantifies channel-level sensitivity to style shifts, using this prior knowledge to re-calibrate feature responses and selectively whiten domain-sensitive feature correlations.

Listen

Deep learning models for visual semantic segmentation achieve strong results when evaluated on familiar environments, but their accuracy degrades sharply when deployed in new, unseen environments. In safety-critical fields such as autonomous driving, training models on every conceivable real-world scenario—such as adverse weather, lighting shifts, and varied geographic locations—is practically impossible. Most existing approaches to this challenge overlook an essential property of internal representations: certain visual features are inherently sensitive to superficial environmental style, while others reliably capture underlying domain-invariant content.

The article develops and evaluates Domain-Invariant Representation Learning, a framework designed to enhance single-source domain generalization in semantic segmentation. The primary objective is to measure feature sensitivity to domain styles and use that sensitivity as a guide to both suppress style-dependent features and selectively eliminate sensitive feature correlations.

The approach introduces three integrated modules without requiring target-domain data during training. First, a Sensitivity-aware Prior Module estimates channel sensitivity by calculating feature differences between original images and style-altered versions generated via photometric transformations. Second, a Prior Guided Attention Module trains channel-wise attention weights to penalize sensitive features and reward robust ones. Third, Guided Feature Whitening decouples feature covariances using the sensitivity prior and removes the correlations most vulnerable to domain shifts. The framework was evaluated across standard benchmarks transferring from synthetic datasets (GTAV, Synthia) to diverse real-world urban datasets (Cityscapes, BDD, Mapillary) using multiple deep backbones, including ResNet-50, ShuffleNet, and MobileNet.

The findings show substantial and consistent performance gains across all evaluated settings. When trained on synthetic GTAV data and tested across Cityscapes, BDD, and Mapillary, the ResNet-50 model achieved a mean intersection-over-union score of 40.60%, outperforming the standard baseline at 27.42% and the strongest prior selective whitening method at 37.37%. Similar consistent improvements occurred on lightweight architectures, raising ShuffleNet performance from 25.44% to 33.52% and MobileNet from 26.03% to 33.92%. Models trained on clean urban datasets maintained reasonable segmentation predictions when exposed to severe unseen conditions like night driving and heavy rain. Crucially, the computational overhead remained minimal, increasing ResNet-50 inference time by only 0.22 milliseconds per frame (from 10.71 ms to 10.93 ms on an NVIDIA V100 GPU).

These results demonstrate that explicitly quantifying feature sensitivity enables networks to decouple style from content more cleanly than previous statistical whitening methods. By delivering superior generalization without extra data collection or noticeable latency costs, this technique lowers operational risk and improves perception safety in unpredictable real-world operating conditions.

Engineering and research teams deploying computer vision in autonomous driving or robotics should integrate sensitivity-guided attention and whitening into early convolutional layers of existing architectures. Future development should focus on expanding validation across a wider variety of outdoor operational domains and testing the framework alongside broader multi-sensor perception stacks.

A primary limitation is that sensitivity modeling relies on simulated style transformations such as color jitter and blurring, which may not capture all real-world environmental variations like structural or sensor-level anomalies. Nonetheless, the consistent empirical gains across multiple datasets and network backbones provide high confidence in the framework's effectiveness for single-source domain generalization.

Cover for DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation

Abstract

Model generalization to the unseen scenes is crucial to real-world applications, such as autonomous driving, which requires robust vision systems. To enhance the model generalization, domain generalization through learning the domain-invariant representation has been widely studied. However, most existing works learn the shared feature space within multi-source domains but ignore the characteristic of the feature itself (e.g., the feature sensitivity to the domain-specific style). Therefore, we propose the Domain-invariant Representation Learning (DIRL) for domain generalization which utilizes the feature sensitivity as the feature prior to guide the enhancement of the model generalization capability. The guidance reflects in two folds: 1) Feature re-calibration that introduces the Prior Guided Attention Module (PGAM) to emphasize the insensitive features and suppress the sensitive features. 2) Feature whiting that proposes the Guided Feature Whiting (GFW) to remove the feature correlations which are sensitive to the domain-specific style. We construct the domain-invariant representation which suppresses the effect of the domain-specific style on the quality and correlation of the features. As a result, our method is simple yet effective, and can enhance the robustness of various backbone networks with little computational cost. Extensive experiments over multiple domains generalizable segmentation tasks show the superiority of our approach to other methods.

Table of Contents

  • Introduction
  • Related Works
  • Domain Generalization
  • Domain Adaptation for Semantic Segmentation
  • Model Interpretability
  • Method
  • Sensitivity-aware Prior Module (SAPM)
  • Prior Guided Attention Module (PGAM)
  • Guided Feature Whiting (GFW)
  • Network Architecture in DIRL
  • Experiments
  • Model and Dateset
  • Training Details
  • Ablation Experiments
  • Comparisons with State-of-the-Art Methods
  • Qualitative Analysis
  • Conclusion
  • References

Knowls

  1. Knowl 1 — Domain-Invariant Representation Learning (DIRL) Framework

    model/method

    Domain-Invariant Representation Learning (DIRL) is a single-source domain generalization framework for semantic segmentation that uses feature sensitivity to domain-specific style shifts as an explicit prior to guide feature learning. In deep neural networks, different feature channels exhibit varying levels of sensitivity to domain shifts: channels corresponding to texture and color are typically domain-sensitive, whereas channels corresponding to high-level semantic objects and parts are domain-insensitive.

    DIRL leverages this observation via three main steps:

    1. Sensitivity Quantification: A Sensitivity-aware Prior Module (SAPM) computes a sensitive guiding vector ss by comparing intermediate feature activations on an original image versus a photo-metrically perturbed version.
    2. Feature Re-calibration: A Prior Guided Attention Module (PGAM) applies channel-wise attention weights learned under a Sensitivity Guidance loss (LSGL_{SG}), suppressing style-sensitive channels while emphasizing domain-invariant channels.
    3. Guided Covariance Whitening: A Guided Feature Whitening (GFW) module computes a Sensitive Prior Map to identify and selectively zero out the covariance terms between style-sensitive feature dimensions via a loss LGFWL_{GFW}.

    These components are inserted after the first three convolution groups of the backbone encoder to eliminate style sensitivity in early feature representations.

  2. Knowl 2 — Sensitivity-Aware Prior Module (SAPM)

    model/method

    The Sensitivity-aware Prior Module (SAPM) calculates a channel-wise style sensitivity vector s∈[0,1]C×1×1s \in [0, 1]^{C \times 1 \times 1} without requiring any trainable parameters.

    Given an input image, a style-shifted variant is generated using photo-metric transformations (color jittering and Gaussian blurring). Let Mb,Mb′∈R1×C×H×WM_b, M'_b \in \mathbb{R}^{1 \times C \times H \times W} denote the intermediate feature representations of the original and transformed images at a given layer, where CC is the number of channels and H×WH \times W is the spatial resolution.

    The feature difference vector d∈RC×1×1d \in \mathbb{R}^{C \times 1 \times 1} is computed by spatial Euclidean distance and global average pooling (GAP): d=GAP(L2(Mb−Mb′))d = \mathrm{GAP}\left(L_2(M_b - M'_b)\right)

    The channel difference vector dd is then min-max normalized across channels to form the sensitive guiding vector ss: s=d−min⁡(d)max⁡(d)−min⁡(d)∈[0,1]C×1×1s = \frac{d - \min(d)}{\max(d) - \min(d)} \in [0, 1]^{C \times 1 \times 1} where min⁡(⋅)\min(\cdot) and max⁡(⋅)\max(\cdot) operate over the channel dimension. Each element sjs_j acts as a guiding factor for channel jj, where higher values indicate greater sensitivity to domain style variations.

  3. Knowl 3 — Prior Guided Attention Module (PGAM) and Sensitivity Guidance Loss

    model/method

    The Prior Guided Attention Module (PGAM) dynamically recalibrates channel activations in the feature map M∈RC×H×WM \in \mathbb{R}^{C \times H \times W} using learned channel-wise attention weights w∈[0,1]C×1×1w \in [0, 1]^{C \times 1 \times 1}.

    The attention vector ww is generated by applying global average pooling, a 1×11 \times 1 convolution, and a sigmoid activation σ\sigma: w=σ(Conv1×1(GAP(M)))w = \sigma\left(\mathrm{Conv}_{1\times 1}(\mathrm{GAP}(M))\right)

    To enforce that the learned attention weights inversely reflect style sensitivity during training, the network is supervised by the Sensitivity Guidance loss LSGL_{SG}: LSG=∥log⁡(w)⊙log⁡(s)−1∥2L_{SG} = \left\| \log(w) \odot \log(s) - 1 \right\|_2 where s∈[0,1]C×1×1s \in [0, 1]^{C \times 1 \times 1} is the sensitive guiding vector from the Sensitivity-aware Prior Module (SAPM). This loss penalizes configurations where a style-sensitive channel (sj→1s_j \to 1) receives high attention (wj→1w_j \to 1), constraining wj→0w_j \to 0 when sj→1s_j \to 1.

    During inference, PGAM operates with only a single image input, bypassing the two-image forward pass of SAPM while maintaining channel recalibration and modeling channel interdependencies.

  4. Knowl 4 — Guided Feature Whitening (GFW) and Covariance Suppression Loss

    model/method

    Guided Feature Whitening (GFW) selectively eliminates style-sensitive feature correlations while preserving domain-invariant semantic correlations.

    GFW consists of three sequential operations:

    1. Covariance Matrix Estimation: An input feature map M∈RC×H×WM \in \mathbb{R}^{C \times H \times W} is normalized via instance normalization (without affine parameters) to yield standardized features Ms∈RC×H×WM_s \in \mathbb{R}^{C \times H \times W}. The sample covariance matrix Σs∈RC×C\Sigma_s \in \mathbb{R}^{C \times C} is computed as: Σs=1HW(Ms)(Ms)T\Sigma_s = \frac{1}{HW}(M_s)(M_s)^T

    2. Sensitive Prior Map and Selective Masking: Using the sensitive guiding vector s∈[0,1]C×1s \in [0, 1]^{C \times 1} (flattened from SAPM), a Sensitive Prior Map (SPM) is constructed: SPM=ssT∈RC×C\mathrm{SPM} = s s^T \in \mathbb{R}^{C \times C} A binary Selective Mask SM∈{0,1}C×C\mathrm{SM} \in \{0, 1\}^{C \times C} is defined over the strictly upper-triangular elements of Σs\Sigma_s. The mask assigns SMj,k=1\mathrm{SM}_{j, k} = 1 to the top α\alpha proportion of largest entries in SPM\mathrm{SPM} (with default selective ratio α=0.3\alpha = 0.3) and 00 elsewhere, identifying the most style-sensitive channel pairs.

    3. Guided Feature Whitening Loss: The correlations identified by SM\mathrm{SM} are suppressed via the loss: LGFW=E[∥Σs⊙SM∥]L_{GFW} = \mathbb{E}\left[ \|\Sigma_s \odot \mathrm{SM}\| \right] where ⊙\odot denotes the element-wise Hadamard product and E[⋅]\mathbb{E}[\cdot] denotes the arithmetic mean over the selected non-zero elements in SM\mathrm{SM}.

  5. Knowl 5 — Network Architecture and Combined Objective in DIRL

    model/method

    In DIRL, instance normalization and Prior Guided Attention Modules (PGAM) are inserted after each of the first N=3N=3 convolutional groups of the segmentation backbone encoder (e.g., ResNet-50, ShuffleNetV2, MobileNetV2 with DeepLabV3+).

    The overall optimization objective is given by: Ltotal=Lseg+λ1(1N∑i=1NLSGi)+λ2(1N∑i=1NLGFWi)L_{total} = L_{seg} + \lambda_1 \left(\frac{1}{N}\sum_{i=1}^N L_{SG}^i\right) + \lambda_2 \left(\frac{1}{N}\sum_{i=1}^N L_{GFW}^i\right) where:

    • LSGiL_{SG}^i is the Sensitivity Guidance loss computed at layer ii.
    • LGFWiL_{GFW}^i is the Guided Feature Whitening loss computed on the instance-normalized covariance matrix at layer ii.
    • λ1\lambda_1 and λ2\lambda_2 are loss balancing hyperparameters set to λ1=0.8\lambda_1 = 0.8 and λ2=0.6\lambda_2 = 0.6.
    • LsegL_{seg} is the standard pixel-level cross-entropy segmentation loss: Lseg=−∑m=1M∑j=1H×W∑c=1Cclsymjclog⁡(pmjc)L_{seg} = -\sum_{m=1}^M \sum_{j=1}^{H \times W} \sum_{c=1}^{C_{cls}} y_{mjc} \log(p_{mjc}) where MM is the number of training images, H×WH \times W is the spatial resolution, CclsC_{cls} is the number of semantic categories, ymjc∈{0,1}y_{mjc} \in \{0, 1\} is the one-hot ground-truth label, and pmjcp_{mjc} is the predicted probability.
  6. Knowl 6 — Synthetic-to-Real Domain Generalization Performance

    data/table

    The domain generalization performance of DIRL was evaluated by training DeepLabV3+ models on the synthetic GTAV dataset (19 shared classes) and testing on three real-world driving datasets: Cityscapes (C), BDD100K (B), and Mapillary Vistas (M). Results are reported in mean Intersection-over-Union (mIoU, %).

    Backbone Method C B M Mean
    ResNet50 Baseline 28.95 25.12 28.18 27.42
    SW 29.91 27.48 29.71 29.03
    IBN-Net 33.85 32.30 37.75 34.63
    IterNorm 31.81 32.70 33.88 32.80
    DPRC 37.42 32.14 34.12 34.56
    IW 33.21 32.67 37.35 34.41
    IRW 33.57 33.18 38.42 35.06
    ISW 36.58 35.20 40.33 37.37
    DIRL 41.04 39.15 41.60 40.60
    ShuffleNet Baseline 25.56 22.17 28.60 25.44
    IBN-Net 27.10 31.82 34.89 31.27
    ISW 30.98 32.06 35.31 32.78
    DIRL 31.88 32.57 36.12 33.52
    MobileNet Baseline 25.92 25.73 26.45 26.03
    IBN-Net 30.14 27.66 27.07 28.29
    ISW 30.86 30.05 30.67 30.53
    DIRL 34.67 32.78 34.31 33.92

    DIRL achieves the highest generalization mIoU across all target domains and backbones, outperforming the previous state-of-the-art method ISW by +3.23% mean mIoU on ResNet-50, +0.74% on ShuffleNet, and +3.39% on MobileNet.

  7. Knowl 7 — Real-to-Synthetic and Adverse Conditions Domain Generalization Performance

    data/table

    The domain generalization capability of DIRL was evaluated by training on the real-world Cityscapes dataset and testing on unseen target domains: BDD100K (B), Synthia (S, 16 classes), and GTAV (G). The target domains encompass synthetic environments as well as adverse conditions (e.g., rainy and nighttime scenes in BDD100K). Results are measured in mIoU (%).

    Backbone Method B S G Mean
    ResNet50 Baseline 44.96 23.29 42.55 36.93
    SW 48.49 26.10 44.87 39.82
    IBN-Net 48.56 26.14 45.06 39.92
    IterNorm 49.23 25.98 45.73 40.31
    IW 48.19 25.81 45.21 39.74
    IRW 48.67 26.05 45.64 40.12
    ISW 50.73 26.20 45.00 40.64
    DIRL 51.80 26.50 46.52 41.61
    ShuffleNet Baseline 38.09 21.25 36.45 31.93
    IBN-Net 41.89 22.99 40.91 35.26
    ISW 41.94 22.82 40.17 34.98
    DIRL 42.55 23.74 41.23 35.84
    MobileNet Baseline 40.13 21.64 37.32 33.03
    IBN-Net 44.97 23.23 41.13 36.44
    ISW 45.17 22.91 41.17 36.42
    DIRL 47.55 23.29 41.43 37.42

    DIRL achieves the highest performance across all backbones, reaching 41.61% mean mIoU on ResNet-50 (+0.97% over ISW) and maintaining superior robustness on both synthetic targets (Synthia, GTAV) and adverse driving scenes (BDD).

  8. Knowl 8 — Ablation of Feature Re-calibration and Whitening Components in DIRL

    data/table

    An ablation study evaluated the individual contributions of DIRL components on the GTAV →\to {Cityscapes (C), BDD (B), Mapillary (M)} task using DeepLabV3+ with a ResNet-50 backbone.

    Method LSGL_{SG} LISWL_{ISW} LGFWL_{GFW} C B M
    Baseline 28.95 25.14 28.18
    DA 30.81 26.32 29.05
    DU 33.57 28.74 30.24
    DIRL variants ✓ 36.60 30.66 33.55
    ✓ ✓ 40.20 38.10 40.79
    ✓ ✓ 41.04 39.15 41.60

    Key observations:

    • 'DA' (direct addition of PGAM without sensitivity guidance loss LSGL_{SG}) and 'DU' (directly scaling features by the fixed inverse sensitivity values (1−s)(1-s)) underperform learnable recalibration guided by LSGL_{SG} (36.60% vs. 30.81% and 33.57% on Cityscapes).
    • Combining LSGL_{SG} with Guided Feature Whitening (LGFWL_{GFW}) outperforms combining LSGL_{SG} with Instance Selective Whitening (LISWL_{ISW}) across all datasets (41.04% vs. 40.20% on Cityscapes), confirming the benefit of sensitivity-guided covariance decoupling over variance-statistic decoupling.
  9. Knowl 9 — Hyperparameter Sensitivity Analysis for DIRL

    data/table

    The performance of DIRL on GTAV →\to {Cityscapes, BDD, Mapillary} with ResNet-50 was evaluated across varying values of the selective mask ratio α\alpha, the Sensitivity Guidance loss weight λ1\lambda_1, and the Guided Feature Whitening loss weight λ2\lambda_2. Performance is reported as mean mIoU (%) across the three target domains.

    Choice of α\alpha 0.2 0.3 0.4 0.5
    Mean mIoU 38.30 40.60 39.07 38.11
    Choice of λ1\lambda_1 0.4 0.6 0.8 1.0
    Mean mIoU 37.85 39.77 40.60 39.50
    Choice of λ2\lambda_2 0.2 0.4 0.6 0.8
    Mean mIoU 38.79 38.94 40.60 39.57

    The sensitivity curves for α\alpha, λ1\lambda_1, and λ2\lambda_2 all follow a bell-shaped profile. Optimal performance (40.60% mean mIoU) is achieved at α=0.3\alpha = 0.3, λ1=0.8\lambda_1 = 0.8, and λ2=0.6\lambda_2 = 0.6.

  10. Knowl 10 — Computational and Inference Efficiency of DIRL

    data/table

    The model complexity and inference speed of DIRL were compared against Baseline (DeepLabV3+), IBN-Net, and ISW using a ResNet-50 backbone. Inference benchmarking was conducted at a resolution of 2048×10242048 \times 1024 on a single NVIDIA V100 GPU, averaged over 500 runs.

    Models Params (M) GFLOPS Inference Time (ms)
    Baseline 45.082 554.31 10.71
    IBN-Net 45.083 554.31 10.18
    ISW 45.081 554.31 10.50
    DIRL 45.414 554.98 10.93

    DIRL introduces only 0.332M additional parameters, 0.67 GFLOPS of additional compute, and an inference latency increase of 0.22 ms relative to the baseline, confirming that the module additions do not add meaningful computational overhead at test time.

Coverage note — None was omitted; all key architectural modules, loss formulations, benchmark domain generalization results, ablation studies, hyperparameter analyses, and computational efficiency metrics are covered.

References

  1. 1.Bau, A.; Belinkov, Y.; Sajjad, H.; Durrani, N.; Dalvi, F.; and Glass, J. 2018. Identifying and controlling important neurons in neural machine translation. arXiv preprint arXiv:1811.01157.
  2. 2.Bau, D.; Zhou, B.; Khosla, A.; Oliva, A.; and Torralba, A. 2017. Network dissection: Quantifying interpretability of deep visual representations. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6541–6549.
  3. 3.Bau, D.; Zhu, J.-Y.; Strobelt, H.; Lapedriza, A.; Zhou, B.; and Torralba, A. 2020. Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences, 117(48): 30071–30078.
  4. 4.Chen, L.-C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; and Yuille, A. L. 2017. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 40(4): 834–848.
  5. 5.Choi, S.; Jung, S.; Yun, H.; Kim, J. T.; Kim, S.; and Choo, J. 2021. RobustNet: Improving Domain Generalization in Urban-Scene Segmentation via Instance Selective Whitening. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 11580–11590.
  6. 6.Chu, W.; Hung, W.-C.; Tsai, Y.-H.; Cai, D.; and Yang, M.-H. 2019. Weakly-supervised caricature face parsing through domain adaptation. In The IEEE International Conference on Image Processing (ICIP), 3282–3286.
  7. 7.Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; and Schiele, B. 2016. The cityscapes dataset for semantic urban scene understanding. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3213–3223.
  8. 8.Dou, Q.; Coelho de Castro, D.; Kamnitsas, K.; and Glocker, B. 2019. Domain generalization via model-agnostic learning of semantic features. Advances in Neural Information Processing Systems (NeurIPS), 32: 6450–6461.
  9. 9.Ghifary, M.; Kleijn, W. B.; Zhang, M.; and Balduzzi, D. 2015. Domain generalization for object recognition with multi-task autoencoders. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2551–2559.
  10. 10.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep residual learning for image recognition. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770–778.
  11. 11.Hu, J.; Shen, L.; and Sun, G. 2018. Squeeze-and-excitation networks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 7132–7141.
  12. 12.Hu, X.; Fu, C.-W.; Zhu, L.; and Heng, P.-A. 2019. Depth-attentional features for single-image rain removal. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 8022–8031.
  13. 13.Huang, J.; Guan, D.; Xiao, A.; and Lu, S. 2021. Fsdr: Frequency space domain randomization for domain generalization. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6891–6902.
  14. 14.Huang, L.; Zhou, Y.; Zhu, F.; Liu, L.; and Shao, L. 2019. Iterative normalization: Beyond standardization towards efficient whitening. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4874–4883.
  15. 15.Lee, S.; Hyun, J.; Seong, H.; and Kim, E. 2020. Unsupervised Domain Adaptation for Semantic Segmentation by Content Transfer. arXiv preprint arXiv:2012.12545.
  16. 16.Li, D.; Yang, Y.; Song, Y.-Z.; and Hospedales, T. M. 2018a. Learning to generalize: Meta-learning for domain generalization. In Thirty-Second AAAI Conference on Artificial Intelligence (AAAI).
  17. 17.Li, D.; Zhang, J.; Yang, Y.; Liu, C.; Song, Y.-Z.; and Hospedales, T. M. 2019. Episodic training for domain generalization. In The Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 1446–1455.
  18. 18.Li, H.; Pan, S. J.; Wang, S.; and Kot, A. C. 2018b. Domain generalization with adversarial feature learning. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 5400–5409.
  19. 19.Li, Y.; Yuan, L.; and Vasconcelos, N. 2019. Bidirectional learning for domain adaptation of semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6936–6945.
  20. 20.Lin, G.; Milan, A.; Shen, C.; and Reid, I. 2017. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 1925–1934.
  21. 21.Liu, W.; Rabinovich, A.; and Berg, A. C. 2015. Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579.
  22. 22.Ma, N.; Zhang, X.; Zheng, H.-T.; and Sun, J. 2018. Shufflenet v2: Practical guidelines for efficient cnn architecture design. In Proceedings of the European conference on computer vision (ECCV), 116–131.
  23. 23.Mancini, M.; Bulo, S. R.; Caputo, B.; and Ricci, E. 2018. Best sources forward: domain generalization through source-specific nets. In The IEEE international conference on image processing (ICIP), 1353–1357. IEEE.
  24. 24.Muandet, K.; Balduzzi, D.; and Schölkopf, B. 2013. Domain generalization via invariant feature representation. In International Conference on Machine Learning (ICML), 10–18.
  25. 25.Neuhold, G.; Ollmann, T.; Rota Bulo, S.; and Kontschieder, P. 2017. The mapillary vistas dataset for semantic understanding of street scenes. In The Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 4990–4999.
  26. 26.Olah, C.; Satyanarayan, A.; Johnson, I.; Carter, S.; Schubert, L.; Ye, K.; and Mordvintsev, A. 2018. The building blocks of interpretability. Distill, 3(3): e10.
  27. 27.Pan, X.; Luo, P.; Shi, J.; and Tang, X. 2018. Two at once: Enhancing learning and generalization capacities via ibn-net. In The Proceedings of the European Conference on Computer Vision (ECCV), 464–479.
  28. 28.Pan, X.; Zhan, X.; Shi, J.; Tang, X.; and Luo, P. 2019. Switchable whitening for deep representation learning. In The Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 1863–1871.
  29. 29.Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems (NeurIPS), 32: 8026–8037.
  30. 30.Qiao, F.; Zhao, L.; and Peng, X. 2020. Learning to learn single domain generalization. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 12556–12565.
  31. 31.Richter, S. R.; Vineet, V.; Roth, S.; and Koltun, V. 2016. Playing for data: Ground truth from computer games. In The Proceedings of the European Conference on Computer Vision (ECCV), 102–118. Springer.
  32. 32.Ros, G.; Sellart, L.; Materzynska, J.; Vazquez, D.; and Lopez, A. M. 2016. The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3234–3243.
  33. 33.Roy, S.; Siarohin, A.; Sangineto, E.; Bulo, S. R.; Sebe, N.; and Ricci, E. 2019. Unsupervised domain adaptation using feature-whitening and consensus loss. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 9471–9480.
  34. 34.Saito, K.; Watanabe, K.; Ushiku, Y.; and Harada, T. 2018. Maximum classifier discrepancy for unsupervised domain adaptation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3723–3732.
  35. 35.Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; and Chen, L.-C. 2018. Mobilenetv2: Inverted residuals and linear bottlenecks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4510–4520.
  36. 36.Seo, S.; Suh, Y.; Kim, D.; Kim, G.; Han, J.; and Han, B. 2020. Learning to optimize domain specific normalization for domain generalization. In The European Conference on Computer Vision (ECCV), 68–83. Springer.
  37. 37.Tobin, J.; Fong, R.; Ray, A.; Schneider, J.; Zaremba, W.; and Abbeel, P. 2017. Domain randomization for transferring deep neural networks from simulation to the real world. In The international conference on intelligent robots and systems (IROS), 23–30. IEEE.
  38. 38.Tsai, Y.-H.; Hung, W.-C.; Schulter, S.; Sohn, K.; Yang, M.-H.; and Chandraker, M. 2018. Learning to adapt structured output space for semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 7472–7481.
  39. 39.Ulyanov, D.; Vedaldi, A.; and Lempitsky, V. 2016. Instance normalization: The missing ingredient for fast stylization. arXiv preprint arXiv:1607.08022.
  40. 40.Valada, A.; Vertens, J.; Dhall, A.; and Burgard, W. 2017. AdapNet: Adaptive semantic segmentation in adverse environmental conditions. In The IEEE International Conference on Robotics and Automation (ICRA).
  41. 41.Volpi, R.; Namkoong, H.; Sener, O.; Duchi, J.; Murino, V.; and Savarese, S. 2018. Generalizing to unseen domains via adversarial data augmentation. arXiv preprint arXiv:1805.12018.
  42. 42.Vu, T.-H.; Jain, H.; Bucher, M.; Cord, M.; and Pérez, P. 2019. Advent: Adversarial entropy minimization for domain adaptation in semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2517–2526.
  43. 43.Wu, A.; Han, Y.; Zhu, L.; and Yang, Y. 2021. Instance-Invariant Domain Adaptive Object Detection via Progressive Disentanglement. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI).
  44. 44.Yang, Y.; and Soatto, S. 2020. Fda: Fourier domain adaptation for semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4085–4095.
  45. 45.Yu, F.; Chen, H.; Wang, X.; Xian, W.; Chen, Y.; Liu, F.; Madhavan, V.; and Darrell, T. 2020. Bdd100k: A diverse driving dataset for heterogeneous multitask learning. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2636–2645.
  46. 46.Yu, F.; Zhang, M.; Dong, H.; Hu, S.; Dong, B.; and Zhang, L. 2021. DAST: Unsupervised Domain Adaptation in Semantic Segmentation Based on Discriminator Attention and Self-Training. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 10754–10762.
  47. 47.Yue, X.; Zhang, Y.; Zhao, S.; Sangiovanni-Vincentelli, A.; Keutzer, K.; and Gong, B. 2019. Domain randomization and pyramid consistency: Simulation-to-real generalization without accessing target domain data. In The Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2100–2110.
  48. 48.Zhang, P.; Zhang, B.; Zhang, T.; Chen, D.; Wang, Y.; and Wen, F. 2021. Prototypical pseudo label denoising and target structure learning for domain adaptive semantic segmentation. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 12414–12424.
  49. 49.Zheng, S.; Lu, J.; Zhao, H.; Zhu, X.; Luo, Z.; Wang, Y.; Fu, Y.; Feng, J.; Xiang, T.; Torr, P. H.; et al. 2021. Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6881–6890.
  50. 50.Zou, Y.; Yu, Z.; Liu, X.; Kumar, B.; and Wang, J. 2019. Confidence regularized self-training. In The IEEE International Conference on Computer Vision (ICCV), 5982–5991.
  51. 51.Zou, Y.; Yu, Z.; Vijaya Kumar, B.; and Wang, J. 2018. Unsupervised domain adaptation for semantic segmentation via class-balanced self-training. In The European Conference on Computer Vision (ECCV), 289–305.

Citation

MLA
Xu, Q., et al. “DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, 2022, pp. 2884–92, https://doi.org/10.1609/AAAI.V36I3.20193.
APA
Xu, Q., Yao, L., Jiang, Z., Jiang, G., Chu, W., Han, W., Zhang, W., Wang, C., & Tai, Y. (2022). DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation. Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), 2884–2892. https://doi.org/10.1609/AAAI.V36I3.20193
Chicago
Xu, Q., L. Yao, Z. Jiang, et al. 2022. “DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation”. Proceedings of the AAAI Conference on Artificial Intelligence 36 (3): 2884–92. https://doi.org/10.1609/AAAI.V36I3.20193.
Harvard
Xu, Q. et al. (2022) “DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation”, Proceedings of the AAAI Conference on Artificial Intelligence, 36(3), pp. 2884–2892. Available at: https://doi.org/10.1609/AAAI.V36I3.20193.
Vancouver
1. Xu Q, Yao L, Jiang Z, Jiang G, Chu W, Han W, Zhang W, Wang C, Tai Y (2022) DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation. Proceedings of the AAAI Conference on Artificial Intelligence 36:2884–2892

BibTeX

@article{Xu_2022, title={DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation}, volume={36}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/AAAI.V36I3.20193}, DOI={10.1609/aaai.v36i3.20193}, number={3}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Xu, Qi and Yao, Liang and Jiang, Zhengkai and Jiang, Guannan and Chu, Wenqing and Han, Wenhui and Zhang, Wei and Wang, Chengjie and Tai, Ying}, year={2022}, month=June, pages={2884–2892} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF