DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation
Qi XuLiang YaoZhengkai JiangGuannan JiangWenqing ChuWenhui HanWei ZhangChengjie WangYing Tai
Proposes a domain generalization framework for semantic segmentation that quantifies channel-level sensitivity to style shifts, using this prior knowledge to re-calibrate feature responses and selectively whiten domain-sensitive feature correlations.
Deep learning models for visual semantic segmentation achieve strong results when evaluated on familiar environments, but their accuracy degrades sharply when deployed in new, unseen environments. In safety-critical fields such as autonomous driving, training models on every conceivable real-world scenario—such as adverse weather, lighting shifts, and varied geographic locations—is practically impossible. Most existing approaches to this challenge overlook an essential property of internal representations: certain visual features are inherently sensitive to superficial environmental style, while others reliably capture underlying domain-invariant content.
The article develops and evaluates Domain-Invariant Representation Learning, a framework designed to enhance single-source domain generalization in semantic segmentation. The primary objective is to measure feature sensitivity to domain styles and use that sensitivity as a guide to both suppress style-dependent features and selectively eliminate sensitive feature correlations.
The approach introduces three integrated modules without requiring target-domain data during training. First, a Sensitivity-aware Prior Module estimates channel sensitivity by calculating feature differences between original images and style-altered versions generated via photometric transformations. Second, a Prior Guided Attention Module trains channel-wise attention weights to penalize sensitive features and reward robust ones. Third, Guided Feature Whitening decouples feature covariances using the sensitivity prior and removes the correlations most vulnerable to domain shifts. The framework was evaluated across standard benchmarks transferring from synthetic datasets (GTAV, Synthia) to diverse real-world urban datasets (Cityscapes, BDD, Mapillary) using multiple deep backbones, including ResNet-50, ShuffleNet, and MobileNet.
The findings show substantial and consistent performance gains across all evaluated settings. When trained on synthetic GTAV data and tested across Cityscapes, BDD, and Mapillary, the ResNet-50 model achieved a mean intersection-over-union score of 40.60%, outperforming the standard baseline at 27.42% and the strongest prior selective whitening method at 37.37%. Similar consistent improvements occurred on lightweight architectures, raising ShuffleNet performance from 25.44% to 33.52% and MobileNet from 26.03% to 33.92%. Models trained on clean urban datasets maintained reasonable segmentation predictions when exposed to severe unseen conditions like night driving and heavy rain. Crucially, the computational overhead remained minimal, increasing ResNet-50 inference time by only 0.22 milliseconds per frame (from 10.71 ms to 10.93 ms on an NVIDIA V100 GPU).
These results demonstrate that explicitly quantifying feature sensitivity enables networks to decouple style from content more cleanly than previous statistical whitening methods. By delivering superior generalization without extra data collection or noticeable latency costs, this technique lowers operational risk and improves perception safety in unpredictable real-world operating conditions.
Engineering and research teams deploying computer vision in autonomous driving or robotics should integrate sensitivity-guided attention and whitening into early convolutional layers of existing architectures. Future development should focus on expanding validation across a wider variety of outdoor operational domains and testing the framework alongside broader multi-sensor perception stacks.
A primary limitation is that sensitivity modeling relies on simulated style transformations such as color jitter and blurring, which may not capture all real-world environmental variations like structural or sensor-level anomalies. Nonetheless, the consistent empirical gains across multiple datasets and network backbones provide high confidence in the framework's effectiveness for single-source domain generalization.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). Provides a comprehensive foundational survey of domain generalization principles and data manipulation versus representation learning paradigms upon which DIRL is built.
- Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). Establishes core domain generalization concepts and taxonomy, clarifying the distinction between domain-invariant representation learning and data manipulation techniques.
- Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). Establishes standardized benchmark methodologies and empirical baselines for out-of-distribution evaluation that contextualize the evaluation protocol used in DIRL.
- Paper: The SYNTHIA Dataset: A Large Collection of Synthetic Images for Semantic Segmentation of Urban Scenes, Germán Ros et al. (2016). Introduces the SYNTHIA dataset, which serves as one of the primary synthetic source domains used to benchmark DIRL's cross-domain urban segmentation performance.
- Paper: Fully convolutional networks for semantic segmentation, Jonathan Long et al. (2015). Presents the foundational fully convolutional network paradigm that underlies modern visual semantic segmentation models adapted by DIRL.
- Paper: Learning Robust Global Representations by Penalizing Local Predictive Power, Haohan Wang et al. (2019). Introduces the concept of suppressing superficial, domain-sensitive local feature cues in early layers to encourage invariant global representations.
- Paper: Adversarial Style Augmentation for Domain Generalized Urban-Scene Segmentation, Zhun Zhong et al. (2022). Extends single-source domain generalized urban segmentation by dynamically generating adversarial visual styles during training rather than relying solely on static photometric transformations.
- Paper: Decompose, Adjust, Compose: Effective Normalization by Playing with Frequency for Domain Generalization, Sangrok Lee et al. (2023). Continues the study of feature normalization and domain-invariant representations by using frequency-domain phase and amplitude decomposition to prevent content distortion.
- Paper: Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation, Zixian Su et al. (2023). Applies single-source domain generalization principles to dense visual segmentation in medical imaging by using saliency-guided feature transformations.
- Paper: Improved Test-Time Adaptation for Domain Generalization, Liang Chen et al. (2023). Extends generalization beyond fixed training-time representations by incorporating test-time adaptation with lightweight tunable parameters during deployment.
- Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). Explores an alternative paradigm for robust cross-domain semantic segmentation by reformulating dense prediction as generative mask modeling.
