Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation
Zixian SuKai YaoXi YangKaizhu HuangQiufeng WangJie Sun
Proposes a class-level location-scale data augmentation framework paired with gradient-guided saliency balancing to guarantee bounded generalization risk and improve medical image segmentation across unseen domains.
Deep learning models deployed for medical image segmentation often fail when encountering data from different hospitals, imaging protocols, or scanner vendors. This distribution shift poses significant clinical and operational risks, as models trained on a single source dataset struggle to maintain accuracy on unseen target data. The article addresses the challenge of single-source domain generalization, where an automated segmentation model must be trained using data from only one source domain to perform reliably across new, unobserved imaging environments without retraining.
The main objective of the article is to demonstrate and theoretically validate a novel data augmentation framework, termed Saliency-balancing Location-scale Augmentation (SLAug), designed to improve the generalization capability of medical image segmentation models. The approach addresses the limitations of conventional global and random augmentations by combining global image adjustments with localized, organ-level transformations, while using model gradient cues (saliency maps) to guide how these images are blended.
To evaluate this framework, the authors conducted experiments across two challenging cross-domain benchmarks: abdominal organ segmentation across computed tomography (CT) and magnetic resonance imaging (MRI) modalities, and cardiac structure segmentation across different MRI imaging sequences. The method was implemented using a standard U-Net architecture with an EfficientNet backbone. The empirical evaluation compared the proposed framework against a standard training baseline and seven state-of-the-art domain generalization techniques using overlap accuracy (Dice score) as the primary evaluation metric.
The findings show substantial improvements in model robustness. First, SLAug outperformed all baseline and competing methods across both benchmarks, narrowing the generalization gap by an average of 47.77% relative to the prior leading method. Second, in cross-modality abdominal segmentation, SLAug achieved the highest average Dice scores of 88.63% (CT-to-MRI) and 83.05% (MRI-to-CT), surpassing the strongest existing baseline by approximately 2.49 percentage points. Third, the framework performed consistently across cardiac segmentation tasks, achieving top scores of 86.69% and 87.67% across differing sequence directions. Fourth, ablation studies confirmed that pairing global and local location-scale transformations with saliency-balancing fusion was critical; replacing saliency guidance with random image blending degraded cross-domain accuracy by 2.9 percentage points and lowered source-domain accuracy.
These results indicate that combining anatomically aware, class-level transformations with gradient-informed blending generates realistic, high-diversity training samples without inducing severe over-generalization. For healthcare organizations and technology developers, adopting this approach can improve clinical safety, lower deployment costs, and minimize the operational need to collect and annotate expensive new datasets for every imaging vendor or clinical site.
Decision-makers should consider integrating this plug-and-play augmentation module into existing medical segmentation training pipelines to improve baseline model generalizability. However, before deploying the technique in live diagnostic environments, technical teams should conduct validation across broader anatomical targets, larger multi-center clinical cohorts, and diverse imaging modalities. While confidence in the reported experimental and theoretical gains is high, real-world deployment should still account for potential boundary conditions where source and target modalities exhibit extreme structural divergence.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). Provides a comprehensive taxonomy and theoretical foundation of domain generalization strategies, establishing the conceptual backdrop for addressing single-source domain shift.
- Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). Surveys core methodologies in representation learning and data manipulation for generalizing to unseen domains, directly motivating the source paper's augmentation-centric focus.
- Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). Demonstrates the critical role of strong data augmentation and rigorous empirical risk minimization baselines when evaluating domain generalization algorithms.
- Paper: AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty, Dan Hendrycks et al. (2019). Introduces foundational principles of mixing augmented image chains to enhance model robustness against out-of-distribution shifts.
- Paper: U-Net: Convolutional Networks for Biomedical Image Segmentation, Olaf Ronneberger et al. (2015). Establishes the foundational U-Net architecture and basic data augmentation practices widely used as the structural baseline for medical image segmentation.
- Paper: The Medical Segmentation Decathlon, M. Antonelli et al. (2021). Details the multi-task medical segmentation benchmarks and standardization protocols that contextualize cross-modality evaluation in healthcare imaging.
- Paper: Attention Gated Networks: Learning to Leverage Salient Regions in Medical Images, Jo Schlemper et al. (2018). Explains how to leverage spatial saliency and attention mechanisms within convolutional networks to highlight critical anatomical regions in medical scans.
- Paper: Domain Generalization via Invariant Feature Representation, Krikamol Muandet et al. (2013). Introduces theoretical foundations and error bounds for learning invariant representations to generalize to unseen target domains without retraining.
- Paper: Improved Test-Time Adaptation for Domain Generalization, Liang Chen et al. (2023). Extends domain generalization from training-time data augmentation to test-time adaptation, dynamically updating models during deployment on unseen target distributions.
- Paper: Decompose, Adjust, Compose: Effective Normalization by Playing with Frequency for Domain Generalization, Sangrok Lee et al. (2023). Advances domain generalization techniques by analyzing content-style balance in the frequency domain to normalize features without content distortion.
- Paper: Generative Semantic Segmentation, Jiaqi Chen et al. (2023). Explores an alternative generative formulation of semantic segmentation to improve cross-domain generalization beyond standard discriminative training paradigms.
- Paper: Segment anything in medical images, Jun Ma et al. (2023). Builds upon universal medical segmentation by evaluating how large foundation models generalize across clinical modalities using interactive prompts.
