Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation

Zixian SuKai YaoXi YangKaizhu HuangQiufeng WangJie Sun

article2023AAAI119 citations

Proposes a class-level location-scale data augmentation framework paired with gradient-guided saliency balancing to guarantee bounded generalization risk and improve medical image segmentation across unseen domains.

Listen

Deep learning models deployed for medical image segmentation often fail when encountering data from different hospitals, imaging protocols, or scanner vendors. This distribution shift poses significant clinical and operational risks, as models trained on a single source dataset struggle to maintain accuracy on unseen target data. The article addresses the challenge of single-source domain generalization, where an automated segmentation model must be trained using data from only one source domain to perform reliably across new, unobserved imaging environments without retraining.

The main objective of the article is to demonstrate and theoretically validate a novel data augmentation framework, termed Saliency-balancing Location-scale Augmentation (SLAug), designed to improve the generalization capability of medical image segmentation models. The approach addresses the limitations of conventional global and random augmentations by combining global image adjustments with localized, organ-level transformations, while using model gradient cues (saliency maps) to guide how these images are blended.

To evaluate this framework, the authors conducted experiments across two challenging cross-domain benchmarks: abdominal organ segmentation across computed tomography (CT) and magnetic resonance imaging (MRI) modalities, and cardiac structure segmentation across different MRI imaging sequences. The method was implemented using a standard U-Net architecture with an EfficientNet backbone. The empirical evaluation compared the proposed framework against a standard training baseline and seven state-of-the-art domain generalization techniques using overlap accuracy (Dice score) as the primary evaluation metric.

The findings show substantial improvements in model robustness. First, SLAug outperformed all baseline and competing methods across both benchmarks, narrowing the generalization gap by an average of 47.77% relative to the prior leading method. Second, in cross-modality abdominal segmentation, SLAug achieved the highest average Dice scores of 88.63% (CT-to-MRI) and 83.05% (MRI-to-CT), surpassing the strongest existing baseline by approximately 2.49 percentage points. Third, the framework performed consistently across cardiac segmentation tasks, achieving top scores of 86.69% and 87.67% across differing sequence directions. Fourth, ablation studies confirmed that pairing global and local location-scale transformations with saliency-balancing fusion was critical; replacing saliency guidance with random image blending degraded cross-domain accuracy by 2.9 percentage points and lowered source-domain accuracy.

These results indicate that combining anatomically aware, class-level transformations with gradient-informed blending generates realistic, high-diversity training samples without inducing severe over-generalization. For healthcare organizations and technology developers, adopting this approach can improve clinical safety, lower deployment costs, and minimize the operational need to collect and annotate expensive new datasets for every imaging vendor or clinical site.

Decision-makers should consider integrating this plug-and-play augmentation module into existing medical segmentation training pipelines to improve baseline model generalizability. However, before deploying the technique in live diagnostic environments, technical teams should conduct validation across broader anatomical targets, larger multi-center clinical cohorts, and diverse imaging modalities. While confidence in the reported experimental and theoretical gains is high, real-world deployment should still account for potential boundary conditions where source and target modalities exhibit extreme structural divergence.

Cover for Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation

Abstract

Single-source domain generalization (SDG) in medical image segmentation is a challenging yet essential task as domain shifts are quite common among clinical image datasets. Previous attempts most conduct global-only/random augmentation. Their augmented samples are usually insufficient in diversity and informativeness, thus failing to cover the possible target domain distribution. In this paper, we rethink the data augmentation strategy for SDG in medical image segmentation. Motivated by the class-level representation invariance and style mutability of medical images, we hypothesize that unseen target data can be sampled from a linear combination of C (the class number) random variables, where each variable follows a location-scale distribution at the class level. Accordingly, data augmented can be readily made by sampling the random variables through a general form. On the empirical front, we implement such strategy with constrained Bézier transformation on both global and local (i.e. class-level) regions, which can largely increase the augmentation diversity. A Saliency-balancing Fusion mechanism is further proposed to enrich the informativeness by engaging the gradient information, guiding augmentation with proper orientation and magnitude. As an important contribution, we prove theoretically that our proposed augmentation can lead to an upper bound of the generalization risk on the unseen target domain, thus confirming our hypothesis. Combining the two strategies, our Saliency-balancing Location-scale Augmentation (SLAug) exceeds the state-of-the-art works by a large margin in two challenging SDG tasks. Code is available at https://github.com/Kaiseem/SLAug.

Table of Contents

  • Introduction
  • Related Work
  • Main Methodology
  • Preliminary
  • Location-scale Augmentation
  • Theoretical Evidence
  • Saliency-balancing Fusion
  • Experiments and Results
  • Datasets and Preprocessing
  • Network Architecture and Training Configurations
  • Results and Comparative Analysis
  • Analytical Experiments
  • Conclusion
  • References
  • A Appendix
  • Additive Experimental Details
  • Hyper-parameter Analysis
  • More Visualization

Knowls

  1. Knowl 1 — Saliency-Balancing Location-Scale Augmentation (SLAug) Algorithm

    algorithm

    The Saliency-balancing Location-scale Augmentation (SLAug) framework generates augmented training images for single-source domain generalization in medical image segmentation by combining global transformations, class-level local transformations, and gradient-based saliency fusion.

    Input: Training image x∈[0,1]W×Hx \in [0, 1]^{W \times H}, groundtruth segmentation mask m∈{1,…,C}W×Hm \in \{1, \dots, C\}^{W \times H}, segmentation network fθf_\theta, segmentation loss function L\mathcal{L}, location-scale hyperparameters (σ1,σ2)(\sigma_1, \sigma_2), downsampling grid size gg
    Output: Fused augmented image x~fused\tilde{x}_{fused}
    1: Compute global location-scale augmented image:
       xg←αF0(x)+βx^g \leftarrow \alpha F_0(x) + \beta, where α∼TN(1,σ1)\alpha \sim \mathcal{TN}(1, \sigma_1), β∼TN(0,σ2)\beta \sim \mathcal{TN}(0, \sigma_2)
    2: Compute local location-scale augmented image:
       xl←∑c=1C(αcFpc(mc⊙x)+βc)x^l \leftarrow \sum_{c=1}^C (\alpha_c F_{p_c}(m^c \odot x) + \beta_c), where αc∼TN(1,σ1)\alpha_c \sim \mathcal{TN}(1, \sigma_1), βc∼TN(0,σ2)\beta_c \sim \mathcal{TN}(0, \sigma_2), p1=1p_1 = 1, and pc>1=0.5p_{c > 1} = 0.5
    3: Apply standard spatial and intensity augmentations FF to generate paired views:
       (x~g,x~l,m~)←F(xg,xl,m)(\tilde{x}^g, \tilde{x}^l, \tilde{m}) \leftarrow F(x^g, x^l, m)
    4: Calculate pixel-level gradient w.r.t. the GLA sample:
       Grad←∇x~gL(fθ(x~g),m~)Grad \leftarrow \nabla_{\tilde{x}^g} \mathcal{L}(f_\theta(\tilde{x}^g), \tilde{m})
    5: Compute normalized saliency map ss:
       sraw←∥Grad∥2 across channelss_{raw} \leftarrow \|Grad\|_2 \text{ across channels}
       sdown←Downsample sraw to g×gs_{down} \leftarrow \text{Downsample } s_{raw} \text{ to } g \times g
       ssmooth←Interpolate sdown to W×H using quadratic B-spliness_{smooth} \leftarrow \text{Interpolate } s_{down} \text{ to } W \times H \text{ using quadratic B-splines}
       s←ssmooth−min⁡(ssmooth)max⁡(ssmooth)−min⁡(ssmooth)s \leftarrow \frac{s_{smooth} - \min(s_{smooth})}{\max(s_{smooth}) - \min(s_{smooth})}
    6: Harmonize features into a fused training sample:
       x~fused←s⊙x~g+(1−s)⊙x~l\tilde{x}_{fused} \leftarrow s \odot \tilde{x}^g + (1 - s) \odot \tilde{x}^l
    7: return x~fused\tilde{x}_{fused}

    The segmentation network fθf_\theta is trained jointly on the global location-scale augmented images x~g\tilde{x}^g and the saliency-fused images x~fused\tilde{x}_{fused} using the total objective L(x,m)=Lce(x,m)+Ldice(x,m)\mathcal{L}(x, m) = \mathcal{L}_{ce}(x, m) + \mathcal{L}_{dice}(x, m), where Lce\mathcal{L}_{ce} denotes standard cross-entropy loss and Ldice\mathcal{L}_{dice} denotes Dice loss.

  2. Knowl 2 — Generalization Error Bound under Location-Scale Augmentation

    theoretical result

    Let XS={(xi,S,mi,S)}i=1nS\mathcal{X}_S = \{(x_{i,S}, m_{i,S})\}_{i=1}^{n_S} denote data from NSN_S source domains/distributions {ϕSi}i=1NS\{\phi_S^i\}_{i=1}^{N_S} with labeling functions fSif_S^i, and let XU={(xi,U,mi,U)}i=1nU\mathcal{X}_U = \{(x_{i,U}, m_{i,U})\}_{i=1}^{n_U} be an unseen target domain distribution ϕU\phi_U with labeling function fUf_U. Let ΛS={ϕˉ:ϕˉ(⋅)=∑i=1NSπiϕSi(⋅),π∈ΔNS−1}\Lambda_S = \{\bar{\phi} : \bar{\phi}(\cdot) = \sum_{i=1}^{N_S} \pi_i \phi_S^i(\cdot), \pi \in \Delta_{N_S - 1}\} define the convex hull of source distributions, and let ϕˉU=arg⁡min⁡πdH[ϕU,∑i=1NSπiϕSi]\bar{\phi}_U = \arg\min_{\pi} d_\mathcal{H}[\phi_U, \sum_{i=1}^{N_S} \pi_i \phi_S^i] be the mixture in ΛS\Lambda_S closest to ϕU\phi_U with divergence distance γ=dH[ϕˉU,ϕU]\gamma = d_\mathcal{H}[\bar{\phi}_U, \phi_U].

    Assume the following:

    1. Target Representation: Unseen target data xUx_U can be represented as a linear combination of class-level random variables following the location-scale distribution of the source, xU=∑c=1C(αcxSc+βc)x_U = \sum_{c=1}^C (\alpha_c x_S^c + \beta_c), placing the unseen distribution entirely within the convex hull of the augmented source mixture distributions (i.e., ϕU∈ΛS\phi_U \in \Lambda_S, yielding γ=0\gamma = 0).
    2. Bounded Augmentation Divergence: The pairwise H\mathcal{H}-divergence between distinct augmented distributions ϕSaug′\phi_S^{aug'} and ϕSaug′′\phi_S^{aug''} is upper-bounded by δ=max⁡i,jdH[ϕSaugi,ϕSaugj]\delta = \max_{i,j} d_\mathcal{H}[\phi_S^{aug_i}, \phi_S^{aug_j}].
    3. Covariate Shift: The true labeling function is domain-invariant (fSπ=fUf_{S\pi} = f_U).

    If the empirical source risk on each source domain ii is bounded by L(yi,y)=ϵi≤σL(y^i, y) = \epsilon_i \le \sigma, then the risk RU[h]R_U[h] of any hypothesis h∈Hh \in \mathcal{H} on the unseen domain is upper-bounded by:

    RU[h]≤∑i=1NSπiRSi[h]+δ≤σ+δR_U[h] \le \sum_{i=1}^{N_S} \pi_i R_S^i[h] + \delta \le \sigma + \delta

    where σ\sigma represents the source risk minimized during training, and δ\delta is the maximum discrepancy between augmented views constrained by semantic consistency.

  3. Knowl 3 — Global and Local Location-Scale Augmentation with Constrained Bézier Transformations

    model/method

    Location-scale augmentation models cross-domain medical image shifts through monotonic intensity transformations parameterized by location and scale factors applied at both global and class-level resolutions.

    The monotonic intensity transfer function is formulated via a cubic Bézier curve over the pixel intensity range [vlow,vhigh][v_{low}, v_{high}] of non-blank foreground regions:

    Beˊzier(t)=∑k=03(3k)(1−t)3−ktkPk,t∈[vlow,vhigh]B\acute{e}zier(t) = \sum_{k=0}^3 \binom{3}{k} (1 - t)^{3-k} t^k P_k, \quad t \in [v_{low}, v_{high}]

    where P0=(vlow,vlow)P_0 = (v_{low}, v_{low}) and P3=(vhigh,vhigh)P_3 = (v_{high}, v_{high}) anchor the minimum and maximum values, while intermediate control points P1,P2P_1, P_2 are randomly sampled from [vlow,vhigh][v_{low}, v_{high}]. An inverse mapping swaps the anchors to P0=(vlow,vhigh)P_0 = (v_{low}, v_{high}) and P3=(vhigh,vlow)P_3 = (v_{high}, v_{low}). The operator Fp(⋅)F_p(\cdot) denotes the Bézier transformation with inversion probability pp.

    1. Global Location-scale Augmentation (GLA): Transforms the normalized image x∈[0,1]x \in [0, 1] globally: GLA(x)=αF0(x)+βGLA(x) = \alpha F_0(x) + \beta where α∼TN(1,σ1)\alpha \sim \mathcal{TN}(1, \sigma_1) and β∼TN(0,σ2)\beta \sim \mathcal{TN}(0, \sigma_2) are sampled from truncated Gaussian distributions, and p=0p = 0 prevents full contrast inversion to preserve source-like appearance.

    2. Local Location-scale Augmentation (LLA): Uses the ground truth semantic mask m=∑c=1Cmcm = \sum_{c=1}^C m^c to segment class-level components xc=mc⊙xx^c = m^c \odot x, applying independent transforms before linear recombination: LLA(x,m)=∑c=1C(αcFpc(xc)+βc)LLA(x, m) = \sum_{c=1}^C \left( \alpha_c F_{p_c}(x^c) + \beta_c \right) where αc∼TN(1,σ1)\alpha_c \sim \mathcal{TN}(1, \sigma_1), βc∼TN(0,σ2)\beta_c \sim \mathcal{TN}(0, \sigma_2), p1=1.0p_1 = 1.0 (inverting background to force dissimilarity from GLA), and pc>1=0.5p_{c > 1} = 0.5. Hyperparameters are set to σ1=0.1\sigma_1 = 0.1 and σ2=0.5\sigma_2 = 0.5.

  4. Knowl 4 — Saliency-Balancing Fusion Mechanism

    model/method

    Saliency-Balancing Fusion (SBF) blends a globally augmented image x~g\tilde{x}^g with a locally augmented image x~l\tilde{x}^l using the network's loss gradients as a spatial sensitivity weight, preventing over-generalization while expanding domain coverage.

    Given the globally augmented image x~g\tilde{x}^g, ground truth mask m~\tilde{m}, and segmentation model fθf_\theta, the gradient is computed as:

    Grad=∇x~gL(fθ(x~g),m~)Grad = \nabla_{\tilde{x}^g} \mathcal{L}(f_\theta(\tilde{x}^g), \tilde{m})

    The channel-wise ℓ2\ell_2-norm of GradGrad identifies large-gradient regions where the model is sensitive or prone to error under slight distribution shifts. To remove high-frequency gradient noise, the raw saliency map is downsampled to a spatial grid size of g×gg \times g (e.g., g=3g=3 for abdominal CT/MRI, g=18g=18 for cardiac MR) and subsequently upsampled back to the original image dimensions (W,H)(W, H) via quadratic B-spline interpolation smoothing. After min-max normalization to [0,1][0, 1], the smoothed saliency map ss guides convex combination:

    x~fused=s⊙x~g+(1−s)⊙x~l\tilde{x}_{fused} = s \odot \tilde{x}^g + (1 - s) \odot \tilde{x}^l

    This retains the source-like GLA appearance in high-gradient sensitive regions (s≈1s \approx 1) while introducing high-diversity LLA transformations in low-gradient background/stable areas (s≈0s \approx 0).

  5. Knowl 5 — Class-Level Location-Scale Hypothesis for Medical Image Domain Shift

    assumption

    In cross-modality and cross-sequence medical image segmentation, unseen target domain images xUx_U can be modeled as a linear combination of CC semantic class regions, where each class region is generated via a location-scale transformation of the corresponding source domain region xSc=mc⊙xSx_S^c = m^c \odot x_S:

    xU=∑c=1CxUc=∑c=1C(αcxSc+βc)x_U = \sum_{c=1}^C x_U^c = \sum_{c=1}^C (\alpha^c x_S^c + \beta^c)

    Here, CC is the number of semantic segmentation classes, mc∈{0,1}W×Hm^c \in \{0, 1\}^{W \times H} is the binary mask for class cc, αc>0\alpha^c > 0 is a class-specific scale parameter governing local contrast/variance, and βc∈R\beta^c \in \mathbb{R} is a class-specific location parameter governing local mean intensity shift. This hypothesis formalizes two empirical properties of medical scans: class-level representation invariance (structural pixel-level similarity within an organ across modalities) and style mutability (inter-domain variance manifesting as class-level brightness, contrast, and clarity shifts).

  6. Knowl 6 — Cross-Modality and Cross-Sequence Segmentation Benchmark Performance of SLAug

    data/table

    SLAug achieves state-of-the-art single-source domain generalization performance across abdominal CT-to-MRI, abdominal MRI-to-CT, cardiac bSSFP-to-LGE, and cardiac LGE-to-bSSFP segmentation benchmarks, evaluated using the Dice similarity coefficient (%):

    Abdominal CT →\to MRI Cardiac bSSFP →\to LGE
    Method Liver R-Kidney L-Kidney Spleen Average LVC MYO RVC Average
    Supervised 91.30 92.43 89.86 89.83 90.85 92.04 83.11 89.30 88.15
    ERM 78.03 78.11 78.45 74.65 77.31 86.06 66.98 74.94 75.99
    Cutout 79.80 82.32 82.14 76.24 80.12 88.35 69.06 79.19 78.87
    RSC 76.40 75.79 76.60 67.56 74.09 87.06 69.77 75.69 77.51
    MixStyle 77.63 78.41 78.03 77.12 77.80 85.78 64.23 75.61 75.21
    AdvBias 78.54 81.70 80.69 79.73 80.17 88.23 70.29 80.32 79.62
    RandConv 73.63 79.69 85.89 83.43 80.66 89.88 75.60 85.70 83.73
    CSDG 86.62 87.48 86.88 84.27 86.31 90.35 77.82 86.87 85.01
    SLAug (ours) 90.08 89.23 87.54 87.67 88.63 91.53 80.65 87.90 86.69
    Abdominal MRI →\to CT Cardiac LGE →\to bSSFP
    Method Liver R-Kidney L-Kidney Spleen Average LVC MYO RVC Average
    Supervised 98.87 92.11 91.75 88.55 89.74 91.16 82.93 90.39 88.16
    ERM 87.90 40.44 65.17 55.90 62.35 90.16 78.59 87.04 85.26
    Cutout 86.99 63.66 73.74 57.60 70.50 90.88 79.14 87.74 85.92
    RSC 88.10 46.60 75.94 53.61 66.07 90.21 78.63 87.96 85.60
    MixStyle 86.66 48.26 65.20 55.68 63.95 91.22 79.64 88.16 86.34
    AdvBias 87.63 52.48 68.28 50.95 64.84 91.20 79.50 88.10 86.27
    RandConv 84.14 76.81 77.99 67.32 76.56 91.98 80.92 88.83 87.24
    CSDG 85.62 80.02 80.42 75.56 80.40 91.37 80.43 89.16 86.99
    SLAug (ours) 89.26 80.98 82.05 79.93 83.05 91.92 81.49 89.61 87.67

    SLAug improves upon the strongest prior baseline (CSDG) by 2.32% on Abdominal CT →\to MRI, 2.65% on Abdominal MRI →\to CT, 1.68% on Cardiac bSSFP →\to LGE, and 0.68% on Cardiac LGE →\to bSSFP, narrowing the domain gap to the supervised upper bound by 47.77% across all four transfer tasks.

  7. Knowl 7 — Ablation Study of SLAug Components

    data/table

    An ablation study isolates the contributions of Global Location-scale Augmentation (GLA), Local Location-scale Augmentation (LLA), and Saliency-Balancing Fusion (SBF) on Abdominal CT →\to MRI (Abd.) and Cardiac bSSFP →\to LGE (Card.) domain generalization:

    Methods GLA LLA SBF Abd. (%) Card. (%) Avg. (%)
    ERM - - - 77.31 75.99 76.65
    Variant 1 ✓ - - 80.28 74.21 77.25
    Variant 2 - ✓ - 75.43 85.12 80.28
    Variant 3 ✓ ✓ - 85.48 85.55 85.52
    Variant 4 ✓ - ✓ 83.57 82.39 82.98
    Variant 5 - ✓ ✓ 75.42 85.00 80.21
    SLAug ✓ ✓ ✓ 88.63 86.69 87.66

    Key observations include:

    1. GLA vs. LLA Complementarity: GLA alone (Variant 1) only improves the abdominal task (80.28%), while LLA alone (Variant 2) primarily boosts the cardiac task (85.12%). Combining them (Variant 3) achieves strong performance (85.52% avg) by bridging both global and local distribution shifts.
    2. Saliency Source Selection: Saliency maps computed from GLA samples (Variant 4 vs. Variant 1) yield substantial performance gains (+5.73% avg) by identifying valid decision boundary sensitivities, whereas saliency maps computed from LLA samples (Variant 5 vs. Variant 2) degrade performance slightly (-0.07% avg) due to noisy gradients on heavily perturbed representations.
    3. Full System Superiority: Combining GLA, LLA, and SBF achieves the highest average Dice score of 87.66%.
  8. Knowl 8 — Impact of Saliency Guidance on Target Generalization and Source Risk Reduction

    empirical result

    In the Abdominal CT →\to MRI segmentation task, Saliency-Balancing Fusion (SBF) simultaneously enhances cross-domain generalization and reduces empirical source-domain error compared to unguided fusion methods:

    1. Cross-Domain Target Performance (CT →\to MRI):

      • SBF achieves a Dice score of 88.6%.
      • Random map fusion achieves 85.7%.
      • No fusion (directly feeding both augmented images into training) achieves 85.5%.
      • Baseline ERM achieves 77.3%.
    2. In-Domain Source Preservation (CT →\to CT test set):

      • SBF achieves a source Dice score of 93.7%.
      • No fusion achieves 92.9%.
      • Random fusion decreases source Dice score to 92.5% (a 0.4% drop compared to no fusion).
      • Baseline ERM achieves 89.7%.

    SBF improves source performance by 0.8% over no fusion and 4.0% over ERM. Because the theoretical generalization bound depends directly on the source risk σ\sigma, lowering source empirical risk through SBF leads directly to a tighter generalization bound on unseen domains.

  9. Knowl 9 — Experimental Setup for Medical SDG Evaluation

    experimental setup

    The single-source domain generalization (SDG) framework is evaluated across two medical segmentation domains:

    1. Datasets:

      • Abdominal Multi-Organ Cross-Modality: Combined CT images from the MICCAI Multi-Atlas Labeling Beyond Cranial Vault Challenge (13 abdominal organs, evaluating Liver, Right Kidney, Left Kidney, and Spleen) and T2-SPIR MRI images from the CHAOS Challenge.
      • Cardiac Cross-Sequence: Multi-sequence cardiac MRI dataset (MS-CMRSeg challenge), evaluating Left Ventricle Cavity (LVC), Myocardium (MYO), and Right Ventricle Cavity (RVC) across balanced Steady-State Free Precession (bSSFP) and Late Gadolinium Enhancement (LGE) sequences.
    2. Model Architecture & Optimization:

      • Base architecture: 2D U-Net backbone with an EfficientNet-b2 encoder, trained completely from scratch.
      • Optimizer: Adam with initial learning rate 3×10−43 \times 10^{-4} and weight decay 3×10−53 \times 10^{-5}.
      • Learning rate schedule: Constant for the first 50 epochs, followed by linear decay to zero over 1,950 epochs (total 2,000 epochs).
      • Batch size: 32 on a single NVIDIA GeForce RTX 3090 GPU (24GB memory).
      • Saliency grid size hyperparameter gg: Set empirically to g=3g = 3 for the abdominal dataset and g=18g = 18 for the cardiac dataset.
      • Base common augmentations: Affine, Elastic deformation, Brightness, Contrast, Gamma, and Additive Gaussian Noise.

Coverage note — None omitted; all core contributions, assumptions, mathematical definitions, algorithms, theoretical bounds, empirical results, and ablation studies from the paper are represented.

References

  1. 1.Albuquerque, I.; Monteiro, J.; Darvishi, M.; Falk, T. H.; and Mitliagkas, I. 2019. Generalizing to unseen domains via distribution matching. arXiv:1911.00804.
  2. 2.Buzug, T. M. 2011. Computed tomography. In Springer handbook of medical technology, 311–342. Springer.
  3. 3.Chen, C.; Dou, Q.; Chen, H.; Qin, J.; and Heng, P.-A. 2019. Synergistic image and feature adaptation: Towards crossmodality domain adaptation for medical image segmentation. In AAAI Conference on Artificial Intelligence, volume 33, 865–872.
  4. 4.Chen, C.; Qin, C.; Qiu, H.; Ouyang, C.; Wang, S.; Chen, L.; Tarroni, G.; Bai, W.; and Rueckert, D. 2020. Realistic adversarial data augmentation for MR image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention, 667–677. Springer.
  5. 5.Chen, Y.; Ouyang, X.; Zhu, K.; and Agam, G. 2021. Maskbased data augmentation for semi-supervised semantic segmentation. arXiv:2101.10156.
  6. 6.David, S. B.; Lu, T.; Luu, T.; and Pal, D. 2010. Impossibility theorems for domain adaptation. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, volume 9, 129–136.
  7. 7.DeVries, T.; and Taylor, G. W. 2017. Improved regularization of convolutional neural networks with cutout. arXiv:1708.04552.
  8. 8.Forbes, G. B. 2012. Human body composition: growth, aging, nutrition, and activity. Springer Science & Business Media.
  9. 9.Hou, X.; and Zhang, L. 2007. Saliency detection: A spectral residual approach. In 2007 IEEE Conference on computer vision and pattern recognition, 1–8. Ieee.
  10. 10.Huang, Z.; Wang, H.; Xing, E. P.; and Huang, D. 2020. Selfchallenging improves cross-domain generalization. In European Conference on Computer Vision, 124–140. Springer.
  11. 11.Kavur, A. E.; Gezer, N. S.; Barış, M.; Aslan, S.; Conze, P.-H.; Groza, V.; Pham, D. D.; Chatterjee, S.; Ernst, P.; Özkan, S.; et al. 2021. CHAOS challenge-combined (CTMR) healthy abdominal organ segmentation. Medical Image Analysis, 69: 101950.
  12. 12.Kim, J.-H.; Choo, W.; and Song, H. O. 2020. Puzzle mix: Exploiting saliency and local statistics for optimal mixup. In International Conference on Machine Learning, 5275–5285. PMLR.
  13. 13.Kingma, D. P.; and Ba, J. 2014. Adam: A Method for Stochastic Optimization. arXiv:1412.6980.
  14. 14.Landman, B.; Xu, Z.; Igelsias, J.; Styner, M.; Langerak, T.; and Klein, A. 2015. Miccai multi-atlas labeling beyond the cranial vault-workshop and challenge. In Proc. MICCAI Multi-Atlas Labeling Beyond Cranial Vault—Workshop Challenge, volume 5, 12.
  15. 15.Milletari, F.; Navab, N.; and Ahmadi, S.-A. 2016. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In International Conference on 3D Vision, 565–571. IEEE.
  16. 16.Mortenson, M. E. 1999. Mathematics for computer graphics applications. Industrial Press Inc.
  17. 17.Olsson, V.; Tranheden, W.; Pinto, J.; and Svensson, L. 2021. Classmix: Segmentation-based data augmentation for semisupervised learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 1369–1378.
  18. 18.Ouyang, C.; Biffi, C.; Chen, C.; Kart, T.; Qiu, H.; and Rueckert, D. 2020. Self-supervision with superpixels: Training few-shot medical image segmentation without annotation. In European Conference on Computer Vision, 762–780. Springer.
  19. 19.Ouyang, C.; Chen, C.; Li, S.; Li, Z.; Qin, C.; Bai, W.; and Rueckert, D. 2021. Causality-inspired Single-source Domain Generalization for Medical Image Segmentation. arXiv:2111.12525.
  20. 20.Simonyan, K.; Vedaldi, A.; and Zisserman, A. 2014. Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps. In 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Workshop Track Proceedings.
  21. 21.Valanarasu, J. M. J.; Oza, P.; Hacihaliloglu, I.; and Patel, V. M. 2021. Medical Transformer: Gated Axial-Attention for Medical Image Segmentation. In de Bruijne, M.; Cattin, P. C.; Cotin, S.; Padoy, N.; Speidel, S.; Zheng, Y.; and Essert, C., eds., Medical Image Computing and Computer Assisted Intervention, volume 12901, 36–46.
  22. 22.Volk, G.; Müller, S.; Von Bernuth, A.; Hospach, D.; and Bringmann, O. 2019. Towards robust CNN-based object detection through augmentation with synthetic rain variations. In 2019 IEEE Intelligent Transportation Systems Conference, 285–292.
  23. 23.Wang, B.; and Dudek, P. 2014. A fast self-tuning background subtraction algorithm. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 395–398.
  24. 24.Wang, S.; Yu, L.; Li, K.; Yang, X.; Fu, C.-W.; and Heng, P.A. 2020. Dofe: Domain-oriented feature embedding for generalizable fundus image segmentation on unseen datasets. IEEE Transactions on Medical Imaging, 39(12): 4237–4248.
  25. 25.Wei, Y.; Feng, J.; Liang, X.; Cheng, M.-M.; Zhao, Y.; and Yan, S. 2017. Object region mining with adversarial erasing: A simple classification to semantic segmentation approach. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1568–1576.
  26. 26.Xu, Z.; Liu, D.; Yang, J.; Raffel, C.; and Niethammer, M. 2021. Robust and Generalizable Visual Representation Learning via Random Convolutions. In International Conference on Learning Representations.
  27. 27.Yao, K.; Su, Z.; Huang, K.; Yang, X.; Sun, J.; Hussain, A.; and Coenen, F. 2022. A novel 3D unsupervised domain adaptation framework for cross-modality medical image segmentation. IEEE Journal of Biomedical and Health Informatics.
  28. 28.Zhang, J.; Zhang, Y.; and Xu, X. 2021. Objectaug: objectlevel data augmentation for semantic image segmentation. In 2021 International Joint Conference on Neural Networks (IJCNN), 1–8. IEEE.
  29. 29.Zhao, R.; Ouyang, W.; Li, H.; and Wang, X. 2015. Saliency detection by multi-context deep learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1265–1274.
  30. 30.Zhou, K.; Yang, Y.; Qiao, Y.; and Xiang, T. 2021. Domain Generalization with MixStyle. In International Conference on Learning Representations.
  31. 31.Zhou, Z.; Qi, L.; Yang, X.; Ni, D.; and Shi, Y. 2022. Generalizable Cross-modality Medical Image Segmentation via Style Augmentation and Dual Normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20856–20865.
  32. 32.Zhou, Z.; Sodha, V.; Rahman Siddiquee, M. M.; Feng, R.; Tajbakhsh, N.; Gotway, M. B.; and Liang, J. 2019. Models genesis: Generic autodidactic models for 3d medical image analysis. In International conference on medical image computing and computer-assisted intervention, 384–393. Springer.
  33. 33.Zhuang, X.; Xu, J.; Luo, X.; Chen, C.; Ouyang, C.; Rueckert, D.; Campello, V. M.; Lekadir, K.; Vesal, S.; RaviKumar, N.; et al. 2020. Cardiac segmentation on late gadolinium enhancement MRI: a benchmark study from multi-sequence cardiac MR segmentation challenge. arXiv:2006.12434.

Citation

MLA
Su, Z., et al. “Rethinking Data Augmentation for Single-source Domain Generalization in Medical Image Segmentation”. arXiv, 2022, http://arxiv.org/abs/2211.14805v1.
APA
Su, Z., Yao, K., Yang, X., Wang, Q., Sun, J., & Huang, K. (2022). Rethinking Data Augmentation for Single-source Domain Generalization in Medical Image Segmentation. arXiv. http://arxiv.org/abs/2211.14805v1
Chicago
Su, Z., K. Yao, X. Yang, Q. Wang, J. Sun, and K. Huang. 2022. “Rethinking Data Augmentation for Single-source Domain Generalization in Medical Image Segmentation”. arXiv. http://arxiv.org/abs/2211.14805v1.
Harvard
Su, Z. et al. (2022) “Rethinking Data Augmentation for Single-source Domain Generalization in Medical Image Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.14805v1.
Vancouver
1. Su Z, Yao K, Yang X, Wang Q, Sun J, Huang K (2022) Rethinking Data Augmentation for Single-source Domain Generalization in Medical Image Segmentation. arXiv

BibTeX

@article{su2022rethinking,
  title = {Rethinking Data Augmentation for Single-source Domain Generalization in Medical Image Segmentation},
  author = {Su, Zixian and Yao, Kai and Yang, Xi and Wang, Qiufeng and Sun, Jie and Huang, Kaizhu},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.14805v1},
  eprint = {2211.14805}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF