Improving Out-of-Distribution Robustness via Selective Augmentation
Huaxiu YaoYu WangSai LiLinjun ZhangWeixin LiangJames ZouChelsea Finn
Develops LISA, a selective data interpolation method that pairs samples sharing labels across different domains or domains across different labels, achieving superior out-of-distribution generalization and theoretically reduced worst-group error across diverse shift benchmarks.
Machine learning systems frequently experience severe performance drops when deployed in real-world settings where operating conditions deviate from training conditions. This challenge typically arises from subpopulation shifts, such as class imbalances where models exploit misleading spurious correlations, or domain shifts, where models encounter entirely new environments like new hospitals or camera sensors. Standard techniques attempt to enforce stability across environments by adding complex mathematical penalties during training, but these regularizers often restrict model flexibility, prove difficult to optimize, and behave inconsistently across different applications.
The article evaluates a straightforward data augmentation method called LISA (Learning Invariant Predictors with Selective Augmentation) designed to improve model robustness under distribution shifts. The objective is to demonstrate that selectively interpolating data samples enables neural networks to ignore misleading domain attributes and rely solely on robust predictive features, without requiring artificial constraints on the model's internal structure.
To evaluate this framework, the authors conducted empirical benchmarks across nine standard image and natural language datasets, including the multi-domain WILDS and MetaShift benchmarks, alongside theoretical analysis using linear discriminant models. The approach linearly blends input features and target labels under two distinct strategies: intra-label augmentation, which pairs data points sharing the same class but coming from different domains to cancel domain-specific noise, and intra-domain augmentation, which pairs samples within the same domain possessing different labels to force the model to look beyond environmental artifacts.
The findings establish three primary outcomes. First, LISA consistently matched or outperformed seven leading robustness baselines across all nine datasets. For instance, in worst-case subpopulation accuracy, LISA achieved 89.3% on facial attribute recognition and 72.6% on toxic comment detection, clearly surpassing prior invariant-learning baselines. Second, in domain-shift environments, LISA raised accuracy on medical tumor identification to 77.1% compared to 70.3% for standard training, while also delivering leading performance on satellite imagery and cellular perturbation benchmarks. Third, controlled tests confirmed that these improvements stemmed directly from neutralizing spurious correlations rather than generic data expansion, showing the largest relative performance gains when the gap between training and testing distributions was widest.
These results carry significant practical implications for deploying reliable machine learning in safety-critical and high-risk domains such as medical diagnosis and content moderation. LISA mitigates the risk of catastrophic failures caused by environmental shifts while avoiding the hyperparameter fragility and training instability associated with prior regularization algorithms. By simply adjusting how training batches are constructed, teams can achieve superior reliability with minimal overhead.
Organizations developing machine learning models for shifting environments should consider integrating selective data blending into existing training pipelines as an alternative to complex loss regularizers. For tasks characterized by strong spurious correlations, practitioners should balance intra-label and intra-domain strategies, whereas intra-label mixing alone provides the greatest benefit when domain shifts are natural and widespread.
A primary operational limitation of LISA is its reliance on pairing samples with identical labels, which restricts its immediate utility in complex tasks such as object detection or generative modeling. Additionally, the approach assumes access to domain or group metadata during training, although empirical evidence indicates intra-label blending can succeed even without explicit domain tags when domain correlations are weak. Confidence in the underlying method remains high due to consistent empirical gains paired with formal theoretical validation.
- Paper: mixup: Beyond Empirical Risk Minimization, Hongyi Zhang et al. (2017). Introduces the foundational mixup data interpolation principle that LISA directly adapts into a selective cross-domain and cross-label augmentation strategy.
- Paper: Invariant Risk Minimization, Martin Arjovsky et al. (2019). Establishes the invariant prediction framework across multiple environments that LISA aims to achieve without imposing explicit representation penalties.
- Paper: Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization, Shiori Sagawa et al. (2019). Formalizes the problem of subpopulation shifts and worst-group error minimization that serves as a primary evaluation target and theoretical baseline in LISA.
- Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). Introduces the benchmark suite and evaluation standards for real-world subpopulation and domain shifts used extensively to assess LISA.
- Paper: In Search of Lost Domain Generalization, Ishaan Gulrajani et al. (2020). Demonstrates the critical importance of strong empirical baselines and data augmentation in domain generalization benchmarks against which LISA is compared.
- Paper: Manifold Mixup: Better Representations by Interpolating Hidden States, Vikas Verma et al. (2018). Extends mixup interpolation into hidden representation spaces, motivating representation-level mixing dynamics relevant to invariant predictor design.
- Paper: Return of Frustratingly Easy Domain Adaptation, Baochen Sun et al. (2015). Presents a foundational correlation-alignment approach for domain shift mitigation that LISA compares against in its empirical evaluation.
- Paper: Generalizing to Unseen Domains: A Survey on Domain Generalization, Jindong Wang et al. (2021). Provides a comprehensive taxonomy of domain generalization methods, contextualizing representation regularization versus data manipulation strategies like LISA.
- Paper: Change is Hard: A Closer Look at Subpopulation Shift, Yuzhe Yang et al. (2023). Provides an in-depth benchmark and diagnostic taxonomy that systematically tests modern subpopulation-shift algorithms across diverse real-world shift mechanisms.
- Paper: Using Mixup as a Regularizer Can Surprisingly Improve Accuracy & Out-of-Distribution Robustness, Francesco Pinto et al. (2022). Investigates regularized mixup formulations to simultaneously optimize out-of-distribution robustness and in-distribution accuracy.
- Paper: Rethinking Data Augmentation for Single-Source Domain Generalization in Medical Image Segmentation, Zixian Su et al. (2023). Develops targeted saliency-guided augmentation strategies for single-source domain generalization in complex medical segmentation tasks.
- Paper: Improved Test-Time Adaptation for Domain Generalization, Liang Chen et al. (2023). Complements training-time domain generalization by introducing adaptive consistency objectives and lightweight parameter updates during test time.
- Paper: Balancing Discriminability and Transferability for Source-Free Domain Adaptation, Jogendra Nath Kundu et al. (2022). Applies intermediate mixup domain concepts to source-free domain adaptation settings to preserve discriminability and transferability.
