Generalizing to Unseen Domains: A Survey on Domain Generalization

Jindong WangCuiling LanChang LiuYidong OuyangTao Qin

article2021TKDE1,881 citations

Systematizes out-of-distribution machine learning by establishing theoretical foundations for domain generalization, classifying existing methods into a clear three-part taxonomy, and providing standardized benchmark datasets with an open-source codebase for fair evaluation.

Listen

Standard machine learning systems rely on the assumption that operational data will closely match the training data. In real-world deployments, however, models routinely encounter unfamiliar environments, altered image conditions, or evolving language patterns. When these domain shifts occur, model accuracy often drops sharply. Because collecting labeled training data from every potential deployment scenario is prohibitively expensive or impossible, building models that generalize reliably to unseen target environments has become an essential operational challenge.

The article provides a comprehensive survey to define the core principles of domain generalization, categorize existing algorithmic methodologies, review common datasets and applications, and outline open research directions for the field.

The authors conducted a structured literature review across key machine learning subfields, organizing technical approaches into three overarching categories: data manipulation, representation learning, and learning strategies. Data manipulation expands the diversity and volume of training samples through techniques like domain randomization and generative modeling. Representation learning maps inputs into feature spaces that isolate domain-invariant attributes or disentangle shared signals from domain-specific noise. Learning strategies use broader frameworks, such as ensemble methods and meta-learning, which simulate domain shifts during training to improve generalizability.

The survey highlights four primary findings. First, representation learning remains the dominant and theoretically grounded paradigm for domain generalization, utilizing mathematical alignments, kernel methods, and disentanglement frameworks. Second, data manipulation serves as a simple, cost-effective method to boost generalization by generating synthetic diversity, though it lacks firm theoretical guarantees regarding generalization risk. Third, adversarial training methods—while effective in settings where target data is accessible—showed limited and inconsistent performance gains when applied to domain generalization tasks with completely unseen targets. Fourth, while initially concentrated on basic image classification, domain generalization has successfully expanded into high-stakes domains including medical imaging, speech recognition, face anti-spoofing, and industrial fault diagnosis.

These findings have direct strategic implications for reducing the risks, costs, and safety concerns associated with deploying machine learning models in critical production workflows. Rather than continually funding expensive data collection and retraining cycles when deploying to new environments, organizations can deploy generalized models that perform consistently out-of-the-box. Moreover, understanding that adversarial training yields diminishing returns in unseen settings helps technical teams allocate development budgets toward more effective approaches, such as explicit feature alignment and data augmentation.

Organizations should adopt tailored combinations of these complementary approaches, such as pairing lightweight data augmentation with domain-invariant representation learning. Looking forward, technical leaders and researchers should direct efforts toward emerging operational frontiers: continuous domain generalization to handle streaming data without catastrophic forgetting, zero-shot generalization across entirely new task categories, interpretable feature design using causal reasoning, and leveraging large-scale pre-trained models. Because existing theoretical guarantees for data manipulation remain limited and standard algorithms assume static label spaces, decision-makers should exercise caution and conduct thorough stress-testing when deploying systems into dynamically shifting operational environments.

  • Paper: A theory of learning from different domains, Shai Ben-David et al. (2010). Provides the foundational statistical learning theory and discrepancy distance bounds between source and target distributions that underpin the theoretical foundations of domain generalization surveyed in this paper.
  • Paper: Invariant Risk Minimization, Martin Arjovsky et al. (2019). Establishes invariant risk minimization across training environments to discover causal mechanisms robust to unseen test environments, serving as a core learning strategy reviewed in the survey.
  • Paper: Analysis of Representations for Domain Adaptation, Shai Ben-David et al. (2006). Presents early theoretical analysis on learning invariant representations to bound out-of-distribution error, laying the groundwork for domain generalization theory.
  • Book: Domain-Adversarial Training of Neural Networks, Yaroslav Ganin et al. (2016). Introduces adversarial domain alignment using gradient reversal to learn invariant features, a foundational representation learning paradigm extensively covered in the survey.
  • Paper: Moment Matching for Multi-Source Domain Adaptation, Xingchao Peng et al. (2018). Introduces moment matching across multiple source distributions as well as the DomainNet benchmark, directly informing multi-source representation learning techniques and evaluation benchmarks.
  • Paper: Deep CORAL: Correlation Alignment for Deep Domain Adaptation, Baochen Sun et al. (2016). Presents correlation alignment for feature covariance matching, which is a principal discrepancy-based representation learning method discussed in the survey.
  • Paper: A Survey on Transfer Learning, Sinno Jialin Pan et al. (2010). Offers the classic taxonomy and definitions of transfer learning and domain shift settings that contextualize and motivate domain generalization.
  • Paper: WILDS: A Benchmark of in-the-Wild Distribution Shifts, Pang Wei Koh et al. (2020). Constructs benchmark datasets for out-of-distribution shifts arising naturally in the wild, providing the empirical foundation for testing domain generalization algorithms.
  • Paper: Unbiased look at dataset bias, A. Torralba et al. (2011). Exposes the phenomenon of dataset bias and cross-domain degradation in vision systems that domain generalization methodologies are explicitly designed to overcome.
Cover for Generalizing to Unseen Domains: A Survey on Domain Generalization

Abstract

Machine learning systems generally assume that the training and testing distributions are the same. To this end, a key requirement is to develop models that can generalize to unseen distributions. Domain generalization (DG), i.e., out-of-distribution generalization, has attracted increasing interests in recent years. Domain generalization deals with a challenging setting where one or several different but related domain(s) are given, and the goal is to learn a model that can generalize to an unseen test domain. Great progress has been made in the area of domain generalization for years. This paper presents the first review of recent advances in this area. First, we provide a formal definition of domain generalization and discuss several related fields. We then thoroughly review the theories related to domain generalization and carefully analyze the theory behind generalization. We categorize recent algorithms into three classes: data manipulation, representation learning, and learning strategy, and present several popular algorithms in detail for each category. Third, we introduce the commonly used datasets, applications, and our open-sourced codebase for fair evaluation. Finally, we summarize existing literature and present some potential research topics for the future.

Table of Contents

  • I Introduction
  • II Background
  • II-A Formalization of Domain Generalization
  • II-B Related Research Areas
  • III Theory
  • III-A Domain Adaptation
  • III-B Domain Generalization
  • III-B1 Average risk estimation error bound
  • III-B2 Generalization risk bound
  • IV Methodology
  • IV-A Data Manipulation
  • IV-A1 Data augmentation-based DG
  • IV-A2 Data generation-based DG
  • IV-B Representation Learning
  • IV-B1 Domain-invariant representation-based DG
  • IV-B2 Feature disentanglement-based DG
  • IV-C Learning Strategy
  • IV-C1 Ensemble learning-based DG
  • IV-C2 Meta-learning-based DG
  • IV-C3 Gradient operation-based DG
  • IV-C4 Distributionally robust optimization-based DG
  • IV-C5 Self-supervised learning-based DG
  • IV-C6 Other learning strategy for DG
  • V Other Domain Generalization Research Areas
  • V-A Single-source Domain Generalization
  • V-B Semi-supervised Domain Generalization
  • V-C Federated Learning with Domain Generalization
  • V-D Other DG Settings
  • VI Applications
  • VII Datasets, Evaluation, and Benchmark
  • VII-A Datasets
  • VII-B Evaluation
  • VII-C Benchmark
  • VIII Discussion
  • VIII-A Summary of Existing Literature
  • VIII-B Future Research Challenges
  • VIII-B1 Continuous domain generalization
  • VIII-B2 Domain generalization to novel categories
  • VIII-B3 Interpretable domain generalization
  • VIII-B4 Large-scale pre-training/self-learning and DG
  • VIII-B5 Test-time Generalization
  • VIII-B6 Performance evaluation for DG
  • IX Conclusion
  • References

Knowls

  1. Knowl 1 — Formal Definition of Domain and Domain Generalization

    definition

    A domain is defined over a nonempty input space X⊂Rd\mathcal{X} \subset \mathbb{R}^d and an output space Y⊂R\mathcal{Y} \subset \mathbb{R} as a dataset S={(xi,yi)}i=1n∼PXYS = \{(x_i, y_i)\}_{i=1}^n \sim P_{XY} sampled from a joint probability distribution PXYP_{XY} governing the input random variable XX and label random variable YY.

    In domain generalization (DG), a model is provided with MM distinct training (source) domains: Strain={Si∣i=1,…,M}S_{train} = \{S^i \mid i = 1, \dots, M\} where each Si={(xji,yji)}j=1ni∼PXYiS^i = \{(x_j^i, y_j^i)\}_{j=1}^{n_i} \sim P_{XY}^i, and the joint distributions between any two distinct training domains differ, such that PXYi≠PXYjP_{XY}^i \neq P_{XY}^j for all 1≤i≠j≤M1 \le i \neq j \le M.

    The objective of domain generalization is to learn a generalizable predictive function h:X→Yh: \mathcal{X} \to \mathcal{Y} using solely the MM source domains to minimize the expected prediction error on an unseen test domain Stest={(x,y)}∼PXYtestS_{test} = \{(x, y)\} \sim P_{XY}^{test}, where StestS_{test} is inaccessible during training and PXYtest≠PXYiP_{XY}^{test} \neq P_{XY}^i for all i∈{1,…,M}i \in \{1, \dots, M\}: min⁡hE(x,y)∼Stest[ℓ(h(x),y)]\min_h \mathbb{E}_{(x,y) \sim S_{test}} [\ell(h(x), y)] where ℓ(⋅,⋅)\ell(\cdot, \cdot) denotes the loss function and E\mathbb{E} denotes the mathematical expectation.

  2. Knowl 2 — Distinctions Between Domain Generalization and Related Learning Paradigms

    definition

    Domain generalization (DG) differs from other related machine learning paradigms in target domain accessibility, task formulation, and distribution shifts:

    • Multi-Task Learning / Multi-Domain Learning: Jointly optimizes models across multiple related tasks to improve performance on those original tasks; it does not aim to generalize to unobserved test domains.
    • Transfer Learning: Typically pre-trains on a source task and fine-tunes on a target task where target domain data is accessible during training and tasks may differ. In DG, the target domain is unseen during training, and the training and test tasks share the same label space under different distributions.
    • Domain Adaptation (DA): Aims to maximize performance on a specific target domain using labeled source data and accessible (labeled or unlabeled) target domain data during training. DG does not observe any target domain data during training.
    • Meta-Learning: Optimizes a learning-to-learn algorithm across varying tasks. In DG, the underlying classification or regression task is identical, though meta-learning strategies can be applied to simulate domain shifts across source domain splits.
    • Lifelong / Continual Learning: Learns continuously over sequential domains/tasks over time while preventing catastrophic forgetting, accessing target domains at each step without explicitly handling cross-domain distribution shifts prior to observation.
    • Zero-Shot Learning (ZSL): Learns from seen classes to classify unseen object classes using auxiliary semantic descriptions, whereas standard DG classifies samples from seen classes under unseen data distributions.
  3. Knowl 3 — Three-Part Taxonomy of Domain Generalization Methods

    model/method

    Domain generalization methods are structured into three main categories based on their operational level:

    1. Data Manipulation: Augments or generates input samples to increase diversity and volume, enabling models to learn generalizable representations. It encompasses:
      • Data Augmentation: Standard transformations, domain randomization (procedural variations in texture, illumination, context, and camera views), and adversarial data augmentation (perturbing inputs along directions of maximal domain shift while preserving label semantics).
      • Data Generation: Generates novel samples using generative architectures (such as Variational Autoencoders and Generative Adversarial Networks) or interpolation techniques (such as Mixup in input or feature space).
    2. Representation Learning: Constructs representations that are transferable and robust across domains. It encompasses:
      • Domain-Invariant Representation Learning: Minimizes domain distribution divergence using kernel methods (e.g., Domain-Invariant Component Analysis), domain-adversarial training (minimax games between feature extractors and domain discriminators), or explicit feature alignment (matching moments, Maximum Mean Discrepancy, correlation alignments, or Wasserstein distances).
      • Feature Disentanglement: Decomposes latent features into domain-shared (invariant) and domain-specific components through multi-component network structures or generative models.
    3. Learning Strategy: Leverages general learning paradigms to handle domain shifts. It encompasses:
      • Ensemble Learning: Combines domain-specific sub-networks, classifiers, or normalization statistics weighted dynamically for unseen inputs.
      • Meta-Learning: Splits source domains into meta-train and meta-test subsets per iteration to mimic domain shifts during parameter optimization.
      • Other Strategies: Incorporates self-supervised auxiliary tasks (such as solving jigsaw puzzles), episodic training, or self-challenging dominant-feature masking.
  4. Knowl 4 — Optimization Formulation for Data Manipulation-Based Domain Generalization

    equation

    Data manipulation techniques in domain generalization expand the diversity and volume of training samples via a manipulation function mani(⋅)\text{mani}(\cdot). The learning objective is formulated as: min⁡hEx,y[ℓ(h(x),y)]+Ex′,y[ℓ(h(x′),y)]\min_h \mathbb{E}_{x,y} [\ell(h(x), y)] + \mathbb{E}_{x',y} [\ell(h(x'), y)] where:

    • h:X→Yh: \mathcal{X} \to \mathcal{Y} is the predictive hypothesis.
    • (x,y)∼Strain(x, y) \sim S_{train} denotes the original training sample and label pairs drawn from the source domains.
    • x′=mani(x)x' = \text{mani}(x) denotes the manipulated data instance produced by the manipulation operator mani(⋅)\text{mani}(\cdot).
    • ℓ(⋅,⋅)\ell(\cdot, \cdot) is the task loss function.
    • E\mathbb{E} denotes mathematical expectation over the corresponding data distributions.

    The operator mani(⋅)\text{mani}(\cdot) can be instantiated via random geometric/photometric augmentations, domain randomization, adversarial gradient ascent perturbations, generative synthesis (VAEs, GANs), or Mixup linear interpolations.

  5. Knowl 5 — Optimization Formulation for Domain-Invariant Representation Learning

    equation

    In representation learning for domain generalization, the predictive hypothesis hh is decomposed into h=f∘gh = f \circ g, where g:X→Zg: \mathcal{X} \to \mathcal{Z} denotes a feature representation extractor mapping inputs to a latent feature space Z\mathcal{Z}, and f:Z→Yf: \mathcal{Z} \to \mathcal{Y} denotes the classifier mapping latent features to labels. The optimization objective is formulated as: min⁡f,gEx,y[ℓ(f(g(x)),y)]+λℓreg\min_{f, g} \mathbb{E}_{x,y} [\ell(f(g(x)), y)] + \lambda \ell_{reg} where:

    • ℓ(⋅,⋅)\ell(\cdot, \cdot) is the task loss function computed over source domain samples (x,y)∼Strain(x, y) \sim S_{train}.
    • ℓreg\ell_{reg} is a domain-invariance regularization loss that penalizes discrepancies across source domain distributions in the latent space Z\mathcal{Z} (such as Maximum Mean Discrepancy, adversarial domain classification loss, covariance correlation alignment, or optimal transport / Wasserstein distance).
    • λ≥0\lambda \ge 0 is a regularization tradeoff parameter balancing classification empirical risk and domain invariance.
  6. Knowl 6 — Optimization Formulation for Feature Disentanglement in Domain Generalization

    equation

    Feature disentanglement methods decompose the representation of an input xx into a domain-shared feature vector gc(x)g_c(x) and a domain-specific feature vector gs(x)g_s(x). The optimization objective is formulated as: min⁡gc,gs,fEx,y[ℓ(f(gc(x)),y)]+λℓreg+μℓrecon([gc(x),gs(x)],x)\min_{g_c, g_s, f} \mathbb{E}_{x,y} [\ell(f(g_c(x)), y)] + \lambda \ell_{reg} + \mu \ell_{recon}([g_c(x), g_s(x)], x) where:

    • gc:X→Zcg_c: \mathcal{X} \to \mathcal{Z}_c is the domain-shared (domain-invariant) feature extraction mapping.
    • gs:X→Zsg_s: \mathcal{X} \to \mathcal{Z}_s is the domain-specific feature extraction mapping.
    • f:Zc→Yf: \mathcal{Z}_c \to \mathcal{Y} is the classifier operating exclusively on the domain-shared representation gc(x)g_c(x).
    • [gc(x),gs(x)][g_c(x), g_s(x)] denotes the combination or integration of the two feature representations.
    • ℓ(⋅,⋅)\ell(\cdot, \cdot) is the supervised task prediction loss.
    • ℓreg\ell_{reg} is a regularization loss enforcing mutual independence or separation between domain-shared and domain-specific features.
    • ℓrecon\ell_{recon} is a reconstruction loss ensuring that the combined features [gc(x),gs(x)][g_c(x), g_s(x)] retain complete input information without loss.
    • λ≥0\lambda \ge 0 and μ≥0\mu \ge 0 are hyperparameter weights balancing disentanglement regularization and reconstruction fidelity.
  7. Knowl 7 — Optimization Formulation for Meta-Learning in Domain Generalization

    equation

    To implement meta-learning for domain generalization, source domains are dynamically partitioned at each training step into a meta-train split SmtrnS_{mtrn} and a meta-test split SmteS_{mte} to simulate out-of-distribution domain shift. Parameter optimization is formulated as: θ←θ−α∂(ℓ(Smte;θ)+βℓ(Smtrn;ϕ))∂θ\theta \leftarrow \theta - \alpha \frac{\partial (\ell(S_{mte}; \theta) + \beta \ell(S_{mtrn}; \phi))}{\partial \theta} where:

    • θ\theta denotes the global model parameters.
    • Smtrn⊂StrainS_{mtrn} \subset S_{train} and Smte⊂StrainS_{mte} \subset S_{train} denote the meta-train and meta-test domain subsets, respectively (Smtrn∩Smte=∅S_{mtrn} \cap S_{mte} = \emptyset).
    • ϕ\phi denotes the task-specific parameters updated via inner-loop optimization on SmtrnS_{mtrn} starting from θ\theta.
    • ℓ(S;⋅)\ell(S; \cdot) is the empirical loss computed on data split SS.
    • α>0\alpha > 0 is the outer-loop (meta) learning rate.
    • β≥0\beta \ge 0 is the weighting coefficient for the meta-train loss.
  8. Knowl 8 — Ensemble and Auxiliary Learning Strategies for Domain Generalization

    model/method

    Domain generalization exploits architectural ensembles and auxiliary training paradigms to improve robustness without explicit divergence minimization:

    • Domain-Specific Ensembles and Normalization: Maintains dedicated classifiers, network layers, or batch normalization (BN) parameters for each source domain while sharing core representations. Test predictions are computed as a linear combination of domain-specific models, with aggregation weights predicted by an auxiliary domain classifier or by measuring the distance between test instance normalization statistics and domain population statistics.
    • Low-Rank Parameter Decomposition (UndoBias): Decomposes domain-specific model parameters wiw_i for the ii-th domain into a shared base parameter vector w0w_0 and a domain-specific offset Δi\Delta_i, such that wi=w0+Δiw_i = w_0 + \Delta_i.
    • Self-Supervised Pretext Tasks: Trains the feature extractor on auxiliary self-supervised tasks (such as solving jigsaw puzzles on tiled image patches) jointly with task classification to learn spatial and semantic structures that generalize across visual styles.
    • Episodic and Self-Challenging Training: Employs minimax iterative training between classifiers and feature extractors to prepare for worst-case shifts, or iteratively masks the most dominant activated features during training to force the network to utilize diverse, label-correlated secondary features.
  9. Knowl 9 — Standard Benchmarks and Application Domains in Domain Generalization

    model/method

    Domain generalization methods are evaluated across multiple standardized benchmarks and real-world modalities:

    • Standard Image Classification Benchmarks:
      • Rotated MNIST: Hand-written digits rotated at varying angles (e.g., 0∘,15∘,30∘,45∘,60∘,75∘0^\circ, 15^\circ, 30^\circ, 45^\circ, 60^\circ, 75^\circ).
      • PACS: 7 object classes across 4 distinct visual domains: Photo, Art painting, Cartoon, and Sketch.
      • VLCS: 5 object categories sampled across 4 distinct photographic datasets: VOC2007, LabelMe, Caltech-101, and SUN09.
      • Office-Home: 65 categories across 4 visual styles: Art, Clipart, Product, and Real-world.
    • Computer Vision Applications: Satellite image classification under geographical and sensor shifts, semantic segmentation, action recognition, face presentation attack detection (anti-spoofing), and person re-identification (ReID).
    • Natural Language Processing: Cross-domain sentiment classification on the multi-domain Amazon Review dataset, semantic parsing, and stance detection.
    • Healthcare and Engineering: Medical image segmentation, chest X-ray diagnosis, Parkinson's disease detection, and rotating machinery fault diagnostics under sensor distribution shifts.
  10. Knowl 10 — Key Open Research Challenges in Domain Generalization

    limitation

    Four major research challenges remain in domain generalization:

    1. Continuous Domain Generalization: Standard DG assumes a static set of source domains, but practical streaming applications involve continuous, non-stationary data streams. Models must continuously adapt to incoming domains without suffering catastrophic forgetting of earlier domain knowledge.
    2. Domain Generalization to Novel Categories (Zero-Shot DG): Standard DG assumes a closed, identical label space Y\mathcal{Y} across all source and target domains. Generalizing simultaneously to unseen domains and unseen categories (zero-shot domain generalization) requires unifying domain and task transfer mechanisms.
    3. Interpretable Domain Generalization via Causality: Deep DG models often lack semantic interpretability regarding which learned features represent invariant causal mechanisms versus spurious statistical correlations. Structural causal models offer a principled framework to disentangle invariant causal features from context-dependent confounders.
    4. Integration with Large-Scale Pre-Training: Pre-trained foundation models provide rich baseline representations that improve transfer performance. Designing efficient, domain-generalizable pre-training and fine-tuning objectives specifically optimized for out-of-distribution robustness remains an open problem.

Coverage note — None was omitted; all major contributions of this survey paper—including formal definitions, related-field comparisons, categorized methodology formulations, benchmarks, and future challenges—are covered in the knowls.

References

  1. 1.[Anoosheh et al., 2018] Asha Anoosheh, Eirikur Agustsson, Radu Timofte, and Luc Van Gool. Combogan: Unrestrained scalability for image domain translation. In CVPR Workshop, pages 783–790, 2018.
  2. 2.[Balaji et al., 2018] Yogesh Balaji, Swami Sankaranarayanan, and Rama Chellappa. Metareg: Towards domain generalization using meta-regularization. In NeurIPS, pages 998–1008, 2018.
  3. 3.[Ben-David et al., 2007] Shai Ben-David, John Blitzer, Koby Crammer, Fernando Pereira, et al. Analysis of representations for domain adaptation. In NIPS, volume 19, page 137, 2007.
  4. 4.[Biesialska et al., 2020] Magdalena Biesialska, Katarzyna Biesialska, and Marta R Costa-jussa. Continual lifelong learning in natural language processing: A survey. arXiv preprint:2012.09823, 2020.
  5. 5.[Blanchard et al., 2011] Gilles Blanchard, Gyemin Lee, and Clayton Scott. Generalizing from several related classification tasks to a new unlabeled sample. In NeurIPS, pages 2178–2186, 2011.
  6. 6.[Blanchard et al., 2017] Gilles Blanchard, Aniket Anand Deshmukh, Urun Dogan, Gyemin Lee, and Clayton Scott. Domain generalization by marginal transfer learning. arXiv preprint arXiv:1711.07910, 2017.
  7. 7.[Brown et al., 2020] Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020.
  8. 8.[Carlucci et al., 2019] Fabio M Carlucci, Antonio D'Innocente, Silvia Bucci, Barbara Caputo, and Tatiana Tommasi. Domain generalization by solving jigsaw puzzles. In CVPR, 2019.
  9. 9.[Caruana, 1997] Rich Caruana. Multitask learning. Machine learning, 28(1):41–75, 1997.
  10. 10.[Chen et al., 2020] Keyu Chen, Di Zhuang, and J Morris Chang. Discriminative adversarial domain generalization with meta-learning based cross-domain validation. arXiv preprint arXiv:2011.00444, 2020.
  11. 11.[Deshmukh et al., 2019] Aniket Anand Deshmukh, Yunwen Lei, Srinagesh Sharma, Urun Dogan, James W Cutler, and Clayton Scott. A generalization error bound for multiclass domain generalization. arXiv:1905.10392, 2019.
  12. 12.[Devlin et al., 2018] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  13. 13.[Ding and Fu, 2017] Zhengming Ding and Yun Fu. Deep domain generalization with structured low-rank constraint. IEEE TIP, 27(1):304–313, 2017.
  14. 14.[Dou et al., 2019] Qi Dou, Daniel Coelho de Castro, Konstantinos Kamnitsas, and Ben Glocker. Domain generalization via model-agnostic learning of semantic features. In NeurIPS, 2019.
  15. 15.[Du and others, 2020] Yingjun Du et al. Learning to learn with variational information bottleneck for domain generalization. In ECCV, 2020.
  16. 16.[D'Innocente and Caputo, 2018] Antonio D'Innocente and Barbara Caputo. Domain generalization with domain-specific aggregation modules. In German Conference on Pattern Recognition, pages 187–198. Springer, 2018.
  17. 17.[Erfani et al., 2016] Sarah Erfani, Mahsa Baktashmotlagh, Masoud Moshtaghi, Vinh Nguyen, Christopher Leckie, James Bailey, and Ramamohanarao Kotagiri. Robust domain generalisation by enforcing distribution invariance. In AAAI, pages 1455–1461, 2016.
  18. 18.[Fang et al., 2013] Chen Fang, Ye Xu, and Daniel N Rockmore. Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias. In ICCV, 2013.
  19. 19.[Finn et al., 2017] Chelsea Finn, P. Abbeel, and S. Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML, 2017.
  20. 20.[Gan et al., 2016] Chuang Gan, Tianbao Yang, and Boqing Gong. Learning attributes equals multi-source domain generalization. In CVPR, pages 87–97, 2016.
  21. 21.[Ganin and Lempitsky, 2015] Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In ICML, pages 1180–1189, 2015.
  22. 22.[Ganin et al., 2016] Yaroslav Ganin, E. Ustinova, Hana Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky. Domain-adversarial training of neural networks. J. Mach. Learn. Res., 17:59:1–59:35, 2016.
  23. 23.[Garg et al., 2020] Vikas K Garg, Adam Kalai, Katrina Ligett, and Zhiwei Steven Wu. Learn to expect the unexpected: Probably approximately correct domain generalization. arXiv preprint arXiv:2002.05660, 2020.
  24. 24.[Ghifary et al., 2015] Muhammad Ghifary, W. Kleijn, M. Zhang, and D. Balduzzi. Domain generalization for object recognition with multi-task autoencoders. ICCV, pages 2551–2559, 2015.
  25. 25.[Ghifary et al., 2016] Muhammad Ghifary, David Balduzzi, W Bastiaan Kleijn, and Mengjie Zhang. Scatter component analysis: A unified framework for domain adaptation and domain generalization. IEEE TPAMI, 39(7):1414–1430, 2016.
  26. 26.[Gong et al., 2019] Rui Gong, Wen Li, Yuhua Chen, and Luc Van Gool. Dlow: Domain flow for adaptation and generalization. In CVPR, pages 2477–2486, 2019.
  27. 27.[Goodfellow et al., 2014] Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. In NIPS, 2014.
  28. 28.[Gretton et al., 2012] Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Scholkopf, and Alexander Smola. A kernel two-sample test. Journal of Machine Learning Research, 13(1):723–773, 2012.
  29. 29.[Grubinger et al., 2015] Thomas Grubinger, Adriana Birlutiu, Holger Schoner, Thomas Natschlager, and Tom Heskes. Domain generalization based on transfer component analysis. In International Work-Conference on Artificial Neural Networks, pages 325–334, 2015.
  30. 30.[Hu et al., 2019] Shoubo Hu, Kun Zhang, Zhitang Chen, and Laiwan Chan. Domain generalization via multidomain discriminant analysis. In UAI, volume 35, 2019.
  31. 31.[Huang et al., 2020] Zeyi Huang, Haohan Wang, Eric P Xing, and Dong Huang. Self-challenging improves cross-domain generalization. In ECCV, volume 2, 2020.
  32. 32.[Ilse et al., 2020] Maximilian Ilse, Jakub M Tomczak, Christos Louizos, and Max Welling. Diva: Domain invariant variational autoencoders. In Proceedings of the Third Conference on Medical Imaging with Deep Learning, 2020.
  33. 33.[Jia et al., 2020] Yunpei Jia, Jie Zhang, Shiguang Shan, and Xilin Chen. Single-side domain generalization for face anti-spoofing. In CVPR, pages 8484–8493, 2020.
  34. 34.[Jin et al., 2020a] Xin Jin, Cuiling Lan, Wenjun Zeng, and Zhibo Chen. Feature alignment and restoration for domain generalization and adaptation. In NeurIPS, 2020.
  35. 35.[Jin et al., 2020b] Xin Jin, Cuiling Lan, Wenjun Zeng, Zhibo Chen, and Li Zhang. Style normalization and restitution for generalizable person re-identification. In CVPR, pages 3143–3152, 2020.
  36. 36.[Jin et al., 2021] Xin Jin, Cuiling Lan, Wenjun Zeng, and Zhibo Chen. Style normalization and restitution for domain generalization and adaptation. arXiv preprint arXiv:2101.00588, 2021.
  37. 37.[Jing and Tian, 2020] Longlong Jing and Yingli Tian. Self-supervised visual feature learning with deep neural networks: A survey. IEEE TPAMI, 2020.
  38. 38.[Khirodkar et al., 2019] Rawal Khirodkar, Donghyun Yoo, and Kris Kitani. Domain randomization for scene-specific car detection and pose estimation. In WACV, 2019.
  39. 39.[Khosla et al., 2012] Aditya Khosla, Tinghui Zhou, Tomasz Malisiewicz, Alexei A Efros, and Antonio Torralba. Undoing the damage of dataset bias. In ECCV, 2012.
  40. 40.[Kingma and Welling, 2013] Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv:1312.6114, 2013.
  41. 41.[Li et al., 2017a] Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales. Deeper, broader and artier domain generalization. In ICCV, pages 5542–5550, 2017.
  42. 42.[Li et al., 2017b] Wen Li, Zheng Xu, Dong Xu, Dengxin Dai, and Luc Van Gool. Domain generalization and adaptation using low rank exemplar svms. IEEE TPAMI, 40(5):1114–1127, 2017.
  43. 43.[Li et al., 2018a] Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales. Learning to generalize: Meta-learning for domain generalization. In AAAI, 2018.
  44. 44.[Li et al., 2018b] Haoliang Li, Sinno Jialin Pan, S. Wang, and A. Kot. Domain generalization with adversarial feature learning. In CVPR, pages 5400–5409, 2018.
  45. 45.[Li et al., 2018c] Ya Li, Mingming Gong, Xinmei Tian, Tongliang Liu, and Dacheng Tao. Domain generalization via conditional invariant representations. In AAAI, 2018.
  46. 46.[Li et al., 2018d] Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao. Deep domain generalization via conditional invariant adversarial networks. In ECCV, pages 624–639, 2018.
  47. 47.[Li et al., 2019a] Da Li, Jianshu Zhang, Yongxin Yang, Cong Liu, Yi-Zhe Song, and Timothy M Hospedales. Episodic training for domain generalization. In CVPR, pages 1446–1455, 2019.
  48. 48.[Li et al., 2019b] Yiying Li, Yongxin Yang, Wei Zhou, and Timothy M Hospedales. Feature-critic networks for heterogeneous domain generalization. In ICML, 2019.
  49. 49.[Li et al., 2020] Xiang Li, Wei Zhang, Hui Ma, Zhong Luo, and Xu Li. Domain generalization in rotating machinery fault diagnostics using deep neural networks. Neurocomputing, 403:409–420, 2020.
  50. 50.[Liu et al., 2020] Chang Liu, Xinwei Sun, Jindong Wang, Tao Li, Tao Qin, Wei Chen, and Tie-Yan Liu. Learning causal semantic representation for out-of-distribution prediction. arXiv preprint arXiv:2011.01681, 2020.
  51. 51.[Mahajan et al., 2020] Divyat Mahajan, S. Tople, and Amit Sharma. Domain generalization using causal matching. In ICML Workshop, 2020.
  52. 52.[Mancini et al., 2018] Massimiliano Mancini, Samuel Rota Bulo, Barbara Caputo, and Elisa Ricci. Best sources forward: domain generalization through source-specific nets. In ICIP, pages 1353–1357, 2018.
  53. 53.[Maniyar et al., 2020] Udit Maniyar, Aniket Anand Deshmukh, Urun Dogan, Vineeth N Balasubramanian, et al. Zero shot domain generalization. arXiv preprint arXiv:2008.07443, 2020.
  54. 54.[Motiian and others, 2017] Saeid Motiian et al. Unified deep supervised domain adaptation and generalization. In ICCV, 2017.
  55. 55.[Muandet et al., 2013] Krikamol Muandet, David Balduzzi, and Bernhard Scholkopf. Domain generalization via invariant feature representation. In ICML, 2013.
  56. 56.[Niu et al., 2015] Li Niu, Wen Li, and Dong Xu. Multi-view domain generalization for visual recognition. In ICCV, pages 4193–4201, 2015.
  57. 57.[Pan and Yang, 2010] S. Pan and Qiang Yang. A survey on transfer learning. IEEE TKDE, 22:1345–1359, 2010.
  58. 58.[Pan et al., 2011] Sinno Jialin Pan, I. Tsang, James T. Kwok, and Qiang Yang. Domain adaptation via transfer component analysis. IEEE TNN, 22:199–210, 2011.
  59. 59.[Patel et al., 2015] Vishal M Patel, Raghuraman Gopalan, Ruonan Li, and Rama Chellappa. Visual domain adaptation: A survey of recent advances. IEEE signal process mag, 32(3):53–69, 2015.
  60. 60.[Peng et al., 2019] Xingchao Peng, Zijun Huang, Ximeng Sun, and Kate Saenko. Domain agnostic learning with disentangled representations. In ICML, 2019.
  61. 61.[Piratla et al., 2020] Vihari Piratla, Praneeth Netrapalli, and Sunita Sarawagi. Efficient domain generalization via common-specific low-rank decomposition. In ICML, pages 7728–7738, 2020.
  62. 62.[Prakash and others, 2019] Aayush Prakash et al. Structured domain randomization: Bridging the reality gap by context-aware synthetic data. In ICRA, pages 7249–7255, 2019.
  63. 63.[Qiao et al., 2020] Fengchun Qiao, Long Zhao, and Xi Peng. Learning to learn single domain generalization. In CVPR, pages 12556–12565, 2020.
  64. 64.[Rahman et al., 2019] Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktashmotlagh, and Sridha Sridharan. Multi-component image translation for deep domain generalization. In WACV, pages 579–588. IEEE, 2019.
  65. 65.[Rahman et al., 2020] Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktashmotlagh, and Sridha Sridharan. Correlation-aware adversarial domain adaptation and generalization. Pattern Recognition, 100:107124, 2020.
  66. 66.[Ryu et al., 2019] Jongbin Ryu, Gitaek Kwon, Ming-Hsuan Yang, and Jongwoo Lim. Generalized convolutional forest networks for domain generalization and visual recognition. In ICLR, 2019.
  67. 67.[Segu` et al., 2020] Mattia Segu, Alessio Tonioni, and Federico Tombari. Batch normalization embeddings for deep domain generalization. arXiv:2011.12672, 2020.
  68. 68.[Shankar et al., 2018] Shiv Shankar, Vihari Piratla, Soumen Chakrabarti, Siddhartha Chaudhuri, Preethi Jyothi, and Sunita Sarawagi. Generalizing across domains via crossgradient training. In ICLR, 2018.
  69. 69.[Shao et al., 2019] Rui Shao, Xiangyuan Lan, Jiawei Li, and Pong C Yuen. Multi-adversarial discriminative deep domain generalization for face presentation attack detection. In CVPR, pages 10023–10031, 2019.
  70. 70.[Sharifi-Noghabi et al., 2020] Hossein Sharifi-Noghabi, Hossein Asghari, Nazanin Mehrasa, and Martin Ester. Domain generalization via semi-supervised meta learning. arXiv:2009.12658, 2020.
  71. 71.[Somavarapu et al., 2020] Nathan Somavarapu, Chih-Yao Ma, and Zsolt Kira. Frustratingly simple domain generalization via image stylization. arXiv:2006.11207, 2020.
  72. 72.[Sun and Saenko, 2016] Baochen Sun and Kate Saenko. Deep coral: Correlation alignment for deep domain adaptation. In ECCV, pages 443–450, 2016.
  73. 73.[Tolstikhin et al., 2017] Ilya Tolstikhin, Olivier Bousquet, Sylvain Gelly, and Bernhard Schoelkopf. Wasserstein auto-encoders. arXiv preprint arXiv:1711.01558, 2017.
  74. 74.[Vanschoren, 2018] Joaquin Vanschoren. Meta-learning: A survey. arXiv preprint arXiv:1810.03548, 2018.
  75. 75.[Venkateswara et al., 2017] Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and S. Panchanathan. Deep hashing network for unsupervised domain adaptation. CVPR, pages 5385–5394, 2017.
  76. 76.[Vilalta and Drissi, 2002] Ricardo Vilalta and Youssef Drissi. A perspective view and survey of meta-learning. Artificial intelligence review, 18(2):77–95, 2002.
  77. 77.[Volpi et al., 2018] Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese. Generalizing to unseen domains via adversarial data augmentation. In NeurIPS, pages 5334–5344, 2018.
  78. 78.[Wang et al., 2018] Jindong Wang, Wenjie Feng, Yiqiang Chen, Han Yu, Meiyu Huang, and Philip S Yu. Visual domain adaptation with manifold embedded distribution alignment. In ACMMM, pages 402–410, 2018.
  79. 79.[Wang et al., 2019] Jindong Wang, Yiqiang Chen, Han Yu, Meiyu Huang, and Qiang Yang. Easy transfer learning by exploiting intra-domain structures. In ICME, pages 1210–1215, 2019.
  80. 80.[Wang et al., 2020a] Bailin Wang, Mirella Lapata, and Ivan Titov. Meta-learning for domain generalization in semantic parsing. arXiv:2010.11988, 2020.
  81. 81.[Wang et al., 2020b] Hao Wang, Hao He, and Dina Katabi. Continuously indexed domain adaptation. arXiv preprint arXiv:2007.01807, 2020.
  82. 82.[Wang et al., 2020c] Jindong Wang, Yiqiang Chen, Wenjie Feng, Han Yu, Meiyu Huang, and Qiang Yang. Transfer learning with dynamic distribution adaptation. ACM TIST, 11(1):1–25, 2020.
  83. 83.[Wang et al., 2020d] Wenhao Wang, Shengcai Liao, Fang Zhao, Cuicui Kang, and Ling Shao. Domainmix: Learning generalizable person re-identification without human annotations. arXiv:2011.11953, 2020.
  84. 84.[Wang et al., 2020e] Yufei Wang, Haoliang Li, and Alex C Kot. Heterogeneous domain generalization via domain mixup. In ICASSP, pages 3622–3626, 2020.
  85. 85.[Wang et al., 2020f] Zhen Wang, Qiansheng Wang, Chengguo Lv, Xue Cao, and Guohong Fu. Unseen target stance detection with adversarial domain generalization. In IJCNN, pages 1–8, 2020.
  86. 86.[Wang, 2018] Jindong Wang. Everything about transfer learning and domain adapation. https://github.com/jindongwang/transferlearning, 2018.
  87. 87.[Zhang et al., 2018] Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. In ICLR, 2018.
  88. 88.[Zhao et al., 2020a] Shanshan Zhao, Mingming Gong, Tongliang Liu, Huan Fu, and Dacheng Tao. Domain generalization via entropy regularization. In NeurIPS, volume 33, 2020.
  89. 89.[Zhao et al., 2020b] Yuyang Zhao, Zhun Zhong, Fengxiang Yang, Zhiming Luo, Yaojin Lin, Shaozi Li, and Nicu Sebe. Learning to generalize unseen domains via memory-based multi-source meta-learning for person re-identification. arXiv preprint arXiv:2012.00417, 2020.
  90. 90.[Zhou et al., 2020a] Fan Zhou, Zhuqing Jiang, Changjian Shui, B. Wang, and B. Chaib-draa. Domain generalization with optimal transport and metric learning. ArXiv, abs/2007.10573, 2020.
  91. 91.[Zhou et al., 2020b] Kaiyang Zhou, Yongxin Yang, Timothy Hospedales, and Tao Xiang. Deep domain-adversarial image generation for domain generalisation. In AAAI, 2020.
  92. 92.[Zhou et al., 2020c] Kaiyang Zhou, Yongxin Yang, Timothy M. Hospedales, and Tao Xiang. Learning to generate novel domains for domain generalization. In ECCV, 2020.
  93. 93.[Zhou et al., 2021] Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang. Domain generalization with mixstyle. In ICLR, 2021.
  94. 94.[Zhu et al., 2020] Yongchun Zhu, Fuzhen Zhuang, Jindong Wang, Guolin Ke, Jingwu Chen, Jiang Bian, Hui Xiong, and Qing He. Deep subdomain adaptation network for image classification. IEEE TNNLS, 2020.
  95. 95.[Zhuang et al., 2020] Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He. A comprehensive survey on transfer learning. Proceedings of the IEEE, 109(1):43–76, 2020.

Citation

MLA
Wang, J., et al. “Generalizing to Unseen Domains: A Survey on Domain Generalization”. IEEE Transactions on Knowledge and Data Engineering, 2022, pp. 1–1, https://doi.org/10.1109/TKDE.2022.3178128.
APA
Wang, J., Lan, C., Liu, C., Ouyang, Y., Qin, T., Lu, W., Chen, Y., Zeng, W., & Yu, P. (2022). Generalizing to Unseen Domains: A Survey on Domain Generalization. IEEE Transactions on Knowledge and Data Engineering, 1–1. https://doi.org/10.1109/TKDE.2022.3178128
Chicago
Wang, J., C. Lan, C. Liu, et al. 2022. “Generalizing to Unseen Domains: A Survey on Domain Generalization”. IEEE Transactions on Knowledge and Data Engineering, 1–1. https://doi.org/10.1109/TKDE.2022.3178128.
Harvard
Wang, J. et al. (2022) “Generalizing to Unseen Domains: A Survey on Domain Generalization”, IEEE Transactions on Knowledge and Data Engineering, pp. 1–1. Available at: https://doi.org/10.1109/TKDE.2022.3178128.
Vancouver
1. Wang J, Lan C, Liu C, Ouyang Y, Qin T, Lu W, Chen Y, Zeng W, Yu P (2022) Generalizing to Unseen Domains: A Survey on Domain Generalization. IEEE Transactions on Knowledge and Data Engineering 1–1

BibTeX

@article{Wang_2022, title={Generalizing to Unseen Domains: A Survey on Domain Generalization}, ISSN={2326-3865}, url={http://dx.doi.org/10.1109/TKDE.2022.3178128}, DOI={10.1109/tkde.2022.3178128}, journal={IEEE Transactions on Knowledge and Data Engineering}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Wang, Jindong and Lan, Cuiling and Liu, Chang and Ouyang, Yidong and Qin, Tao and Lu, Wang and Chen, Yiqiang and Zeng, Wenjun and Yu, Philip}, year={2022}, pages={1–1} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: https://creativecommons.org/licenses/by/4.0/