MADG: Margin-based Adversarial Learning for Domain Generalization

Aveen DayalVimal K. B.Linga Reddy CenkeramaddiC. Krishna MohanAbhinav KumarVineeth N. Balasubramanian

article2023NeurIPS102 citations

Proposes a margin-based adversarial learning framework for domain generalization backed by Rademacher complexity bounds that achieves consistent state-of-the-art performance across standard DomainBed benchmarks.

Listen

Deep learning models frequently suffer severe performance drops when deployed in real-world environments because operational data often differs from training data. While traditional techniques require access to target operational data during model development, many practical applications demand models that can generalize directly to entirely new, unseen target environments without prior exposure. Most existing adversarial learning approaches attempt this by minimizing zero-one loss divergence metrics between training sources, but these metrics yield loose theoretical error bounds and are difficult to optimize efficiently.

The main objective of the article is to develop a theoretically grounded adversarial domain generalization framework using margin-based discrepancy metrics and to demonstrate its superior accuracy and consistency on unseen target environments.

To achieve this, the article establishes a mathematical error bound for unseen target domains by combining margin loss, the convex hull of source distributions, and statistical complexity measures. Guided by this theoretical foundation, the authors design Margin-based Adversarial learning for Domain Generalization (MADG). The algorithm utilizes a primary task classifier alongside multiple auxiliary classifiers that measure pairwise margin disparity discrepancies across source domains, simultaneously training a shared feature extractor to align distributions. The framework was evaluated across five standard image classification benchmark datasets—VLCS, PACS, OfficeHome, TerraIncognita, and DomainNet—using standard evaluation protocols with pre-trained ResNet-50 backbones.

The key findings demonstrate that MADG consistently outperforms existing baseline approaches. First, MADG achieves the highest overall average accuracy of 66.0% across all five standard benchmark datasets, surpassing strong empirical baselines such as standard Empirical Risk Minimization (65.0%) and prior adversarial methods such as DANN (63.8%) and CDANN (64.1%). Second, MADG achieves the best median rank and significantly reduces performance divergence from top-performing models across datasets, showing a geometric mean difference of only 0.3 compared to 0.7–5.3 in prior methods. Third, MADG provides marked performance gains on specific datasets, such as achieving 71.3% accuracy on OfficeHome (an improvement of roughly 3% over most baselines) and 65.6% on Colored MNIST compared to 57.8% for standard empirical risk minimization. Finally, empirical analyses confirm that computing margin discrepancies across all source pairs is essential, as reducing the number of pairwise classifiers degrades accuracy.

These results demonstrate that margin-based alignment offers a more informative and reliable path toward out-of-distribution robustness. For leadership and engineering teams deploying machine learning in safety-critical and high-variability operations, adopting margin-based adversarial alignment reduces the risk of model failure on new operational data. Importantly, MADG achieves these gains without substantial operational overhead, maintaining hardware memory and training time demands comparable to standard baseline approaches.

Organizations developing computer vision and machine learning models for unpredictable target environments should consider adopting margin-based adversarial loss functions over legacy zero-one divergence objectives. When implementing this methodology, technical teams should optimize pairwise discrepancy across all available source datasets and calibrate the margin parameter, as moderate margin values yield the best balance between decision boundary tightness and model stability. Further pilot evaluations on specific proprietary target workflows are recommended before enterprise-scale deployment to confirm performance in distinct operational settings.

The primary limitation of the study is that its theoretical guarantees depend on the degree to which an unseen target distribution relates to the convex hull of the training sources; if a target domain is exceptionally distant or completely dissimilar from the available training domains, performance guarantees weaken. Within the benchmark image recognition domains tested, confidence in the reported improvements remains high due to consistent multi-trial testing across diverse datasets.

arXiv: 2311.08503

No sufficiently relevant recommendations were found.

Cover for MADG: Margin-based Adversarial Learning for Domain Generalization

Abstract

Domain Generalization (DG) techniques have emerged as a popular approach to address the challenges of domain shift in Deep Learning (DL), with the goal of generalizing well to the target domain unseen during the training. In recent years, numerous methods have been proposed to address the DG setting, among which one popular approach is the adversarial learning-based methodology. The main idea behind adversarial DG methods is to learn domain-invariant features by minimizing a discrepancy metric. However, most adversarial DG methods use 0-1 loss based HΔH divergence metric. In contrast, the margin loss-based discrepancy metric has the following advantages: more informative, tighter, practical, and efficiently optimizable. To mitigate this gap, this work proposes a novel adversarial learning DG algorithm, MADG, motivated by a margin loss-based discrepancy metric. The proposed MADG model learns domain-invariant features across all source domains and uses adversarial training to generalize well to the unseen target domain. We also provide a theoretical analysis of the proposed MADG model based on the unseen target error bound. Specifically, we construct the link between the source and unseen domains in the real-valued hypothesis space and derive the generalization bound using margin loss and Rademacher complexity. We extensively experiment with the MADG model on popular real-world DG datasets, VLCS, PACS, OfficeHome, DomainNet, and TerraIncognita. We evaluate the proposed algorithm on DomainBed's benchmark and observe consistent performance across all the datasets.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminaries
  • 4 Margin-based Approach to Domain Generalization: Theory
  • 5 MADG: Methodology
  • 6 Experiments
  • 7 More Empirical Analysis and Ablation Studies
  • 8 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Margin disparity discrepancy for domain comparison

    definition

    MADG operates in a multiclass setting with input space X\mathcal X, label space Y\mathcal Y of size KK, scoring-function class F\mathcal F where each f:X→RKf:\mathcal X\to\mathbb R^K, and classifier hf(x)=arg⁡max⁡y∈Yf(x,y)h_f(x)=\arg\max_{y\in\mathcal Y}f(x,y). For margin threshold ρ>0\rho>0, the margin of ff on labeled example (x,y)(x,y) is

    ρf(x,y)=12(f(x,y)−max⁡y′≠yf(x,y′)).\rho_f(x,y)=\frac{1}{2}\left(f(x,y)-\max_{y'\neq y}f(x,y')\right).

    The paper uses the ramp margin loss

    Φρ(t)={0,ρ≤t,1−t/ρ,0≤t≤ρ,1,t≤0,\Phi_\rho(t)= \begin{cases} 0,&\rho\le t,\\ 1-t/\rho,&0\le t\le\rho,\\ 1,&t\le 0, \end{cases}

    and margin error err⁡D(ρ)(f)=E(x,y)∼D[Φρ(ρf(x,y))]\operatorname{err}^{(\rho)}_D(f)=\mathbb E_{(x,y)\sim D}[\Phi_\rho(\rho_f(x,y))]. The margin disparity of a scoring function f′f' relative to reference classifier ff on domain DD is

    disp⁡D(ρ)(f′,f)=E(x,y)∼D[Φρ(ρf′(x,hf(x)))].\operatorname{disp}^{(\rho)}_D(f',f)=\mathbb E_{(x,y)\sim D}\left[\Phi_\rho\bigl(\rho_{f'}(x,h_f(x))\bigr)\right].

    For two domains Di,DkD_i,D_k, the margin disparity discrepancy (MDD) is the classifier-aware quantity

    df,F(ρ)(Di,Dk)=sup⁡f′∈F(disp⁡Dk(ρ)(f′,f)−disp⁡Di(ρ)(f′,f)).d^{(\rho)}_{f,\mathcal F}(D_i,D_k)=\sup_{f'\in\mathcal F}\left(\operatorname{disp}^{(\rho)}_{D_k}(f',f)-\operatorname{disp}^{(\rho)}_{D_i}(f',f)\right).

    Its empirical version replaces each expectation by an average over the corresponding sample. Unlike a 00–11 disagreement, MDD compares real-valued margins, keeps ff fixed while optimizing over f′f', and uses a smooth thresholded loss. The margin geometry illustrated on page 4 shows that the margin-based agreement region is smaller than the corresponding 00–11 agreement region.

  2. Knowl 2 — Unseen-domain error bounds from source-domain mixtures

    theoretical result

    Let DS1,…,DSNsD_{S_1},\ldots,D_{S_{N_s}} be labeled source distributions and let DUD_U be an unseen target distribution. For a convex combination of Ns−1N_s-1 labeled sources, DˉS=∑i=1Ns−1αiDSi\bar D_S=\sum_{i=1}^{N_s-1}\alpha_iD_{S_i} with αi≥0\alpha_i\ge0 and ∑iαi=1\sum_i\alpha_i=1, the error on an unlabeled source domain DSTD_{S_T} satisfies

    err⁡DST(hf)≤∑i=1Ns−1αi(err⁡DSi(ρ)(f)+df,F(ρ)(DSi,DST))+λ^,\operatorname{err}_{D_{S_T}}(h_f)\le \sum_{i=1}^{N_s-1}\alpha_i\left(\operatorname{err}^{(\rho)}_{D_{S_i}}(f)+d^{(\rho)}_{f,\mathcal F}(D_{S_i},D_{S_T})\right)+\hat\lambda,

    where λ^=min⁡f∗∈F[∑i=1Ns−1αierr⁡DSi(ρ)(f∗)+err⁡DST(ρ)(f∗)]\hat\lambda=\min_{f^*\in\mathcal F}\left[\sum_{i=1}^{N_s-1}\alpha_i\operatorname{err}^{(\rho)}_{D_{S_i}}(f^*)+\operatorname{err}^{(\rho)}_{D_{S_T}}(f^*)\right] is independent of ff.

    For the domain-generalization setting, define the convex hull of the source distributions as ΛS={∑i=1NsπiDSi:πi≥0,∑iπi=1}\Lambda_S=\{\sum_{i=1}^{N_s}\pi_iD_{S_i}:\pi_i\ge0,\sum_i\pi_i=1\}. Let DˉU=∑i=1NsπiDSi\bar D_U=\sum_{i=1}^{N_s}\pi_iD_{S_i} be the source-hull projection minimizing MDD to DUD_U, and define

    γ=df,F(ρ)(DU,DˉU).\gamma=d^{(\rho)}_{f,\mathcal F}(D_U,\bar D_U).

    If every pair of source domains has MDD at most ϵ\epsilon, then the unseen-domain error is bounded by

    err⁡DU(hf)≤∑i=1Nsπierr⁡DSi(ρ)(f)+ϵ+γ+λˉ,\operatorname{err}_{D_U}(h_f)\le \sum_{i=1}^{N_s}\pi_i\operatorname{err}^{(\rho)}_{D_{S_i}}(f)+\epsilon+\gamma+\bar\lambda,

    where

    λˉ=min⁡f∗∈F[∑i=1Nsπierr⁡DSi(ρ)(f∗)+err⁡DU(ρ)(f∗)].\bar\lambda=\min_{f^*\in\mathcal F}\left[\sum_{i=1}^{N_s}\pi_i\operatorname{err}^{(\rho)}_{D_{S_i}}(f^*)+\operatorname{err}^{(\rho)}_{D_U}(f^*)\right].

    The paper identifies ϵ\epsilon with the largest pairwise source MDD. If the target lies in the source convex hull, then γ=0\gamma=0; otherwise γ>0\gamma>0 measures the remaining target-to-source-hull discrepancy. Thus, source diversity can reduce the projection term, while source classification error and pairwise MDD remain explicit contributors to the target bound.

  3. Knowl 3 — Rademacher-complexity generalization bound

    theoretical result

    For a function class A\mathcal A with functions mapping into [a,b][a,b], the paper uses the distributional Rademacher complexity

    Rn,D(A)=ED^∼DnEσ[sup⁡g∈A1n∑r=1nσrg(xr,yr)],\mathfrak R_{n,D}(\mathcal A)=\mathbb E_{\hat D\sim D^n}\mathbb E_{\sigma}\left[\sup_{g\in\mathcal A}\frac{1}{n}\sum_{r=1}^{n}\sigma_rg(x_r,y_r)\right],

    where D^={(xr,yr)}r=1n\hat D=\{(x_r,y_r)\}_{r=1}^n and the independent Rademacher variables satisfy σr∈{−1,+1}\sigma_r\in\{-1,+1\}. Define

    ΠHF={x↦f(x,h(x)):h∈H,f∈F},\Pi_{\mathcal H}\mathcal F=\{x\mapsto f(x,h(x)):h\in\mathcal H,f\in\mathcal F\},

    and

    Π1F={x↦f(x,y):y∈Y,f∈F}.\Pi_1\mathcal F=\{x\mapsto f(x,y):y\in\mathcal Y,f\in\mathcal F\}.

    Let (i′,k′)(i',k') be a source-domain pair attaining the maximum empirical MDD, with source sample sizes ni′n_{i'} and nk′n_{k'}. For any δ>0\delta>0, with probability at least 1−3δ1-3\delta over the labeled source samples, the unseen-domain error of every f∈Ff\in\mathcal F is bounded by

    err⁡DU(hf)≤∑i=1Nsπierr⁡D^Si(ρ)(f)+df,F(ρ)(D^Si′,D^Sk′)+γ+λˉ+KρRni′,DSi′(ΠHF)+KρRnk′,DSk′(ΠHF)+log⁡(2/δ)2ni′+log⁡(2/δ)2nk′+∑i=1Nsπi[2K2ρRni,DSi(Π1F)+log⁡(2/δ)2ni].\begin{aligned} \operatorname{err}_{D_U}(h_f)\le{}& \sum_{i=1}^{N_s}\pi_i\operatorname{err}^{(\rho)}_{\hat D_{S_i}}(f) +d^{(\rho)}_{f,\mathcal F}(\hat D_{S_{i'}},\hat D_{S_{k'}})+\gamma+\bar\lambda\\ &+\frac{K}{\rho}\mathfrak R_{n_{i'},D_{S_{i'}}}(\Pi_{\mathcal H}\mathcal F) +\frac{K}{\rho}\mathfrak R_{n_{k'},D_{S_{k'}}}(\Pi_{\mathcal H}\mathcal F)\\ &+\sqrt{\frac{\log(2/\delta)}{2n_{i'}}} +\sqrt{\frac{\log(2/\delta)}{2n_{k'}}}\\ &+\sum_{i=1}^{N_s}\pi_i\left[ \frac{2K^2}{\rho}\mathfrak R_{n_i,D_{S_i}}(\Pi_1\mathcal F) +\sqrt{\frac{\log(2/\delta)}{2n_i}} \right]. \end{aligned}

    Here K=∣Y∣K=|\mathcal Y|, err⁡D^(ρ)(f)\operatorname{err}^{(\rho)}_{\hat D}(f) is empirical margin error, γ=df,F(ρ)(DU,DˉU)\gamma=d^{(\rho)}_{f,\mathcal F}(D_U,\bar D_U) is the target-to-source-hull discrepancy, and λˉ\bar\lambda is the ideal source-plus-target margin loss defined in the source-mixture bound. The result makes the margin threshold explicit: smaller ρ\rho increases the complexity coefficients, whereas an excessively large ρ\rho can increase empirical margin error because it produces a weak classifier.

  4. Knowl 4 — MADG minimizes source error and pairwise source discrepancy

    model/method

    MADG is a margin-based adversarial domain-generalization method for NsN_s labeled source domains. Its theoretical training target is to minimize a weighted source margin-error term together with the sum of empirical MDDs over every unordered source-domain pair:

    min⁡f∈F∑i=1Nsπierr⁡D^Si(ρ)(f)+∑i=1Ns−1∑k=i+1Nsdf,F(ρ)(D^Si,D^Sk),\min_{f\in\mathcal F} \sum_{i=1}^{N_s}\pi_i\operatorname{err}^{(\rho)}_{\hat D_{S_i}}(f) +\sum_{i=1}^{N_s-1}\sum_{k=i+1}^{N_s} d^{(\rho)}_{f,\mathcal F}(\hat D_{S_i},\hat D_{S_k}),

    where πi≥0\pi_i\ge0 and ∑iπi=1\sum_i\pi_i=1. The sum of pairwise discrepancies is used as an efficiently optimized surrogate for the maximum source-pair MDD, which upper-bounds discrepancies within the convex hull of the source domains.

    To implement the supremum in each MDD, MADG introduces a feature extractor GG and j=(Ns2)j=\binom{N_s}{2} adversarial scoring functions f1′,…,fj′f'_1,\ldots,f'_j. Each fl′f'_l is assigned to one unordered source pair (il,kl)(i_l,k_l) and attempts to maximize the discrepancy for that pair, while the main classifier ff and feature extractor GG attempt to minimize source classification loss and the resulting transfer loss. Thus the method learns features that remain discriminative for all source labels while reducing margin-based disagreement between source domains.

  5. Knowl 5 — Cross-entropy surrogate and adversarial MADG objective

    model/method

    Because the empirical MDD supremum is inconvenient for direct stochastic-gradient optimization, MADG uses a pairwise surrogate. For a source pair (DSi,DSk)(D_{S_i},D_{S_k}) and adversarial classifier fl′f'_l, define

    E(D^Si)=E(x,y)∼D^Si[L(f(G(x)),y)],\mathcal E(\hat D_{S_i})= \mathbb E_{(x,y)\sim\hat D_{S_i}} \left[L\bigl(f(G(x)),y\bigr)\right],

    and

    D(ρ^,l)(D^Si,D^Sk)=E(x,y)∼D^Sk[L′(fl′(G(x)),f(G(x)))]−ρ^ E(x,y)∼D^Si[L(fl′(G(x)),f(G(x)))],\mathcal D^{(\hat\rho,l)}(\hat D_{S_i},\hat D_{S_k})= \mathbb E_{(x,y)\sim\hat D_{S_k}} \left[L'\bigl(f'_l(G(x)),f(G(x))\bigr)\right] -\hat\rho\,\mathbb E_{(x,y)\sim\hat D_{S_i}} \left[L\bigl(f'_l(G(x)),f(G(x))\bigr)\right],

    where ρ^=exp⁡(ρ)\hat\rho=\exp(\rho), LL is cross-entropy, and L′L' is the adversarial complement loss. With softmax σw(z)=exp⁡(zw)/∑q=1Kexp⁡(zq)\sigma_w(z)=\exp(z_w)/\sum_{q=1}^{K}\exp(z_q), the losses used by the paper are

    L(f(G(x)),y)=−log⁡[σy(f(G(x)))],L(f(G(x)),y)=-\log\left[\sigma_y(f(G(x)))\right], L′(fl′(G(x)),f(G(x)))=log⁡[1−σhf(G(x))(fl′(G(x)))].L'\bigl(f'_l(G(x)),f(G(x))\bigr)= \log\left[1-\sigma_{h_f(G(x))}\bigl(f'_l(G(x))\bigr)\right].

    The resulting minimax problem is

    min⁡f,Gmax⁡f1′,…,fj′{∑i=1NsπiE(D^Si)+∑l=1jD(ρ^,l)(D^Sil,D^Skl)}.\min_{f,G}\max_{f'_1,\ldots,f'_j} \left\{ \sum_{i=1}^{N_s}\pi_i\mathcal E(\hat D_{S_i}) +\sum_{l=1}^{j}\mathcal D^{(\hat\rho,l)}(\hat D_{S_{i_l}},\hat D_{S_{k_l}}) \right\}.

    The feature extractor receives gradient-reversed transfer gradients through a Gradient Reversal Layer, while the main classifier is trained for source classification. The page-7 architecture diagram depicts one shared feature extractor, one classification head, and one adversarial MDD head for every source pair.

  6. Knowl 6 — Two-stage MADG training algorithm

    algorithm

    MADG takes NsN_s labeled source domains, a margin parameter, learning-rate and optimization settings, a fixed pairing of all j=(Ns2)j=\binom{N_s}{2} source pairs, and numbers of epochs and minibatches. It outputs the trained feature extractor GG and classifier ff.

    Input: Ns labeled source domains, all unordered source pairs, total_epochs, total_batches
    Initialize feature extractor G, classifier f, and adversarial classifiers f'_1,...,f'_j
    for epoch from 1 to total_epochs do
        for batch from 1 to total_batches do
            Draw a minibatch from every source domain
            Update f and G to reduce source cross-entropy and the current transfer objective
            Compute source predictions y_hat_i = f(G(x_i)) for every source minibatch
            For each pair l with source domains (i_l, k_l), compute the surrogate MDD D_l
            Set transfer loss to the sum of all pair losses D_l
            Update adversarial classifiers f'_1,...,f'_j to increase the transfer objective
            Update G through gradient reversal so that source-domain transfer loss decreases
        end for
    end for
    return f and G

    The paper alternates the parameter updates rather than updating all heads jointly: the main classifier and feature extractor are updated first, then a forward pass produces the reference predictions, after which the adversarial classifiers and feature extractor are updated using the transfer loss. An ablation reported by the paper finds this alternating strategy better than a joint update.

  7. Knowl 7 — Domain-generalization experimental protocol

    experimental setup

    MADG was evaluated for image classification on five DomainBed benchmark datasets: VLCS, PACS, OfficeHome, TerraIncognita, and DomainNet. The feature extractor was an ImageNet-pretrained ResNet-50, optimization used stochastic gradient descent with momentum, and each minibatch contained 32 samples from every source domain. Hyperparameters were selected with the test-domain-validation procedure. Each experiment was run for three trials, and the paper reports mean accuracy with standard deviation.

    In addition to per-dataset accuracy and average accuracy, the evaluation used three consistency metrics. Median rank (M) is the median rank of a method across datasets, with lower values preferred. Arithmetic mean of differences (AD) is the arithmetic mean, across datasets, of the difference between the best accuracy on that dataset and the method's accuracy. Geometric mean of differences (GD) is the corresponding geometric mean; lower AD and GD indicate smaller gaps from the best method and therefore more consistent performance.

  8. Knowl 8 — MADG achieves the strongest aggregate benchmark performance

    data/table

    The benchmark comparison measures domain-generalization accuracy across five datasets and consistency using AD, GD, and median rank. The representative entries below reproduce the paper's reported values; accuracy values are percentages and are reported as mean ±\pm standard deviation. MADG has the best aggregate accuracy, AD, and GD among the compared methods, although another method can be better on an individual dataset.

    Could not parse LaTeX table

    MADG improves average accuracy by about 11 percentage point over ERM and about 22 percentage points over DANN and CDANN. Its OfficeHome accuracy is approximately 33 percentage points higher than the other reported methods, and its lower AD and GD indicate better cross-dataset consistency. Among theory-based methods reporting all five datasets, MADG exceeds MTL by 1.21.2 percentage points in average accuracy, while reducing AD by 1.31.3 and GD by 1.71.7.

  9. Knowl 9 — Margin and pair-selection ablations validate MADG design choices

    empirical result

    The paper studied the practical margin ρ^=exp⁡(ρ)\hat\rho=\exp(\rho) on PACS and VLCS. The best tested value was ρ^=1.5\hat\rho=1.5 on both datasets, showing the predicted trade-off between obtaining a larger margin and increasing the surrogate loss for a weak classifier.

    Could not parse LaTeX table

    MADG normally uses one adversarial classifier for every source pair, j=(Ns2)j=\binom{N_s}{2}. Replacing this with only Ns−1N_s-1 classifiers, each comparing one fixed source to the other sources, reduced accuracy on every reported dataset and especially hurt TerraIncognita:

    Could not parse LaTeX table

    These results support using the full set of pairwise MDD terms rather than a restricted set of source comparisons.

  10. Knowl 10 — Colored MNIST supports the source-convex-hull interpretation

    empirical result

    The paper approximates the target-to-source-hull term γ\gamma with pairwise Jensen–Shannon divergence and evaluates this interpretation on Colored MNIST. The three domains are labeled +90+90, +80+80, and −90-90. Their pairwise divergences are small, and the corresponding approximate γ\gamma values are also small:

    Could not parse LaTeX table

    On this dataset, MADG achieved 65.6%65.6\% average accuracy versus 57.8%57.8\% for ERM, indicating that the margin-based adversarial representation learned more transferable features. Per-domain accuracies were:

    Could not parse LaTeX table

    The particularly large improvement on the −90-90 domain is consistent with the paper's claim that small source-to-hull discrepancy makes domain-invariant learning more effective.

Coverage note — Lower-priority computational-cost results, weighted-MDD experiments, and the additional OfficeHome comparison against MIRO and SD were omitted because they support rather than define the paper's main theoretical and algorithmic contributions.

References

  1. 1.K. Aggarwal, S. K. Singh, M. Chopra, S. Kumar, and F. Colace, “Deep learning in robotics for strengthening industry 4.0.: opportunities, challenges and future directions,” Robotics and AI for Cybersecurity and Critical Infrastructure in Smart Cities, pp. 1–19, 2022.
  2. 2.M. Tsuneki, “Deep learning models in medical image analysis,” Journal of Oral Biosciences, vol. 64, no. 3, pp. 312–320, 2022.
  3. 3.T. Ayoub Shaikh, T. Rasool, and F. Rasheed Lone, “Towards leveraging the role of machine learning and artificial intelligence in precision agriculture and smart farming,” Computers and Electronics in Agriculture, vol. 198, p. 107119, 2022.
  4. 4.J. Quionero-Candela, M. Sugiyama, A. Schwaighofer, and N. D. Lawrence, Dataset Shift in Machine Learning. The MIT Press, 2009.
  5. 5.J. Wang, C. Lan, C. Liu, Y. Ouyang, T. Qin, W. Lu, Y. Chen, W. Zeng, and P. Yu, “Generalizing to unseen domains: A survey on domain generalization,” IEEE Transactions on Knowledge and Data Engineering, 2022.
  6. 6.Y. Zhang, “A survey of unsupervised domain adaptation for visual recognition,” arXiv preprint arXiv:2112.06745, 2021.
  7. 7.X. Liu, C. Yoo, F. Xing, H. Oh, G. El Fakhri, J.-W. Kang, J. Woo et al., “Deep unsupervised domain adaptation: A review of recent advances and perspectives,” APSIPA Transactions on Signal and Information Processing, vol. 11, no. 1, 2022.
  8. 8.G. Blanchard, G. Lee, and C. Scott, “Generalizing from several related classification tasks to a new unlabeled sample,” in Advances in Neural Information Processing Systems, vol. 24, 2011.
  9. 9.T. Matsuura and T. Harada, “Domain generalization using a mixture of multiple latent domains,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 11 749–11 756.
  10. 10.S. Lin, C.-T. Li, and A. C. Kot, “Multi-domain adversarial feature generalization for person re-identification,” IEEE Transactions on Image Processing, vol. 30, pp. 1596–1607, 2020.
  11. 11.Z. Deng, F. Ding, C. Dwork, R. Hong, G. Parmigiani, P. Patil, and P. Sur, “Representation via representations: Domain generalization via adversarially learned invariant representations,” arXiv preprint arXiv:2006.11478, 2020.
  12. 12.S. Zhao, M. Gong, T. Liu, H. Fu, and D. Tao, “Domain generalization via entropy regularization,” Advances in Neural Information Processing Systems, vol. 33, pp. 16 096–16 107, 2020.
  13. 13.K. Akuzawa, Y. Iwasawa, and Y. Matsuo, “Adversarial invariant feature learning with accuracy constraint for domain generalization,” in Machine Learning and Knowledge Discovery in Databases: European Conference. Springer, 2020, pp. 315–331.
  14. 14.E. Rosenfeld, P. Ravikumar, and A. Risteski, “An online learning approach to interpolation and extrapolation in domain generalization,” in International Conference on Artificial Intelligence and Statistics. PMLR, 2022, pp. 2641–2657.
  15. 15.M.-H. Bui, T. Tran, A. Tran, and D. Phung, “Exploiting domain-specific features to enhance domain generalization,” Advances in Neural Information Processing Systems, vol. 34, pp. 21 189–21 201, 2021.
  16. 16.V. Narayanan, A. A. Deshmukh, U. Dogan, and V. N. Balasubramanian, “On challenges in unsupervised domain generalization,” in NeurIPS 2021 Workshop on Pre-registration in Machine Learning. PMLR, 2022, pp. 42–58.
  17. 17.Y. Shu, Z. Cao, C. Wang, J. Wang, and M. Long, “Open domain generalization with domain-augmented meta-learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 9624–9633.
  18. 18.J. Yuan, X. Ma, D. Chen, K. Kuang, F. Wu, and L. Lin, “Label-efficient domain generalization via collaborative exploration and generalization,” in Proceedings of the 30th ACM International Conference on Multimedia, 2022, pp. 2361–2370.
  19. 19.S. Paul, T. Dutta, and S. Biswas, “Universal cross-domain retrieval: Generalizing across classes and domains,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 12 056–12 064.
  20. 20.Q. Xu, R. Zhang, Y. Zhang, Y. Wang, and Q. Tian, “A fourier-based framework for domain generalization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14 383–14 392.
  21. 21.Z. Li, Z. Cui, S. Wang, Y. Qi, X. Ouyang, Q. Chen, Y. Yang, Z. Xue, D. Shen, and J.-Z. Cheng, “Domain generalization for mammography detection via multi-style and multi-view contrastive learning,” in Medical Image Computing and Computer Assisted Intervention–MICCAI Proceedings, 2021, pp. 98–108.
  22. 22.D. Kim, Y. Yoo, S. Park, J. Kim, and J. Lee, “Selfreg: Self-supervised contrastive regularization for domain generalization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 9619–9628.
  23. 23.Y. Chen, Y. Wang, Y. Pan, T. Yao, X. Tian, and T. Mei, “A style and semantic memory mechanism for domain generalization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 9164–9173.
  24. 24.D. Li, H. Gouk, and T. Hospedales, “Finding lost dg: Explaining domain generalization via model complexity,” arXiv preprint arXiv:2202.00563, 2022.
  25. 25.M. Zhang, H. Marklund, N. Dhawan, A. Gupta, S. Levine, and C. Finn, “Adaptive risk minimization: Learning to adapt to domain shift,” Advances in Neural Information Processing Systems, vol. 34, pp. 23 664–23 678, 2021.
  26. 26.A. Nasery, S. Thakur, V. Piratla, A. De, and S. Sarawagi, “Training for the future: A simple gradient interpolation loss to generalize along time,” Advances in Neural Information Processing Systems, vol. 34, pp. 19 198–19 209, 2021.
  27. 27.J. Cha, S. Chun, K. Lee, H.-C. Cho, S. Park, Y. Lee, and S. Park, “Swad: Domain generalization by seeking flat minima,” Advances in Neural Information Processing Systems, vol. 34, pp. 22 405–22 418, 2021.
  28. 28.A. Robey, G. J. Pappas, and H. Hassani, “Model-based domain generalization,” Advances in Neural Information Processing Systems, vol. 34, pp. 20 210–20 229, 2021.
  29. 29.J. Cha, K. Lee, S. Park, and S. Chun, “Domain generalization by mutual-information regularization with pre-trained models,” in Proceedings of the European Conference on Computer Vision (ECCV), 2022, pp. 440–457.
  30. 30.C. Fang, Y. Xu, and D. N. Rockmore, “Unbiased metric learning: On the utilization of multiple datasets and web images for softening bias,” in 2013 IEEE International Conference on Computer Vision, 2013, pp. 1657–1664.
  31. 31.D. Li, Y. Yang, Y.-Z. Song, and T. M. Hospedales, “Deeper, broader and artier domain generalization,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 5542–5550.
  32. 32.H. Venkateswara, J. Eusebio, S. Chakraborty, and S. Panchanathan, “Deep hashing network for unsupervised domain adaptation,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 5018–5027.
  33. 33.S. Beery, G. Van Horn, and P. Perona, “Recognition in terra incognita,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 456–473.
  34. 34.X. Peng, Q. Bai, X. Xia, Z. Huang, K. Saenko, and B. Wang, “Moment matching for multi-source domain adaptation,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1406–1415.
  35. 35.I. Gulrajani and D. Lopez-Paz, “In search of lost domain generalization,” in 9th International Conference on Learning Representations, ICLR, 2021.
  36. 36.A. Gretton, A. Smola, J. Huang, M. Schmittfull, K. Borgwardt, and B. Schölkopf, “Covariate Shift by Kernel Mean Matching,” in Dataset Shift in Machine Learning. The MIT Press, 2008.
  37. 37.J. Lin, Y. Tang, J. Wang, and W. Zhang, “Mitigating both covariate and conditional shift for domain generalization,” in 2022 IEEE 8th International Conference on Cloud Computing and Intelligent Systems (CCIS). IEEE, 2022, pp. 437–443.
  38. 38.S. Ben-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan, “A theory of learning from different domains,” Machine learning, vol. 79, no. 1, pp. 151–175, 2010.
  39. 39.V. Koltchinskii and D. Panchenko, “Empirical margin distributions and bounding the generalization error of combined classifiers,” The Annals of Statistics, vol. 30, no. 1, pp. 1–50, 2002.
  40. 40.Y. Zhang, T. Liu, M. Long, and M. Jordan, “Bridging theory and algorithm for domain adaptation,” in Proceedings of the 36th International Conference on Machine Learning, vol. 97, 2019, pp. 7404–7413.
  41. 41.P. L. Bartlett and S. Mendelson, “Rademacher and gaussian complexities: Risk bounds and structural results,” Journal of Machine Learning Research, vol. 3, no. Nov, pp. 463–482, 2002.
  42. 42.M. M. Rahman, C. Fookes, M. Baktashmotlagh, and S. Sridharan, “Correlation-aware adversarial domain adaptation and generalization,” Pattern Recognition, vol. 100, p. 107124, 2020.
  43. 43.G. Blanchard, G. Lee, and C. Scott, “Generalizing from several related classification tasks to a new unlabeled sample,” Advances in neural information processing systems, vol. 24, 2011.
  44. 44.K. Muandet, D. Balduzzi, and B. Schölkopf, “Domain generalization via invariant feature representation,” in International conference on machine learning. PMLR, 2013, pp. 10–18.
  45. 45.A. A. Deshmukh, Y. Lei, S. Sharma, U. Dogan, J. W. Cutler, and C. Scott, “A generalization error bound for multi-class domain generalization,” arXiv preprint arXiv:1905.10392, 2019.
  46. 46.S. Hu, K. Zhang, Z. Chen, and L. Chan, “Domain generalization via multidomain discriminant analysis,” in Uncertainty in Artificial Intelligence. PMLR, 2020, pp. 292–302.
  47. 47.G. Blanchard, A. A. Deshmukh, Ü. Dogan, G. Lee, and C. Scott, “Domain generalization by marginal transfer learning,” The Journal of Machine Learning Research, vol. 22, no. 1, pp. 46–100, 2021.
  48. 48.I. Albuquerque, J. Monteiro, M. Darvishi, T. H. Falk, and I. Mitliagkas, “Generalizing to unseen domains via distribution matching,” arXiv preprint arXiv:1911.00804, 2019.
  49. 49.H. Ye, C. Xie, T. Cai, R. Li, Z. Li, and L. Wang, “Towards a theoretical framework of out-of-distribution generalization,” Advances in Neural Information Processing Systems, vol. 34, pp. 23 519–23 531, 2021.
  50. 50.R. Vedantam, D. Lopez-Paz, and D. J. Schwab, “An empirical investigation of domain generalization with empirical risk minimizers,” Advances in Neural Information Processing Systems, vol. 34, pp. 28 131–28 143, 2021.
  51. 51.M. Mohri, A. Rostamizadeh, and A. Talwalkar, Foundations of machine learning. MIT press, 2018.
  52. 52.Y. Ganin and V. Lempitsky, “Unsupervised domain adaptation by backpropagation,” in International conference on machine learning. PMLR, 2015, pp. 1180–1189.
  53. 53.A. Rame, C. Dancette, and M. Cord, “Fishr: Invariant gradient variances for out-of-distribution generalization,” in International Conference on Machine Learning. PMLR, 2022, pp. 18 347–18 377.
  54. 54.K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
  55. 55.O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein et al., “Imagenet large scale visual recognition challenge,” International journal of computer vision, vol. 115, pp. 211–252, 2015.
  56. 56.N. Qian, “On the momentum term in gradient descent learning algorithms,” Neural networks, vol. 12, no. 1, pp. 145–151, 1999.
  57. 57.V. N. Vapnik, “An overview of statistical learning theory,” IEEE transactions on neural networks, vol. 10, no. 5, pp. 988–999, 1999.
  58. 58.B. Sun and K. Saenko, “Deep coral: Correlation alignment for deep domain adaptation,” in Computer Vision–ECCV 2016 Workshops: Amsterdam, The Netherlands, October 8-10 and 15-16, 2016, Proceedings, Part III 14. Springer, 2016, pp. 443–450.
  59. 59.Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky, “Domain-adversarial training of neural networks,” The journal of machine learning research, vol. 17, no. 1, pp. 2096–2030, 2016.
  60. 60.Y. Li, X. Tian, M. Gong, Y. Liu, T. Liu, K. Zhang, and D. Tao, “Deep domain generalization via conditional invariant adversarial networks,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 624–639.
  61. 61.D. Li, Y. Yang, Y.-Z. Song, and T. Hospedales, “Learning to generalize: Meta-learning for domain generalization,” in Proceedings of the AAAI conference on artificial intelligence, vol. 32, no. 1, 2018.
  62. 62.H. Li, S. J. Pan, S. Wang, and A. C. Kot, “Domain generalization with adversarial feature learning,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 5400–5409.
  63. 63.M. Arjovsky, L. Bottou, I. Gulrajani, and D. Lopez-Paz, “Invariant risk minimization,” arXiv preprint arXiv:1907.02893, 2019.
  64. 64.S. Sagawa, P. W. Koh, T. B. Hashimoto, and P. Liang, “Distributionally robust neural networks for group shifts: On the importance of regularization for worst-case generalization,” arXiv preprint arXiv:1911.08731, 2019.
  65. 65.S. Yan, H. Song, N. Li, L. Zou, and L. Ren, “Improve unsupervised domain adaptation with mixup training,” arXiv preprint arXiv:2001.00677, 2020.
  66. 66.M. Zhang, H. Marklund, A. Gupta, S. Levine, and C. Finn, “Adaptive risk minimization: A meta-learning approach for tackling group shift,” arXiv preprint arXiv:2007.02931, vol. 8, p. 9, 2020.
  67. 67.Z. Huang, H. Wang, E. P. Xing, and D. Huang, “Self-challenging improves cross-domain generalization,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. Springer, 2020, pp. 124–140.
  68. 68.H. Nam, H. Lee, J. Park, W. Yoon, and D. Yoo, “Reducing domain gap by reducing style bias,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 8690–8699.
  69. 69.D. Krueger, E. Caballero, J.-H. Jacobsen, A. Zhang, J. Binas, D. Zhang, R. Le Priol, and A. Courville, “Out-of-distribution generalization via risk extrapolation (rex),” in International Conference on Machine Learning. PMLR, 2021, pp. 5815–5826.
  70. 70.G. Parascandolo, A. Neitz, A. Orvieto, L. Gresele, and B. Schölkopf, “Learning explanations that are hard to vary,” in 9th International Conference on Learning Representations, ICLR, 2021.
  71. 71.Y. Shi, J. Seely, P. H. Torr, N. Siddharth, N. Hannun, N. Usunier, and G. Synnaeve, “Gradient matching for domain generalization,” arXiv preprint arXiv:2104.09937, 2021.
  72. 72.S. Shahtalebi, J.-C. Gagnon-Audet, T. Laleh, M. Faramarzi, K. Ahuja, and I. Rish, “Sandmask: An enhanced gradient masking strategy for the discovery of invariances in domain generalization,” ArXiv, vol. abs/2106.02266, 2021.
  73. 73.G. Zhang, H. Zhao, Y. Yu, and P. Poupart, “Quantifying and improving transferability in domain generalization,” Advances in Neural Information Processing Systems, vol. 34, pp. 10 957–10 970, 2021.
  74. 74.M. Pezeshki, O. Kaba, Y. Bengio, A. C. Courville, D. Precup, and G. Lajoie, “Gradient starvation: A learning proclivity in neural networks,” Advances in Neural Information Processing Systems, vol. 34, pp. 1256–1272, 2021.

Citation

MLA
Dayal, A., et al. “MADG: Margin-based Adversarial Learning for Domain Generalization”. arXiv, 2023, http://arxiv.org/abs/2311.08503v1.
APA
Dayal, A., B., V. K., Cenkeramaddi, L. R., Mohan, C. K., Kumar, A., & Balasubramanian, V. N. (2023). MADG: Margin-based Adversarial Learning for Domain Generalization. arXiv. http://arxiv.org/abs/2311.08503v1
Chicago
Dayal, A., V. K. B., L. R. Cenkeramaddi, C. K. Mohan, A. Kumar, and V. N. Balasubramanian. 2023. “MADG: Margin-based Adversarial Learning for Domain Generalization”. arXiv. http://arxiv.org/abs/2311.08503v1.
Harvard
Dayal, A. et al. (2023) “MADG: Margin-based Adversarial Learning for Domain Generalization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2311.08503v1.
Vancouver
1. Dayal A, B. VK, Cenkeramaddi LR, Mohan CK, Kumar A, Balasubramanian VN (2023) MADG: Margin-based Adversarial Learning for Domain Generalization. arXiv

BibTeX

@article{dayal2023madg,
  title = {MADG: Margin-based Adversarial Learning for Domain Generalization},
  author = {Dayal, Aveen and B., Vimal K. and Cenkeramaddi, Linga Reddy and Mohan, C. Krishna and Kumar, Abhinav and Balasubramanian, Vineeth N},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2311.08503v1},
  eprint = {2311.08503}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors