Federated Domain Generalization with Generalization Adjustment

Ruipeng ZhangQinwei XuJiangchao YaoYa ZhangQi TianYanfeng Wang

article2023CVPR121 citations

Proposes Generalization Adjustment, a server-side aggregation strategy that dynamically recalibrates client weights using a variance-reduction regularizer on generalization gaps to improve model performance on unseen target domains without sharing private multi-domain data.

Listen

Modern organizations frequently need to train artificial intelligence models collaboratively across distributed data sites without sharing sensitive private data. While federated learning enables such decentralized training, deployed models often perform poorly when introduced to entirely new environments or client sites with different data distributions—a challenge known as federated domain generalization. Standard centralized techniques that address data distribution shifts require simultaneous access to multi-site data in shared batches, which directly violates core privacy requirements.

The article develops and evaluates a server-side framework, Generalization Adjustment, designed to train global models that generalize effectively to unseen target environments without exchanging raw client data or spectra. The approach introduces a global fairness objective that minimizes the variance of generalization gaps across training sites by dynamically calibrating model aggregation weights on the central server.

To validate the framework, the authors conducted theoretical analyses alongside extensive leave-one-domain-out empirical evaluations on four standard benchmark image datasets: PACS, OfficeHome, TerraInc, and DomainNet. The experimental setup tested the integration of Generalization Adjustment across nine existing federated learning and domain generalization base algorithms using standard neural network backbones across multiple communication rounds.

The analysis yielded several key findings. First, Generalization Adjustment consistently improved out-of-domain performance when paired with every baseline method across all four benchmarks. For instance, applying the method to standard federated averaging boosted overall average accuracy from 63.10% to 65.06% across PACS, OfficeHome, and TerraInc, and from 57.29% to 60.24% on DomainNet. Second, the largest gains occurred on highly challenging datasets with severe distribution shifts, such as TerraInc, where test accuracy increased by more than 3 percentage points on average. Third, the dynamic weighting mechanism significantly narrowed both the mean and the variance of generalization gaps among participating sites, driving global and local models into flatter, more robust regions of the optimization loss surface.

These findings demonstrate that central aggregation weighting is a powerful, low-overhead mechanism for building robust distributed models. Relying purely on local client optimizations often fails to prevent global models from overfitting to participating sites. By adjusting aggregation weights solely based on scalar generalization gap values, organizations can improve cross-site model reliability and compliance without introducing new data privacy vulnerabilities or communication bottlenecks.

Organizations deploying federated learning systems across diverse institutions should adopt dynamic aggregation weighting in place of static, sample-size-proportional averaging. Because the Generalization Adjustment technique operates strictly at the central server during model aggregation, technical teams can readily integrate it into existing local training pipelines without altering edge hardware or client algorithms.

While the empirical results are robust across benchmark vision datasets, the authors acknowledge that the current method relies on an approximation of optimal generalization weights rather than direct closed-form optimization. The weighting scheme may also exhibit over-confidence when encountering source sites with conflicting gradients or easily fitted data distributions. Further operational pilots are advisable when expanding the framework beyond controlled image benchmarks to complex, multimodal production environments.

Cover for Federated Domain Generalization with Generalization Adjustment

Abstract

Federated Domain Generalization (FedDG) attempts to learn a global model in a privacy-preserving manner that generalizes well to new clients possibly with domain shift. Recent exploration mainly focuses on designing an unbiased training strategy within each individual domain. However, without the support of multi-domain data jointly in the mini-batch training, almost all methods cannot guarantee the generalization under domain shift. To overcome this problem, we propose a novel global objective incorporating a new variance reduction regularizer to encourage fairness. A novel FL-friendly method named Generalization Adjustment (GA) is proposed to optimize the above objective by dynamically calibrating the aggregation weights. The theoretical analysis of GA demonstrates the possibility to achieve a tighter generalization bound with an explicit re-weighted aggregation, substituting the implicit multi-domain data sharing that is only applicable to the conventional DG settings. Besides, the proposed algorithm is generic and can be combined with any local client training-based methods. Extensive experiments on several benchmark datasets have shown the effectiveness of the proposed method, with consistent improvements over several FedDG algorithms when used in combination. The source code is released at https://github.com/MediaBrain-SJTU/FedDG-GA

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Domain Generalization
  • 2.2. Federated Learning
  • 2.3. Federated Domain Generalization
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Motivation
  • 3.3. Generalization Adjustment for FedDG
  • 3.4. Theoretical analysis
  • 4. Experimental Results
  • 4.1. Dataset and implementation details
  • 4.2. Main results
  • 4.3. Ablation studies
  • 5. Conclusion and Discussion
  • References

Knowls

  1. Knowl 1 — Federated domain generalization training objective

    definition

    Federated Domain Generalization (FedDG) treats each of MM isolated clients as a distinct source domain. Client ii has an empirical dataset D^i={(xji,yji)}j=1Ni\widehat D_i=\{(x_j^i,y_j^i)\}_{j=1}^{N_i} with NiN_i examples, a model f(x;θ)f(x;\theta) with parameters θ\theta, and an empirical risk

    E^D^i(θ)=1Ni∑j=1NiL(f(xji;θ),yji),\widehat{\mathcal E}_{\widehat D_i}(\theta)=\frac{1}{N_i}\sum_{j=1}^{N_i}\mathcal L\bigl(f(x_j^i;\theta),y_j^i\bigr),

    where L\mathcal L is the task loss. The usual federated empirical-risk objective uses fixed sample-proportional weights pi=Ni/∑k=1MNkp_i=N_i/\sum_{k=1}^M N_k:

    min⁡θ  ∑i=1MpiE^D^i(θ).\min_{\theta}\;\sum_{i=1}^{M}p_i\widehat{\mathcal E}_{\widehat D_i}(\theta).

    Unlike ordinary federated learning, the client distributions in FedDG are different domains rather than heterogeneous samples from one overall distribution. The learned global model must therefore perform well on an unseen target domain TT whose distribution is unavailable during training, while clients cannot jointly construct multi-domain mini-batches.

  2. Knowl 2 — Fairness-regularized global objective

    model/method

    The proposed objective learns both local model parameters and aggregation weights to reduce domain bias in the global model. For client/domain ii, let θi⋆\theta_i^\star be its locally optimal parameter vector and let θ=∑i=1Maiθi\theta=\sum_{i=1}^M a_i\theta_i be the aggregated global parameter vector, where ai≥0a_i\ge 0 and ∑iai=1\sum_i a_i=1. The empirical generalization gap on source domain ii is

    GD^i(θ)=E^D^i(θ)−E^D^i(θi⋆).G_{\widehat D_i}(\theta)=\widehat{\mathcal E}_{\widehat D_i}(\theta)-\widehat{\mathcal E}_{\widehat D_i}(\theta_i^\star).

    The proposed global objective is

    min⁡θ1,…,θM,a  ∑i=1MaiE^D^i(θ)+β Var⁡ ⁣({GD^i(θ)}i=1M)\min_{\theta_1,\ldots,\theta_M,\mathbf a}\;\sum_{i=1}^{M}a_i\widehat{\mathcal E}_{\widehat D_i}(\theta)+\beta\,\operatorname{Var}\!\left(\left\{G_{\widehat D_i}(\theta)\right\}_{i=1}^{M}\right)

    subject to θ=∑i=1Maiθi\theta=\sum_{i=1}^{M}a_i\theta_i, ai≥0a_i\ge 0, and ∑i=1Mai=1\sum_{i=1}^{M}a_i=1. Here β≥0\beta\ge0 controls the trade-off between aggregate empirical risk and equality of the source-domain generalization gaps. The variance term is intended to make the global solution similarly flat across domains rather than favoring clients whose local distributions are easiest to fit. The paper identifies β=0\beta=0 with the FedAvg objective, while very large β\beta prioritizes equal gaps over aggregate risk.

  3. Knowl 3 — Generalization Adjustment aggregation algorithm

    algorithm

    Generalization Adjustment (GA) is a server-side procedure that can be attached to an arbitrary local client-training algorithm. Its inputs are an initial global model θ0\theta^0, MM source clients with datasets D^i\widehat D_i, local epoch count EE, total communication rounds RR, and step size d∈(0,1)d\in(0,1). It initializes ai0=1/Ma_i^0=1/M and uses the local model from the previous round to estimate each client’s gap.

    At round rr, client ii evaluates the current global model and computes

    Gir=E^D^i(θr)−E^D^i(θir−1′),G_i^r=\widehat{\mathcal E}_{\widehat D_i}(\theta^r)-\widehat{\mathcal E}_{\widehat D_i}(\theta_i^{r-1\prime}),

    where θir−1′\theta_i^{r-1\prime} is the client’s locally trained model from the preceding round; the initial local model is the initialized global model. It then trains locally for EE epochs using any algorithm Alg⁡\operatorname{Alg}, producing θir′=Alg⁡(θr,D^i,E)\theta_i^{r\prime}=\operatorname{Alg}(\theta^r,\widehat D_i,E), and sends the updated model and scalar gap to the server. The server computes μr=M−1∑iGir\mu^r=M^{-1}\sum_iG_i^r and the decayed adjustment size dr=(1−r/R)dd^r=(1-r/R)d, then updates each weight by

    a~ir=air−1+(Gir−μr)drmax⁡j(Gjr−μr).\widetilde a_i^r=a_i^{r-1}+\frac{(G_i^r-\mu^r)d^r}{\max_j(G_j^r-\mu^r)}.

    The server clips negative provisional weights to zero, normalizes them as air=a~ir/∑ka~kra_i^r=\widetilde a_i^r/\sum_k\widetilde a_k^r, and aggregates

    θr+1=∑i=1Mairθir′.\theta^{r+1}=\sum_{i=1}^{M}a_i^r\theta_i^{r\prime}.

    The new global model is broadcast as the initialization for the next round. Thus, clients with larger-than-average gaps receive larger aggregation weights, while the linear decay prevents late-round weight instability. GA communicates only the scalar generalization gaps and model updates; it does not require exchanging raw data, domain statistics, or jointly formed mini-batches. The method leaves local training unrestricted and is therefore compatible with different federated or domain-generalization training algorithms.

  4. Knowl 4 — Unseen-domain generalization bound

    theoretical result

    Let TT be an unseen target domain, let ET(θ)\mathcal E_T(\theta) be the expected target risk, and let θT⋆\theta_T^\star minimize target-domain risk. For each source domain D^i\widehat D_i, let GD^i(θ)=E^D^i(θ)−E^D^i(θi⋆)G_{\widehat D_i}(\theta)=\widehat{\mathcal E}_{\widehat D_i}(\theta)-\widehat{\mathcal E}_{\widehat D_i}(\theta_i^\star), where θi⋆\theta_i^\star is the source-domain local optimum. Let aia_i be nonnegative aggregation weights summing to one, NiN_i the number of source examples, dHΔH(D^i,T)d_{\mathcal H\Delta\mathcal H}(\widehat D_i,T) the HΔH\mathcal H\Delta\mathcal H divergence between source domain ii and target domain TT, dd the hypothesis-complexity quantity appearing in the concentration term, and λ\lambda the residual joint-error term. Then, for any δ∈(0,1)\delta\in(0,1), with probability at least 1−δ1-\delta over the sampled source data, the learned global model after federated training satisfies

    ET(θ)−ET(θT⋆)≤∑i=1Mai(GD^i(θ)+dHΔH(D^i,T)+log⁡(d/δ)+log⁡(Md/δ)2Ni)+λ.\mathcal E_T(\theta)-\mathcal E_T(\theta_T^\star) \le \sum_{i=1}^{M}a_i\left( G_{\widehat D_i}(\theta) +d_{\mathcal H\Delta\mathcal H}(\widehat D_i,T) +\frac{\sqrt{\log(d/\delta)}+\sqrt{\log(Md/\delta)}}{\sqrt{2N_i}} \right)+\lambda.

    The bound separates three source-weighted effects: the flatness/generalization gap on each source domain, its divergence from the unseen target, and a finite-sample term. The result shows why changing aggregation weights can tighten the FedDG bound even when clients cannot perform explicit multi-domain alignment.

  5. Knowl 5 — Flatness variance as a surrogate for optimal domain weighting

    theoretical result

    The ideal aggregation weights would minimize the weighted source-to-target domain divergence, but the unseen target distribution and the derivatives of each client’s risk with respect to its aggregation weight are unavailable in FedDG. The paper therefore uses a flatness-based surrogate: under the approximation that the global model behaves like a robustly perturbed local solution, it proposes

    arg⁡min⁡a∑i=1Mai dHΔH(D^i,T)  ≈  arg⁡min⁡aVar⁡ ⁣({GD^i(θ)}i=1M),\arg\min_{\mathbf a}\sum_{i=1}^{M}a_i\,d_{\mathcal H\Delta\mathcal H}(\widehat D_i,T) \;\approx\; \arg\min_{\mathbf a}\operatorname{Var}\!\left(\left\{G_{\widehat D_i}(\theta)\right\}_{i=1}^{M}\right),

    with ai≥0a_i\ge0 and ∑iai=1\sum_i a_i=1. If the variance-minimizing solution makes all source gaps equal to a constant CC, then ∑iaiGD^i(θ)=∑ipiGD^i(θ)\sum_i a_iG_{\widehat D_i}(\theta)=\sum_i p_iG_{\widehat D_i}(\theta) for the fixed sample-proportional FedAvg weights pip_i. Because the selected weights minimize the surrogate divergence objective, they also satisfy

    min⁡a∑iaidHΔH(D^i,T)≤min⁡p∑ipidHΔH(D^i,T).\min_{\mathbf a}\sum_i a_i d_{\mathcal H\Delta\mathcal H}(\widehat D_i,T) \le \min_{\mathbf p}\sum_i p_i d_{\mathcal H\Delta\mathcal H}(\widehat D_i,T).

    Combining equalized gaps with this divergence inequality gives a generalization bound no looser, and potentially tighter, than the bound obtained with fixed FedAvg weights. GA is the practical momentum-based approximation to these variance-minimizing weights.

  6. Knowl 6 — Benchmark protocol for evaluating GA

    experimental setup

    The method was evaluated in leave-one-domain-out FedDG experiments on four image domain-generalization benchmarks: PACS (9,991 images and four domains), OfficeHome (15,588 images and four domains), TerraInc (24,788 images and four domains), and DomainNet (569,010 images and six domains). In each run, one domain was held out entirely as the unseen target client, while the remaining domains served as isolated source clients; the reported benchmark score averages the target-domain accuracies over all held-out domains.

    The local backbones were ImageNet-pretrained ResNet-18 for PACS and OfficeHome, ResNet-50 for TerraInc, and AlexNet for DomainNet. Local training used batch size 1616, learning rate 0.0010.001, E=5E=5 local epochs per communication round, and R=40R=40 rounds. GA used default step size d=0.05d=0.05, and reported results were averaged over three runs. The method was tested on top of FedAvg, FedCSA, FedNova, FedProx, FedSAM, HarmoFL, Scaffold, Amplitude Mix, and RSC.

  7. Knowl 7 — Consistent accuracy gains across FedDG benchmarks

    data/table

    The benchmark results compare each base aggregation or local-training method with the same method augmented by GA. Entries are mean classification accuracies in percent; PACS, OfficeHome, and TerraInc are averages over four held-out target domains, and DomainNet is an average over six held-out target domains. Across all listed methods and datasets, adding GA improves the reported average accuracy.

    Could not parse LaTeX table

    Each cell reports base method / the same method with GA. The improvements occur for conventional federated optimizers, reweighting methods, and domain-generalization local-training methods, supporting the claim that GA is a generic aggregation mechanism rather than a method tied to one local optimizer.

  8. Knowl 8 — GA reduces source-domain gap mean and variance

    empirical result

    At the end of training, the experiments compared the source-domain generalization gaps GD^i(θ)G_{\widehat D_i}(\theta) produced by each base method with those produced after GA aggregation. On PACS, OfficeHome, and TerraInc, the plotted comparisons show that GA reduces both the mean gap across source domains and the standard deviation of the gaps for the evaluated methods. The effect is especially pronounced on TerraInc, where the largest gap reductions coincide with the largest accuracy improvements.

    The target-domain loss-surface visualizations on PACS show the same mechanism geometrically: with GA, the global model and the local models lie in a flatter low-loss region of every held-out target domain, and the distances between global and local solutions are shorter than with fixed FedAvg weights. These observations support the intended interpretation that equalizing source-domain gaps reduces domain bias and improves consistency of the global solution under domain shift.

  9. Knowl 9 — Step-size and decay ablation

    data/table

    Using FedAvg as the backbone, the ablation varies the GA step size dd and compares linear decay of the step size with no decay. Each entry is mean accuracy in percent, reported as linear decay / no decay. “Fix” denotes the original sample-proportional FedAvg weights rather than GA; d=0d=0 produces uniform domain weights and therefore makes decay irrelevant.

    Could not parse LaTeX table

    The results indicate that performance is not highly sensitive to the step size and is generally stable with or without decay. The best overall decayed result is 65.0665.06 at d=0.05d=0.05. Uniform weighting at d=0d=0 still outperforms fixed sample-proportional weighting, showing that sample-count weights are not necessarily suitable when each client represents a separate domain.

  10. Knowl 10 — Limitations of the approximate weighting strategy

    limitation

    The optimal generalization weights cannot currently be obtained by directly optimizing the unseen-domain bound because the target-domain divergence and the required weight derivatives are unavailable under privacy constraints. GA therefore estimates the weights indirectly through source-domain gaps and their variance, so its weights may not be optimal. The authors further caution that GA can become over-confident in domains that are easy to fit but have inconsistent gradients. This leaves the direct evaluation and optimization of generalization weights as an open problem.

Coverage note — No substantial contributed material was omitted; background, related work, proof derivations, and acknowledgements were excluded while the method, theory, experiments, ablations, and stated limitation were retained.

References

  1. 1.Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019. 3
  2. 2.Yogesh Balaji et al. Metareg: Towards domain generalization using meta-regularization. In NeurIPS, pages 998–1008, 2018. 2
  3. 3.Sara Beery, Grant Van Horn, and Pietro Perona. Recognition in terra incognita. In ECCV, pages 456–473, 2018. 5
  4. 4.Gilles Blanchard, Gyemin Lee, and Clayton Scott. Generalizing from several related classification tasks to a new unlabeled sample. NeurIPS, 24, 2011. 1
  5. 5.Fabio M Carlucci, Antonio D’Innocente, Silvia Bucci, Barbara Caputo, and Tatiana Tommasi. Domain generalization by solving jigsaw puzzles. In CVPR, pages 2229–2238, 2019. 2
  6. 6.Junbum Cha, Sanghyuk Chun, Kyungjae Lee, Han-Cheol Cho, Seunghyun Park, Yunsung Lee, and Sungrae Park. Swad: Domain generalization by seeking flat minima. NeurIPS, 34, 2021. 3, 5
  7. 7.Jiahua Dong, Lixu Wang, Zhen Fang, Gan Sun, Shichao Xu, Xiao Wang, and Qi Zhu. Federated class-incremental learning. In CVPR, pages 10164–10173, 2022. 2
  8. 8.Qi Dou, Daniel Coelho de Castro, Konstantinos Kamnitsas, and Ben Glocker. Domain generalization via model-agnostic learning of semantic features. In NeurIPS, pages 6450–6461, 2019. 2
  9. 9.Ziqing Fan, Yanfeng Wang, Jiangchao Yao, Lingjuan Lyu, Ya Zhang, and Qi Tian. Fedskip: Combatting statistical heterogeneity with federated skip aggregation. In ICDM, pages 131–140, 2022. 2
  10. 10.Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur. Sharpness-aware minimization for efficiently improving generalization. In ICLR, 2021. 3
  11. 11.Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew Gordon Wilson. Loss surfaces, mode connectivity, and fast ensembling of dnns. In NeurIPS, 2018. 8
  12. 12.Muhammad Ghifary, David Balduzzi, W Bastiaan Kleijn, and Mengjie Zhang. Scatter component analysis: A unified framework for domain adaptation and domain generalization. IEEE transactions on pattern analysis and machine intelligence, 39(7):1414–1430, 2016. 1
  13. 13.Ishaan Gulrajani and David Lopez-Paz. In search of lost domain generalization. In ICLR, 2021. 5
  14. 14.Yuanxiong Guo, Ying Sun, Rui Hu, and Yanmin Gong. Hybrid local SGD for federated learning with heterogeneous communications. In ICLR, 2022. 2
  15. 15.Jiaxing Huang, Dayan Guan, Aoran Xiao, and Shijian Lu. Fsdr: Frequency space domain randomization for domain generalization. In CVPR, pages 6891–6902, 2021. 2
  16. 16.Zeyi Huang, Haohan Wang, E. Xing, and D. Huang. Self-challenging improves cross-domain generalization. In ECCV, 2020. 2, 5, 6
  17. 17.Meirui Jiang, Zirui Wang, and Qi Dou. Harmofl: Harmonizing local and global drifts in federated learning on heterogeneous medical images. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 1087–1095, 2022. 2, 5, 6
  18. 18.Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh. Scaffold: Stochastic controlled averaging for federated learning. In ICML, pages 5132–5143. PMLR, 2020. 1, 2, 6
  19. 19.Aditya Khosla, Tinghui Zhou, Tomasz Malisiewicz, Alexei A Efros, and Antonio Torralba. Undoing the damage of dataset bias. In ECCV, pages 158–171. Springer, 2012. 1
  20. 20.David Krueger, Ethan Caballero, Joern-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Remi Le Priol, and Aaron Courville. Out-of-distribution generalization via risk extrapolation (rex). In ICML, pages 5815–5826. PMLR, 2021. 3
  21. 21.Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M Hospedales. Deeper, broader and artier domain generalization. In ICCV, pages 5542–5550, 2017. 5
  22. 22.Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M. Hospedales. Learning to generalize: Meta-learning for domain generalization. In AAAI, 2018. 2
  23. 23.Da Li, Jianshu Zhang, Yongxin Yang, Cong Liu, Yi-Zhe Song, and Timothy M Hospedales. Episodic training for domain generalization. In ICCV, pages 1446–1455, 2019. 2
  24. 24.Haoliang Li, Sinno Jialin Pan, Shiqi Wang, and Alex C Kot. Domain generalization with adversarial feature learning. In CVPR, pages 5400–5409, 2018. 2
  25. 25.Qinbin Li, Yiqun Diao, Quan Chen, and Bingsheng He. Federated learning on non-iid data silos: An experimental study. CoRR, abs/2102.02079, 2021. 2
  26. 26.Shenghui Li, Edith Ngai, Fanghua Ye, and Thiemo Voigt. Auto-weighted robust federated learning with corrupted data sources. ACM Transactions on Intelligent Systems and Technology (TIST), 2022. 2, 6
  27. 27.Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems, 2:429–450, 2020. 1, 2, 6
  28. 28.Xiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp, and Qi Dou. Fed{bn}: Federated learning on non-{iid} features via local batch normalization. In ICLR, 2021. 5
  29. 29.Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao. Deep domain generalization via conditional invariant adversarial networks. In ECCV, pages 624–639, 2018. 2
  30. 30.Quande Liu, Cheng Chen, Jing Qin, Qi Dou, and Pheng-Ann Heng. Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous frequency space. CVPR, 2021. 1, 2, 6
  31. 31.Zezhong Ma, Mengying Zhao, Xiaojun Cai, and Zhiping Jia. Fast-convergent federated learning with class-weighted aggregation. Journal of Systems Architecture, 117:102125, 2021. 2, 6
  32. 32.Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics, pages 1273–1282. PMLR, 2017. 1, 2, 3, 4, 6
  33. 33.Krikamol Muandet, David Balduzzi, Bernhard Schölkopf, et al. Domain generalization via invariant feature representation. In ICML, pages 10–18, 2013. 2
  34. 34.Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In ICCV, pages 1406–1415, 2019. 5
  35. 35.Fengchun Qiao, Long Zhao, and Xi Peng. Learning to learn single domain generalization. In CVPR, pages 12556–12565, 2020. 2
  36. 36.Zhe Qu, Xingyu Li, Rui Duan, Yao Liu, Bo Tang, and Zhuo Lu. Generalized federated learning via sharpness aware minimization. In ICML, 2022. 2, 3, 5, 6
  37. 37.Rui Shao, Xiangyuan Lan, Jiawei Li, and Pong C Yuen. Multi-adversarial discriminative deep domain generalization for face presentation attack detection. In CVPR, pages 10023–10031, 2019. 2
  38. 38.Zebang Shen, Juan Cervino, Hamed Hassani, and Alejandro Ribeiro. An agnostic approach to federated learning with class imbalance. In ICLR, 2022. 2
  39. 39.Yuge Shi, Jeffrey Seely, Philip Torr, Siddharth N, Awni Hannun, Nicolas Usunier, and Gabriel Synnaeve. Gradient matching for domain generalization. In ICLR, 2022. 2
  40. 40.Baochen Sun and Kate Saenko. Deep coral: Correlation alignment for deep domain adaptation. In ECCV, pages 443–450. Springer, 2016. 1, 2
  41. 41.Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, and Sethuraman Panchanathan. Deep hashing network for unsupervised domain adaptation. In CVPR, pages 5018–5027, 2017. 5
  42. 42.Riccardo Volpi, Hongseok Namkoong, Ozan Sener, John C Duchi, Vittorio Murino, and Silvio Savarese. Generalizing to unseen domains via adversarial data augmentation. In NeurIPS, pages 5334–5344, 2018. 2
  43. 43.Jingye Wang, Ruoyi Du, Dongliang Chang, and Zhanyu Ma. Domain generalization via frequency-based feature disentanglement and interaction. CoRR, abs/2201.08029, 2022. 6
  44. 44.Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H Vincent Poor. Tackling the objective inconsistency problem in heterogeneous federated optimization. NeurIPS, 33:7611–7623, 2020. 1, 2, 6
  45. 45.Shujun Wang, Lequan Yu, Caizi Li, Chi-Wing Fu, and Pheng-Ann Heng. Learning from extrinsic and intrinsic supervisions for domain generalization. In ECCV, 2020. 2
  46. 46.An Xu, Wenqi Li, Pengfei Guo, Dong Yang, Holger Roth, Ali Hatamizadeh, Can Zhao, Daguang Xu, Heng Huang, and Ziyue Xu. Closing the generalization gap of cross-silo federated medical image segmentation. arXiv preprint arXiv:2203.10144, 2022. 3
  47. 47.Qinwei Xu, Ruipeng Zhang, Ya Zhang, Yanfeng Wang, and Qi Tian. A fourier-based framework for domain generalization. In CVPR, pages 14383–14392, 2021. 2, 5, 6
  48. 48.Rui Ye, Zhenyang Ni, Chenxin Xu, Jianyu Wang, Siheng Chen, and Yonina C Eldar. Fedfm: Anchor-based feature matching for data heterogeneity in federated learning. arXiv preprint arXiv:2210.07615, 2022. 2
  49. 49.Jaehong Yoon, Wonyong Jeong, Giwoong Lee, Eunho Yang, and Sung Ju Hwang. Federated continual learning with weighted inter-client transfer. In ICML, pages 12073–12086. PMLR, 2021. 2
  50. 50.Honglin Yuan, Warren Richard Morningstar, Lin Ning, and Karan Singhal. What do we mean by generalization in federated learning? In ICLR, 2022. 2
  51. 51.Ruipeng Zhang, Qinwei Xu, Chaoqin Huang, Ya Zhang, and Yanfeng Wang. Semi-supervised domain generalization for medical image analysis. In 2022 IEEE 19th International Symposium on Biomedical Imaging (ISBI), pages 1–5. IEEE, 2022. 6
  52. 52.Kaiyang Zhou, Yongxin Yang, Timothy Hospedales, and Tao Xiang. Learning to generate novel domains for domain generalization. In ECCV, 2020. 2
  53. 53.Kaiyang Zhou, Yongxin Yang, Timothy M Hospedales, and Tao Xiang. Deep domain-adversarial image generation for domain generalisation. In AAAI, pages 13025–13032, 2020. 2
  54. 54.Kaiyang Zhou, Yongxin Yang, Yu Qiao, and Tao Xiang. Domain generalization with mixstyle. In ICLR, 2021. 2

Citation

MLA
Zhang, R., et al. “Federated Domain Generalization with Generalization Adjustment”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 3954–63, https://doi.org/10.1109/CVPR52729.2023.00385.
APA
Zhang, R., Xu, Q., Yao, J., Zhang, Y., Tian, Q., & Wang, Y. (2023). Federated Domain Generalization with Generalization Adjustment. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3954–3963. https://doi.org/10.1109/CVPR52729.2023.00385
Chicago
Zhang, R., Q. Xu, J. Yao, Y. Zhang, Q. Tian, and Y. Wang. 2023. “Federated Domain Generalization with Generalization Adjustment”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3954–63. https://doi.org/10.1109/CVPR52729.2023.00385.
Harvard
Zhang, R. et al. (2023) “Federated Domain Generalization with Generalization Adjustment”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 3954–3963. Available at: https://doi.org/10.1109/CVPR52729.2023.00385.
Vancouver
1. Zhang R, Xu Q, Yao J, Zhang Y, Tian Q, Wang Y (2023) Federated Domain Generalization with Generalization Adjustment. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 3954–3963

BibTeX

@inproceedings{Zhang_2023, title={Federated Domain Generalization with Generalization Adjustment}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00385}, DOI={10.1109/cvpr52729.2023.00385}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Zhang, Ruipeng and Xu, Qinwei and Yao, Jiangchao and Zhang, Ya and Tian, Qi and Wang, Yanfeng}, year={2023}, month=June, pages={3954–3963} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE