Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation

An XuWenqi LiPengfei GuoDong YangHolger RothAli HatamizadehCan ZhaoDaguang XuHeng HuangZiyue Xu

article2022CVPR76 citations

Proposes the FedSM framework alongside the SoftPull optimization method to eliminate client drift and close the performance gap between federated learning and centralized training in medical image segmentation by dynamically selecting among personalized and global models at inference time.

Listen

Training deep learning models for medical image segmentation across multiple healthcare institutions requires addressing data scarcity while maintaining strict patient privacy. Federated learning addresses privacy by enabling institutions to train artificial intelligence models collaboratively without sharing raw patient data. However, differences in imaging equipment, patient demographics, and protocols across clinical sites create a phenomenon known as client drift. This discrepancy causes standard federated models to diverge during local training and underperform compared to centralized training, where all patient data is pooled into a single repository. The article presents and evaluates a new collaborative framework called Federated Super Model (FedSM) and an optimization method called SoftPull, designed to eliminate the performance gap between privacy-preserving federated learning and centralized training.

To evaluate this framework, the authors conducted experiments across real-world medical imaging benchmarks, including retinal optic disc and cup segmentation across six clinical datasets and prostate segmentation across six magnetic resonance imaging sources. Rather than attempting to force a single model to fit every diverse clinical site, FedSM develops a collection of personalized models alongside a global baseline and trains an automated model selector. During inference, this selector identifies the specific institutional data profile that most closely matches a new patient image and directs the image to the appropriate personalized model, or falls back to the global model if confidence is low. The authors supported this approach with theoretical mathematical convergence guarantees and compared its performance against standard federated algorithms, single-site local models, and centralized training.

The findings show that FedSM is the first federated learning framework to completely match and occasionally exceed centralized training performance in medical image segmentation. In retinal segmentation benchmarks with high diversity among sites, FedSM achieved an average client accuracy score of approximately 0.891 and an overall global score of 0.903, surpassing standard federated learning by roughly two percentage points and matching centralized training. In prostate segmentation benchmarks with higher inter-site consistency, FedSM similarly matched centralized benchmark levels. Furthermore, the SoftPull personalization method proved superior to existing federated personalization techniques by avoiding both the overfitting common in purely local training and the excessive smoothing found in standard federated averaging.

These results demonstrate that healthcare consortia can deploy high-performing clinical segmentation models without compromising patient privacy or transferring proprietary institutional data. This approach protects compliance, reduces data aggregation risks, and provides smaller clinical sites with enterprise-grade models that they could not train independently. Stakeholders pursuing multi-institutional AI initiatives should evaluate flexible model-selection architectures like FedSM rather than relying on one-size-fits-all federated models. While the framework requires additional communication bandwidth during training and relies on confidence threshold tuning for unseen data, the underlying evidence is robust across diverse datasets. Future implementation efforts should focus on real-world clinical pilot testing and assessing performance when onboarding entirely new healthcare institutions.

arXiv: 2203.10144
Cover for Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation

Abstract

Cross-silo federated learning (FL) has attracted much attention in medical imaging analysis with deep learning in recent years as it can resolve the critical issues of insufficient data, data privacy, and training efficiency. However, there can be a generalization gap between the model trained from FL and the one from centralized training. This important issue comes from the non-iid data distribution of the local data in the participating clients and is well-known as client drift. In this work, we propose a novel training framework FedSM to avoid the client drift issue and successfully close the generalization gap compared with the centralized training for medical image segmentation tasks for the first time. We also propose a novel personalized FL objective formulation and a new method SoftPull to solve it in our proposed framework FedSM. We conduct rigorous theoretical analysis to guarantee its convergence for optimizing the non-convex smooth objective function. Real-world medical image segmentation experiments using deep FL validate the motivations and effectiveness of our proposed method.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Methodology
  • 3.1. New Framework: FedSM
  • 3.1.1 Ensemble
  • 3.1.2 FedSM-extra
  • 3.1.3 FedSM
  • 3.2. New Personalization: SoftPull
  • 3.3. All Together
  • 4. Experiments
  • 4.1. General Results
  • 4.2. Validate Motivation
  • 4.3. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — FedSM super-model framework for avoiding client drift

    model/method

    FedSM replaces the goal of learning one model that fits all client distributions with a federated super model containing three components for KK clients: a global model with weights wgw_g, personalized models with weights wp,1,…,wp,Kw_{p,1},\ldots,w_{p,K}, and a model selector with weights wsw_s. For an input image xx, the segmentation model ff produces h0=f(wg,x)h_0=f(w_g,x) from the global model and hk=f(wp,k,x)h_k=f(w_{p,k},x) from personalized model kk.

    The selector produces normalized scores y^s,k\hat y_{s,k} for a candidate set Ω⊆{0,1,…,K}\Omega\subseteq\{0,1,\ldots,K\}, and the final segmentation output is

    h=∑k∈Ωy^s,khk,∑k∈Ωy^s,k=1. h=\sum_{k\in\Omega}\hat y_{s,k}h_k,\qquad \sum_{k\in\Omega}\hat y_{s,k}=1.

    The practical FedSM training target for an example from client kk is ys=one_hot⁡(k)y_s=\operatorname{one\_hot}(k), so the selector learns to associate each local data distribution with its personalized model. The global model is not assigned a selector class during training. FedSM therefore avoids requiring local models to converge toward a single global solution, which is the source of client drift under strongly non-iid data.

    The paper also evaluates FedSM-extra. After independently training the global and personalized models, FedSM-extra assigns each training example to the model with the lowest segmentation loss and trains the selector with cross-entropy against that one-hot assignment. This can provide a more direct selector target than the client-index target, but requires extra federated communication rounds because the candidate models must be trained first.

  2. Knowl 2 — SoftPull personalized optimization

    model/method

    SoftPull is the personalized federated optimization method used to produce the models in FedSM. Let Dk\mathcal D_k be the data at client kk, let LDk(w)L_{\mathcal D_k}(w) be its loss for model weights ww, and let

    wk∗=arg⁡min⁡wLDk(w)w_k^*=\arg\min_w L_{\mathcal D_k}(w)

    be the client-local optimum. For an interpolation coefficient λ∈[1/K,1]\lambda\in[1/K,1], SoftPull defines the desired personalized optimum for client kk by mixing the local optimum of client kk with the personalized optima of all other clients:

    wp,k∗=λwk∗+(1−λ)1K−1∑k′=1, k′≠kKwp,k′∗.w_{p,k}^*=\lambda w_k^*+(1-\lambda)\frac{1}{K-1}\sum_{k'=1,\,k'\ne k}^{K}w_{p,k'}^*.

    This construction yields the explicit joint objective

    F({wp,k}k=1K)=∑k=1KLDk(1λwp,k−1−λλ1K−1∑k′=1, k′≠kKwp,k′).F(\{w_{p,k}\}_{k=1}^{K})=\sum_{k=1}^{K}L_{\mathcal D_k}\left(\frac{1}{\lambda}w_{p,k}-\frac{1-\lambda}{\lambda}\frac{1}{K-1}\sum_{k'=1,\,k'\ne k}^{K}w_{p,k'}\right).

    After each federated training round, the server applies the SoftPull update

    wp,k←λwp,k+(1−λ)1K−1∑k′=1, k′≠kKwp,k′.w_{p,k}\leftarrow\lambda w_{p,k}+(1-\lambda)\frac{1}{K-1}\sum_{k'=1,\,k'\ne k}^{K}w_{p,k'}.

    At λ=1/K\lambda=1/K, this update becomes the common average used by FedAvg; larger λ\lambda preserves more client-specific information. The paper argues that larger λ\lambda is appropriate when client distributions are less similar, whereas smaller λ\lambda allows a client with limited data to benefit more from the other clients.

  3. Knowl 3 — FedSM training and inference procedures

    algorithm

    FedSM jointly trains the global model, personalized models, and selector in each federated round. The client weight is pk=nk/np_k=n_k/n, where nkn_k is the number of examples at client kk and n=∑k=1Knkn=\sum_{k=1}^{K}n_k. The base optimizer is denoted by OPT⁡\operatorname{OPT}; η\eta and ηs\eta_s are the learning rates for the segmentation models and selector, respectively.

    Input: local datasets Dk\mathcal D_k, rounds RR, clients KK, learning rates η\eta and ηs\eta_s, coefficient λ\lambda, client weights pkp_k
    Initialize global model wgw_g, personalized models wp,kw_{p,k}, selector wsw_s, and optimizer OPT⁡\operatorname{OPT}
    for round r=1,…,Rr=1,\ldots,R do
        Server sends wgw_g, wp,kw_{p,k}, and wsw_s to client kk
        for every client k=1,…,Kk=1,\ldots,K in parallel do
            Set local global copy wg,k←wgw_{g,k}\leftarrow w_g and selector copy ws,k←wsw_{s,k}\leftarrow w_s
            for every minibatch (x,y)(x,y) in Dk\mathcal D_k do
                Update wg,kw_{g,k} with OPT⁡(wg,k,η,∇L(f(wg,k;x),y))\operatorname{OPT}(w_{g,k},\eta,\nabla L(f(w_{g,k};x),y))
                Update wp,kw_{p,k} with OPT⁡(wp,k,η,∇L(f(wp,k;x),y))\operatorname{OPT}(w_{p,k},\eta,\nabla L(f(w_{p,k};x),y))
                Set selector target ys←one_hot⁡(k)y_s\leftarrow\operatorname{one\_hot}(k)
                Update ws,kw_{s,k} with OPT⁡(ws,k,ηs,∇Ls(fs(ws,k;x),ys))\operatorname{OPT}(w_{s,k},\eta_s,\nabla L_s(f_s(w_{s,k};x),y_s))
            end for
            Send wg,kw_{g,k}, wp,kw_{p,k}, and ws,kw_{s,k} to the server
        end for
        Server sets wg←∑k=1Kpkwg,kw_g\leftarrow\sum_{k=1}^{K}p_kw_{g,k} and ws←∑k=1Kpkws,kw_s\leftarrow\sum_{k=1}^{K}p_kw_{s,k}
        for every client kk do
            Apply SoftPull: wp,k←λwp,k+(1−λ)1K−1∑k′≠kwp,k′w_{p,k}\leftarrow\lambda w_{p,k}+(1-\lambda)\frac{1}{K-1}\sum_{k'\ne k}w_{p,k'}
        end for
    end for
    Output: wgw_g, all wp,kw_{p,k}, and wsw_s

    For an inference image xx, the selector computes y^s=fs(ws;x)\hat y_s=f_s(w_s;x). If max⁡ky^s,k>γ\max_k\hat y_{s,k}>\gamma, where γ\gamma is a tuned confidence threshold, FedSM chooses k=arg⁡max⁡ky^s,kk=\arg\max_k\hat y_{s,k} and returns f(wp,k;x)f(w_{p,k};x). Otherwise it returns the global prediction f(wg;x)f(w_g;x). Thus, low-confidence inputs that do not resemble any local distribution are assigned to the global model. Each training round communicates the global model, the selector, and personalized-model information, with the paper summarizing the cost relative to model size as 2wg+ws2w_g+w_s; after training, the complete super model is sent to each client once for inference.

  4. Knowl 4 — Convergence guarantee for SoftPull

    theoretical result

    The paper analyzes SoftPull for the non-convex objective

    F({wp,k}k=1K)=∑k=1KLDk(1λwp,k−1−λλ1K−1∑k′=1, k′≠kKwp,k′),F(\{w_{p,k}\}_{k=1}^{K})=\sum_{k=1}^{K}L_{\mathcal D_k}\left(\frac{1}{\lambda}w_{p,k}-\frac{1-\lambda}{\lambda}\frac{1}{K-1}\sum_{k'=1,\,k'\ne k}^{K}w_{p,k'}\right),

    where KK is the number of clients, wp,k∈Rdw_{p,k}\in\mathbb R^d is the personalized-model parameter vector for client kk, and λ∈[1/K,1]\lambda\in[1/K,1]. The guarantee assumes that every client loss LDkL_{\mathcal D_k} is LL-smooth, meaning

    ∥∇LDk(w1)−∇LDk(w2)∥22≤L∥w1−w2∥22,\|\nabla L_{\mathcal D_k}(w_1)-\nabla L_{\mathcal D_k}(w_2)\|_2^2\le L\|w_1-w_2\|_2^2,

    that the stochastic gradient has bounded variance

    E[∥∇LDk(w,x)−∇LDk(w)∥22]≤σ2,\mathbb E\left[\|\nabla L_{\mathcal D_k}(w,x)-\nabla L_{\mathcal D_k}(w)\|_2^2\right]\le\sigma^2,

    and that the client gradient is bounded by ∥∇LDk(w)∥22≤G2\|\nabla L_{\mathcal D_k}(w)\|_2^2\le G^2. Here xx is a random example from client kk, η\eta is the local learning rate, RR is the number of federated rounds, MM is the number of local update steps per round, and E\mathbb E denotes expectation over stochastic sampling.

    Under these assumptions, the average squared gradient of FF over all clients, rounds, and local steps satisfies the paper's bound

    1KRM∑r=0R−1∑m=0M−1∑k=1KE[∥∇wp,kr,mF∥22]=O(1ηRMλ2+M∑r=0R−1(1−λ)2Rλ2+M2η2∑r=0R−1(1−λ)2Rλ4).\frac{1}{KRM}\sum_{r=0}^{R-1}\sum_{m=0}^{M-1}\sum_{k=1}^{K}\mathbb E\left[\|\nabla_{w_{p,k}^{r,m}}F\|_2^2\right] =\mathcal O\left(\frac{1}{\eta RM\lambda^2}+\frac{M\sum_{r=0}^{R-1}(1-\lambda)^2}{R\lambda^2}+\frac{M^2\eta^2\sum_{r=0}^{R-1}(1-\lambda)^2}{R\lambda^4}\right).

    With η=O((RM)−1/2)\eta=\mathcal O((RM)^{-1/2}) and M=O(R1/3)M=\mathcal O(R^{1/3}), the convergence rate is O((RM)−1/2)\mathcal O((RM)^{-1/2}), with a convergence-error term of order O ⁣(M∑r=0R−1(1−λ)2/(Rλ2))\mathcal O\!\left(M\sum_{r=0}^{R-1}(1-\lambda)^2/(R\lambda^2)\right). The bound explains why stronger pulling toward other clients, corresponding to smaller λ\lambda, can increase the optimization error even though it may improve generalization by reducing local overfitting.

  5. Knowl 5 — Medical federated segmentation experimental protocol

    experimental setup

    The evaluation covers retinal optic-disc-and-cup segmentation from 2D fundus images and prostate segmentation from 3D magnetic-resonance data represented as 2D slices. Each task uses six clients. The retinal clients are Drishti-GS1, RIGA BinRushed, RIGA Magrabia, RIGA MESSIDOR, RIM-ONE, and REFUGE; their train/validation/test counts are respectively (50,25,26)(50,25,26), (98,49,48)(98,49,48), (47,24,23)(47,24,23), (230,115,115)(230,115,115), (80,40,39)(80,40,39), and (400,200,200)(400,200,200), for global totals of (905,453,451)(905,453,451). The prostate clients are I2CVB, MSD, NCI ISBI 3T, NCI ISBI DX, Promise12, and ProstateX; their corresponding counts are (153,77,61)(153,77,61), (404,215,245)(404,215,245), (464,219,198)(464,219,198), (361,162,150)(361,162,150), (609,289,329)(609,289,329), and (1179,582,532)(1179,582,532), for global totals of (3170,1544,1515)(3170,1544,1515).

    All data are randomly split into training, validation, and test sets with proportions 0.5/0.25/0.250.5/0.25/0.25 and resized to 256×256256\times256. The global and personalized segmentation models are 2D U-Nets, while the selector is VGG-11. Local training lasts one epoch per round for 150 rounds; FedSM-extra uses 100 rounds for the global and personalized models followed by 50 additional rounds for the selector. The loss is Dice loss, the test metric is the Dice coefficient, and the base optimizer is Adam with β=(0.9,0.999)\beta=(0.9,0.999). Learning rates and the FedSM confidence threshold are tuned for each method. Each experiment is run three times and the reported values are means. The retinal task has lower inter-client similarity, with differences in position, color, brightness, and background; the prostate task has higher similarity and primarily differs in brightness. Comparisons include centralized training, single-client local training, FedAvg, FedProx, Scaffold, FedSM, and FedSM-extra.

  6. Knowl 6 — Retinal segmentation results under strong non-iid heterogeneity

    data/table

    The retinal experiment compares models on each client's test set, the average of the six client Dice scores, and the Dice score on the union of all client test sets. Dice values below are the reported averages of optic-disc and optic-cup Dice coefficients. FedSM achieves a client-average Dice of 0.89140.8914 and global Dice of 0.90280.9028, slightly exceeding centralized training at 0.88980.8898 and 0.90140.9014, respectively. FedAvg reaches only 0.87100.8710 client-average and 0.89230.8923 global Dice, while FedProx and Scaffold are lower. FedSM-extra is close to FedSM, showing that the simpler client-index selector target is effective.

    Could not parse LaTeX table
  7. Knowl 7 — Prostate segmentation results under higher data similarity

    data/table

    The prostate experiment uses the same per-client, client-average, and global Dice comparisons. The higher similarity between clients reduces the gap between federated and centralized training. FedSM obtains the best client-average Dice among the federated methods, 0.87630.8763, and the best federated global Dice, 0.86920.8692; centralized training gives 0.87370.8737 and 0.86510.8651. FedSM-extra remains competitive at 0.87360.8736 client-average and 0.86730.8673 global Dice.

    Could not parse LaTeX table
  8. Knowl 8 — Model-selector behavior on unseen clients

    empirical result

    To test whether the selector identifies a similar training distribution rather than simply memorizing a participating client, the paper trains FedSM on five retinal clients and evaluates it on the sixth client, rotating the held-out client. With γ=0\gamma=0, the selector must choose among personalized models, so the global model is never selected. The observed selection frequencies are:

    Could not parse LaTeX table

    Here GM denotes the global model and PMjj denotes personalized model jj; N/A indicates that the held-out client's own personalized model is absent. For example, unseen client 6 is assigned mainly to personalized models 5 and 3, consistent with their stronger cross-client performance on client 6. Similar correspondences occur for the other held-out clients, supporting the claim that selector features capture inter-client distribution similarity. Increasing γ\gamma lets low-confidence examples fall back to the global model and improves the held-out-client Dice for clients 5 and 6 by approximately 3%.

  9. Knowl 9 — SoftPull personalization and coefficient ablations

    empirical result

    On the retinal task, SoftPull is compared with local fine-tuning (FT), APFL, Per-FedAvg, and Per-FedMe inside the FedSM framework. SoftPull gives the highest client-average Dice, 0.89140.8914, and global Dice, 0.90280.9028. Its global Dice exceeds APFL's 0.89660.8966 by 0.00620.0062 and the best competing method, FT at 0.89840.8984, by $0.0044.

    Could not parse LaTeX table

    The retinal coefficient sweep gives:

    Could not parse LaTeX table

    The best retinal coefficient is λ=0.7\lambda=0.7, closer to client-specific training in the lower-similarity setting. In the higher-similarity prostate task, the best coefficient is λ=0.3\lambda=0.3, closer to the FedAvg endpoint 1/K=1/61/K=1/6. Loss-surface visualizations further show that local training produces a sharp, overfit optimum, FedAvg produces an excessively flat and underfit optimum, and SoftPull provides tunable intermediate flatness that improves generalization.

  10. Knowl 10 — Threshold fallback is a practical limitation of FedSM

    limitation

    FedSM's selector is trained with client indices 1,…,K1,\ldots,K and has no explicit class for the global model. Consequently, inference needs the heuristic confidence threshold γ\gamma: high-confidence inputs use the highest-scoring personalized model, while low-confidence inputs use the global model. The threshold must be tuned, and the selector cannot directly learn a global-model label from the standard FedSM training target. The paper reports that an appropriate γ\gamma can ensure performance no worse than always using the FedAvg global model, but this is conditional on selecting a suitable threshold. FedSM-extra avoids this missing global class by selecting among trained models using their losses, but requires additional communication rounds.

Coverage note — No substantial contributed component was omitted; lower-level proof details and auxiliary visual examples were excluded because they do not add standalone reconstruction value beyond the stated convergence result and experiments.

References

  1. 1.Nci isbi dataset. https://www.cancerimagingarchive.net/. 5
  2. 2.Promise12 dataset. https://promise12.grand-challenge.org/. 5
  3. 3.Prostatex dataset. https://prostatex.grand-challenge.org/. 5
  4. 4.Refuge dataset. https://refuge.grand-challenge.org/details/. 5
  5. 5.Durmus Alp Emre Acar, Yue Zhao, Ramon Matas, Matthew Mattina, Paul Whatmough, and Venkatesh Saligrama. Federated learning based on dynamic regularization. In International Conference on Learning Representations, 2020. 2
  6. 6.Ahmed Almazroa, Sami Alodhayb, Essameldin Osman, Eslam Ramadan, Mohammed Hummadi, Mohammed Dlaim, Muhannad Alkatee, Kaamran Raahemifar, and Vasudevan Lakshminarayanan. Retinal fundus images for glaucoma analysis: the riga dataset. In Medical Imaging 2018: Imaging Informatics for Healthcare, Research, and Applications, volume 10579, page 105790B. International Society for Optics and Photonics, 2018. 5
  7. 7.Michela Antonelli, Annika Reinke, Spyridon Bakas, Keyvan Farahani, Bennett A Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M Summers, Bram van Ginneken, et al. The medical segmentation decathlon. arXiv preprint arXiv:2106.05735, 2021. 5
  8. 8.Aaron Defazio, Francis Bach, and Simon Lacoste-Julien. Saga: A fast incremental gradient method with support for non-strongly convex composite objectives. Advances in neural information processing systems, 27:1646–1654, 2014. 2
  9. 9.Aaron Defazio and Leon Bottou. On the ineffectiveness ´of variance reduced optimization for deep learning. arXiv preprint arXiv:1812.04529, 2018. 2
  10. 10.Yuyang Deng, Mohammad Mahdi Kamani, and Mehrdad Mahdavi. Adaptive personalized federated learning. arXiv preprint arXiv:2003.13461, 2020. 4, 8
  11. 11.Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar. Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach. Advances in Neural Information Processing Systems, 33, 2020. 2, 8
  12. 12.Francisco Fumero, Silvia Alayon, Jos ´ e L Sanchez, Jose ´ Sigut, and M Gonzalez-Hernandez. Rim-one: An open retinal image database for optic nerve evaluation. In 2011 24th international symposium on computer-based medical systems (CBMS), pages 1–6. IEEE, 2011. 5
  13. 13.Hongchang Gao, An Xu, and Heng Huang. On the convergence of communication-efficient local sgd for federated learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, pages 18–19, 2021. 1
  14. 14.Jonas Geiping, Hartmut Bauermeister, Hannah Droge, and ¨Michael Moeller. Inverting gradients–how easy is it to break privacy in federated learning? arXiv preprint arXiv:2003.14053, 2020. 1
  15. 15.Avishek Ghosh, Jichan Chung, Dong Yin, and Kannan Ramchandran. An efficient framework for clustered federated learning. Advances in Neural Information Processing Systems, 33, 2020. 2
  16. 16.Bin Gu, An Xu, Zhouyuan Huo, Cheng Deng, and Heng Huang. Privacy-preserving asynchronous vertical federated learning algorithms for multiparty collaborative learning. IEEE Transactions on Neural Networks and Learning Systems, 2021. 1
  17. 17.Pengfei Guo, Puyang Wang, Jinyuan Zhou, Shanshan Jiang, and Vishal M Patel. Multi-institutional collaborations for improving deep learning-based magnetic resonance image reconstruction using federated learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2423–2432, 2021. 1
  18. 18.Haowei He, Gao Huang, and Yang Yuan. Asymmetric valleys: Beyond sharp and flat local minima. arXiv preprint arXiv:1902.00744, 2019. 8
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 1
  20. 20.Kevin Hsieh, Amar Phanishayee, Onur Mutlu, and Phillip Gibbons. The non-iid data quagmire of decentralized machine learning. In International Conference on Machine Learning, pages 4387–4398. PMLR, 2020. 1
  21. 21.Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019. 1
  22. 22.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International conference on machine learning, pages 448–456. PMLR, 2015. 2
  23. 23.Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson. Averaging weights leads to wider optima and better generalization. arXiv preprint arXiv:1803.05407, 2018. 8
  24. 24.Yihan Jiang, Jakub Konecnˇ y, Keith Rush, and Sreeram Kannan. Improving federated learning personalization via model agnostic meta learning. arXiv preprint arXiv:1909.12488, 2019. 2
  25. 25.Rie Johnson and Tong Zhang. Accelerating stochastic gradient descent using predictive variance reduction. Advances in neural information processing systems, 26:315–323, 2013. 2
  26. 26.Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurelien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al. Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977, 2019. 1
  27. 27.Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh. Scaffold: Stochastic controlled averaging for federated learning. In International Conference on Machine Learning, pages 5132–5143. PMLR, 2020. 2, 5, 6
  28. 28.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 2
  29. 29.Jakub Konecnˇ y, H Brendan McMahan, Felix X Yu, Peter Richtarik, Ananda Theertha Suresh, and Dave Bacon. Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492, 2016. 1
  30. 30.Guillaume Lemaˆıtre, Robert Mart´ı, Jordi Freixenet, Joan C Vilanova, Paul M Walker, and Fabrice Meriaudeau. Computer-aided detection and diagnosis for prostate cancer based on mono and multi-parametric mri: a review. Computers in biology and medicine, 60:8–31, 2015. 5
  31. 31.Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su. Scaling distributed machine learning with the parameter server. In 11th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 14), pages 583–598, 2014. 1
  32. 32.Tian Li, Shengyuan Hu, Ahmad Beirami, and Virginia Smith. Ditto: Fair and robust federated learning through personalization. In International Conference on Machine Learning, pages 6357–6368. PMLR, 2021. 2
  33. 33.Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018. 2, 6
  34. 34.Tian Li, Maziar Sanjabi, Ahmad Beirami, and Virginia Smith. Fair resource allocation in federated learning. In International Conference on Learning Representations, 2020. 2
  35. 35.Xianfeng Liang, Shuheng Shen, Jingchang Liu, Zhen Pan, Enhong Chen, and Yifei Cheng. Variance reduced local sgd with lower communication complexity. arXiv preprint arXiv:1912.12844, 2019. 2
  36. 36.Quande Liu, Cheng Chen, Jing Qin, Qi Dou, and Pheng-Ann Heng. Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous frequency space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1013–1023, 2021. 1, 2
  37. 37.Yishay Mansour, Mehryar Mohri, Jae Ro, and Ananda Theertha Suresh. Three approaches for personalization with applications to federated learning. arXiv preprint arXiv:2002.10619, 2020. 3, 4
  38. 38.Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics, pages 1273–1282. PMLR, 2017. 1, 6
  39. 39.Mehryar Mohri, Gary Sivek, and Ananda Theertha Suresh. Agnostic federated learning. In International Conference on Machine Learning, pages 4615–4625, 2019. 2
  40. 40.Sashank J Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konecnˇ y, Sanjiv Kumar, and Hugh Brendan McMahan. Adaptive federated optimization. In International Conference on Learning Representations, 2020. 4
  41. 41.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015. 1, 5
  42. 42.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 1, 5
  43. 43.J. Sivaswamy, S. R. Krishnadas, G. Datt Joshi, M. Jain, and A. U. Syed Tabish. Drishti-gs: Retinal image dataset for optic nerve head(onh) segmentation. In 2014 IEEE 11th International Symposium on Biomedical Imaging (ISBI), pages 53–56, April 2014. 5
  44. 44.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014. 2
  45. 45.Canh T Dinh, Nguyen Tran, and Tuan Dung Nguyen. Personalized federated learning with moreau envelopes. Advances in Neural Information Processing Systems, 33, 2020. 2, 8
  46. 46.Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H. Vincent Poor. Tackling the objective inconsistency problem in heterogeneous federated optimization. In Advances in Neural Information Processing Systems, volume 33, pages 7611–7623, 2020. 2
  47. 47.Kangkang Wang, Rajiv Mathews, Chloe Kiddon, Hubert Eichner, Françoise Beaufays, and Daniel Ramage. Federated evaluation of on-device personalization. arXiv preprint arXiv:1910.10252, 2019. 2, 8
  48. 48.An Xu and Heng Huang. Double momentum sgd for federated learning. arXiv preprint arXiv:2102.03970, 2021. 1
  49. 49.An Xu, Zhouyuan Huo, and Heng Huang. On the acceleration of deep learning model parallelism with staleness. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2088–2097, 2020. 1
  50. 50.An Xu, Zhouyuan Huo, and Heng Huang. Step-ahead error feedback for distributed training with compressed gradient. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10478–10486, 2021. 1
  51. 51.Guandao Yang, Tianyi Zhang, Polina Kirichenko, Junwen Bai, Andrew Gordon Wilson, and Chris De Sa. Swalp: Stochastic weight averaging in low precision training. In International Conference on Machine Learning, pages 7015–7024. PMLR, 2019. 8
  52. 52.Hongxu Yin, Arun Mallya, Arash Vahdat, Jose M Alvarez, Jan Kautz, and Pavlo Molchanov. See through gradients: Image batch recovery via gradinversion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16337–16346, 2021. 1
  53. 53.Hao Yu, Rong Jin, and Sen Yang. On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization. In International Conference on Machine Learning, pages 7184–7193. PMLR, 2019. 2
  54. 54.Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. idlg: Improved deep leakage from gradients. arXiv preprint arXiv:2001.02610, 2020. 1
  55. 55.Ligeng Zhu, Zhijian Liu, and Song Han. Deep leakage from gradients. Advances in Neural Information Processing Systems, 32:14774–14784, 2019. 1

Citation

MLA
Xu, A., et al. “Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation”. arXiv, 2022, http://arxiv.org/abs/2203.10144v2.
APA
Xu, A., Li, W., Guo, P., Yang, D., Roth, H., Hatamizadeh, A., Zhao, C., Xu, D., Huang, H., & Xu, Z. (2022). Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation. arXiv. http://arxiv.org/abs/2203.10144v2
Chicago
Xu, A., W. Li, P. Guo, et al. 2022. “Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation”. arXiv. http://arxiv.org/abs/2203.10144v2.
Harvard
Xu, A. et al. (2022) “Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.10144v2.
Vancouver
1. Xu A, Li W, Guo P, Yang D, Roth H, Hatamizadeh A, Zhao C, Xu D, Huang H, Xu Z (2022) Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation. arXiv

BibTeX

@article{xu2022closing,
  title = {Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation},
  author = {Xu, An and Li, Wenqi and Guo, Pengfei and Yang, Dong and Roth, Holger and Hatamizadeh, Ali and Zhao, Can and Xu, Daguang and Huang, Heng and Xu, Ziyue},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.10144v2},
  eprint = {2203.10144}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE