Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation
An XuWenqi LiPengfei GuoDong YangHolger RothAli HatamizadehCan ZhaoDaguang XuHeng HuangZiyue Xu
Proposes the FedSM framework alongside the SoftPull optimization method to eliminate client drift and close the performance gap between federated learning and centralized training in medical image segmentation by dynamically selecting among personalized and global models at inference time.
Training deep learning models for medical image segmentation across multiple healthcare institutions requires addressing data scarcity while maintaining strict patient privacy. Federated learning addresses privacy by enabling institutions to train artificial intelligence models collaboratively without sharing raw patient data. However, differences in imaging equipment, patient demographics, and protocols across clinical sites create a phenomenon known as client drift. This discrepancy causes standard federated models to diverge during local training and underperform compared to centralized training, where all patient data is pooled into a single repository. The article presents and evaluates a new collaborative framework called Federated Super Model (FedSM) and an optimization method called SoftPull, designed to eliminate the performance gap between privacy-preserving federated learning and centralized training.
To evaluate this framework, the authors conducted experiments across real-world medical imaging benchmarks, including retinal optic disc and cup segmentation across six clinical datasets and prostate segmentation across six magnetic resonance imaging sources. Rather than attempting to force a single model to fit every diverse clinical site, FedSM develops a collection of personalized models alongside a global baseline and trains an automated model selector. During inference, this selector identifies the specific institutional data profile that most closely matches a new patient image and directs the image to the appropriate personalized model, or falls back to the global model if confidence is low. The authors supported this approach with theoretical mathematical convergence guarantees and compared its performance against standard federated algorithms, single-site local models, and centralized training.
The findings show that FedSM is the first federated learning framework to completely match and occasionally exceed centralized training performance in medical image segmentation. In retinal segmentation benchmarks with high diversity among sites, FedSM achieved an average client accuracy score of approximately 0.891 and an overall global score of 0.903, surpassing standard federated learning by roughly two percentage points and matching centralized training. In prostate segmentation benchmarks with higher inter-site consistency, FedSM similarly matched centralized benchmark levels. Furthermore, the SoftPull personalization method proved superior to existing federated personalization techniques by avoiding both the overfitting common in purely local training and the excessive smoothing found in standard federated averaging.
These results demonstrate that healthcare consortia can deploy high-performing clinical segmentation models without compromising patient privacy or transferring proprietary institutional data. This approach protects compliance, reduces data aggregation risks, and provides smaller clinical sites with enterprise-grade models that they could not train independently. Stakeholders pursuing multi-institutional AI initiatives should evaluate flexible model-selection architectures like FedSM rather than relying on one-size-fits-all federated models. While the framework requires additional communication bandwidth during training and relies on confidence threshold tuning for unseen data, the underlying evidence is robust across diverse datasets. Future implementation efforts should focus on real-world clinical pilot testing and assessing performance when onboarding entirely new healthcare institutions.
- Paper: SCAFFOLD: Stochastic Controlled Averaging for Federated Learning, Sai Praneeth Karimireddy et al. (2019). This paper establishes the foundational concept and mechanics of client drift under non-IID data distributions, which the source directly aims to counteract.
- Paper: Personalized Federated Learning with Moreau Envelopes, Canh T. Dinh et al. (2020). It provides the mathematical formulation for personalized federated learning objectives using regularized bi-level optimization that informs the source's objective design.
- Paper: Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach, Alireza Fallah et al. (2020). This work introduces key convergence analysis frameworks and formulations for personalized federated learning under non-convex smooth objectives.
- Paper: FedBN: Federated Learning on Non-IID Features via Local Batch Normalization, Xiaoxiao Li et al. (2021). It demonstrates how non-IID feature shifts across institutions impact federated medical image analysis and shows the value of local parameter specialization.
- Paper: The future of digital health with federated learning, Nicola Rieke et al. (2020). This survey motivates the specific cross-silo federated learning domain in healthcare and highlights the real-world challenge of institutional data silos.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). It introduces proximal regularization to stabilize local updates under statistical heterogeneity, serving as a core baseline for addressing client drift.
- Paper: Towards Personalized Federated Learning, Alysa Ziying Tan et al. (2021). This survey provides a taxonomy of personalized federated learning paradigms used to bridge performance gaps under heterogeneous client distributions.
- Paper: Fair Federated Medical Image Segmentation via Client Contribution Estimation, Meirui Jiang et al. (2023). This paper builds upon federated medical image segmentation under cross-silo heterogeneity by incorporating client contribution estimation to optimize collaboration and performance fairness.
- Paper: CD2-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning, Yiqing Shen et al. (2022). It advances model personalization on heterogeneous medical and vision benchmarks through cyclic distillation and channel decoupling across network layers.
- Paper: Federated Domain Generalization with Generalization Adjustment, Ruipeng Zhang et al. (2023). It extends the focus on client distribution shifts by introducing variance reduction regularization to ensure federated domain generalization to unseen clients.
- Paper: FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning, Jianqing Zhang et al. (2024). This work broadens personalized and heterogeneous federated optimization by utilizing trainable global prototypes with adaptive margins.
