Fair Federated Medical Image Segmentation via Client Contribution Estimation
Meirui JiangHolger R. RothWenqi LiDong YangCan ZhaoVishwesh NathDaguang XuQi DouZiyue Xu
Proposes a federated learning framework, FedCE, that simultaneously optimizes collaboration and performance fairness across heterogeneous medical institutions by estimating client contributions through gradient direction differences and auxiliary model prediction errors.
Collaborative healthcare research increasingly relies on federated learning, a decentralized machine learning technique that enables multiple hospitals to jointly train artificial intelligence models without sharing sensitive patient data. However, medical consortia struggle to maintain long-term participation because institutions lack fair incentives for their contributions, known as collaboration fairness. Concurrently, wide variations in imaging protocols and patient volume across sites cause standard models to perform poorly on underrepresented institutions, violating performance fairness. While existing research addresses these challenges separately, practical deployment requires solving both simultaneously.
The article develops and demonstrates a novel framework called Federated Training via Contribution Estimation (FedCE). This approach aims to accurately quantify each hospital's true marginal value and directly apply these estimates to dynamically balance model updates, achieving fair reward allocation alongside uniform, high-accuracy diagnostic performance across all participating sites.
The researchers designed an efficient contribution estimation mechanism inspired by cooperative game theory. Rather than computing exhaustive combinations of participants, the method evaluates each site against the collective group across two dimensions: optimization direction differences in gradient space and local validation error on an auxiliary model in data space. These metrics are combined using summation or multiplication to dynamically weight client updates during training without requiring raw data sharing or separate validation sets. The framework was evaluated across two multi-institution medical imaging benchmarks—a 6-site retinal fundus dataset and a 6-institution prostate magnetic resonance imaging dataset—and validated through formal convergence and distribution-shift proofs.
The empirical evaluation revealed four central findings. First, FedCE improved overall segmentation accuracy across institutions by over 2% while reducing client performance variance by up to 5.1 points compared to standard federated averaging. Second, the method significantly boosted accuracy for outlier clients with distinct imaging distributions from roughly 40% under standard baselines to 54–57%. Third, the computed client values closely mirrored empirical benchmark contributions, achieving a 93% to 96% correlation with ground-truth leave-one-out experiments. Finally, the framework successfully exposed simulated free riders—clients attempting to obtain global models using duplicated data—within the first 10 training rounds while maintaining estimation stability within 1% during global distribution shifts.
These findings demonstrate that linking fair credit assignment directly to model aggregation improves both equity and core predictive utility. By allocating greater aggregation weight to clinically unique or rare datasets, consortia can eliminate underrepresentation risks and provide transparent, auditable incentive structures. This reduces the risk of institutional attrition, lowers administrative overhead for multi-center initiatives, and enhances diagnostic safety for patients at smaller or specialized clinics.
Decision-makers establishing medical imaging consortia should deploy contribution-based aggregation methods to automatically govern participation rewards and prevent free riding. Leaders can choose between multiplicative or additive combinations based on heterogeneity levels, with additive variants offering slightly greater stability across uniform datasets. Before widespread clinical deployment, consortia should conduct pilot studies on additional imaging modalities and expand the framework to handle corrupt or adversarial client updates.
The study’s primary limitation lies in its focus on clean, controlled benchmark distributions across six institutions per task, leaving extreme client volumes and corrupted or noisy labels for future investigation. Nonetheless, because the underlying theoretical proofs and empirical results align closely across multiple distinct medical imaging tasks, stakeholders can maintain high confidence in the framework's effectiveness for multi-institutional deployments.
- Paper: Closing the Generalization Gap of Cross-silo Federated Medical Image Segmentation, An Xu et al. (2022). This study articulates the core challenge of client drift and the cross-silo generalization gap in multi-institutional medical image segmentation that FedCE directly aims to solve.
- Paper: The future of digital health with federated learning, Nicola Rieke et al. (2020). This foundational review details the collaborative requirements, regulatory dynamics, and data-silo challenges inherent to deploying federated learning across healthcare consortia.
- Paper: FedBN: Federated Learning on Non-IID Features via Local Batch Normalization, Xiaoxiao Li et al. (2021). This paper establishes how inter-scanner feature shifts undermine distributed medical imaging models, providing essential context for handling non-IID client distributions.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). This seminal paper introduces the Federated Averaging algorithm, which serves as the fundamental baseline that FedCE adapts and enhances via contribution-weighted aggregation.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). This work formulates the optimization hurdles and client drift that arise under heterogeneous network settings, framing the performance disparities FedCE addresses.
- Paper: Advances and Open Problems in Federated Learning, Peter Kairouz et al. (2019). This comprehensive survey outlines the open problems of performance fairness, collaboration fairness, and incentive allocation across cross-silo federated learning.
No sufficiently relevant recommendations were found.
