Fair Federated Medical Image Segmentation via Client Contribution Estimation

Meirui JiangHolger R. RothWenqi LiDong YangCan ZhaoVishwesh NathDaguang XuQi DouZiyue Xu

article2023CVPR105 citations

Proposes a federated learning framework, FedCE, that simultaneously optimizes collaboration and performance fairness across heterogeneous medical institutions by estimating client contributions through gradient direction differences and auxiliary model prediction errors.

Listen

Collaborative healthcare research increasingly relies on federated learning, a decentralized machine learning technique that enables multiple hospitals to jointly train artificial intelligence models without sharing sensitive patient data. However, medical consortia struggle to maintain long-term participation because institutions lack fair incentives for their contributions, known as collaboration fairness. Concurrently, wide variations in imaging protocols and patient volume across sites cause standard models to perform poorly on underrepresented institutions, violating performance fairness. While existing research addresses these challenges separately, practical deployment requires solving both simultaneously.

The article develops and demonstrates a novel framework called Federated Training via Contribution Estimation (FedCE). This approach aims to accurately quantify each hospital's true marginal value and directly apply these estimates to dynamically balance model updates, achieving fair reward allocation alongside uniform, high-accuracy diagnostic performance across all participating sites.

The researchers designed an efficient contribution estimation mechanism inspired by cooperative game theory. Rather than computing exhaustive combinations of participants, the method evaluates each site against the collective group across two dimensions: optimization direction differences in gradient space and local validation error on an auxiliary model in data space. These metrics are combined using summation or multiplication to dynamically weight client updates during training without requiring raw data sharing or separate validation sets. The framework was evaluated across two multi-institution medical imaging benchmarks—a 6-site retinal fundus dataset and a 6-institution prostate magnetic resonance imaging dataset—and validated through formal convergence and distribution-shift proofs.

The empirical evaluation revealed four central findings. First, FedCE improved overall segmentation accuracy across institutions by over 2% while reducing client performance variance by up to 5.1 points compared to standard federated averaging. Second, the method significantly boosted accuracy for outlier clients with distinct imaging distributions from roughly 40% under standard baselines to 54–57%. Third, the computed client values closely mirrored empirical benchmark contributions, achieving a 93% to 96% correlation with ground-truth leave-one-out experiments. Finally, the framework successfully exposed simulated free riders—clients attempting to obtain global models using duplicated data—within the first 10 training rounds while maintaining estimation stability within 1% during global distribution shifts.

These findings demonstrate that linking fair credit assignment directly to model aggregation improves both equity and core predictive utility. By allocating greater aggregation weight to clinically unique or rare datasets, consortia can eliminate underrepresentation risks and provide transparent, auditable incentive structures. This reduces the risk of institutional attrition, lowers administrative overhead for multi-center initiatives, and enhances diagnostic safety for patients at smaller or specialized clinics.

Decision-makers establishing medical imaging consortia should deploy contribution-based aggregation methods to automatically govern participation rewards and prevent free riding. Leaders can choose between multiplicative or additive combinations based on heterogeneity levels, with additive variants offering slightly greater stability across uniform datasets. Before widespread clinical deployment, consortia should conduct pilot studies on additional imaging modalities and expand the framework to handle corrupt or adversarial client updates.

The study’s primary limitation lies in its focus on clean, controlled benchmark distributions across six institutions per task, leaving extreme client volumes and corrupted or noisy labels for future investigation. Nonetheless, because the underlying theoretical proofs and empirical results align closely across multiple distinct medical imaging tasks, stakeholders can maintain high confidence in the framework's effectiveness for multi-institutional deployments.

No sufficiently relevant recommendations were found.

Cover for Fair Federated Medical Image Segmentation via Client Contribution Estimation

Abstract

How to ensure fairness is an important topic in federated learning (FL). Recent studies have investigated how to reward clients based on their contribution (collaboration fairness), and how to achieve uniformity of performance across clients (performance fairness). Despite achieving progress on either one, we argue that it is critical to consider them together, in order to engage and motivate more diverse clients joining FL to derive a high-quality global model. In this work, we propose a novel method to optimize both types of fairness simultaneously. Specifically, we propose to estimate client contribution in gradient and data space. In gradient space, we monitor the gradient direction differences of each client with respect to others. And in data space, we measure the prediction error on client data using an auxiliary model. Based on this contribution estimation, we propose a FL method, federated training via contribution estimation (FedCE), i.e., using estimation as global model aggregation weights. We have theoretically analyzed our method and empirically evaluated it on two real-world medical datasets. The effectiveness of our approach has been validated with significant performance improvements, better collaboration fairness, better performance fairness, and comprehensive analytical studies. Code is available at https://nvidia.github.io/NVFlare/research/fed-ce

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Fairness in Federated Learning
  • 2.2. Shapley Value based Client Valuation
  • 3. Methods
  • 3.1. Preliminary
  • 3.2. Client Contribution Estimation
  • 3.3. Federated Training via Contribution Estimation - FedCE
  • 3.4. Theoretical Analysis for FedCE
  • 4. Experiment
  • 4.1. Experimental Settings
  • 4.2. Experimental Results
  • 4.3. Analytical Studies
  • 5. Conclusion
  • Acknowledgement
  • References

Knowls

  1. Knowl 1 — Dual-Space Client Contribution Estimation for Federated Learning

    model/method

    Exact Shapley Value (SV) calculation requires evaluating all 2N2^N client coalitions, which is computationally prohibitive in federated learning (FL). To approximate SV efficiently while capturing both optimization dynamics and generalization, client contribution is approximated by evaluating the marginal value of each client i∈[N]i \in [N] directly against the coalition of all other remaining clients N∖{i}N \setminus \{i\} across gradient space and data space.

    In gradient space, the optimization direction difference is measured at communication round kk by computing the cosine similarity between local client gradient ∇Fi(wk,i)\nabla F_i(w_{k,i}) (the difference between local updated parameter wk,iκiw_{k,i}^{\kappa_i} and received global parameter wkw_k) and the aggregated gradient excluding client ii, ∇F(wk−i)\nabla F(w_k^{-i}):

    ∇F(wk−i)=∇F(wk)−pi∇F(wk,i)1−pi,\nabla F(w_k^{-i}) = \frac{\nabla F(w_k) - p_i \nabla F(w_{k,i})}{1 - p_i},

    Γk,i(cos)=1−cos⁡(∇Fi(wk,i),∇F(wk−i)),\Gamma_{k,i}(\text{cos}) = 1 - \cos\left(\nabla F_i(w_{k,i}), \nabla F(w_k^{-i})\right),

    where ∇F(wk)=∑j=1Npj∇Fj(wk,j)\nabla F(w_k) = \sum_{j=1}^N p_j \nabla F_j(w_{k,j}), ∑j=1Npj=1\sum_{j=1}^N p_j = 1, and pi≥0p_i \ge 0 is the baseline importance weight (such as local sample proportion). Γk,i(cos)\Gamma_{k,i}(\text{cos}) is normalized across clients to sum to 1. A client with a different gradient direction receives a higher score.

    In data space, the aggregated model parameters excluding client ii are constructed as wk−i=wk−piwk,i1−piw_k^{-i} = \frac{w_k - p_i w_{k,i}}{1 - p_i} and evaluated on client ii's empirical validation distribution D^i\hat{\mathcal{D}}_i:

    Γk,i(err)=E(D^i;wk−i)=1−Score(D^i;wk−i),\Gamma_{k,i}(\text{err}) = E(\hat{\mathcal{D}}_i; w_k^{-i}) = 1 - \text{Score}(\hat{\mathcal{D}}_i; w_k^{-i}),

    which is also normalized across clients to sum to 1. A higher error indicates that the remaining coalition lacks the representation present in client ii's data.

    The combined round-level contribution is computed using either a multiplication-based (Γk,im\Gamma_{k,i}^m) or summation-based (Γk,is\Gamma_{k,i}^s) mechanism:

    Γk,im=Γk,i(cos)×Γk,i(err),\Gamma_{k,i}^m = \Gamma_{k,i}(\text{cos}) \times \Gamma_{k,i}(\text{err}),

    Γk,is=Γk,i(cos)+Γk,i(err).\Gamma_{k,i}^s = \Gamma_{k,i}(\text{cos}) + \Gamma_{k,i}(\text{err}).

  2. Knowl 2 — FedCE: Federated Training via Contribution Estimation

    algorithm

    Federated Training via Contribution Estimation (FedCE) utilizes dynamically accumulated dual-space client contribution estimations as aggregation weights for global model updates, replacing static sample-size-based weighting to simultaneously promote collaboration fairness (rewarding high contributors) and performance fairness (upweighting unique, underrepresented client data distributions).

    Input: Communication rounds KK, number of clients NN, local datasets {D^i}i=1N\{\hat{\mathcal{D}}_i\}_{i=1}^N, learning rate η\eta, local update steps {κi}i=1N\{\kappa_i\}_{i=1}^N.
    Output: Final global model wKw_K, accumulated client contributions {ρK,i}i=1N\{\rho_{K,i}\}_{i=1}^N.
    Initialize server model w0w_0
    for k=1,…,K−1k = 1, \dots, K-1 do
        Server distributes global model wkw_k to all clients: wk,i0←wkw_{k,i}^0 \leftarrow w_k
        for Client i=1,…,Ni = 1, \dots, N in parallel do
            Compute overall previous update ∇F(wk)=wk−wk−1\nabla F(w_k) = w_k - w_{k-1}
            for j=1,…,κij = 1, \dots, \kappa_i do
                wk,ij←wk,ij−1−η∇Fi(wk,ij−1)w_{k,i}^{j} \leftarrow w_{k,i}^{j-1} - \eta \nabla F_i(w_{k,i}^{j-1})
            end for
            Compute local client gradient ∇Fi(wk,i)=wk,iκi−wk,i0\nabla F_i(w_{k,i}) = w_{k,i}^{\kappa_i} - w_{k,i}^0
            Compute excluded gradient ∇F(wk−i)=∇F(wk)−ρ(k−1,i)∇Fi(wk,i)1−ρ(k−1,i)\nabla F(w_k^{-i}) = \frac{\nabla F(w_k) - \rho_{(k-1,i)} \nabla F_i(w_{k,i})}{1 - \rho_{(k-1,i)}}
            Compute gradient metric Γk,i(cos)=1−cos⁡(∇Fi(wk,i),∇F(wk−i))\Gamma_{k,i}(\text{cos}) = 1 - \cos(\nabla F_i(w_{k,i}), \nabla F(w_k^{-i}))
            Construct excluded model wk−i=wk−ρ(k−1,i)wk,iκi1−ρ(k−1,i)w_k^{-i} = \frac{w_k - \rho_{(k-1,i)} w_{k,i}^{\kappa_i}}{1 - \rho_{(k-1,i)}}
            Compute validation error Γk,i(err)=E(D^i;wk−i)\Gamma_{k,i}(\text{err}) = E(\hat{\mathcal{D}}_i; w_k^{-i})
            Combine metrics into Γk,i=Γk,im\Gamma_{k,i} = \Gamma_{k,i}^m or Γk,is\Gamma_{k,i}^s
            Compute normalized historical weight ρk,i=∑r=1kΓr,i∑j=1N∑r=1kΓr,j\rho_{k,i} = \frac{\sum_{r=1}^k \Gamma_{r,i}}{\sum_{j=1}^N \sum_{r=1}^k \Gamma_{r,j}}
            Send local updated model wk,iκiw_{k,i}^{\kappa_i} and contribution ρk,i\rho_{k,i} to server
        end for
        Server aggregates global model: wk+1←∑i=1Nρk,iwk,iκiw_{k+1} \leftarrow \sum_{i=1}^N \rho_{k,i} w_{k,i}^{\kappa_i}
    eend for
    return wK,{ρK,i}i=1Nw_K, \{\rho_{K,i}\}_{i=1}^N
  3. Knowl 3 — Robustness Upper Bound of Client Contribution Valuation Under Distribution Shift

    theoretical result

    Let Γ\Gamma be a BB-Lipschitz stable valuation function with respect to sample space Z=X×Y\mathcal{Z} = \mathcal{X} \times \mathcal{Y}. Let Ds\mathcal{D}_s and Dt\mathcal{D}_t be two probability distributions supported on Z\mathcal{Z}. For any number of clients N∈NN \in \mathbb{N} and any client i∈[N]i \in [N], the difference in approximated client contribution valuation ν^\hat{\nu} between distribution Ds\mathcal{D}_s and distribution Dt\mathcal{D}_t satisfies:

    ∣ν^(i;Γ,Ds,N)−ν^(i;Γ,Dt,N)∣≤2NB⋅W1(Ds,Dt),|\hat{\nu}(i; \Gamma, \mathcal{D}_s, N) - \hat{\nu}(i; \Gamma, \mathcal{D}_t, N)| \le 2 N B \cdot W_1(\mathcal{D}_s, \mathcal{D}_t),

    where W1(Ds,Dt)W_1(\mathcal{D}_s, \mathcal{D}_t) is the first Wasserstein distance between Ds\mathcal{D}_s and Dt\mathcal{D}_t. This confirms that small distributional variations lead to bounded changes in the estimated client contribution.

  4. Knowl 4 — Convergence Descent Bound and Aggregation Weight Condition for FedCE

    theoretical result

    Assume the federated objective function FF is LL-Lipschitz smooth and local gradient variances are bounded. At communication round k∈[0,K−1]k \in [0, K-1], with learning rate η\eta satisfying ηL−1<0\eta L - 1 < 0, local steps κ(k,i)\kappa_{(k,i)}, and client variance factors β(k,i),δ(k,i)\beta_{(k,i)}, \delta_{(k,i)}, the global objective decrease is bounded by:

    F(wk+1)−F(wk)≤(η4(2ηL+∑i=1Npiρ(k,i)ηA(k,i))−η)∥∇F(wk)∥2,F(w_{k+1}) - F(w_k) \le \left( \frac{\eta}{4} \left(2\eta L + \sum_{i=1}^N p_i \rho_{(k,i)} \eta A_{(k,i)}\right) - \eta \right) \|\nabla F(w_k)\|^2,

    where A(k,i)=ηκ(k,i)(κ(k,i)−1)β(k,i)2δ(k,i)A_{(k,i)} = \eta \sqrt{\kappa_{(k,i)}(\kappa_{(k,i)} - 1)} \beta_{(k,i)}^2 \delta_{(k,i)}, and δ(k,i)\delta_{(k,i)} measures the proportion of local client gradients relative to the global gradient norm:

    δ(k,i)≈∥∑s=0κ−1∇Fi(w(k,i)s)∥2∑s=0κ−1∥∇F(wk)∥2.\delta_{(k,i)} \approx \frac{\|\sum_{s=0}^{\kappa-1} \nabla F_i(w_{(k,i)}^s)\|^2}{\sum_{s=0}^{\kappa-1} \|\nabla F(w_k)\|^2}.

    Under A(k,i)A_{(k,i)}-dominated convergence, the global model converges when client aggregation weights satisfy:

    ρ(k,i)≤4A(k,i)1/2A(k,i)−2L=O(1A(k,i)).\rho_{(k,i)} \le \frac{4 A_{(k,i)}^{1/2}}{A_{(k,i)} - 2L} = \mathcal{O}\left(\frac{1}{\sqrt{A_{(k,i)}}}\right).

    To promote convergence when heterogeneous or boundary clients are under-represented (low δ(k,i)\delta_{(k,i)}), FedCE assigns higher weights ρ(k,i)\rho_{(k,i)} to clients presenting divergent gradient directions or high local exclusion errors.

  5. Knowl 5 — Medical Image Segmentation Performance and Inter-Client Uniformity

    data/table

    FedCE was evaluated against standard federated averaging (FedAvg), standalone client training, and fairness-oriented FL baselines (q-FedAvg, CFFL, FedCI, CGSV) across 6-client splits on two multi-center medical segmentation benchmarks: Retinal Fundus Image Segmentation and Prostate MRI Segmentation. Local update epoch was set to 1, Adam optimizer with initial learning rate 1×10−31\times 10^{-3}, batch size 8, trained for 200 communication rounds using Dice loss. Client 5 in the retinal fundus dataset exhibits severe domain shift compared to other clients.

    Task Retinal Fundus Segmentation Prostate MRI Segmentation
    Client 1 2 3 4 5 6 Avg. Std. 1 2 3 4 5 6 Avg. Std.
    Standalone 86.69 85.51 86.21 89.91 79.77 90.98 86.51 3.95 91.23 84.59 87.57 87.37 86.70 89.25 87.79 2.26
    FedAvg 81.34 85.21 83.28 88.16 40.81 90.79 78.27 18.66 91.10 84.59 89.02 89.09 83.87 89.27 87.82 2.90
    q-FedAvg 86.24 86.97 87.37 89.13 44.68 90.72 80.85 17.80 90.94 85.60 89.28 89.18 84.27 88.67 87.99 2.52
    CFFL 85.72 86.29 86.96 88.62 41.12 90.16 79.81 19.02 91.01 85.49 89.24 88.98 82.11 88.17 87.50 3.20
    FedCI 87.02 86.93 87.35 88.53 40.99 90.22 80.17 19.24 91.21 85.40 89.49 88.37 83.96 88.49 87.82 2.68
    CGSV 83.46 85.57 85.47 88.48 33.79 91.01 77.96 21.80 91.15 84.90 89.27 88.09 83.47 89.16 87.67 2.91
    FedCE (Multi.) 86.73 87.45 87.51 89.26 57.30 90.25 83.08 12.70 91.43 85.79 89.21 89.13 85.68 88.62 88.31 2.22
    FedCE (Sum.) 87.22 87.36 87.93 89.66 54.42 90.92 82.92 14.03 91.18 85.54 89.59 89.22 84.99 88.79 88.22 2.43

    On the highly heterogeneous Retinal Fundus task, existing FL methods degrade significantly on the outlier client (Client 5: 33.79%--44.68% Dice, std up to 21.80), while FedCE improves Client 5's performance to 57.30% (Multi.) and 54.42% (Sum.), boosting overall average Dice to 83.08% and lowering standard deviation across clients to 12.70. On Prostate MRI, FedCE achieves the highest overall accuracy (88.31%) and lowest inter-client standard deviation (2.22).

  6. Knowl 6 — Performance Fairness Metrics Comparison Across FL Methods

    data/table

    Performance fairness evaluates the uniformity and fidelity of federated client performance relative to individual standalone model performance. Pearson correlation (↑\uparrow, higher is better) and Euclidean distance (↓\downarrow, lower is better) between local test Dice scores of standalone models and federated models quantify performance fairness.

    Task Retinal Fundus Segmentation Prostate MRI Segmentation
    Metric Pearson Correlation Euclidean Distance Pearson Correlation Euclidean Distance
    FedAvg 88.82 (1.8×10−21.8 \times 10^{-2}) 38.94 88.67 (1.9×10−21.9 \times 10^{-2}) 3.31
    q-FedAvg 87.02 (2.4×10−22.4 \times 10^{-2}) 38.96 91.69 (1.0×10−21.0 \times 10^{-2}) 2.61
    CFFL 84.53 (3.4×10−13.4 \times 10^{-1}) 38.96 92.47 (8.3×10−38.3 \times 10^{-3}) 2.46
    FedCI 85.69 (2.9×10−22.9 \times 10^{-2}) 48.42 93.57 (6.1×10−36.1 \times 10^{-3}) 2.25
    CGSV 87.47 (2.3×10−22.3 \times 10^{-2}) 49.47 87.71 (2.2×10−22.2 \times 10^{-2}) 3.45
    FedCE (Multi.) 88.94 (1.8×10−21.8 \times 10^{-2}) 24.57 98.25 (4.6×10−44.6 \times 10^{-4}) 1.09
    FedCE (Sum.) 89.11 (1.7×10−21.7 \times 10^{-2}) 24.69 97.15 (1.2×10−31.2 \times 10^{-3}) 1.41

    Values in parentheses denote the statistical pp-value. FedCE achieves the lowest Euclidean distance to standalone performance on both tasks (24.57 vs. 38.94--49.47 on Retinal Fundus; 1.09 vs. 2.25--3.45 on Prostate MRI) and achieves the highest Pearson correlation on both datasets (89.11 on Retinal Fundus; 98.25 on Prostate MRI, p<0.001p < 0.001), confirming improved performance fairness across participating clients.

  7. Knowl 7 — Collaboration Fairness and Contribution Estimation Accuracy

    data/table

    Collaboration fairness ensures that client reward allocations accurately reflect actual contributions to the global model. Valuation accuracy was benchmarked against leave-one-out data valuation ground truth (measuring performance degradation when a specific client is omitted from training). Dynamic client aggregation weights produced by each FL method were compared to leave-one-out values using Pearson correlation (↑\uparrow), Euclidean distance (↓\downarrow), and Cosine similarity (↑\uparrow).

    Task Retinal Fundus Segmentation Prostate MRI Segmentation
    Metric Pearson Corr. Euclidean Dist. Cosine Sim. Pearson Corr. Euclidean Dist. Cosine Sim.
    FedAvg -39.76 (4.4×10−14.4 \times 10^{-1}) 0.55 0.26 3.01 (9.5×10−19.5 \times 10^{-1}) 0.62 0.52
    q-FedAvg 63.28 (1.8×10−11.8 \times 10^{-1}) 0.31 0.57 63.35 (1.8×10−11.8 \times 10^{-1}) 0.59 0.57
    CFFL 0.90 (9.9×10−19.9 \times 10^{-1}) 0.45 0.47 75.44 (8.3×10−28.3 \times 10^{-2}) 0.49 0.74
    FedCI -12.36 (8.2×10−18.2 \times 10^{-1}) 0.37 0.53 -0.31 (1.0×1001.0 \times 10^0) 0.61 0.53
    CGSV -44.50 (3.8×10−13.8 \times 10^{-1}) 0.57 0.22 -1.85 (9.7×10−19.7 \times 10^{-1}) 0.63 0.50
    FedCE (Multi.) 94.93 (3.8×10−33.8 \times 10^{-3}) 0.17 0.82 93.12 (6.9×10−36.9 \times 10^{-3}) 0.49 0.75
    FedCE (Sum.) 96.34 (2.0×10−32.0 \times 10^{-3}) 0.22 0.73 93.53 (6.1×10−36.1 \times 10^{-3}) 0.53 0.69

    Values in parentheses denote the pp-value. FedCE achieves strong positive correlations (>93>93% with p<0.01p < 0.01) and high cosine similarities (0.730.73--0.820.82 on Retinal Fundus, 0.690.69--0.750.75 on Prostate MRI) relative to ground-truth leave-one-out impact, whereas standard FedAvg and baseline valuation methods exhibit poor or negative correlation with true marginal utility.

  8. Knowl 8 — Early Free-Rider Detection in Federated Learning

    model/method

    A 'free rider' in federated learning is a client that lacks sufficient genuine local data and attempts to obtain the trained consensus model by participating with repeated copies of a single image or low-quality data. In FedCE, free riders are detected at early communication rounds without additional validation data by computing a free-rider detection index from variables generated during local contribution estimation:

    Indexfree-rider(i)=cos⁡(∇Fi(wk,i),∇F(wk))×∣E(D^i;wk,i)−E(D^i;wk)∣.\text{Index}_{\text{free-rider}}(i) = \cos\left(\nabla F_i(w_{k,i}), \nabla F(w_k)\right) \times \left| E(\hat{\mathcal{D}}_i; w_{k,i}) - E(\hat{\mathcal{D}}_i; w_k) \right|.

    Because a free rider's local training on duplicated data fails to produce useful gradient steps while suffering distinct error patterns, the detection value spikes significantly. Empirical testing across 6 clients shows that free riders are detected within the first 10 communication rounds, where the designated free rider exhibits values exceeding genuine clients by over 5 times, enabling early exclusion or penalty assignment before training concludes.

  9. Knowl 9 — Ablation of Contribution Estimation Components

    empirical result

    An ablation study evaluated the individual and combined effects of the gradient-space metric Γ(cos)\Gamma(\text{cos}) and data-space validation metric Γ(err)\Gamma(\text{err}) on final segmentation performance across the Retinal Fundus and Prostate MRI benchmarks:

    1. Retinal Fundus Image Segmentation:

      • Γ(cos)\Gamma(\text{cos}) only: 81.61% Dice
      • Γ(err)\Gamma(\text{err}) only: 82.12% Dice
      • FedCE (Multiplication): 82.62% Dice
      • FedCE (Summation): 83.35% Dice
    2. Prostate MRI Segmentation:

      • Γ(cos)\Gamma(\text{cos}) only: 88.17% Dice
      • Γ(err)\Gamma(\text{err}) only: 88.16% Dice
      • FedCE (Multiplication): 88.15% Dice
      • FedCE (Summation): 88.35% Dice

    Solely relying on either gradient direction divergence or data space exclusion error results in lower performance than combining both components. The two metrics serve complementary roles: gradient difference monitors update direction uniqueness, while data space error checks whether the client's data distribution is already adequately covered by the remaining client coalition.

  10. Knowl 10 — Empirical Valuation Invariance Under Client Coalition Shift

    empirical result

    To evaluate the sensitivity of FedCE contribution estimation when the client coalition changes, relative contribution weights were measured on the Retinal Fundus dataset under two setups: the 'original distribution' containing all 6 clients, and the 'distribution excluding Client 5' (the most divergent client) containing the remaining 5 clients.

    After re-normalizing the relative weights among the 5 common clients (Client 1, Client 2, Client 3, Client 4, Client 6), the shift in assigned contribution percentages was consistently under 1% for all clients:

    • Client 1 shift: Δ=0.61%\Delta = 0.61\%
    • Client 2 shift: Δ=0.13%\Delta = 0.13\%
    • Client 3 shift: Δ=0.59%\Delta = 0.59\%
    • Client 4 shift: Δ=0.005%\Delta = 0.005\%
    • Client 6 shift: Δ=0.12%\Delta = 0.12\%

    This empirical result confirms the theoretical Wasserstein-distance stability bound, showing that FedCE provides consistent and reliable credit assignment even when participants join or leave the federated coalition.

Coverage note — All major contributions—including dual-space contribution modeling, the FedCE algorithm, stability and convergence theorems, comprehensive empirical comparisons on collaboration and performance fairness, free-rider identification, and ablation studies—are included; intermediate lemma proofs from the appendix are excluded per guidelines.

References

  1. 1.Nicola Rieke, Jonny Hancox, Wenqi Li, Fausto Milletari, Holger R Roth, Shadi Albarqouni, Spyridon Bakas, et al. The future of digital health with federated learning. NPJ digital medicine, 3(1):1–7, 2020.
  2. 2.Chuhan Wu, Fangzhao Wu, et al. Communication-efficient federated learning via knowledge distillation. Nature communications, 13(1):1–8, 2022.
  3. 3.Micah J Sheller, Brandon Edwards, G Anthony Reina, Jason Martin, Sarthak Pati, et al. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Scientific reports, 10(1):1–12, 2020.
  4. 4.Ittai Dayan, Holger R Roth, et al. Federated learning for predicting clinical outcomes in patients with covid-19. Nature medicine, 27(10):1735–1743, 2021.
  5. 5.Qi Dou, Tiffany Y So, Meirui Jiang, Quande Liu, et al. Federated deep learning for detecting covid-19 lung abnormalities in ct: a privacy-preserving multinational validation study. NPJ digital medicine, 4(1):1–11, 2021.
  6. 6.Sarthak Pati, Ujjwal Baid, Brandon Edwards, Micah Sheller, Shih-Han Wang, G Anthony Reina, Patrick Foley, et al. Federated learning enables big data for rare cancer boundary detection. arXiv preprint arXiv:2204.10836, 2022.
  7. 7.Liangqiong Qu, Yuyin Zhou, Paul Pu Liang, Yingda Xia, Feifei Wang, et al. Rethinking architecture design for tackling data heterogeneity in federated learning. In CVPR, pages 10061–10071, 2022.
  8. 8.Pengfei Guo, Puyang Wang, Jinyuan Zhou, Shanshan Jiang, and Vishal M. Patel. Multi-institutional collaborations for improving deep learning-based magnetic resonance image reconstruction using federated learning. In CVPR, pages 2423–2432, June 2021.
  9. 9.Zirui Zhou, Lingyang Chu, Changxin Liu, Lanjun Wang, Jian Pei, and Yong Zhang. Towards fair federated learning. In KDD, pages 4100–4101, 2021.
  10. 10.Marc Aubreville, Christof A Bertram, Taryn A Donovan, Christian Marzahl, Andreas Maier, and Robert Klopfleisch. A completely annotated whole slide image dataset of canine breast cancer to aid human breast cancer research. Scientific data, 7(1):1–10, 2020.
  11. 11.Xiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp, and Qi Dou. FedBN: Federated learning on non-IID features via local batch normalization. In ICLR, 2021.
  12. 12.Jie Xu, Benjamin S Glicksberg, Chang Su, Peter Walker, Jiang Bian, and Fei Wang. Federated learning for healthcare informatics. Journal of Healthcare Informatics Research, 5(1):1–19, 2021.
  13. 13.Lin Zhang, Li Shen, Liang Ding, Dacheng Tao, and Ling-Yu Duan. Fine-tuning global model via data-free knowledge distillation for non-iid federated learning. In CVPR, pages 10174–10183, June 2022.
  14. 14.Tian Li, Maziar Sanjabi, Ahmad Beirami, and Virginia Smith. Fair resource allocation in federated learning. In ICLR, 2020.
  15. 15.Mehryar Mohri, Gary Sivek, and Ananda Theertha Suresh. Agnostic federated learning. In ICML, pages 4615–4625. PMLR, 2019.
  16. 16.Xiaoxuan Liu, Ben Glocker, Melissa M McCradden, Marzyeh Ghassemi, Alastair K Denniston, and Lauren Oakden-Rayner. The medical algorithmic audit. The Lancet Digital Health, 2022.
  17. 17.Jiawen Kang, Zehui Xiong, Dusit Niyato, Shengli Xie, and Junshan Zhang. Incentive mechanism for reliable federated learning: A joint optimization approach to combining reputation and contract theory. IEEE Internet of Things Journal, 6(6):10700–10714, 2019.
  18. 18.Lingjuan Lyu, Xinyi Xu, Qian Wang, and Han Yu. Collaborative fairness in federated learning. In Federated Learning, pages 189–204. Springer, 2020.
  19. 19.Xinyi Xu, Lingjuan Lyu, Xingjun Ma, Chenglin Miao, Chuan Sheng Foo, and Bryan Kian Hsiang Low. Gradient driven rewards to guarantee fairness in collaborative machine learning. NeurIPS, 34:16104–16117, 2021.
  20. 20.Zeou Hu, Kiarash Shaloudegi, Guojun Zhang, and Yaoliang Yu. Federated learning meets multi-objective optimization. IEEE TNSE, 2022.
  21. 21.Tian Li, Shengyuan Hu, Ahmad Beirami, and Virginia Smith. Ditto: Fair and robust federated learning through personalization. In ICML, pages 6357–6368, 2021.
  22. 22.Lloyd S Shapley. A value for n-person games. Classics in game theory, 69, 1997.
  23. 23.Shuyue Wei, Yongxin Tong, et al. Efficient and fair data valuation for horizontal federated learning. In Federated Learning, pages 139–152. Springer, 2020.
  24. 24.Tianshu Song, Yongxin Tong, and Shuyue Wei. Profit allocation for federated learning. In IEEE Big Data, pages 2577–2586. IEEE, 2019.
  25. 25.Zelei Liu, Yuanyuan Chen, Han Yu, et al. Gtg-shapley: Efficient and accurate participant contribution evaluation in federated learning. ACM TIST, 13(4):1–21, 2022.
  26. 26.Solon Barocas, Moritz Hardt, and Arvind Narayanan. Fairness in machine learning. Nips tutorial, 1:2, 2017.
  27. 27.Shira Mitchell, Eric Potash, et al. Prediction-based decisions and fairness: A catalogue of choices, assumptions, and definitions. arXiv preprint arXiv:1811.07867, 2018.
  28. 28.Yuchen Zeng, Hongxu Chen, and Kangwook Lee. Improving fairness via federated learning. arXiv preprint arXiv:2110.15545, 2021.
  29. 29.Wei Du, Depeng Xu, Xintao Wu, and Hanghang Tong. Fairness-aware agnostic federated learning. In SDM, pages 181–189. SIAM, 2021.
  30. 30.Borja Rodríguez Gálvez, Filip Granqvist, Rogier van Dalen, and Matt Seigel. Enforcing fairness in private federated learning via the modified method of differential multipliers. In NeurIPS Workshop Privacy in Machine Learning, 2021.
  31. 31.Sen Cui, Weishen Pan, Jian Liang, et al. Addressing algorithmic disparity and performance inconsistency in federated learning. NeurIPS, 34:26091–26102, 2021.
  32. 32.Lingyang Chu, Lanjun Wang, Yanjie Dong, et al. Fedfair: Training fair models in cross-silo federated learning. arXiv preprint arXiv:2109.05662, 2021.
  33. 33.Daniel Yue Zhang, Ziyi Kou, and Dong Wang. Fairfl: A fair federated learning approach to reducing demographic bias in privacy-sensitive classification models. In Big Data, pages 1051–1060. IEEE, 2020.
  34. 34.Yuyang Deng, Mohammad Mahdi Kamani, and Mehrdad Mahdavi. Distributionally robust federated averaging. NeurIPS, 33:15111–15122, 2020.
  35. 35.Zhuan Shi, Lan Zhang, Zhenyu Yao, et al. Fedfaim: A model performance-based fair incentive mechanism for federated learning. IEEE Transactions on Big Data, 2022.
  36. 36.Mehryar Mohri, Gary Sivek, and Ananda Theertha Suresh. Agnostic federated learning. In ICML, pages 4615–4625, 2019.
  37. 37.Roger B Myerson. Game theory: analysis of conflict. Harvard university press, 1997.
  38. 38.Amirata Ghorbani and James Zou. Data shapley: Equitable valuation of data for machine learning. In ICML, pages 2242–2251. PMLR, 2019.
  39. 39.Siyi Tang, Amirata Ghorbani, et al. Data valuation for medical imaging using shapley value and application to a large-scale chest x-ray dataset. Scientific reports, 11, 2021.
  40. 40.Ruoxi Jia, David Dao, Boxin Wang, et al. Towards efficient data valuation based on the shapley value. In AISTATS, pages 1167–1176. PMLR, 2019.
  41. 41.Sourav Kumar, A Lakshminarayanan, et al. Towards more efficient data valuation in healthcare federated learning using ensembling. In DeCaF, FAIR workshops, pages 119–129. Springer, 2022.
  42. 42.Tianhao Wang, Johannes Rausch, et al. A principled approach to data valuation for federated learning. In Federated Learning, pages 153–167. Springer, 2020.
  43. 43.Amirata Ghorbani, Michael Kim, and James Zou. A distributional framework for data valuation. In ICML, pages 3535–3544. PMLR, 2020.
  44. 44.Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In AISTATS, pages 1273–1282, 2017.
  45. 45.Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. On the convergence of fedavg on non-iid data. In ICLR, 2019.
  46. 46.Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. In MLSys, 2020.
  47. 47.Sashank J. Reddi, Zachary Charles, Manzil Zaheer, et al. Adaptive federated optimization. In ICLR, 2021.
  48. 48.Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, et al. SCAFFOLD: Stochastic controlled averaging for federated learning. In ICML, 2020.
  49. 49.Qianqian Tong, Guannan Liang, and Jinbo Bi. Effective federated adaptive gradient methods with non-iid decentralized data. arXiv preprint arXiv:2009.06557, 2020.
  50. 50.Jose Ignacio Orlando, Huazhu Fu, João Barbosa Breda, et al. Refuge challenge: A unified framework for evaluating automated methods for glaucoma assessment from fundus photographs. MedIA, 59:101570, 2020.
  51. 51.Geert Litjens, Robert Toth, Wendy van de Ven, Caroline Hoeks, Sjoerd Kerkstra, Bram van Ginneken, Graham Vincent, Gwenael Guillard, Neil Birbeck, Jindang Zhang, et al. Evaluation of prostate segmentation algorithms for mri: the promise12 challenge. MedIA, 18(2):359–373, 2014.
  52. 52.Quande Liu, Qi Dou, Lequan Yu, and Pheng Ann Heng. Ms-net: Multi-site network for improving prostate segmentation with heterogeneous mri data. IEEE TMI, 2020.
  53. 53.Guillaume Lemaître, Robert Martí, Jordi Freixenet, Joan C Vilanova, Paul M Walker, and Fabrice Meriaudeau. Computer-aided detection and diagnosis for prostate cancer based on mono and multi-parametric mri: a review. Computers in biology and medicine, 60:8–31, 2015.
  54. 54.Bloch Nicholas, Madabhushi Anant, Huisman Henkjan, Freymann John, Kirby Justin, et al. Nci-proc. ieee-isbi conf. 2013 challenge: Automated segmentation of prostate structures. The Cancer Imaging Archive, 2015.
  55. 55.Francisco Fumero, Silvia Alayón, José L Sanchez, Jose Sigut, and M Gonzalez-Hernandez. Rim-one: An open retinal image database for optic nerve evaluation. In 2011 24th international symposium on computer-based medical systems (CBMS), pages 1–6. IEEE, 2011.
  56. 56.Jayanthi Sivaswamy, S Krishnadas, Arunava Chakravarty, G Joshi, A Syed Tabish, et al. A comprehensive retinal image dataset for the assessment of glaucoma from the optic nerve head analysis. JSM Biomedical Imaging Data Papers, 2(1):1004, 2015.
  57. 57.Ahmed Almazroa, Sami Alodhayb, Essameldin Osman, Eslam Ramadan, Mohammed Hummadi, Mohammed Dlaim, Muhannad Alkatee, Kaamran Raahemifar, and Vasudevan Lakshminarayanan. Retinal fundus images for glaucoma analysis: the riga dataset. In Medical Imaging 2018: Imaging Informatics for Healthcare, Research, and Applications, volume 10579, pages 55–62. SPIE, 2018.
  58. 58.Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV), pages 565–571. IEEE, 2016.
  59. 59.Abraham Savitzky and Marcel JE Golay. Smoothing and differentiation of data by simplified least squares procedures. Analytical chemistry, 36(8):1627–1639, 1964.
  60. 60.Lingjuan Lyu, Han Yu, and Qiang Yang. Threats to federated learning: A survey. arXiv preprint arXiv:2003.02133, 2020.
  61. 61.Olivier Bousquet and André Elisseeff. Stability and generalization. The Journal of Machine Learning Research, 2:499–526, 2002.

Citation

MLA
Jiang, M., et al. “Fair Federated Medical Image Segmentation via Client Contribution Estimation”. arXiv, 2023, http://arxiv.org/abs/2303.16520v1.
APA
Jiang, M., Roth, H. R., Li, W., Yang, D., Zhao, C., Nath, V., Xu, D., Dou, Q., & Xu, Z. (2023). Fair Federated Medical Image Segmentation via Client Contribution Estimation. arXiv. http://arxiv.org/abs/2303.16520v1
Chicago
Jiang, M., H. R. Roth, W. Li, et al. 2023. “Fair Federated Medical Image Segmentation via Client Contribution Estimation”. arXiv. http://arxiv.org/abs/2303.16520v1.
Harvard
Jiang, M. et al. (2023) “Fair Federated Medical Image Segmentation via Client Contribution Estimation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.16520v1.
Vancouver
1. Jiang M, Roth HR, Li W, Yang D, Zhao C, Nath V, Xu D, Dou Q, Xu Z (2023) Fair Federated Medical Image Segmentation via Client Contribution Estimation. arXiv

BibTeX

@article{jiang2023fair,
  title = {Fair Federated Medical Image Segmentation via Client Contribution Estimation},
  author = {Jiang, Meirui and Roth, Holger R and Li, Wenqi and Yang, Dong and Zhao, Can and Nath, Vishwesh and Xu, Daguang and Dou, Qi and Xu, Ziyue},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.16520v1},
  eprint = {2303.16520}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE