Federated Learning on Non-IID Data Silos: An Experimental Study

Qinbin LiYiqun DiaoQuan ChenBingsheng He

article2021IEEE International Conference on Data Engineering1,534 citations

Presents a systematic experimental benchmark and comprehensive data partitioning framework to evaluate state-of-the-art federated learning algorithms across heterogeneous data silos, revealing that no single method consistently outperforms others under non-IID conditions.

Listen

Data privacy regulations and cross-organizational boundaries increasingly trap valuable datasets in isolated distributed silos. Federated learning enables organizations to collaboratively train machine learning models without transferring raw data across borders or corporate firewalls. However, local data across these silos is rarely identical in distribution. This non-independent and identically distributed (non-IID) data creates training drift and degrades model performance. Until now, systematic evaluations of federated algorithms under realistic, diverse data skews have been lacking due to overly simplistic testing setups.

The article introduces a comprehensive benchmarking framework, NIID-Bench, to systematically evaluate how well leading horizontal federated learning algorithms handle diverse data skews across distributed data silos. It evaluates and compares the accuracy, stability, and operational efficiency of four prominent algorithms: Federated Averaging (FedAvg), FedProx, SCAFFOLD, and FedNova.

To establish a credible evaluation, the article defines six distinct data partitioning strategies covering label distribution skew, feature distribution skew, and quantity skew. The benchmark tests four state-of-the-art algorithms across nine diverse datasets, encompassing six image recognition datasets and three tabular datasets. The baseline evaluations primarily simulate setups across multiple parties, evaluating overall classification accuracy, local epoch sensitivity, party sampling effects, scaling bottlenecks, and computational overhead across controlled communication rounds.

The findings show that no single federated learning algorithm consistently outperforms the others across all non-IID scenarios. Label distribution skew poses the most severe challenge, causing dramatic accuracy collapses across all algorithms when participants hold only a single data category. In contrast, feature distribution skews and pure quantity skews exhibit minimal accuracy degradation relative to homogeneous benchmarks. Across specific settings, FedProx achieves the best accuracy in most label and quantity skew environments, whereas SCAFFOLD excels in feature skew settings. However, algorithms exhibit critical operational trade-offs: FedProx incurs roughly 1.5 to 3 times the computational time per round compared to FedAvg, SCAFFOLD doubles the required communication bandwidth per round and suffers severe performance drops during partial client participation, and standard layer averaging techniques introduce instability into deeper architectures.

These results demonstrate that organizations cannot rely on a single default federated algorithm or assume modern methods universally solve data heterogeneity. For decision-makers, selecting the wrong algorithm risks significant communication and computing expenses, extended development timelines, or outright model failure in deployment. Understanding the dominant type of local data skew is vital before deploying federated infrastructure, as matching algorithm mechanics to specific data imbalances directly governs operational costs and final system accuracy.

Organizations should adopt tailored decision logic: deploy FedProx or standard FedAvg when dealing primarily with quantity imbalances or distribution-based label skews, and apply SCAFFOLD when feature skews dominate and full network participation is guaranteed. Hyperparameters, such as the number of local training epochs, require careful tuning to prevent model drift. Further industry-wide work is recommended to design skew-resistant client sampling techniques, automate parameter tuning, and develop lightweight profiling methods that identify data distributions prior to training without compromising privacy.

A primary limitation of the study is its reliance on synthetic partitions of existing centralized datasets rather than natively generated, complex real-world federated deployments. Additionally, client scaling was evaluated primarily in smaller cohorts and a 100-client simulation. Nevertheless, the experimental rigor and broad range of scenarios provide high confidence in the relative strengths, weaknesses, and structural bottlenecks identified across the tested algorithms.

  • Paper: Model-Contrastive Federated Learning, Qinbin Li et al. (2021). This paper introduces model-contrastive learning (MOON) to directly mitigate local model drift under the non-IID data distributions characterized in the source benchmark.
Cover for Federated Learning on Non-IID Data Silos: An Experimental Study

Abstract

Due to the increasing privacy concerns and data regulations, training data have been increasingly fragmented, forming distributed databases of multiple "data silos" (e.g., within different organizations and countries). To develop effective machine learning services, there is a must to exploit data from such distributed databases without exchanging the raw data. Recently, federated learning (FL) has been a solution with growing interests, which enables multiple parties to collaboratively train a machine learning model without exchanging their local data. A key and common challenge on distributed databases is the heterogeneity of the data distribution among the parties. The data of different parties are usually non-independently and identically distributed (i.e., non-IID). There have been many FL algorithms to address the learning effectiveness under non-IID data settings. However, there lacks an experimental study on systematically understanding their advantages and disadvantages, as previous studies have very rigid data partitioning strategies among parties, which are hardly representative and thorough. In this paper, to help researchers better understand and study the non-IID data setting in federated learning, we propose comprehensive data partitioning strategies to cover the typical non-IID data cases. Moreover, we conduct extensive experiments to evaluate state-of-the-art FL algorithms. We find that non-IID does bring significant challenges in learning accuracy of FL algorithms, and none of the existing state-of-the-art FL algorithms outperforms others in all cases. Our experiments provide insights for future studies of addressing the challenges in "data silos".

Table of Contents

  • I Introduction
  • II Preliminaries
  • II-A Notations
  • II-B FedAvg
  • II-C Effect of Non-IID Data
  • III FL Algorithms on Non-IID Data
  • III-A FedProx
  • III-B FedNova
  • III-C SCAFFOLD
  • III-D Other Studies
  • III-E Motivation of this study
  • IV Simulating Non-IID Data Setting
  • IV-A Research Problems
  • IV-B Label Distribution Skew
  • IV-C Feature Distribution Skew
  • IV-D Quantity Skew
  • IV-E Experiments in Existing Studies
  • V Experiments
  • V-A Overall Accuracy Comparison
  • V-A1 Comparison among different non-IID settings
  • V-A2 Comparison among different algorithms
  • V-A3 Comparison among different tasks
  • V-B Communication Efficiency
  • V-C Robustness to Local Updates
  • V-D Party Sampling
  • V-E Scalability
  • V-F Efficiency
  • V-G Mixed Types of Skew
  • V-H Insights on the Experimental Results
  • VI Future Directions
  • VI-A Opportunities for data management
  • VI-B Opportunities for better FL design
  • VII Related Work
  • VIII Conclusion
  • References
  • -A Training Curves
  • -B Number of Local Epochs
  • -C Party Sampling
  • -D Batch Size
  • -E Model Architectures

Knowls

  1. Knowl 1 — Non-IID Partitioning Strategies for Federated Learning Benchmarking

    model/method

    NIID-Bench defines six primary data partitioning strategies across three main non-IID categories to simulate data heterogeneity across NN distributed clients (parties):

    1. Label Distribution Skew:

      • Quantity-based label imbalance (#C=k\#C = k): Each party is randomly assigned samples from exactly kk distinct classes. The total dataset instances for each class are evenly and disjointly divided among the parties assigned to that class.
      • Distribution-based label imbalance (pk∼DirN(β)p_k \sim \text{Dir}_N(\beta)): For each class kk, a proportion vector pk=(pk,1,…,pk,N)p_k = (p_{k,1}, \dots, p_{k,N}) is sampled from a Dirichlet distribution DirN(β)\text{Dir}_N(\beta) parameterized by concentration β>0\beta > 0. Party jj receives a proportion pk,jp_{k,j} of the total instances of class kk. Smaller β\beta values induce higher label skew.
    2. Feature Distribution Skew:

      • Noise-based feature imbalance (x^∼Gau(σ)\hat{x} \sim \text{Gau}(\sigma)): The global dataset is partitioned uniformly among all NN parties. For party PiP_i (i∈{1,…,N}i \in \{1, \dots, N\}), zero-mean Gaussian noise with client-dependent variance is added to each feature vector: x^=x+ϵ\hat{x} = x + \epsilon, where ϵ∼N(0,σ2⋅iNI)\epsilon \sim \mathcal{N}(0, \sigma^2 \cdot \frac{i}{N} \mathbf{I}) and σ\sigma is a global noise parameter.
      • Synthetic feature imbalance (FCUBE): Features in R3\mathbb{R}^3 are partitioned geometrically across parties such that spatial feature distributions vary while label distributions remain balanced.
      • Real-world feature imbalance (e.g., FEMNIST): Natural partition where data is assigned to parties according to data source identities (e.g., different human writers for handwritten characters).
    3. Quantity Skew (q∼DirN(β)q \sim \text{Dir}_N(\beta)):

      • While the underlying feature and label conditional distributions remain identical across clients, client dataset sizes ∣Di∣|D_i| vary. A distribution vector q=(q1,…,qN)∼DirN(β)q = (q_1, \dots, q_N) \sim \text{Dir}_N(\beta) determines the fraction of the total dataset assigned to client jj (∣Dj∣=qj∣D∣|D_j| = q_j |D|).
    4. Mixed Skews:

      • Label and Feature Skew: Data is first partitioned via distribution-based label imbalance (pk∼DirN(β)p_k \sim \text{Dir}_N(\beta)), followed by client-dependent Gaussian noise injection (x^∼Gau(σ)\hat{x} \sim \text{Gau}(\sigma)).
      • Feature and Quantity Skew: Data is first partitioned via Dirichlet quantity skew (q∼DirN(β)q \sim \text{Dir}_N(\beta)), followed by client-dependent Gaussian noise injection (x^∼Gau(σ)\hat{x} \sim \text{Gau}(\sigma)).
  2. Knowl 2 — Synthetic FCUBE Feature Imbalance Dataset

    model/method

    The FCUBE dataset is a 3D synthetic benchmark designed to isolate feature distribution skew without introducing label distribution imbalance or quantity imbalance across N=4N = 4 federated parties.

    Data Generation and Partitioning Construction:

    1. Let the feature domain be a 3D cube centered at the origin, with points denoted by (x1,x2,x3)∈[−1,1]3(x_1, x_2, x_3) \in [-1, 1]^3.
    2. The binary classification label y∈{0,1}y \in \{0, 1\} is determined solely by the separating plane x1=0x_1 = 0: y={0if x1<01if x1≥0y = \begin{cases} 0 & \text{if } x_1 < 0 \\ 1 & \text{if } x_1 \ge 0 \end{cases}
    3. The coordinate planes x1=0x_1 = 0, x2=0x_2 = 0, and x3=0x_3 = 0 divide the cube into 8 symmetric sub-cubes (octants).
    4. Each of the N=4N=4 parties is assigned exactly two sub-cubes that are point-symmetric with respect to the origin (0,0,0)(0, 0, 0):
      • Party 1: Octants (+,+,+)(+, +, +) and (−,−,−)(-, -, -)
      • Party 2: Octants (+,+,−)(+, +, -) and (−,−,+)(-, -, +)
      • Party 3: Octants (+,−,+)(+, -, +) and (−,+,−)(-, +, -)
      • Party 4: Octants (+,−,−)(+, -, -) and (−,+,+)(-, +, +)

    Because one octant in each pair has x1>0x_1 > 0 (label 1) and the other has x1<0x_1 < 0 (label 0) with equal volume and uniform point density, every party has an exact 1:1 balance of class 0 and class 1 instances, and identical sample counts ∣Di∣=∣D∣4|D_i| = \frac{|D|}{4}. However, the local feature distributions P(x∣Pi)P(x \mid P_i) are mutually disjoint and distinct across all parties.

  3. Knowl 3 — Unified Local Training and Server Aggregation for FedAvg, FedProx, and FedNova

    algorithm

    The federated optimization algorithms FedAvg, FedProx, and FedNova share a common local training and server coordination structure, differentiated by the local loss objective and the server aggregation weighting.

    Let NN be the total number of parties, DiD^i the local dataset of party PiP_i, TT the total communication rounds, EE the number of local epochs, η\eta the learning rate, and St⊆{1,…,N}S_t \subseteq \{1, \dots, N\} the subset of parties sampled in round tt.

    Input: Local datasets D1,…,DND^1, \dots, D^N, total parties NN, communication rounds TT, local epochs EE, learning rate η\eta, FedProx parameter μ≥0\mu \ge 0
    Output: Global model wTw^T
    Server executes:
      initialize global model w0w^0
      for t=0,1,…,T−1t = 0, 1, \dots, T-1 do
        Sample active party subset St⊆{1,…,N}S_t \subseteq \{1, \dots, N\}
        n←∑i∈St∣Di∣n \leftarrow \sum_{i \in S_t} |D^i|
        for each party i∈Sti \in S_t in parallel do
          Δwit,τi←LocalTraining(i,wt)\Delta w_i^t, \tau_i \leftarrow \text{LocalTraining}(i, w^t)
        end for
        if using FedAvg or FedProx then
          wt+1←wt−η∑i∈St∣Di∣nΔwitw^{t+1} \leftarrow w^t - \eta \sum_{i \in S_t} \frac{|D^i|}{n} \Delta w_i^t
        else if using FedNova then
          τeff←∑i∈St∣Di∣nτi\tau_{\text{eff}} \leftarrow \sum_{i \in S_t} \frac{|D^i|}{n} \tau_i
          wt+1←wt−η⋅τeff∑i∈St∣Di∣nτiΔwitw^{t+1} \leftarrow w^t - \eta \cdot \tau_{\text{eff}} \sum_{i \in S_t} \frac{|D^i|}{n \tau_i} \Delta w_i^t
        end if
      end for
      return wTw^T
    LocalTraining(i,wti, w^t):
      wi←wtw_i \leftarrow w^t
      τi←0\tau_i \leftarrow 0
      for epoch k=1,…,Ek = 1, \dots, E do
        for each mini-batch b={x,y}⊆Dib = \{x, y\} \subseteq D^i do
          if using FedAvg or FedNova then
            L(wi;b)=∑(x,y)∈bℓ(wi;x,y)L(w_i; b) = \sum_{(x,y) \in b} \ell(w_i; x, y)
          else if using FedProx then
            L(wi;b)=∑(x,y)∈bℓ(wi;x,y)+μ2∥wi−wt∥2L(w_i; b) = \sum_{(x,y) \in b} \ell(w_i; x, y) + \frac{\mu}{2} \|w_i - w^t\|^2
          end if
          wi←wi−η∇L(wi;b)w_i \leftarrow w_i - \eta \nabla L(w_i; b)
          τi←τi+1\tau_i \leftarrow \tau_i + 1
        end for
      end for
      Δwit←wt−wiη\Delta w_i^t \leftarrow \frac{w^t - w_i}{\eta}
      return Δwit,τi\Delta w_i^t, \tau_i

    In FedProx, μ\mu controls the proximal regularization penalty pulling local updates toward the initial global model wtw^t. In FedNova, τi\tau_i is the number of local SGD steps executed by party ii, and scaling updates by τeffτi\frac{\tau_{\text{eff}}}{\tau_i} eliminates objective inconsistency arising from heterogeneous local computation steps.

  4. Knowl 4 — SCAFFOLD Variance-Reduced Federated Optimization

    algorithm

    SCAFFOLD mitigates client drift under non-IID data by introducing control variates that estimate the drift between the global update direction and local update directions.

    Let cc denote the server control variate, cic_i denote the local control variate for party PiP_i, and St⊆{1,…,N}S_t \subseteq \{1, \dots, N\} denote the sampled parties at round tt.

    Input: Local datasets D1,…,DND^1, \dots, D^N, total parties NN, communication rounds TT, local epochs EE, learning rate η\eta
    Output: Global model wTw^T
    Server executes:
      initialize global model w0w^0, global control variate c0←0c^0 \leftarrow 0
      for t=0,1,…,T−1t = 0, 1, \dots, T-1 do
        Sample active party subset St⊆{1,…,N}S_t \subseteq \{1, \dots, N\}
        n←∑i∈St∣Di∣n \leftarrow \sum_{i \in S_t} |D^i|
        for each party i∈Sti \in S_t in parallel do
          Δwit,Δci←LocalTraining(i,wt,ct)\Delta w_i^t, \Delta c_i \leftarrow \text{LocalTraining}(i, w^t, c^t)
        end for
        wt+1←wt−η∑i∈St∣Di∣nΔwitw^{t+1} \leftarrow w^t - \eta \sum_{i \in S_t} \frac{|D^i|}{n} \Delta w_i^t
        ct+1←ct+1N∑i∈StΔcic^{t+1} \leftarrow c^t + \frac{1}{N} \sum_{i \in S_t} \Delta c_i
      end for
      return wTw^T
    LocalTraining(i,wt,cti, w^t, c^t):
      wi←wtw_i \leftarrow w^t
      τi←0\tau_i \leftarrow 0
      for epoch k=1,…,Ek = 1, \dots, E do
        for each mini-batch b={x,y}⊆Dib = \{x, y\} \subseteq D^i do
          g(wi;b)=∇∑(x,y)∈bℓ(wi;x,y)g(w_i; b) = \nabla \sum_{(x,y) \in b} \ell(w_i; x, y)
          wi←wi−η(g(wi;b)−ci+ct)w_i \leftarrow w_i - \eta (g(w_i; b) - c_i + c^t)
          τi←τi+1\tau_i \leftarrow \tau_i + 1
        end for
      end for
      Δwit←wt−wiη\Delta w_i^t \leftarrow \frac{w^t - w_i}{\eta}
      ci+←ci−ct+1τiη(wt−wi)c_i^+ \leftarrow c_i - c^t + \frac{1}{\tau_i \eta} (w^t - w_i)
      Δci←ci+−ci\Delta c_i \leftarrow c_i^+ - c_i
      ci←ci+c_i \leftarrow c_i^+
      return Δwit,Δci\Delta w_i^t, \Delta c_i

    Because both the model update Δwit\Delta w_i^t and control variate update Δci\Delta c_i must be transmitted between clients and the server in each round, SCAFFOLD incurs exactly double the per-round communication data volume of FedAvg.

  5. Knowl 5 — Experimental Benchmarking Setup on Non-IID Data Silos

    experimental setup

    The experimental evaluation uses 9 datasets (6 image and 3 tabular datasets) with standardized models, optimizers, and training parameters:

    1. Datasets:

      • MNIST: 60,000 train / 10,000 test, 784 features, 10 classes.
      • FMNIST: 60,000 train / 10,000 test, 784 features, 10 classes.
      • CIFAR-10: 50,000 train / 10,000 test, 1,024 features (32×3232\times32 RGB), 10 classes.
      • SVHN: 73,257 train / 26,032 test, 1,024 features, 10 classes.
      • adult: 32,561 train / 16,281 test, 123 features, 2 classes.
      • rcv1: 15,182 train / 5,060 test, 47,236 features, 2 classes.
      • covtype: 435,759 train / 145,253 test, 54 features, 2 classes.
      • FCUBE: 4,000 train / 1,000 test, 3 features, 2 classes.
      • FEMNIST: 341,873 train / 40,832 test, 784 features, 10 classes.
    2. Model Architectures:

      • Image datasets: A Convolutional Neural Network (CNN) consisting of two 5×55\times 5 convolution layers (6 channels, then 16 channels) each followed by 2×22\times 2 max-pooling, and two fully-connected layers with ReLU activations (120 units, then 84 units).
      • Tabular datasets: A Multi-Layer Perceptron (MLP) with three hidden layers containing 32, 16, and 8 units, respectively.
      • Extended models tested on CIFAR-10 include VGG-9 and ResNet-50.
    3. Hyperparameters:

      • Total parties N=10N = 10 by default (except N=4N=4 for FCUBE). Full client participation (St={1,…,N}S_t = \{1, \dots, N\}) per round by default.
      • Default communication rounds T=50T = 50. Local epochs E=10E = 10. Mini-batch size B=64B = 64.
      • Optimizer: SGD with momentum 0.9. Learning rate η=0.1\eta = 0.1 for rcv1 and η=0.01\eta = 0.01 for all other datasets.
      • FedProx parameter μ\mu tuned over {0.001,0.01,0.1,1.0}\{0.001, 0.01, 0.1, 1.0\}.
  6. Knowl 6 — Accuracy Comparison of FL Algorithms across Non-IID Partitions

    data/table

    The table below reports the top-1 test accuracy (mean ±\pm standard deviation over three random trials) of FedAvg, FedProx, SCAFFOLD, and FedNova across diverse non-IID partitioning strategies and homogeneous partitions on standard benchmarks.

    Category Dataset Partitioning FedAvg FedProx SCAFFOLD FedNova
    Label Skew MNIST pk∼Dir(0.5)p_k \sim \text{Dir}(0.5) 98.9%±0.1%98.9\% \pm 0.1\% 98.9%±0.1%98.9\% \pm 0.1\% 99.0%±0.1%99.0\% \pm 0.1\% 98.9%±0.1%98.9\% \pm 0.1\%
    MNIST #C=1\#C = 1 29.8%±7.9%29.8\% \pm 7.9\% 40.9%±23.1%40.9\% \pm 23.1\% 9.9%±0.2%9.9\% \pm 0.2\% 39.2%±22.1%39.2\% \pm 22.1\%
    MNIST #C=2\#C = 2 97.0%±0.4%97.0\% \pm 0.4\% 96.4%±0.3%96.4\% \pm 0.3\% 95.9%±0.3%95.9\% \pm 0.3\% 94.5%±1.5%94.5\% \pm 1.5\%
    MNIST #C=3\#C = 3 98.0%±0.2%98.0\% \pm 0.2\% 97.9%±0.4%97.9\% \pm 0.4\% 96.6%±1.5%96.6\% \pm 1.5\% 98.0%±0.3%98.0\% \pm 0.3\%
    FMNIST pk∼Dir(0.5)p_k \sim \text{Dir}(0.5) 88.1%±0.6%88.1\% \pm 0.6\% 88.1%±0.9%88.1\% \pm 0.9\% 88.4%±0.5%88.4\% \pm 0.5\% 88.5%±0.5%88.5\% \pm 0.5\%
    FMNIST #C=1\#C = 1 11.2%±2.0%11.2\% \pm 2.0\% 28.9%±3.9%28.9\% \pm 3.9\% 12.8%±4.8%12.8\% \pm 4.8\% 14.8%±5.9%14.8\% \pm 5.9\%
    FMNIST #C=2\#C = 2 77.3%±4.9%77.3\% \pm 4.9\% 74.9%±2.6%74.9\% \pm 2.6\% 42.8%±28.7%42.8\% \pm 28.7\% 70.4%±5.1%70.4\% \pm 5.1\%
    FMNIST #C=3\#C = 3 80.7%±1.9%80.7\% \pm 1.9\% 82.5%±1.9%82.5\% \pm 1.9\% 77.7%±3.8%77.7\% \pm 3.8\% 78.9%±3.0%78.9\% \pm 3.0\%
    CIFAR-10 pk∼Dir(0.5)p_k \sim \text{Dir}(0.5) 68.2%±0.7%68.2\% \pm 0.7\% 67.9%±0.7%67.9\% \pm 0.7\% 69.8%±0.7%69.8\% \pm 0.7\% 66.8%±1.5%66.8\% \pm 1.5\%
    CIFAR-10 #C=1\#C = 1 10.0%±0.0%10.0\% \pm 0.0\% 12.3%±2.0%12.3\% \pm 2.0\% 10.0%±0.0%10.0\% \pm 0.0\% 10.0%±0.0%10.0\% \pm 0.0\%
    CIFAR-10 #C=2\#C = 2 49.8%±3.3%49.8\% \pm 3.3\% 50.7%±1.7%50.7\% \pm 1.7\% 49.1%±1.7%49.1\% \pm 1.7\% 46.5%±3.5%46.5\% \pm 3.5\%
    CIFAR-10 #C=3\#C = 3 58.3%±1.2%58.3\% \pm 1.2\% 57.1%±1.2%57.1\% \pm 1.2\% 57.8%±1.4%57.8\% \pm 1.4\% 54.4%±1.1%54.4\% \pm 1.1\%
    SVHN pk∼Dir(0.5)p_k \sim \text{Dir}(0.5) 86.1%±0.7%86.1\% \pm 0.7\% 86.6%±0.9%86.6\% \pm 0.9\% 86.8%±0.3%86.8\% \pm 0.3\% 86.4%±0.6%86.4\% \pm 0.6\%
    SVHN #C=1\#C = 1 11.1%±0.0%11.1\% \pm 0.0\% 19.6%±0.0%19.6\% \pm 0.0\% 6.7%±0.0%6.7\% \pm 0.0\% 10.6%±0.8%10.6\% \pm 0.8\%
    SVHN #C=2\#C = 2 80.2%±0.8%80.2\% \pm 0.8\% 79.3%±0.9%79.3\% \pm 0.9\% 62.7%±11.6%62.7\% \pm 11.6\% 75.4%±4.8%75.4\% \pm 4.8\%
    SVHN #C=3\#C = 3 82.0%±0.7%82.0\% \pm 0.7\% 82.1%±1.0%82.1\% \pm 1.0\% 77.2%±2.0%77.2\% \pm 2.0\% 80.5%±1.2%80.5\% \pm 1.2\%
    adult pk∼Dir(0.5)p_k \sim \text{Dir}(0.5) 78.4%±0.9%78.4\% \pm 0.9\% 80.5%±0.7%80.5\% \pm 0.7\% 76.4%±0.0%76.4\% \pm 0.0\% 52.3%±26.7%52.3\% \pm 26.7\%
    adult #C=1\#C = 1 82.5%±2.2%82.5\% \pm 2.2\% 76.4%±0.0%76.4\% \pm 0.0\% 23.6%±0.0%23.6\% \pm 0.0\% 50.8%±0.9%50.8\% \pm 0.9\%
    rcv1 pk∼Dir(0.5)p_k \sim \text{Dir}(0.5) 48.2%±0.7%48.2\% \pm 0.7\% 70.3%±13.3%70.3\% \pm 13.3\% 64.4%±24.3%64.4\% \pm 24.3\% 49.3%±2.1%49.3\% \pm 2.1\%
    rcv1 #C=1\#C = 1 51.8%±0.7%51.8\% \pm 0.7\% 51.8%±0.7%51.8\% \pm 0.7\% 51.8%±0.7%51.8\% \pm 0.7\% 51.8%±0.7%51.8\% \pm 0.7\%
    covtype pk∼Dir(0.5)p_k \sim \text{Dir}(0.5) 77.2%±7.4%77.2\% \pm 7.4\% 70.9%±0.7%70.9\% \pm 0.7\% 67.7%±14.9%67.7\% \pm 14.9\% 74.8%±12.9%74.8\% \pm 12.9\%
    covtype #C=1\#C = 1 48.8%±0.1%48.8\% \pm 0.1\% 59.1%±2.1%59.1\% \pm 2.1\% 49.6%±1.4%49.6\% \pm 1.4\% 50.4%±1.4%50.4\% \pm 1.4\%
    Feature Skew MNIST x^∼Gau(0.1)\hat{x} \sim \text{Gau}(0.1) 99.1%±0.1%99.1\% \pm 0.1\% 99.1%±0.1%99.1\% \pm 0.1\% 99.1%±0.1%99.1\% \pm 0.1\% 99.1%±0.1%99.1\% \pm 0.1\%
    FMNIST x^∼Gau(0.1)\hat{x} \sim \text{Gau}(0.1) 89.1%±0.3%89.1\% \pm 0.3\% 89.0%±0.2%89.0\% \pm 0.2\% 89.3%±0.0%89.3\% \pm 0.0\% 89.0%±0.1%89.0\% \pm 0.1\%
    CIFAR-10 x^∼Gau(0.1)\hat{x} \sim \text{Gau}(0.1) 68.9%±0.3%68.9\% \pm 0.3\% 69.3%±0.2%69.3\% \pm 0.2\% 70.1%±0.2%70.1\% \pm 0.2\% 68.5%±1.3%68.5\% \pm 1.3\%
    SVHN x^∼Gau(0.1)\hat{x} \sim \text{Gau}(0.1) 88.1%±0.5%88.1\% \pm 0.5\% 88.1%±0.2%88.1\% \pm 0.2\% 88.1%±0.4%88.1\% \pm 0.4\% 88.1%±0.4%88.1\% \pm 0.4\%
    FCUBE synthetic 99.8%±0.2%99.8\% \pm 0.2\% 99.8%±0.0%99.8\% \pm 0.0\% 99.7%±0.3%99.7\% \pm 0.3\% 99.7%±0.1%99.7\% \pm 0.1\%
    FEMNIST real-world 99.4%±0.0%99.4\% \pm 0.0\% 99.3%±0.1%99.3\% \pm 0.1\% 99.4%±0.1%99.4\% \pm 0.1\% 99.3%±0.1%99.3\% \pm 0.1\%
    Quantity Skew MNIST q∼Dir(0.5)q \sim \text{Dir}(0.5) 99.2%±0.1%99.2\% \pm 0.1\% 99.2%±0.1%99.2\% \pm 0.1\% 99.1%±0.1%99.1\% \pm 0.1\% 99.1%±0.1%99.1\% \pm 0.1\%
    FMNIST q∼Dir(0.5)q \sim \text{Dir}(0.5) 89.4%±0.1%89.4\% \pm 0.1\% 89.7%±0.3%89.7\% \pm 0.3\% 88.8%±0.4%88.8\% \pm 0.4\% 86.1%±2.9%86.1\% \pm 2.9\%
    CIFAR-10 q∼Dir(0.5)q \sim \text{Dir}(0.5) 72.0%±0.3%72.0\% \pm 0.3\% 71.2%±0.6%71.2\% \pm 0.6\% 62.4%±4.1%62.4\% \pm 4.1\% 10.0%±0.0%10.0\% \pm 0.0\%
    SVHN q∼Dir(0.5)q \sim \text{Dir}(0.5) 88.3%±1.0%88.3\% \pm 1.0\% 88.4%±0.4%88.4\% \pm 0.4\% 11.0%±7.4%11.0\% \pm 7.4\% 41.3%±21.1%41.3\% \pm 21.1\%
    adult q∼Dir(0.5)q \sim \text{Dir}(0.5) 82.2%±0.1%82.2\% \pm 0.1\% 84.8%±0.2%84.8\% \pm 0.2\% 81.6%±4.5%81.6\% \pm 4.5\% 43.2%±33.9%43.2\% \pm 33.9\%
    rcv1 q∼Dir(0.5)q \sim \text{Dir}(0.5) 96.7%±0.3%96.7\% \pm 0.3\% 96.8%±0.4%96.8\% \pm 0.4\% 49.0%±1.9%49.0\% \pm 1.9\% 51.8%±0.7%51.8\% \pm 0.7\%
    covtype q∼Dir(0.5)q \sim \text{Dir}(0.5) 88.1%±0.2%88.1\% \pm 0.2\% 84.6%±0.2%84.6\% \pm 0.2\% 63.2%±20.8%63.2\% \pm 20.8\% 51.2%±3.2%51.2\% \pm 3.2\%
    Homogeneous MNIST IID 99.1%±0.1%99.1\% \pm 0.1\% 99.1%±0.1%99.1\% \pm 0.1\% 99.2%±0.0%99.2\% \pm 0.0\% 99.1%±0.1%99.1\% \pm 0.1\%
    FMNIST IID 89.6%±0.3%89.6\% \pm 0.3\% 89.5%±0.2%89.5\% \pm 0.2\% 89.7%±0.2%89.7\% \pm 0.2\% 89.4%±0.2%89.4\% \pm 0.2\%
    CIFAR-10 IID 70.4%±0.2%70.4\% \pm 0.2\% 70.2%±0.1%70.2\% \pm 0.1\% 71.5%±0.3%71.5\% \pm 0.3\% 69.5%±1.0%69.5\% \pm 1.0\%
    SVHN IID 88.5%±0.5%88.5\% \pm 0.5\% 88.5%±0.8%88.5\% \pm 0.8\% 88.0%±0.8%88.0\% \pm 0.8\% 88.4%±0.5%88.4\% \pm 0.5\%
    FCUBE IID 99.7%±0.1%99.7\% \pm 0.1\% 99.6%±0.2%99.6\% \pm 0.2\% 99.8%±0.1%99.8\% \pm 0.1\% 99.9%±0.1%99.9\% \pm 0.1\%
    FEMNIST IID 99.3%±0.1%99.3\% \pm 0.1\% 99.4%±0.1%99.4\% \pm 0.1\% 99.4%±0.0%99.4\% \pm 0.0\% 99.3%±0.0%99.3\% \pm 0.0\%
    adult IID 82.6%±0.4%82.6\% \pm 0.4\% 84.8%±0.2%84.8\% \pm 0.2\% 83.8%±2.5%83.8\% \pm 2.5\% 82.6%±0.0%82.6\% \pm 0.0\%
    rcv1 IID 96.8%±0.4%96.8\% \pm 0.4\% 96.6%±0.6%96.6\% \pm 0.6\% 80.9%±27.8%80.9\% \pm 27.8\% 96.6%±0.4%96.6\% \pm 0.4\%
    covtype IID 87.9%±0.1%87.9\% \pm 0.1\% 85.2%±0.0%85.2\% \pm 0.0\% 88.0%±2.3%88.0\% \pm 2.3\% 87.9%±0.2%87.9\% \pm 0.2\%

    Key Empirical Inferences:

    • Extreme Label Skew (#C=1\#C=1): Induces catastrophic accuracy collapse across all algorithms (e.g., reaching random guessing ≈10%\approx 10\% on CIFAR-10 and SVHN). FedProx achieves the highest resilience in single-class partitions on MNIST (40.9%40.9\%), FMNIST (28.9%28.9\%), and SVHN (19.6%19.6\%).
    • Algorithm Dominance: No algorithm universally outperforms the rest. Across label skew, FedProx achieves the best accuracy in 11 settings, FedAvg in 8, SCAFFOLD in 4, and FedNova in 3. Across feature skew, SCAFFOLD leads in 5 settings, FedAvg in 4, FedProx in 3, and FedNova in 2. Across quantity skew, FedProx leads in 5 and FedAvg in 3, while SCAFFOLD and FedNova suffer severe degradation on CIFAR-10, SVHN, adult, and covtype.
  7. Knowl 7 — Decision Tree for Federated Algorithm Selection under Non-IID Data

    model/method

    Based on empirical evaluation across diverse datasets and non-IID conditions, the optimal federated learning algorithm can be selected using a systematic decision tree structured around the prevailing data skew:

    1. Feature Distribution Skew:

      • Select SCAFFOLD (consistently provides superior tracking of client feature shifts using control variates).
    2. Quantity Skew:

      • Select FedProx or FedAvg (SCAFFOLD and FedNova suffer severe optimization degradation and instability under disparate client dataset sizes).
    3. Label Distribution Skew:

      • Quantity-based label imbalance (#C=k\#C = k):
        • Select FedAvg or FedProx (FedProx provides superior stabilization when k=1k=1).
      • Distribution-based label imbalance (pk∼Dir(β)p_k \sim \text{Dir}(\beta)):
        • For Image Datasets: Select SCAFFOLD.
        • For Tabular Datasets: Select FedProx.

    When prior knowledge of client data distributions is unavailable, identifying the non-IID skew profile remains a critical prerequisite for algorithm selection.

  8. Knowl 8 — Failure Modes of SCAFFOLD under Client Partial Participation

    empirical result

    When only a random subset of data silos participates in each federated training round (partial client participation, simulated with N=100N = 100 total clients and a sampling fraction of 0.10.1, i.e., 10 clients per round), the stability of federated algorithms changes substantially compared to full participation:

    1. SCAFFOLD Control-Variate Stagnation:

      • SCAFFOLD fails completely or achieves severely depressed accuracy across all non-IID partitioning settings (e.g., pk∼Dir(0.5)p_k \sim \text{Dir}(0.5), q∼Dir(0.5)q \sim \text{Dir}(0.5), #C∈{1,2,3}\#C \in \{1, 2, 3\}).
      • Mechanism: Each party's local control variate cic_i is updated only when that party is sampled. Under small participation rates (10%10\%), the elapsed time between consecutive updates for a single client is large (1010 rounds on average). The local drift estimates ci−cc_i - c become stale and inaccurate relative to the rapidly evolving global model wtw^t, injecting incorrect gradient corrections that destabilize training.
    2. Variance of FedAvg, FedProx, and FedNova:

      • While FedAvg, FedProx, and FedNova achieve higher accuracy than SCAFFOLD, they experience high round-to-round test accuracy variance under partial participation, driven by shifting sample composition and disparate local update trajectories.
  9. Knowl 9 — Per-Round Computation Time and Communication Volume

    data/table

    The table below reports the per-round computation wall-clock time (in seconds) and network communication payload (in megabytes) per party across representative image and tabular benchmark datasets.

    Metric / Algorithm MNIST CIFAR-10 adult rcv1
    Computation Time (s)
    FedAvg 73s 193s 15s 66s
    FedProx 133s 233s 44s 76s
    SCAFFOLD 77s 197s 14s 66s
    FedNova 73s 189s 17s 65s
    Communication Size (MB)
    FedAvg 1.95 MB 2.73 MB 0.20 MB 66.54 MB
    FedProx 1.95 MB 2.73 MB 0.20 MB 66.54 MB
    SCAFFOLD 3.91 MB 5.46 MB 0.41 MB 133.08 MB
    FedNova 1.95 MB 2.73 MB 0.20 MB 66.54 MB

    Key Resource Characteristics:

    • Computational Overhead: FedProx increases computation time significantly (by 1.2×1.2\times to 2.9×2.9\times over FedAvg) because the proximal term μ2∥w−wt∥2\frac{\mu}{2} \|w - w^t\|^2 must be evaluated and differentiated during every mini-batch gradient descent step. FedNova and SCAFFOLD introduce negligible local compute overhead beyond standard SGD.
    • Communication Overhead: SCAFFOLD doubles the per-round network communication volume (2.0×2.0\times) relative to FedAvg, FedProx, and FedNova across all datasets (e.g., 5.46 MB vs. 2.73 MB on CIFAR-10, 133.08 MB vs. 66.54 MB on rcv1) due to transmitting both model parameters Δwit\Delta w_i^t and control variates Δci\Delta c_i.
  10. Knowl 10 — Impact of Mixed Skews, Local Epochs, and Batch Normalization on FL Stability

    empirical result

    Federated optimization under non-IID data exhibits several critical interactions with training hyperparameters, architectural components, and compound skews:

    1. Performance under Compound/Mixed Skews (CIFAR-10):

      • When combining distribution-based label skew (pk∼Dir(0.5)p_k \sim \text{Dir}(0.5)) with noise-based feature skew (x^∼Gau(0.1)\hat{x} \sim \text{Gau}(0.1)), accuracy drops below any single skew setting: FedAvg falls from 68.2%68.2\% (label skew) and 68.9%68.9\% (feature skew) to 66.1%66.1\%; FedProx drops to 64.8%64.8\%; SCAFFOLD drops to 67.8%67.8\%; FedNova drops to 65.9%65.9\%.
      • When combining feature skew with quantity skew (q∼Dir(0.5)q \sim \text{Dir}(0.5)), FedAvg (69.1%69.1\%) and FedProx (69.2%69.2\%) track pure feature skew performance, while SCAFFOLD (62.2%62.2\%) and FedNova (10.0%10.0\%) suffer severe collapse due to quantity skew vulnerability.
    2. Sensitivity to Local Epoch Count EE:

      • Increasing local epochs (from E∈{10,20,40,80}E \in \{10, 20, 40, 80\}) severely impairs convergence under extreme label skew (#C=1,2\#C = 1, 2). For instance, on CIFAR-10 with #C=2\#C = 2, increasing EE from 20 to 80 drops accuracy across all algorithms. Optimal local epoch count depends strongly on skew type and dataset complexity.
    3. Instability of Batch Normalization Layers:

      • In deeper models containing Batch Normalization (e.g., ResNet-50 vs. VGG-9 on CIFAR-10), training exhibits high instability across communication rounds. Simple server-side averaging of running mean and running variance statistics fails because local statistics reflect heterogeneous client data distributions rather than the global distribution.
    4. Invariance of Batch Size Dynamics:

      • Varying local batch size B∈{16,32,64,128,256}B \in \{16, 32, 64, 128, 256\} produces monotonic scaling behavior identical to centralized deep learning across all four algorithms: smaller batch sizes accelerate convergence per round while larger batch sizes slow training, demonstrating that non-IID data skew does not alter the fundamental role of mini-batch size.

Coverage note — None was omitted; all major empirical findings, partitioning definitions, algorithms, resource evaluations, and stability analyses from the paper are represented.

References

  1. 1.S. AbdulRahman, H. Tout, A. Mourad, and C. Talhi. Fedmccs: multicriteria client selection model for optimal iot federated learning. IEEE Internet of Things Journal, 8(6):4723–4735, 2020.
  2. 2.D. A. E. Acar, Y. Zhao, R. Matas, M. Mattina, P. Whatmough, and V. Saligrama. Federated learning based on dynamic regularization. In International Conference on Learning Representations, 2021.
  3. 3.R. Agrawal and R. Srikant. Privacy-preserving data mining. In Proceedings of the 2000 ACM SIGMOD international conference on Management of data, pages 439–450, 2000.
  4. 4.M. Andreux, J. O. du Terrail, C. Beguier, and E. W. Tramel. Siloed federated learning for multi-centric histopathology datasets. In Domain Adaptation and Representation Transfer, and Distributed and Collaborative Learning, pages 129–139. Springer, 2020.
  5. 5.K. Bonawitz, H. Eichner, W. Grieskamp, D. Huba, A. Ingerman, V. Ivanov, C. M. Kiddon, J. Konečn ý, S. Mazzocchi, B. McMahan, T. V. Overveldt, D. Petrou, D. Ramage, and J. Roselander. Towards federated learning at scale: System design. In SysML, 2019.
  6. 6.S. Caldas, S. M. K. Duddu, P. Wu, T. Li, J. Konečn ý, H. B. McMahan, V. Smith, and A. Talwalkar. Leaf: A benchmark for federated settings. arXiv preprint arXiv:1812.01097, 2018.
  7. 7.S. Chaudhuri, R. Motwani, and V. Narasayya. Random sampling for histogram construction: How much is enough? ACM SIGMOD Record, 27(2):436–447, 1998.
  8. 8.G. Cohen, S. Afshar, J. Tapson, and A. Van Schaik. Emnist: Extending mnist to handwritten letters. In 2017 International Joint Conference on Neural Networks (IJCNN), pages 2921–2926. IEEE, 2017.
  9. 9.Z. Dai, B. K. H. Low, and P. Jaillet. Federated bayesian optimization via thompson sampling. Advances in Neural Information Processing Systems, 33, 2020.
  10. 10.Y. Deng, M. M. Kamani, and M. Mahdavi. Distributionally robust federated averaging. Advances in Neural Information Processing Systems, 33, 2020.
  11. 11.Diemert Eustache, Meynet Julien, P. Galland, and D. Lefortier. Attribution modeling increases efficiency of bidding in display advertising. In Proceedings of the AdKDD and TargetAd Workshop, KDD, Halifax, NS, Canada, August, 14, 2017, page To appear. ACM, 2017.
  12. 12.J. Ding, U. F. Minhas, J. Yu, C. Wang, J. Do, Y. Li, H. Zhang, B. Chandramouli, J. Gehrke, D. Kossmann, D. Lomet, and T. Kraska. Alex: An updatable adaptive learned index. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, SIGMOD ’20, page 969–984, New York, NY, USA, 2020. Association for Computing Machinery.
  13. 13.C. T. Dinh, N. H. Tran, and T. D. Nguyen. Personalized federated learning with moreau envelopes. Advances in Neural Information Processing Systems, 2020.
  14. 14.C. Dwork. Differential privacy. Encyclopedia of Cryptography and Security, pages 338–340, 2011.
  15. 15.A. Fallah, A. Mokhtari, and A. Ozdaglar. Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach. Advances in Neural Information Processing Systems, 33, 2020.
  16. 16.X. Y. Felix, A. S. Rawat, A. K. Menon, and S. Kumar. Federated learning with only positive labels. arXiv preprint arXiv:2004.10342, 2020.
  17. 17.M. Fredrikson, S. Jha, and T. Ristenpart. Model inversion attacks that exploit confidence information and basic countermeasures. In Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security, pages 1322–1333. ACM, 2015.
  18. 18.S. Ganguly, P. B. Gibbons, Y. Matias, and A. Silberschatz. Bifocal sampling for skew-resistant join size estimation. In Proceedings of the 1996 ACM SIGMOD International Conference on Management of Data, SIGMOD ’96, page 271–281, New York, NY, USA, 1996. Association for Computing Machinery.
  19. 19.R. C. Geyer, T. Klein, and M. Nabi. Differentially private federated learning: A client level perspective. arXiv preprint arXiv:1712.07557, 2017.
  20. 20.A. C. Gilbert, S. Guha, P. Indyk, Y. Kotidis, S. Muthukrishnan, and M. J. Strauss. Fast, small-space algorithms for approximate histogram maintenance. In Proceedings of the thiry-fourth annual ACM symposium on Theory of computing, pages 389–398, 2002.
  21. 21.N. Guha, A. Talwlkar, and V. Smith. One-shot federated learning. arXiv preprint arXiv:1902.11175, 2019.
  22. 22.F. Hanzely, S. Hanzely, S. Horvath, and P. Richt árik. Lower bounds and optimal algorithms for personalized federated learning. Advances in Neural Information Processing Systems, 2020.
  23. 23.A. Hard, K. Rao, R. Mathews, S. Ramaswamy, F. Beaufays, S. Augenstein, H. Eichner, C. Kiddon, and D. Ramage. Federated learning for mobile keyboard prediction. arXiv preprint arXiv:1811.03604, 2018.
  24. 24.S. Hasan, S. Thirumuruganathan, J. Augustine, N. Koudas, and G. Das. Deep learning models for selectivity estimation of multi-attribute queries. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, SIGMOD ’20, page 1035–1050, New York, NY, USA, 2020. Association for Computing Machinery.
  25. 25.C. He, M. Annavaram, and S. Avestimehr. Group knowledge transfer: Federated learning of large cnns at the edge. Advances in Neural Information Processing Systems, 33, 2020.
  26. 26.C. He, S. Li, J. So, M. Zhang, H. Wang, X. Wang, P. Vepakomma, A. Singh, H. Qiu, L. Shen, et al. Fedml: A research library and benchmark for federated machine learning. arXiv preprint arXiv:2007.13518, 2020.
  27. 27.K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  28. 28.T.-M. H. Hsu, H. Qi, and M. Brown. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019.
  29. 29.S. Hu, Y. Li, X. Liu, Q. Li, Z. Wu, and B. He. The oarf benchmark suite: Characterization and implications for federated learning systems. arXiv preprint arXiv:2006.07856, 2020.
  30. 30.J. Huang. Maximum likelihood estimation of dirichlet distribution parameters. CMU Technique Report, 2005.
  31. 31.N. Hynes, D. Dao, D. Yan, R. Cheng, and D. Song. A demonstration of sterling: A privacy-preserving data marketplace. Proceedings of the VLDB Endowment, 11(12):2086–2089, 2018.
  32. 32.R. Johnson and T. Zhang. Accelerating stochastic gradient descent using predictive variance reduction. Advances in neural information processing systems, 26:315–323, 2013.
  33. 33.P. Kairouz, H. B. McMahan, B. Avent, A. Bellet, M. Bennis, A. N. Bhagoji, K. Bonawitz, Z. Charles, G. Cormode, R. Cummings, et al. Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977, 2019.
  34. 34.G. A. Kaissis, M. R. Makowski, D. Ruckert, and R. F. Braren. Secure, privacy-preserving and federated machine learning in medical imaging. Nature Machine Intelligence, pages 1–7, 2020.
  35. 35.S. P. Karimireddy, S. Kale, M. Mohri, S. J. Reddi, S. U. Stich, and A. T. Suresh. Scaffold: Stochastic controlled averaging for on-device federated learning. In Proceedings of the 37th International Conference on Machine Learning. PMLR, 2020.
  36. 36.A. Krizhevsky, G. Hinton, et al. Learning multiple layers of features from tiny images. 2009.
  37. 37.Y. Kwon, M. Balazinska, B. Howe, and J. Rolia. Skew-resistant parallel processing of feature-extracting scientific user-defined functions. In Proceedings of the 1st ACM Symposium on Cloud Computing, SoCC ’10, page 75–86, New York, NY, USA, 2010. Association for Computing Machinery.
  38. 38.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  39. 39.Q. Li, Y. Diao, Q. Chen, and B. He. Federated learning on non-iid data silos: An experimental study. arXiv preprint arXiv:2102.02079, 2021.
  40. 40.Q. Li, B. He, and D. Song. Model-contrastive federated learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021.
  41. 41.Q. Li, B. He, and D. Song. Practical one-shot federated learning for cross-silo setting. IJCAI, 2021.
  42. 42.Q. Li, Z. Wen, and B. He. Practical federated gradient boosting decision trees. In AAAI, pages 4642–4649, 2020.
  43. 43.Q. Li, Z. Wen, Z. Wu, S. Hu, N. Wang, and B. He. A survey on federated learning systems: Vision, hype and reality for data privacy and protection. arXiv preprint arXiv:1907.09693, 2019.
  44. 44.T. Li, A. K. Sahu, A. Talwalkar, and V. Smith. Federated learning: Challenges, methods, and future directions. arXiv preprint arXiv:1908.07873, 2019.
  45. 45.T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith. Federated optimization in heterogeneous networks. In MLSys, 2020.
  46. 46.T. Li, J. Zhong, J. Liu, W. Wu, and C. Zhang. Ease.ml: Towards multi-tenant resource sharing for machine learning workloads. 11(5):607–620, Jan. 2018.
  47. 47.X. Li, K. Huang, W. Yang, S. Wang, and Z. Zhang. On the convergence of fedavg on non-iid data. In International Conference on Learning Representations, 2020.
  48. 48.X. Li, M. JIANG, X. Zhang, M. Kamp, and Q. Dou. Fed{bn}: Federated learning on non-{iid} features via local batch normalization. In International Conference on Learning Representations, 2021.
  49. 49.Y. Liang, Y. Guo, Y. Gong, C. Luo, J. Zhan, and Y. Huang. An isolated data island benchmark suite for federated learning. arXiv preprint arXiv:2008.07257, 2020.
  50. 50.T. Lin, L. Kong, S. U. Stich, and M. Jaggi. Ensemble distillation for robust model fusion in federated learning. Advances in Neural Information Processing Systems, 33, 2020.
  51. 51.L. Liu, F. Zhang, J. Xiao, and C. Wu. Evaluation framework for large-scale federated learning. arXiv preprint arXiv:2003.01575, 2020.
  52. 52.Y. Liu, Y. Kang, C. Xing, T. Chen, and Q. Yang. A secure federated transfer learning framework. IEEE Intelligent Systems, 2020.
  53. 53.L. v. d. Maaten and G. Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(Nov):2579–2605, 2008.
  54. 54.R. Marcus, A. Kipf, A. van Renen, M. Stoian, S. Misra, A. Kemper, T. Neumann, and T. Kraska. Benchmarking learned indexes. Proc. VLDB Endow., 14(1):1–13, Sept. 2020.
  55. 55.R. Marcus, P. Negi, H. Mao, C. Zhang, M. Alizadeh, T. Kraska, O. Papaemmanouil, and N. Tatbul. Neo: A learned query optimizer. Proc. VLDB Endow., 12(11):1705–1718, July 2019.
  56. 56.H. B. McMahan, E. Moore, D. Ramage, S. Hampson, et al. Communication-efficient learning of deep networks from decentralized data. arXiv preprint arXiv:1602.05629, 2016.
  57. 57.M. Mohri, G. Sivek, and A. T. Suresh. Agnostic federated learning. In International Conference on Machine Learning, pages 4615–4625. PMLR, 2019.
  58. 58.Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng. Reading digits in natural images with unsupervised feature learning. 2011.
  59. 59.J. Neyman. On the two different aspects of the representative method: the method of stratified sampling and the method of purposive selection. In Breakthroughs in statistics, pages 123–150. Springer, 1992.
  60. 60.C. Niu, Z. Zheng, F. Wu, X. Gao, and G. Chen. Trading data in good faith: Integrating truthfulness and privacy preservation in data markets. In 2017 IEEE 33rd International Conference on Data Engineering (ICDE), pages 223–226. IEEE, 2017.
  61. 61.X. Peng, Z. Huang, Y. Zhu, and K. Saenko. Federated adversarial domain adaptation. In International Conference on Learning Representations, 2020.
  62. 62.S. Reddi, Z. Charles, M. Zaheer, Z. Garrett, K. Rush, J. Konečn ý, S. Kumar, and H. B. McMahan. Adaptive federated optimization. arXiv preprint arXiv:2003.00295, 2020.
  63. 63.A. Reisizadeh, F. Farnia, R. Pedarsani, and A. Jadbabaie. Robust federated learning: The case of affine distribution shifts. Advances in Neural Information Processing Systems, 2020.
  64. 64.S. J. Rizvi and J. R. Haritsa. Maintaining data privacy in association rule mining. In VLDB’02: Proceedings of the 28th International Conference on Very Large Databases, pages 682–693. Elsevier, 2002.
  65. 65.M. Schmidt, N. Le Roux, and F. Bach. Minimizing finite sums with the stochastic average gradient. Mathematical Programming, 162(1-2):83–112, 2017.
  66. 66.S. Shastri, V. Banakar, M. Wasserman, A. Kumar, and V. Chidambaram. Understanding and benchmarking the impact of gdpr on database systems. Proc. VLDB Endow., 13(7):1064–1077, Mar. 2020.
  67. 67.A. P. Sheth and J. A. Larson. Federated database systems for managing distributed, heterogeneous, and autonomous databases. ACM Computing Surveys (CSUR), 22(3):183–236, 1990.
  68. 68.R. Shokri, M. Stronati, C. Song, and V. Shmatikov. Membership inference attacks against machine learning models. In 2017 IEEE Symposium on Security and Privacy (SP), pages 3–18. IEEE, 2017.
  69. 69.M. J. Smith, C. Sala, J. M. Kanter, and K. Veeramachaneni. The machine learning bazaar: Harnessing the ml ecosystem for effective system development. In Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data, SIGMOD ’20, page 785–800, New York, NY, USA, 2020. Association for Computing Machinery.
  70. 70.P. Voigt and A. Von dem Bussche. The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing, 2017.
  71. 71.H. Wang, M. Yurochkin, Y. Sun, D. Papailiopoulos, and Y. Khazaeni. Federated learning with matched averaging. In International Conference on Learning Representations, 2020.
  72. 72.J. Wang, Q. Liu, H. Liang, G. Joshi, and H. V. Poor. Tackling the objective inconsistency problem in heterogeneous federated optimization. Advances in Neural Information Processing Systems, 33, 2020.
  73. 73.L. Wang, S. Xu, X. Wang, and Q. Zhu. Addressing class imbalance in federated learning. In AAAI, 2021.
  74. 74.W. Wang, J. Gao, M. Zhang, S. Wang, G. Chen, T. K. Ng, B. C. Ooi, J. Shao, and M. Reyad. Rafiki: Machine learning as an analytics service system. Proc. VLDB Endow., 12(2):128–140, Oct. 2018.
  75. 75.Y. Wu, S. Cai, X. Xiao, G. Chen, and B. C. Ooi. Privacy preserving vertical federated learning for tree-based models. Proceedings of the VLDB Endowment, 2020.
  76. 76.H. Xiao, K. Rasul, and R. Vollgraf. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  77. 77.Q. Yang, Y. Liu, T. Chen, and Y. Tong. Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology (TIST), 10(2):1–19, 2019.
  78. 78.M. Yurochkin, M. Agarwal, S. Ghosh, K. Greenewald, N. Hoang, and Y. Khazaeni. Bayesian nonparametric federated learning of neural networks. In Proceedings of the 36th International Conference on Machine Learning. PMLR, 2019.
  79. 79.K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising. IEEE transactions on image processing, 26(7):3142–3155, 2017.

Citation

MLA
Li, Q., et al. “Federated Learning on Non-IID Data Silos: An Experimental Study”. arXiv, 2021, http://arxiv.org/abs/2102.02079v4.
APA
Li, Q., Diao, Y., Chen, Q., & He, B. (2021). Federated Learning on Non-IID Data Silos: An Experimental Study. arXiv. http://arxiv.org/abs/2102.02079v4
Chicago
Li, Q., Y. Diao, Q. Chen, and B. He. 2021. “Federated Learning on Non-IID Data Silos: An Experimental Study”. arXiv. http://arxiv.org/abs/2102.02079v4.
Harvard
Li, Q. et al. (2021) “Federated Learning on Non-IID Data Silos: An Experimental Study”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2102.02079v4.
Vancouver
1. Li Q, Diao Y, Chen Q, He B (2021) Federated Learning on Non-IID Data Silos: An Experimental Study. arXiv

BibTeX

@article{li2021federated,
  title = {Federated Learning on Non-IID Data Silos: An Experimental Study},
  author = {Li, Qinbin and Diao, Yiqun and Chen, Quan and He, Bingsheng},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2102.02079v4},
  eprint = {2102.02079}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF