Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification

Tzu-Ming Harry HsuHang QiMatthew Brown

article2019arXiv1,705 citations

Demonstrates how non-identical client data distributions degrade federated visual classification and introduces a server momentum technique that improves model accuracy from 30.1% to 76.9% in severely skewed settings.

Listen

Federated learning enables organizations to train artificial intelligence models across distributed edge devices, such as mobile phones, while keeping user data private and localized. However, real-world deployments face a significant challenge: unlike centralized data centers where data is uniformly distributed, edge devices naturally collect highly skewed and non-identical datasets. As visual recognition systems are increasingly deployed on edge hardware, understanding and mitigating the performance degradation caused by these data disparities has become a critical operational priority.

The article evaluates the impact of non-identical client data distributions on visual classification within the standard Federated Averaging framework and demonstrates an optimization method to preserve model accuracy under severe data skew.

To conduct this evaluation, the researchers synthesized a continuous spectrum of data distributions across a simulated population of 100 clients using the CIFAR-10 image benchmark. By varying a concentration parameter in a Dirichlet statistical distribution, the study modeled settings ranging from perfectly balanced client data to extreme cases where clients held images from only a single class. The evaluation analyzed model performance, communication rounds, client participation rates, and local training epochs, comparing standard federated learning against an enhanced method incorporating server-side momentum.

The findings show that standard Federated Averaging experiences severe performance drops as client data becomes more skewed, with accuracy plummeting from nearly 84% in balanced settings to below 30% under extreme skew when client participation is low. Furthermore, higher data heterogeneity increases training volatility and narrows the viable window of learning rates, making models difficult to tune. Increasing the fraction of participating clients yields diminishing returns on balanced data but is vital for non-identical settings. Crucially, introducing server momentum—termed FedAvgM—substantially mitigates these issues, lifting visual classification accuracy from 30.1% up to 76.9% in the most skewed environments and closely tracking centralized performance baselines of 86.0%.

These results demonstrate that unmitigated data heterogeneity poses a major technical risk to edge-based machine learning, potentially leading to model instability or failure in production. Incorporating server momentum provides a high-impact, practical mitigation that stabilizes training across distributed environments without requiring raw data sharing. However, using server momentum introduces additional hyperparameter complexity; when only a few devices report per round, engineering teams must carefully pair low client learning rates with high momentum to prevent model divergence.

For practical implementation, organizations deploying federated learning systems across diverse edge devices should adopt server-side momentum to protect against performance collapse. Engineering teams must invest in tuning effective learning rates and maintain adequate client reporting fractions where communication budgets permit. Further investigation using larger real-world datasets and broader network architectures is recommended to validate these hyperparameter configurations prior to wide-scale deployment.

arXiv: 1909.06335
Cover for Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification

Abstract

Federated Learning enables visual models to be trained in a privacy-preserving way using real-world data from mobile devices. Given their distributed nature, the statistics of the data across these devices is likely to differ significantly. In this work, we look at the effect such non-identical data distributions has on visual classification via Federated Learning. We propose a way to synthesize datasets with a continuous range of identicalness and provide performance measures for the Federated Averaging algorithm. We show that performance degrades as distributions differ more, and propose a mitigation strategy via server momentum. Experiments on CIFAR-10 demonstrate improved classification performance over a range of non-identicalness, with classification accuracy improved from 30.1% to 76.9% in the most skewed settings.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Synthetic Non-Identical Client Data
  • 4 Experiments and Results
  • 4.1 Classification Performance with Non-Identical Distributions
  • 4.2 Accumulating Model Updates with Momentum
  • References

Knowls

  1. Knowl 1 — Dirichlet-Based Synthesis of Non-Identical Client Data Partitions

    model/method

    To model non-identical (non-IID) data distributions across federated clients along a continuous spectrum, client class label distributions are synthesized using a Dirichlet distribution.

    Let NN denote the total number of target classes, and let p=(p1,p2,…,pN)\mathbf{p} = (p_1, p_2, \dots, p_N) be a prior probability distribution over the NN classes such that ∑i=1Npi=1\sum_{i=1}^N p_i = 1 and pi≥0p_i \ge 0. For each client kk, a local class distribution vector q(k)=(q1(k),…,qN(k))\mathbf{q}^{(k)} = (q_1^{(k)}, \dots, q_N^{(k)}) is sampled from a Dirichlet distribution parameterized by αp\alpha \mathbf{p}:

    q(k)∼Dir(αp)\mathbf{q}^{(k)} \sim \text{Dir}(\alpha \mathbf{p})

    where α>0\alpha > 0 is a concentration parameter governing the degree of identicalness across clients:

    • As α→∞\alpha \to \infty, the sampled distributions converge to q(k)→p\mathbf{q}^{(k)} \to \mathbf{p} for all clients, corresponding to identical (IID) class distributions across the population.
    • As α→0\alpha \to 0, each vector q(k)\mathbf{q}^{(k)} approaches a one-hot distribution, meaning each client holds training instances belonging almost exclusively to a single class chosen at random.

    Each client kk is then assigned local training examples drawn according to class proportions specified by q(k)\mathbf{q}^{(k)}.

  2. Knowl 2 — Federated Averaging with Server Momentum (FedAvgM)

    algorithm

    Federated Averaging with Server Momentum (FedAvgM) applies momentum accumulation at the central server to dampen update oscillations caused by client data heterogeneity and client sampling noise.

    In standard FedAvg, the server updates global weights ww by a weighted average of client updates Δw=∑k=1KnknΔwk\Delta w = \sum_{k=1}^K \frac{n_k}{n} \Delta w_k, setting w←w−Δww \leftarrow w - \Delta w. FedAvgM maintains a server-side velocity state vv and applies Nesterov accelerated gradient updates at the server level.

    Input: total communication rounds TT, client population size MM, reporting fraction C∈(0,1]C \in (0, 1], local epoch count EE, local batch size BB, local learning rate η\eta, server momentum parameter β∈[0,1)\beta \in [0, 1)
    Output: final global model weights wTw_T
    Initialize global model weights w0w_0, server momentum buffer v0=0v_0 = 0
    for each round t=0,1,…,T−1t = 0, 1, \dots, T-1 do
        Sample a subset StS_t of max⁡(⌊C⋅M⌋,1)\max(\lfloor C \cdot M \rfloor, 1) clients uniformly at random from the population
        for each client k∈Stk \in S_t in parallel do
            Set local model wt,0(k)=wtw_{t, 0}^{(k)} = w_t
            Split client kk's dataset DkD_k (size nkn_k) into batches of size BB
            for each local epoch e=1,…,Ee = 1, \dots, E do
                for each batch b⊂Dkb \subset D_k do
                    wt,step(k)←wt,step(k)−η∇L(wt,step(k);b)w_{t, \text{step}}^{(k)} \leftarrow w_{t, \text{step}}^{(k)} - \eta \nabla \mathcal{L}(w_{t, \text{step}}^{(k)}; b)
            Compute client update Δwt(k)=wt−wt,final(k)\Delta w_t^{(k)} = w_t - w_{t, \text{final}}^{(k)}
        Compute total sampled dataset size nt=∑k∈Stnkn_t = \sum_{k \in S_t} n_k
        Aggregate average model update Δwt=∑k∈StnkntΔwt(k)\Delta w_t = \sum_{k \in S_t} \frac{n_k}{n_t} \Delta w_t^{(k)}
        Update server velocity vt+1=βvt+Δwtv_{t+1} = \beta v_t + \Delta w_t
        Update server model weights wt+1=wt−(βvt+1+Δwt)w_{t+1} = w_t - (\beta v_{t+1} + \Delta w_t)
    return wTw_T
  3. Knowl 3 — Server Momentum Mitigation of Non-IID Performance Degradation

    empirical result

    Incorporating server-side momentum (FedAvgM) substantially mitigates the performance degradation of federated visual classification caused by non-identical client data distributions on CIFAR-10.

    • In the most skewed setting (Dirichlet concentration α→0\alpha \to 0, where each client holds images from only one class), test classification accuracy improves from 30.1%30.1\% with vanilla FedAvg to 76.9%76.9\% with FedAvgM.
    • In challenging regimes with low client participation (C=0.05C = 0.05, corresponding to 5 participating clients per round) and E=1E = 1, vanilla FedAvg test accuracy degrades rapidly to approximately 35%–40%35\%\text{--}40\%, whereas FedAvgM maintains a stable test accuracy exceeding 75%75\%, approaching the centralized baseline accuracy (86.0%86.0\%).
    • The stabilizing effect is attributed to momentum buffering and smoothing out the conflicting, high-variance local gradient directions generated by clients holding sparse subsets of class labels.
  4. Knowl 4 — FedAvg Accuracy under Varying Non-IID Dirichlet Skew, Client Participation, and Local Epochs

    data/table

    On CIFAR-10 partitioned into 100 clients (500 images per client) over 10,000 communication rounds with local batch size B=64B=64, the test accuracy of vanilla FedAvg depends strongly on the Dirichlet concentration parameter α\alpha, the client reporting fraction CC, and the number of local epochs EE. Each entry is evaluated using the best tuned client learning rate averaged over 5 random population seeds.

    Setting α=100.0\alpha=100.0 α=10.0\alpha=10.0 α=1.0\alpha=1.0 α=0.5\alpha=0.5 α=0.2\alpha=0.2 α=0.1\alpha=0.1 α=0.05\alpha=0.05 α=0.00\alpha=0.00
    (a) Local Epochs E=1E = 1
    C=0.40C = 0.40 0.839 0.833 0.789 0.770 0.746 0.732 0.703 0.659
    C=0.20C = 0.20 0.841 0.830 0.785 0.761 0.745 0.721 0.701 0.650
    C=0.10C = 0.10 0.836 0.823 0.781 0.759 0.728 0.689 0.660 0.479
    C=0.05C = 0.05 0.829 0.817 0.768 0.699 0.651 0.579 0.526 0.407
    (b) Local Epochs E=5E = 5
    C=0.40C = 0.40 0.847 0.829 0.784 0.761 0.732 0.709 0.647 0.481
    C=0.20C = 0.20 0.845 0.829 0.779 0.749 0.715 0.674 0.629 0.359
    C=0.10C = 0.10 0.841 0.831 0.773 0.738 0.705 0.642 0.584 0.293
    C=0.05C = 0.05 0.835 0.819 0.745 0.697 0.638 0.580 0.520 0.268

    The table demonstrates:

    1. Accuracy drops severely as α\alpha decreases toward 0, especially for α≤0.1\alpha \le 0.1.
    2. Higher reporting fraction CC yields diminishing returns on IID data (large α\alpha) but provides critical accuracy gains on non-IID data (small α\alpha).
    3. More local epochs (E=5E=5) slightly improve accuracy in near-IID settings (e.g., 0.8470.847 vs 0.8390.839 at α=100.0,C=0.40\alpha=100.0, C=0.40), but cause severe degradation under non-IID distributions (e.g., falling to 0.2680.268 vs 0.4070.407 at α=0.00,C=0.05\alpha=0.00, C=0.05).
  5. Knowl 5 — Experimental Benchmark Setup for Federated CIFAR-10 Classification

    experimental setup

    The federated visual classification benchmark uses the CIFAR-10 dataset (50,000 training images, 10,000 testing images across 10 classes) under the following protocol:

    • Client Population: 100 clients, each holding exactly 500 images. The prior class distribution p\mathbf{p} is uniform (pi=0.1p_i = 0.1 for each of the 10 classes), matching the test set distribution.
    • Data Heterogeneity: Dirichlet concentration parameter α∈{100.0,10.0,1.0,0.5,0.2,0.1,0.05,0.0}\alpha \in \{100.0, 10.0, 1.0, 0.5, 0.2, 0.1, 0.05, 0.0\} (where α=0.0\alpha=0.0 represents strictly 1 class per client).
    • Model Architecture: Standard CNN architecture from McMahan et al. with a fixed weight decay of 0.0040.004 and no learning rate decay schedule.
    • Training Hyperparameters: Total communication rounds T=10,000T = 10{,}000; client batch size B=64B = 64; local epochs E∈{1,5}E \in \{1, 5\}; reporting fraction C∈{0.05,0.1,0.2,0.4}C \in \{0.05, 0.1, 0.2, 0.4\} (5, 10, 20, 40 clients per round).
    • Hyperparameter Grid: Client learning rate η∈{10−4,3×10−4,10−3,3×10−3,10−2,3×10−2,10−1,3×10−1}\eta \in \{10^{-4}, 3\times 10^{-4}, 10^{-3}, 3\times 10^{-3}, 10^{-2}, 3\times 10^{-2}, 10^{-1}, 3\times 10^{-1}\}. For FedAvgM, server momentum parameter β∈{0,0.7,0.9,0.97,0.99,0.997}\beta \in \{0, 0.7, 0.9, 0.97, 0.99, 0.997\} with server learning rate fixed at 1.0.
  6. Knowl 6 — Hyperparameter Sensitivity and Effective Learning Rate in Federated Optimization

    empirical result

    The difficulty of hyperparameter tuning in federated learning is directly dictated by data heterogeneity α\alpha, reporting fraction CC, and local computation EE:

    1. Vanilla FedAvg Sensitivity: When data are identical (large α\alpha), a broad range of client learning rates η\eta spanning approximately two orders of magnitude achieves near-optimal test accuracy. When data distributions are non-identical (low α\alpha) and reporting fraction CC is small, the viable learning rate window narrows substantially, and slight deviations lead to failure or chance-level accuracy (10%10\%).
    2. Effective Learning Rate in FedAvgM: For FedAvgM, hyperparameter selection can be analyzed via the effective learning rate:

    ηeff=η1−β\eta_{\text{eff}} = \frac{\eta}{1 - \beta}

    where η\eta is the client learning rate and β\beta is the server momentum coefficient. At large reporting fractions (C=0.40C = 0.40), ηeff\eta_{\text{eff}} can be selected across two orders of magnitude. At low reporting fractions (C=0.05C = 0.05), the viable window for ηeff\eta_{\text{eff}} shrinks to a single order of magnitude. 3. Mitigating Divergence: To avoid divergence when CC is low or when local epochs EE are high (which increases local update variance and noise), optimal performance requires combining a small absolute learning rate η\eta with a high server momentum β\beta.

Coverage note — None was omitted; all core contributions, including the Dirichlet non-IID partitioning method, the FedAvgM algorithm, the full CIFAR-10 empirical benchmark results, and the effective learning rate analysis, are fully captured.

References

  1. 1.Sebastian Caldas, Peter Wu, Tian Li, Jakub Konečnỳ, H Brendan McMahan, Virginia Smith, and Ameet Talwalkar. Leaf: A benchmark for federated settings. arXiv preprint arXiv:1812.01097, 2018.
  2. 2.Gregory Cohen, Saeed Afshar, Jonathan Tapson, and André van Schaik. EMNIST: an extension of MNIST to handwritten letters. arXiv preprint arXiv:1702.05373, 2017.
  3. 3.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. Technical report, Citeseer, 2009.
  4. 4.Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. On the convergence of FedAvg on non-IID data. arXiv preprint arXiv:1907.02189, 2019.
  5. 5.Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics, pages 1273–1282, 2017.
  6. 6.Yu Nesterov. Gradient methods for minimizing composite objective function. 2007.
  7. 7.Anit Kumar Sahu, Tian Li, Maziar Sanjabi, Manzil Zaheer, Ameet Talwalkar, and Virginia Smith. On the convergence of federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018.
  8. 8.Felix Sattler, Simon Wiedemann, Klaus-Robert Müller, and Wojciech Samek. Robust and communication-efficient federated learning from non-IID data. arXiv preprint arXiv:1903.02891, 2019.
  9. 9.Christopher J Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl. Measuring the effects of data parallelism on neural network training. arXiv preprint arXiv:1811.03600, 2018.
  10. 10.TensorFlow. Advanced convolutional neural networks. URL https://www.tensorflow.org/tutorials/images/deep_cnn.
  11. 11.Mikhail Yurochkin, Mayank Agarwal, Soumya Ghosh, Kristjan Greenewald, Nghia Hoang, and Yasaman Khazaeni. Bayesian nonparametric federated learning of neural networks. In International Conference on Machine Learning, pages 7252–7261, 2019.
  12. 12.Yue Zhao, Meng Li, Liangzhen Lai, Naveen Suda, Damon Civin, and Vikas Chandra. Federated learning with non-IID data. arXiv preprint arXiv:1806.00582, 2018.

Citation

MLA
Hsu, T.-M. H., et al. “Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification”. arXiv, 2019, http://arxiv.org/abs/1909.06335v1.
APA
Hsu, T.-M. H., Qi, H., & Brown, M. (2019). Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification. arXiv. http://arxiv.org/abs/1909.06335v1
Chicago
Hsu, T.-M. H., H. Qi, and M. Brown. 2019. “Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification”. arXiv. http://arxiv.org/abs/1909.06335v1.
Harvard
Hsu, T.-M.H., Qi, H. and Brown, M. (2019) “Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1909.06335v1.
Vancouver
1. Hsu T-MH, Qi H, Brown M (2019) Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification. arXiv

BibTeX

@article{hsu2019measuring,
  title = {Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification},
  author = {Hsu, Tzu-Ming Harry and Qi, Hang and Brown, Matthew},
  year = {2019},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1909.06335v1},
  eprint = {1909.06335}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission