Federated Learning on Non-IID Data Silos: An Experimental Study
Qinbin LiYiqun DiaoQuan ChenBingsheng He
Presents a systematic experimental benchmark and comprehensive data partitioning framework to evaluate state-of-the-art federated learning algorithms across heterogeneous data silos, revealing that no single method consistently outperforms others under non-IID conditions.
Data privacy regulations and cross-organizational boundaries increasingly trap valuable datasets in isolated distributed silos. Federated learning enables organizations to collaboratively train machine learning models without transferring raw data across borders or corporate firewalls. However, local data across these silos is rarely identical in distribution. This non-independent and identically distributed (non-IID) data creates training drift and degrades model performance. Until now, systematic evaluations of federated algorithms under realistic, diverse data skews have been lacking due to overly simplistic testing setups.
The article introduces a comprehensive benchmarking framework, NIID-Bench, to systematically evaluate how well leading horizontal federated learning algorithms handle diverse data skews across distributed data silos. It evaluates and compares the accuracy, stability, and operational efficiency of four prominent algorithms: Federated Averaging (FedAvg), FedProx, SCAFFOLD, and FedNova.
To establish a credible evaluation, the article defines six distinct data partitioning strategies covering label distribution skew, feature distribution skew, and quantity skew. The benchmark tests four state-of-the-art algorithms across nine diverse datasets, encompassing six image recognition datasets and three tabular datasets. The baseline evaluations primarily simulate setups across multiple parties, evaluating overall classification accuracy, local epoch sensitivity, party sampling effects, scaling bottlenecks, and computational overhead across controlled communication rounds.
The findings show that no single federated learning algorithm consistently outperforms the others across all non-IID scenarios. Label distribution skew poses the most severe challenge, causing dramatic accuracy collapses across all algorithms when participants hold only a single data category. In contrast, feature distribution skews and pure quantity skews exhibit minimal accuracy degradation relative to homogeneous benchmarks. Across specific settings, FedProx achieves the best accuracy in most label and quantity skew environments, whereas SCAFFOLD excels in feature skew settings. However, algorithms exhibit critical operational trade-offs: FedProx incurs roughly 1.5 to 3 times the computational time per round compared to FedAvg, SCAFFOLD doubles the required communication bandwidth per round and suffers severe performance drops during partial client participation, and standard layer averaging techniques introduce instability into deeper architectures.
These results demonstrate that organizations cannot rely on a single default federated algorithm or assume modern methods universally solve data heterogeneity. For decision-makers, selecting the wrong algorithm risks significant communication and computing expenses, extended development timelines, or outright model failure in deployment. Understanding the dominant type of local data skew is vital before deploying federated infrastructure, as matching algorithm mechanics to specific data imbalances directly governs operational costs and final system accuracy.
Organizations should adopt tailored decision logic: deploy FedProx or standard FedAvg when dealing primarily with quantity imbalances or distribution-based label skews, and apply SCAFFOLD when feature skews dominate and full network participation is guaranteed. Hyperparameters, such as the number of local training epochs, require careful tuning to prevent model drift. Further industry-wide work is recommended to design skew-resistant client sampling techniques, automate parameter tuning, and develop lightweight profiling methods that identify data distributions prior to training without compromising privacy.
A primary limitation of the study is its reliance on synthetic partitions of existing centralized datasets rather than natively generated, complex real-world federated deployments. Additionally, client scaling was evaluated primarily in smaller cohorts and a 100-client simulation. Nevertheless, the experimental rigor and broad range of scenarios provide high confidence in the relative strengths, weaknesses, and structural bottlenecks identified across the tested algorithms.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). This seminal paper introduces the core FederatedAveraging algorithm and initial empirical evaluations on non-IID data that define the baseline paradigm tested in the experimental study.
- Paper: Federated Learning with Non-IID Data, Yue Zhao et al. (2018). This paper analyzes the mathematical causes and severe accuracy loss of federated learning under non-IID client partitions, framing the foundational data heterogeneity problems examined by the source.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). This work introduces FedProx to handle statistical and system heterogeneity, serving as one of the central baseline algorithms systematically benchmarked in the source study.
- Paper: SCAFFOLD: Stochastic Controlled Averaging for Federated Learning, Sai Praneeth Karimireddy et al. (2019). This work formulates the SCAFFOLD algorithm to correct client drift under heterogeneous local data distributions, representing a key state-of-the-art method evaluated in the benchmark.
- Paper: Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification, Tzu-Ming Harry Hsu et al. (2019). This paper establishes the use of Dirichlet distributions to model continuous spectra of non-IID client skews, providing the methodological basis for modern federated data partitioning strategies.
- Paper: Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization, Jianyu Wang et al. (2020). This paper introduces FedNova to solve objective inconsistency during heterogeneous local updates, providing another essential comparative algorithm analyzed in the experimental study.
- Paper: Adaptive Federated Optimization, Sashank Reddi et al. (2020). This paper develops adaptive server-side optimization techniques for non-identical federated data, which form part of the evaluated algorithmic spectrum.
- Paper: LEAF: A Benchmark for Federated Settings, Sebastian Caldas et al. (2018). This paper presents the LEAF benchmark framework for federated learning datasets and evaluation metrics, establishing standard non-IID experimental setups.
- Paper: Federated Machine Learning, Qiang Yang et al. (2019). This foundational survey defines cross-silo federated learning and the regulatory challenges of isolated databases that motivate the source study's focus on data silos.
- Paper: On the Convergence of FedAvg on Non-IID Data, Xiang Li et al. (2019). This study provides formal convergence bounds for Federated Averaging on non-IID partitions, clarifying the theoretical trade-offs between local steps and data skew.
- Paper: Model-Contrastive Federated Learning, Qinbin Li et al. (2021). This paper introduces model-contrastive learning (MOON) to directly mitigate local model drift under the non-IID data distributions characterized in the source benchmark.
