An Aggregation-Free Federated Learning for Tackling Data Heterogeneity
Yuan WangHuazhu FuRenuga KanagaveluQingsong WeiYong LiuRick Siow Mong Goh
Proposes FedAF, an aggregation-free federated learning framework where clients generate synthetic condensed data and soft labels via Sliced Wasserstein Distance regularization, eliminating client drift and significantly boosting convergence speed and accuracy on heterogeneous datasets.
Federated learning enables multiple distributed clients to train a shared global artificial intelligence model without sharing their raw, private datasets. However, real-world deployments frequently suffer from significant data heterogeneity, where client data distributions differ sharply in class labels or visual features. In standard frameworks, this diversity causes client drift and catastrophic forgetting, as local updates diverge from the central learning goal, severely degrading overall model accuracy and slowing convergence.
The article introduces FedAF, an aggregation-free federated learning framework designed to eliminate client drift across heterogeneous networks. FedAF replaces the standard cycle of averaging local model weights with a framework where clients condense their local datasets into compact synthetic samples and share them, along with predictive soft labels, for the central server to train the global model directly.
To evaluate this framework, the authors conducted extensive benchmark experiments across image classification datasets, including Fashion-MNIST, CIFAR-10, CIFAR-100, and DomainNet. They simulated harsh label-skew conditions using varying Dirichlet distribution parameters across ten clients and evaluated feature-skew across six distinct visual domains. FedAF was benchmarked against traditional aggregate-then-adapt methods, such as FedAvg, FedProx, FedDyn, and MOON, as well as an existing aggregation-free condensation method, FedDM.
The findings show that FedAF consistently outperforms existing methods in both final model accuracy and training efficiency. In label-skew settings, FedAF improved accuracy by up to 25.44% on CIFAR-10, 17.91% on CIFAR-100, and 31.03% on Fashion-MNIST compared to standard baseline approaches. It also improved accuracy by up to 4.87% over existing aggregation-free methods. In terms of training speed, FedAF accelerated convergence by up to 80%, reaching target accuracy thresholds in two to three rounds where prior methods required ten to fifteen rounds. Furthermore, under domain-level feature skew, FedAF achieved the highest overall average accuracy across six domains, ranking first or second in every individual domain.
These results demonstrate that shifting to a data condensation paradigm allows organizations to train robust global models across non-uniform data environments without sacrificing communication efficiency or privacy. By sharing condensed synthetic data rather than vulnerable gradient updates, organizations mitigate model divergence while maintaining resistance to data inference attacks. The framework also retains historical knowledge better between training rounds, delivering predictable model stabilization.
For engineering and research teams looking to deploy this framework, the article recommends configuring synthetic dataset sizes to around 50 images per class, as this provides an optimal balance among accuracy, communication overhead, and privacy preservation before returns diminish. Teams should also implement parameter re-sampling around a 0.9 retention ratio to prevent overfitting during local data synthesis.
Confidence in these findings is high for standard image classification benchmarks across simulated client splits. However, decision-makers should note that evaluations were conducted in simulated research settings with ten or fewer clients. Further testing on production-scale edge hardware, real-world network latency, and non-image modalities is advised before wide operational deployment.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). Introduces the foundational aggregate-then-adapt FederatedAveraging (FedAvg) paradigm that FedAF directly modifies and seeks to replace to avoid client drift.
- Paper: SCAFFOLD: Stochastic Controlled Averaging for Federated Learning, Sai Praneeth Karimireddy et al. (2019). Establishes the theoretical framework and lower bounds for client drift caused by data heterogeneity under standard local-update aggregation methods.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). Formulates the canonical proximal-regularization approach to mitigate client drift in heterogeneous networks, serving as an essential baseline.
- Paper: Federated Learning with Non-IID Data, Yue Zhao et al. (2018). Analyzes the mathematical mechanisms of weight divergence and performance loss under label and feature distribution skews in federated optimization.
- Paper: Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization, Jianyu Wang et al. (2020). Provides a comprehensive theoretical treatment of objective inconsistency and solution bias in heterogeneous federated aggregation.
- Paper: Ensemble Distillation for Robust Model Fusion in Federated Learning, Tao Lin et al. (2020). Pioneers alternative knowledge distillation and data-free server-side fusion techniques to bypass naive parameter averaging under heterogeneous data.
- Paper: Federated Learning on Non-IID Data Silos: An Experimental Study, Qinbin Li et al. (2021). Categorizes and benchmarks distinct non-IID regimes, including label skew and feature skew, which define the evaluation criteria used in FedAF.
- Paper: FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning, Jianqing Zhang et al. (2024). Extends non-aggregating and representation-sharing approaches by learning trainable global prototypes via contrastive learning under extreme model and data heterogeneity.
