Improving the Model Consistency of Decentralized Federated Learning
Yifan ShiLi ShenKang WeiYan SunBo YuanXueqian WangDacheng Tao
Proposes two decentralized federated learning algorithms that combine Sharpness Aware Minimization with multiple gossip communication steps to reduce client model inconsistency, providing non-convex convergence guarantees and matching centralized performance on heterogeneous data.
Decentralized federated learning enables distributed participants to collaboratively train machine learning models by communicating directly with neighboring devices rather than through a central coordinating server. This structure improves data privacy and reduces central communication bottlenecks across edge networks, healthcare systems, and connected devices. However, because local data is often non-uniform and network connections can be sparse, participating devices frequently develop highly inconsistent models, causing severe performance drops and poor generalization compared to traditional centralized approaches.
The main objective of the article is to develop and evaluate two new decentralized optimization algorithms, named DFedSAM and DFedSAM-MGS, designed to eliminate model inconsistency and achieve performance on par with centralized training. The proposed methods achieve this by having each local device seek flatter, more robust loss regions and by accelerating neighbor-to-neighbor consensus through multiple communication exchanges per training round.
The authors evaluated the approaches using both mathematical convergence proofs and extensive empirical experiments. The experimental evaluation tested image classification tasks across 100 simulated clients on standard benchmarks under both uniform and highly skewed data distributions. The testing examined diverse communication network topologies, ranging from sparsely connected rings and grids to fully connected structures, comparing results against leading centralized and decentralized baselines.
The findings demonstrate substantial performance and consistency improvements. First, the proposed DFedSAM and DFedSAM-MGS consistently outperformed existing decentralized methods across all benchmark tasks and data distributions. On benchmark datasets, DFedSAM-MGS achieved test accuracies comparable to state-of-the-art centralized baselines, such as reaching 84.26% to 86.47% accuracy under varying data skew on standard image tasks. Second, the methods demonstrated superior resilience on sparse network topologies; in restrictive ring networks, DFedSAM-MGS improved accuracy by 8.0 percentage points over standard decentralized momentum approaches. Third, Hessian eigenvalue analyses verified that the algorithms successfully produce flatter loss landscapes, reducing model overfitting across local nodes. Finally, theoretical analysis confirmed that increasing consensus communication steps effectively neutralizes the negative convergence impact of sparse network designs.
These results show that decentralized networks can match centralized model quality without requiring central server infrastructure, significantly lowering operational risks and bandwidth costs. Organizations deploying edge machine learning in privacy-sensitive or infrastructure-constrained environments can adopt these methods to achieve reliable model accuracy across heterogeneous participants.
Decision-makers should consider adopting flatness-aware training and multi-step neighbor communication when deploying peer-to-peer federated systems. Practitioners should tune the number of consensus steps to balance network communication volume against target model accuracy, as a moderate consensus setting of four communication steps provided the best operational trade-off in the study. Before full deployment, teams should conduct pilot tests to optimize hyperparameter settings for their specific network hardware and data heterogeneity levels.
- Paper: Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent, Xiangru Lian et al. (2017). This paper establishes how neighbor-to-neighbor decentralized SGD can match centralized optimization, providing the network and consensus foundations for the source’s decentralized algorithms.
- Paper: Multi-Consensus Decentralized Accelerated Gradient Descent, Haishan Ye et al. (2023). Its multi-consensus analysis develops the communication-step mechanisms that the source builds on to improve convergence over sparse decentralized networks.
- Paper: DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training, Rong Dai et al. (2022). This decentralized federated method exposes the same topology and data-heterogeneity challenges that motivate the source’s effort to improve consistency and accuracy.
No sufficiently relevant recommendations were found.
