DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training
Rong DaiLi ShenFengxiang HeXinmei TianDacheng Tao
Proposes a peer-to-peer personalized federated learning framework that applies client-tailored sparse masks throughout training and communication, cutting bottleneck bandwidth and local compute costs while improving accuracy across heterogeneous edge devices.
Modern edge and mobile devices generate massive amounts of decentralized data, making distributed training essential for preserving data privacy. However, classical centralized federated learning faces significant communication bottlenecks and severe vulnerability to single-point server failures or attacks. In addition, real-world deployments suffer from high data heterogeneity across clients and wide variations in hardware, memory, and computing capacities, which centralized and dense-model architectures struggle to accommodate efficiently.
The main objective of the article is to develop and evaluate Dis-PFL, a personalized federated learning framework operating over a fully decentralized, peer-to-peer communication network using customized sparse local models to address both data and device heterogeneity.
The authors evaluated the framework through theoretical generalization analysis and extensive empirical simulations across three standard image classification benchmarks: CIFAR-10, CIFAR-100, and Tiny-ImageNet. The evaluation spanned 100 client nodes under two non-identical data distribution scenarios, evaluated multiple network topologies such as ring, time-varying dynamic, and fully connected graphs, and benchmarked against leading centralized and decentralized baselines.
The key findings demonstrate that Dis-PFL consistently outperforms existing centralized and decentralized baselines in model accuracy while reducing resource overhead. First, Dis-PFL achieved higher test accuracy across all benchmark datasets, reaching up to 85.70% on Dirichlet-partitioned CIFAR-10 compared to centralized federated learning at 78.07% and fine-tuned decentralized parallel stochastic gradient descent at 83.90%. Second, Dis-PFL cut the peak communication burden of the busiest node by roughly 50% compared to standard dense approaches and reduced local floating-point computing operations by 15% to 40% compared to dense and fine-tuned baselines. Third, the framework converged significantly faster, requiring roughly 20% to 50% fewer communication rounds than alternative approaches to achieve target accuracy levels. Finally, when deployed in heterogeneous environments where devices varied in capacity from 20% to 100% of the dense model size, Dis-PFL maintained robust performance and adapted dynamically without restricting the entire network to the weakest device's limitations.
These findings indicate that decentralized sparse training offers an effective, cost-efficient path to scaling private collaborative learning across resource-constrained edge hardware. Eliminating reliance on central servers significantly reduces systemic security risks and network infrastructure costs while accelerating training timelines.
Organizations deploying distributed machine learning across diverse edge hardware should consider adopting decentralized sparse model architectures to balance device performance and network bandwidth. However, decision-makers must carefully tune the sparsity ratio to prevent performance degradation caused by overly sparse parameter masks or insufficient mask diversity. Practitioners should also conduct pilot evaluations, as the article relies on simulated edge environments, and further investigation is needed to explore real-world network latency and the deeper structural relationships between local data distributions and the generated model masks.
- Paper: Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent, Xiangru Lian et al. (2017). This paper establishes the foundational theory and empirical viability of decentralized parallel SGD over peer-to-peer topologies, which DisPFL directly adopts and sparsifies.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). This seminal paper introduces the fundamental federated averaging paradigm and data heterogeneity problem that DisPFL seeks to decentralize and personalize.
- Paper: Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training, Yujun Lin et al. (2018). This foundational work on extreme gradient sparsification provides the principles of communication-efficient sparse training utilized in DisPFL's localized mask generation.
- Paper: Federated Learning with Personalization Layers, Manoj Ghuhan Arivazhagan et al. (2019). This work introduces architecture-level parameter personalization to combat client heterogeneity, establishing a primary conceptual baseline that DisPFL advances via decentralized sparse sub-networks.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). This paper formalizes optimization under systems and data heterogeneity in federated settings, providing essential context for DisPFL's evaluation against heterogeneous client capabilities.
- Paper: Towards Personalized Federated Learning, Alysa Ziying Tan et al. (2021). This comprehensive survey categorizes the personalization strategies and system efficiency trade-offs that DisPFL unifies into a decentralized framework.
- Paper: Ditto: Fair and Robust Federated Learning Through Personalization, Tian Li et al. (2020). This work explores personalized federated optimization as a means to resolve statistical client disparities, motivating the local model adaptation mechanisms explored in DisPFL.
- Paper: FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning, Jianqing Zhang et al. (2024). This work extends personalized federated learning under simultaneous model and data heterogeneity by utilizing trainable global prototypes and contrastive learning rather than sparse subnetwork masks.
- Paper: CD2-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning, Yiqing Shen et al. (2022). This paper advances model personalization across heterogeneous clients through cyclic distillation and channel decoupling across network layers without requiring manual layer selection.
- Paper: DAdaQuant: Doubly-adaptive quantization for communication-efficient Federated Learning, Robert Hönig et al. (2022). This work provides an alternative communication-efficient federated learning mechanism using doubly-adaptive parameter quantization across heterogeneous edge devices.
- Paper: Personalized Federated Learning via Variational Bayesian Inference, Xu Zhang et al. (2022). This paper explores personalized federated learning under limited client data by utilizing variational Bayesian inference to capture uncertainty and prevent local overfitting.
