DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training

Rong DaiLi ShenFengxiang HeXinmei TianDacheng Tao

article2022ICML161 citations

Proposes a peer-to-peer personalized federated learning framework that applies client-tailored sparse masks throughout training and communication, cutting bottleneck bandwidth and local compute costs while improving accuracy across heterogeneous edge devices.

Listen

Modern edge and mobile devices generate massive amounts of decentralized data, making distributed training essential for preserving data privacy. However, classical centralized federated learning faces significant communication bottlenecks and severe vulnerability to single-point server failures or attacks. In addition, real-world deployments suffer from high data heterogeneity across clients and wide variations in hardware, memory, and computing capacities, which centralized and dense-model architectures struggle to accommodate efficiently.

The main objective of the article is to develop and evaluate Dis-PFL, a personalized federated learning framework operating over a fully decentralized, peer-to-peer communication network using customized sparse local models to address both data and device heterogeneity.

The authors evaluated the framework through theoretical generalization analysis and extensive empirical simulations across three standard image classification benchmarks: CIFAR-10, CIFAR-100, and Tiny-ImageNet. The evaluation spanned 100 client nodes under two non-identical data distribution scenarios, evaluated multiple network topologies such as ring, time-varying dynamic, and fully connected graphs, and benchmarked against leading centralized and decentralized baselines.

The key findings demonstrate that Dis-PFL consistently outperforms existing centralized and decentralized baselines in model accuracy while reducing resource overhead. First, Dis-PFL achieved higher test accuracy across all benchmark datasets, reaching up to 85.70% on Dirichlet-partitioned CIFAR-10 compared to centralized federated learning at 78.07% and fine-tuned decentralized parallel stochastic gradient descent at 83.90%. Second, Dis-PFL cut the peak communication burden of the busiest node by roughly 50% compared to standard dense approaches and reduced local floating-point computing operations by 15% to 40% compared to dense and fine-tuned baselines. Third, the framework converged significantly faster, requiring roughly 20% to 50% fewer communication rounds than alternative approaches to achieve target accuracy levels. Finally, when deployed in heterogeneous environments where devices varied in capacity from 20% to 100% of the dense model size, Dis-PFL maintained robust performance and adapted dynamically without restricting the entire network to the weakest device's limitations.

These findings indicate that decentralized sparse training offers an effective, cost-efficient path to scaling private collaborative learning across resource-constrained edge hardware. Eliminating reliance on central servers significantly reduces systemic security risks and network infrastructure costs while accelerating training timelines.

Organizations deploying distributed machine learning across diverse edge hardware should consider adopting decentralized sparse model architectures to balance device performance and network bandwidth. However, decision-makers must carefully tune the sparsity ratio to prevent performance degradation caused by overly sparse parameter masks or insufficient mask diversity. Practitioners should also conduct pilot evaluations, as the article relies on simulated edge environments, and further investigation is needed to explore real-world network latency and the deeper structural relationships between local data distributions and the generated model masks.

arXiv: 2206.00187
Cover for DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training

Abstract

Personalized federated learning is proposed to handle the data heterogeneity problem amongst clients by learning dedicated tailored local models for each user. However, existing works are often built in a centralized way, leading to high communication pressure and high vulnerability when a failure or an attack on the central server occurs. In this work, we propose a novel personalized federated learning framework in a decentralized (peer-to-peer) communication protocol named Dis-PFL, which employs personalized sparse masks to customize sparse local models on the edge. To further save the communication and computation cost, we propose a decentralized sparse training technique, which means that each local model in Dis-PFL only maintains a fixed number of active parameters throughout the whole local training and peer-to-peer communication process. Comprehensive experiments demonstrate that Dis-PFL significantly saves the communication bottleneck for the busiest node among all clients and, at the same time, achieves higher model accuracy with less computation cost and communication rounds. Furthermore, we demonstrate that our method can easily adapt to heterogeneous local clients with varying computation complexities and achieves better personalized performances.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Dis-PFL algorithm
  • 3.1. Problem formulation
  • 3.2. Algorithm
  • 3.3. Generalization analysis
  • 4. Experiments
  • 4.1. Experiment Setup
  • 4.2. Main experiments evaluation
  • 4.3. Experiments on client heterogeneous setting
  • 4.4. Empirical analysis of the learned sparse masks.
  • 4.5. Discussion of the sparsity ratio
  • 5. Conclusion
  • Acknowledgement
  • References
  • Appendix
  • A. More details on algorithm implementation
  • A.1. Algorithm 2
  • B. Experiments
  • B.1. Datasets
  • B.2. Model Architectures
  • B.3. Hyper-Parameters
  • B.4. More details about baselines
  • B.5. More experiments results
  • B.5.1. CONVERGENCE SPEED
  • B.5.2. DIFFERENT TOPOLOGY
  • B.6. Extended experiments on random clients dropping settings
  • C. Proof of Theorem 1
  • C.1. Notations and Preliminaries
  • C.2. Key Lemmas
  • C.3. Main proof

Knowls

  1. Knowl 1 — Personalized Federated Learning Formulation with Binary Sparse Masks

    model/method

    In personalized federated learning (PFL), the goal is to customize dedicated models for each client to handle data heterogeneity. Dis-PFL formulates this objective by learning a shared dense parameter base w∈Rdw \in \mathbb{R}^d along with personalized binary sparse masks mk∈{0,1}dm_k \in \{0, 1\}^d for each client k∈{1,…,K}k \in \{1, \dots, K\}:

    min⁡w,m1,…,mKf(w,m1,…,mK)=1K∑k=1KFk(w⊙mk)\min_{w, m_1, \dots, m_K} f(w, m_1, \dots, m_K) = \frac{1}{K} \sum_{k=1}^K F_k(w \odot m_k)

    where ⊙\odot denotes the element-wise Hadamard product, KK is the total number of clients, and Fk(w⊙mk)F_k(w \odot m_k) is the true risk over the local data distribution Dk\mathcal{D}_k of client kk, defined as:

    Fk(w⊙mk):=E(x,y)∼Dk[L(w⊙mk;(x,y))]F_k(w \odot m_k) := \mathbb{E}_{(x, y) \sim \mathcal{D}_k} [\mathcal{L}(w \odot m_k; (x, y))]

    Here, L(⋅;⋅)\mathcal{L}(\cdot; \cdot) is the loss function on sample (x,y)(x, y). A mask value (mk)i=1(m_k)_i = 1 indicates that coordinate ii of the global parameter vector is active in the personalized model for client kk, while (mk)i=0(m_k)_i = 0 keeps it dormant. Each client's personalized model is thus a sub-network of the global model ww, tailored to local resources and data.

  2. Knowl 2 — Intersection-Weighted Gossip Parameter Aggregation

    model/method

    In a decentralized peer-to-peer federated learning system where each client kk maintains a sparse local model wk,t∈Rdw_{k,t} \in \mathbb{R}^d constrained by a personalized binary mask mk,t∈{0,1}dm_{k,t} \in \{0, 1\}^d, clients aggregate model parameters received from their neighborhood set Sk,t\mathcal{S}_{k,t} by averaging exclusively over the intersection of active weights.

    At communication round tt, client kk computes the intermediate aggregated model wk,t+1/2w_{k, t+1/2} according to:

    wk,t+1/2=(wk,t+∑j∈Sk,twj,tmk,t+∑j∈Sk,tmj,t)⊙mk,tw_{k, t+1/2} = \left( \frac{w_{k,t} + \sum_{j \in \mathcal{S}_{k,t}} w_{j,t}}{m_{k,t} + \sum_{j \in \mathcal{S}_{k,t}} m_{j,t}} \right) \odot m_{k,t}

    where division is performed element-wise for each coordinate where the denominator is non-zero (otherwise evaluating to 0), and ⊙\odot denotes the Hadamard product. This operation performs a weighted average of active parameters across intersecting neighbors while masking out coordinates outside client kk's current active sub-network mk,tm_{k,t}.

  3. Knowl 3 — Dis-PFL Algorithm for Decentralized Sparse Personalized Federated Learning

    algorithm

    Dis-PFL executes personalized federated learning over a decentralized network by maintaining fixed-sparsity models on edge devices, combining intersection-weighted gossip averaging, local sparse SGD updates, and dynamic mask adjustments based on gradient information.

    Input: Total number of clients KK; client capacity constraints {ck}k=1K\{c_k\}_{k=1}^K; total communication rounds TT; local SGD steps NN; learning rate η\eta; initial pruning rate α0\alpha_0
    Initialization: For each client kk, initialize sparse model wk,0w_{k,0} and binary mask mk,0m_{k,0} using the Erdős-Rényi Kernel (ERK) according to capacity ckc_k
    Output: Personalized sparse local models {wk,T}k=1K\{w_{k,T}\}_{k=1}^K
    for t=0t = 0 to T−1T - 1 do
        for each client k∈{1,…,K}k \in \{1, \dots, K\} in parallel do
            Receive neighbors' sparse models wj,tw_{j,t} and masks mj,tm_{j,t} from neighborhood set Sk,t\mathcal{S}_{k,t}
            Compute intermediate model:
                wk,t+1/2=(wk,t+∑j∈Sk,twj,tmk,t+∑j∈Sk,tmj,t)⊙mk,tw_{k, t+1/2} = \left( \frac{w_{k,t} + \sum_{j \in \mathcal{S}_{k,t}} w_{j,t}}{m_{k,t} + \sum_{j \in \mathcal{S}_{k,t}} m_{j,t}} \right) \odot m_{k,t}
            w~k,t,0=wk,t+1/2\widetilde{w}_{k,t,0} = w_{k, t+1/2}
            for τ=0\tau = 0 to N−1N - 1 do
                Sample mini-batch ξk,t,τ\xi_{k,t,\tau} from local dataset
                Compute stochastic gradient: gk,t,τ(w~k,t,τ)=∇w~L(w~k,t,τ;ξk,t,τ)g_{k,t,\tau}(\widetilde{w}_{k,t,\tau}) = \nabla_{\widetilde{w}} \mathcal{L}(\widetilde{w}_{k,t,\tau}; \xi_{k,t,\tau})
                Update sparse parameters: w~k,t,τ+1=w~k,t,τ−η⋅(mk,t⊙gk,t,τ(w~k,t,τ))\widetilde{w}_{k,t,\tau+1} = \widetilde{w}_{k,t,\tau} - \eta \cdot (m_{k,t} \odot g_{k,t,\tau}(\widetilde{w}_{k,t,\tau}))
            end for
            wk,t+1=w~k,t,Nw_{k,t+1} = \widetilde{w}_{k,t,N}
            Compute pruning rate αt\alpha_t via cosine annealing from α0\alpha_0
            Compute dense gradient g(wk,t+1)g(w_{k,t+1}) on a local mini-batch
            for each layer jj do
                Prune αt\alpha_t-proportion of weights with smallest magnitude in layer jj to produce mk,t+1/2jm_{k, t+1/2}^j
                Regrow αt\alpha_t-proportion of dormant weights with largest gradient magnitudes ∣g(wk,t+1)∣|g(w_{k,t+1})| to obtain mk,t+1jm_{k,t+1}^j
            end for
        end for
    end for
  4. Knowl 4 — Generalization Bound for Decentralized Sparse Personalized Federated Learning

    theoretical result

    Let D~\widetilde{\mathcal{D}} be the union distribution over the client local distributions D1,…,DK\mathcal{D}_1, \dots, \mathcal{D}_K, and let SS be the training sample set of total size NN. Define the expected risk R(A(S))=Ex∼D~[L(w;x)]R(A(S)) = \mathbb{E}_{x \sim \widetilde{\mathcal{D}}}[\mathcal{L}(w; x)] and empirical risk R^S(A(S))=1K∑k=1K1nk∑i=1nkL(w;xi)\widehat{R}_S(A(S)) = \frac{1}{K}\sum_{k=1}^K \frac{1}{n_k}\sum_{i=1}^{n_k} \mathcal{L}(w; x_i) for the hypothesis A(S)=wA(S) = w produced by Dis-PFL.

    Assume the loss function satisfies ∥L∥∞≤1\|\mathcal{L}\|_\infty \le 1, the gradient space has maximum diameter Dg:=max⁡W,z,z′∥∇ℓ(z,W)−∇ℓ(z′,W)∥D_g := \max_{W, z, z'} \|\nabla \ell(z, W) - \nabla \ell(z', W)\|, mini-batch size is τ\tau, training iterations is TT, Gaussian noise variance is σ2\sigma^2, and β∈(0,1]\beta \in (0, 1] represents the proportion of remaining non-zero parameters across the aggregated client masks (the complement of sparsity).

    If the training sample size satisfies N≥2ε′2ln⁡(16e−ε′δ′)N \ge \frac{2}{\varepsilon'^2} \ln\left(\frac{16}{e^{-\varepsilon'}\delta'}\right), then for any distribution D~\widetilde{\mathcal{D}}, with probability at least 1−e−ε′δ′ε′ln⁡(2ε′)1 - \frac{e^{-\varepsilon'}\delta'}{\varepsilon'}\ln\left(\frac{2}{\varepsilon'}\right):

    ∣R^S(A(S))−R(A(S))∣<9ε′|\widehat{R}_S(A(S)) - R(A(S))| < 9\varepsilon'

    where ε′\varepsilon' and δ′\delta' are functions of TT, ε~\widetilde{\varepsilon}, and δ\delta, with ε~\widetilde{\varepsilon} defined as:

    ε~=log⁡(N−τN+τNexp⁡(2βDgστlog⁡1δ+β2Dg22τ2σ2))\widetilde{\varepsilon} = \log\left( \frac{N - \tau}{N} + \frac{\tau}{N} \exp\left( \frac{\sqrt{2}\beta D_g \sigma}{\tau} \sqrt{\log\frac{1}{\delta}} + \frac{\beta^2 D_g^2}{2\tau^2 \sigma^2} \right) \right)

    Because ε~\widetilde{\varepsilon} is strictly monotonically increasing with respect to β\beta, a sparser network (smaller remaining parameter ratio β\beta) yields a smaller ε~\widetilde{\varepsilon}, leading to a tighter generalization bound between empirical and expected risk.

  5. Knowl 5 — Dis-PFL Performance Benchmark Across Non-IID Partitions

    data/table

    Performance of Dis-PFL compared against centralized and decentralized baselines on CIFAR-10, CIFAR-100, and Tiny-ImageNet using ResNet-18 (with Group Normalization) over 100 clients. Two non-IID data partitions are evaluated: Dirichlet partition (extDir(α) ext{Dir}(\alpha) with α=0.3\alpha=0.3 for CIFAR-10, α=0.2\alpha=0.2 for CIFAR-100 and Tiny-ImageNet) and Pathological partition (2 classes per client for CIFAR-10, 10 for CIFAR-100, 20 for Tiny-ImageNet). The baseline D-PSGD is extended to decentralized federated settings with multi-epoch local training, and FT indicates local fine-tuning on the consensus model.

    Task Methods Dir Part Acc (%) Path Part Acc (%) Comm (MB) FLOPs (101210^{12})
    CIFAR-10 Local 61.55±0.261.55 \pm 0.2 86.48±0.286.48 \pm 0.2 - 8.3
    FedAvg 78.07±0.578.07 \pm 0.5 54.53±0.654.53 \pm 0.6 446.9 8.3
    FedAvg-FT 81.20±0.581.20 \pm 0.5 84.96±0.284.96 \pm 0.2 446.9 8.3
    D-PSGD 79.02±0.479.02 \pm 0.4 58.07±0.558.07 \pm 0.5 446.9 8.3
    D-PSGD-FT 83.90±0.283.90 \pm 0.2 90.87±0.290.87 \pm 0.2 446.9 8.3
    Ditto 74.68±0.274.68 \pm 0.2 87.73±0.187.73 \pm 0.1 446.9 8.3
    FOMO 64.68±0.264.68 \pm 0.2 88.24±0.188.24 \pm 0.1 446.9 8.3
    SubFedAvg 76.70±0.276.70 \pm 0.2 88.30±0.288.30 \pm 0.2 278.8 4.7
    Dis-PFL 85.70±0.2\mathbf{85.70 \pm 0.2} 91.05±0.2\mathbf{91.05 \pm 0.2} 223.4 7.0
    CIFAR-100 Local 29.23±0.229.23 \pm 0.2 52.46±0.252.46 \pm 0.2 - 8.3
    FedAvg 41.72±0.541.72 \pm 0.5 33.24±0.633.24 \pm 0.6 448.7 8.3
    FedAvg-FT 49.19±0.549.19 \pm 0.5 63.53±0.763.53 \pm 0.7 448.7 8.3
    D-PSGD 41.87±0.441.87 \pm 0.4 35.42±0.235.42 \pm 0.2 448.7 8.3
    D-PSGD-FT 51.42±0.451.42 \pm 0.4 67.24±0.167.24 \pm 0.1 448.7 8.3
    Ditto 38.26±0.238.26 \pm 0.2 54.02±0.354.02 \pm 0.3 448.7 8.3
    FOMO 28.39±0.128.39 \pm 0.1 52.74±0.152.74 \pm 0.1 448.7 8.3
    SubFedAvg 43.91±0.243.91 \pm 0.2 60.67±0.160.67 \pm 0.1 346.6 5.7
    Dis-PFL 53.48±0.3\mathbf{53.48 \pm 0.3} 68.64±0.4\mathbf{68.64 \pm 0.4} 224.3 7.0
    Tiny-Imagenet Local 6.76±0.26.76 \pm 0.2 17.68±0.317.68 \pm 0.3 - 66.6
    FedAvg 12.30±0.312.30 \pm 0.3 10.40±0.310.40 \pm 0.3 450.7 66.6
    FedAvg-FT 14.80±0.214.80 \pm 0.2 28.30±0.228.30 \pm 0.2 450.7 66.6
    D-PSGD 12.13±0.512.13 \pm 0.5 16.50±0.416.50 \pm 0.4 450.7 66.6
    D-PSGD-FT 15.50±0.315.50 \pm 0.3 28.60±0.328.60 \pm 0.3 450.7 66.6
    Ditto 15.69±0.215.69 \pm 0.2 24.55±0.324.55 \pm 0.3 450.7 66.6
    FOMO 5.20±0.45.20 \pm 0.4 9.39±0.39.39 \pm 0.3 450.7 66.6
    SubFedAvg 12.18±0.412.18 \pm 0.4 19.73±0.519.73 \pm 0.5 290.9 40.2
    Dis-PFL 16.95±0.4\mathbf{16.95 \pm 0.4} 31.71±0.4\mathbf{31.71 \pm 0.4} 225.3 54.5

    Comm denotes the maximum communication traffic (download/upload in MB) of the busiest node across all clients (or the central server in centralized methods) per communication round, and FLOPs denotes the total floating-point operations per client during the local training phase of one round. Dis-PFL (set to 0.5 sparsity) attains higher personalized test accuracy across all tasks while reducing peak communication volume by approximately 50%50\% relative to dense decentralized/centralized baselines.

  6. Knowl 6 — Dis-PFL Convergence Speed Across Non-IID Partitions

    empirical result

    Dis-PFL converges in substantially fewer global communication rounds to achieve target local test accuracy thresholds compared to both centralized and decentralized baselines under 100-client Dirichlet and Pathological non-IID partitions on ResNet-18:

    • CIFAR-10 (Dirichlet partition): Dis-PFL reaches 60%60\%, 70%70\%, and 80%80\% test accuracy in 59, 144, and 301 rounds respectively, whereas FedAvg, D-PSGD, Ditto, FOMO, and SubFedAvg fail to reach 80%80\% accuracy within 500 rounds.
    • CIFAR-10 (Pathological partition): Dis-PFL reaches 50%50\%, 80%80\%, and 85%85\% accuracy in 3, 33, and 81 rounds, compared to SubFedAvg (33, 99, 181 rounds) and Ditto (21, 138, 256 rounds).
    • CIFAR-100 (Dirichlet partition): Dis-PFL reaches 25%25\%, 40%40\%, and 50%50\% accuracy in 64, 231, and 393 rounds; all other baselines fail to achieve 50%50\% accuracy within 500 rounds.
    • CIFAR-100 (Pathological partition): Dis-PFL reaches 30%30\%, 50%50\%, and 60%60\% accuracy in 29, 159, and 293 rounds; the next best baseline, SubFedAvg, requires 148, 202, and 377 rounds.
    • Tiny-ImageNet (Dirichlet partition): Dis-PFL reaches 5%5\%, 10%10\%, and 15%15\% accuracy in 13, 58, and 195 rounds; the closest baseline, Ditto, requires 39, 137, and 261 rounds, while all others fail to reach 15%15\% within 300 rounds.
  7. Knowl 7 — Impact of Network Topology on Decentralized Federated Learning Performance

    data/table

    Comparison of local test accuracy, maximum communication per round (Comm), and local compute FLOPs on CIFAR-10 across different communication topologies: Isolated Local Training, Ring topology (each node connects to 2 fixed neighbors), and Fully-Connected topology (FC, each node connects to all 99 neighbors) with 100 clients.

    Topology Method Dir Part Acc (%) Path Part Acc (%) Comm (MB) FLOPs (101210^{12})
    Separate Local 61.66±0.261.66 \pm 0.2 86.48±0.286.48 \pm 0.2 - 8.3
    Ring D-PSGD 49.46±0.249.46 \pm 0.2 24.42±0.524.42 \pm 0.5 89.4 8.3
    D-PSGD-FT 67.80±0.367.80 \pm 0.3 86.68±0.286.68 \pm 0.2 89.4 8.3
    Dis-PFL 67.81±0.2\mathbf{67.81 \pm 0.2} 86.70±0.2\mathbf{86.70 \pm 0.2} 44.6 7.0
    Fully-connected D-PSGD 79.56±0.279.56 \pm 0.2 60.45±0.360.45 \pm 0.3 4423.9 8.3
    D-PSGD-FT 84.57±0.284.57 \pm 0.2 90.58±0.390.58 \pm 0.3 4423.9 8.3
    Dis-PFL 86.71±0.2\mathbf{86.71 \pm 0.2} 91.14±0.1\mathbf{91.14 \pm 0.1} 2211.4 7.0

    While dense consensus methods like D-PSGD suffer severe performance collapses in sparse topologies such as Ring under non-IID data (49.46%49.46\% Dirichlet, 24.42%24.42\% Pathological), Dis-PFL maintains high personalization accuracy across topologies while consuming exactly half the communication bandwidth of dense methods in every topology configuration.

  8. Knowl 8 — Dis-PFL Performance under Client Hardware Heterogeneity

    data/table

    Evaluation of Dis-PFL and decentralized baselines on 100 clients under two hardware constraint settings on CIFAR-10 with ResNet-18 and VGG-11 architectures:

    • Setting (i) Homogeneous constraint: All clients can compute, store, and transmit at most 50%50\% of the dense model parameters.
    • Setting (ii) Heterogeneous constraint: Clients are divided into 5 groups of 20 clients each, possessing resource budgets of 20%20\%, 40%40\%, 60%60\%, 80%80\%, and 100%100\% parameter capacity. Dis-PFL assigns corresponding sparsities {0.8,0.6,0.4,0.2,0.0}\{0.8, 0.6, 0.4, 0.2, 0.0\} to each group, whereas baseline D-PSGD/D-PSGD-FT must truncate the global model to 20%20\% capacity to accommodate the weakest device.
    ResNet-18 VGG-11
    Methods Dir Acc Path Acc Comm FLOPs Dir Acc Path Acc Comm FLOPs
    (%) (%) (MB) (101110^{11}) (%) (%) (MB) (101110^{11})
    D-PSGD (20% params) 72.24 59.83 89.4 20.9 66.00 59.60 73.8 5.7
    D-PSGD-FT (20% params) 78.46 87.74 89.4 20.9 74.71 87.76 73.8 5.7
    D-PSGD (50% params) 79.02 62.12 223.4 45.8 76.71 73.60 184.6 12.6
    D-PSGD-FT (50% params) 83.77 88.76 223.4 45.8 82.60 91.94 184.6 12.6
    Dis-PFL (Setting ii) 84.33 90.84 268.0 71.3 83.04 91.02 221.4 18.0
    Dis-PFL (Setting i) 84.85 91.27 223.4 70.4 84.95 92.11 184.6 17.3

    Comm denotes average communication cost (MB) per node per round. Dis-PFL natively accommodates varying client compute and memory constraints by assigning personalized sparsity ratios, significantly outperforming capacity-constrained dense baselines.

  9. Knowl 9 — Effect of Model Sparsity Ratio on Accuracy and Efficiency

    data/table

    Evaluation of Dis-PFL on CIFAR-10 with 100 clients under Dirichlet partition (α=0.3\alpha = 0.3) across different uniform sparsity ratios s∈{0.2,0.4,0.5,0.6,0.8}s \in \{0.2, 0.4, 0.5, 0.6, 0.8\} (where remaining parameter proportion is β=1−s\beta = 1 - s):

    Sparsity 0.8 0.6 0.5 0.4 0.2
    Test Accuracy (%) 83.27 84.08 85.70 84.22 84.10
    Comm (MB) 89.30 178.69 223.4 268.1 357.5
    FLOPs (101210^{12}) 4.6 6.5 7.0 7.5 8.2

    Communication (Comm) and computation FLOPs decrease monotonically as sparsity increases. However, personalization accuracy exhibits an empirical trade-off: a high sparsity ratio (0.80.8) limits mask overlap between neighbors and causes slight accuracy drops (83.27%83.27\%), while a low sparsity ratio (0.20.2) restricts mask differentiation across clients (84.10%84.10\%). An intermediate sparsity of 0.50.5 provides the optimal balance between information sharing and client specialization, achieving the highest accuracy (85.70%85.70\%).

  10. Knowl 10 — Correlation Between Learned Sparse Masks and Client Label Distributions

    empirical result

    In a 20-node CIFAR-10 experiment where clients are partitioned into 4 distinct groups sharing similar Dirichlet label distributions (extDir(0.3) ext{Dir}(0.3)), Dis-PFL produces sparse masks whose pairwise distances align with the clients' underlying data distributions.

    Task similarity between any two clients is measured by the cosine similarity of their label distribution vectors, while mask distance is measured by the aligned Hamming distance between their learned binary sparse masks mim_i and mjm_j. Clients belonging to the same cluster (sharing high label distribution cosine similarity) consistently converge to sparse masks with small Hamming distances, while clients in disparate clusters learn distinct masks with large Hamming distances. This demonstrates that dynamic decentralized sparse training autonomously discovers personalized sub-networks that reflect client-level data similarities.

Coverage note — None was omitted; all key theoretical bounds, algorithmic designs, benchmarks across multiple datasets and topologies, hardware heterogeneity analyses, and mask correlation studies were converted into standalone knowls. Intermediate steps of the differential privacy proofs were excluded per instructions.

References

  1. 1.Acar, D. A. E., Zhao, Y., Navarro, R. M., Mattina, M., Whatmough, P. N., and Saligrama, V. Federated learning based on dynamic regularization. In International Conference on Learning Representations, 2021.
  2. 2.Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M. Qsgd: Communication-efficient sgd via gradient quantization and encoding. Advances in Neural Information Processing Systems, 30:1709–1720, 2017.
  3. 3.Allen-Zhu, Z., Li, Y., and Song, Z. A convergence theory for deep learning via over-parameterization. In International Conference on Machine Learning, pp. 242–252. PMLR, 2019.
  4. 4.Arivazhagan, M. G., Aggarwal, V., Singh, A. K., and Choudhary, S. Federated learning with personalization layers. arXiv preprint arXiv:1912.00818, 2019.
  5. 5.Balle, B., Barthe, G., and Gaboardi, M. Privacy amplification by subsampling: Tight analyses via couplings and divergences. In NeurIPS, 2018.
  6. 6.Basu, D., Data, D., Karakus, C., and Diggavi, S. Qsparselocal-sgd: Distributed sgd with quantization, sparsification and local computations. Advances in Neural Information Processing Systems, 32, 2019.
  7. 7.Bibikar, S., Vikalo, H., Wang, Z., and Chen, X. Federated dynamic sparse training: Computing less, communicating less, yet learning better. arXiv preprint arXiv:2112.09824, 2021.
  8. 8.Blot, M., Picard, D., Cord, M., and Thome, N. Gossip training for deep learning. arXiv preprint arXiv:1611.09726, 2016.
  9. 9.Chen, C., Shen, L., Huang, H., Liu, W., and Luo, Z.-Q. Efficient-adam: Communication-efficient distributed adam with complexity analysis. 2020.
  10. 10.Chen, C., Shen, L., Huang, H., and Liu, W. Quantized adam with error feedback. ACM Transactions on Intelligent Systems and Technology (TIST), 12(5):1–26, 2021a.
  11. 11.Chen, C., Zhang, J., Shen, L., Zhao, P., and Luo, Z. Communication efficient primal-dual algorithm for nonconvex nonsmooth distributed optimization. In International Conference on Artificial Intelligence and Statistics, pp. 1594–1602. PMLR, 2021b.
  12. 12.Chen, H.-Y. and Chao, W.-L. Fedbe: Making bayesian model ensemble applicable to federated learning. In International Conference on Learning Representations, 2020.
  13. 13.Chen, H.-Y. and Chao, W.-L. On bridging generic and personalized federated learning. arXiv preprint arXiv:2107.00778, 2021.
  14. 14.Cheng, G., Chadha, K., and Duchi, J. Fine-tuning is fine in federated learning. arXiv preprint arXiv:2108.07313, 2021.
  15. 15.Collins, L., Hassani, H., Mokhtari, A., and Shakkottai, S. Exploiting shared representations for personalized federated learning. arXiv preprint arXiv:2102.07078, 2021.
  16. 16.Deng, Y., Kamani, M. M., and Mahdavi, M. Adaptive personalized federated learning. arXiv preprint arXiv:2003.13461, 2020.
  17. 17.Diao, E., Ding, J., and Tarokh, V. Heterofl: Computation and communication efficient federated learning for heterogeneous clients. In International Conference on Learning Representations, 2020.
  18. 18.Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E. Rigging the lottery: Making all tickets winners. In International Conference on Machine Learning, pp. 2943–2952. PMLR, 2020.
  19. 19.Fallah, A., Mokhtari, A., and Ozdaglar, A. Personalized federated learning: A meta-learning approach. arXiv preprint arXiv:2002.07948, 2020.
  20. 20.Fraboni, Y., Vidal, R., Kameni, L., and Lorenzi, M. Clustered sampling: Low-variance and improved representativity for clients selection in federated learning. In International Conference on Machine Learning, pp. 3407–3416. PMLR, 2021.
  21. 21.Frankle, J. and Carbin, M. The lottery ticket hypothesis: Finding sparse, trainable neural networks. In International Conference on Learning Representations, 2018.
  22. 22.Han, S., Pool, J., Tran, J., and Dally, W. Learning both weights and connections for efficient neural network. Advances in Neural Information Processing Systems, 28, 2015.
  23. 23.Hanzely, F., Zhao, B., and Kolar, M. Personalized federated learning: A unified framework and universal optimization techniques. arXiv preprint arXiv:2102.09743, 2021.
  24. 24.He, F. and Tao, D. Recent advances in deep learning theory. arXiv preprint arXiv:2012.10931, 2020.
  25. 25.He, F., Liu, T., and Tao, D. Control batch size and learning rate to generalize well: Theoretical and empirical evidence. In Advances in Neural Information Processing Systems, pp. 1143–1152, 2019.
  26. 26.He, F., Wang, B., and Tao, D. Tighter generalization bounds for iterative differentially private learning algorithms. In Uncertainty in Artificial Intelligence, pp. 802–812. PMLR, 2021.
  27. 27.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  28. 28.Hong, J., Wang, H., Wang, Z., and Zhou, J. Efficient splitmix federated learning for on-demand and in-situ customization. arXiv preprint arXiv:2203.09747, 2022.
  29. 29.Horvath, S., Laskaridis, S., Almeida, M., Leontiadis, I., Venieris, S., and Lane, N. Fjord: Fair and accurate federated learning under heterogeneous targets with ordered dropout. Advances in Neural Information Processing Systems, 34, 2021.
  30. 30.Hsieh, K., Phanishayee, A., Mutlu, O., and Gibbons, P. The non-iid data quagmire of decentralized machine learning. In International Conference on Machine Learning, pp. 4387–4398. PMLR, 2020.
  31. 31.Hsu, T.-M. H., Qi, H., and Brown, M. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019.
  32. 32.Huang, T., Lin, W., Shen, L., Li, K., and Zomaya, A. Y. Stochastic client selection for federated learning with volatile clients. IEEE Internet of Things Journal, 2022a.
  33. 33.Huang, T., Liu, S., Shen, L., He, F., Lin, W., and Tao, D. Achieving personalized federated learning with sparse local models. arXiv preprint arXiv:2201.11380, 2022b.
  34. 34.Ivkin, N., Rothchild, D., Ullah, E., Stoica, I., Arora, R., et al. Communication-efficient distributed sgd with sketching. Advances in Neural Information Processing Systems, 32: 13144–13154, 2019.
  35. 35.Jiang, Y., Konečny, J., Rush, K., and Kannan, S. Improving federated learning personalization via model agnostic meta learning. arXiv preprint arXiv:1909.12488, 2019.
  36. 36.Jiang, Z., Balu, A., Hegde, C., and Sarkar, S. Collaborative deep learning in fixed topology networks. Advances in Neural Information Processing Systems, 30, 2017.
  37. 37.Khan, L. U., Saad, W., Han, Z., Hossain, E., and Hong, C. S. Federated learning for internet of things: Recent advances, taxonomy, and open challenges. IEEE Communications Surveys & Tutorials, 2021.
  38. 38.Koloskova, A., Stich, S., and Jaggi, M. Decentralized stochastic optimization and gossip algorithms with compressed communication. In International Conference on Machine Learning, pp. 3478–3487. PMLR, 2019.
  39. 39.Krizhevsky, A. et al. Learning multiple layers of features from tiny images. 2009.
  40. 40.Lalitha, A., Shekhar, S., Javidi, T., and Koushanfar, F. Fully decentralized federated learning. In Third workshop on Bayesian Deep Learning (NeurIPS), 2018.
  41. 41.Lalitha, A., Kilinc, O. C., Javidi, T., and Koushanfar, F. Peer-to-peer federated learning on graphs. arXiv preprint arXiv:1901.11173, 2019.
  42. 42.Li, A., Sun, J., Wang, B., Duan, L., Li, S., Chen, Y., and Li, H. Lotteryfl: Personalized and communication-efficient federated learning with lottery ticket hypothesis on non-iid datasets. arXiv preprint arXiv:2008.03371, 2020a.
  43. 43.Li, A., Sun, J., Zeng, X., Zhang, M., Li, H., and Chen, Y. Fedmask: Joint computation and communication-efficient personalized federated learning via heterogeneous masking. In Proceedings of the 19th ACM Conference on Embedded Networked Sensor Systems, pp. 42–55, 2021a.
  44. 44.Li, S., Zhou, T., Tian, X., and Tao, D. Learning to collaborate in decentralized learning of personalized models. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2022.
  45. 45.Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V. Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems, 2:429–450, 2020b.
  46. 46.Li, T., Hu, S., Beirami, A., and Smith, V. Ditto: Fair and robust federated learning through personalization. In International Conference on Machine Learning, pp. 6357–6368. PMLR, 2021b.
  47. 47.Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J. Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent. Advances in Neural Information Processing Systems, 2017.
  48. 48.Liang, P. P., Liu, T., Ziyin, L., Allen, N. B., Auerbach, R. P., Brent, D., Salakhutdinov, R., and Morency, L.-P. Think locally, act globally: Federated learning with local and global representations. arXiv preprint arXiv:2001.01523, 2020.
  49. 49.Lim, H., Andersen, D. G., and Kaminsky, M. 3lc: Lightweight and effective traffic compression for distributed machine learning. Proceedings of Machine Learning and Systems, 1:53–64, 2019.
  50. 50.Lin, T., Kong, L., Stich, S. U., and Jaggi, M. Ensemble distillation for robust model fusion in federated learning. In NeurIPS, 2020.
  51. 51.Lin, T., Karimireddy, S. P., Stich, S. U., and Jaggi, M. Quasi-global momentum: Accelerating decentralized deep learning on heterogeneous data. In International Conference on Machine Learning, pp. 6654–6665. PMLR, 2021.
  52. 52.Liu, S., Chen, T., Chen, X., Atashgahi, Z., Yin, L., Kou, H., Shen, L., Pechenizkiy, M., Wang, Z., and Mocanu, D. C. Sparse training via boosting pruning plasticity with neuroregeneration. Advances in Neural Information Processing Systems, 34, 2021a.
  53. 53.Liu, S., Yin, L., Mocanu, D. C., and Pechenizkiy, M. Do we actually need dense over-parameterization? in-time over-parameterization in sparse training. In International Conference on Machine Learning, pp. 6989–7000. PMLR, 2021b.
  54. 54.Liu, S., Chen, T., Chen, X., Shen, L., Mocanu, D. C., Wang, Z., and Pechenizkiy, M. The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training. In International Conference on Learning Representations, 2022a.
  55. 55.Liu, S., Tian, Y., Chen, T., and Shen, L. Don’t be so dense: Sparse-to-sparse gan training without sacrificing performance. arXiv preprint arXiv:2203.02770, 2022b.
  56. 56.Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T. Rethinking the value of network pruning. In International Conference on Learning Representations, 2018.
  57. 57.McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A. Communication-efficient learning of deep networks from decentralized data. In Artificial intelligence and statistics, pp. 1273–1282. PMLR, 2017.
  58. 58.Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A. Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nature communications, 9 (1):1–12, 2018.
  59. 59.Mohri, M., Rostamizadeh, A., and Talwalkar, A. Foundations of machine learning. 2018.
  60. 60.Mostafa, H. and Wang, X. Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization. In International Conference on Machine Learning, pp. 4646–4655. PMLR, 2019.
  61. 61.Nguyen, D. C., Ding, M., Pathirana, P. N., Seneviratne, A., Li, J., and Poor, H. V. Federated learning for internet of things: A comprehensive survey. IEEE Communications Surveys & Tutorials, 2021.
  62. 62.Nishio, T. and Yonetani, R. Client selection for federated learning with heterogeneous resources in mobile edge. In ICC 2019-2019 IEEE international conference on communications (ICC), pp. 1–7. IEEE, 2019.
  63. 63.Shamsian, A., Navon, A., Fetaya, E., and Chechik, G. Personalized federated learning using hypernetworks. arXiv preprint arXiv:2103.04628, 2021.
  64. 64.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, May 2015.
  65. 65.Sun, T., Li, D., and Wang, B. Decentralized federated averaging. arXiv preprint arXiv:2104.11375, 2021.
  66. 66.T Dinh, C., Tran, N., and Nguyen, T. D. Personalized federated learning with moreau envelopes. Advances in Neural Information Processing Systems, 33, 2020.
  67. 67.Vahidian, S., Morafah, M., and Lin, B. Personalized federated learning by structured and unstructured pruning under data heterogeneity. arXiv preprint arXiv:2105.00562, 2021.
  68. 68.Voigt, P. and Von dem Bussche, A. The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing, 10:3152676, 2017.
  69. 69.Wang, H., Yurochkin, M., Sun, Y., Papailiopoulos, D., and Khazaeni, Y. Federated learning with matched averaging. In International Conference on Learning Representations, 2020.
  70. 70.Warnat-Herresthal, S., Schultze, H., Shastry, K. L., Manamohan, S., Mukherjee, S., Garg, V., Sarveswara, R., Händler, K., Pickkers, P., Aziz, N. A., et al. Swarm learning for decentralized and confidential clinical machine learning. Nature, 594(7862):265–270, 2021.
  71. 71.Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H. Terngrad: Ternary gradients to reduce communication in distributed deep learning. Advances in Neural Information Processing Systems, 30, 2017.
  72. 72.Wu, Y. and He, K. Group normalization. In Proceedings of the European conference on computer vision (ECCV), pp. 3–19, 2018.
  73. 73.Yu, Y., Wu, J., and Huang, J. Exploring fast and communication-efficient algorithms in large-scale distributed networks. In The 22nd International Conference on Artificial Intelligence and Statistics, pp. 674–683. PMLR, 2019.
  74. 74.Yuan, Y., Chen, R., Sun, C., Wang, M., Hua, F., Yi, X., Yang, T., and Liu, J. Defed: A principled decentralized and privacy-preserving federated learning algorithm. arXiv preprint arXiv:2107.07171, 2021.
  75. 75.Zhang, L., Shen, L., Ding, L., Tao, D., and Duan, L.-Y. Fine-tuning global model via data-free knowledge distillation for non-iid federated learning. arXiv preprint arXiv:2203.09249, 2022.
  76. 76.Zhang, M., Sapra, K., Fidler, S., Yeung, S., and Alvarez, J. M. Personalized federated learning with first order model optimization. In International Conference on Learning Representations, 2020.
  77. 77.Zhu, T., He, F., Zhang, L., Niu, Z., Song, M., and Tao, D. Topology-aware generalization of decentralized sgd. In International Conference on Machine Learning, 2022.
  78. 78.Zou, D. and Gu, Q. An improved analysis of training overparameterized deep neural networks. Advances in Neural Information Processing Systems, 32:2055–2064, 2019.

Citation

MLA
Dai, R., et al. “DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training”. International Conference on Machine Learning, vol. 162, 2022, pp. 4587–604, https://proceedings.mlr.press/v162/dai22b.html.
APA
Dai, R., Shen, L., He, F., Tian, X., & Tao, D. (2022). DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training. International Conference on Machine Learning, 162, 4587–4604. https://proceedings.mlr.press/v162/dai22b.html
Chicago
Dai, R., L. Shen, F. He, X. Tian, and D. Tao. 2022. “DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training”. International Conference on Machine Learning 162: 4587–4604. https://proceedings.mlr.press/v162/dai22b.html.
Harvard
Dai, R. et al. (2022) “DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training”, International Conference on Machine Learning. PMLR, pp. 4587–4604. Available at: https://proceedings.mlr.press/v162/dai22b.html.
Vancouver
1. Dai R, Shen L, He F, Tian X, Tao D (2022) DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training. In: International Conference on Machine Learning. PMLR, pp 4587–4604

BibTeX

@InProceedings{pmlr-v162-dai22b,
  title = 	 {{D}is{PFL}: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training},
  author =       {Dai, Rong and Shen, Li and He, Fengxiang and Tian, Xinmei and Tao, Dacheng},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {4587--4604},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/dai22b/dai22b.pdf},
  url = 	 {https://proceedings.mlr.press/v162/dai22b.html},
  abstract = 	 {Personalized federated learning is proposed to handle the data heterogeneity problem amongst clients by learning dedicated tailored local models for each user. However, existing works are often built in a centralized way, leading to high communication pressure and high vulnerability when a failure or an attack on the central server occurs. In this work, we propose a novel personalized federated learning framework in a decentralized (peer-to-peer) communication protocol named DisPFL, which employs personalized sparse masks to customize sparse local models on the edge. To further save the communication and computation cost, we propose a decentralized sparse training technique, which means that each local model in DisPFL only maintains a fixed number of active parameters throughout the whole local training and peer-to-peer communication process. Comprehensive experiments demonstrate that DisPFL significantly saves the communication bottleneck for the busiest node among all clients and, at the same time, achieves higher model accuracy with less computation cost and communication rounds. Furthermore, we demonstrate that our method can easily adapt to heterogeneous local clients with varying computation complexities and achieves better personalized performances.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/