Improving the Model Consistency of Decentralized Federated Learning

Yifan ShiLi ShenKang WeiYan SunBo YuanXueqian WangDacheng Tao

article2023ICML103 citations

Proposes two decentralized federated learning algorithms that combine Sharpness Aware Minimization with multiple gossip communication steps to reduce client model inconsistency, providing non-convex convergence guarantees and matching centralized performance on heterogeneous data.

Listen

Decentralized federated learning enables distributed participants to collaboratively train machine learning models by communicating directly with neighboring devices rather than through a central coordinating server. This structure improves data privacy and reduces central communication bottlenecks across edge networks, healthcare systems, and connected devices. However, because local data is often non-uniform and network connections can be sparse, participating devices frequently develop highly inconsistent models, causing severe performance drops and poor generalization compared to traditional centralized approaches.

The main objective of the article is to develop and evaluate two new decentralized optimization algorithms, named DFedSAM and DFedSAM-MGS, designed to eliminate model inconsistency and achieve performance on par with centralized training. The proposed methods achieve this by having each local device seek flatter, more robust loss regions and by accelerating neighbor-to-neighbor consensus through multiple communication exchanges per training round.

The authors evaluated the approaches using both mathematical convergence proofs and extensive empirical experiments. The experimental evaluation tested image classification tasks across 100 simulated clients on standard benchmarks under both uniform and highly skewed data distributions. The testing examined diverse communication network topologies, ranging from sparsely connected rings and grids to fully connected structures, comparing results against leading centralized and decentralized baselines.

The findings demonstrate substantial performance and consistency improvements. First, the proposed DFedSAM and DFedSAM-MGS consistently outperformed existing decentralized methods across all benchmark tasks and data distributions. On benchmark datasets, DFedSAM-MGS achieved test accuracies comparable to state-of-the-art centralized baselines, such as reaching 84.26% to 86.47% accuracy under varying data skew on standard image tasks. Second, the methods demonstrated superior resilience on sparse network topologies; in restrictive ring networks, DFedSAM-MGS improved accuracy by 8.0 percentage points over standard decentralized momentum approaches. Third, Hessian eigenvalue analyses verified that the algorithms successfully produce flatter loss landscapes, reducing model overfitting across local nodes. Finally, theoretical analysis confirmed that increasing consensus communication steps effectively neutralizes the negative convergence impact of sparse network designs.

These results show that decentralized networks can match centralized model quality without requiring central server infrastructure, significantly lowering operational risks and bandwidth costs. Organizations deploying edge machine learning in privacy-sensitive or infrastructure-constrained environments can adopt these methods to achieve reliable model accuracy across heterogeneous participants.

Decision-makers should consider adopting flatness-aware training and multi-step neighbor communication when deploying peer-to-peer federated systems. Practitioners should tune the number of consensus steps to balance network communication volume against target model accuracy, as a moderate consensus setting of four communication steps provided the best operational trade-off in the study. Before full deployment, teams should conduct pilot tests to optimize hyperparameter settings for their specific network hardware and data heterogeneity levels.

arXiv: 2302.04083

No sufficiently relevant recommendations were found.

Cover for Improving the Model Consistency of Decentralized Federated Learning

Abstract

To mitigate the privacy leakages and communication burdens of Federated Learning (FL), decentralized FL (DFL) discards the central server and each client only communicates with its neighbors in a decentralized communication network. However, existing DFL suffers from high inconsistency among local clients, which results in severe distribution shift and inferior performance compared with centralized FL (CFL), especially on heterogeneous data or sparse communication topologies. To alleviate this issue, we propose two DFL algorithms named DFedSAM and DFedSAM-MGS to improve the performance of DFL. Specifically, DFedSAM leverages gradient perturbation to generate local flat models via Sharpness Aware Minimization (SAM), which searches for models with uniformly low loss values. DFedSAM-MGS further boosts DFedSAM by adopting Multiple Gossip Steps (MGS) for better model consistency, which accelerates the aggregation of local flat models and better balances communication complexity and generalization. Theoretically, we present improved convergence rates O(1KT+1T+1K1/2T3/2(1−λ)2)\mathcal{O}\left(\frac{1}{\sqrt{K T}}+\frac{1}{T}+\frac{1}{K^{1 / 2} T^{3 / 2}(1-\lambda)^2}\right) and O(1KT+1T+λQ+1K1/2T3/2(1−λQ)2)\mathcal{O}\left(\frac{1}{\sqrt{K T}}+\frac{1}{T}+\frac{\lambda^{Q+1}}{K^{1 / 2} T^{3 / 2}\left(1-\lambda^Q\right)^2}\right) in non-convex setting for DFedSAM and DFedSAM-MGS, respectively, where 1−λ1 - \lambda is the spectral gap of gossip matrix and QQ is the number of MGS. Empirically, our methods can achieve competitive performance compared with CFL methods and outperform existing DFL methods.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Methodology
  • 3.1. Problem Setting
  • 3.2. DFedSAM and DFedSAM-MG
  • 4. Convergence Analysis
  • 5. Experiments
  • 5.1. Experiment Setup
  • 5.2. Performance Evaluation
  • 5.3. Topology-aware Performance
  • 5.4. Ablation Study
  • 6. Conclusions and Future Work
  • Acknowledgements
  • References
  • A. Communication Network Topologies
  • B. More Details on Algorithm Implementation
  • B.1. Datasets and backbones
  • B.2. More details about baselines
  • B.3. Hyperparameters
  • B.4. Communication configurations
  • C. More Details on Experiments under the Pathological Partition
  • D. Convergence Analysis for DFedSAM and DFedSAM-MGS
  • D.1. Preliminary Lemmas
  • D.2. Proof of convergence results for DFedSAM
  • D.3. Proof of convergence results for DFedSAM-MGS

Knowls

  1. Knowl 1 — DFedSAM uses sharpness-aware local updates before neighbor averaging

    model/method

    DFedSAM is a decentralized federated learning method in which each client performs KK local sharpness-aware updates before communicating with neighbors. At client ii, for local parameter yy and minibatch sample [?]—more conventionally denoted [?]—the method computes an unperturbed stochastic gradient g=∇Fi(y;ξ)g=\nabla F_i(y;\xi), forms a perturbation δ=ρg/∥g∥2\delta=\rho g/\|g\|_2 with radius ρ>0\rho>0, then updates y←y−η∇Fi(y+δ;ξ)y\leftarrow y-\eta\nabla F_i(y+\delta;\xi). Here FiF_i is the client’s sample loss, ξ\xi is a local minibatch, and η\eta is the learning rate; the same minibatch is used for the two gradient evaluations. After KK such updates, each client averages its resulting model with its own and its neighbors’ models using the gossip weights wiℓw_{i\ell}. The intended local objective is to seek parameters with low loss throughout a radius-ρ\rho neighborhood, rather than only at the unperturbed parameter. This use of local flat models is intended to reduce overfitting and inconsistency among client models.

  2. Knowl 2 — DFedSAM-MGS applies repeated gossip averaging to local flat models

    model/method

    DFedSAM-MGS adds QQ successive neighbor-averaging steps after the KK local DFedSAM updates. Let zi(0)z_i^{(0)} be client ii’s model after local training. At gossip step q∈{0,…,Q−1}q\in\{0,\ldots,Q-1\}, every client computes zi(q+1)=∑ℓ∈N(i)wiℓzℓ(q)z_i^{(q+1)}=\sum_{\ell\in\mathcal N(i)}w_{i\ell}z_\ell^{(q)}, where N(i)\mathcal N(i) includes the clients whose models client ii can receive and wiℓw_{i\ell} are entries of the network’s mixing matrix. The model after the QQth step is used for the next communication round. Thus, Q=1Q=1 gives the single averaging step of DFedSAM, while larger QQ spends additional neighbor communications to accelerate consensus among client models. The method is designed to trade communication volume for consistency and generalization; the paper uses Q=4Q=4 in its main comparisons.

  3. Knowl 3 — DFL objective and communication-network model

    definition

    The paper formulates decentralized federated learning as stochastic non-convex empirical risk minimization over mm clients and parameters x∈Rdx\in\mathbb R^d:

    min⁡x∈Rdf(x),f(x)=1m∑i=1mfi(x),fi(x)=Eξ∼Di[Fi(x;ξ)].\min_{x\in\mathbb R^d} f(x),\qquad f(x)=\frac{1}{m}\sum_{i=1}^{m}f_i(x),\qquad f_i(x)=\mathbb E_{\xi\sim D_i}[F_i(x;\xi)].

    Here DiD_i is client ii’s data distribution, which may differ across clients, and Fi(x;ξ)F_i(x;\xi) is its sample loss. The clients communicate over an undirected connected graph and there is no central server; each client exchanges models only with its neighbors. A gossip matrix W=[wij]∈[0,1]m×mW=[w_{ij}]\in[0,1]^{m\times m} assigns the averaging weights and respects the graph’s edges. For the convergence analysis, WW is symmetric, has consensus null space null⁡(I−W)=span⁡{1}\operatorname{null}(I-W)=\operatorname{span}\{\mathbf 1\}, and satisfies I⪰W≻−II\succeq W\succ -I. If its eigenvalues are ordered by magnitude, define λ=max⁡{∣λ2(W)∣,∣λm(W)∣}\lambda=\max\{|\lambda_2(W)|,|\lambda_m(W)|\}; the spectral gap is 1−λ∈(0,1]1-\lambda\in(0,1], with larger λ\lambda corresponding to weaker connectivity.

  4. Knowl 4 — Non-convex stationarity rates depend on heterogeneity and gossip connectivity

    theoretical result

    For DFedSAM and DFedSAM-MGS, the paper bounds the expected squared gradient norm of the average client iterate in the stochastic non-convex setting. The result assumes each client objective is differentiable and LL-smooth, local stochastic-gradient variance is bounded by σl2\sigma_l^2, and inter-client gradient heterogeneity is bounded by β=max⁡isup⁡x∥∇fi(x)−∇f(x)∥\beta=\max_i\sup_x\|\nabla f_i(x)-\nabla f(x)\|; the global gradient variance is also bounded and no greater than β2\beta^2. The gossip matrix and its spectral parameter λ\lambda are as defined for the decentralized network. With local stepsize η=O(1/(LKT))\eta=O(1/(L\sqrt{KT})), subject to η≤1/(10KL)\eta\leq1/(10KL), and perturbation radius ρ=O(T−1/2)\rho=O(T^{-1/2}), the stated rates are:

    min⁡1≤t≤TE∥∇f(xˉt)∥2=O ⁣(f(xˉ1)−f⋆+σl2KT+K(β2+σl2)T+L2K T3/2+β2+σl2K T3/2(1−λ)2)\min_{1\leq t\leq T}\mathbb E\|\nabla f(\bar x^t)\|^2 =O\!\left(\frac{f(\bar x^1)-f^\star+\sigma_l^2}{\sqrt{KT}}+\frac{K(\beta^2+\sigma_l^2)}{T}+\frac{L^2}{\sqrt K\,T^{3/2}}+\frac{\beta^2+\sigma_l^2}{\sqrt K\,T^{3/2}(1-\lambda)^2}\right)

    for DFedSAM, and

    min⁡1≤t≤TE∥∇f(xˉt)∥2=O ⁣(f(xˉ1)−f⋆+σl2KT+K(β2+σl2)T+L2K T3/2+Φ(λ,m,Q)β2+σl2K T3/2),\min_{1\leq t\leq T}\mathbb E\|\nabla f(\bar x^t)\|^2 =O\!\left(\frac{f(\bar x^1)-f^\star+\sigma_l^2}{\sqrt{KT}}+\frac{K(\beta^2+\sigma_l^2)}{T}+\frac{L^2}{\sqrt K\,T^{3/2}}+\Phi(\lambda,m,Q)\frac{\beta^2+\sigma_l^2}{\sqrt K\,T^{3/2}}\right),

    for DFedSAM-MGS, where TT is the number of communication rounds, KK is the number of local updates per round, xˉt\bar x^t is the mean of the mm client models, and f⋆f^\star is the minimum value of ff. The topology-and-gossip factor is

    Φ(λ,m,Q)=λQ+1(1−λ)2m2(Q−1)+λQ+1(1−λQ)2.\Phi(\lambda,m,Q)=\frac{\lambda^Q+1}{(1-\lambda)^2m^{2(Q-1)}}+\frac{\lambda^Q+1}{(1-\lambda^Q)^2}.

    The MGS bound therefore replaces the single-gossip topology penalty with a factor that also depends on the number of gossip steps and clients; the paper’s analysis indicates that additional gossip can reduce the effect of sparse connectivity.

  5. Knowl 5 — Main benchmark protocol compares decentralized and centralized learning

    experimental setup

    The main image-classification experiments evaluate CIFAR-10 and CIFAR-100 under IID data and Dirichlet non-IID partitions with concentration parameters α=0.3\alpha=0.3 and α=0.6\alpha=0.6. VGG-11 is used for CIFAR-10 and ResNet-18 for CIFAR-100; batch-normalization layers are replaced with group normalization. Experiments use 100 clients, minibatches of 128, 1,000 communication rounds, and a local learning rate initialized to 0.1 and decayed by 0.998 per round. The local training epoch count is 5 for all compared methods except D-PSGD, which uses 1. DFedSAM uses ρ=0.01\rho=0.01, and the main DFedSAM-MGS comparison uses Q=4Q=4. To make communication comparisons more comparable, decentralized experiments use time-varying neighbor connections with at most 10 neighbors per client; centralized methods use a 10% client participation ratio. The methods compared include centralized FedAvg and FedSAM, and decentralized D-PSGD, DFedAvg, DFedAvgM, DisPFL, DFedSAM, and DFedSAM-MGS.

  6. Knowl 6 — DFedSAM-MGS improves benchmark results, especially on CIFAR-10

    data/table

    The reported benchmark table gives validation accuracy and generalization error (both in percent). The entries below compare the centralized SAM baseline, DFedSAM, and DFedSAM-MGS; the latter is decentralized and uses Q=4Q=4 and ρ=0.01\rho=0.01. Across the three data partitions, DFedSAM-MGS has lower reported generalization error than the centralized FedSAM baseline on CIFAR-10, while on CIFAR-100 it remains competitive but has higher reported error than FedSAM.

    Dataset Method Dirichlet 0.3 Dirichlet 0.6 IID
    Validation Gen. error Validation Gen. error Validation Gen. error
    CIFAR-10 FedSAM 82.49 16.26 85.04 14.85 84.98 15.00
    DFedSAM 82.04 17.37 84.38 15.06 85.30 14.14
    DFedSAM-MGS 84.26 15.27 85.14 14.51 86.47 13.22
    CIFAR-100 FedSAM 52.98 47.01 55.88 44.11 59.60 40.39
    DFedSAM 48.66 51.21 52.70 47.15 53.12 46.85
    DFedSAM-MGS 52.37 47.55 54.91 45.04 56.15 43.82

    The broader comparison reports that both proposed methods outperform the decentralized baselines overall, and that DFedSAM-MGS approaches centralized performance. Its largest clear advantage in the selected results is CIFAR-10 validation accuracy: it reaches 84.26%, 85.14%, and 86.47% for Dirichlet 0.3, Dirichlet 0.6, and IID data, respectively.

  7. Knowl 7 — Multiple gossip steps help most on sparse communication topologies

    data/table

    On CIFAR-10, the paper compares decentralized methods across ring, grid, exponential, and fully connected communication graphs, ordered from sparser to denser connectivity. DFedSAM-MGS obtains the highest reported accuracy in every topology and shows its largest gain over DFedAvgM on the ring: 80.55% versus 72.55%. The benchmark uses the decentralized setting in which clients can communicate with their neighbors in the same round.

    Method Ring Grid Exponential Fully connected
    D-PSGD 68.96 74.36 74.90 75.35
    DFedAvg 69.95 80.17 83.13 83.48
    DFedAvgM 72.55 85.24 86.94 87.50
    DFedSAM 73.19 85.28 87.44 88.05
    DFedSAM-MGS 80.55 87.39 88.06 88.20

    The results support the claim that additional gossip can mitigate the performance cost of sparse connectivity: DFedSAM-MGS lifts ring accuracy by 8.00 percentage points relative to DFedAvgM, while its advantage is smaller on fully connected graphs.

  8. Knowl 8 — Hessian measurements indicate flatter solutions from SAM-based DFL

    empirical result

    To compare loss-surface flatness, the paper measures the largest Hessian eigenvalue after convergence on CIFAR-10 partitioned with a Dirichlet distribution of α=0.6\alpha=0.6, using VGG-11. The reported values are 182.80 for FedAvg, 578.15 for DFedAvg, 282.22 for DFedSAM, and 181.60 for DFedSAM-MGS. Since the paper uses a smaller largest eigenvalue as its flatness indicator, the measurements show DFedAvg’s solution to be much sharper than FedAvg’s, while DFedSAM-MGS is close to FedAvg by this measure. This is empirical evidence consistent with, but does not itself prove, the proposed flatness-based explanation.

  9. Knowl 9 — Severe class-partition heterogeneity limits shared-model accuracy

    data/table

    A more severe non-IID test assigns each of 100 clients two randomly selected class partitions from 200 sorted data partitions, and evaluates after 500 communication rounds. The reported test accuracies (%) show that the proposed methods remain stronger than the decentralized baselines, but do not eliminate the difficulty of fitting one shared model: CIFAR-100 accuracy remains low for every method, and centralized FedSAM exceeds DFedSAM-MGS on both datasets. The paper identifies this setting as evidence that a single all-client model may be unsuitable under especially strong statistical heterogeneity.

    Method CIFAR-10 CIFAR-100
    FedAvg 63.23 9.66
    FedSAM 65.61 13.07
    D-PSGD 39.53 5.60
    DisPFL 47.61 5.69
    DFedAvg 45.67 5.35
    DFedAvgM 51.64 7.65
    DFedSAM 60.58 9.91
    DFedSAM-MGS 64.45 12.07

    In particular, DFedSAM-MGS is close to FedSAM on CIFAR-10 (64.45% versus 65.61%) and CIFAR-100 (12.07% versus 13.07%), but the absolute CIFAR-100 result illustrates the remaining limitation under this partition.

  10. Knowl 10 — Ablations show accuracy–communication and hyperparameter trade-offs

    empirical result

    The ablations find that increasing the number of gossip steps QQ generally improves model consistency and test performance but increases local communication; the balance point varies by topology, and the authors select Q=4Q=4 for the main DFedSAM-MGS comparisons. In local-update experiments, increasing the training epochs per round among 1, 2, 3, and 5 accelerates convergence in communication rounds. In tests with 50, 100, and 150 participating clients, the smaller client counts (50 or 100) converge better and attain higher test accuracy than 150 under the tested settings; the authors associate this with more local data per client. Perturbation-radius experiments over {0.01,0.025,0.05,0.1,0.2,0.4,0.6,0.8,1.0}\{0.01,0.025,0.05,0.1,0.2,0.4,0.6,0.8,1.0\} identify ρ=0.01\rho=0.01 as a satisfactory convergence–performance trade-off. These are empirical observations from the paper’s CIFAR-10 ablations, not universal parameter-selection guarantees.

Coverage note — No substantial contributed material is deliberately omitted; proof-only intermediate steps and background or related-work material are excluded.

References

  1. 1.Abbas, M., Xiao, Q., Chen, L., Chen, P., and Chen, T. Sharpmaml: Sharpness-aware model-agnostic meta learning. In International Conference on Machine Learning, ICML, pp. 10–32, 2022.
  2. 2.Acar, D. A. E., Zhao, Y., Navarro, R. M., Mattina, M., Whatmough, P. N., and Saligrama, V. Federated learning based on dynamic regularization. In 9th International Conference on Learning Representations, ICLR, 2021.
  3. 3.Andriushchenko, M. and Flammarion, N. Towards understanding sharpness-aware minimization. In International Conference on Machine Learning, ICML, Proceedings of Machine Learning Research, pp. 639–668. PMLR, 2022.
  4. 4.Beltran, E. T. M., P´erez, M. Q., S´anchez, P. M. S., Bernal, S. L., Bovet, G., Perez, M. G., P´erez, G. M., and Celdr´an, A. H. Decentralized federated learning: Fundamentals, state-of-the-art, frameworks, trends, and challenges. arXiv preprint arXiv:2211.08413, 2022.
  5. 5.Bottou, L. Large-scale machine learning with stochastic gradient descent. In Proceedings of COMPSTAT’2010, pp. 177–186. Springer, 2010.
  6. 6.Bottou, L., Curtis, F. E., and Nocedal, J. Optimization methods for large-scale machine learning. Siam Review, 60(2):223–311, 2018.
  7. 7.Caldarola, D., Caputo, B., and Ciccone, M. Improving generalization in federated learning by seeking flat minima. CoRR, abs/2203.11834, 2022.
  8. 8.Chen, C., Zhang, J., Shen, L., Zhao, P., and Luo, Z. Communication efficient primal-dual algorithm for nonconvex nonsmooth distributed optimization. In International Conference on Artificial Intelligence and Statistics, pp. 1594–1602. PMLR, 2021.
  9. 9.Chen, H.-Y. and Chao, W.-L. On bridging generic and personalized federated learning for image classification. arXiv preprint arXiv:2107.00778, 2021.
  10. 10.Dai, R., Shen, L., He, F., Tian, X., and Tao, D. Dispfl: Towards communication-efficient personalized federated learning via decentralized sparse training. In International Conference on Machine Learning, ICML, Proceedings of Machine Learning Research, pp. 4587–4604. PMLR, 2022.
  11. 11.Deng, Y., Kamani, M. M., and Mahdavi, M. Adaptive personalized federated learning. arXiv preprint arXiv:2003.13461, 2020.
  12. 12.Du, J., Yan, H., Feng, J., Zhou, J. T., Zhen, L., Goh, R. S. M., and Tan, V. Efficient sharpness-aware minimization for improved training of neural networks. In International Conference on Learning Representations, 2021.
  13. 13.Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. Sharpness-aware minimization for efficiently improving generalization. In International Conference on Learning Representations, 2021.
  14. 14.Ghadimi, S. and Lan, G. Stochastic first-and zeroth-order methods for nonconvex stochastic programming. SIAM Journal on Optimization, 23(4):2341–2368, 2013.
  15. 15.Hashemi, A., Acharya, A., Das, R., Vikalo, H., Sanghavi, S., and Dhillon, I. On the benefits of multiple gossip steps in communication-constrained decentralized federated learning. IEEE Transactions on Parallel and Distributed Systems, TPDS, pp. 2727–2739, 2022.
  16. 16.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  17. 17.Hsu, T.-M. H., Qi, H., and Brown, M. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019.
  18. 18.Huang, Y., Chu, L., Zhou, Z., Wang, L., Liu, J., Pei, J., and Zhang, Y. Personalized cross-silo federated learning on non-iid data. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pp. 7865–7873, 2021.
  19. 19.Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., et al. Advances and open problems in federated learning. Foundations and Trends in Machine Learning, pp. 1–210, 2021.
  20. 20.Kang, J., Ye, D., Nie, J., Xiao, J., Deng, X., Wang, S., Xiong, Z., Yu, R., and Niyato, D. Blockchain-based federated learning for industrial metaverses: Incentive scheme with optimal aoi. In 2022 IEEE International Conference on Blockchain (Blockchain), pp. 71–78. IEEE, 2022.
  21. 21.Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A. T. Scaffold: Stochastic controlled averaging for federated learning. In International Conference on Machine Learning, pp. 5132–5143. PMLR, 2020.
  22. 22.Koloskova, A., Loizou, N., Boreiri, S., Jaggi, M., and Stich, S. A unified theory of decentralized sgd with changing topology and local updates. In International Conference on Machine Learning, pp. 5381–5393. PMLR, 2020.
  23. 23.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  24. 24.Kwon, J., Kim, J., Park, H., and Choi, I. K. Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks. In International Conference on Machine Learning, pp. 5905–5914. PMLR, 2021.
  25. 25.Lalitha, A., Shekhar, S., Javidi, T., and Koushanfar, F. Fully decentralized federated learning. In Third workshop on Bayesian Deep Learning (NeurIPS), 2018.
  26. 26.Lalitha, A., Kilinc, O. C., Javidi, T., and Koushanfar, F. Peer-to-peer federated learning on graphs. arXiv preprint arXiv:1901.11173, 2019.
  27. 27.Li, B., Cen, S., Chen, Y., and Chi, Y. Communication-efficient distributed optimization in networks with gradient tracking and variance reduction. Journal of Machine Learning Research, JMLR, pp. 180:1–180:51, 2020a.
  28. 28.Li, H., Xu, Z., Taylor, G., Studer, C., and Goldstein, T. Visualizing the loss landscape of neural nets. Advances in neural information processing systems, 31, 2018.
  29. 29.Li, J., Shao, Y., Wei, K., Ding, M., Ma, C., Shi, L., Han, Z., and Poor, H. V. Blockchain assisted decentralized federated learning (blade-fl): Performance analysis and resource allocation. IEEE Transactions on Parallel and Distributed Systems, 33(10):2401–2415, 2022. doi: 10.1109/TPDS.2021.3138848.
  30. 30.Li, T., Sahu, A. K., Talwalkar, A., and Smith, V. Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, pp. 50–60, 2020b.
  31. 31.Li, T., Sahu, A. K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V. Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems, pp. 429–450, 2020c.
  32. 32.Li, T., Hu, S., Beirami, A., and Smith, V. Ditto: Fair and robust federated learning through personalization. In International Conference on Machine Learning, pp. 6357–6368. PMLR, 2021.
  33. 33.Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J. Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent. In Advances in Neural Information Processing Systems, pp. 5330–5340, 2017.
  34. 34.Liu, Y., Mai, S., Chen, X., Hsieh, C.-J., and You, Y. Towards efficient and scalable sharpness-aware minimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12360–12370, 2022.
  35. 35.Mcmahan, H. B., Moore, E., Ramage, D., Hampson, S., and Arcas, B. A. Y. Communication-efficient learning of deep networks from decentralized data. pp. 1273–1282, 2017.
  36. 36.Mi, P., Shen, L., Ren, T., Zhou, Y., Sun, X., Ji, R., and Tao, D. Make sharpness-aware minimization stronger: A sparsified perturbation approach. arXiv preprint arXiv:2210.05177, 2022.
  37. 37.Nedic, A. and Ozdaglar, A. Distributed subgradient methods for multi-agent optimization. IEEE Transactions on Automatic Control, 54(1):48–61, 2009.
  38. 38.Nguyen, T., Dakka, M., Diakiw, S., VerMilyea, M., Perugini, M., Hall, J., and Perugini, D. A novel decentralized federated learning approach to train on globally distributed, poor quality, and protected private medical data. Scientific Reports, 12(1):8888, 2022.
  39. 39.Qu, Z., Li, X., Duan, R., Liu, Y., Tang, B., and Lu, Z. Generalized federated learning via sharpness aware minimization. In International Conference on Machine Learning, ICML, pp. 18250–18280, 2022.
  40. 40.Reddi, S. J., Charles, Z., Zaheer, M., Garrett, Z., Rush, K., Konecnˇy, J., Kumar, S., and McMahan, H. B. Adaptive federated optimization. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=LkFG3lB13U5.
  41. 41.Roy, A. G., Siddiqui, S., Polsterl, S., Navab, N., and Wachinger, C. Braintorrent: A peer-to-peer environment for decentralized federated learning. arXiv preprint arXiv:1905.06731, 2019.
  42. 42.Sahu, A. K., Li, T., Sanjabi, M., Zaheer, M., Talwalkar, A., and Smith, V. On the convergence of federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127, pp. 3, 2018.
  43. 43.Shi, Y., Liu, Y., Wei, K., Shen, L., Wang, X., and Tao, D. Make landscape flatter in differentially private federated learning. arXiv preprint arXiv:2303.11242, 2023.
  44. 44.Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  45. 45.Sun, T., Li, D., and Wang, B. Decentralized federated averaging. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  46. 46.Sun, Y., Shen, L., Huang, T., Ding, L., and Tao, D. Fedspeed: Larger local interval, less communication round, and higher generalization accuracy. arXiv preprint arXiv:2302.10429, 2023.
  47. 47.Wang, H., Marella, S., and Anderson, J. Fedadmm: A federated primal-dual algorithm allowing partial participation. arXiv preprint arXiv:2203.15104, 2022a.
  48. 48.Wang, J., Charles, Z., Xu, Z., Joshi, G., McMahan, H. B., Al-Shedivat, M., Andrew, G., Avestimehr, S., Daly, K., Data, D., et al. A field guide to federated optimization. arXiv preprint arXiv:2107.06917, 2021.
  49. 49.Wang, L., Xu, Y., Xu, H., Chen, M., and Huang, L. Accelerating decentralized federated learning in heterogeneous edge computing. IEEE Transactions on Mobile Computing, 2022b.
  50. 50.Wang, Y., Su, Z., Zhang, N., and Benslimane, A. Learning in the air: Secure federated learning for uav-assisted crowdsensing. IEEE Transactions on network science and engineering, 8(2):1055–1069, 2020.
  51. 51.Warnat-Herresthal, S., Schultze, H., Shastry, K. L., Manamohan, S., Mukherjee, S., Garg, V., Sarveswara, R., Handler, K., Pickkers, P., Aziz, N. A., et al. Swarm learning for decentralized and confidential clinical machine learning. Nature, pp. 265–270, 2021.
  52. 52.Xiao, H., Rasul, K., and Vollgraf, R. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017.
  53. 53.Yang, H., Fang, M., and Liu, J. Achieving linear speedup with partial worker participation in non-IID federated learning. In International Conference on Learning Representations, 2021.
  54. 54.Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W. Pyhessian: Neural networks through the lens of the hessian. In 2020 IEEE international conference on big data (Big data), pp. 581–590. IEEE, 2020.
  55. 55.Ye, H. and Zhang, T. Deepca: Decentralized exact pca with linear convergence rate. J. Mach. Learn. Res., 22(238):1–27, 2021.
  56. 56.Ye, H., Zhou, Z., Luo, L., and Zhang, T. Decentralized accelerated proximal gradient descent. Advances in Neural Information Processing Systems, 33:18308–18317, 2020.
  57. 57.Yu, H., Yang, S., and Zhu, S. Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning. In Proceedings of the AAAI Conference on Artificial Intelligence, pp. 5693–5700, 2019.
  58. 58.Yu, Z., Hu, J., Min, G., Xu, H., and Mills, J. Proactive content caching for internet-of-vehicles based on peer-to-peer federated learning. In 2020 IEEE 26th International Conference on Parallel and Distributed Systems (ICPADS), pp. 601–608. IEEE, 2020.
  59. 59.Yuan, Y., Chen, R., Sun, C., Wang, M., Hua, F., Yi, X., Yang, T., and Liu, J. Defed: A principled decentralized and privacy-preserving federated learning algorithm. arXiv preprint arXiv:2107.07171, 2021.
  60. 60.Zhang, M., Sapra, K., Fidler, S., Yeung, S., and Alvarez, J. M. Personalized federated learning with first order model optimization. arXiv preprint arXiv:2012.08565, 2020.
  61. 61.Zhang, X., Fang, M., Liu, Z., Yang, H., Liu, J., and Zhu, Z. Net-fleet: Achieving linear convergence speedup for fully decentralized federated learning with heterogeneous data. arXiv preprint arXiv:2208.08490, 2022.
  62. 62.Zhao, Y., Zhang, H., and Hu, X. Penalizing gradient norm for efficiently improving generalization in deep learning. In International Conference on Machine Learning, ICML, pp. 26982–26992. PMLR, 2022.
  63. 63.Zhong, Q., Ding, L., Shen, L., Mi, P., Liu, J., Du, B., and Tao, D. Improving sharpness-aware minimization with fisher mask for better generalization on language models. arXiv preprint arXiv:2210.05497, 2022.
  64. 64.Zhu, T., He, F., Zhang, L., Niu, Z., Song, M., and Tao, D. Topology-aware generalization of decentralized sgd. In International Conference on Machine Learning, ICML, pp. 27479–27503. PMLR, 2022.

Citation

MLA
Shi, Y., et al. “Improving the Model Consistency of Decentralized Federated Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 31269–91, https://proceedings.mlr.press/v202/shi23d.html.
APA
Shi, Y., Shen, L., Wei, K., Sun, Y., Yuan, B., Wang, X., & Tao, D. (2023). Improving the Model Consistency of Decentralized Federated Learning. International Conference on Machine Learning, 202, 31269–31291. https://proceedings.mlr.press/v202/shi23d.html
Chicago
Shi, Y., L. Shen, K. Wei, et al. 2023. “Improving the Model Consistency of Decentralized Federated Learning”. International Conference on Machine Learning 202: 31269–91. https://proceedings.mlr.press/v202/shi23d.html.
Harvard
Shi, Y. et al. (2023) “Improving the Model Consistency of Decentralized Federated Learning”, International Conference on Machine Learning. PMLR, pp. 31269–31291. Available at: https://proceedings.mlr.press/v202/shi23d.html.
Vancouver
1. Shi Y, Shen L, Wei K, Sun Y, Yuan B, Wang X, Tao D (2023) Improving the Model Consistency of Decentralized Federated Learning. In: International Conference on Machine Learning. PMLR, pp 31269–31291

BibTeX

@InProceedings{pmlr-v202-shi23d,
  title = 	 {Improving the Model Consistency of Decentralized Federated Learning},
  author =       {Shi, Yifan and Shen, Li and Wei, Kang and Sun, Yan and Yuan, Bo and Wang, Xueqian and Tao, Dacheng},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {31269--31291},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/shi23d/shi23d.pdf},
  url = 	 {https://proceedings.mlr.press/v202/shi23d.html},
  abstract = 	 {To mitigate the privacy leakages and communication burdens of Federated Learning (FL), decentralized FL (DFL) discards the central server and each client only communicates with its neighbors in a decentralized communication network. However, existing DFL suffers from high inconsistency among local clients, which results in severe distribution shift and inferior performance compared with centralized FL (CFL), especially on heterogeneous data or sparse communication topologies. To alleviate this issue, we propose two DFL algorithms named DFedSAM and DFedSAM-MGS to improve the performance of DFL. Specifically, DFedSAM leverages gradient perturbation to generate local flat models via Sharpness Aware Minimization (SAM), which searches for models with uniformly low loss values. DFedSAM-MGS further boosts DFedSAM by adopting Multiple Gossip Steps (MGS) for better model consistency, which accelerates the aggregation of local flat models and better balances communication complexity and generalization. Theoretically, we present improved convergence rates $\small \mathcal{O}\big(\frac{1}{\sqrt{KT}}+\frac{1}{T}+\frac{1}{K^{1/2}T^{3/2}(1-\lambda)^2}\big)$ and $\small \mathcal{O}\big(\frac{1}{\sqrt{KT}}+\frac{1}{T}+\frac{\lambda^Q+1}{K^{1/2}T^{3/2}(1-\lambda^Q)^2}\big)$ in non-convex setting for DFedSAM and DFedSAM-MGS, respectively, where $1-\lambda$ is the spectral gap of gossip matrix and $Q$ is the number of MGS. Empirically, our methods can achieve competitive performance compared with CFL methods and outperform existing DFL methods.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/