Accelerated Federated Learning with Decoupled Adaptive Optimization

Jiayin JinJiaxiang RenYang ZhouLingjuan LyuJi LiuDejing Dou

article2022ICML64 citations

Proposes a principled ordinary differential equation decomposition framework and a decoupled adaptive optimization algorithm, FedDA, that accelerates federated learning convergence by accurately distributing centralized momentum updates across local client iterations.

Listen

Federated learning enables multiple edge devices, such as mobile phones, to collaboratively train a shared machine learning model without centralizing private user data. However, deploying this technology in real-world environments faces significant bottlenecks: excessive communication overhead between devices and the central server, and data heterogeneity across devices, which causes local models to drift away from the global objective. While centralized machine learning relies on adaptive optimization methods—such as stochastic gradient descent with momentum, Adam, and AdaGrad—to accelerate training, directly adapting these methods to federated environments often causes severe gradient deviations, training instability, and inconsistent convergence.

The article establishes a rigorous mathematical foundation for federated adaptive optimization and introduces FEDDA, a momentum-decoupling adaptive optimization framework designed to accelerate federated training while preserving model stability and accuracy.

The researchers developed their method by framing centralized optimization algorithms as continuous ordinary differential equation systems and analyzing how to decompose these systems across decentralized clients. Building on this framework, FEDDA decouples the calculation of global momentum from local parameter updates. This allows local devices to track momentum linearly without distorting local training steps. The aggregated momentum is then applied at the server level. To address late-stage training instability and convergence inconsistency, the approach shifts to single-iteration, full-batch client updates near the end of training to closely match centralized optimization behavior. The method was evaluated against nine state-of-the-art baselines across standard image classification (CIFAR-100, EMNIST) and natural language processing (Stack Overflow) benchmarks.

The evaluation produced several key findings. First, FEDDA significantly accelerated training convergence, achieving average convergence speed improvements of 34.3% on CIFAR-100, 22.6% on EMNIST, and 75.4% on Stack Overflow compared to baseline methods. Second, the decoupled global momentum reduced the theoretical momentum error growth from an exponential rate seen in standard local momentum approaches to a much slower algebraic rate. Third, FEDDA delivered superior final model accuracy across all tested optimizers, reaching over 51.8% accuracy on CIFAR-100 and up to 86.8% on EMNIST, consistently outperforming standard federated learning algorithms and existing adaptive frameworks. Finally, tests on Stack Overflow showed an average accuracy improvement of 22.3% when utilizing momentum-based stochastic gradient descent over competing methods.

These results demonstrate that federated learning systems do not need to sacrifice training speed or model quality to maintain communication efficiency and data privacy. By using decoupled momentum tracking and terminal stabilization, organizations can reduce the total training rounds required to reach target performance, directly lowering communication bandwidth demands, device energy consumption, and infrastructure costs.

Engineering and deployment teams implementing federated learning systems should consider adopting decoupled momentum mechanisms and late-stage full-batch stabilization to accelerate model updates. When configuring deployments, practitioners must carefully tune hyperparameters, particularly client and server learning rates and the number of local iterations per round, as excessively high local iterations can cause over-fitting on simple local datasets and degrade global model performance.

While the theoretical analysis and empirical results across standard benchmarks provide strong confidence in the method's effectiveness, the experimental findings are based on simulated federated benchmarks under controlled conditions. Decision-makers should validate performance through targeted pilot studies on their specific edge hardware configurations, network environments, and production data distributions before executing full-scale rollouts.

Jin et al (2022).pdf

No sufficiently relevant recommendations were found.

Cover for Accelerated Federated Learning with Decoupled Adaptive Optimization

Abstract

The federated learning (FL) framework enables edge clients to collaboratively learn a shared inference model while keeping privacy of training data on clients. Recently, many heuristics efforts have been made to generalize centralized adaptive optimization methods, such as SGDM, Adam, AdaGrad, etc., to federated settings for improving convergence and accuracy. However, there is still a paucity of theoretical principles on where to and how to design and utilize adaptive optimization methods in federated settings. This work aims to develop novel adaptive optimization methods for FL from the perspective of dynamics of ordinary differential equations (ODEs). First, an analytic framework is established to build a connection between federated optimization methods and decompositions of ODEs of corresponding centralized optimizers. Second, based on this analytic framework, a momentum decoupling adaptive optimization method, FedDA, is developed to fully utilize the global momentum on each local iteration and accelerate the training convergence. Last but not least, full batch gradients are utilized to mimic centralized optimization in the end of the training process to ensure the convergence and overcome the possible inconsistency caused by adaptive optimization methods.

Table of Contents

  • 1. Introduction
  • 2. Preliminaries and Notations
  • 3. Decomposition of ODEs and FL
  • 4. Momentum Decoupling Adaptive Optimization: FedDA+SGDM
  • 5. Experiments
  • 6. Conclusions
  • References
  • A. Appendix
  • A.1. Related Work
  • A.2. Momentum Decoupling Adaptive Optimization: FedDA+Adam and FedDA+AdaGrad
  • A.3. Advantage of Global Momentum in FedDA
  • A.4. Additional Experiments
  • A.5. Experimental Details

Knowls

  1. Knowl 1 — ODE decomposition links centralized optimization to federated averaging

    model/method

    Let MM clients have local objectives Fi(w)F_i(w) and sample counts NiN_i, with total sample count N=∑i=1MNiN=\sum_{i=1}^M N_i and weights pi=Ni/Np_i=N_i/N. The global objective is F(w)=∑i=1MpiFi(w)F(w)=\sum_{i=1}^M p_iF_i(w). The paper uses the gradient-flow ODE for centralized gradient descent as a design framework for federated optimization:

    dwdτ=−∇F(w).\frac{dw}{d\tau}=-\nabla F(w).

    A precise decomposition evolves client parameters wiw_i according to dwi/dτ=−∇Fi(w)dw_i/d\tau=-\nabla F_i(w), with the shared argument w=∑ipiwiw=\sum_i p_iw_i. The weighted aggregate then follows the centralized gradient flow, but evaluating each local gradient at the shared parameter requires communication at every step. To permit multiple local steps between aggregations, the local dynamics instead use each client’s own parameter, dwi/dτ=−∇Fi(wi)dw_i/d\tau=-\nabla F_i(w_i). Periodically averaging these independently evolved parameters is the ODE-based interpretation of FedAvg: it approximates centralized training, with the approximation becoming less reliable as local trajectories drift farther from the shared trajectory. The same design principle—find a suitable decomposition of a centralized optimizer’s ODE—motivates the paper’s adaptive federated method.

  2. Knowl 2 — FedDA decouples local parameter updates from global momentum

    model/method

    FedDA+SGDM maintains a global momentum by aggregating client momentum states, while keeping that momentum out of the clients’ parameter updates. For client ii, let wi,tw_{i,t} be its local parameter, mi,tm_{i,t} its momentum state, and gi(wi,t)g_i(w_{i,t}) a mini-batch gradient. At round rr, each participating client starts from the global parameter wrw_r and global momentum mrm_r, and initializes an accumulator Pi,0=0P_{i,0}=0. For TT local steps it computes

    wi,t+1=wi,t−ηgi(wi,t),mi,t+1=βmi,t+(1−β)gi(wi,t),Pi,t+1=Pi,t+mi,t+1,w_{i,t+1}=w_{i,t}-\eta g_i(w_{i,t}),\qquad m_{i,t+1}=\beta m_{i,t}+(1-\beta)g_i(w_{i,t}),\qquad P_{i,t+1}=P_{i,t}+m_{i,t+1},

    where η\eta is the client learning rate and β\beta is the momentum coefficient. Thus local parameter motion depends on the local gradient, not on the momentum state, but every local gradient contributes to momentum accumulated for the global update. After the local steps, the server aggregates the clients’ accumulated momenta Pi,TP_{i,T} and final states mi,Tm_{i,T} using sample-weighted aggregation (normalized over participating clients when only a subset participates), obtaining PrP_r and mr+1m_{r+1}. It updates

    wr+1=wr−αηPr,w_{r+1}=w_r-\alpha\eta P_r,

    where α\alpha scales the server update. Summing the successive momentum states in PrP_r makes the global model use momentum contributions from each local step, rather than applying momentum only once to an aggregate of multiple-step client updates.

  3. Knowl 3 — Momentum deviation bounds distinguish FedDA from local momentum

    theoretical result

    Under the assumptions that every client gradient is Lipschitz continuous with a common constant LgL_g and that client gradient norms are bounded, the paper compares momentum trajectories with centralized SGDM trajectories initialized consistently. In the ordinary local-momentum decomposition, a client’s momentum deviation from centralized momentum is characterized by an exponential-rate bound O(eλt)O(e^{\lambda t}), where tt is the time parameter in the ODE comparison and λ>0\lambda>0 is determined by the comparison dynamics. In FedDA, the aggregated global momentum’s deviation from centralized momentum is bounded at an algebraic rate O(t2)O(t^2). The result concerns the stated smoothness and bounded-gradient setting; it does not assert that FedDA’s local parameters follow centralized trajectories exactly.

  4. Knowl 4 — One full-batch local step is used to stabilize late training

    model/method

    The paper proposes switching to full-batch client gradients and a single local iteration per round near the end of training, when the global parameter is close to a local minimum. In the paper’s ODE-decomposition framework, one local step using full-batch gradients makes FedDA agree with centralized training for that round. The authors use this late-training mode to reduce gradient error, address convergence inconsistency associated with adaptive federated updates, and stabilize convergence. This is a proposed training strategy and conditional agreement claim, rather than a general convergence theorem for arbitrary multi-step stochastic FedDA training.

  5. Knowl 5 — FedDA adapts the decoupled momentum construction to Adam, AdaGrad, and FedProx

    model/method

    For FedDA+Adam and FedDA+AdaGrad, clients perform local gradient updates while updating and accumulating first-moment momentum states as in FedDA+SGDM. The server aggregates the accumulated and final momentum quantities and forms an effective gradient G=(P−β1mr)/(1−β1)G=(P-\beta_1m_r)/(1-\beta_1), where PP is the aggregate of the within-round momentum accumulators, mrm_r is the previous global first-moment state, and β1\beta_1 is the first-moment coefficient. The server feeds GG to its adaptive optimizer: Adam updates first- and second-moment states using coefficients β1\beta_1 and β2\beta_2 and applies bias correction; AdaGrad accumulates G2G^2 in its second-moment state. The second-moment state is kept at the server because the square of an aggregate gradient generally cannot be precisely decomposed into a weighted aggregate of squared client gradients. The paper also describes a FedProx-compatible version: clients add the proximal term μ(wi,t−wr)\mu(w_{i,t}-w_r) to their local gradient before updating their parameters and momentum, while retaining the same global aggregation rule; μ\mu is the proximal regularization coefficient.

  6. Knowl 6 — Evaluation covers heterogeneous image and text federated tasks

    experimental setup

    The experiments compare FedDA using SGDM, Adam, and AdaGrad with federated-learning and federated-optimization baselines on three tasks. CIFAR-100 uses 50,000 training and 10,000 test examples distributed across 500 training and 100 test clients; clients are formed using a partitioning strategy intended to reflect label heterogeneity, and the model is ResNet-18 with group normalization replacing batch normalization. EMNIST-62 has 671,585 training and 77,483 test examples across 3,400 training and 3,400 test clients; its character-recognition model is a CNN. Stack Overflow has 135,818,730 training and 16,586,035 test examples, with 342,477 training clients and 204,088 test clients; the task is tag prediction using bag-of-words logistic regression with the 10,000 most frequent words and 500 most frequent tags in a one-versus-rest setup. Training runs for 4,000 rounds on CIFAR-100 and 1,500 rounds on EMNIST and Stack Overflow. Clients are sampled uniformly at random without replacement within a round and with replacement across rounds.

  7. Knowl 7 — Reported convergence speedups vary across the three benchmarks

    empirical result

    Using the rounds at which accuracy curves reach their convergence points, the authors report average convergence improvements for FedDA of 34.3% on CIFAR-100 and 22.6% on EMNIST relative to the compared federated methods. On Stack Overflow, they report a 75.4% convergence improvement. The experiments run for 4,000 rounds on CIFAR-100 and 1,500 rounds on each of EMNIST and Stack Overflow. The paper also reports that, in most tested optimizer settings, FedDA converges faster than the alternatives; on EMNIST, its loss is reported to converge within fewer than 200 rounds in the highlighted experiments. These are reported comparisons from the plotted training curves, not guarantees for other tasks or settings.

  8. Knowl 8 — Final classification accuracy shows strong but nonuniform FedDA results

    data/table

    The table gives final test accuracy for each method with SGDM, Adam, and AdaGrad on CIFAR-100, EMNIST, and Stack Overflow. It shows that FedDA leads all listed methods for each optimizer on CIFAR-100 and EMNIST, but its advantage is not uniform on Stack Overflow: it trails CLIMB with SGDM, ties FedOpt with Adam, and trails FedOpt with AdaGrad.

    CIFAR-100 EMNIST Stack Overflow
    Method SGDM Adam AdaGrad SGDM Adam AdaGrad SGDM Adam AdaGrad
    FedLocal 0.384 0.009 0.113 0.834 0.055 0.806 0.152 0.576 0.229
    SCAFFOLD 0.010 0.010 0.010 0.794 0.794 0.794 0.250 0.250 0.250
    FedLin 0.440 0.440 0.440 0.805 0.805 0.805 0.224 0.224 0.224
    FedAvg 0.324 0.324 0.324 0.850 0.850 0.850 0.252 0.252 0.252
    MFL 0.293 0.346 0.135 0.848 0.055 0.047 0.134 0.101 0.215
    CLIMB 0.010 0.010 0.010 0.843 0.843 0.843 0.302 0.302 0.302
    STEM 0.014 0.014 0.014 0.051 0.051 0.051 0.196 0.196 0.196
    MimeLite 0.427 0.009 0.009 0.835 0.851 0.821 0.271 0.211 0.078
    FedOpt 0.425 0.443 0.301 0.838 0.847 0.840 0.225 0.642 0.691
    FedDA 0.518 0.510 0.488 0.860 0.853 0.868 0.273 0.642 0.674
  9. Knowl 9 — Increasing local iterations can help before causing client drift

    empirical result

    In FedDA experiments varying the number of local iterations over 11, 55, 1010, and 2020 on CIFAR-100, EMNIST, and Stack Overflow, final accuracy generally rises at first and then levels off or falls as local work increases. This indicates a trade-off: too few local steps can slow local and global progress, whereas too many can move client models toward client-specific minima that do not match the global objective. On EMNIST, the paper specifically reports that SGDM and AdaGrad accuracy becomes close to zero when the local iteration count exceeds 1010, attributing this to clients converging toward local minima distant from the global minimum.

  10. Knowl 10 — FedDA also improves EMNIST autoencoder reconstruction error

    data/table

    As an additional EMNIST autoencoder evaluation, the paper reports final mean squared error (MSE) for six methods under three optimizer choices. Lower MSE is better. FedDA has the lowest reported error for SGDM, Adam, and AdaGrad, including an AdaGrad MSE of 0.01320.0132 compared with 0.01450.0145 for FedOpt.

    Method SGDM Adam AdaGrad
    FedLocal 0.0169 0.0289 0.0168
    FedAvg 0.0171 0.0171 0.0171
    MFL 0.0168 0.0290 0.0291
    MimeLite 0.0183 0.0307 0.0287
    FedOpt 0.0175 0.0173 0.0145
    FedDA 0.0167 0.0166 0.0132

Coverage note — The detailed client/server learning-rate sweeps and compute-environment specifications are omitted as secondary sensitivity and reproducibility details; the main benchmark setup and local-iteration sensitivity are retained.

References

  1. 1.Acar, D. A. E., Zhao, Y., Matas, R., Mattina, M., Whatmough, P., and Saligrama, V. Federated learning based on dynamic regularization. In Int. Conf. on Learning Representations (ICLR), pp. 1–36, 2021.
  2. 2.Afonin, A. and Karimireddy, S. P. Towards model agnostic federated learning using knowledge distillation. arXiv preprint arXiv:2110.15210, 2021.
  3. 3.Anonymous. Acceleration of federated learning with alleviated forgetting in local training. In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–19, 2022a. under review.
  4. 4.Anonymous. An agnostic approach to federated learning with class imbalance. In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–12, 2022b. under review.
  5. 5.Anonymous. AQUILA: Communication efficient federated learning with adaptive quantization of lazily-aggregated gradients. In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–23, 2022c. under review.
  6. 6.Anonymous. Hybrid local SGD for federated learning with heterogeneous communications. In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–41, 2022d. under review.
  7. 7.Anonymous. Improving federated learning face recognition via privacy-agnostic clusters. In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–21, 2022e. under review.
  8. 8.Anonymous. Privacy-preserving task-agnostic vision transformer for image processing. In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–28, 2022f. under review.
  9. 9.Anonymous. Recycling model updates in federated learning: Are gradient subspaces low-rank? In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–71, 2022g. under review.
  10. 10.Anonymous. Unsupervised federated learning is possible. In Submitted to Int. Conf. on Learning Representations (ICLR), pp. 1–22, 2022h. under review.
  11. 11.Balunovic, M., Dimitrov, D. I., Staab, R., and Vechev, M. ´ Bayesian framework for gradient leakage. arXiv preprint arXiv:2111.04706, 2021.
  12. 12.Bao, X., Liu, L., Xiao, N., Zhou, Y., and Zhang, Q. Policy-driven autonomic configuration management for nosql. In Proceedings of the 2015 IEEE International Conference on Cloud Computing (CLOUD’15), pp. 245–252, New York, NY, June 27-July 2 2015.
  13. 13.Caldas, S., Konecnˇ y, J., McMahan, H. B., and Talwalkar, A. ´ Expanding the reach of federated learning by reducing client resource requirements. CoRR, abs/1812.07210, 2018a.
  14. 14.Caldas, S., Wu, P., Li, T., Konecnˇ y, J., McMahan, H. B., ´ Smith, V., and Talwalkar, A. LEAF: A benchmark for federated settings. CoRR, abs/1812.01097, 2018b.
  15. 15.Chen, H.-Y. and Chao, W.-L. On bridging generic and personalized federated learning. arXiv preprint arXiv:2107.00778, 2021.
  16. 16.Chen, M., Poor, H. V., Saad, W., and Cui, S. Convergence time minimization of federated learning over wireless networks. In IEEE Int. Conf. on Communications (ICC), pp. 1–6, 2020.
  17. 17.Dai, Z., Low, B. K. H., and Jaillet, P. Federated bayesian optimization via thompson sampling. In Annual Conf. on Neural Information Processing Systems (NeurIPS), pp. 1–13, 2020.
  18. 18.Duchi, J., Hazan, E., and Singer, Y. Adaptive subgradient methods for online learning and stochastic optimization. Journal of machine learning research, 12(7), 2011a.
  19. 19.Duchi, J. C., Hazan, E., and Singer, Y. Adaptive subgradient methods for online learning and stochastic optimization. J. Mach. Learn. Res., 12:2121–2159, 2011b.
  20. 20.Fowl, L., Geiping, J., Czaja, W., Goldblum, M., and Goldstein, T. Robbing the fed: Directly obtaining private data in federated learning with modified models. arXiv preprint arXiv:2110.13057, 2021.
  21. 21.Gao, H., Xu, A., and Huang, H. On the convergence of communication-efficient local sgd for federated learning. In AAAI Conf. on Artificial Intelligence, volume 35, pp. 7510–7518, 2021.
  22. 22.Goswami, S., Pokhrel, A., Lee, K., Liu, L., Zhang, Q., and Zhou, Y. Graphmap: Scalable iterative graph processing using nosql. The Journal of Supercomputing (TJSC), 76(9):6619–6647, 2020.
  23. 23.Guimu Guo, Da Yan, L. Y. J. K. C. L. Z. J. and Zhou, Y. Maximal directed quasi-clique mining. In Proceedings of the 38th IEEE International Conference on Data Engineering (ICDE’22), Kuala Lumpur, Malaysia, May 9-12 2022.
  24. 24.Hamer, J., Mohri, M., and Suresh, A. T. Fedboost: A communication-efficient algorithm for federated learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, pp. 3973–3983, 2020a.
  25. 25.Hamer, J., Mohri, M., and Suresh, A. T. FedBoost: A communication-efficient algorithm for federated learning. In Int. Conf. on Machine Learning (ICML), volume 119, pp. 3973–3983, 2020b.
  26. 26.He, C., Annavaram, M., and Avestimehr, S. Group knowledge transfer: Federated learning of large cnns at the edge. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  27. 27.Hong, J., Wang, H., Wang, Z., and Zhou, J. Federated robustness propagation: Sharing adversarial robustness in federated learning. arXiv preprint arXiv:2106.10196, 2021.
  28. 28.Hsieh, K., Phanishayee, A., Mutlu, O., and Gibbons, P. B. The non-iid data quagmire of decentralized machine learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 4387–4398. PMLR, 2020.
  29. 29.Hyeon-Woo, N., Ye-Bin, M., and Oh, T.-H. Fedpara: Low-rank hadamard product for communication-efficient federated learning. arXiv preprint arXiv:2108.06098, 2021.
  30. 30.Ingerman, A. and Ostrowski, K. Introducing tensorflow federated. https://medium.com/tensorflow/introducing-tensorflow-federated, 2019.
  31. 31.Jiang, Y., Perng, C.-S., Sailer, A., Silva-Lepe, I., Zhou, Y., and Li, T. Csm: A cloud service marketplace for complex service acquisition. ACM Transactions on Intelligent Systems and Technology (TIST), 8(1):1–25, 2016.
  32. 32.Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., et al. Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977, 2019.
  33. 33.Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K. A., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Eichner, H., Rouayheb, S. E., Evans, D., Gardner, J., Garrett, Z., Gascon, A., Ghazi, B., Gibbons, P. B., Gruteser, M., Harchaoui, Z., He, C., He, L., Huo, Z., Hutchinson, B., Hsu, J., Jaggi, M., Javidi, T., Joshi, G., Khodak, M., Konecnˇ y, J., Korolova, A., Koushanfar, F., Koyejo, S., Lepoint, T., Liu, Y., Mittal, P., Mohri, M., Nock, R., Ozg ¨ ur, A., Pagh, R., Qi, H., Ramage, D., Raskar, R., Raykova, M., Song, D., Song, W., Stich, S. U., Sun, Z., Suresh, A. T., Tramer, F., Vepakomma, P., Wang, J., Xiong, L., Xu, Z., Yang, Q., Yu, F. X., Yu, H., and Zhao, S. Advances and open problems in federated learning. Found. Trends Mach. Learn., 14(1-2):1–210, 2021.
  34. 34.Karimireddy, S. P., He, L., and Jaggi, M. Byzantine-robust learning on heterogeneous datasets via bucketing. arXiv preprint arXiv:2006.09365, 2020a.
  35. 35.Karimireddy, S. P., Jaggi, M., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T. Mime: Mimicking centralized stochastic algorithms in federated learning. arXiv preprint arXiv:2008.03606, 2020b.
  36. 36.Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A. T. Scaffold: Stochastic controlled averaging for federated learning. In Int. Conf. on Machine Learning (ICML), pp. 5132–=–5143, 2020c.
  37. 37.Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T. SCAFFOLD: stochastic controlled averaging for federated learning. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pp. 5132–5143. PMLR, 2020d.
  38. 38.Karimireddy, S. P., Jaggi, M., Kale, S., Mohri, M., Reddi, S. J., Stich, S. U., and Suresh, A. T. Mime: Mimicking centralized stochastic algorithms in federated learning. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, 2021.
  39. 39.Khanduri, P., Sharma, P., Yang, H., Hong, M., Liu, J., Rajawat, K., and Varshney, P. K. Stem: A stochastic two-sided momentum algorithm achieving near-optimal sample and communication complexities for federated learning. In Annual Conf. on Neural Information Processing Systems (NeurIPS), pp. 1–12, 2021.
  40. 40.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  41. 41.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. In Bengio, Y. and LeCun, Y. (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, 2015.
  42. 42.Konecnˇ y, J., McMahan, H. B., Ramage, D., and Richt ´ arik, P. Federated optimization: Distributed machine learning for on-device intelligence. CoRR, abs/1610.02527, 2016a.
  43. 43.Konecnˇ y, J., McMahan, H. B., Yu, F. X., Richt ´ arik, P., Suresh, A. T., and Bacon, D. Federated learning: Strategies for improving communication efficiency. In NIPS Workshop on Private Multi-Party Machine Learning, 2016b.
  44. 44.Krizhevsky, A. Learning multiple layers of features from tiny images. Technical Report, 2009.
  45. 45.Lee, K., Liu, L., Tang, Y., Zhang, Q., and Zhou, Y. Efficient and customizable data partitioning framework for distributed big rdf data processing in the cloud. In Proceedings of the 2013 IEEE International Conference on Cloud Computing (CLOUD’13), pp. 327–334, Santa Clara, CA, June 27-July 2 2013.
  46. 46.Lee, K., Liu, L., Schwan, K., Pu, C., Zhang, Q., Zhou, Y., Yigitoglu, E., and Yuan, P. Scaling iterative graph computations with graphmap. In Proceedings of the 27th IEEE international conference for High Performance Computing, Networking, Storage and Analysis (SC’15), pp. 57:1–57:12, Austin, TX, November 15-20 2015.
  47. 47.Lee, K., Liu, L., Ganti, R. L., Srivatsa, M., Zhang, Q., Zhou, Y., and Wang, Q. Lightwieight indexing and querying services for big spatial data. IEEE Transactions on Services Computing (TSC), 12(3):343–355, 2019.
  48. 48.Leroy, D., Coucke, A., Lavril, T., Gisselbrecht, T., and Dureau, J. Federated learning for keyword spotting. In IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP), pp. 6341–6345, 2019.
  49. 49.Li, W. and McCallum, A. Pachinko allocation: Dag-structured mixture models of topic correlations. In Cohen, W. W. and Moore, A. W. (eds.), Machine Learning, Proceedings of the Twenty-Third International Conference (ICML 2006), Pittsburgh, Pennsylvania, USA, June 25-29, 2006, volume 148 of ACM International Conference Proceeding Series, pp. 577–584. ACM, 2006.
  50. 50.Li, Z., Kovalev, D., Qian, X., and Richtarik, P. Acceleration for compressed gradient descent in distributed and federated optimization. In Int. Conf. on Machine Learning (ICML), volume 119, pp. 5895–5904, 2020.
  51. 51.Liu, J., Huang, J., Zhou, Y., Li, X., Ji, S., Xiong, H., and Dou, D. From distributed machine learning to federated learning: A survey. arXiv preprint arXiv:2104.14362, 2021.
  52. 52.Liu, J., Huang, J., Zhou, Y., Li, X., Ji, S., Xiong, H., and Dou, D. From distributed machine learning to federated learning: A survey. Knowledge and Information Systems (KAIS), 64(4):885–917, 2022.
  53. 53.Liu, W., Chen, L., Chen, Y., and Zhang, W. Accelerating federated learning via momentum gradient descent. IEEE Trans. Parallel Distributed Syst., 31(8):1754–1766, 2020a.
  54. 54.Liu, W., Chen, L., Chen, Y., and Zhang, W. Accelerating federated learning via momentum gradient descent. IEEE Transactions on Parallel and Distributed Systems (TPDS), 31(8):1754–1766, 2020b.
  55. 55.Liu, Y., Kang, Y., Zhang, X., Li, L., Cheng, Y., Chen, T., Hong, M., and Yang, Q. A communication efficient vertical federated learning framework. CoRR, abs/1912.11187, 2019.
  56. 56.McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A. Communication-efficient learning of deep networks from decentralized data. In Singh, A. and Zhu, X. J. (eds.), Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017, 20-22 April 2017, Fort Lauderdale, FL, USA, volume 54 of Proceedings of Machine Learning Research, pp. 1273–1282. PMLR, 2017a.
  57. 57.McMahan, B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A. Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics, pp. 1273–1282, 2017b.
  58. 58.McMahan, H. B. and Streeter, M. J. Adaptive bound optimization for online convex optimization. In Kalai, A. T. and Mohri, M. (eds.), COLT 2010 - The 23rd Conference on Learning Theory, Haifa, Israel, June 27-29, 2010, pp. 244–256. Omnipress, 2010.
  59. 59.McMahan, H. B., Moore, E., Ramage, D., and y Arcas, B. A. Federated learning of deep networks using model averaging. CoRR, abs/1602.05629, 2016.
  60. 60.Mills, J., Hu, J., and Min, G. Communication-efficient federated learning for wireless edge intelligence in iot. IEEE Internet of Things Journal, 7(7):5986–5994, 2019.
  61. 61.Mills, J., Hu, J., Min, G., Jin, R., Zheng, S., and Wang, J. Accelerating federated learning with a global biased optimiser. arXiv preprint arXiv:2108.09134, 2021.
  62. 62.Mitra, A., Jaafar, R., Pappas, G., and Hassani, H. Linear convergence in federated learning: Tackling client heterogeneity and sparse gradients. In The 35th Conference on Neural Information Processing Systems, (NeurIPS’21), Online, December 6-14 2021a.
  63. 63.Mitra, A., Jaafar, R., Pappas, G. J., and Hassani, H. Linear convergence in federated learning: Tackling client heterogeneity and sparse gradients. In Annual Conf. on Neural Information Processing Systems (NeurIPS), pp. 1–14, 2021b.
  64. 64.Oh, J., Kim, S., and Yun, S.-Y. Fedbabu: Towards enhanced representation for federated image classification. arXiv preprint arXiv:2106.06042, 2021.
  65. 65.Ozfatura, E., Ozfatura, K., and Gund ¨ uz, D. Fedadc: Accelerated federated learning with drift control. In IEEE Int. Symposium on Information Theory (ISIT), pp. 467–472, 2021.
  66. 66.Palanisamy, B., Liu, L., Lee, K., Meng, S., Tang, Y., and Zhou, Y. Anonymizing continuous queries with delay-tolerant mix-zones over road networks. Distributed and Parallel Databases (DAPD), 32(1):91–118, 2014.
  67. 67.Palanisamy, B., Liu, L., Zhou, Y., and Wang, Q. Privacy-preserving publishing of multilevel utility-controlled graph datasets. ACM Transactions on Internet Technology (TOIT), 18(2):24:1–24:21, 2018.
  68. 68.Pathak, R. and Wainwright, M. J. Fedsplit: an algorithmic framework for fast federated optimization. In Annual Conf. on Neural Information Processing Systems (NeurIPS), volume 33, pp. 7057–7066, 2020.
  69. 69.Qian, N. On the momentum term in gradient descent learning algorithms. Neural Networks, 12(1):145–151, 1999.
  70. 70.Reddi, S., Zaheer, M., Sachan, D., Kale, S., and Kumar, S. Adaptive methods for nonconvex optimization. In Annual Conf. on Neural Information Processing Systems (NeurIPS), pp. 1–17, 2018.
  71. 71.Reddi, S. J., Charles, Z., Zaheer, M., Garrett, Z., Rush, K., Konecnˇ y, J., Kumar, S., and McMahan, H. B. Adaptive federated optimization. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021, 2021a.
  72. 72.Reddi, S. J., Charles, Z., Zaheer, M., Garrett, Z., Rush, K., Konecnˇ y, J., Kumar, S., and McMahan, H. B. Adaptive federated optimization. In Int. Conf. on Learning Representations (ICLR), 2021b.
  73. 73.Rothchild, D., Panda, A., Ullah, E., Ivkin, N., Stoica, I., Braverman, V., Gonzalez, J., and Arora, R. Fetchsgd: Communication-efficient federated learning with sketching. In Int. Conf. on Machine Learning (ICML), pp. 8253–8265, 2020.
  74. 74.Rumelhart, D. E., Hinton, G. E., and Williams, R. J. Learning internal representations by error propagation. In Rumelhart, D. E. and McClelland, J. L. (eds.), Parallel Distributed Processing. MIT Press, 1986.
  75. 75.Sharma, A., Chen, W., Zhao, J., Qiu, Q., Chaterji, S., and Bagchi, S. Tesseract: Gradient flip score to secure federated learning against model poisoning attacks. arXiv preprint arXiv:2110.10108, 2021.
  76. 76.Su, Z., Liu, L., Li, M., Fan, X., and Zhou, Y. Servicetrust: Trust management in service provision networks. In Proceedings of the 10th IEEE International Conference on Services Computing (SCC’13), pp. 272–279, Santa Clara, CA, June 27-July 2 2013.
  77. 77.Su, Z., Liu, L., Li, M., Fan, X., and Zhou, Y. Reliable and resilient trust management in distributed service provision networks. ACM Transactions on the Web (TWEB), 9(3):1–37, 2015.
  78. 78.Sutskever, I., Martens, J., Dahl, G. E., and Hinton, G. E. On the importance of initialization and momentum in deep learning. In Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013, volume 28 of JMLR Workshop and Conference Proceedings, pp. 1139–1147, 2013.
  79. 79.TensorFlow. Tensorflow federated stack overflow dataset. https://www.tensorflow.org/federated/api_docs/python/tff/simulation/datasets/stackoverflow/load_data, 2019.
  80. 80.Triastcyn, A., Reisser, M., and Louizos, C. Dp-rec: Private & communication-efficient federated learning. arXiv preprint arXiv:2111.05454, 2021.
  81. 81.Wang, H.-P., Stich, S. U., He, Y., and Fritz, M. Progfed: Effective, communication, and computation efficient federated learning by progressive training. arXiv preprint arXiv:2110.05323, 2021a.
  82. 82.Wang, J., Tantia, V., Ballas, N., and Rabbat, M. Slowmo: Improving communication-efficient distributed sgd with slow momentum. In Int. Conf. on Learning Representations (ICLR), pp. 1–27, 2020.
  83. 83.Wang, J., Xu, Z., Garrett, Z., Charles, Z., Liu, L., and Joshi, G. Local adaptivity in federated learning: Convergence and consistency. In The International Workshop on Federated Learning for User Privacy and Data Confidentiality in Conjunction with ICML 2021, (FL-ICML’21), Online, December 6-14 2021b.
  84. 84.Wang, J., Xu, Z., Garrett, Z., Charles, Z., Liu, L., and Joshi, G. Local adaptivity in federated learning: Convergence and consistency. arXiv preprint arXiv:2106.02305, 2021c.
  85. 85.Woodworth, B. E., Patel, K. K., and Srebro, N. Minibatch vs local SGD for heterogeneous distributed learning. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020.
  86. 86.Wu, H. and Wang, P. Fast-convergent federated learning with adaptive weighting. IEEE Transactions on Cognitive Communications and Networking, 2021.
  87. 87.Wu, W., He, L., Lin, W., and Mao, R. Accelerating federated learning over reliability-agnostic clients in mobile edge computing systems. IEEE Transactions on Parallel and Distributed Systems (TPDS), 32(7):1539–1551, 2020.
  88. 88.Wu, Y. and He, K. Group normalization. Int. J. Comput. Vis., 128(3):742–755, 2020.
  89. 89.Xia, S., Zhu, J., Yang, Y., Zhou, Y., Shi, Y., and Chen, W. Fast convergence algorithm for analog federated learning. In IEEE Int. Conf. on Communications (ICC), pp. 1–6, 2021.
  90. 90.Xie, C., Koyejo, O., Gupta, I., and Lin, H. Local adaalter: Communication-efficient stochastic gradient descent with adaptive learning rates. CoRR, abs/1911.09030, 2019.
  91. 91.Yan, D., Qu, W., Guo, G., Wang, X., and Zhou, Y. Prefixfpm: A parallel framework for general-purpose mining of frequent and closed patterns. The VLDB Journal (VLDBJ), 31(2):253–286, 2022a.
  92. 92.Yan, D., Zhou, Y., Guo, G., and Liu, H. Parallel graph processing. In Sherif Sakr, A. Y. Z. and Taheri, J. (eds.), Encyclopedia of Big Data Technologies. Springer, 2022b.
  93. 93.Yang, H. H-fl: A hierarchical communication-efficient and privacy-protected architecture for federated learning. In Int. Joint Conf. on Artificial Intelligence (IJCAI), pp. 479–485, 2021.
  94. 94.Yapp, A. Z. H., Koh, H. S. N., Lai, Y. T., Kang, J., Li, X., Ng, J. S., Jiang, H., Lim, W. Y. B., Xiong, Z., and Niyato1, D. Communication-efficient and scalable decentralized federated edge learning. In Int. Joint Conf. on Artificial Intelligence (IJCAI), pp. 5032–5035, 2021.
  95. 95.Yuan, H., Morningstar, W., Ning, L., and Singhal, K. What do we mean by generalization in federated learning? arXiv preprint arXiv:2110.14216, 2021a.
  96. 96.Yuan, H., Zaheer, M., and Reddi, S. Federated composite optimization. In Int. Conf. on Machine Learning (ICML), volume 139, pp. 12253–12266, 2021b.
  97. 97.Yun, C., Rajput, S., and Sra, S. Minibatch vs local sgd with shuffling: Tight convergence bounds and beyond. arXiv preprint arXiv:2110.10342, 2021.
  98. 98.Zhang, H., Liu, J., Jia, J., Zhou, Y., Dai, H., and Dou, D. Fedduap: Federated learning with dynamic update and adaptive pruning using shared data on the server. In Proceedings of the 31st International Joint Conference on Artificial Intelligence (IJCAI’22), Messe Wien, Vienna, Austria, July 23-29 2022.
  99. 99.Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S. J., Kumar, S., and Sra, S. Why ADAM beats SGD for attention models. CoRR, abs/1912.03194, 2019.
  100. 100.Zhang, M., Sapra, K., Fidler, S., Yeung, S., and Alvarez, J. M. Personalized federated learning with first order model optimization. In Int. Conf. on Learning Representations (ICLR), 2021.
  101. 101.Zhang, Q., Liu, L., Ren, Y., Lee, K., Tang, Y., Zhao, X., and Zhou, Y. Residency aware inter-vm communication in virtualized cloud: Performance measurement and analysis. In Proceedings of the 2013 IEEE International Conference on Cloud Computing (CLOUD’13), pp. 204–211, Santa Clara, CA, June 27-July 2 2013.
  102. 102.Zhang, Q., Liu, L., Lee, K., Zhou, Y., Singh, A., Mandagere, N., Gopisetty, S., and Alatorre, G. Improving hadoop service provisioning in a geographically distributed cloud. In Proceedings of the 2014 IEEE International Conference on Cloud Computing (CLOUD’14), pp. 432–439, Anchorage, AK, June 27-July 2 2014.
  103. 103.Zhang, Z., Jin, J., Zhang, Z., Zhou, Y., Zhao, X., Ren, J., Liu, J., Wu, L., Jin, R., and Dou, D. Validating the lottery ticket hypothesis with inertial manifold theory. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021 (NeurIPS’21), Virtual.
  104. 104.Zhou, C., Liu, J., Jia, J., Zhou, J., Zhou, Y., Dai, H., and Dou, D. Efficient device scheduling with multi-job federated learning. In Proceedings of the 36th AAAI Conference on Artificial Intelligence (AAAI’22), Vancouver, Canada, February 22-March 1 2022a.
  105. 105.Zhou, C., Liu, J., Jia, J., Zhou, J., Zhou, Y., Dai, H., and Dou, D. Efficient device scheduling with multi-job federated learning. In AAAI Conf. on Artificial Intelligence, pp. 1–14, 2022b.
  106. 106.Zhou, Y. Innovative Mining, Processing, and Application of Big Graphs. PhD thesis, Georgia Institute of Technology, Atlanta, GA, USA, 2017.
  107. 107.Zhou, Y. and Liu, L. Approximate deep network embedding for mining large-scale graphs. In Proceedings of the 2019 IEEE International Conference on Cognitive Machine Intelligence (CogMI’19), pp. 53–60, Los Angeles, CA, December 12-14 2019.
  108. 108.Zhou, Y., Liu, L., Lee, K., Pu, C., and Zhang, Q. Fast iterative graph computation with resource aware graph parallel abstractions. In Proceedings of the 24th ACM Symposium on High-Performance Parallel and Distributed Computing (HPDC’15), pp. 179–190, Portland, OR, June 15-19 2015a.
  109. 109.Zhou, Y., Liu, L., Lee, K., and Zhang, Q. Graphtwist: Fast iterative graph computation with two-tier optimizations. Proceedings of the VLDB Endowment (PVLDB), 8(11):1262–1273, 2015b.
  110. 110.Zhou, Y., Amimeur, A., Jiang, C., Dou, D., Jin, R., and Wang, P. Density-aware local siamese autoencoder network embedding with autoencoder graph clustering. In Proceedings of the 2018 IEEE International Conference on Big Data (BigData’18), pp. 1162–1167, Seattle, WA, December 10-13 2018a.
  111. 111.Zhou, Y., Wu, S., Jiang, C., Zhang, Z., Dou, D., Jin, R., and Wang, P. Density-adaptive local edge representation learning with generative adversarial network multi-label edge classification. In Proceedings of the 18th IEEE International Conference on Data Mining (ICDM’18), pp. 1464–1469, Singapore, November 17-20 2018b.
  112. 112.Zhou, Y., Ye, Q., and Lv, J. Communication-efficient federated learning with compensated overlap-fedavg. IEEE Transactions on Parallel and Distributed Systems (TPDS), 33(1):192–205, 2022c.

Citation

MLA
Jin, J., et al. “Accelerated Federated Learning with Decoupled Adaptive Optimization”. International Conference on Machine Learning, vol. 162, 2022, pp. 10298–322, https://proceedings.mlr.press/v162/jin22e.html.
APA
Jin, J., Ren, J., Zhou, Y., Lyu, L., Liu, J., & Dou, D. (2022). Accelerated Federated Learning with Decoupled Adaptive Optimization. International Conference on Machine Learning, 162, 10298–10322. https://proceedings.mlr.press/v162/jin22e.html
Chicago
Jin, J., J. Ren, Y. Zhou, L. Lyu, J. Liu, and D. Dou. 2022. “Accelerated Federated Learning with Decoupled Adaptive Optimization”. International Conference on Machine Learning 162: 10298–322. https://proceedings.mlr.press/v162/jin22e.html.
Harvard
Jin, J. et al. (2022) “Accelerated Federated Learning with Decoupled Adaptive Optimization”, International Conference on Machine Learning. PMLR, pp. 10298–10322. Available at: https://proceedings.mlr.press/v162/jin22e.html.
Vancouver
1. Jin J, Ren J, Zhou Y, Lyu L, Liu J, Dou D (2022) Accelerated Federated Learning with Decoupled Adaptive Optimization. In: International Conference on Machine Learning. PMLR, pp 10298–10322

BibTeX

@InProceedings{pmlr-v162-jin22e,
  title = 	 {Accelerated Federated Learning with Decoupled Adaptive Optimization},
  author =       {Jin, Jiayin and Ren, Jiaxiang and Zhou, Yang and Lyu, Lingjuan and Liu, Ji and Dou, Dejing},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {10298--10322},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/jin22e/jin22e.pdf},
  url = 	 {https://proceedings.mlr.press/v162/jin22e.html},
  abstract = 	 {The federated learning (FL) framework enables edge clients to collaboratively learn a shared inference model while keeping privacy of training data on clients. Recently, many heuristics efforts have been made to generalize centralized adaptive optimization methods, such as SGDM, Adam, AdaGrad, etc., to federated settings for improving convergence and accuracy. However, there is still a paucity of theoretical principles on where to and how to design and utilize adaptive optimization methods in federated settings. This work aims to develop novel adaptive optimization methods for FL from the perspective of dynamics of ordinary differential equations (ODEs). First, an analytic framework is established to build a connection between federated optimization methods and decompositions of ODEs of corresponding centralized optimizers. Second, based on this analytic framework, a momentum decoupling adaptive optimization method, FedDA, is developed to fully utilize the global momentum on each local iteration and accelerate the training convergence. Last but not least, full batch gradients are utilized to mimic centralized optimization in the end of the training process to ensure the convergence and overcome the possible inconsistency caused by adaptive optimization methods.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/