Accelerated Federated Learning with Decoupled Adaptive Optimization
Jiayin JinJiaxiang RenYang ZhouLingjuan LyuJi LiuDejing Dou
Proposes a principled ordinary differential equation decomposition framework and a decoupled adaptive optimization algorithm, FedDA, that accelerates federated learning convergence by accurately distributing centralized momentum updates across local client iterations.
Federated learning enables multiple edge devices, such as mobile phones, to collaboratively train a shared machine learning model without centralizing private user data. However, deploying this technology in real-world environments faces significant bottlenecks: excessive communication overhead between devices and the central server, and data heterogeneity across devices, which causes local models to drift away from the global objective. While centralized machine learning relies on adaptive optimization methods—such as stochastic gradient descent with momentum, Adam, and AdaGrad—to accelerate training, directly adapting these methods to federated environments often causes severe gradient deviations, training instability, and inconsistent convergence.
The article establishes a rigorous mathematical foundation for federated adaptive optimization and introduces FEDDA, a momentum-decoupling adaptive optimization framework designed to accelerate federated training while preserving model stability and accuracy.
The researchers developed their method by framing centralized optimization algorithms as continuous ordinary differential equation systems and analyzing how to decompose these systems across decentralized clients. Building on this framework, FEDDA decouples the calculation of global momentum from local parameter updates. This allows local devices to track momentum linearly without distorting local training steps. The aggregated momentum is then applied at the server level. To address late-stage training instability and convergence inconsistency, the approach shifts to single-iteration, full-batch client updates near the end of training to closely match centralized optimization behavior. The method was evaluated against nine state-of-the-art baselines across standard image classification (CIFAR-100, EMNIST) and natural language processing (Stack Overflow) benchmarks.
The evaluation produced several key findings. First, FEDDA significantly accelerated training convergence, achieving average convergence speed improvements of 34.3% on CIFAR-100, 22.6% on EMNIST, and 75.4% on Stack Overflow compared to baseline methods. Second, the decoupled global momentum reduced the theoretical momentum error growth from an exponential rate seen in standard local momentum approaches to a much slower algebraic rate. Third, FEDDA delivered superior final model accuracy across all tested optimizers, reaching over 51.8% accuracy on CIFAR-100 and up to 86.8% on EMNIST, consistently outperforming standard federated learning algorithms and existing adaptive frameworks. Finally, tests on Stack Overflow showed an average accuracy improvement of 22.3% when utilizing momentum-based stochastic gradient descent over competing methods.
These results demonstrate that federated learning systems do not need to sacrifice training speed or model quality to maintain communication efficiency and data privacy. By using decoupled momentum tracking and terminal stabilization, organizations can reduce the total training rounds required to reach target performance, directly lowering communication bandwidth demands, device energy consumption, and infrastructure costs.
Engineering and deployment teams implementing federated learning systems should consider adopting decoupled momentum mechanisms and late-stage full-batch stabilization to accelerate model updates. When configuring deployments, practitioners must carefully tune hyperparameters, particularly client and server learning rates and the number of local iterations per round, as excessively high local iterations can cause over-fitting on simple local datasets and degrade global model performance.
While the theoretical analysis and empirical results across standard benchmarks provide strong confidence in the method's effectiveness, the experimental findings are based on simulated federated benchmarks under controlled conditions. Decision-makers should validate performance through targeted pilot studies on their specific edge hardware configurations, network environments, and production data distributions before executing full-scale rollouts.
- Paper: Adaptive Federated Optimization, Sashank Reddi et al. (2020). Read this first to understand the server-side adaptive federated optimizers that FEDDA develops by decoupling momentum from local updates.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). Its introduction of FedAvg establishes the local-update and server-aggregation framework whose communication and client-drift limits motivate FEDDA.
- Paper: SCAFFOLD: Stochastic Controlled Averaging for Federated Learning, Sai Praneeth Karimireddy et al. (2019). SCAFFOLD explains how local client drift disrupts federated convergence, clarifying the optimization problem FEDDA addresses with decoupled momentum.
- Paper: Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization, Jianyu Wang et al. (2020). FedNova's analysis of objective inconsistency under heterogeneous local work provides useful grounding for FEDDA's focus on stable convergence across clients.
No sufficiently relevant recommendations were found.
