DAdaQuant: Doubly-adaptive quantization for communication-efficient Federated Learning
Robert HönigYiren ZhaoRobert Mullins
Proposes DAdaQuant, a communication-efficient federated learning algorithm that dynamically adjusts parameter quantization levels across training rounds and individual clients to achieve up to 2.8× higher uplink compression without sacrificing accuracy.
Federated learning enables organizations to train artificial intelligence models across decentralized edge devices, such as smartphones and sensors, while keeping raw user data private. However, repeatedly transmitting large model updates over networks with restricted upload bandwidth creates severe communication bottlenecks, driving up energy consumption and prolonging training cycles. While parameter quantization reduces message size by rounding model numbers into discrete bins, conventional approaches apply a single, fixed precision level across all devices and throughout the entire training process, leaving substantial efficiency gains untapped.
The article introduces and evaluates DAdaQuant, a novel compression algorithm designed to drastically reduce client-to-server communication in federated learning without sacrificing model accuracy or convergence speed. The framework achieves this through doubly-adaptive quantization, dynamically adjusting precision over time and across individual participating clients.
To assess the method, the authors developed a baseline called Federated QSGD—adapting stochastic fixed-point quantization with difference coding and lossless compression for federated systems—and integrated DAdaQuant into standard federated optimization workflows. They evaluated performance across five diverse benchmark datasets and model architectures, including logistic regression, convolutional neural networks, and recurrent neural networks, across image and natural language tasks while simulating device compute heterogeneity.
The evaluation demonstrates that DAdaQuant significantly improves communication efficiency, outperforming the strongest non-adaptive quantization baselines by up to 2.8 times across various tasks while maintaining target accuracy. The client-adaptive component alone delivers substantial gains on imbalanced datasets by allocating higher precision to heavily weighted clients and coarser precision to smaller ones, matching analytical variance bounds. Concurrently, the time-adaptive component saves bandwidth by starting with coarse precision and progressively increasing resolution as training stabilizes, adding negligible computational overhead of approximately 1%.
These findings indicate that communication bottlenecks in federated learning can be significantly mitigated through dynamic precision scheduling rather than static compression schemes. By cutting total data transfer by orders of magnitude compared to uncompressed baselines, DAdaQuant enables organizations to reduce mobile data costs, lower energy footprints on edge devices, and accelerate distributed training cycles.
Engineering teams deploying federated learning systems should adopt doubly-adaptive quantization strategies for client-to-server updates, particularly in heterogeneous edge environments where client data sizes vary widely. Further work should explore applying DAdaQuant principles to emerging vector quantizers and evaluating end-to-end performance in real-world cellular deployments with unstable connectivity.
Confidence in these findings is supported by rigorous mathematical proofs of optimality for fixed-point quantization and consistent empirical results across diverse model architectures. Readers should note that bandwidth savings depend partly on local dataset size variations, and tasks requiring high initial precision may require conservative starting quantization settings to prevent early convergence slowdowns.
- Paper: Federated Learning: Strategies for Improving Communication Efficiency, Jakub Konečný et al. (2016). This foundational paper introduces communication-reduction strategies for federated learning, including structured updates and probabilistic quantization, providing essential background for adaptive quantization schemes.
- Paper: QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding, Dan Alistarh et al. (2016). It establishes the theoretical principles and convergence guarantees of randomized gradient quantization and encoding in distributed SGD that underpin federated model compression.
- Paper: Robust and Communication-Efficient Federated Learning From Non-i.i.d. Data, Felix Sattler et al. (2019). It analyzes the vulnerability of standard gradient quantization methods under non-IID client distributions, motivating the need for robust, client-adaptive compression protocols.
- Paper: signSGD: compressed optimisation for non-convex problems, Jeremy Bernstein et al. (2018). It analyzes 1-bit sign-based gradient compression and theoretical convergence in distributed non-convex optimization, serving as a core non-adaptive compression baseline.
- Paper: Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training, Yujun Lin et al. (2018). It details gradient compression and error-feedback techniques to reduce communication bandwidth during distributed neural network training.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). It establishes the foundational FederatedAveraging framework and defines the communication-efficiency challenges in decentralized optimization across edge clients.
- Paper: FedNL: Making Newton-Type Methods Applicable to Federated Learning, Mher Safaryan et al. (2022). It extends communication-efficient federated learning to second-order Newton-type methods by compressing curvature matrices and adapting to heterogeneous edge clients.
