FedALA: Adaptive Local Aggregation for Personalized Federated Learning
Jianqing ZhangYang HuaHao WangTao SongZhengui XueRuhui MaHaibing Guan
Proposes an adaptive local aggregation framework for personalized federated learning that learns element-wise blending weights between global and local models to tailor client initialization to local objectives without incurring extra communication overhead.
Federated learning enables multiple distributed clients to train machine learning models collaboratively while keeping private data local. However, real-world data across clients is often highly unbalanced and non-identically distributed. This statistical heterogeneity causes traditional federated methods to produce a single global model that performs poorly on individual client tasks. While personalized federated learning techniques attempt to address this issue, existing approaches often require transferring multiple models between clients—incurring high communication overhead and privacy risks—or rely on coarse, rule-based aggregation methods that fail to capture the exact parameters an individual client needs.
The article develops and evaluates a new method called Federated Learning with Adaptive Local Aggregation, along with its core aggregation module. The main objective is to adaptively merge downloaded global models and local client models at an element-wise level to improve local model personalization without increasing baseline communication overhead.
To demonstrate this approach, the researchers conducted extensive empirical evaluations across five computer vision and natural language processing benchmarks, testing both shallow neural networks and deep architectures such as ResNet-18 across 20 to 100 clients. They compared the proposed method against eleven state-of-the-art baselines under both extreme class partitioning and realistic distributed data conditions, measuring model accuracy, training runtimes, and network transfer volume.
The analysis reveals several key findings. First, the proposed approach consistently outperforms all eleven benchmark methods, achieving up to a 3.27 percentage point increase in test accuracy over the strongest baseline on complex image datasets. Second, the standalone aggregation module is highly modular and elevates existing baseline methods when integrated, improving their test accuracy by up to 24.19 percentage points. Third, the system achieves these gains with zero additional network communication overhead relative to standard baseline protocols, requiring clients to transmit only a single model per iteration. Finally, restricting adaptive aggregation to the highest network layer reduces trainable parameters significantly while preserving strong accuracy, adding only about 0.34 minutes of computation time per iteration compared to standard approaches.
These results demonstrate that selective, element-level model aggregation successfully preserves valuable generic features from the global model while filtering out conflicting updates that misdirect local training. For organizations deploying distributed machine learning, this framework offers higher task performance and robust client scalability without escalating bandwidth costs or exacerbating client privacy vulnerabilities.
Based on these findings, teams managing distributed machine learning pipelines should consider incorporating the adaptive local aggregation module into their existing federated frameworks. Practitioners should prioritize applying the module primarily to higher model layers and utilize a moderate local data sample size to optimize the balance between computational runtime and final model accuracy.
The article's findings are supported by consistent, multi-run experiments across standard academic benchmarks. However, evaluation is limited to simulated distributed environments with up to 100 clients and fixed model architectures. Decision-makers should validate the framework in operational pilot deployments with real-world network latency and diverse edge hardware before executing large-scale organizational rollouts.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). It introduces the foundational FederatedAveraging (FedAvg) algorithm and the challenge of statistical data heterogeneity that FedALA directly addresses.
- Paper: Towards Personalized Federated Learning, Alysa Ziying Tan et al. (2021). It provides a systematic taxonomy and literature survey of personalized federated learning strategies, framing the exact problem setting and baseline approaches evaluated in FedALA.
- Paper: Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach, Alireza Fallah et al. (2020). It establishes Per-FedAvg, a core meta-learning benchmark for personalized federated learning that motivates the need for element-wise local adaptation without excessive computation.
- Paper: Personalized Federated Learning with Moreau Envelopes, Canh T. Dinh et al. (2020). It presents pFedMe, an influential bi-level personalized federated learning baseline that FedALA builds upon and benchmarks against.
- Paper: Federated Learning with Personalization Layers, Manoj Ghuhan Arivazhagan et al. (2019). It introduces FedPer's split-layer architecture for personalization, which FedALA directly compares against and refines through adaptive element-wise and layer-specific local aggregation.
- Paper: Ditto: Fair and Robust Federated Learning Through Personalization, Tian Li et al. (2020). It formulates Ditto as a localized regularization framework for personalized federated learning, serving as a primary state-of-the-art comparative baseline for FedALA.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). It develops the FedProx framework to handle data and system heterogeneity via proximal regularization, establishing a baseline mechanism for constraining local client drift.
- Paper: FedBN: Federated Learning on Non-IID Features via Local Batch Normalization, Xiaoxiao Li et al. (2021). It introduces FedBN's parameter-decoupling approach for local client adaptation, highlighting the benefits of selectively retaining local network parameters under non-IID conditions.
- Paper: Clustered Federated Learning: Model-Agnostic Distributed Multitask Optimization Under Privacy Constraints, Felix Sattler et al. (2019). It introduces clustered federated learning to handle incongruent local data distributions, offering essential context on multi-model aggregation techniques under non-IID data.
- Paper: SCAFFOLD: Stochastic Controlled Averaging for Federated Learning, Sai Praneeth Karimireddy et al. (2019). It formalizes client drift caused by non-IID data distributions and introduces control variates to correct gradient divergence during local training.
- Paper: An Aggregation-Free Federated Learning for Tackling Data Heterogeneity, Yuan Wang et al. (2024). It moves beyond the traditional aggregate-then-adapt framework of personalized federated learning methods like FedALA by proposing an entirely aggregation-free collaborative scheme.
- Paper: FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning, Jianqing Zhang et al. (2024). It extends personalized federated learning under data heterogeneity to also support model-heterogeneous architectures using contrastive trainable global prototypes.
