FedALA: Adaptive Local Aggregation for Personalized Federated Learning

Jianqing ZhangYang HuaHao WangTao SongZhengui XueRuhui MaHaibing Guan

article2023AAAI534 citations

Proposes an adaptive local aggregation framework for personalized federated learning that learns element-wise blending weights between global and local models to tailor client initialization to local objectives without incurring extra communication overhead.

Listen

Federated learning enables multiple distributed clients to train machine learning models collaboratively while keeping private data local. However, real-world data across clients is often highly unbalanced and non-identically distributed. This statistical heterogeneity causes traditional federated methods to produce a single global model that performs poorly on individual client tasks. While personalized federated learning techniques attempt to address this issue, existing approaches often require transferring multiple models between clients—incurring high communication overhead and privacy risks—or rely on coarse, rule-based aggregation methods that fail to capture the exact parameters an individual client needs.

The article develops and evaluates a new method called Federated Learning with Adaptive Local Aggregation, along with its core aggregation module. The main objective is to adaptively merge downloaded global models and local client models at an element-wise level to improve local model personalization without increasing baseline communication overhead.

To demonstrate this approach, the researchers conducted extensive empirical evaluations across five computer vision and natural language processing benchmarks, testing both shallow neural networks and deep architectures such as ResNet-18 across 20 to 100 clients. They compared the proposed method against eleven state-of-the-art baselines under both extreme class partitioning and realistic distributed data conditions, measuring model accuracy, training runtimes, and network transfer volume.

The analysis reveals several key findings. First, the proposed approach consistently outperforms all eleven benchmark methods, achieving up to a 3.27 percentage point increase in test accuracy over the strongest baseline on complex image datasets. Second, the standalone aggregation module is highly modular and elevates existing baseline methods when integrated, improving their test accuracy by up to 24.19 percentage points. Third, the system achieves these gains with zero additional network communication overhead relative to standard baseline protocols, requiring clients to transmit only a single model per iteration. Finally, restricting adaptive aggregation to the highest network layer reduces trainable parameters significantly while preserving strong accuracy, adding only about 0.34 minutes of computation time per iteration compared to standard approaches.

These results demonstrate that selective, element-level model aggregation successfully preserves valuable generic features from the global model while filtering out conflicting updates that misdirect local training. For organizations deploying distributed machine learning, this framework offers higher task performance and robust client scalability without escalating bandwidth costs or exacerbating client privacy vulnerabilities.

Based on these findings, teams managing distributed machine learning pipelines should consider incorporating the adaptive local aggregation module into their existing federated frameworks. Practitioners should prioritize applying the module primarily to higher model layers and utilize a moderate local data sample size to optimize the balance between computational runtime and final model accuracy.

The article's findings are supported by consistent, multi-run experiments across standard academic benchmarks. However, evaluation is limited to simulated distributed environments with up to 100 clients and fixed model architectures. Decision-makers should validate the framework in operational pilot deployments with real-world network latency and diverse edge hardware before executing large-scale organizational rollouts.

Cover for FedALA: Adaptive Local Aggregation for Personalized Federated Learning

Abstract

A key challenge in federated learning (FL) is the statistical heterogeneity that impairs the generalization of the global model on each client. To address this, we propose a method Federated learning with Adaptive Local Aggregation (FedALA) by capturing the desired information in the global model for client models in personalized FL. The key component of FedALA is an Adaptive Local Aggregation (ALA) module, which can adaptively aggregate the downloaded global model and local model towards the local objective on each client to initialize the local model before training in each iteration. To evaluate the effectiveness of FedALA, we conduct extensive experiments with five benchmark datasets in computer vision and natural language processing domains. FedALA outperforms eleven state-of-the-art baselines by up to 3.27% in test accuracy. Furthermore, we also apply ALA module to other federated learning methods and achieve up to 24.19% improvement in test accuracy. Code is available at https://github.com/TsingZ0/FedALA.

Table of Contents

  • Introduction
  • Related Work
  • Traditional Federated Learning
  • Personalized Federated Learning
  • Method
  • Problem Statement
  • Adaptive Local Aggregation (ALA)
  • Analysis of ALA
  • Experiments
  • Experimental Setup
  • Effect of Hyperparameters
  • Performance Comparison and Analysis
  • Applicability of ALA
  • Update Direction Correction
  • Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Adaptive Local Aggregation (ALA) Formulation

    model/method

    In personalized federated learning, clients possess statistically heterogeneous local datasets D1,…,DND_1, \dots, D_N. Rather than overwriting the local model Θit−1\Theta_i^{t-1} of client ii with the broadcast global model Θt−1\Theta^{t-1} at communication round tt (as in FedAvg), Adaptive Local Aggregation (ALA) initializes the local model Θ^it\hat{\Theta}_i^t via element-wise weighted aggregation toward the client's local empirical loss objective:

    Θ^it:=Θit−1+(Θt−1−Θit−1)⊙[1∣Θi∣−p;Wip]\hat{\Theta}_i^t := \Theta_i^{t-1} + (\Theta^{t-1} - \Theta_i^{t-1}) \odot [\mathbf{1}_{|\Theta_i|-p}; W_i^p]

    where ⊙\odot is the Hadamard (element-wise) product, ∣Θi∣|\Theta_i| denotes the total number of layers (or residual blocks) of the model, and p∈{1,…,∣Θi∣}p \in \{1, \dots, |\Theta_i|\} is a hyperparameter specifying the number of higher layers subjected to adaptive weight learning. The vector 1∣Θi∣−p\mathbf{1}_{|\Theta_i|-p} is a constant all-ones vector having the dimension of the lower ∣Θi∣−p|\Theta_i|-p layers, which directly incorporates the generic feature representations learned by the global model. The aggregation weight vector WipW_i^p corresponds to the remaining pp higher layers, with elements constrained to the range [0,1][0, 1] via the clipping function σ(w)=max⁡(0,min⁡(1,w))\sigma(w) = \max(0, \min(1, w)).

    To optimize WipW_i^p, client ii freezes the model weights Θit−1\Theta_i^{t-1} and Θt−1\Theta^{t-1}, samples a random subset Dis,t⊆DiD_i^{s,t} \subseteq D_i comprising s%s\% of its local data, and updates WipW_i^p via gradient descent with learning rate η\eta:

    Wip←Wip−η∇WipL(Θ^it,Dis,t;Θt−1)W_i^p \leftarrow W_i^p - \eta \nabla_{W_i^p} \mathcal{L}(\hat{\Theta}_i^t, D_i^{s,t}; \Theta^{t-1})

  2. Knowl 2 — FedALA Algorithm

    algorithm

    The Federated Learning with Adaptive Local Aggregation (FedALA) algorithm orchestrates distributed training across NN clients with personalized local initialization.

    Input: NN clients, client joining ratio ρ\rho, loss function L\mathcal{L}, initial global model Θ0\Theta^0, local learning rate α\alpha, ALA learning rate η\eta, data subsample ratio s%s\%, ALA depth range pp, clipping operator σ(⋅)\sigma(\cdot)
    Output: Client personalized local models Θ^1,…,Θ^N\hat{\Theta}_1, \dots, \hat{\Theta}_N
    Server broadcasts Θ0\Theta^0 to all clients
    Each client i∈[N]i \in [N] initializes WipW_i^p to all ones
    for iteration t=1,…,Tt = 1, \dots, T do
        Server samples active client subset ItI^t of size ρN\rho N
        Server sends Θt−1\Theta^{t-1} to clients in ItI^t
        for client i∈Iti \in I^t in parallel do
            Client ii samples s%s\% local data subset Dis,t⊂DiD_i^{s,t} \subset D_i
            if t=2t = 2 then
                while WipW_i^p has not converged do
                    Wip←Wip−η∇WipL(Θ^it,Dis,t;Θt−1)W_i^p \leftarrow W_i^p - \eta \nabla_{W_i^p} \mathcal{L}(\hat{\Theta}_i^t, D_i^{s,t}; \Theta^{t-1})
                    Wip←σ(Wip)W_i^p \leftarrow \sigma(W_i^p)
            else if t>2t > 2 then
                Wip←Wip−η∇WipL(Θ^it,Dis,t;Θt−1)W_i^p \leftarrow W_i^p - \eta \nabla_{W_i^p} \mathcal{L}(\hat{\Theta}_i^t, D_i^{s,t}; \Theta^{t-1})
                Wip←σ(Wip)W_i^p \leftarrow \sigma(W_i^p)
            Client ii computes Θ^it=Θit−1+(Θt−1−Θit−1)⊙[1∣Θi∣−p;Wip]\hat{\Theta}_i^t = \Theta_i^{t-1} + (\Theta^{t-1} - \Theta_i^{t-1}) \odot [\mathbf{1}_{|\Theta_i|-p}; W_i^p]
            Client ii trains local model: Θit←Θ^it−α∇Θ^itL(Θ^it,Di;Θt−1)\Theta_i^t \leftarrow \hat{\Theta}_i^t - \alpha \nabla_{\hat{\Theta}_i^t} \mathcal{L}(\hat{\Theta}_i^t, D_i; \Theta^{t-1})
            Client ii uploads Θit\Theta_i^t to the server
        Server aggregates global model: Θt←∑i∈Itki∑j∈ItkjΘit\Theta^t \leftarrow \sum_{i \in I^t} \frac{k_i}{\sum_{j \in I^t} k_j} \Theta_i^t, where ki=∣Di∣k_i = |D_i|
    return Θ^1,…,Θ^N\hat{\Theta}_1, \dots, \hat{\Theta}_N

    In round t=1t=1, ALA is bypassed because Θi0=Θ0\Theta^0_i = \Theta^0. In round t=2t=2, WipW_i^p is trained over multiple epochs until convergence; for all rounds t>2t > 2, WipW_i^p is fine-tuned for only a single epoch, allowing weight reuse while reducing local computation.

  3. Knowl 3 — Effective Parameter Update and Direction Correction in ALA

    theoretical result

    When applying Adaptive Local Aggregation (ALA) over the entire network (p=∣Θi∣p = |\Theta_i|) without clipping constraints, the gradient of the local objective Lit=L(Θ^it,Dis,t;Θt−1)\mathcal{L}_i^t = \mathcal{L}(\hat{\Theta}_i^t, D_i^{s,t}; \Theta^{t-1}) with respect to the weight vector WiW_i is:

    ∇WiLit=(Θt−1−Θit−1)⊙∇Θ^itLit\nabla_{W_i} \mathcal{L}_i^t = (\Theta^{t-1} - \Theta_i^{t-1}) \odot \nabla_{\hat{\Theta}_i^t} \mathcal{L}_i^t

    Applying a gradient descent update Wi←Wi−η∇WiLitW_i \leftarrow W_i - \eta \nabla_{W_i} \mathcal{L}_i^t and substituting into the local initialization Θ^it=Θit−1+(Θt−1−Θit−1)⊙Wi\hat{\Theta}_i^t = \Theta_i^{t-1} + (\Theta^{t-1} - \Theta_i^{t-1}) \odot W_i yields an equivalent parameter update on the initialized local model Θ^it\hat{\Theta}_i^t:

    Θ^it←Θ^it−η(Θt−1−Θit−1)⊙(Θt−1−Θit−1)⊙∇Θ^itLit\hat{\Theta}_i^t \leftarrow \hat{\Theta}_i^t - \eta (\Theta^{t-1} - \Theta_i^{t-1}) \odot (\Theta^{t-1} - \Theta_i^{t-1}) \odot \nabla_{\hat{\Theta}_i^t} \mathcal{L}_i^t

    This demonstrates that updating the aggregation weights WiW_i is equivalent to scaling the local loss gradient ∇Θ^itLit\nabla_{\hat{\Theta}_i^t} \mathcal{L}_i^t element-wise by the squared difference between the received global model and the client's previous local model, (Θt−1−Θit−1)⊙2(\Theta^{t-1} - \Theta_i^{t-1})^{\odot 2}. The difference dynamic (Θt−1−Θit−1)(\Theta^{t-1} - \Theta_i^{t-1}) incorporates global generic information directly into the local trajectory, correcting the parameter initialization toward the local loss reduction basin rather than being misdirected by conflicting global model updates.

  4. Knowl 4 — Performance Comparison on Pathological and Practical Non-IID Benchmarks

    data/table

    Test accuracy across computer vision and natural language processing benchmarks under pathological non-IID settings (clients receive data from 2 out of 10 classes for MNIST and CIFAR-10, and 10 out of 100 classes for CIFAR-100) and practical non-IID settings (Dirichlet distribution Dir(β)\text{Dir}(\beta) with default β=0.1\beta = 0.1). Models evaluated include a 4-layer CNN for MNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet (TINY); ResNet-18 for Tiny-ImageNet (TINY*); and fastText for AG News. Experiments use 20 clients with client participation ratio ρ=1\rho = 1, 2000 iterations, batch size 10, 1 local epoch, and evaluate local test accuracy across 5 independent runs.

    Settings Pathological heterogeneous setting Practical heterogeneous setting
    Methods MNIST Cifar10 Cifar100 Cifar10 Cifar100 TINY TINY* AG News
    FedAvg 97.93±\pm0.05 55.09±\pm0.83 25.98±\pm0.13 59.16±\pm0.47 31.89±\pm0.47 19.46±\pm0.20 19.45±\pm0.13 79.57±\pm0.17
    FedProx 98.01±\pm0.09 55.06±\pm0.75 25.94±\pm0.16 59.21±\pm0.40 31.99±\pm0.41 19.37±\pm0.22 19.27±\pm0.23 79.35±\pm0.23
    FedAvg-C 99.79±\pm0.00 92.13±\pm0.03 66.17±\pm0.03 90.34±\pm0.01 51.80±\pm0.02 30.67±\pm0.08 36.94±\pm0.10 95.89±\pm0.25
    FedProx-C 99.80±\pm0.04 92.12±\pm0.03 66.07±\pm0.08 90.33±\pm0.01 51.84±\pm0.07 30.77±\pm0.13 38.78±\pm0.52 96.10±\pm0.22
    Per-FedAvg 99.63±\pm0.02 89.63±\pm0.23 56.80±\pm0.26 87.74±\pm0.19 44.28±\pm0.33 25.07±\pm0.07 21.81±\pm0.54 93.27±\pm0.25
    FedRep 99.77±\pm0.03 91.93±\pm0.14 67.56±\pm0.31 90.40±\pm0.24 52.39±\pm0.35 37.27±\pm0.20 39.95±\pm0.61 96.28±\pm0.14
    pFedMe 99.75±\pm0.02 90.11±\pm0.10 58.20±\pm0.14 88.09±\pm0.32 47.34±\pm0.46 26.93±\pm0.19 33.44±\pm0.33 91.41±\pm0.22
    Ditto 99.81±\pm0.00 92.39±\pm0.06 67.23±\pm0.07 90.59±\pm0.01 52.87±\pm0.64 32.15±\pm0.04 35.92±\pm0.43 95.45±\pm0.17
    FedAMP 99.76±\pm0.02 90.79±\pm0.16 64.34±\pm0.37 88.70±\pm0.18 47.69±\pm0.49 27.99±\pm0.11 29.11±\pm0.15 94.18±\pm0.09
    FedPHP 99.73±\pm0.00 90.01±\pm0.00 63.09±\pm0.04 88.92±\pm0.02 50.52±\pm0.16 35.69±\pm3.26 29.90±\pm0.51 94.38±\pm0.12
    FedFomo 99.83±\pm0.00 91.85±\pm0.02 62.49±\pm0.22 88.06±\pm0.02 45.39±\pm0.45 26.33±\pm0.22 26.84±\pm0.11 95.84±\pm0.15
    APPLE 99.75±\pm0.01 90.97±\pm0.05 65.80±\pm0.08 89.37±\pm0.11 53.22±\pm0.20 35.04±\pm0.47 39.93±\pm0.52 95.63±\pm0.21
    PartialFed 99.86±\pm0.01 89.60±\pm0.13 61.39±\pm0.12 87.38±\pm0.08 48.81±\pm0.20 35.26±\pm0.18 37.50±\pm0.16 85.20±\pm0.16
    FedALA 99.88±\pm0.01 92.44±\pm0.02 67.83±\pm0.06 90.67±\pm0.03 55.92±\pm0.03 40.54±\pm0.02 41.94±\pm0.05 96.52±\pm0.08

    FedALA achieves the highest test accuracy across all eight benchmark tasks. In the practical non-IID setting on TINY (4-layer CNN), FedALA exceeds the closest baseline (FedRep) by 3.27% accuracy (40.54% vs. 37.27%).

  5. Knowl 5 — Computational and Communication Complexity of FedALA

    data/table

    Computational cost (total training time to convergence and time per iteration) measured on Tiny-ImageNet with ResNet-18 (TINY*) using eight NVIDIA 2080 Ti GPUs, along with per-client communication overhead (number of transmitted parameters per iteration). Σ\Sigma is the total parameter count of the model backbone, αf<1\alpha_f < 1 is the proportion of parameters in the feature extractor, and MM is the number of neighbor client models downloaded per round by FedFomo and APPLE (set to M=20M=20).

    Methods Total time Time/iter. Param./iter.
    FedAvg 365 min 1.59 min 2∗Σ2 * \Sigma
    FedProx 325 min 1.99 min 2∗Σ2 * \Sigma
    FedAvg-C 607 min 24.28 min 2∗Σ2 * \Sigma
    FedProx-C 711 min 28.44 min 2∗Σ2 * \Sigma
    Per-FedAvg 121 min 3.56 min 2∗Σ2 * \Sigma
    FedRep 471 min 4.09 min 2∗αf∗Σ2 * \alpha_f * \Sigma
    pFedMe 1157 min 10.24 min 2∗Σ2 * \Sigma
    Ditto 318 min 11.78 min 2∗Σ2 * \Sigma
    FedAMP 92 min 1.53 min 2∗Σ2 * \Sigma
    FedPHP 264 min 4.06 min 2∗Σ2 * \Sigma
    FedFomo 193 min 2.72 min (1+M)∗Σ(1 + M) * \Sigma
    APPLE 132 min 2.93 min (1+M)∗Σ(1 + M) * \Sigma
    PartialFed 693 min 2.13 min 2∗Σ2 * \Sigma
    FedALA 7+116 min 1.93 min 2∗Σ2 * \Sigma

    FedALA maintains standard FedAvg communication cost (2Σ2\Sigma) by exchanging only one global and one local model per round. In computation, FedALA requires 1.93 minutes per iteration (only 0.34 min above FedAvg) and converges in 123 total minutes (7 min7\text{ min} for iteration 2 initial weight convergence plus 116 min116\text{ min} for subsequent rounds), outperforming fine-tuning baselines (FedAvg-C: 607 min) and dual-model optimization methods (pFedMe: 1157 min, Ditto: 318 min).

  6. Knowl 6 — Plug-in Applicability of Adaptive Local Aggregation to Existing FL Methods

    empirical result

    Because Adaptive Local Aggregation (ALA) modifies only the local model initialization step without altering downstream local training or server aggregation logic, it can be integrated into existing federated learning algorithms. Evaluating ALA (s=80,p=1s = 80, p = 1) on Tiny-ImageNet (4-layer CNN) and CIFAR-100 in the practical non-IID setting (extDir(0.1) ext{Dir}(0.1)) yields the following performance gains:

    • Traditional FL algorithms gain large accuracy boosts: FedAvg increases from 19.46% to 40.54% (+21.08%) on Tiny-ImageNet and 31.89% to 55.92% (+24.03%) on CIFAR-100. FedProx increases from 19.37% to 40.53% (+21.16%) on Tiny-ImageNet and 31.99% to 56.18% (+24.19%) on CIFAR-100.
    • Personalized FL methods also improve: Per-FedAvg improves by +5.83% (Tiny-ImageNet) and +4.40% (CIFAR-100); Ditto improves by +8.60% (Tiny-ImageNet) and +3.46% (CIFAR-100); FedPHP improves by +4.47% (Tiny-ImageNet) and +3.76% (CIFAR-100).
    • Some pFL algorithms exhibit marginal gains: FedRep improves by +0.62%/+0.63% because its representation layer is already frozen during local fine-tuning; pFedMe improves by +0.37%/+0.57% because batch-level Moreau envelope optimization dominates; FedAMP improves by +0.19%/+0.34% because attention aggregation discards generic model information; PartialFed improves by +0.14%/+0.18% because batch-level parameter selection re-aggregates parameters, overriding the ALA initialization.
  7. Knowl 7 — Parameter Sensitivity to Aggregation Range and Data Subsampling in ALA

    empirical result

    Evaluation of hyperparameter sensitivity on Tiny-ImageNet with ResNet-18 (TINY*) in the practical non-IID setting (Dir(0.1)\text{Dir}(0.1)) highlights the trade-off between aggregation range pp (number of higher layers subjected to adaptive weights), subsampling ratio s%s\% (percentage of local data used to train WipW_i^p), test accuracy, and parameter overhead:

    • Aggregation Range (pp) with s=80s = 80: When varying pp from 6 down to 1, test accuracy remains stable: p=6p=6 (41.71%, 11.182M trainable parameters), p=5p=5 (41.54%, 11.172M parameters), p=4p=4 (41.62%, 11.024M parameters), p=3p=3 (41.86%, 10.499M parameters), p=2p=2 (42.47%, 8.399M parameters), and p=1p=1 (41.94%, 0.005M parameters). Restricting adaptive weight learning to the final layer (p=1p=1) reduces the number of learnable parameters by over 99.9%99.9\% (from 11.182M to 5,000) with negligible accuracy decline, confirming that lower layers capture generic features that can be directly overwritten with global model weights.
    • Data Subsampling (s%s\%) with p=1p = 1: Increasing ss from 5% to 100% improves accuracy: s=5%s=5\% (39.53%), s=10%s=10\% (40.62%), s=20%s=20\% (40.02%), s=40%s=40\% (40.23%), s=60%s=60\% (41.11%), s=80%s=80\% (41.94%), and s=100%s=100\% (42.11%). Subsampling s=80%s = 80\% achieves nearly the same performance as full local data (s=100%s=100\%) while reducing the gradient computation during weight optimization.
  8. Knowl 8 — Robustness to Data Heterogeneity and Client Scalability

    empirical result

    FedALA exhibits strong robustness across varying degrees of statistical heterogeneity and varying numbers of participating clients:

    • Heterogeneity Levels: Under severe non-IID distribution (Dir(0.01)\text{Dir}(0.01) on Tiny-ImageNet), FedALA achieves 55.75% test accuracy (compared to FedRep's 55.43% and Ditto's 50.62%). Under moderate non-IID distribution (Dir(0.5)\text{Dir}(0.5) on Tiny-ImageNet), baseline pFL methods that excessively emphasize personalization degrade below FedAvg (21.14%), including FedAMP (12.48%), FedFomo (11.59%), Per-FedAvg (16.36%), and FedRep (16.74%). In contrast, FedALA achieves 27.85% accuracy, maintaining superiority across diverse heterogeneity levels.
    • Client Scalability: On CIFAR-100 with a fixed overall dataset, increasing the client pool from 20 to 50 and 100 clients reduces the data available per client. At 50 clients, FedALA attains 55.61% accuracy (Ditto: 54.22%, FedPHP: 52.44%). At 100 clients, FedALA drops by less than 1% to 54.68% accuracy, while baselines degrade substantially (Per-FedAvg drops to 36.07%, FedFomo to 38.91%, PartialFed to 39.31%, FedRep to 44.61%, and APPLE to 52.81%).

Coverage note — None was omitted; all key contributions including algorithm design, theoretical analysis of ALA update dynamics, comprehensive benchmark tables, computational/communication complexity, hyperparameter sensitivity, plug-in generality, heterogeneity scaling, and trajectory analysis are covered.

References

  1. 1.Arjovsky, M.; Chintala, S.; and Bottou, L. 2017. Wasserstein generative adversarial networks. In ICML.
  2. 2.Chrabaszcz, P.; Loshchilov, I.; and Hutter, F. 2017. A Down-sampled Variant of Imagenet as an Alternative to the Cifar Datasets. arXiv preprint arXiv:1707.08819.
  3. 3.Collins, L.; Hassani, H.; Mokhtari, A.; and Shakkottai, S. 2021. Exploiting Shared Representations for Personalized Federated Learning. In ICML.
  4. 4.Courbariaux, M.; Hubara, I.; Soudry, D.; El-Yaniv, R.; and Bengio, Y. 2016. Binarized Neural Networks: Training Deep Neural Networks with Weights and Activations Constrained to +1 or -1. arXiv preprint arXiv:1602.02830.
  5. 5.Fallah, A.; Mokhtari, A.; and Ozdaglar, A. 2020. Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach. In NeurIPS.
  6. 6.Finn, C.; Abbeel, P.; and Levine, S. 2017. Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In ICML.
  7. 7.He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep Residual Learning for Image Recognition. In CVPR.
  8. 8.Huang, Y.; Chu, L.; Zhou, Z.; Wang, L.; Liu, J.; Pei, J.; and Zhang, Y. 2021. Personalized Cross-Silo Federated Learning on Non-IID Data. In AAAI.
  9. 9.Joulin, A.; Grave, E.; Bojanowski, P.; and Mikolov, T. 2017. Bag of Tricks for Efficient Text Classification”. In EACL.
  10. 10.Kairouz, P.; McMahan, H. B.; Avent, B.; Bellet, A.; Bennis, M.; Bhagoji, A. N.; Bonawitz, K.; Charles, Z.; Cormode, G.; Cummings, R.; et al. 2019. Advances and Open Problems in Federated Learning. arXiv preprint arXiv:1912.04977.
  11. 11.Krizhevsky, A.; and Geoffrey, H. 2009. Learning Multiple Layers of Features From Tiny Images. Citeseer.
  12. 12.LeCun, Y.; Bengio, Y.; and Hinton, G. 2015. Deep learning. nature, 521(7553): 436–444.
  13. 13.LeCun, Y.; Bottou, L.; Bengio, Y.; and Haffner, P. 1998. Gradient-based Learning Applied to Document Recognition. Proceedings of the IEEE, 86(11): 2278–2324.
  14. 14.Li, H.; Xu, Z.; Taylor, G.; Studer, C.; and Goldstein, T. 2018. Visualizing the Loss Landscape of Neural Nets. In NeurIPS.
  15. 15.Li, Q.; He, B.; and Song, D. 2021. Model-Contrastive Federated Learning. In CVPR.
  16. 16.Li, T.; Hu, S.; Beirami, A.; and Smith, V. 2021a. Ditto: Fair and Robust Federated Learning Through Personalization. In ICML.
  17. 17.Li, T.; Sahu, A. K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V. 2020. Federated Optimization in Heterogeneous Networks. In MLSys.
  18. 18.Li, X.-C.; Zhan, D.-C.; Shao, Y.; Li, B.; and Song, S. 2021b. FedPHP: Federated Personalization with Inherited Private Models. In ECML PKDD.
  19. 19.Lin, T.; Kong, L.; Stich, S. U.; and Jaggi, M. 2020. Ensemble Distillation for Robust Model Fusion in Federated Learning. In NeurIPS.
  20. 20.Luo, J.; and Wu, S. 2021. Adapt to Adaptation: Learning Personalization for Cross-Silo Federated Learning. arXiv preprint arXiv:2110.08394.
  21. 21.Luo, L.; Xiong, Y.; Liu, Y.; and Sun, X. 2018. Adaptive Gradient Methods with Dynamic Bound of Learning Rate. In ICLR.
  22. 22.McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and y Arcas, B. A. 2017. Communication-Efficient Learning of Deep Networks from Decentralized Data. In AISTATS.
  23. 23.Reisizadeh, A.; Mokhtari, A.; Hassani, H.; Jadbabaie, A.; and Pedarsani, R. 2020. Fedpaq: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization. In AISTATS.
  24. 24.Shamsian, A.; Navon, A.; Fetaya, E.; and Chechik, G. 2021. Personalized Federated Learning using Hypernetworks. In ICML.
  25. 25.Sun, B.; Huo, H.; Yang, Y.; and Bai, B. 2021. PartialFed: Cross-Domain Personalized Federated Learning via Partial Initialization. In NeurIPS.
  26. 26.T Dinh, C.; Tran, N.; and Nguyen, T. D. 2020. Personalized Federated Learning with Moreau Envelopes. In NeurIPS.
  27. 27.Van Hasselt, H.; Guez, A.; and Silver, D. 2016. Deep reinforcement learning with double q-learning. In AAAI.
  28. 28.Wang, H.; Kaplan, Z.; Niu, D.; and Li, B. 2020a. Optimizing Federated Learning on Non-IID Data with Reinforcement Learning. In InfoComm.
  29. 29.Wang, H.; Yurochkin, M.; Sun, Y.; Papailiopoulos, D.; and Khazaeni, Y. 2020b. Federated learning with matched averaging. arXiv preprint arXiv:2002.06440.
  30. 30.Wang, J.; Liu, Q.; Liang, H.; Joshi, G.; and Poor, H. V. 2020c. Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization. In NeurIPS.
  31. 31.Yosinski, J.; Clune, J.; Bengio, Y.; and Lipson, H. 2014. How Transferable Are Features in Deep Neural Networks? In NeurIPS.
  32. 32.Zhang, M.; Sapra, K.; Fidler, S.; Yeung, S.; and Alvarez, J. M. 2020. Personalized Federated Learning with First Order Model Optimization. In ICLR.
  33. 33.Zhang, X.; Zhao, J.; and LeCun, Y. 2015. Character-level Convolutional Networks for Text Classification. In NeurIPS.
  34. 34.Zhao, Y.; Li, M.; Lai, L.; Suda, N.; Civin, D.; and Chandra, V. 2018. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582.
  35. 35.Zhu, L.; Liu, Z.; and Han, S. 2019. Deep leakage from gradients. In NeurIPS.
  36. 36.Zhu, Z.; Hong, J.; and Zhou, J. 2021. Data-Free Knowledge Distillation for Heterogeneous Federated Learning. In ICML.

Citation

MLA
Zhang, J., et al. “FedALA: Adaptive Local Aggregation for Personalized Federated Learning”. Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 9, 2023, pp. 11237–44, https://doi.org/10.1609/aaai.v37i9.26330.
APA
Zhang, J., Hua, Y., Wang, H., Song, T., Xue, Z., Ma, R., & Guan, H. (2023). FedALA: Adaptive Local Aggregation for Personalized Federated Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 37(9), 11237–11244. https://doi.org/10.1609/aaai.v37i9.26330
Chicago
Zhang, J., Y. Hua, H. Wang, et al. 2023. “FedALA: Adaptive Local Aggregation for Personalized Federated Learning”. Proceedings of the AAAI Conference on Artificial Intelligence 37 (9): 11237–44. https://doi.org/10.1609/aaai.v37i9.26330.
Harvard
Zhang, J. et al. (2023) “FedALA: Adaptive Local Aggregation for Personalized Federated Learning”, Proceedings of the AAAI Conference on Artificial Intelligence, 37(9), pp. 11237–11244. Available at: https://doi.org/10.1609/aaai.v37i9.26330.
Vancouver
1. Zhang J, Hua Y, Wang H, Song T, Xue Z, Ma R, Guan H (2023) FedALA: Adaptive Local Aggregation for Personalized Federated Learning. Proceedings of the AAAI Conference on Artificial Intelligence 37:11237–11244

BibTeX

@article{Zhang_2023, title={FedALA: Adaptive Local Aggregation for Personalized Federated Learning}, volume={37}, ISSN={2159-5399}, url={http://dx.doi.org/10.1609/aaai.v37i9.26330}, DOI={10.1609/aaai.v37i9.26330}, number={9}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, publisher={Association for the Advancement of Artificial Intelligence (AAAI)}, author={Zhang, Jianqing and Hua, Yang and Wang, Hao and Song, Tao and Xue, Zhengui and Ma, Ruhui and Guan, Haibing}, year={2023}, month=June, pages={11237–11244} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF