Towards Personalized Federated Learning

Alysa Ziying TanHan YuLizhen CuiQiang Yang

article2021IEEE Transactions on Neural Networks and Learning Systems1,349 citations

Presents a structured taxonomy of personalized federated learning strategies designed to overcome data heterogeneity across private devices, identifying critical open challenges in architecture design, trustworthy learning, and realistic benchmarking.

Listen

Modern artificial intelligence deployments increasingly rely on data distributed across personal edge devices and institutional silos. Concurrently, strict privacy regulations such as the General Data Protection Regulation and commercial considerations prevent pooling raw data into centralized repositories. While federated learning enables collaborative model training across distributed clients without directly sharing raw records, standard approaches train a single shared model aimed at the average participant. In real-world environments, local data distributions differ significantly across users, causing standard federated optimization to suffer from client drift, poor convergence, and inadequate local accuracy.

The article systematically reviews personalized federated learning frameworks designed to overcome data heterogeneity. It establishes a structured taxonomy that evaluates how different personalization strategies improve model accuracy and system efficiency on diverse, non-identical data distributions.

The review synthesizes current research across two overarching strategies: global model personalization and learning individual personalized models. Global model personalization retains a single shared model base and adapts it via data manipulation (such as augmentation or client selection) or model optimization (including regularization, meta-learning, and transfer learning). Conversely, learning personalized models customizes model architectures or aggregates clients based on mathematical similarity, utilizing techniques such as parameter decoupling, knowledge distillation, multi-task learning, model interpolation, and clustering.

The article's primary findings identify distinct trade-offs across these techniques. First, model-based personalization using meta-learning and regularization significantly mitigates client drift, but computing higher-order gradients introduces heavy computational overhead on edge devices. Second, architecture-based approaches using knowledge distillation allow resource-constrained clients to run customized, lightweight local architectures, though they often depend on shared proxy datasets that introduce data collection challenges. Third, similarity-based methods, particularly clustering and multi-task learning, deliver high local accuracy when natural subgroups exist, but they multiply communication overhead by broadcasting multiple group models and scale poorly in large federated networks. Fourth, current evaluation across the literature relies heavily on artificial, single-dimension data skew simulations rather than realistic multi-modal datasets, while largely overlooking holistic trustworthy metrics such as algorithmic fairness, explainability, and adversarial robustness.

These findings indicate that no single personalization strategy suits all enterprise deployments. Relying solely on standard federated averaging introduces operational and compliance risks when local user behavior varies widely. Organizations deploying privacy-preserving machine learning must balance model accuracy against hardware constraints, bandwidth limits, and potential privacy leakages associated with proxy data or explainability interfaces.

Decision-makers preparing to implement federated systems should evaluate participant hardware and data distributions before selecting a personalization framework. Knowledge distillation and parameter decoupling are best suited for heterogeneous edge environments with differing compute capacities, whereas clustering and regularized loss functions are preferable for institutional cross-silo settings with stable infrastructure. Prior to large-scale deployment, organizations should conduct multi-dimensional pilot tests that measure compute latency, communication payload sizes, and fairness across underrepresented user segments.

While the article provides high confidence in the algorithmic categorization and identified trade-offs, readers should note that existing empirical results largely stem from simulated data splits on benchmark vision and text tasks. Further validation on complex, non-stationary enterprise data streams is necessary to confirm long-term production stability.

  • Paper: Federated Learning on Non-IID Data: A Survey, Hangyu Zhu et al. (2021). This survey provides a comprehensive analysis of non-IID data distributions across parametric and non-parametric federated models, extending the personalization taxonomy into specific mitigation strategies.
  • Paper: Model-Contrastive Federated Learning, Qinbin Li et al. (2021). This work introduces model-contrastive learning (MOON) to correct local drift under heterogeneous data, presenting a concrete algorithmic strategy aligned with PFL principles.
  • Paper: Federated Learning on Non-IID Data Silos: An Experimental Study, Qinbin Li et al. (2021). It delivers an extensive experimental benchmarking study (NIID-Bench) across distinct non-IID partition types, putting personalized and heterogeneous federated learning techniques to practical test.
Cover for Towards Personalized Federated Learning

Abstract

In parallel with the rapid adoption of Artificial Intelligence (AI) empowered by advances in AI research, there have been growing awareness and concerns of data privacy. Recent significant developments in the data regulation landscape have prompted a seismic shift in interest towards privacy-preserving AI. This has contributed to the popularity of Federated Learning (FL), the leading paradigm for the training of machine learning models on data silos in a privacy-preserving manner. In this survey, we explore the domain of Personalized FL (PFL) to address the fundamental challenges of FL on heterogeneous data, a universal characteristic inherent in all real-world datasets. We analyze the key motivations for PFL and present a unique taxonomy of PFL techniques categorized according to the key challenges and personalization strategies in PFL. We highlight their key ideas, challenges and opportunities and envision promising future trajectories of research towards new PFL architectural design, realistic PFL benchmarking, and trustworthy PFL approaches.

Table of Contents

  • I Introduction
  • I-A Categorization of Federated Learning
  • I-B Motivations for Personalized Federated Learning
  • I-B1 Poor Convergence on Heterogeneous Data
  • I-B2 Lack of Solution Personalization
  • I-C Contributions
  • II Strategies for Personalized Federated Learning
  • III Strategy I: Global Model Personalization
  • III-A Data-based Approaches
  • III-B Model-based Approaches
  • III-B1 Between Global and Local Models
  • III-B2 Between Historical Local Model Snapshots
  • IV Strategy II: Learning Personalized Models
  • IV-A Architecture-based Approaches
  • IV-B Similarity-based Approaches
  • V PFL Benchmark & Evaluation Metrics
  • V-1 Quantity Skew
  • V-2 Feature Distribution Skew
  • V-3 Label Distribution Skew
  • V-4 Label Preference Skew
  • VI Promising Future Research Directions
  • VI-A Opportunities for PFL Architectural Design
  • VI-B Opportunities for PFL Benchmarking
  • VI-C Opportunities for Trustworthy PFL
  • VII Conclusions
  • References

Knowls

  1. Knowl 1 — Taxonomy of Personalized Federated Learning Strategies

    definition

    Personalized Federated Learning (PFL) methods are structured into a two-level hierarchical taxonomy based on the core challenge and personalization mechanism:

    1. Strategy I: Global Model Personalization (Data Heterogeneity Focus) Addresses client drift and poor convergence when training on non-independent and identically distributed (non-IID) data across distributed clients. It employs a two-stage paradigm ("FL training + local adaptation"), where a shared global model is trained and subsequently adapted to each client's local dataset. This strategy is subdivided into:

      • Data-based Approaches: Techniques that reduce statistical heterogeneity across client datasets prior to or during global model aggregation. Subcategories include Data Augmentation (synthetic data generation, GANs, feature-level oversampling) and Client Selection (reinforcement learning or bandit-based selection of balanced client subsets, tier-based scheduling).
      • Model-based Approaches: Techniques that train a robust global model initialization for downstream local adaptation or regularize local adaptation updates. Subcategories include Regularized Local Loss (proximal terms, parameter importance penalization, contrastive representation learning, variance reduction), Meta-Learning (optimization-based learning-to-learn formulations for fast few-step local adaptation), and Transfer Learning (domain adaptation, feature alignment layers).
    2. Strategy II: Learning Personalized Models (Solution Personalization Focus) Addresses the limitation of deploying a single global model for diverse local tasks by learning distinct, customized models for individual clients via modified aggregation protocols. This strategy is subdivided into:

      • Architecture-based Approaches: Tailor individual neural network architectures to client-specific capacities and tasks. Subcategories include Parameter Decoupling (splitting parameters into federated base layers and private local layers, or private user embeddings) and Knowledge Distillation (transferring logit/representation knowledge across heterogeneous student-teacher architectures to clients, to the server, bi-directionally, or peer-to-peer).
      • Similarity-based Approaches: Exploit explicit or implicit relationships across client data distributions to build customized models. Subcategories include Multi-Task Learning (joint optimization of client-specific tasks or attentional pairwise collaborations), Model Interpolation (combining local and global models via parameterized convex combinations), and Clustering (grouping clients into homogeneous sub-federations via post-hoc bi-partitioning, agglomerative clustering, iterative multi-model assignment, or distance-based multi-center optimization).
  2. Knowl 2 — Optimization Formulations of Standard FL, Local Learning, and PFL

    equation

    In horizontal federated learning with CC participating clients, the standard Federated Learning (FL) optimization problem minimizes a single global objective function:

    min⁡w∈RdF(w):=1C∑c=1Cfc(w)\min_{w \in \mathbb{R}^d} F(w) := \frac{1}{C} \sum_{c=1}^C f_c(w)

    where w∈Rdw \in \mathbb{R}^d denotes the shared global model parameters, and fc(w):=E(x,y)∼Dc[fc(w;x,y)]f_c(w) := \mathbb{E}_{(x,y) \sim \mathcal{D}_c}[f_c(w; x, y)] represents the expected empirical loss over the local private data distribution Dc\mathcal{D}_c of client cc.

    In contrast, purely local learning isolates each client without communication:

    min⁡θ1,…,θC∈RdF(θ):=1C∑c=1Cfc(θc)\min_{\theta_1, \dots, \theta_C \in \mathbb{R}^d} F(\theta) := \frac{1}{C} \sum_{c=1}^C f_c(\theta_c)

    where θc∈Rd\theta_c \in \mathbb{R}^d denotes the private model parameters of client cc.

    Personalized Federated Learning (PFL) lies between these two extremes. Standard FL provides high generalization via cross-client collaboration but fails to fit non-IID local distributions due to client drift. Local learning yields fully customized models but lacks generalization guarantees due to local data scarcity. PFL optimizes individualized client parameters θc\theta_c (or personalized adaptations of a global parameter vector ww) by balancing local task customization with collaborative knowledge sharing across the federation.

  3. Knowl 3 — Model-Based Global Model Personalization Formulations

    model/method

    Model-based Global Model Personalization optimizes the global model to facilitate rapid local adaptation or regularizes client updates during federated training:

    1. Regularized Local Loss Formulation Each client cc minimizes a composite local objective to mitigate client drift and representation divergence:

      min⁡θ∈Rdhc(θ;w):=fc(θ)+lreg(θ;w)\min_{\theta \in \mathbb{R}^d} h_c(\theta; w) := f_c(\theta) + l_{\text{reg}}(\theta; w)

      where fc(θ)f_c(\theta) is the empirical loss on client cc's data, ww is the current global model parameter vector, and lreg(θ;w)l_{\text{reg}}(\theta; w) is a regularization penalty. Key realizations include:

      • FedProx: Employs an ℓ2\ell_2 proximal term lreg(θ;w)=μ2∥θ−w∥2l_{\text{reg}}(\theta; w) = \frac{\mu}{2} \|\theta - w\|^2.
      • FedCL: Uses Elastic Weight Consolidation (EWC) with parameter importance estimated via the Fisher information matrix computed on a server proxy dataset.
      • SCAFFOLD: Adds a variance-reduction correction term based on the difference between global and local control variates (v−vc)(v - v_c).
      • MOON: Uses model-contrastive loss to minimize representation distance between the local model and the global model while maximizing distance to the previous local model snapshot.
    2. Federated Meta-Learning Formulation (Per-FedAvg) Rather than optimizing a global model for direct deployment, Per-FedAvg optimizes an initial model ww such that a small number of local gradient descent steps yield strong performance on any client's local distribution:

      min⁡w∈RdF(w):=1C∑c=1Cfc(w−α∇fc(w))\min_{w \in \mathbb{R}^d} F(w) := \frac{1}{C} \sum_{c=1}^C f_c\left(w - \alpha \nabla f_c(w)\right)

      where α>0\alpha > 0 is the local adaptation step size. To address the computational cost of computing the exact second-order gradients (Hessian-vector products), first-order MAML (FO-MAML, dropping the Hessian) or Hessian-free MAML (HF-MAML, approximating the Hessian-vector product via gradient differences) are used.

  4. Knowl 4 — Comparison of Global Model Personalization Techniques

    data/table

    Global model personalization techniques retain a single global model during federated training and subsequently adapt it locally. Their respective advantages and disadvantages are summarized below:

    Method Advantages Disadvantages
    Data Augmentation Easy to implement; builds directly on general FL training workflows Risks privacy leakage from sharing data samples or statistical distributions; often requires a representative proxy dataset
    Client Selection Only modifies client sampling in standard FL aggregation rounds Computational overhead from reinforcement learning/bandit subset optimization; may require server-side proxy datasets
    Regularization Straightforward implementation via slight modifications to local loss functions in FedAvg Constrained to a single global model architecture; does not personalize model capacity
    Meta-Learning Explicitly optimizes the global parameter initialization for rapid few-step local adaptation Computationally expensive due to second-order gradient calculations; constrained to a single global model setup
    Transfer Learning Reduces domain discrepancy between global model features and target client data via adaptation layers Constrained to a single global model setup; lower layers remain rigid

    These approaches retain the assumption that all clients share a single neural network architecture with the FL parameter server, limiting their application when edge devices have heterogeneous compute and memory capacities.

  5. Knowl 5 — Architecture-Based Personalized Federated Learning

    model/method

    Architecture-based PFL methods enable heterogeneous or customized model structures across clients using two primary mechanisms:

    1. Parameter Decoupling Decouples private local model parameters from federated shared parameters. Private parameters are trained locally and never transmitted to the FL server, allowing the model to fit client-specific representations:

      • Layer-wise Partitioning: The network is divided into shared base layers (lower layers that extract low-level, generic feature representations across the federation) and private personalized layers (upper classification layers retained on each client).
      • Representation Decoupling (e.g., LG-FedAvg): Private user-specific embeddings or specialized local encoders are learned locally to capture domain-specific feature distributions (e.g., text, images), while shared feature extractors/classifiers are aggregated across clients. Fair, unbiased representations can also be enforced via local adversarial objectives.
      • Distinction from Split Learning (SL): While both decouple layers, SL splits networks sequentially between client and server (transmitting activations and gradients layer-wise), whereas parameter decoupling performs parallel client-side training where base weights are aggregated at the server and personal weights remain fully local.
    2. Knowledge Distillation (KD) Enables clients with diverse computational capacities to train different model architectures (e.g., ResNet, MobileNet) by exchanging knowledge via soft label predictions (logits) or representations rather than raw parameter weights across four communication topologies:

      • Distillation to clients (e.g., FedGen): A server-side generator distills inductive bias to regularize client feature representations without proxy datasets.
      • Distillation to server (e.g., FedDF): Client prototype models produce ensemble logit outputs on an unlabelled public dataset to train a server-side student model.
      • Bi-directional distillation (e.g., FedGKT): Small client models send extracted features/logits to a large server model, and the server distills predictions back via Kullback-Leibler (KL) divergence loss, shifting computational burden away from edge devices.
      • Peer-to-peer/distributed distillation (e.g., D-Distillation): Edge devices broadcast soft decisions to local neighbors, reaching consensus to regularize local loss without a central parameter server.
  6. Knowl 6 — Similarity-Based Personalized Federated Learning Formulations

    model/method

    Similarity-based PFL models inter-client relationships to learn distinct personalized models for related clients through three frameworks:

    1. Multi-Task Learning (MTL) and Attentive Collaboration (FedAMP) Treats each client as an individual task. In FedAMP, the parameter server maintains a personalized cloud model ucu_c for each client cc, computed as an attention-weighted linear combination of client models: uc=∑m∈Cξc,mθm,subject to ∑m∈Cξc,m=1u_c = \sum_{m \in \mathcal{C}} \xi_{c,m} \theta_m, \quad \text{subject to } \sum_{m \in \mathcal{C}} \xi_{c,m} = 1 Client cc then optimizes its local parameter vector θc\theta_c using an attentional proximal objective: θc∗=arg⁡min⁡θ∈Rdfc(θ)+μ2α∥θ−uc∥2\theta_c^* = \arg\min_{\theta \in \mathbb{R}^d} f_c(\theta) + \frac{\mu}{2\alpha} \|\theta - u_c\|^2 where α\alpha is the gradient step size, μ\mu is a regularization coefficient, and ξc,m\xi_{c,m} reflects data similarity between clients cc and mm.

    2. Model Interpolation Balances global generalization and client personalization by learning a client-specific convex combination of local and global models. A mixing coefficient λc∈[0,1]\lambda_c \in [0, 1] is adaptively optimized: θcpersonalized=λcθc+(1−λc)w\theta_c^{\text{personalized}} = \lambda_c \theta_c + (1 - \lambda_c) w where higher λc\lambda_c assigns more weight to the local model when client cc's local distribution diverges strongly from the global average.

    3. Clustering-Based Federated Learning (Multi-Center Loss) Partitions the federation into KK clusters such that clients within cluster kk collaboratively learn a cluster model w(k)w^{(k)}. The multi-center clustering objective is: ℓ=1C∑k=1K∑c=1Crc(k)Dist(θc,w(k))\ell = \frac{1}{C} \sum_{k=1}^K \sum_{c=1}^C r_c^{(k)} \text{Dist}\left(\theta_c, w^{(k)}\right) where rc(k)∈{0,1}r_c^{(k)} \in \{0, 1\} is a binary cluster assignment indicator (rc(k)=1r_c^{(k)} = 1 if k=arg⁡min⁡jDist(θc,w(j))k = \arg\min_j \text{Dist}(\theta_c, w^{(j)}) and 00 otherwise), solved via Expectation-Maximization (EM) or iterative federated clustering (IFCA).

  7. Knowl 7 — Comparison of Learning Personalized Models Techniques

    data/table

    Techniques under Strategy II directly optimize individualized models by modifying the aggregation process. Their comparative trade-offs are summarized below:

    Method Advantages Disadvantages
    Parameter Decoupling Simple formulation; allows layer-wise architectural flexibility per client Research challenge to determine the optimal boundary between private and federated layers
    Knowledge Distillation High architectural flexibility; communication-efficient; handles hardware/resource heterogeneity Difficult to design optimal student-teacher capacity gaps; often relies on unlabelled proxy datasets
    Multi-Task Learning Captures pairwise client relationships to co-train models across related clients Computationally intensive for large networks; sensitive to clients with corrupted or poor-quality data
    Model Interpolation Simple formulation balancing local and global model mixtures via adaptive weights Relies on a single global model as the anchor; degrades under severe non-IID conditions
    Clustering Naturally fits scenarios with inherent multi-modal population distributions High communication/computation costs in cluster discovery; requires extra cluster management infrastructure
  8. Knowl 8 — Taxonomy and Simulation of Non-IID Statistical Heterogeneity

    definition

    In PFL experimental evaluation, non-IID data distributions among clients are categorized into four distinct statistical skew types:

    1. Quantity Skew Local dataset sizes ∣Dc∣|\mathcal{D}_c| vary substantially across clients c∈{1,…,C}c \in \{1, \dots, C\}, while the underlying data distribution remains identical. Simulated in benchmarks by assigning sample counts according to a power-law distribution.

    2. Feature Distribution Skew (Covariate Shift) The marginal feature distribution Pc(x)P_c(x) varies across clients, but the conditional label distribution P(y∣x)P(y|x) is shared (Pc(y∣x)=P(y∣x)P_c(y|x) = P(y|x) for all cc). Simulated by partitioning data according to natural user identifiers (e.g., sensor/health monitoring users) or by applying client-specific feature transformations such as image rotations.

    3. Label Distribution Skew (Prior Probability Shift) The marginal label distribution Pc(y)P_c(y) varies across clients, but the conditional feature distribution P(x∣y)P(x|y) remains invariant (Pc(x∣y)=P(x∣y)P_c(x|y) = P(x|y) for all cc). Simulated by:

      • Allocating a fixed small number of classes kk to each client.
      • Drawing class allocation probabilities from a Dirichlet distribution Dir(α)\text{Dir}(\alpha), where α→0\alpha \to 0 generates extreme non-IID label skew (each client holds single-class data) and α→∞\alpha \to \infty (e.g., α=100\alpha = 100) approaches an IID allocation.
    4. Label Preference Skew (Concept Shift) The conditional distribution Pc(x∣y)P_c(x|y) varies across clients while the marginal label distribution P(y)P(y) is identical across clients. Simulated by swapping a fraction of ground-truth labels on specific subsets of clients to model subjective label divergence.

  9. Knowl 9 — Three-Dimensional Evaluation Metric Taxonomy for PFL Systems

    definition

    PFL benchmarking encompasses three core metric dimensions:

    1. Model Performance Metrics

      • Accuracy: Assessed via the mean test accuracy of personalized local models across clients, accuracy distributions (histogram profiling, variance metrics), and client-level accuracy delta (comparing local model accuracy before and after personalization).
      • Convergence: Measured by local training loss curves, the number of communication rounds to reach target accuracy, local epochs per round, and theoretical asymptotic convergence bounds.
    2. System Performance Metrics

      • Communication Efficiency: Quantified by total communication rounds, total transmitted parameters per round, and compressed message payload sizes.
      • Computational Efficiency: Evaluated via client-side Floating Point Operations (FLOPs) and client/server execution runtime.
      • System Heterogeneity & Scalability: Tested by simulating non-uniform local training epochs, variable client CPU capacities, diverse model architectures, total elapsed time, and memory footprints under large client populations (C≫100C \gg 100).
      • Fault Tolerance: Robustness to client dropout ratios and straggler delays during aggregation.
    3. Trustworthy AI Metrics

      • Robustness: Model performance under adversarial data poisoning, model poisoning, and inference privacy attacks.
      • Fairness: Distribution of performance across clients, measuring whether PFL disproportionately degrades accuracy for minority clients or low-resource devices.
      • Explainability: Interpretability of client-specific predictions while preventing inadvertent gradient-based data leakage.
  10. Knowl 10 — Open Challenges and Future Trajectories in Personalized Federated Learning

    limitation

    Several open research challenges limit the practical deployment of PFL systems:

    1. Privacy-Preserving Data Heterogeneity Analytics Calculating statistical distance metrics (e.g., Earth Mover's Distance, 1-Wasserstein, Total Variation) to quantify local data non-IID-ness currently requires access to raw client data distributions. Measuring statistical heterogeneity in a strictly privacy-preserving manner remains an open problem.

    2. Spatial and Temporal Adaptability

      • Spatial Adaptability (Cold-Start & Catastrophic Forgetting): Existing PFL algorithms generally assume a fixed client set during training. When new clients join midway, PFL models face cold-start issues, and integrating new clients can induce catastrophic forgetting of existing representations due to the stability-plasticity dilemma.
      • Temporal Adaptability (Concept Drift): Real-world client data streams are non-stationary. Incorporating automated drift detection (e.g., Change Detection Techniques), drift understanding, and adaptive local updating into PFL remains underdeveloped.
    3. Trustworthy PFL Trade-offs

      • Explainability vs. Privacy: Generating interpretable representations for personalized models often relies on gradient or feature attribution methods that risk leaking private local data.
      • Incentive Mechanism Design: Sustaining open collaboration across self-interested data owners requires game-theoretic and auction-based incentive schemes to reward high-quality personalization contributions.

Coverage note — None was omitted. The knowls cover all core contributed aspects: the hierarchical PFL taxonomy, mathematical problem formulations, Strategy I (data- and model-based), Strategy II (architecture- and similarity-based), benchmark skews, evaluation metrics, and open challenges.

References

  1. 1.G. A. Kaissis, M. R. Makowski, D. Ruckert, and R. F. Braren, ¨ “Secure, privacy-preserving and federated machine learning in medical imaging,” Nat. Mach. Intell., vol. 2, no. 6, pp. 305–311, 2020.
  2. 2.S. Warnat-Herresthal, H. Schultze, K. L. Shastry, S. Manamohan, S. Mukherjee, V. Garg, R. Sarveswara, K. Handler, P. Pickkers, N. A. ¨ Aziz et al., “Swarm learning for decentralized and confidential clinical machine learning,” Nature, vol. 594, no. 7862, pp. 265–270, 2021.
  3. 3.M. J. Sheller, B. Edwards, G. A. Reina, J. Martin, S. Pati, A. Kotrotsou, M. Milchenko, W. Xu, D. Marcus, R. R. Colen et al., “Federated learn- ing in medicine: facilitating multi-institutional collaborations without sharing patient data,” Sci. Rep., vol. 10, no. 1, pp. 1–12, 2020.
  4. 4.I. Dayan, H. R. Roth, A. Zhong, A. Harouni, A. Gentili, A. Z. Abidin, A. Liu, A. B. Costa, B. J. Wood, C.-S. Tsai et al., “Federated learning for predicting clinical outcomes in patients with covid-19,” Nat. Med., pp. 1–9, 2021.
  5. 5.P. Voigt and A. von dem Bussche, The EU General Data Protection Regulation (GDPR). Springer International Publishing, 2017.
  6. 6.Y. Cheng, Y. Liu, T. Chen, and Q. Yang, “Federated learning for privacy-preserving ai,” Commun. ACM, vol. 63, no. 12, pp. 33–36, 2020.
  7. 7.Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated Machine Learning: Concept and Applications,” ACM TIST, vol. 10, no. 2, pp. 1–19, 2019.
  8. 8.P. Kairouz, H. B. McMahan, and B. Avent et al., “Advances and Open Problems in Federated Learning,” Found. Trends Mach. Learn., vol. 14, no. 1-2, pp. 1–210, 2021.
  9. 9.H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-Efficient Learning of Deep Networks from Decentralized Data,” in AISTATS, 2017, pp. 1273–1282.
  10. 10.G. Drainakis, K. V. Katsaros, P. Pantazopoulos, V. Sourlas, and A. Amditis, “Federated vs. centralized machine learning under privacy- elastic users: A comparative analysis,” in IEEE NCA, 2020, pp. 1–8.
  11. 11.W. Y. B. Lim, N. C. Luong, D. T. Hoang, Y. Jiao, Y.-C. Liang, Q. Yang, D. Niyato, and C. Miao, “Federated learning in mobile edge networks: A comprehensive survey,” IEEE Commun. Surv. Tutor., vol. 22, no. 3, pp. 2031–2063, 2020.
  12. 12.S. P. Karimireddy, S. Kale, M. Mohri, S. J. Reddi, S. U. Stich, and A. T. Suresh, “SCAFFOLD: Stochastic Controlled Averaging for Federated Learning,” in ICML, 2020, pp. 5132–5143.
  13. 13.X. Li, K. Huang, W. Yang, S. Wang, and Z. Zhang, “On the Conver- gence of FedAvg on Non-IID Data,” in ICLR, 2020.
  14. 14.T. Li, A. K. Sahu, A. Talwalkar, and V. Smith, “Federated Learning: Challenges, Methods, and Future Directions,” IEEE Signal Process. Mag., vol. 37, no. 3, pp. 50–60, 2020.
  15. 15.V. Mothukuri, R. M. Parizi, S. Pouriyeh, Y. Huang, A. Dehghantanha, and G. Srivastava, “A survey on security and privacy of federated learning,” FGCS, vol. 115, pp. 619–640, 2021.
  16. 16.L. Lyu, H. Yu, X. Ma, L. Sun, J. Zhao, Q. Yang, and P. S. Yu, “Privacy and robustness in federated learning: Attacks and defenses,” arXiv:2012.06337, 2021.
  17. 17.Y. Mansour, M. Mohri, J. Ro, and A. T. Suresh, “Three Ap- proaches for Personalization with Applications to Federated Learning,” arXiv:2002.10619, 2020.
  18. 18.N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” JAIR, vol. 16, pp. 321–357, 2002.
  19. 19.H. He, Y. Bai, E. A. Garcia, and S. Li, “ADASYN: Adaptive synthetic sampling approach for imbalanced learning,” in IJCNN, 2008, pp. 1322–1328.
  20. 20.M. Kubat and S. Matwin, “Addressing the Curse of Imbalanced Training Sets: One-Sided Selection,” in ICML, 1997, pp. 179–186.
  21. 21.Y. Zhao, M. Li, L. Lai, N. Suda, D. Civin, and V. Chandra, “Federated Learning with Non-IID Data,” arXiv:1806.00582, 2018.
  22. 22.E. Jeong, S. Oh, H. Kim, J. Park, M. Bennis, and S.-L. Kim, “Communication-efficient on-device machine learning: Fed- erated distillation and augmentation under non-iid private data,” arXiv:1811.11479, 2018.
  23. 23.M. Duan, D. Liu, X. Chen, R. Liu, Y. Tan, and L. Liang, “Self- Balancing Federated Learning With Global Imbalanced Data in Mobile Systems,” IEEE TPDS, vol. 32, no. 1, pp. 59–71, 2021.
  24. 24.Q. Wu, X. Chen, Z. Zhou, and J. Zhang, “Fedhome: Cloud-edge based personalized federated learning for in-home health monitoring,” IEEE TMC, 2020.
  25. 25.H. Wang, Z. Kaplan, D. Niu, and B. Li, “Optimizing Federated Learning on Non-IID Data with Reinforcement Learning,” in IEEE INFOCOM, 2020, pp. 1698–1707.
  26. 26.M. Yang, X. Wang, H. Zhu, H. Wang, and H. Qian, “Federated learning with class imbalance reduction,” in IEEE EUSIPCO, 2021, pp. 2174– 2178.
  27. 27.Z. Chai, A. Ali, S. Zawad, S. Truex, A. Anwar, N. Baracaldo, Y. Zhou, H. Ludwig, F. Yan, and Y. Cheng, “Tifl: A tier-based federated learning system,” in ACM HPDC, 2020, pp. 125–136.
  28. 28.L. Li, M. Duan, D. Liu, Y. Zhang, A. Ren, X. Chen, Y. Tan, and C. Wang, “Fedsae: A novel self-adaptive federated learning framework in heterogeneous systems,” in IJCNN, 2021.
  29. 29.T. Li, A. K. Sahu, M. Zaheer, and et al., “Federated Optimization in Heterogeneous Networks,” MLSys, vol. 2, pp. 429–450, 2020.
  30. 30.X. Yao and L. Sun, “Continual Local Training for Better Initialization of Federated Models,” in IEEE ICIP, 2020, pp. 1736–1740.
  31. 31.J. Kirkpatrick, R. Pascanu, and N. Rabinowitz et al., “Overcoming catastrophic forgetting in neural networks,” PNAS, vol. 114, no. 13, pp. 3521–3526, 2017.
  32. 32.Q. Li, B. He, and D. Song, “Model-contrastive federated learning,” in CVPR, 2021, pp. 10 713–10 722.
  33. 33.T. Hospedales, A. Antoniou, P. Micaelli, and A. Storkey, “Meta- learning in neural networks: A survey,” IEEE TPAMI, no. 01, pp. 1–1, 2020.
  34. 34.C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in ICML, 2017, pp. 1126–1135.
  35. 35.A. Nichol, J. Achiam, and J. Schulman, “On First-Order Meta-Learning Algorithms,” arXiv:1803.02999, 2018.
  36. 36.Y. Jiang, J. Konecnˇ y, K. Rush, and S. Kannan, “Improving Feder- ´ ated Learning Personalization via Model Agnostic Meta Learning,” arXiv:1909.12488, 2019.
  37. 37.A. Fallah, A. Mokhtari, and A. Ozdaglar, “Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach,” in NeurIPS, vol. 33, 2020, pp. 3557–3568.
  38. 38.——, “On the convergence theory of gradient-based model-agnostic meta-learning algorithms,” in AISTATS, 2020, pp. 1082–1092.
  39. 39.C. T. Dinh, N. H. Tran, and T. D. Nguyen, “Personalized federated learning with moreau envelopes,” in NeurIPS, vol. 33, 2020, pp. 21 394–21 405.
  40. 40.M. Khodak, M.-F. Balcan, and A. Talwalkar, “Adaptive Gradient-Based Meta-Learning Methods,” in NeurIPS, vol. 32, 2019, pp. 5917–5928.
  41. 41.S. J. Pan and Q. Yang, “A survey on transfer learning,” IEEE TKDE, vol. 22, no. 10, pp. 1345–1359, 2009.
  42. 42.D. Li and J. Wang, “Fedmd: Heterogenous federated learning via model distillation,” arXiv:1910.03581, 2019.
  43. 43.Y. Chen, X. Qin, J. Wang, C. Yu, and W. Gao, “Fedhealth: A federated transfer learning framework for wearable healthcare,” IEEE Intell. Syst., vol. 35, no. 4, pp. 83–93, 2020.
  44. 44.H. Yang, H. He, W. Zhang, and X. Cao, “FedSteg: A Federated Transfer Learning Framework for Secure Image Steganalysis,” IEEE TNSE, vol. 8, no. 2, pp. 1084–1094, 2020.
  45. 45.B. Sun, J. Feng, and K. Saenko, “Return of frustratingly easy domain adaptation,” in AAAI, 2016, pp. 2058–2065.
  46. 46.M. G. Arivazhagan, V. Aggarwal, A. K. Singh, and S. Choudhary, “Federated Learning with Personalization Layers,” arXiv:1912.00818, 2019.
  47. 47.D. Bui, K. Malik, J. Goetz, H. Liu, S. Moon, A. Kumar, and K. G. Shin, “Federated User Representation Learning,” arXiv:1909.12535, 2019.
  48. 48.P. P. Liang, T. Liu, and L. Ziyin et al., “Think locally, act globally: Federated learning with local and global representations,” arXiv:2001.01523, 2019.
  49. 49.O. Gupta and R. Raskar, “Distributed learning of deep neural network over multiple agents,” J. Netw. Comput. Appl., vol. 116, pp. 1–8, 2018.
  50. 50.P. Vepakomma, O. Gupta, T. Swedish, and R. Raskar, “Split learning for health: Distributed deep learning without sharing raw patient data,” arXiv:1812.00564, 2018.
  51. 51.C. Thapa, M. A. P. Chamikara, S. Camtepe, and L. Sun, “Splitfed: When federated learning meets split learning,” arXiv:2004.12088, 2020.
  52. 52.Y. Gao, M. Kim, S. Abuadbba, Y. Kim, C. Thapa, K. Kim, S. A. Camtep, H. Kim, and S. Nepal, “End-to-end evaluation of federated learning and split learning for internet of things,” in IEEE SRDS, 2020, pp. 91–100.
  53. 53.Q. Yang, Y. Liu, T. Chen, and Y. Tong, “Federated machine learning: Concept and applications,” ACM TIST, vol. 10, no. 2, pp. 1–19, 2019.
  54. 54.G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv:1503.02531, 2015.
  55. 55.Z. Zhu, J. Hong, and J. Zhou, “Data-free knowledge distillation for heterogeneous federated learning,” in ICML, 2021.
  56. 56.T. Lin, L. Kong, S. U. Stich, and M. Jaggi, “Ensemble distillation for robust model fusion in federated learning,” in NeurIPS, vol. 33, 2020, pp. 2351–2363.
  57. 57.C. He, M. Annavaram, and S. Avestimehr, “Group knowledge transfer: Federated learning of large cnns at the edge,” in NeurIPS, vol. 33, 2020, pp. 14 068–14 080.
  58. 58.I. Bistritz, A. Mann, and N. Bambos, “Distributed distillation for on- device learning,” in NeurIPS, vol. 33, 2020, pp. 22 593–22 604.
  59. 59.V. Smith, C.-K. Chiang, M. Sanjabi, and A. Talwalkar, “Federated Multi-Task Learning,” in NeurIPS, vol. 30, 2017, pp. 4427–4437.
  60. 60.L. Corinzia and J. M. Buhmann, “Variational Federated Multi-Task Learning,” arXiv:1906.06268, 2019.
  61. 61.Y. Huang, L. Chu, Z. Zhou, L. Wang, J. Liu, J. Pei, and Y. Zhang, “Personalized cross-silo federated learning on non-iid data,” in AAAI, vol. 35, no. 9, 2021, pp. 7865–7873.
  62. 62.N. Shoham, T. Avidor, A. Keren, N. Israel, D. Benditkis, L. Mor-Yosef, and I. Zeitak, “Overcoming forgetting in federated learning on non-iid data,” arXiv:1910.07796, 2019.
  63. 63.F. Hanzely and P. Richtarik, “Federated Learning of a Mixture of ´ Global and Local Models,” arXiv:2002.05516, 2020.
  64. 64.Y. Deng, M. M. Kamani, and M. Mahdavi, “Adaptive Personalized Federated Learning,” arXiv:2003.13461, 2020.
  65. 65.E. Diao, J. Ding, and V. Tarokh, “Heterofl: Computation and communi- cation efficient federated learning for heterogeneous clients,” in ICLR, 2021.
  66. 66.F. Sattler, K.-R. Muller, and W. Samek, “Clustered federated learn- ¨ ing: Model-agnostic distributed multitask optimization under privacy constraints,” IEEE TNNLS, vol. 32, no. 8, pp. 3710–3722, 2020.
  67. 67.C. Briggs, Z. Fan, and P. Andras, “Federated learning with hierarchical clustering of local updates to improve training on non-IID data,” in IJCNN, 2020, pp. 1–9.
  68. 68.A. Ghosh, J. Chung, D. Yin, and K. Ramchandran, “An efficient framework for clustered federated learning,” in NeurIPS, vol. 33, 2020, pp. 19 586–19 597.
  69. 69.L. Huang, A. L. Shea, H. Qian, A. Masurkar, H. Deng, and D. Liu, “Patient clustering improves efficiency of federated machine learning to predict mortality and hospital stay time using distributed electronic medical records,” J. Biomed. Inform., vol. 99, p. 103291, 2019.
  70. 70.M. Duan, D. Liu, X. Ji, R. Liu, L. Liang, X. Chen, and Y. Tan, “Fedgroup: Efficient federated learning via decomposed similarity- based clustering,” in IEEE ISPA, 2021, pp. 228–237.
  71. 71.S. Vassilvitskii and D. Arthur, “k-means++: The advantages of careful seeding,” in ACM-SIAM, 2006, pp. 1027–1035.
  72. 72.M. Xie, G. Long, T. Shen, T. Zhou, X. Wang, and J. Jiang, “Multi- Center Federated Learning,” arXiv:2005.01026, 2020.
  73. 73.Y. Liu, X. Jia, M. Tan, R. Vemulapalli, Y. Zhu, B. Green, and X. Wang, “Search to distill: Pearls are everywhere but not the eyes,” in CVPR, 2020, pp. 7539–7548.
  74. 74.C. Li, J. Peng, L. Yuan, G. Wang, X. Liang, L. Lin, and X. Chang, “Block-wisely supervised neural architecture search with knowledge distillation,” in CVPR, 2020, pp. 1989–1998.
  75. 75.Y. Liang, Y. Guo, Y. Gong, C. Luo, J. Zhan, and Y. Huang, “Flbench: A benchmark suite for federated learning,” in FICC, 2020, pp. 166–176.
  76. 76.T. Hao, Y. Huang, X. Wen, W. Gao, F. Zhang, C. Zheng, L. Wang, H. Ye, K. Hwang, Z. Ren et al., “Edge aibench: towards comprehensive end-to-end edge computing benchmarking,” in Bench, 2018, pp. 23–30.
  77. 77.S. Hu, Y. Li, X. Liu, Q. Li, Z. Wu, and B. He, “The oarf bench- mark suite: Characterization and implications for federated learning systems,” arXiv:2006.07856, 2020.
  78. 78.C. He, K. Balasubramanian, E. Ceyani, C. Yang, H. Xie, L. Sun, L. He, L. Yang, P. S. Yu, Y. Rong et al., “Fedgraphnn: A federated learning system and benchmark for graph neural networks,” arXiv:2104.07145, 2021.
  79. 79.S. Caldas, S. M. K. Duddu, P. Wu, T. Li, J. Konecnˇ y, H. B. McMahan, ´ V. Smith, and A. Talwalkar, “LEAF: A Benchmark for Federated Settings,” arXiv:1812.01097, 2019.
  80. 80.G. Cohen, S. Afshar, J. Tapson, and A. Van Schaik, “Emnist: Extending mnist to handwritten letters,” in IJCNN, 2017, pp. 2921–2926.
  81. 81.Z. Liu, P. Luo, X. Wang, and X. Tang, “Deep learning face attributes in the wild,” in ICCV, 2015, pp. 3730–3738.
  82. 82.J. Luo, X. Wu, Y. Luo, A. Huang, Y. Huang, Y. Liu, and Q. Yang, “Real-world image datasets for federated learning,” arXiv:1910.11089, 2019.
  83. 83.T.-M. H. Hsu, H. Qi, and M. Brown, “Federated visual classification with real-world data distribution,” in ECCV, 2020, pp. 76–92.
  84. 84.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proc. IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
  85. 85.A. Krizhevsky, “Learning multiple layers of features from tiny images,” MIT and NYU, Tech. Rep., 2009.
  86. 86.T.-M. H. Hsu, H. Qi, and M. Brown, “Measuring the effects of non-identical data distribution for federated visual classification,” arXiv:1909.06335, 2019.
  87. 87.K. Wang, R. Mathews, C. Kiddon, H. Eichner, F. Beaufays, and D. Ramage, “Federated evaluation of on-device personalization,” arXiv:1910.10252, 2019.
  88. 88.T. Li, S. Hu, A. Beirami, and V. Smith, “Ditto: Fair and robust federated learning through personalization,” in ICML, 2021, pp. 6357–6368.
  89. 89.S. Divi, Y.-S. Lin, H. Farrukh, and Z. B. Celik, “New metrics to evaluate the performance and fairness of personalized federated learning,” arXiv:2107.13173, 2021.
  90. 90.D. Sui, Y. Chen, J. Zhao, Y. Jia, Y. Xie, and W. Sun, “Feded: Federated learning via ensemble distillation for medical relation extraction,” in EMNLP, 2020, pp. 2118–2128.
  91. 91.P. Xiao, S. Cheng, V. Stankovic, and D. Vukobratovic, “Averaging is probably not the optimum way of aggregating parameters in federated learning,” Entropy, vol. 22, no. 3, p. 314, 2020.
  92. 92.H. Wang, M. Yurochkin, Y. Sun, D. Papailiopoulos, and Y. Khazaeni, “Federated learning with matched averaging,” in ICLR, 2020.
  93. 93.H. Zhu, H. Zhang, and Y. Jin, “From federated learning to federated neural architecture search: a survey,” Complex Intell. Syst., vol. 7, no. 2, pp. 639–657, 2021.
  94. 94.Q. Wu, K. He, and X. Chen, “Personalized federated learning for intelligent IoT applications: A cloud-edge based framework,” IEEE OJ- CS, vol. 1, pp. 35–44, 2020.
  95. 95.R. Kemker, M. McClure, A. Abitino, T. Hayes, and C. Kanan, “Mea- suring catastrophic forgetting in neural networks,” in AAAI, 2018, pp. 3390–3398.
  96. 96.M. Delange, R. Aljundi, M. Masana, S. Parisot, X. Jia, A. Leonardis, G. Slabaugh, and T. Tuytelaars, “A continual learning survey: Defying forgetting in classification tasks,” IEEE TPAMI, 2021.
  97. 97.Y. Lin, S. Han, H. Mao, Y. Wang, and W. Dally, “Deep Gradient Compression: Reducing the communication bandwidth for distributed training,” in ICLR, 2018.
  98. 98.Y. Chen, X. Sun, and Y. Jin, “Communication-efficient federated deep learning with layerwise asynchronous model update and temporally weighted aggregation,” IEEE TNNLS, vol. 31, no. 10, pp. 4229–4238, 2019.
  99. 99.J. Lu, A. Liu, F. Dong, F. Gu, J. Gama, and G. Zhang, “Learning under concept drift: A review,” IEEE TKDE, vol. 31, no. 12, pp. 2346–2363, 2018.
  100. 100.F. E. Casado, D. Lema, R. Iglesias, C. V. Regueiro, and S. Barro, “Concept drift detection and adaptation for robotics and mobile devices in federated and continual settings,” in WAF, 2021, pp. 79–93.
  101. 101.S. Zheng, Y. Cao, and M. Yoshikawa, “Incentive mechanism for privacy-preserving federated learning,” arXiv:2106.04384, 2021.
  102. 102.Y. Zhan, J. Zhang, Z. Hong, L. Wu, P. Li, and S. Guo, “A survey of incentive mechanism design for federated learning,” IEEE Trans. Emerg. Topics Comput., 2021.
  103. 103.N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM CSUR, vol. 54, no. 6, pp. 1–35, 2021.
  104. 104.K. Holstein, J. W. Vaughan, H. Daume III, M. Dud ´ ´ık, and H. Wallach, “Improving fairness in machine learning systems: What do industry practitioners need?” in CHI, 2019, pp. 1–16.
  105. 105.M. Mohri, G. Sivek, and A. T. Suresh, “Agnostic Federated Learning,” in ICML, 2019, pp. 4615–4625.
  106. 106.T. Li, M. Sanjabi, A. Beirami, and V. Smith, “Fair resource allocation in federated learning,” in ICLR, 2020.
  107. 107.J. Zhang, C. Li, A. Robles-Kelly, and M. Kankanhalli, “Hierarchically Fair Federated Learning,” arXiv:2004.10386, 2020.
  108. 108.L. Lyu, J. Yu, K. Nandakumar, Y. Li, X. Ma, J. Jin, H. Yu, and K. S. Ng, “Towards fair and privacy-preserving federated deep models,” IEEE TPDS, vol. 31, no. 11, pp. 2524–2541, 2020.
  109. 109.A. Adadi and M. Berrada, “Peeking inside the black-box: A survey on explainable artificial intelligence (XAI),” IEEE Access, vol. 6, pp. 52 138–52 160, 2018.
  110. 110.A. B. Arrieta, N. D´ıaz-Rodr´ıguez, and J. Del Ser et al., “Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai,” Inf. Fusion, vol. 58, pp. 82–115, 2020.
  111. 111.S. Tonekaboni, S. Joshi, M. D. McCradden, and A. Goldenberg, “What clinicians want: contextualizing explainable machine learning for clinical end use,” in MLHC, 2019, pp. 359–380.
  112. 112.R. Shokri, M. Strobel, and Y. Zick, “On the privacy risks of model explanations,” in AIES, 2021, pp. 231–241.

Citation

MLA
Tan, A. Z., et al. “Towards Personalized Federated Learning”. IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 12, 2023, pp. 9587–603, https://doi.org/10.1109/TNNLS.2022.3160699.
APA
Tan, A. Z., Yu, H., Cui, L., & Yang, Q. (2023). Towards Personalized Federated Learning. IEEE Transactions on Neural Networks and Learning Systems, 34(12), 9587–9603. https://doi.org/10.1109/TNNLS.2022.3160699
Chicago
Tan, A. Z., H. Yu, L. Cui, and Q. Yang. 2023. “Towards Personalized Federated Learning”. IEEE Transactions on Neural Networks and Learning Systems 34 (12): 9587–9603. https://doi.org/10.1109/TNNLS.2022.3160699.
Harvard
Tan, A.Z. et al. (2023) “Towards Personalized Federated Learning”, IEEE Transactions on Neural Networks and Learning Systems, 34(12), pp. 9587–9603. Available at: https://doi.org/10.1109/TNNLS.2022.3160699.
Vancouver
1. Tan AZ, Yu H, Cui L, Yang Q (2023) Towards Personalized Federated Learning. IEEE Transactions on Neural Networks and Learning Systems 34:9587–9603

BibTeX

@article{Tan_2023, title={Towards Personalized Federated Learning}, volume={34}, ISSN={2162-2388}, url={http://dx.doi.org/10.1109/TNNLS.2022.3160699}, DOI={10.1109/tnnls.2022.3160699}, number={12}, journal={IEEE Transactions on Neural Networks and Learning Systems}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Tan, Alysa Ziying and Yu, Han and Cui, Lizhen and Yang, Qiang}, year={2023}, month=Dec, pages={9587–9603} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/