Personalized Federated Learning via Variational Bayesian Inference
Xu ZhangYinchuan LiWenpeng LiKaiyang GuoYunfeng Shao
Develops pFedBayes, a personalized federated learning framework that uses variational Bayesian neural networks to prevent client overfitting on non-i.i.d. data while guaranteeing minimax-optimal generalization error bounds.
Modern machine learning applications increasingly rely on federated learning to train artificial intelligence models across distributed devices while keeping private data localized. However, real-world deployments face two major challenges: private data across individual devices is statistically diverse and non-identical, and local data volumes are often too small to train complex neural networks effectively. These constraints frequently result in model overfitting and severe performance degradation, particularly in high-stakes fields such as finance, healthcare, and autonomous systems.
The article addresses these dual challenges by developing and evaluating a personalized federated learning framework named pFedBayes. The main objective is to demonstrate that integrating Bayesian neural networks—where model weights are treated as probability distributions rather than fixed values—into a two-level optimization process can simultaneously prevent overfitting on limited data and achieve tailored local personalization with quantifiable prediction uncertainty.
To establish this framework, the authors formulated a two-level variational Bayesian inference model where local clients use the aggregated global distribution as an informed prior rather than relying on arbitrary assumptions. The method was rigorously tested through theoretical analysis of generalization error bounds and empirical simulations across standard benchmark image datasets (MNIST, Fashion-MNIST, and CIFAR-10). The experimental setup evaluated performance across 10 to 20 clients under varying data volumes (small, medium, and large) and compared pFedBayes against seven state-of-the-art global and personalized federated learning baselines.
The findings show that pFedBayes consistently outperforms existing baseline methods, especially when local data is scarce. On the complex CIFAR-10 dataset with small sample sizes, pFedBayes surpassed other state-of-the-art algorithms by 11.71% in personalized model accuracy and by 3.47% in global model accuracy. It also demonstrated top performance on MNIST and Fashion-MNIST, outperforming competitors by 1.25% and 0.42% on small subsets, respectively. Additionally, the algorithm exhibited rapid and stable convergence within roughly 50 iterations on small datasets while maintaining the unique ability to quantify predictive uncertainty as training progressed.
These results demonstrate that treating network parameters probabilistically effectively controls overfitting while allowing flexible local adaptation. The ability to measure output uncertainty provides critical decision-support value for safety-critical and regulated applications where understanding model confidence is necessary. While pFedBayes delivered clear superiority in limited and medium data regimes, its global model showed reduced performance advantages on large datasets, reflecting known scaling characteristics of Bayesian neural networks that require specialized aggregation techniques.
Organizations deploying federated learning in data-constrained or heterogeneous environments should consider adopting Bayesian personalized architectures to improve local model accuracy and capture reliable uncertainty metrics. Practitioners must carefully calibrate the regularization parameter balancing personalization and global aggregation, as well as set modest local learning rates to ensure stable convergence. Future work should focus on piloting the approach in operational environments with varied client connectivity and integrating specialized scaling techniques to maintain performance when data volumes expand.
- Paper: Weight Uncertainty in Neural Network, Charles Blundell et al. (2015). Introduces Bayes by Backprop for variational inference over neural network weights, providing the foundational Bayesian neural network formulation that the source adapts for decentralized federated optimization.
- Paper: Personalized Federated Learning with Moreau Envelopes, Canh T. Dinh et al. (2020). Formulates bi-level optimization for personalized federated learning via Moreau envelopes, establishing the core bi-level regularized objective structure that the source extends into a variational Bayesian framework.
- Paper: Personalized Federated Learning with Theoretical Guarantees: A Model-Agnostic Meta-Learning Approach, Alireza Fallah et al. (2020). Pioneers model-agnostic meta-learning for personalized federated learning, establishing key baseline concepts for adapting global shared models to heterogeneous local clients.
- Paper: Ditto: Fair and Robust Federated Learning Through Personalization, Tian Li et al. (2020). Develops a bi-level multi-task objective that regularizes local personalized models toward a global model, providing a fundamental baseline for personalized federated learning objectives.
- Paper: Federated Learning with Personalization Layers, Manoj Ghuhan Arivazhagan et al. (2019). Introduces layer-based personalization in federated deep neural networks to mitigate statistical data heterogeneity across clients.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). Introduces proximal regularization to stabilize federated optimization under statistical heterogeneity, forming an essential foundation for regularized client optimization.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). Introduces the standard FederatedAveraging algorithm that serves as the universal foundation for global federated aggregation.
- Paper: Federated Multi-Task Learning, Virginia Smith et al. (2017). Presents a multi-task learning formulation for federated systems to address non-identical client distributions, laying the groundwork for personalized federated learning.
- Paper: Personalized Federated Learning through Local Memorization, Othmane Marfoq et al. (2022). Extends personalized federated learning under client heterogeneity by leveraging localized spatial memorization within neural feature extractors.
- Paper: CD2-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning, Yiqing Shen et al. (2022). Advances personalized federated learning by introducing channel decoupling and cyclic knowledge distillation across layers for heterogeneous client domains.
- Paper: FedTGP: Trainable Global Prototypes with Adaptive-Margin-Enhanced Contrastive Learning for Data and Model Heterogeneity in Federated Learning, Jianqing Zhang et al. (2024). Builds on client personalization challenges under statistical heterogeneity by developing trainable global prototype representations via margin-enhanced contrastive learning.
