Federated Learning with Personalization Layers
Manoj Ghuhan ArivazhaganVinay AggarwalAaditya Kumar SinghSunav Choudhary
Introduces FedPer, a federated learning framework that splits deep neural networks into shared base layers and local personalization layers to overcome performance degradation caused by statistical data heterogeneity across edge devices.
Modern edge devices generate massive volumes of personalized user data, but strict privacy constraints prevent centralizing this raw information. Federated learning enables devices to train machine learning models collaboratively on-device without sharing raw data. However, individual users naturally possess diverse preferences and behaviors, creating severe statistical differences in local data distributions. Standard federated learning techniques attempt to train a single global model for every device, which causes severe performance degradation and unfairness when tasks require personalized outputs.
The article introduces and evaluates FedPer, a novel training approach for deep feedforward neural networks that splits models into shared base layers and private personalization layers. The main objective is to demonstrate that this split architecture can overcome the performance failures of standard federated learning on heterogeneous personalization tasks while preserving data privacy.
To evaluate the method, the authors conducted extensive simulated experiments using two standard deep neural network architectures: ResNet-34 and MobileNet-v1. The evaluations tested the framework across 100 global aggregation rounds on benchmark image classification datasets (CIFAR-10 and CIFAR-100 across 10 clients with varying degrees of non-identical data splits) and a real-world personalized image aesthetics dataset (FLICKR-AES across 30 clients). Under the proposed framework, shared base layers are aggregated across devices via a central server using federated averaging, while private personalization layers remain strictly local and are updated solely using local data.
The findings show that FedPer delivers superior performance, converging faster and achieving significantly higher steady-state accuracy than standard federated averaging under severe data heterogeneity. On the real-world personalized aesthetics task, standard federated learning completely failed—performing no better than random guessing—whereas FedPer successfully learned distinct user preferences. Additionally, FedPer substantially reduced the variation in model accuracy across individual devices, ensuring fairer outcomes. Control experiments confirmed that both components are essential: shared base layers successfully extract shared visual features that purely local training cannot learn due to limited data, while at least one personalization layer is vital to capture unique user preferences.
These results demonstrate that standard federated learning is fundamentally ill-suited for personalized applications where identical inputs receive different user labels. Adopting a split-layer approach allows organizations to deliver highly tailored, high-performing user experiences without transferring sensitive personal data off edge devices. Furthermore, keeping personalization layers private reduces communication bandwidth requirements between client devices and central servers, lowering operational costs and network overhead.
Organizations developing personalized edge applications should adopt split-layer architectures like FedPer rather than enforcing single global models. Implementation teams should select personalization depths based on task complexity; experimental results show that 1 to 2 personal layers provide optimal performance without excessive local computation. Before deploying at scale, practitioners should conduct pilot studies on target hardware to evaluate edge device memory limits and test optional local fine-tuning steps, which improved accuracy on standard classification benchmarks but showed negligible benefits on aesthetic tasks.
The evaluations were conducted under controlled assumptions, including static client datasets, synchronous communications, and all participating devices remaining active throughout training. Because real-world deployments frequently experience intermittent connectivity and dynamic data updates, organizations should validate the approach under realistic network constraints and varied device hardware profiles.
- Paper: Communication-Efficient Learning of Deep Networks from Decentralized Data, H. B. McMahan et al. (2016). Introduces the foundational FederatedAveraging (FedAvg) algorithm and decentralized edge training framework that FedPer extends by decoupling base layers from personalized layers.
- Paper: Federated Learning with Non-IID Data, Yue Zhao et al. (2018). Analyzes weight divergence and severe performance degradation in standard federated learning caused by non-IID data distributions, establishing the core problem that FedPer directly addresses.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). Formulates the mathematical challenge of statistical heterogeneity in federated networks and serves as a key optimization baseline for mitigating non-IID client drift.
- Paper: Federated Multi-Task Learning, Virginia Smith et al. (2017). Provides the foundational multi-task learning formulation for edge devices that underpins the paradigm of learning personalized client models rather than a single global model.
- Paper: Federated Optimization: Distributed Machine Learning for On-Device Intelligence, Jakub Konečný et al. (2016). Formulates the foundational distributed optimization constraints and on-device intelligence settings that necessitate personalized federated architectures.
- Paper: Towards Personalized Federated Learning, Alysa Ziying Tan et al. (2021). Surveys the broader landscape and taxonomy of personalized federated learning, categorizing parameter-decoupling approaches like FedPer alongside meta-learning and distillation strategies.
- Paper: Personalized Federated Learning with Moreau Envelopes, Canh T. Dinh et al. (2020). Advances personalized federated learning by formulating bi-level optimization via Moreau envelopes as an alternative regularization approach to layer-split personalization.
- Paper: Ditto: Fair and Robust Federated Learning Through Personalization, Tian Li et al. (2020). Extends client-level personalization to simultaneously achieve robustness against adversarial attacks and fairness across heterogeneous client devices.
- Paper: Federated Learning on Non-IID Data: A Survey, Hangyu Zhu et al. (2021). Provides a comprehensive taxonomy and evaluation of non-IID mitigation strategies, contextualizing local-layer personalization methods within parametric federated learning.
- Paper: Federated Learning on Non-IID Data Silos: An Experimental Study, Qinbin Li et al. (2021). Presents an experimental benchmark across distinct non-IID partition types to systematically test algorithm robustness against the label and feature skews targeted by personalization.
- Paper: Model-Contrastive Federated Learning, Qinbin Li et al. (2021). Proposes model-contrastive learning to correct local representation drift in heterogeneous settings, building on the feature-representation perspective of federated deep networks.
