CD2-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning

Yiqing ShenYuyin ZhouLequan Yu

article2022CVPR87 citations

Proposes a personalized federated learning framework that decouples model parameters at the channel dimension rather than across layers, pairing this design with cyclic knowledge distillation to address diverse non-IID data distributions across natural and medical image benchmarks.

Listen

Federated learning allows multiple organizations or devices to collaboratively train artificial intelligence models without sharing private, decentralized data. However, real-world deployments encounter severe performance drops due to statistical data heterogeneity across participants, including variations in features, label distributions, and user preferences. Existing personalized federated learning techniques attempt to fix this by splitting models into separate shared and private layers. Because they restrict customization to either the top or bottom of a network, these methods rely on prior assumptions about the data and fail when different forms of data discrepancy co-occur.

The article evaluates and demonstrates CD2-pFed, a framework designed to achieve robust model personalization across diverse data heterogeneity scenarios without requiring manual layer selection. The primary objective is to improve both local customized accuracy and generalizability to unseen data while maintaining data privacy.

The researchers developed an approach that splits model capacity channel by channel across all network layers, allocating private channels for local learning and shared channels for global federation. To align these representations, they introduced a mutual knowledge distillation scheme that regularizes local and global predictions against each other during training, alongside a progressive schedule that increases personalization over time and temporal smoothing for stable updates. The approach was evaluated across four image classification benchmarks covering natural and medical domains, including CIFAR-10, CIFAR-100, the aesthetic assessment dataset FLICKR-AES with REAL-CUR validation, and a multi-center colorectal cancer histology image collection termed HISTO-FED.

The empirical findings demonstrate significant performance advantages. On CIFAR-10 under severe label distribution skew, the method improved local classification accuracy by roughly 32% over standard federated averaging (reaching 91.82%) and outperformed leading layer-wise methods on both local and unseen data tests. On CIFAR-100, local test accuracy increased by about 29% over standard federation, outperforming competing personalized frameworks by 1.2% to 4.7%. On user preference data with concept drift, the approach improved local accuracy from 24.50% to 47.89% and external test accuracy from 20.08% to 45.67%. In medical histology with distinct stain variations, the framework consistently surpassed baseline models across participating centers and maintained the highest accuracy on unseen hospital data (65.66%).

These results demonstrate that channel-level personalization provides a unified and reliable solution to varied data skews, eliminating the need to guess whether top or bottom layers require customization. By bridging the gap between local and global representations, the method delivers faster convergence during training, which reduces communication overhead and network operational costs for distributed systems while maintaining strict data privacy compliance.

Organizations implementing federated systems should consider channel-level decoupling and prediction alignment when managing diverse data endpoints, especially in sensitive sectors such as healthcare and cross-device consumer applications. Before wider operational adoption across complex computer vision systems, technical teams should conduct pilot studies, as the current distillation mechanism is tailored for classification outputs and requires adaptation for tasks like object detection and semantic segmentation.

The primary limitations of this work stem from its reliance on fixed personalization proportions across all layers and the restriction of the distillation scheme to classification tasks. Nonetheless, because the framework was rigorously validated across multiple distinct non-identical data splits and real-world multi-center datasets, confidence in its reported classification advantages remains high.

arXiv: 2204.03880
Cover for CD2-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning

Abstract

Federated learning (FL) is a distributed learning paradigm that enables multiple clients to collaboratively learn a shared global model. Despite the recent progress, it remains challenging to deal with heterogeneous data clients, as the discrepant data distributions usually prevent the global model from delivering good generalization ability on each participating client. In this paper, we propose CD²-pFed, a novel Cyclic Distillation-guided Channel Decoupling framework, to personalize the global model in FL, under various settings of data heterogeneity. Different from previous works which establish layer-wise personalization to overcome the non-IID data across different clients, we make the first attempt at channel-wise assignment for model personalization, referred to as channel decoupling. To further facilitate the collaboration between private and shared weights, we propose a novel cyclic distillation scheme to impose a consistent regularization between the local and global model representations during the federation. Guided by the cyclical distillation, our channel decoupling framework can deliver more accurate and generalized results for different kinds of heterogeneity, such as feature skew, label distribution skew, and concept shift. Comprehensive experiments on four benchmarks, including natural image and medical image analysis tasks, demonstrate the consistent effectiveness of our method on both local and external validations.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Data Heterogeneity
  • 2.2. Personalized Federated Learning
  • 2.3. Knowledge Distillation
  • 3. Methodology
  • 3.1. Problem Formulation
  • 3.2. Channel Decoupling for Model Personalization
  • 3.3. Cyclic Distillation
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Experimental Settings
  • 4.3. Experimental Results on Synthesized Data
  • 4.4. Experimental Results on Real-world Data
  • 4.5. Ablation Analysis
  • 5. Limitations and Conclusions
  • References

Knowls

  1. Knowl 1 — Channel Decoupling for Personalized Federated Learning

    model/method

    In personalized federated learning (PFL), data across KK clients is non-IID (exhibiting feature distribution skew, label distribution skew, or concept drift). The CD2-pFed\text{CD}^2\text{-pFed} framework achieves model personalization via channel-wise parameter decoupling across all layers of the target neural network, rather than partitioning the network layer-wise (such as splitting exclusively into bottom or top layers).

    For each client i∈{1,…,K}i \in \{1, \dots, K\}, each layer of the model is vertically split into two channel subsets:

    1. Private personalized parameters wiw_i: a proportion p∈[0,1]p \in [0, 1] of channels in each layer that is retained locally and updated exclusively on local data Di\mathcal{D}_i, without being shared with or aggregated by the central server.
    2. Shared global parameters w0w_0: the remaining (1−p)(1 - p) proportion of channels in each layer, which is transmitted to the central server and aggregated across all clients using standard federated aggregation (such as Federated Averaging, FedAvg).

    When p=0p = 0, the framework degenerates into standard FedAvg with no personalization; when p=1p = 1, the framework degenerates into independent local training with no federated communication. By decoupling channels uniformly across both shallow and deep layers, the model simultaneously learns personalized low-level feature representations and high-level classification representations.

  2. Knowl 2 — Progressive Personalization Ratio Schedule

    model/method

    In CD2-pFed\text{CD}^2\text{-pFed}, fixing a static personalization ratio pp throughout training can impede early global representation learning. To provide a better global parameter initialization before assigning higher local capacity, the fraction of personalized channels is scaled progressively across communication rounds.

    At global communication round t∈{1,…,T}t \in \{1, \dots, T\}, where TT is the total number of global epochs and pp is the maximum personalization ratio, the active personalization ratio ptp_t is computed via a linear growth schedule:

    pt=p⋅tTp_t = p \cdot \frac{t}{T}

    In early rounds (t≪Tt \ll T), ptp_t is small, allocating the vast majority of channels to the shared global model w0w_0 for fast convergence of shared representations. As tt approaches TT, ptp_t increases linearly toward pp, expanding local parameter capacity for client-specific personalization.

  3. Knowl 3 — Cyclic Distillation Regularization

    model/method

    To prevent representation divergence and semantic discrepancy between the locally trained private channels wiw_i and the globally aggregated shared channels w0w_0, CD2-pFed\text{CD}^2\text{-pFed} uses cyclic self-distillation during local client updates.

    Let gwi,w0g_{w_i, w_0} denote the full model parameterized jointly by private weights wiw_i and shared weights w0w_0, yielding the soft prediction output y~i\widetilde{y}_i. Let gwig_{w_i} denote the sub-network formed by only the private personalized channels, producing soft prediction y~iL\widetilde{y}_i^L, and let gw0g_{w_0} denote the sub-network formed by only the shared global channels, producing soft prediction y~iG\widetilde{y}_i^G.

    Given an input sample xj\mathbf{x}_j with ground truth label yjy_j at client ii, the joint network is supervised via cross-entropy loss LCE(y~i,yj)\mathcal{L}_{CE}(\widetilde{y}_i, y_j), while the local and global sub-networks are regularized via symmetric Kullback-Leibler (KL) divergence:

    LCD=12(KL(y~iL∥y~iG)+KL(y~iG∥y~iL))\mathcal{L}_{CD} = \frac{1}{2}\left( \text{KL}(\widetilde{y}_i^L \parallel \widetilde{y}_i^G) + \text{KL}(\widetilde{y}_i^G \parallel \widetilde{y}_i^L) \right)

    The total local loss L\mathcal{L} minimized at client ii is:

    L=LCE+λ⋅LCD\mathcal{L} = \mathcal{L}_{CE} + \lambda \cdot \mathcal{L}_{CD}

    where λ\lambda is a balancing coefficient (set to λ=1\lambda = 1). This bidirectional distillation aligns the feature predictive distributions between local private channels and shared global channels.

  4. Knowl 4 — Temporal Moving Average for Personalized Weights

    model/method

    To stabilize local optimization dynamics under non-IID data distributions, CD2-pFed\text{CD}^2\text{-pFed} updates the personalized private channel weights wiw_i using an exponential moving average (EMA) across local optimization steps.

    At global round tt and local epoch step l∈{1,…,ηi}l \in \{1, \dots, \eta_i\}, let wi′lw_i^{\prime l} denote the raw gradient update of the personalized weights obtained from minimizing the combined loss function. The EMA-adjusted personalized weights wilw_i^l are computed as:

    wil=βtwi′l+(1−βt)wil−1w_i^l = \beta_t w_i^{\prime l} + (1 - \beta_t) w_i^{l-1}

    where wil−1w_i^{l-1} is the weight state from the preceding local step, and eta_t is a time-dependent smoothing coefficient governed by a ramp-up schedule:

    βt={β⋅exp⁡(−5(1−tt0)2),t≤t0β,t>t0\beta_t = \begin{cases} \beta \cdot \exp\left(-5\left(1 - \frac{t}{t_0}\right)^2\right), & t \le t_0 \\ \beta, & t > t_0 \end{cases}

    In this formulation, β\beta is set to 0.50.5, and t0t_0 is set to 10%10\% of the total federated communication rounds TT. This schedule applies heavier smoothing in initial training rounds and transitions to a constant smoothing rate once global representations stabilize.

  5. Knowl 5 — Local Client Training Algorithm in CD2-pFed

    algorithm

    The local client training routine for CD2-pFed\text{CD}^2\text{-pFed} updates both shared global weights and private personalized weights at client i∈{1,…,K}i \in \{1, \dots, K\} for a local epoch budget ηi\eta_i during global round tt.

    Input: Local data Di\mathcal{D}_i, local epoch number ηi\eta_i, current global round tt, total global rounds TT, maximum personalization ratio pp, smoothing hyperparameter β\beta, ramp-up threshold t0t_0, loss balancing weight λ\lambda.
    Output: Updated shared global parameters w0tw_0^t for central aggregation.
    1: Download shared global weights w0t−1w_0^{t-1} from Central Server
    2: Compute active personalization ratio pt=p⋅(t/T)p_t = p \cdot (t / T)
    3: if t≤t0t \le t_0 then
    4: βt=β⋅exp⁡(−5(1−t/t0)2)\beta_t = \beta \cdot \exp(-5(1 - t / t_0)^2)
    5: else
    6: βt=β\beta_t = \beta
    7: end if
    8: for local epoch l=1,2,…,ηil = 1, 2, \dots, \eta_i do
    9: Sample a mini-batch of data (x,y)(\mathbf{x}, y) from Di\mathcal{D}_i
    10: Forward pass through joint network gwi,w0g_{w_i, w_0}, private sub-network gwig_{w_i}, and global sub-network gw0g_{w_0} to compute predictions y~i,y~iL,y~iG\widetilde{y}_i, \widetilde{y}_i^L, \widetilde{y}_i^G
    11: Compute cross-entropy loss LCE\mathcal{L}_{CE} between y~i\widetilde{y}_i and yy
    12: Compute cyclic distillation loss LCD=12(KL(y~iL∥y~iG)+KL(y~iG∥y~iL))\mathcal{L}_{CD} = \frac{1}{2}(\text{KL}(\widetilde{y}_i^L \parallel \widetilde{y}_i^G) + \text{KL}(\widetilde{y}_i^G \parallel \widetilde{y}_i^L))
    13: Compute total loss L=LCE+λLCD\mathcal{L} = \mathcal{L}_{CE} + \lambda \mathcal{L}_{CD}
    14: Update global parameters w0w_0 and compute raw local weight update wi′lw_i^{\prime l} via gradient descent on L\mathcal{L}
    15: Adjust personalized channel weights: wil=βtwi′l+(1−βt)wil−1w_i^l = \beta_t w_i^{\prime l} + (1 - \beta_t) w_i^{l-1}
    16: end for
    17: Upload updated global parameters w0tw_0^t to Central Server

    The central server aggregates the uploaded global parameters from all clients using FedAvg (w0=∑i=1Kαiw0,itw_0 = \sum_{i=1}^K \alpha_i w_{0,i}^t, where αi=∣Di∣/∑j∣Dj∣\alpha_i = |\mathcal{D}_i| / \sum_j |\mathcal{D}_j|), while private parameters wiw_i remain strictly resident on client ii.

  6. Knowl 6 — Experimental Benchmarks and Evaluation Protocol

    experimental setup

    CD2-pFed\text{CD}^2\text{-pFed} is evaluated on four image classification benchmarks covering distinct non-IID distribution forms:

    1. CIFAR-10 (Label distribution skew): 10 classes, 60,000 images, K=10K = 10 clients. Non-IID partitions allocate at most s∈{2,3,4,5,8,10}s \in \{2, 3, 4, 5, 8, 10\} classes per client (s=2s=2 is highest heterogeneity; s=10s=10 is IID). Backbone: LeNet-5, local epochs ηi=1\eta_i = 1, batch size b=128b = 128, maximum personalization ratio p=0.5p = 0.5.
    2. CIFAR-100 (Label distribution skew): 100 classes, K=30K = 30 clients, each assigned at most s=40s = 40 classes. Backbone: ResNet-34, ηi=4\eta_i = 4, b=128b = 128, p=0.5p = 0.5.
    3. FLICKR-AES & REAL-CUR (Concept shift): Personalized image aesthetics on K=30K = 30 users/clients (80% train / 20% test per user), with REAL-CUR (14 personal albums rated by a single user) serving as an out-of-distribution external test set. Backbone: ResNet-34, ηi=4\eta_i = 4, b=4b = 4, p=0.8p = 0.8.
    4. HISTO-FED (Feature / stain skew): Colorectal cancer histology whole-slide images across K=3K = 3 client centers (Client 1: NCT-CRC-HE-100K with 7,180 patches; Client 2: CRC-VAL-HE-7K with 7,180 patches; Client 3: private center with 7,000 patches) plus an external center with 4,000 patches. Backbone: ResNet-32, ηi=4\eta_i = 4, b=4b = 4, p=0.5p = 0.5.

    All models are trained for T=50T = 50 global rounds from scratch using SGD with Nesterov momentum 0.90.9 and weight decay 5×10−45 \times 10^{-4}. Performance is measured with:

    • Local Test Top-1 Accuracy (%): Evaluated on client-specific test sets using the client's local personalized model (w0,wi)(w_0, w_i).
    • New Test Top-1 Accuracy (%): Evaluated on unindexed test samples using an ensemble of all client models.
    • External Test Top-1 Accuracy (%): Evaluated on an external test set drawn from an unseen distribution.
  7. Knowl 7 — Classification Performance on Synthesized Label-Skewed Datasets

    data/table

    The table below compares CD2-pFed\text{CD}^2\text{-pFed} against FedAvg, Local Training, LG-Fed (bottom-layer personalization), and FedPer (top-layer personalization) on CIFAR-10 (s=2s=2, K=10K=10) and CIFAR-100 (s=40s=40, K=30K=30).

    Methods CIFAR-10 (s=2s=2) CIFAR-100 (s=40s=40)
    Local Acc (%) ↑\uparrow New Acc (%) ↑\uparrow Local Acc (%) ↑\uparrow New Acc (%) ↑\uparrow
    FedAvg 58.99±1.5058.99 \pm 1.50 58.99±1.5058.99 \pm 1.50 29.23±1.7529.23 \pm 1.75 29.23±1.7529.23 \pm 1.75
    Local Train 87.92±2.1487.92 \pm 2.14 10.03±0.0610.03 \pm 0.06 44.59±0.9044.59 \pm 0.90 11.98±0.2211.98 \pm 0.22
    LG-Fed 91.77±0.5691.77 \pm 0.56 60.79±1.4560.79 \pm 1.45 56.77±0.7556.77 \pm 0.75 34.50±1.0234.50 \pm 1.02
    FedPer 83.29±0.9883.29 \pm 0.98 57.77±1.9857.77 \pm 1.98 53.24±2.3353.24 \pm 2.33 30.47±1.7330.47 \pm 1.73
    CD2-pFed\text{CD}^2\text{-pFed} (Ours) 91.82±0.43\mathbf{91.82 \pm 0.43} 61.31±1.53\mathbf{61.31 \pm 1.53} 57.98±0.64\mathbf{57.98 \pm 0.64} 35.15±0.56\mathbf{35.15 \pm 0.56}

    On CIFAR-10 (s=2s=2), CD2-pFed\text{CD}^2\text{-pFed} improves local accuracy over FedAvg by 32.83%32.83\% and surpasses LG-Fed (91.77%91.77\%) and FedPer (83.29%83.29\%). On CIFAR-100 (s=40s=40), CD2-pFed\text{CD}^2\text{-pFed} attains 57.98%57.98\% local accuracy (+28.75%+28.75\% over FedAvg) and 35.15%35.15\% new test accuracy (+5.92%+5.92\% over FedAvg), outperforming LG-Fed by 1.21%1.21\% local / 0.65%0.65\% new accuracy and FedPer by 4.74%4.74\% local / 4.68%4.68\% new accuracy.

  8. Knowl 8 — Performance on Concept-Shifted and Feature-Skewed Real-World Datasets

    data/table

    The performance of CD2-pFed\text{CD}^2\text{-pFed} was evaluated against FedAvg, Local Training, LG-Fed, and FedPer on two real-world datasets: FLICKR-AES / REAL-CUR (concept shift in aesthetic scores) and HISTO-FED (feature/stain skew across 3 client centers and 1 external center).

    Methods FLICKR-AES / REAL-CUR HISTO-FED
    Local Acc (%) ↑\uparrow External Acc (%) ↑\uparrow Client 1 (%) Client 2 (%) Client 3 (%) External (%) ↑\uparrow
    FedAvg 24.50±2.0124.50 \pm 2.01 20.08±1.3420.08 \pm 1.34 65.2365.23 65.3165.31 65.4565.45 60.0360.03
    Local Train – – 75.5375.53 74.8774.87 74.2174.21 34.3134.31
    LG-Fed 25.78±2.4025.78 \pm 2.40 20.98±1.3420.98 \pm 1.34 76.3276.32 76.9076.90 77.0177.01 63.2263.22
    FedPer 43.26±3.2343.26 \pm 3.23 40.55±1.7840.55 \pm 1.78 75.4375.43 75.2175.21 75.5675.56 57.8957.89
    CD2-pFed\text{CD}^2\text{-pFed} (Ours) 47.89±2.03\mathbf{47.89 \pm 2.03} 45.67±1.67\mathbf{45.67 \pm 1.67} 77.39\mathbf{77.39} 77.45\mathbf{77.45} 77.38\mathbf{77.38} 65.66\mathbf{65.66}

    On FLICKR-AES (label/concept skew), FedPer outperforms LG-Fed (43.26%43.26\% vs. 25.78%25.78\%) because top-layer personalization handles label skew better than bottom-layer personalization, but CD2-pFed\text{CD}^2\text{-pFed} surpasses both (47.89%47.89\% local, 45.67%45.67\% external). On HISTO-FED (feature skew via stain variations), LG-Fed outperforms FedPer (76.32–77.01%76.32\text{--}77.01\% vs. 75.21–75.56%75.21\text{--}75.56\% local; 63.22%63.22\% vs. 57.89%57.89\% external), whereas CD2-pFed\text{CD}^2\text{-pFed} achieves the highest performance on all internal clients (77.38%ext−−77.45%77.38\% ext{--}77.45\%) and the external center (65.66%65.66\%). Channel decoupling thus avoids performance drops caused by mismatch between fixed layer assignment and the underlying skew type.

  9. Knowl 9 — Ablation Analysis of CD2-pFed Components

    data/table

    An ablation study on CIFAR-10 (s=2s=2, K=10K=10) analyzes the contribution of the three functional components in CD2-pFed\text{CD}^2\text{-pFed}:

    1. LI: Linear Increment of the personalization ratio pt=p⋅(t/T)p_t = p \cdot (t / T).
    2. TA: Temporal moving Average for personalized channel weights.
    3. CD: Cyclic Distillation loss LCD\mathcal{L}_{CD}.
    LI TA CD Local Accuracy (%) ↑\uparrow New Accuracy (%) ↑\uparrow
    90.31±0.6790.31 \pm 0.67 59.12±0.3259.12 \pm 0.32
    ✓ 90.36±0.6590.36 \pm 0.65 59.14±0.3059.14 \pm 0.30
    ✓ 90.45±0.2190.45 \pm 0.21 59.45±0.5459.45 \pm 0.54
    ✓ 90.58±0.4490.58 \pm 0.44 60.57±0.3260.57 \pm 0.32
    ✓ ✓ 91.67±0.5491.67 \pm 0.54 61.20±2.0361.20 \pm 2.03
    ✓ ✓ 91.00±1.0391.00 \pm 1.03 59.84±1.5359.84 \pm 1.53
    ✓ ✓ 90.81±0.3890.81 \pm 0.38 59.34±0.5659.34 \pm 0.56
    ✓ ✓ ✓ 91.82±0.43\mathbf{91.82 \pm 0.43} 61.31±1.53\mathbf{61.31 \pm 1.53}

    Combining all three components (LI + TA + CD) produces a 1.51%1.51\% increase in local test accuracy over base channel decoupling (91.82%91.82\% vs. 90.31%90.31\%) and improves new test accuracy from 59.12%59.12\% to 61.31%61.31\%. Among individual additions, CD provides the largest accuracy gain (+0.27%+0.27\% local, +1.45%+1.45\% new), while TA and LI reduce variance across runs (standard deviation decreases from 0.670.67 to 0.430.43).

  10. Knowl 10 — Limitations of CD2-pFed

    limitation

    The CD2-pFed\text{CD}^2\text{-pFed} framework has two primary limitations:

    1. Task Restriction to Classification: The cyclic distillation mechanism aligns prediction probability vectors via Kullback-Leibler divergence over class logits, restricting the current formulation to classification tasks. Extending cyclic distillation to dense output vision tasks such as semantic segmentation and object detection requires alternative distillation formulations.
    2. Uniform Personalization Ratio across Layers: The framework sets a uniform personalization ratio pp across all network layers rather than optimizing layer-specific partition ratios, leaving layer-adaptive ratio search as an open direction.

Coverage note — None was omitted; all key contributions including the channel decoupling concept, progressive scheduling, cyclic distillation, temporal averaging, algorithmic workflow, experimental evaluations on four datasets, ablation analysis, and stated limitations are fully covered.

References

  1. 1.George J Annas. Hipaa regulations—a new era of medical-record privacy?, 2003. 1
  2. 2.Manoj Ghuhan Arivazhagan, Vinay Aggarwal, Aaditya Kumar Singh, and Sunav Choudhary. Federated learning with personalization layers. arXiv preprint arXiv:1912.00818, 2019. 1, 2, 3, 4, 5, 6, 7, 8
  3. 3.Keith Bonawitz, Hubert Eichner, Wolfgang Grieskamp, Dzmitry Huba, Alex Ingerman, Vladimir Ivanov, Chloe Kiddon, Jakub Koneˇcn` y, Stefano Mazzocchi, H Brendan McMahan, et al. Towards federated learning at scale: System design. arXiv preprint arXiv:1902.01046, 2019. 1
  4. 4.Xuan Gong, Abhishek Sharma, Srikrishna Karanam, Ziyan Wu, Terrence Chen, David Doermann, and Arun Innanje. Ensemble attention distillation for privacy-preserving federated learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15076–15086, October 2021. 1
  5. 5.Pengfei Guo, Puyang Wang, Jinyuan Zhou, Shanshan Jiang, and Vishal M. Patel. Multi-institutional collaborations for improving deep learning-based magnetic resonance image reconstruction using federated learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2423–2432, June 2021. 1
  6. 6.Filip Hanzely and Peter Richt'arik. Federated learning of a mixture of global and local models. arXiv preprint arXiv:2002.05516, 2020. 1, 2
  7. 7.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 1
  8. 8.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 6
  9. 9.Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531, 2015. 3
  10. 10.Yihan Jiang, Jakub Koneˇcn` y, Keith Rush, and Sreeram Kannan. Improving federated learning personalization via model agnostic meta learning. arXiv preprint arXiv:1909.12488, 2019. 2, 3
  11. 11.Peter Kairouz, H Brendan McMahan, Brendan Avent, Aur'elien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al. Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977, 2019. 2, 3, 4
  12. 12.Jing Ke, Yiqing Shen, and Yizhou Lu. Style normalization in histology with federated learning. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 953–956. IEEE, 2021. 6
  13. 13.Jing Ke, Yiqing Shen, Jason D Wright, Naifeng Jing, Xiaoyao Liang, and Dinggang Shen. Identifying patch-level msi from histological images of colorectal cancer by a knowledge distillation model. In 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), pages 1043–1046. IEEE, 2020. 3
  14. 14.Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 5
  15. 15.Viraj Kulkarni, Milind Kulkarni, and Aniruddha Pant. Survey of personalization techniques for federated learning. In 2020 Fourth World Conference on Smart Trends in Systems, Security and Sustainability (WorldS4), pages 794–797. IEEE, 2020. 3
  16. 16.Samuli Laine and Timo Aila. Temporal ensembling for semisupervised learning. arXiv preprint arXiv:1610.02242, 2016. 5
  17. 17.Yann LeCun, L'eon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. 6
  18. 18.Daliang Li and Junpu Wang. Fedmd: Heterogenous federated learning via model distillation. arXiv preprint arXiv:1910.03581, 2019. 4
  19. 19.Qinbin Li, Bingsheng He, and Dawn Song. Modelcontrastive federated learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10713–10722, June 2021. 1
  20. 20.Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018. 3
  21. 21.Xiaoxiao Li, Meirui JIANG, Xiaofei Zhang, Michael Kamp, and Qi Dou. Fedbn: Federated learning on non-iid features via local batch normalization. In International Conference on Learning Representations, 2020. 3
  22. 22.Paul Pu Liang, Terrance Liu, Liu Ziyin, Nicholas B Allen, Randy P Auerbach, David Brent, Ruslan Salakhutdinov, and Louis-Philippe Morency. Think locally, act globally: Federated learning with local and global representations. arXiv preprint arXiv:2001.01523, 2020. 1, 2, 3, 4, 5, 6, 7, 8
  23. 23.Tao Lin, Lingjing Kong, Sebastian U Stich, and Martin Jaggi. Ensemble distillation for robust model fusion in federated learning. arXiv preprint arXiv:2006.07242, 2020. 4
  24. 24.Quande Liu, Cheng Chen, Jing Qin, Qi Dou, and Pheng-Ann Heng. Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous frequency space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1013–1023, June 2021. 1
  25. 25.Jonathan Long, Evan Shelhamer, and Trevor Darrell. Fully convolutional networks for semantic segmentation. In CVPR, pages 3431–3440, 2015. 1
  26. 26.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 4
  27. 27.Yishay Mansour, Mehryar Mohri, Jae Ro, and Ananda Theertha Suresh. Three approaches for personalization with applications to federated learning. arXiv preprint arXiv:2002.10619, 2020. 3
  28. 28.Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communicationefficient learning of deep networks from decentralized data. In Artificial intelligence and statistics, pages 1273–1282. PMLR, 2017. 1, 2, 4, 5, 6, 7, 8
  29. 29.Joaquin Quinonero-Candela, Masashi Sugiyama, Neil D˜ Lawrence, and Anton Schwaighofer. Dataset shift in machine learning. Mit Press, 2009. 2, 3
  30. 30.Jian Ren, Xiaohui Shen, Zhe Lin, Radomir Mech, and David J Foran. Personalized image aesthetics. In Proceedings of the IEEE international conference on computer vision, pages 638–647, 2017. 3, 6
  31. 31.Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. NeurIPS, 28:91–99, 2015. 1
  32. 32.Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. Fitnets: Hints for thin deep nets. arXiv preprint arXiv:1412.6550, 2014. 3
  33. 33.Johannes Schneider and Michail Vlachos. Personalization of deep learning. arXiv preprint arXiv:1909.02803, 2019. 3
  34. 34.Leslie N Smith. Cyclical learning rates for training neural networks. In 2017 IEEE winter conference on applications of computer vision (WACV), pages 464–472. IEEE, 2017. 4
  35. 35.Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet Talwalkar. Federated multi-task learning. arXiv preprint arXiv:1705.10467, 2017. 2, 3
  36. 36.Jingwei Sun, Ang Li, Binghui Wang, Huanrui Yang, Hai Li, and Yiran Chen. Soteria: Provable defense against privacy leakage in federated learning from representation perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9311–9319, June 2021. 1
  37. 37.Alysa Ziying Tan, Han Yu, Lizhen Cui, and Qiang Yang. Towards personalized federated learning. arXiv preprint arXiv:2103.00710, 2021. 2
  38. 38.Antti Tarvainen and Harri Valpola. Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. arXiv preprint arXiv:1703.01780, 2017. 3
  39. 39.Paul Voigt and Axel Von dem Bussche. The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing, 10:3152676, 2017. 1
  40. 40.Qiong Wu, Kaiwen He, and Xu Chen. Personalized federated learning for intelligent iot applications: A cloud-edge based framework. IEEE Open Journal of the Computer Society, 1:35–44, 2020. 3
  41. 41.Guodong Xu, Ziwei Liu, Xiaoxiao Li, and Chen Change Loy. Knowledge distillation meets self-supervision. In European Conference on Computer Vision, pages 588–604. Springer, 2020. 3
  42. 42.Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim. A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4133–4141, 2017. 3
  43. 43.Jiahui Yu and Thomas S Huang. Universally slimmable networks and improved training techniques. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1803–1811, 2019. 5
  44. 44.Tao Yu, Eugene Bagdasaryan, and Vitaly Shmatikov. Salvaging federated learning by local adaptation. arXiv preprint arXiv:2002.04758, 2020. 3
  45. 45.Tao Yu, Eugene Bagdasaryan, and Vitaly Shmatikov. Salvaging federated learning by local adaptation. arXiv preprint arXiv:2002.04758, 2020. 3
  46. 46.Sukmin Yun, Jongjin Park, Kimin Lee, and Jinwoo Shin. Regularizing class-wise predictions via self-knowledge distillation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 13876–13885, 2020. 3
  47. 47.Lin Zhang, Yong Luo, Yan Bai, Bo Du, and Ling-Yu Duan. Federated learning for non-iid data via unified feature learning and optimization objective alignment. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 4420–4428, October 2021. 1
  48. 48.Zizhao Zhang, Han Zhang, Sercan O Arik, Honglak Lee, and Tomas Pfister. Distilling effective supervision from severe label noise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9294–9303, 2020. 3
  49. 49.Yue Zhao, Meng Li, Liangzhen Lai, Naveen Suda, Damon Civin, and Vikas Chandra. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582, 2018. 2, 5
  50. 50.Weiming Zhuang, Xin Gan, Yonggang Wen, Shuai Zhang, and Shuai Yi. Collaborative unsupervised visual representation learning from decentralized data. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 4912–4921, October 2021. 1

Citation

MLA
Shen, Y., et al. “CD$^2$-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning”. arXiv, 2022, http://arxiv.org/abs/2204.03880v1.
APA
Shen, Y., Zhou, Y., & Yu, L. (2022). CD$^2$-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning. arXiv. http://arxiv.org/abs/2204.03880v1
Chicago
Shen, Y., Y. Zhou, and L. Yu. 2022. “CD$^2$-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning”. arXiv. http://arxiv.org/abs/2204.03880v1.
Harvard
Shen, Y., Zhou, Y. and Yu, L. (2022) “CD$^2$-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2204.03880v1.
Vancouver
1. Shen Y, Zhou Y, Yu L (2022) CD$^2$-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning. arXiv

BibTeX

@article{shen2022pfed,
  title = {CD$^2$-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated Learning},
  author = {Shen, Yiqing and Zhou, Yuyin and Yu, Lequan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2204.03880v1},
  eprint = {2204.03880}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE