Model-Contrastive Federated Learning

Qinbin LiBingsheng HeDawn Song

article2021CVPR1,845 citations

Introduces MOON, a model-contrastive federated learning framework that utilizes representation-level contrastive learning to correct local training drift and substantially improve accuracy on heterogeneous distributed data.

Listen

Federated learning enables multiple organizations or devices to collaboratively train artificial intelligence models without sharing private local data. However, real-world deployments face a significant performance challenge due to data heterogeneity, where each participant holds an unbalanced and non-representative subset of data. The article introduces and evaluates MOON (model-contrastive federated learning), a framework designed to correct individual participant model drift and improve overall model accuracy on complex image classification tasks.

To evaluate this solution, the authors conducted extensive empirical experiments using standard image datasets, including CIFAR-10, CIFAR-100, and Tiny-ImageNet, across various non-uniform data distributions. They compared MOON against standard federated averaging and other leading correction techniques across multiple network architectures and scaling scenarios, ranging up to 100 participating parties.

Key findings show that MOON consistently outperforms existing methods across all tested datasets. Under baseline 10-party settings, MOON achieved top-1 classification accuracies of 69.1% on CIFAR-10, 67.5% on CIFAR-100, and 25.1% on Tiny-ImageNet, outperforming the standard federated averaging baseline by an average of 2.6% in absolute accuracy. In large-scale setups with 100 parties on CIFAR-100, MOON widened its lead, achieving 61.8% accuracy compared to the 55.0% reached by traditional federated averaging. Furthermore, MOON dramatically improved communication efficiency, requiring roughly one-half to one-fourth the number of communication rounds to match the benchmark accuracy of baseline approaches.

These results demonstrate that aligning local model representations with the global model effectively resolves local update drift without incurring heavy communication overhead. Previous optimization methods that rely on weight-distance restrictions often slow down model convergence, whereas MOON maintains rapid learning while achieving higher final accuracy. This allows distributed networks to train higher-performing models faster, directly reducing network operational costs and training timelines.

Organizations implementing federated learning on heterogeneous vision tasks should consider adopting MOON's model-contrastive loss framework during local training. When deploying at larger scales, teams should tune the loss weight parameter to balance initial convergence speed with long-term accuracy gains. Future exploration is recommended to evaluate MOON in combination with server-side aggregation enhancements, contemporary regularization techniques, and non-vision applications. While confidence in these visual classification benchmarks is high, practitioners should conduct pilot testing when applying the framework to distinct non-image data types or fully decentralized environments.

Cover for Model-Contrastive Federated Learning

Abstract

Federated learning enables multiple parties to collaboratively train a machine learning model without communicating their local data. A key challenge in federated learning is to handle the heterogeneity of local data distribution across parties. Although many studies have been proposed to address this challenge, we find that they fail to achieve high performance in image datasets with deep learning models. In this paper, we propose MOON: model-contrastive federated learning. MOON is a simple and effective federated learning framework. The key idea of MOON is to utilize the similarity between model representations to correct the local training of individual parties, i.e., conducting contrastive learning in model-level. Our extensive experiments show that MOON significantly outperforms the other state-of-the-art federated learning algorithms on various image classification tasks.

Table of Contents

  • 1 Introduction
  • 2 Background and Related Work
  • 2.1 Federated Learning
  • 2.2 Contrastive Learning
  • 3 Model-Contrastive Federated Learning
  • 3.1 Problem Statement
  • 3.2 Motivation
  • 3.3 Method
  • 3.3.1 Network Architecture
  • 3.3.2 Local Objective
  • 3.4 Comparisons with Contrastive Learning
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Accuracy Comparison
  • 4.3 Communication Efficiency
  • 4.4 Number of Local Epochs
  • 4.5 Scalability
  • 4.6 Heterogeneity
  • 4.7 Loss Function
  • 5 Conclusion
  • References
  • A More Details of the Datasets
  • B Projection Head
  • C IID Partition
  • D Hyper-Parameters Study
  • D.1 Effect of μ\mu
  • D.2 Effect of temperature and output dimension
  • E Combining with FedAvgM
  • F Computation Cost
  • G Number of Negative Pairs

Knowls

  1. Knowl 1 — Model-Contrastive Federated Learning (MOON) Architecture and Framework

    model/method

    Model-Contrastive Federated Learning (MOON) is a federated learning framework designed to correct client drift and representation degradation caused by non-identically distributed (non-IID) local data when training deep neural networks.

    In MOON, each party PiP_i utilizes a neural network structured into three components:

    1. A base encoder f(⋅)f(\cdot) that extracts feature vectors from raw input samples xx.
    2. A projection head g(⋅)g(\cdot) (a multi-layer perceptron) that maps extracted features into a fixed-dimensional latent representation space. The mapping through the encoder and projection head is denoted Rw(x)=g(f(x))R_w(x) = g(f(x)), parameterized by model weights ww.
    3. An output classification layer that maps Rw(x)R_w(x) to class logits, with the entire network denoted Fw(x)F_w(x).

    During round tt of local training, party PiP_i extracts representations for each input sample xx using three separate models:

    • The updating local model: z=Rwit(x)z = R_{w_i^t}(x)
    • The global model received at the start of round tt: zglob=Rwt(x)z_{glob} = R_{w^t}(x)
    • The party's local model from the previous round: zprev=Rwit−1(x)z_{prev} = R_{w_i^{t-1}}(x)

    MOON performs contrastive learning at the model representation level: it minimizes the distance between the local representation zz and the global representation zglobz_{glob} (positive pair), while maximizing the distance between zz and the previous local representation zprevz_{prev} (negative pair).

  2. Knowl 2 — Model-Contrastive Loss and Local Objective

    equation

    In Model-Contrastive Federated Learning (MOON), party PiP_i with local dataset Di\mathcal{D}_i optimizes a joint loss function during communication round tt. For an input sample (x,y)∼Di(x, y) \sim \mathcal{D}_i, let z=Rwit(x)z = R_{w_i^t}(x), zglob=Rwt(x)z_{glob} = R_{w^t}(x), and zprev=Rwit−1(x)z_{prev} = R_{w_i^{t-1}}(x) denote the latent representations output by the base encoder and projection head Rw(⋅)R_w(\cdot) of the current local model witw_i^t, the global model wtw^t, and the previous local model wit−1w_i^{t-1}, respectively.

    The model-contrastive loss is defined as: ℓcon(wit;wit−1;wt;x)=−log⁡exp⁡(sim(z,zglob)τ)exp⁡(sim(z,zglob)τ)+exp⁡(sim(z,zprev)τ)\ell_{con}(w_i^t; w_i^{t-1}; w^t; x) = -\log \frac{\exp\left(\frac{\text{sim}(z, z_{glob})}{\tau}\right)}{\exp\left(\frac{\text{sim}(z, z_{glob})}{\tau}\right) + \exp\left(\frac{\text{sim}(z, z_{prev})}{\tau}\right)} where sim(u,v)=u⊤v∥u∥2∥v∥2\text{sim}(u, v) = \frac{u^\top v}{\|u\|_2 \|v\|_2} is the cosine similarity between vectors uu and vv, and τ>0\tau > 0 is a temperature hyperparameter.

    The total sample loss combines standard supervised loss ℓsup\ell_{sup} (e.g., cross-entropy) and the model-contrastive loss: ℓ=ℓsup(wit;(x,y))+μ ℓcon(wit;wit−1;wt;x)\ell = \ell_{sup}(w_i^t; (x, y)) + \mu \, \ell_{con}(w_i^t; w_i^{t-1}; w^t; x) where μ≥0\mu \ge 0 is a hyperparameter balancing the contrastive penalty. The local optimization problem solved by party PiP_i is: min⁡witE(x,y)∼Di[ℓsup(wit;(x,y))+μ ℓcon(wit;wit−1;wt;x)]\min_{w_i^t} \mathbb{E}_{(x, y) \sim \mathcal{D}_i} \left[ \ell_{sup}(w_i^t; (x, y)) + \mu \, \ell_{con}(w_i^t; w_i^{t-1}; w^t; x) \right]

  3. Knowl 3 — The MOON Federated Optimization Algorithm

    algorithm

    The MOON framework executes iterative rounds of local model-contrastive training and central parameter aggregation across NN distributed parties over TT communication rounds.

    Input: number of communication rounds TT, number of parties NN, number of local epochs EE, temperature τ\tau, learning rate η\eta, hyperparameter μ\mu, local datasets D1,…,DN\mathcal{D}_1, \dots, \mathcal{D}_N where ∣D∣=∑i=1N∣Di∣|\mathcal{D}| = \sum_{i=1}^N |\mathcal{D}_i|
    Output: The final model wTw^T
    Server executes:
      initialize w0w^0
      for t=0,1,…,T−1t = 0, 1, \dots, T - 1 do
        for i=1,2,…,Ni = 1, 2, \dots, N in parallel do
          send the global model wtw^t to PiP_i
          wit←PartyLocalTraining(i,wt)w_i^t \leftarrow \text{PartyLocalTraining}(i, w^t)
        wt+1←∑k=1N∣Dk∣∣D∣wktw^{t+1} \leftarrow \sum_{k=1}^N \frac{|\mathcal{D}_k|}{|\mathcal{D}|} w_k^t
      return wTw^T
    PartyLocalTraining(i,wti, w^t):
      wit←wtw_i^t \leftarrow w^t
      for epoch =1,2,…,E= 1, 2, \dots, E do
        for each batch b={x,y}b = \{x, y\} of Di\mathcal{D}_i do
          ℓsup←CrossEntropyLoss(Fwit(x),y)\ell_{sup} \leftarrow \text{CrossEntropyLoss}(F_{w_i^t}(x), y)
          z←Rwit(x)z \leftarrow R_{w_i^t}(x)
          zglob←Rwt(x)z_{glob} \leftarrow R_{w^t}(x)
          zprev←Rwit−1(x)z_{prev} \leftarrow R_{w_i^{t-1}}(x)
          ℓcon←−log⁡exp⁡(sim(z,zglob)/τ)exp⁡(sim(z,zglob)/τ)+exp⁡(sim(z,zprev)/τ)\ell_{con} \leftarrow -\log \frac{\exp(\text{sim}(z, z_{glob})/\tau)}{\exp(\text{sim}(z, z_{glob})/\tau) + \exp(\text{sim}(z, z_{prev})/\tau)}
          ℓ←ℓsup+μℓcon\ell \leftarrow \ell_{sup} + \mu \ell_{con}
          wit←wit−η∇ℓw_i^t \leftarrow w_i^t - \eta \nabla \ell
      return witw_i^t to server
  4. Knowl 4 — Degeneration of MOON to FedAvg Under Representation Alignment

    theoretical result

    When a local model in MOON produces representations that match the global model representations (i.e., zglob=zprevz_{glob} = z_{prev} or zz has equal cosine similarity to both anchors), the model-contrastive loss evaluates to a constant: ℓcon=−log⁡(exp⁡(sim(z,zglob)/τ)exp⁡(sim(z,zglob)/τ)+exp⁡(sim(z,zprev)/τ))=−log⁡(12)=log⁡2\ell_{con} = -\log \left( \frac{\exp(\text{sim}(z, z_{glob})/\tau)}{\exp(\text{sim}(z, z_{glob})/\tau) + \exp(\text{sim}(z, z_{prev})/\tau)} \right) = -\log\left(\frac{1}{2}\right) = \log 2

    Because the gradient of a constant loss with respect to parameters witw_i^t is zero (∇witℓcon=0\nabla_{w_i^t} \ell_{con} = 0), the local parameter update reduces strictly to the supervised loss gradient: ∇witℓ=∇witℓsup\nabla_{w_i^t} \ell = \nabla_{w_i^t} \ell_{sup}

    As a consequence, when local data heterogeneity is negligible or after feature representations align with the global model, MOON automatically recovers standard FedAvg dynamics without requiring manual adjustment of μ\mu.

  5. Knowl 5 — Experimental Benchmark Setup for Vision Datasets under Non-IID Skew

    experimental setup

    MOON was evaluated across three image classification benchmarks with non-IID client distributions:

    • Datasets: CIFAR-10 (10 classes), CIFAR-100 (100 classes), and Tiny-ImageNet (100,000 images across 200 classes).
    • Model Architectures:
      • CIFAR-10: CNN encoder with two 5×55\times 5 convolution layers (6 and 16 channels) followed by 2×22\times 2 max pooling, and two fully connected layers (120 and 84 units) with ReLU activations.
      • CIFAR-100 and Tiny-ImageNet: ResNet-50 encoder.
      • All models and baselines use a 2-layer MLP projection head outputting a 256-dimensional representation vector before the final classification layer.
    • Optimization and Training Hyperparameters: SGD optimizer with learning rate η=0.01\eta = 0.01, momentum 0.90.9, weight decay 0.000010.00001, batch size 64, and default temperature τ=0.5\tau = 0.5. Local epochs E=10E = 10 for federated methods and E=300E = 300 for SOLO. Communication rounds T=100T = 100 for CIFAR-10/100 and T=20T = 20 for Tiny-ImageNet.
    • Non-IID Partitioning: Dirichlet distribution DirN(β)\text{Dir}_N(\beta) where party allocation vector pk∼DirN(β)p_k \sim \text{Dir}_N(\beta) sets the proportion of class kk samples assigned to party jj, with default concentration β=0.5\beta = 0.5 and default number of parties N=10N = 10.
  6. Knowl 6 — Top-1 Accuracy on CIFAR-10, CIFAR-100, and Tiny-ImageNet

    data/table

    Top-1 test accuracy was measured on CIFAR-10, CIFAR-100, and Tiny-ImageNet with N=10N = 10 parties, Dirichlet concentration β=0.5\beta = 0.5, and E=10E = 10 local epochs across three trials (mean ±\pm standard deviation reported; SOLO reports mean and standard deviation across local models). Hyperparameter μ\mu was tuned over {0.1,1,5,10}\{0.1, 1, 5, 10\} for MOON (optimal: 5 for CIFAR-10, 1 for CIFAR-100, 1 for Tiny-ImageNet) and over {0.001,0.01,0.1,1}\{0.001, 0.01, 0.1, 1\} for FedProx (optimal: 0.01 for CIFAR-10, 0.001 for CIFAR-100, 0.001 for Tiny-ImageNet).

    Could not parse LaTeX table

    MOON outperforms FedAvg by an average of 2.6% accuracy across all tasks. FedProx provides negligible improvement over FedAvg when using deep architectures, and SCAFFOLD degrades substantially on ResNet-50 models (CIFAR-100 and Tiny-ImageNet).

  7. Knowl 7 — Communication Rounds and Speedup to Reach Target Accuracy

    data/table

    Communication efficiency was evaluated by counting the number of communication rounds required for each method to match the top-1 test accuracy achieved by FedAvg after 100 rounds (CIFAR-10/100) or 20 rounds (Tiny-ImageNet) under default settings (N=10,β=0.5,E=10N = 10, \beta = 0.5, E = 10).

    Could not parse LaTeX table

    MOON requires less than half the communication rounds of FedAvg across all three benchmarks (reaching 3.7×3.7\times speedup on CIFAR-10, 2.3×2.3\times on CIFAR-100, and 1.8×1.8\times on Tiny-ImageNet), demonstrating substantial communication savings in distributed settings.

  8. Knowl 8 — Scalability Under Large Party Counts and Client Sampling

    data/table

    Scalability was evaluated on CIFAR-100 (β=0.5\beta = 0.5) under two scaled-up settings: (1) 50 total parties with 100% participation per round, evaluated at 100 and 200 communication rounds; and (2) 100 total parties with 20% random client sampling per round (20 active parties), evaluated at 250 and 500 communication rounds.

    Could not parse LaTeX table

    With μ=10\mu = 10, MOON achieves 63.2% accuracy at 200 rounds with 50 parties and 61.8% accuracy at 500 rounds with 100 parties (sampling fraction 0.2), outperforming FedAvg and FedProx by approximately 7% in both configurations.

  9. Knowl 9 — Robustness to Varying Dirichlet Heterogeneity Levels

    data/table

    The effect of data heterogeneity was tested on CIFAR-100 across 10 parties by varying the Dirichlet concentration parameter β∈{0.1,0.5,5.0}\beta \in \{0.1, 0.5, 5.0\}, where smaller β\beta values correspond to more severely skewed local class distributions.

    Could not parse LaTeX table

    MOON achieves superior top-1 accuracy across all skew levels. Under mild skew (β=5.0\beta = 5.0), MOON maintains a >2%>2\% advantage over FedAvg, whereas FedProx performs worse than FedAvg (64.9% vs. 65.7%).

  10. Knowl 10 — Ablation: Model-Contrastive Loss vs. Direct $\ell_2$ Representation Regularization

    data/table

    To evaluate the specific mechanism of the model-contrastive loss ℓcon\ell_{con}, MOON was compared against a baseline incorporating direct ℓ2\ell_2-norm regularization on representations into the local objective: ℓ=ℓsup+μ∥z−zglob∥22\ell = \ell_{sup} + \mu \|z - z_{glob}\|_2^2 where μ\mu was tuned over {0.001,0.01,0.1,1,5,10}\{0.001, 0.01, 0.1, 1, 5, 10\} and the highest accuracy reported.

    Could not parse LaTeX table

    Direct ℓ2\ell_2-norm representation penalty degrades accuracy on CIFAR-10 (65.8% vs. 66.3% for FedAvg) and provides smaller gains on CIFAR-100 and Tiny-ImageNet than MOON. MOON's contrastive design outperforms direct ℓ2\ell_2 distance minimization by simultaneously pulling local representations toward the global model and pushing them away from previous local representations.

  11. Knowl 11 — Sensitivity of Federated Methods to Local Epoch Count

    empirical result

    Evaluating federated optimization across local epoch counts E∈{1,5,10,20,40}E \in \{1, 5, 10, 20, 40\} demonstrates the relationship between local computational budget and client representation drift:

    • At E=1E = 1, local updates per round are minimal, slowing training convergence within a fixed communication budget. Under this setting, SCAFFOLD achieves low accuracy (45.3% on CIFAR-10, 20.4% on CIFAR-100, 2.6% on Tiny-ImageNet), and FedProx drops to 1.2% on Tiny-ImageNet.
    • At intermediate values (E=5,10E = 5, 10), performance peaks across methods as communication overhead is amortized.
    • At high values (E=20,40E = 20, 40), performance across all federated algorithms degrades as excessive local updates drive models toward inconsistent local minima.
    • MOON achieves the highest test accuracy across all tested local epoch values on CIFAR-10, CIFAR-100, and Tiny-ImageNet, showing resilience against drift under large local update steps.

Coverage note — None was omitted. The knowl set covers the complete MOON methodology, mathematical formulation, algorithm, theoretical representation alignment property, experimental setup, and all quantitative evaluation results including accuracy benchmarks, communication speedup, scalability with client sampling, heterogeneity robustness, loss function ablations, and local epoch sensitivity.

References

  1. 1.Durmus Alp Emre Acar, Yue Zhao, Ramon Matas, Matthew Mattina, Paul Whatmough, and Venkatesh Saligrama. Federated learning based on dynamic regularization. In International Conference on Learning Representations, 2021.
  2. 2.Sebastian Caldas, Sai Meher Karthik Duddu, Peter Wu, Tian Li, Jakub Konecnˇ y, H Brendan McMahan, Virginia Smith, ` and Ameet Talwalkar. Leaf: A benchmark for federated settings. arXiv preprint arXiv:1812.01097, 2018.
  3. 3.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020.
  4. 4.Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton. Big self-supervised models are strong semi-supervised learners. arXiv preprint arXiv:2006.10029, 2020.
  5. 5.Zhongxiang Dai, Bryan Kian Hsiang Low, and Patrick Jaillet. Federated bayesian optimization via thompson sampling. Advances in Neural Information Processing Systems, 33, 2020.
  6. 6.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009.
  7. 7.Canh T Dinh, Nguyen H Tran, and Tuan Dung Nguyen. Personalized federated learning with moreau envelopes. arXiv preprint arXiv:2006.08848, 2020.
  8. 8.Alireza Fallah, Aryan Mokhtari, and Asuman Ozdaglar. Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach. Advances in Neural Information Processing Systems, 33, 2020.
  9. 9.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin ´ Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.
  10. 10.Filip Hanzely, Slavom´ır Hanzely, Samuel Horvath, and Peter ´ Richtarik. Lower bounds and optimal algorithms for person- ´ alized federated learning. arXiv preprint arXiv:2010.02372, 2020.
  11. 11.Chaoyang He, Songze Li, Jinhyun So, Mi Zhang, Hongyi Wang, Xiaoyang Wang, Praneeth Vepakomma, Abhishek Singh, Hang Qiu, Li Shen, Peilin Zhao, Yan Kang, Yang Liu, Ramesh Raskar, Qiang Yang, Murali Annavaram, and Salman Avestimehr. Fedml: A research library and benchmark for federated machine learning. arXiv preprint arXiv:2007.13518, 2020.
  12. 12.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020.
  13. 13.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016.
  14. 14.Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019.
  15. 15.Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. Federated visual classification with real-world data distribution. arXiv preprint arXiv:2003.08082, 2020.
  16. 16.Sixu Hu, Yuan Li, Xu Liu, Qinbin Li, Zhaomin Wu, and Bingsheng He. The oarf benchmark suite: Characterization and implications for federated learning systems. arXiv preprint arXiv:2006.07856, 2020.
  17. 17.Yutao Huang, Lingyang Chu, Zirui Zhou, Lanjun Wang, Jiangchuan Liu, Jian Pei, and Yong Zhang. Personalized cross-silo federated learning on non-iid data. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021.
  18. 18.Longlong Jing and Yingli Tian. Self-supervised visual feature learning with deep neural networks: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020.
  19. 19.Rie Johnson and Tong Zhang. Accelerating stochastic gradient descent using predictive variance reduction. In Advances in neural information processing systems, pages 315–323, 2013.
  20. 20.Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurelien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith ´ Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al. Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977, 2019.
  21. 21.Georgios A Kaissis, Marcus R Makowski, Daniel Ruckert, ¨ and Rickmer F Braren. Secure, privacy-preserving and federated machine learning in medical imaging. Nature Machine Intelligence, pages 1–7, 2020.
  22. 22.Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh. Scaffold: Stochastic controlled averaging for ondevice federated learning. In Proceedings of the 37th International Conference on Machine Learning. PMLR, 2020.
  23. 23.Rajesh Kumar, Abdullah Aman Khan, Sinmin Zhang, WenYong Wang, Yousif Abuidris, Waqas Amin, and Jay Kumar. Blockchain-federated-learning and deep learning models for covid-19 detection using ct imaging. arXiv preprint arXiv:2007.06537, 2020.
  24. 24.Qinbin Li, Yiqun Diao, Quan Chen, and Bingsheng He. Federated learning on non-iid data silos: An experimental study. arXiv preprint arXiv:2102.02079, 2021.
  25. 25.Qinbin Li, Zeyi Wen, and Bingsheng He. Practical federated gradient boosting decision trees. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 4642– 4649, 2020.
  26. 26.Qinbin Li, Zeyi Wen, Zhaomin Wu, Sixu Hu, Naibo Wang, and Bingsheng He. A survey on federated learning systems: Vision, hype and reality for data privacy and protection. arXiv preprint arXiv:1907.09693, 2019.
  27. 27.Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith. Federated learning: Challenges, methods, and future directions. arXiv preprint arXiv:1908.07873, 2019.
  28. 28.Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. In Third Conference on Machine Learning and Systems (MLSys), 2020.
  29. 29.Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. On the convergence of fedavg on non-iid data. In International Conference on Learning Representations, 2020.
  30. 30.Xiaoxiao Li, Meirui JIANG, Xiaofei Zhang, Michael Kamp, and Qi Dou. Fed{bn}: Federated learning on non-{iid} features via local batch normalization. In International Conference on Learning Representations, 2021.
  31. 31.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence ´ Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014.
  32. 32.Yang Liu, Anbu Huang, Yun Luo, He Huang, Youzhi Liu, Yuanyuan Chen, Lican Feng, Tianjian Chen, Han Yu, and Qiang Yang. Fedvision: An online visual object detection platform powered by federated learning. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 13172– 13179, 2020.
  33. 33.Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(Nov):2579–2605, 2008.
  34. 34.H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, et al. Communication-efficient learning of deep networks from decentralized data. arXiv preprint arXiv:1602.05629, 2016.
  35. 35.Ishan Misra and Laurens van der Maaten. Self-supervised learning of pretext-invariant representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6707–6717, 2020.
  36. 36.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  37. 37.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems, pages 8026–8037, 2019.
  38. 38.Kihyuk Sohn. Improved deep metric learning with multiclass n-pair loss objective. In Advances in neural information processing systems, pages 1857–1865, 2016.
  39. 39.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. arXiv preprint arXiv:1906.05849, 2019.
  40. 40.Paul Voigt and Axel Von dem Bussche. The eu general data protection regulation (gdpr). A Practical Guide, 1st Ed., Cham: Springer International Publishing, 2017.
  41. 41.Hongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris Papailiopoulos, and Yasaman Khazaeni. Federated learning with matched averaging. In International Conference on Learning Representations, 2020.
  42. 42.Jianyu Wang, Qinghua Liu, Hao Liang, Gauri Joshi, and H Vincent Poor. Tackling the objective inconsistency problem in heterogeneous federated optimization. Advances in Neural Information Processing Systems, 33, 2020.
  43. 43.Lixu Wang, Shichao Xu, Xiao Wang, and Qi Zhu. Addressing class imbalance in federated learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2021.
  44. 44.Qiang Yang, Yang Liu, Tianjian Chen, and Yongxin Tong. Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology (TIST), 10(2):1–19, 2019.
  45. 45.Mikhail Yurochkin, Mayank Agarwal, Soumya Ghosh, Kristjan Greenewald, Nghia Hoang, and Yasaman Khazaeni. Bayesian nonparametric federated learning of neural networks. In Proceedings of the 36th International Conference on Machine Learning. PMLR, 2019.
  46. 46.Fengda Zhang, Kun Kuang, Zhaoyang You, Tao Shen, Jun Xiao, Yin Zhang, Chao Wu, Yueting Zhuang, and Xiaolin Li. Federated unsupervised representation learning. arXiv preprint arXiv:2010.08982, 2020.
  47. 47.Michael Zhang, Karan Sapra, Sanja Fidler, Serena Yeung, and Jose M. Alvarez. Personalized federated learning with first order model optimization. In International Conference on Learning Representations, 2021.

Citation

MLA
Li, Q., et al. “Model-Contrastive Federated Learning”. arXiv, 2021, http://arxiv.org/abs/2103.16257v1.
APA
Li, Q., He, B., & Song, D. (2021). Model-Contrastive Federated Learning. arXiv. http://arxiv.org/abs/2103.16257v1
Chicago
Li, Q., B. He, and D. Song. 2021. “Model-Contrastive Federated Learning”. arXiv. http://arxiv.org/abs/2103.16257v1.
Harvard
Li, Q., He, B. and Song, D. (2021) “Model-Contrastive Federated Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2103.16257v1.
Vancouver
1. Li Q, He B, Song D (2021) Model-Contrastive Federated Learning. arXiv

BibTeX

@article{li2021model,
  title = {Model-Contrastive Federated Learning},
  author = {Li, Qinbin and He, Bingsheng and Song, Dawn},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2103.16257v1},
  eprint = {2103.16257}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: IEEE