FedBN: Federated Learning on Non-IID Features via Local Batch Normalization

Xiaoxiao LiMeirui JiangXiaofei ZhangMichael KampQi Dou

article2021ICLR1,316 citations

Introduces FedBN, a federated learning method that retains local batch normalization parameters to resolve feature distribution shifts across clients, achieving faster convergence and higher accuracy than FedAvg and FedProx.

Listen

Federated learning enables multiple distributed clients to collaboratively train deep learning models without sharing private local data. However, standard methods degrade significantly when local data is not independently and identically distributed across participants. While most prior work focuses on discrepancies in label distributions, real-world deployments often encounter feature shift, where input data appearance or sensor statistics vary substantially across sites despite identical label targets—such as different scanners in medical imaging or varying lighting in autonomous driving.

The article's main objective is to introduce and evaluate FedBN, a lightweight federated learning method designed to mitigate feature shift across clients by keeping Batch Normalization layers strictly local while aggregating all other network parameters on a central server.

To demonstrate this approach, the researchers conducted theoretical convergence analyses under over-parameterized neural network regimes and performed extensive empirical evaluations across multiple domains. The experimental setup included a five-source benchmark digit classification suite, natural image datasets with varying camera and style characteristics (Office-Caltech10 and DomainNet), and a real-world medical task diagnosing autism spectrum disorder using functional brain imaging across four clinical institutions.

The findings show that FedBN substantially outperforms standard federated averaging (FedAvg) and non-IID frameworks (FedProx), achieving both faster convergence and higher final accuracy. In benchmark evaluations, FedBN maintained superior performance across varying client data sizes and update frequencies, delivering the largest performance margins when local data was scarce. On real-world natural image benchmarks, standard federated methods frequently underperformed single-site local training due to severe feature drift, whereas FedBN consistently outperformed both isolated training and baseline federated techniques by over 6% to 10% on several tasks. In clinical neuroimaging, FedBN improved diagnostic accuracy across multiple hospital sites, demonstrating strong practical efficacy in heterogeneous medical environments.

These results demonstrate that simply maintaining local normalization layers effectively harmonizes feature distributions across disparate devices without altering existing optimization or aggregation algorithms. For decision-makers, FedBN eliminates the hyperparameter tuning and substantial communication overhead common to other non-IID solutions, lowering the cost, implementation risk, and operational barriers of deploying collaborative machine learning across privacy-regulated institutions.

Organizations deploying federated learning in domains with high sensor or environmental variability should adopt FedBN within their model architectures. Because it requires minimal code modification and no extra communication bandwidth, it can be integrated directly into production frameworks. For evaluating unseen external clients, teams should deploy the shared global network alongside locally estimated normalization statistics.

Confidence in these findings is high across image classification and neuroimaging tasks with fixed label distributions. However, decision-makers should note that the theoretical guarantees assume over-parameterized two-layer networks and zero-mean feature inputs. Further empirical evaluation is advised for non-vision modalities, highly unbalanced label spaces, or extreme network architectures before enterprise-wide deployment.

  • Paper: Federated Learning on Non-IID Data: A Survey, Hangyu Zhu et al. (2021). This comprehensive survey contextualizes feature shift mitigation and layer-wise personalization strategies within the broader landscape of non-IID federated learning algorithms.
  • Paper: Federated Learning on Non-IID Data Silos: An Experimental Study, Qinbin Li et al. (2021). This empirical study systematically evaluates federated learning algorithms across explicit feature and label distribution skews using the standardized NIID-Bench framework.
  • Paper: Model-Contrastive Federated Learning, Qinbin Li et al. (2021). This paper advances feature representation handling under non-IID data by introducing model-contrastive learning to align local client representations with the global model.
  • Paper: Towards Personalized Federated Learning, Alysa Ziying Tan et al. (2021). This survey provides an extensive taxonomy of personalized federated learning methods, analyzing architectural parameter-decoupling approaches alongside global adaptation techniques.
Cover for FedBN: Federated Learning on Non-IID Features via Local Batch Normalization

Abstract

The emerging paradigm of federated learning (FL) strives to enable collaborative training of deep models on the network edge without centrally aggregating raw data and hence improving data privacy. In most cases, the assumption of independent and identically distributed samples across local clients does not hold for federated learning setups. Under this setting, neural network training performance may vary significantly according to the data distribution and even hurt training convergence. Most of the previous work has focused on a difference in the distribution of labels or client shifts. Unlike those settings, we address an important problem of FL, e.g., different scanners/sensors in medical imaging, different scenery distribution in autonomous driving (highway vs. city), where local clients store examples with different distributions compared to other clients, which we denote as feature shift non-iid. In this work, we propose an effective method that uses local batch normalization to alleviate the feature shift before averaging models. The resulting scheme, called FedBN, outperforms both classical FedAvg, as well as the state-of-the-art for non-iid data (FedProx) on our extensive experiments. These empirical results are supported by a convergence analysis that shows in a simplified setting that FedBN has a faster convergence rate than FedAvg. Code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Preliminary
  • 4 Federated Averaging with Local Batch Normalization
  • 4.1 Proposed Method - FedBN
  • 4.2 Problem Setup
  • 4.3 Convergence Analysis
  • 5 Experiments
  • 5.1 Benchmark Experiments
  • 5.2 Experiments on real-world datasets
  • 6 Conclusion and Discussion
  • References
  • A Notation Table
  • B Convergence Proof
  • B.1 Evolution Dynamics
  • B.2 Proof of Lemma
  • B.3 Proof of Corollary
  • C FedBN Algorithm
  • D Experimental details
  • D.1 Visualization of Benchmark Datasets
  • D.2 Model Architecture and Training Details on Benchmark
  • D.3 Model Architecture and Traning Details of Image Classification Task on Office-Caltech10 and DomainNet
  • D.4 ABIDE Dataset and Training Details
  • E More Experimental Results on Benchmark Datasets
  • E.1 Convergence Comparison over FedAvg and FedBN
  • E.2 Detailed Statistics of Figure
  • E.3 Compare FedBN with Centralized Training
  • E.4 Different Combinations of EE and BB
  • E.5 Detailed Statistics of Varying Local Dataset Size Experiment
  • E.6 Training on Unequal Dataset Size
  • F Synthetic Data Experiment
  • G Transfer Learning and Testing on Unknown Domain Client

Knowls

  1. Knowl 1 — Federated Learning with Local Batch Normalization (FedBN)

    algorithm

    FedBN is a federated learning algorithm designed to mitigate feature shift non-IID data across clients without modifying server aggregation rules or introducing hyperparameter tuning. The method updates all network parameters locally via stochastic gradient descent (SGD), but selectively aggregates only non-Batch Normalization (non-BN) layers at communication intervals, preserving client-specific BN layers (scaling parameters γ\gamma, shift parameters β\beta, and running mean/variance statistics) locally.

    Input: Total rounds TT, local update period EE, number of clients KK, initial parameters w0,k(l)w_{0, k}^{(l)} for layers $l \in \{1, \dots, L\}
    Output: Local models with personalized BN layers {wT,k}k=1K\{w_{T, k}\}_{k=1}^K
    for round t=1,…,Tt = 1, \dots, T do
        for each client k∈{1,…,K}k \in \{1, \dots, K\} do
            for each layer l∈{1,…,L}l \in \{1, \dots, L\} do
                wt,k(l)←LocalSGD(wt−1,k(l))w_{t, k}^{(l)} \leftarrow \text{LocalSGD}(w_{t-1, k}^{(l)})
            end for
        end for
        if t(modE)==0t \pmod E == 0 then
            for each layer l∈{1,…,L}l \in \{1, \dots, L\} do
                if layer ll is not a Batch Normalization layer then
                    wglobal(l)←1K∑k=1Kwt,k(l)w_{\text{global}}^{(l)} \leftarrow \frac{1}{K} \sum_{k=1}^K w_{t, k}^{(l)}
                    for each client k∈{1,…,K}k \in \{1, \dots, K\} do
                        wt,k(l)←wglobal(l)w_{t, k}^{(l)} \leftarrow w_{\text{global}}^{(l)}
                    end for
                end if
            end for
        end if
    end for

    By keeping BN parameters local, FedBN harmonizes local feature representations before the shared feature extractor weights are averaged across heterogeneous clients.

  2. Knowl 2 — Definition of Feature Shift Non-IID Data in Federated Learning

    definition

    In federated learning with NN distributed clients, each client i∈{1,…,N}i \in \{1, \dots, N\} possesses local data following a joint probability distribution Pi(x,y)P_i(x, y), where x∈Rdx \in \mathbb{R}^d denotes the input features and y∈Ry \in \mathbb{R} (or discrete label space) denotes the target labels. Feature shift non-IID occurs when data distributions deviate in the input feature space across clients while the label distribution remains uniform, encompassing two primary categories:

    1. Covariate Shift: The marginal input feature distributions vary across clients (Pi(x)≠Pj(x)P_i(x) \neq P_j(x) for i≠ji \neq j), while the conditional label distribution given the input is identical across all clients (Pi(y∣x)=Pj(y∣x)P_i(y \mid x) = P_j(y \mid x)).
    2. Concept Shift: The conditional feature distribution given the label varies across clients (Pi(x∣y)≠Pj(x∣y)P_i(x \mid y) \neq P_j(x \mid y) for i≠ji \neq j), while the marginal label distribution P(y)P(y) is identical across clients.
  3. Knowl 3 — Convergence Acceleration of FedBN over FedAvg in Over-Parameterized Networks

    theoretical result

    Consider training a two-layer neural network with mm hidden neurons using gradient descent with step size η=O(1/∥Λ(t)∥)\eta = O(1/\|\Lambda(t)\|) on NN clients, each having MM samples (xji,yji)(x_j^i, y_j^i) under feature covariance shift. Let the initial weight scale satisfy α>1\alpha > 1, targets satisfy ∥y∥∞=O(1)\|y\|_\infty = O(1), and the hidden width satisfy m=Ω(max⁡{N4M4log⁡(NM/δ)/(α4μ04),N2M2log⁡(NM/δ)/μ02})m = \Omega\left(\max\left\{N^4 M^4 \log(NM/\delta)/(\alpha^4 \mu_0^4), N^2 M^2 \log(NM/\delta)/\mu_0^2\right\}\right) for failure probability δ\delta.

    In the magnitude-dominated Neural Tangent Kernel (NTK) regime, the least eigenvalue of the asymptotic Gram matrix controls the linear convergence rate:

    • FedAvg Linear Convergence: ∥f(t)−y∥22≤(1−ημ02)t∥f(0)−y∥22\|f(t) - y\|_2^2 \le \left(1 - \frac{\eta \mu_0}{2}\right)^t \|f(0) - y\|_2^2 where μ0=λmin⁡(G∞)>0\mu_0 = \lambda_{\min}(G^\infty) > 0.

    • FedBN Linear Convergence: ∥f∗(t)−y∥22≤(1−ημ0∗2)t∥f∗(0)−y∥22\|f^*(t) - y\|_2^2 \le \left(1 - \frac{\eta \mu_0^*}{2}\right)^t \|f^*(0) - y\|_2^2 where μ0∗=λmin⁡(G∗∞)>0\mu_0^* = \lambda_{\min}(G^{*\infty}) > 0.

    Because G∗∞=diag(G1∞,…,GN∞)G^{*\infty} = \text{diag}(G_1^\infty, \dots, G_N^\infty) consists solely of the block-diagonal submatrices of G∞G^\infty, linear algebra dictates that: λmin⁡(Gi∞)≥λmin⁡(G∞)∀i∈{1,…,N}  ⟹  μ0∗=min⁡i∈[N]λmin⁡(Gi∞)≥λmin⁡(G∞)=μ0\lambda_{\min}(G_i^\infty) \ge \lambda_{\min}(G^\infty) \quad \forall i \in \{1, \dots, N\} \implies \mu_0^* = \min_{i \in [N]} \lambda_{\min}(G_i^\infty) \ge \lambda_{\min}(G^\infty) = \mu_0

    Consequently, (1−ημ0∗2)≤(1−ημ02)\left(1 - \frac{\eta \mu_0^*}{2}\right) \le \left(1 - \frac{\eta \mu_0}{2}\right), proving that FedBN converges strictly faster than FedAvg.

  4. Knowl 4 — Feature Shift Data Assumptions and Over-Parameterized Model Parameterization

    model/method

    For theoretical analysis of federated learning under feature shift, the training data and models are formalized under the following settings:

    Data Distribution Assumption: For each client i∈[N]i \in [N], input samples xji∈Rdx_j^i \in \mathbb{R}^d are centered (E[xi]=0\mathbb{E}[x^i] = 0) with client-specific covariance matrices Si=E[xi(xi)⊤]S_i = \mathbb{E}[x^i (x^i)^\top] that are positive definite and not all identity matrices. SiS_i is independent of label yy, and for any distinct samples p≠qp \neq q, xp≠κxqx_p \neq \kappa x_q for all κ∈R∖{0}\kappa \in \mathbb{R} \setminus \{0\}.

    FedBN Model Parameterization: The two-layer ReLU network f∗:Rd→Rf^*: \mathbb{R}^d \to \mathbb{R} parameterized by (V,γ,c)∈Rm×d×Rm×N×Rm(V, \gamma, c) \in \mathbb{R}^{m \times d} \times \mathbb{R}^{m \times N} \times \mathbb{R}^m is defined as: f∗(x;V,γ,c)=1m∑k=1mck∑i=1Nσ(γk,ivk⊤x∥vk∥Si)⋅1{x∈client i}f^*(x; V, \gamma, c) = \frac{1}{\sqrt{m}} \sum_{k=1}^m c_k \sum_{i=1}^N \sigma\left(\gamma_{k,i} \frac{v_k^\top x}{\|v_k\|_{S_i}}\right) \cdot \mathbf{1}\{x \in \text{client } i\} where ∥vk∥Si=vk⊤Sivk\|v_k\|_{S_i} = \sqrt{v_k^\top S_i v_k} is the induced vector norm under covariance SiS_i, σ(s)=max⁡{s,0}\sigma(s) = \max\{s, 0\} is the ReLU activation, γk,i\gamma_{k,i} is the client-specific BN scaling parameter, and ckc_k is the output layer weight.

    Initialization: Parameters are initialized as vk(0)∼N(0,α2I)v_k(0) \sim \mathcal{N}(0, \alpha^2 I), ck∼U{−1,1}c_k \sim \mathcal{U}\{-1, 1\}, and γk,i=∥vk(0)∥2/α\gamma_{k,i} = \|v_k(0)\|_2 / \alpha.

  5. Knowl 5 — Auxiliary NTK Gram Matrices for FedAvg and FedBN

    definition

    Given NN clients with MM data points each (total points NMNM), let ipi_p denote the client identity of the pp-th sample xpx_p. The auxiliary Gram matrices G∞∈RNM×NMG^\infty \in \mathbb{R}^{NM \times NM} and G∗∞∈RNM×NMG^{*\infty} \in \mathbb{R}^{NM \times NM} governing the magnitude evolution dynamics in the neural tangent kernel framework are defined as:

    • For FedAvg (global scaling parameters across clients): Gpq∞=Ev∼N(0,α2I)[σ(v⊤xp)σ(v⊤xq)]G_{pq}^\infty = \mathbb{E}_{v \sim \mathcal{N}(0, \alpha^2 I)} \left[ \sigma(v^\top x_p) \sigma(v^\top x_q) \right]

    • For FedBN (client-specific local scaling parameters): Gpq∗∞=Ev∼N(0,α2I)[σ(v⊤xp)σ(v⊤xq)]⋅1{ip=iq}G_{pq}^{*\infty} = \mathbb{E}_{v \sim \mathcal{N}(0, \alpha^2 I)} \left[ \sigma(v^\top x_p) \sigma(v^\top x_q) \right] \cdot \mathbf{1}\{i_p = i_q\}

    where σ(s)=max⁡{s,0}\sigma(s) = \max\{s, 0\} is the ReLU function. G∗∞G^{*\infty} is a block-diagonal matrix whose ii-th diagonal block is the M×MM \times M Gram matrix Gi∞G_i^\infty restricted to samples within client ii.

  6. Knowl 6 — Empirical Performance on Real-World Heterogeneous Benchmarks

    data/table

    FedBN was evaluated on three real-world multi-domain benchmarks: Office-Caltech10 (10 classes across Amazon [A], Caltech [C], DSLR [D], Webcam [W]), DomainNet (10 classes across Clipart [C], Infograph [I], Painting [P], Quickdraw [Q], Real [R], Sketch [S]), and ABIDE I (binary autism spectrum classification using resting-state fMRI connectomes across NYU, USM, UM, UCLA medical centers). Models used were AlexNet with BN for vision tasks and a 3-layer MLP with BN for ABIDE.

    Method Office-Caltech10 DomainNet ABIDE I (medical)
    A C D W C I P Q R S NYU USM UM UCLA
    SingleSet 54.9 40.2 78.7 86.4 41.0 23.8 36.2 73.1 48.5 34.0 58.0 73.4 64.3 57.3
    FedAvg 54.1 44.8 66.9 85.1 48.8 24.9 36.5 56.1 46.3 36.6 62.7 73.1 70.7 64.7
    FedProx 54.2 44.5 65.0 84.4 48.9 24.9 36.6 54.4 47.8 36.9 63.3 73.0 70.5 64.5
    FedBN 63.0 45.3 83.1 90.5 51.2 26.8 41.5 71.3 54.8 42.1 65.6 75.1 68.6 65.5

    Results (mean accuracy over 5 trials) demonstrate that standard FL algorithms (FedAvg, FedProx) can suffer severe performance degradation compared to SingleSet under severe feature shift (e.g., DomainNet Quickdraw dropping from 73.1% to 56.1%/54.4%), whereas FedBN consistently achieves superior performance by preserving domain-specific feature normalizations.

  7. Knowl 7 — Empirical Performance on Digits Classification Benchmark under Domain Shift

    data/table

    The classification accuracy of FedBN was evaluated across five heterogeneous digit datasets serving as 5 FL clients: SVHN, USPS, SynthDigits, MNIST-M, and MNIST (each truncated to 743 samples per client to isolate distribution shift from class/size imbalance; local update epochs E=1E=1, batch size 32, SGD learning rate 10−210^{-2}).

    Method SVHN USPS SynthDigits MNIST-M MNIST
    SingleSet 65.25 (1.07) 95.16 (0.12) 80.31 (0.38) 77.77 (0.47) 94.38 (0.07)
    FedAvg 62.86 (1.49) 95.56 (0.27) 82.27 (0.44) 76.85 (0.54) 95.87 (0.20)
    FedProx 63.08 (1.62) 95.58 (0.31) 82.34 (0.37) 76.64 (0.55) 95.75 (0.21)
    FedBN 71.04 (0.31) 96.97 (0.32) 83.19 (0.42) 78.33 (0.66) 96.57 (0.13)

    Values represent test accuracy mean (standard deviation) across 5 trials. FedBN outperforms SingleSet, FedAvg, and FedProx on all 5 domains, with the largest improvement observed on SVHN (+8.18% over FedAvg), which exhibits the highest visual discrepancy from other domains.

  8. Knowl 8 — Impact of Local Dataset Size on Federated Collaboration Benefits

    empirical result

    Varying the local data availability from 100% down to 1% across clients in the digits benchmark demonstrates the necessity of collaborative training under feature shift:

    • When local data per client is extremely abundant (100%), training locally on single datasets (SingleSet) achieves performance close to federated methods (e.g., USPS: 98.87% SingleSet vs 98.82% FedBN; MNIST: 98.09% SingleSet vs 98.91% FedBN).
    • When local data shrinks to 10%, 5%, or 1%, SingleSet performance drops precipitously (e.g., on SVHN: SingleSet falls from 85.74% at 100% data to 66.81% at 10% data and 12.06% at 1% data).
    • In low-data regimes, FedBN maintains high performance (SVHN at 1% data: 31.98% for FedBN vs 23.67% for FedAvg and 12.06% for SingleSet; USPS at 1% data: 85.05% for FedBN vs 79.09% for FedAvg and 80.11% for SingleSet).

    The performance gain of FedBN over single-client training and FedAvg increases as local sample size decreases.

  9. Knowl 9 — Out-of-Domain Generalization and Unknown Client Evaluation in FedBN

    model/method

    To generalize a trained FedBN global model to a new, previously unseen client outside the federation without joint retraining, two steps are taken:

    1. The global shared parameters (all non-BN layers) are transferred directly to the new client.
    2. The trainable BN parameters ({γ,β}\{\gamma, \beta\}) for the new client are initialized by computing the element-wise average of the local BN parameters across all participating training clients, while the running activation statistics (mean and variance) are computed directly on the new client's local data.

    On Morpho-MNIST datasets with distinct synthetic distortions evaluated as unseen clients:

    • Morpho-global (thinning/thickening perturbations): FedBN achieves 92.45% accuracy compared to 92.35% (FedProx) and 91.28% (FedAvg).
    • Morpho-local (swelling and fracture perturbations): FedBN achieves 94.61% accuracy compared to 94.31% (FedProx) and 93.55% (FedAvg).

Coverage note — None was omitted; all key theoretical assumptions, NTK convergence bounds, algorithms, benchmark evaluations, ablation studies on data size/local epochs, and transfer learning results are represented.

References

  1. 1.Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song. A convergence theory for deep learning via over-parameterization. In International Conference on Machine Learning, pp. 242–252. PMLR, 2019.
  2. 2.Mathieu Andreux, Jean Ogier du Terrail, Constance Beguier, and Eric W Tramel. Siloed federated learning for multi-centric histopathology datasets. In Domain Adaptation and Representation Transfer, and Distributed and Collaborative Learning, pp. 129–139. Springer, 2020.
  3. 3.Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang. Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks. arXiv preprint arXiv:1901.08584, 2019.
  4. 4.Daniel J Beutel, Taner Topal, Akhil Mathur, Xinchi Qiu, Titouan Parcollet, and Nicholas D Lane. Flower: A friendly federated learning research framework. arXiv preprint arXiv:2007.14390, 2020.
  5. 5.Daniel C Castro, Jeremy Tan, Bernhard Kainz, Ender Konukoglu, and Ben Glocker. Morpho-mnist: Quantitative assessment and diagnostics for representation learning. Journal of Machine Learning Research, 20(178):1–29, 2019.
  6. 6.Woong-Gi Chang, Tackgeun You, Seonguk Seo, Suha Kwak, and Bohyung Han. Domain-specific batch normalization for unsupervised domain adaptation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7354–7362, 2019.
  7. 7.Adriana Di Martino, Chao-Gan Yan, Qingyang Li, Erin Denio, Francisco X Castellanos, Kaat Alaerts, Jeffrey S Anderson, Michal Assaf, Susan Y Bookheimer, Mirella Dapretto, et al. The autism brain imaging data exchange: towards a large-scale evaluation of the intrinsic brain architecture in autism. Molecular psychiatry, 19(6):659, 2014.
  8. 8.Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh. Gradient descent provably optimizes over-parameterized neural networks. arXiv preprint arXiv:1810.02054, 2018.
  9. 9.Yonatan Dukler, Quanquan Gu, and Guido Montúfar. Optimization theory for relu neural networks trained with normalization layers, 2020.
  10. 10.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In International conference on machine learning, pp. 1180–1189. PMLR, 2015.
  11. 11.Boqing Gong, Yuan Shi, Fei Sha, and Kristen Grauman. Geodesic flow kernel for unsupervised domain adaptation. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 2066–2073. IEEE, 2012.
  12. 12.Google. TensorFlow Federated: Machine Learning on Decentralized Data, 2020. https://www.tensorflow.org/federated.
  13. 13.Gregory Griffin, Alex Holub, and Pietro Perona. Caltech-256 object category dataset, 2007.
  14. 14.Chaoyang He, Songze Li, Jinhyun So, Mi Zhang, Hongyi Wang, Xiaoyang Wang, Praneeth Vepakomma, Abhishek Singh, Hang Qiu, Li Shen, et al. Fedml: A research library and benchmark for federated machine learning. arXiv preprint arXiv:2007.13518, 2020.
  15. 15.Kevin Hsieh, Amar Phanishayee, Onur Mutlu, and Phillip B Gibbons. The non-iid data quagmire of decentralized machine learning. arXiv preprint arXiv:1910.00189, 2019.
  16. 16.Tzu-Ming Harry Hsu, Hang Qi, and Matthew Brown. Measuring the effects of non-identical data distribution for federated visual classification. arXiv preprint arXiv:1909.06335, 2019.
  17. 17.Jonathan J. Hull. A database for handwritten text recognition research. IEEE Transactions on pattern analysis and machine intelligence, 16(5):550–554, 1994.
  18. 18.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015.
  19. 19.Arthur Jacot, Franck Gabriel, and Clément Hongler. Neural tangent kernel: Convergence and generalization in neural networks. In Advances in neural information processing systems, pp. 8571–8580, 2018.
  20. 20.Peter Kairouz, H Brendan McMahan, Brendan Avent, Aurélien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Keith Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, et al. Advances and open problems in federated learning. arXiv preprint arXiv:1912.04977, 2019.
  21. 21.Michael Kamp and Linara Adilova. Distributed Learning Platform, 2020. https://github.com/fraunhofer-iais/dlplatform.
  22. 22.Michael Kamp, Linara Adilova, Joachim Sicking, Fabian Hüger, Peter Schlicht, Tim Wirtz, and Stefan Wrobel. Efficient decentralized deep learning by dynamic model averaging. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pp. 393–409. Springer, 2018.
  23. 23.Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank J Reddi, Sebastian U Stich, and Ananda Theertha Suresh. Scaffold: Stochastic controlled averaging for on-device federated learning. arXiv preprint arXiv:1910.06378, 2019.
  24. 24.Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann, Ming Zhou, and Klaus Neymeyr. Exponential convergence rates for batch normalization: The power of length-direction decoupling in non-convex optimization. In The 22nd International Conference on Artificial Intelligence and Statistics, pp. 806–815. PMLR, 2019.
  25. 25.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pp. 1097–1105, 2012.
  26. 26.Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998.
  27. 27.Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith. Federated learning: Challenges, methods, and future directions. IEEE Signal Processing Magazine, 37(3):50–60, 2020a.
  28. 28.Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith. Federated optimization in heterogeneous networks. In Conference on Machine Learning and Systems, 2020a, 2020b.
  29. 29.Xiang Li, Kaixuan Huang, Wenhao Yang, Shusen Wang, and Zhihua Zhang. On the convergence of fedavg on non-iid data. arXiv preprint arXiv:1907.02189, 2019.
  30. 30.Yanghao Li, Naiyan Wang, Jianping Shi, Jiaying Liu, and Xiaodi Hou. Revisiting batch normalization for practical domain adaptation. arXiv preprint arXiv:1603.04779, 2016.
  31. 31.Yanghao Li, Naiyan Wang, Jianping Shi, Xiaodi Hou, and Jiaying Liu. Adaptive batch normalization for practical domain adaptation. Pattern Recognition, 80:109–117, 2018.
  32. 32.Quande Liu, Qi Dou, Lequan Yu, and Pheng Ann Heng. Ms-net: Multi-site network for improving prostate segmentation with heterogeneous mri data. IEEE Transactions on Medical Imaging, 2020.
  33. 33.Ping Luo, Xinjiang Wang, Wenqi Shao, and Zhanglin Peng. Towards understanding regularization in batch normalization. arXiv preprint arXiv:1809.00846, 2018.
  34. 34.Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics, pp. 1273–1282, 2017.
  35. 35.Ari S Morcos, David GT Barrett, Neil C Rabinowitz, and Matthew Botvinick. On the importance of single directions for generalization. arXiv preprint arXiv:1803.06959, 2018.
  36. 36.Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu, and Andrew Y Ng. Reading digits in natural images with unsupervised feature learning, 2011.
  37. 37.Xingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang, Kate Saenko, and Bo Wang. Moment matching for multi-source domain adaptation. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1406–1415, 2019.
  38. 38.Sashank Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and H Brendan McMahan. Adaptive federated optimization. arXiv preprint arXiv:2003.00295, 2020.
  39. 39.Amirhossein Reisizadeh, Farzan Farnia, Ramtin Pedarsani, and Ali Jadbabaie. Robust federated learning: The case of affine distribution shifts. arXiv preprint arXiv:2006.08907, 2020.
  40. 40.Nicola Rieke, Jonny Hancox, Wenqi Li, Fausto Milletari, Holger R Roth, Shadi Albarqouni, Spyridon Bakas, Mathieu N Galtier, Bennett A Landman, Klaus Maier-Hein, et al. The future of digital health with federated learning. NPJ digital medicine, 3(1):1–7, 2020.
  41. 41.Theo Ryffel, Andrew Trask, Morten Dahl, Bobby Wagner, Jason Mancuso, Daniel Rueckert, and Jonathan Passerat-Palmbach. A generic framework for privacy preserving deep learning. arXiv preprint arXiv:1811.04017, 2018.
  42. 42.Kate Saenko, Brian Kulis, Mario Fritz, and Trevor Darrell. Adapting visual category models to new domains. In European conference on computer vision, pp. 213–226. Springer, 2010.
  43. 43.Tim Salimans and Durk P Kingma. Weight normalization: A simple reparameterization to accelerate training of deep neural networks. In Advances in neural information processing systems, pp. 901–909, 2016.
  44. 44.Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry. How does batch normalization help optimization? In Advances in Neural Information Processing Systems, pp. 2483–2493, 2018.
  45. 45.Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein. Training (overparametrized) neural networks in near-linear time. arXiv e-prints, pp. arXiv–2006, 2020.
  46. 46.Hongyi Wang, Mikhail Yurochkin, Yuekai Sun, Dimitris Papailiopoulos, and Yasaman Khazaeni. Federated learning with matched averaging. arXiv preprint arXiv:2002.06440, 2020.
  47. 47.Xinwei Zhang, Mingyi Hong, Sairaj Dhople, Wotao Yin, and Yang Liu. Fedpd: A federated learning framework with optimal rates and adaptivity to non-iid data. arXiv preprint arXiv:2005.11418, 2020.
  48. 48.Yue Zhao, Meng Li, Liangzhen Lai, Naveen Suda, Damon Civin, and Vikas Chandra. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582, 2018.

Citation

MLA
Li, X., et al. “FedBN: Federated Learning on Non-IID Features via Local Batch Normalization”. arXiv, 2021, http://arxiv.org/abs/2102.07623v2.
APA
Li, X., Jiang, M., Zhang, X., Kamp, M., & Dou, Q. (2021). FedBN: Federated Learning on Non-IID Features via Local Batch Normalization. arXiv. http://arxiv.org/abs/2102.07623v2
Chicago
Li, X., M. Jiang, X. Zhang, M. Kamp, and Q. Dou. 2021. “FedBN: Federated Learning on Non-IID Features via Local Batch Normalization”. arXiv. http://arxiv.org/abs/2102.07623v2.
Harvard
Li, X. et al. (2021) “FedBN: Federated Learning on Non-IID Features via Local Batch Normalization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2102.07623v2.
Vancouver
1. Li X, Jiang M, Zhang X, Kamp M, Dou Q (2021) FedBN: Federated Learning on Non-IID Features via Local Batch Normalization. arXiv

BibTeX

@article{li2021fedbn,
  title = {FedBN: Federated Learning on Non-IID Features via Local Batch Normalization},
  author = {Li, Xiaoxiao and Jiang, Meirui and Zhang, Xiaofei and Kamp, Michael and Dou, Qi},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2102.07623v2},
  eprint = {2102.07623}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors