On the Effectiveness of Partial Variance Reduction in Federated Learning with Heterogeneous Data

Bo LiMikkel N. SchmidtTommy S. AlstrømSebastian U. Stich

article2023CVPR61 citations

Proposes FedPVR, a communication-efficient federated learning method that mitigates client drift on non-IID data by applying variance reduction exclusively to the final classification layers, achieving faster convergence and higher accuracy with minimal overhead compared to full-model variance reduction.

Listen

Federated learning enables multiple institutions or edge devices to train machine learning models collaboratively without centralizing sensitive local data. However, in practical applications, data is often distributed heterogeneously across participating clients. Under these conditions, the standard baseline algorithm, Federated Averaging, suffers from client drift, causing slow convergence, poor accuracy, and high communication overhead across network rounds.

The article evaluates how data heterogeneity impacts individual layers of deep neural networks during distributed training. It demonstrates and validates a targeted optimization framework, named Partial Variance Reduction (FedPVR), designed to accelerate model convergence and improve classification accuracy under non-uniform data distributions while minimizing communication costs.

The authors analyzed gradient variance and model representations across deep architectures (VGG-11 and ResNet-8) using standard image benchmark datasets (CIFAR-10 and CIFAR-100). They tracked layer-by-layer update directions and measured representation alignment across clients. Based on empirical findings, they formulated an algorithm that applies statistical control variates to correct drift solely on the final classification layers, while allowing earlier feature extraction layers to update via standard local gradient descent. The framework was evaluated under various degrees of data imbalance across ten participating clients and supported by theoretical convergence proofs.

The investigation produced three main findings. First, gradient disagreement across clients is heavily concentrated in the final classification layers, while earlier feature extraction layers naturally maintain high alignment and extract useful representations even under severe data imbalance. Second, applying variance reduction exclusively to the final layers accelerated training speedup by 1.5 to 6.7 times compared to Federated Averaging to reach target accuracy levels. Third, the proposed method required transmitting only about 2.0 to 2.1 times the base model parameters per communication round, whereas full-model variance reduction methods require 4 times the base parameters, achieving comparable or superior top-1 accuracy at roughly half the extra communication bandwidth. Additionally, pairing the model with conformal prediction enabled users to guarantee high empirical coverage by slightly expanding prediction sets.

These findings indicate that full-network alignment constraints are counterproductive in deep federated learning. Allowing diversity in early feature extraction layers enables client models to learn richer representations, while enforcing uniformity in final classification layers prevents decision bias. This approach substantially lowers the energy and communication bandwidth requirements for resource-constrained edge deployments, mitigating the primary bottlenecks of distributed training.

Organizations implementing federated learning on deep neural networks should prioritize targeted classification-layer alignment over whole-model drift correction. For risk-sensitive domains such as medical diagnosis or hazard detection, decision-makers should integrate conformal prediction to provide certified reliability bounds. Future research should evaluate adaptive procedures for dynamically selecting which layers to align and investigate how network capacity and bottleneck designs influence client drift under real-world, non-participating, or cross-device edge conditions.

Cover for On the Effectiveness of Partial Variance Reduction in Federated Learning with Heterogeneous Data

Abstract

Data heterogeneity across clients is a key challenge in federated learning. Prior works address this by either aligning client and server models or using control variates to correct client model drift. Although these methods achieve fast convergence in convex or simple non-convex problems, the performance in over-parameterized models such as deep neural networks is lacking. In this paper, we first revisit the widely used FedAvg algorithm in a deep neural network to understand how data heterogeneity influences the gradient updates across the neural network layers. We observe that while the feature extraction layers are learned efficiently by FedAvg, the substantial diversity of the final classification layers across clients impedes the performance. Motivated by this, we propose to correct model drift by variance reduction only on the final layers. We demonstrate that this significantly outperforms existing benchmarks at a similar or lower communication cost. We furthermore provide proof for the convergence rate of our algorithm.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 2.1. Federated learning
  • 3. Method
  • 3.1. Problem statement
  • 3.2. Motivation
  • 2.2. Variance reduction
  • 2.3. Conformal prediction
  • 3.3. Classifier variance reduction
  • 3.4. Convergence rate
  • 4. Experimental setup
  • 5. Experimental results
  • 5.1. Communication efficiency and accuracy
  • 5.2. Conformal prediction
  • 5.3. Diversity and uniformity
  • 6. Conclusion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Partial Variance-Reduced Federated Learning (FedPVR) Algorithm

    algorithm

    Partial Variance-Reduced Federated Learning (FedPVR) addresses client drift in heterogeneous federated learning by applying stochastic variance reduction solely to designated layers (such as the final classification head) while updating the remaining layers with standard stochastic gradient descent (SGD).

    Let dd be the total number of model parameters, and let p∈{0,1}dp \in \{0, 1\}^d be a binary layer-selection mask containing v=∑j=1dpjv = \sum_{j=1}^d p_j non-zero entries (v≪dv \ll d). The parameter index set is partitioned into variance-reduced indices Ssvr:={j:pj=1}S_{\text{svr}} := \{j : p_j = 1\} and standard SGD indices Ssgd:={j:pj=0}S_{\text{sgd}} := \{j : p_j = 0\}. The server maintains a global model x∈Rdx \in \mathbb{R}^d and a server control variate c∈Rvc \in \mathbb{R}^v. Each client i∈{1,…,N}i \in \{1, \dots, N\} maintains a local model yi∈Rdy_i \in \mathbb{R}^d and a local control variate ci∈Rvc_i \in \mathbb{R}^v.

    Server initializes global model x∈Rdx \in \mathbb{R}^d, server control variate c=0∈Rvc = \mathbf{0} \in \mathbb{R}^v, global step size ηg\eta_g, and local step size ηl\eta_l
    Mask specification: p∈{0,1}dp \in \{0, 1\}^d, Ssgd←{j:pj=0}S_{\text{sgd}} \leftarrow \{j : p_j = 0\}, Ssvr←{j:pj=1}S_{\text{svr}} \leftarrow \{j : p_j = 1\}
    Initialize client control variates ci←0∈Rvc_i \leftarrow \mathbf{0} \in \mathbb{R}^v for all clients i∈[N]i \in [N]
    for round r=1r = 1 to RR do
        Server broadcasts xx and cc to all clients i∈[N]i \in [N]
        for each client i∈[N]i \in [N] in parallel do
            yi←xy_i \leftarrow x
            for step k=1k = 1 to KK do
                Compute mini-batch stochastic gradient gi(yi)g_i(y_i)
                yi,Ssgd←yi,Ssgd−ηlgi(yi)Ssgdy_{i, S_{\text{sgd}}} \leftarrow y_{i, S_{\text{sgd}}} - \eta_l g_i(y_i)_{S_{\text{sgd}}}
                yi,Ssvr←yi,Ssvr−ηl(gi(yi)Ssvr−ci+c)y_{i, S_{\text{svr}}} \leftarrow y_{i, S_{\text{svr}}} - \eta_l (g_i(y_i)_{S_{\text{svr}}} - c_i + c)
            end for
            ci←ci−c+1Kηl(xSsvr−yi,Ssvr)c_i \leftarrow c_i - c + \frac{1}{K \eta_l} (x_{S_{\text{svr}}} - y_{i, S_{\text{svr}}})
            Client ii sends updated local parameters yiy_i and control variate cic_i to server
        end for
        x←(1−ηg)x+ηgN∑i=1Nyix \leftarrow (1 - \eta_g) x + \frac{\eta_g}{N} \sum_{i=1}^N y_i
        c←1N∑i=1Ncic \leftarrow \frac{1}{N} \sum_{i=1}^N c_i
    end for
  2. Knowl 2 — Drift Diversity Metric in Federated Optimization

    definition

    In federated learning with NN participating clients at communication round rr, let xr−1∈Rdx^{r-1} \in \mathbb{R}^d denote the global server model at the beginning of round rr, and let yi,Kr∈Rdy_{i,K}^r \in \mathbb{R}^d denote the local model on client ii after KK local update steps. The update vector for client ii is defined as mir:=yi,Kr−xr−1m_i^r := y_{i,K}^r - x^{r-1}.

    The drift diversity ξr\xi^r across all NN clients at round rr is defined as:

    ξr:=∑i=1N∥mir∥2∥∑i=1Nmir∥2\xi^r := \frac{\sum_{i=1}^N \|m_i^r\|^2}{\|\sum_{i=1}^N m_i^r\|^2}

    When each client updates its model using KK steps of stochastic gradient descent with mini-batch gradients gi(yi,kr)g_i(y_{i,k}^r), the drift diversity can be expressed equivalently in terms of accumulated stochastic gradients:

    ξr=∑i=1N∥∑k=1Kgi(yi,kr)∥2∥∑i=1N∑k=1Kgi(yi,kr)∥2\xi^r = \frac{\sum_{i=1}^N \|\sum_{k=1}^K g_i(y_{i,k}^r)\|^2}{\|\sum_{i=1}^N \sum_{k=1}^K g_i(y_{i,k}^r)\|^2}

    A larger ξr\xi^r indicates higher directional divergence and orthogonality among client updates, reflecting pronounced client drift, whereas ξr≈1\xi^r \approx 1 indicates that client updates align in direction.

  3. Knowl 3 — Convergence Rates of FedPVR Under Smoothness and Partial Heterogeneity

    theoretical result

    Let each client loss function fi:Rd→Rf_i: \mathbb{R}^d \to \mathbb{R} be β\beta-smooth, stochastic gradients gi(x)=∇fi(x;Di)g_i(x) = \nabla f_i(x; \mathcal{D}_i) be unbiased with bounded variance E∥gi(x)−∇fi(x)∥2≤σ2\mathbb{E}\|g_i(x) - \nabla f_i(x)\|^2 \le \sigma^2, and NN clients perform KK local steps per round. Let ηg=N\eta_g = \sqrt{N} be the global step size, D:=∥x0−x∗∥2D := \|x^0 - x^*\|^2 be the initial distance to the optimal model x∗x^*, and F:=f(x0)−f∗F := f(x^0) - f^* be the initial suboptimality gap. For a binary mask p∈{0,1}dp \in \{0, 1\}^d, let ζ1−p2\zeta_{1-p}^2 and ζ^1−p2\hat{\zeta}_{1-p}^2 denote the un-reduced gradient heterogeneity constants on the parameter indices Ssgd={j:pj=0}S_{\text{sgd}} = \{j : p_j = 0\}.

    To achieve expected error ϵ\epsilon, the number of communication rounds RR required by FedPVR satisfies:

    • Strongly convex setting (parameter μ>0\mu > 0): For local step size ηl≤min⁡(180Kηgβ,2620μKηg)\eta_l \le \min\left(\frac{1}{80 K \eta_g \beta}, \frac{26}{20 \mu K \eta_g}\right),

    R=O~(σ2μNKϵ+ζ1−p2μϵ+βμ)R = \tilde{\mathcal{O}}\left(\frac{\sigma^2}{\mu N K \epsilon} + \frac{\zeta_{1-p}^2}{\mu \epsilon} + \frac{\beta}{\mu}\right)

    • General convex setting (parameter μ=0\mu = 0): For local step size ηl≤180Kηgβ\eta_l \le \frac{1}{80 K \eta_g \beta},

    R=O(σ2DKNϵ2+ζ1−p2Dϵ2+βDϵ+F)R = \mathcal{O}\left(\frac{\sigma^2 D}{K N \epsilon^2} + \frac{\zeta_{1-p}^2 D}{\epsilon^2} + \frac{\beta D}{\epsilon} + F\right)

    • Non-convex setting: For local step size ηl≤126Kηgβ\eta_l \le \frac{1}{26 K \eta_g \beta} and R≥1R \ge 1,

    R=O(βσ2FKNϵ2+βζ^1−p2FNϵ2+βFϵ)R = \mathcal{O}\left(\frac{\beta \sigma^2 F}{K N \epsilon^2} + \frac{\beta \hat{\zeta}_{1-p}^2 F}{N \epsilon^2} + \frac{\beta F}{\epsilon}\right)

    When p=1p = \mathbf{1} (full-model variance reduction), ζ1−p2=0\zeta_{1-p}^2 = 0 and ζ^1−p2=0\hat{\zeta}_{1-p}^2 = 0, recovering the SCAFFOLD convergence rate. When the gradient heterogeneity ζ1−p2\zeta_{1-p}^2 of the feature extraction layers is sufficiently small, FedPVR asymptotically matches the rate of centralized SGD with mini-batch size NKN K.

  4. Knowl 4 — Bounded Partial Gradient Heterogeneity Assumptions for FedPVR

    assumption

    Let f(x):=1N∑i=1Nfi(x)f(x) := \frac{1}{N} \sum_{i=1}^N f_i(x) be the global objective across NN client functions fi:Rd→Rf_i: \mathbb{R}^d \to \mathbb{R}, and let p∈{0,1}dp \in \{0, 1\}^d be a binary layer-selection mask where pj=1p_j = 1 for variance-reduced parameters and pj=0p_j = 0 for standard SGD parameters.

    1. Optimum Heterogeneity (ζ\zeta-Heterogeneity): In convex settings with global minimizer x∗x^*, the overall gradient variance across clients at the optimum is:

    ζ2:=1N∑i=1NE∥∇fi(x∗)∥2\zeta^2 := \frac{1}{N} \sum_{i=1}^N \mathbb{E}\|\nabla f_i(x^*)\|^2

    The partial heterogeneity of the un-reduced parameter block 1−p1-p at the optimum is defined as:

    ζ1−p2:=1N∑i=1N∥(1−p)⊙∇fi(x∗)∥2≤ζ2\zeta_{1-p}^2 := \frac{1}{N} \sum_{i=1}^N \|(1 - p) \odot \nabla f_i(x^*)\|^2 \le \zeta^2

    1. Global Heterogeneity (ζ^\hat{\zeta}-Heterogeneity): In non-convex settings, there exists a constant ζ^≥0\hat{\zeta} \ge 0 such that for all x∈Rdx \in \mathbb{R}^d:

    1N∑i=1NE∥∇fi(x)∥2≤ζ^2\frac{1}{N} \sum_{i=1}^N \mathbb{E}\|\nabla f_i(x)\|^2 \le \hat{\zeta}^2

    The partial heterogeneity of the un-reduced parameter block 1−p1-p satisfies:

    1N∑i=1N∥(1−p)⊙∇fi(x)∥2≤ζ^1−p2≤ζ^2\frac{1}{N} \sum_{i=1}^N \|(1 - p) \odot \nabla f_i(x)\|^2 \le \hat{\zeta}_{1-p}^2 \le \hat{\zeta}^2

  5. Knowl 5 — Layer-wise Representation Drift in Deep Neural Networks Under Data Heterogeneity

    empirical result

    When evaluating FedAvg on deep neural networks (e.g., VGG-11) trained on non-IID client partitions generated via a Dirichlet distribution with concentration parameter α=0.1\alpha = 0.1 versus near-IID data with α=100.0\alpha = 100.0:

    1. Early and intermediate layer alignment: Centered Kernel Alignment (CKA) feature representation similarity across clients for shallow and intermediate layers (e.g., layers 4 and 12) remains high and closely aligns between IID and non-IID conditions. This indicates that FedAvg learns shared feature representations in early feature extraction layers despite client data heterogeneity.
    2. Classification head divergence: CKA similarity drops substantially and drift diversity ξ\xi reaches its highest values in the final classification layers under heterogeneous data (α=0.1\alpha = 0.1). This demonstrates that client drift in deep neural networks is primarily concentrated in the classification layers rather than the feature extractor.
  6. Knowl 6 — Feature Diversity and Classifier Uniformity via Selective Layer Variance Reduction

    empirical result

    Varying the starting layer for stochastic variance reduction in a deep neural network (e.g., VGG-11 on CIFAR-100 with Dirichlet parameter α=1.0\alpha = 1.0) reveals a trade-off between representation diversity and classifier alignment:

    1. Applying variance reduction only to deeper classification layers (e.g., layers 16 to 20 in VGG-11, or the classifier in ResNet-8) aligns the client classifier parameters (lowering classifier drift diversity) while simultaneously increasing the drift diversity of the early feature extraction layers.
    2. The higher gradient diversity in early feature extraction layers allows clients to learn richer, distinct representations, which accelerates overall convergence compared to applying variance reduction across the entire network (SCAFFOLD, layers 0 to 20), where suppressing feature extractor diversity slows learning.
    3. Activating variance reduction only on deeper layers (starting layer index ≥10\ge 10) achieves the fastest convergence speedup, whereas applying no variance reduction (FedAvg) achieves the slowest convergence and lowest accuracy under non-IID data.
  7. Knowl 7 — Communication Round Speedup of FedPVR Compared to Baseline FL Algorithms

    data/table

    The required number of communication rounds (and speedup factor relative to FedAvg in parentheses) to reach target Top-1 test accuracy (66% for CIFAR-10, 44% for CIFAR-100) across 10 clients with Dirichlet data heterogeneity α∈{0.1,0.5,1.0}\alpha \in \{0.1, 0.5, 1.0\} using VGG-11 and ResNet-8 architectures (10 local epochs per round, batch size 256):

    CIFAR10 (66%) CIFAR100 (44%)
    α=0.1\alpha=0.1 α=0.5\alpha=0.5 α=0.1\alpha=0.1 α=1.0\alpha=1.0
    Algorithm VGG-11 ResNet-8 VGG-11 ResNet-8 VGG-11 ResNet-8 VGG-11 ResNet-8
    FedAvg 55 (1.0x) 90 (1.0x) 15 (1.0x) 15 (1.0x) 100+ (1.0x) 100+ (1.0x) 80 (1.0x) 56 (1.0x)
    FedProx 52 (1.1x) 75 (1.2x) 16 (0.9x) 20 (0.8x) 100+ (1.0x) 100+ (1.0x) 80 (1.0x) 59 (0.9x)
    SCAFFOLD 39 (1.4x) 57 (1.6x) 14 (1.0x) 9 (1.7x) 80 (>1.3x) 61 (>1.6x) 36 (2.2x) 25 (2.2x)
    FedDyn 27 (2.0x) 67 (1.3x) 15 (1.0x) 34 (0.4x) 80+ (-) 80+ (-) 24 (3.3x) 51 (1.1x)
    FedPVR (Ours) 27 (2.0x) 50 (1.8x) 9 (1.6x) 5 (3.0x) 37 (>2.7x) 66 (>1.5x) 12 (6.7x) 15 (3.7x)

    FedPVR achieves a 1.5×1.5\times to 6.7×6.7\times round speedup over FedAvg and consistently requires fewer communication rounds than FedProx, SCAFFOLD, and FedDyn to reach the target accuracy.

  8. Knowl 8 — Top-1 Test Accuracy and Per-Round Communication Parameter Multipliers

    data/table

    Top-1 test accuracy (%) achieved after 80 communication rounds (10 local epochs per round, totaling 800 local epochs) on CIFAR-10 and CIFAR-100 across VGG-11 and ResNet-8 architectures with Dirichlet non-IID parameter α∈{0.1,0.5,1.0}\alpha \in \{0.1, 0.5, 1.0\}, alongside the number of parameter copies communicated per round between client and server:

    VGG-11 ResNet-8
    CIFAR10 CIFAR100 Comm. CIFAR10 CIFAR100 Comm.
    Algorithm α=0.1\alpha=0.1 α=0.5\alpha=0.5 α=0.1\alpha=0.1 α=1.0\alpha=1.0 Params α=0.1\alpha=0.1 α=0.5\alpha=0.5 α=0.1\alpha=0.1 α=1.0\alpha=1.0 Params
    Centralised 87.5 - 56.3 - - 83.4 - 56.8 - -
    FedAvg 69.3 80.9 34.3 45.0 2x 64.9 79.1 38.8 47.0 2x
    FedProx 72.1 80.4 35.0 43.2 2x 66.1 77.9 42.0 47.2 2x
    SCAFFOLD 74.1 83.5 43.4 50.6 4x 66.6 80.3 43.8 52.3 4x
    FedDyn 77.4 80.1 43.8 45.2 2x 63.8 72.9 36.4 48.1 2x
    FedPVR (Ours) 78.2 84.9 43.5 58.0 2.1x 69.3 83.6 43.5 52.3 2.02x

    FedPVR achieves top-1 accuracy comparable to or exceeding centralized learning under moderate heterogeneity (alpha=0.5\\alpha=0.5 on CIFAR10, alpha=1.0\\alpha=1.0 on CIFAR100) while transmitting only 2.02×2.02\times to 2.1×2.1\times parameter copies per round, compared to 4×4\times required by full-model variance reduction in SCAFFOLD.

  9. Knowl 9 — Conformal Prediction Integration for Heterogeneous Federated Learning

    model/method

    To compensate for accuracy degradation under extreme data heterogeneity (e.g., Dirichlet α=0.1\alpha = 0.1), conformal prediction is integrated as a post-processing step on the server model without requiring local model retraining or client data sharing.

    Conformal prediction outputs an adaptive predictive set guaranteed to contain the true class label with a user-specified marginal coverage level 1−αconf1 - \alpha_{\text{conf}}. In federated settings with VGG-11 and ResNet-8 on CIFAR-10 and CIFAR-100:

    1. By slightly increasing the average predictive set size (e.g., from 1 to 1.5–2.0 classes on CIFAR-10), the empirical coverage of the federated server model matches the top-1 accuracy of centralized training.
    2. FedPVR reaches centralized-level empirical coverage with smaller average predictive set sizes than FedAvg, FedProx, SCAFFOLD, and FedDyn across both architectures.

Coverage note — Detailed mathematical proofs of Theorem 1 and intermediate lemmas in Appendix B were deliberately omitted in accordance with the rule excluding derivations.

References

  1. 1.Durmus Alp Emre Acar, Yue Zhao, Ramon Matas Navarro, Matthew Mattina, Paul N. Whatmough, and Venkatesh Saligrama. Federated learning based on dynamic regularization. CoRR, abs/2111.04263, 2021. 1, 2, 6, 26
  2. 2.Dan Alistarh, Jerry Li, Ryota Tomioka, and Milan Vojnovic. QSGD: randomized quantization for communication-optimal stochastic gradient descent. CoRR, abs/1610.02132, 2016. 3
  3. 3.Anastasios Nikolas Angelopoulos, Stephen Bates, Michael Jordan, and Jitendra Malik. Uncertainty sets for image classifiers using conformal prediction. In International Conference on Learning Representations, 2021. 2, 3, 6
  4. 4.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton. A simple framework for contrastive learning of visual representations. CoRR, abs/2002.05709, 2020. 3
  5. 5.Liam Collins, Hamed Hassani, Aryan Mokhtari, and Sanjay Shakkottai. Fedavg with fine tuning: Local updates lead to representation learning, 2022. 2
  6. 6.Aaron Defazio, Francis R. Bach, and Simon Lacoste-Julien. SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives. CoRR, abs/1407.0202, 2014. 3
  7. 7.Aaron Defazio and Leon Bottou. On the ineffectiveness of variance reduced optimization for deep learning. CoRR, abs/1812.04529, 2018. 2, 3, 5
  8. 8.Aritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini, and Panos Kalnis. On the discrepancy between the theoretical analysis and practical implementations of compressed communication for distributed deep learning. CoRR, abs/1911.08250, 2019. 3
  9. 9.Liang Gao, Huazhu Fu, Li Li, Yingwen Chen, Ming Xu, and Cheng-Zhong Xu. Feddc: Federated learning with non-iid data via local drift decoupling and correction. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10102–10111. IEEE, 2022. 2
  10. 10.Malka N. Halgamuge, Moshe Zukerman, Kotagiri Ramamohanarao, and Hai Le Vu. An estimation of sensor energy consumption. Progress in Electromagnetics Research B, 12:259–295, 2009. 1, 2, 5
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. CoRR, abs/1512.03385, 2015. 2
  12. 12.Rie Johnson and Tong Zhang. Accelerating stochastic gradient descent using predictive variance reduction. In NIPS, 2013. 3
  13. 13.Peter Kairouz, H. Brendan McMahan, Brendan Avent, Aurelien Bellet, Mehdi Bennis, Arjun Nitin Bhagoji, Kallista A. Bonawitz, Zachary Charles, Graham Cormode, Rachel Cummings, Rafael G. L. D’Oliveira, Salim El Rouayheb, David Evans, Josh Gardner, Zachary Garrett, AdriaGascon, Badih Ghazi, Phillip B. Gibbons, Marco Gruteser, Za¨ıd Harchaoui, Chaoyang He, Lie He, Zhouyuan Huo, Ben Hutchinson, Justin Hsu, Martin Jaggi, Tara Javidi, Gauri Joshi, Mikhail Khodak, Jakub Konecnˇ y, Aleksandra Korolova, Farinaz Koushanfar, Sanmi Koyejo, Tancrede Lepoint, Yang Liu, Prateek Mittal, Mehryar Mohri, Richard Nock, Ayfer Ozg ¨ ur, Rasmus Pagh, Mariana Raykova, Hang ¨ Qi, Daniel Ramage, Ramesh Raskar, Dawn Song, Weikang Song, Sebastian U. Stich, Ziteng Sun, Ananda Theertha Suresh, Florian Tramer, Praneeth Vepakomma, Jianyu Wang, Li Xiong, Zheng Xu, Qiang Yang, Felix X. Yu, Han Yu, and Sen Zhao. Advances and open problems in federated learning. CoRR, abs/1912.04977, 2019. 1, 2
  14. 14.Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh. SCAFFOLD: Stochastic controlled averaging for federated learning. In Hal Daume III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 5132–5143. PMLR, 13–18 Jul 2020. 1, 2, 3, 4, 5, 6, 11, 12, 13, 21
  15. 15.Ahmed Khaled, Konstantin Mishchenko, and Peter Richtarik. Better communication complexity for local SGD. CoRR, abs/1909.04746, 2019. 5
  16. 16.Anastasia Koloskova, Nicolas Loizou, Sadra Boreiri, Martin Jaggi, and Sebastian Stich. A unified theory of decentralized SGD with changing topology and local updates. In Hal Daume III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 5381–5393. PMLR, 13–18 Jul 2020. 5, 11, 13
  17. 17.Jakub Konecnˇ y, H. Brendan McMahan, Daniel Ramage, and Peter Richtarik. Federated optimization: Distributed machine learning for on-device intelligence. CoRR, abs/1610.02527, 2016. 1, 2
  18. 18.Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey E. Hinton. Similarity of neural network representations revisited. CoRR, abs/1905.00414, 2019. 4
  19. 19.Alex Krizhevsky. Learning multiple layers of features from tiny images. 2009. 2, 6
  20. 20.Qinbin Li, Bingsheng He, and Dawn Song. Model-contrastive federated learning. CoRR, abs/2103.16257, 2021. 1, 3
  21. 21.Zhize Li, Dmitry Kovalev, Xun Qian, and Peter Richtarik. Acceleration for compressed gradient descent in distributed and federated optimization, 2020. 3
  22. 22.Tao Lin, Lingjing Kong, Sebastian U. Stich, and Martin Jaggi. Ensemble distillation for robust model fusion in federated learning. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and Hsuan-Tien Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. 6, 25, 26
  23. 23.Ilya Loshchilov and Frank Hutter. SGDR: stochastic gradient descent with restarts. CoRR, abs/1608.03983, 2016. 6
  24. 24.Mi Luo, Fei Chen, Dapeng Hu, Yifan Zhang, Jian Liang, and Jiashi Feng. No fear of heterogeneity: Classifier calibration for federated learning with non-iid data. CoRR, abs/2106.05001, 2021. 1, 2, 3, 4, 6
  25. 25.H. Brendan McMahan, Eider Moore, Daniel Ramage, and Blaise Aguera y Arcas. Federated learning of deep networks ¨ using model averaging. CoRR, abs/1602.05629, 2016. 2, 6
  26. 26.Konstantin Mishchenko, Eduard Gorbunov, Martin Takac, ´ and Peter Richtarik. Distributed learning with compressed ´ gradient differences. CoRR, abs/1901.09269, 2019. 3
  27. 27.Konstantin Mishchenko, Grigory Malinovsky, Sebastian Stich, and Peter Richtarik. ProxSkip: Yes! local gradient steps provably lead to communication acceleration! finally! International Conference on Machine Learning (ICML), 2022. 2
  28. 28.Thao Nguyen, Maithra Raghu, and Simon Kornblith. Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth. CoRR, abs/2010.15327, 2020. 2, 4
  29. 29.Jaehoon Oh, Sangmook Kim, and Se-Young Yun. Fedbabu: Towards enhanced representation for federated image classification. CoRR, abs/2106.06042, 2021. 2
  30. 30.Yaniv Romano, Matteo Sesia, and Emmanuel J. Candes. ` Classification with valid and adaptive coverage. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPS’20, Red Hook, NY, USA, 2020. Curran Associates Inc. 3
  31. 31.Anit Kumar Sahu, Tian Li, Maziar Sanjabi, Manzil Zaheer, Ameet Talwalkar, and Virginia Smith. On the convergence of federated optimization in heterogeneous networks. CoRR, abs/1812.06127, 2018. 1, 2, 6, 26
  32. 32.Ohad Shamir, Nathan Srebro, and Tong Zhang. Communication efficient distributed optimization using an approximate newton-type method. CoRR, abs/1312.7853, 2013. 1, 2, 3
  33. 33.Micah J. Sheller, Brandon Edwards, G. Anthony Reina, Jason Martin, Sarthak Pati, Aikaterini Kotrotsou, Mikhail Milchenko, Weilin Xu, Daniel Marcus, Rivka R. Colen, and Spyridon Bakas. Federated learning in medicine: facilitating multi-institutional collaborations without sharing patient data. Scientific Reports, 10(1):12598, Jul 2020. 1
  34. 34.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In International Conference on Learning Representations, 2015. 2
  35. 35.Sebastian U. Stich. Unified optimal analysis of the (stochastic) gradient method. CoRR, abs/1907.04232, 2019. 5, 11, 14
  36. 36.Sebastian U. Stich, Jean-Baptiste Cordonnier, and Martin Jaggi. Sparsified SGD with memory. In Advances in Neural Information Processing Systems 31 (NeurIPS), pages 4452–4463. Curran Associates, Inc., 2018. 3
  37. 37.Sebastian U. Stich and Sai Praneeth Karimireddy. The error-feedback framework: Better rates for SGD with delayed gradients and compressed communication. CoRR, abs/1909.05350, 2019. 15
  38. 38.Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In Yoshua Bengio and Yann LeCun, editors, 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference Track Proceedings, 2014. 2, 5
  39. 39.F Varno, M Saghayi, L Rafiee, S Gupta, S Matwin, and M Havaei. Minimizing client drift in federated learning via adaptive bias estimation. ArXiv, abs/2204.13170, 2022. 2
  40. 40.Jianyu Wang, Zachary Charles, Zheng Xu, Gauri Joshi, H. Brendan McMahan, Blaise Aguera y Arcas, Maruan ¨ Al-Shedivat, Galen Andrew, Salman Avestimehr, Katharine Daly, Deepesh Data, Suhas N. Diggavi, Hubert Eichner, Advait Gadhikar, Zachary Garrett, Antonious M. Girgis, Filip Hanzely, Andrew Hard, Chaoyang He, Samuel Horvath, Zhouyuan Huo, Alex Ingerman, Martin Jaggi, ´ Tara Javidi, Peter Kairouz, Satyen Kale, Sai Praneeth Karimireddy, Jakub Konecnˇ y, Sanmi Koyejo, Tian Li, ´ Luyang Liu, Mehryar Mohri, Hang Qi, Sashank J. Reddi, Peter Richtarik, Karan Singhal, Virginia Smith, Mahdi ´ Soltanolkotabi, Weikang Song, Ananda Theertha Suresh, Sebastian U. Stich, Ameet Talwalkar, Hongyi Wang, Blake E. Woodworth, Shanshan Wu, Felix X. Yu, Honglin Yuan, Manzil Zaheer, Mi Zhang, Tong Zhang, Chunxiang Zheng, Chen Zhu, and Wennan Zhu. A field guide to federated optimization. CoRR, abs/2107.06917, 2021. 2
  41. 41.Yaodong Yu, Alexander Wei, Sai Praneeth Karimireddy, Yi Ma, and Michael I. Jordan. TCT: Convexifying federated learning using bootstrapped neural tangent kernels, 2022. 2, 5, 6
  42. 42.Haoyu Zhao, Zhize Li, and Peter Richtarik. Fedpage: A fast ´ local stochastic gradient method for communication-efficient federated learning. CoRR, abs/2108.04755, 2021. 2

Citation

MLA
Li, B., et al. “On the Effectiveness of Partial Variance Reduction in Federated Learning with Heterogeneous Data”. arXiv, 2022, http://arxiv.org/abs/2212.02191v2.
APA
Li, B., Schmidt, M. N., Alstrøm, T. S., & Stich, S. U. (2022). On the effectiveness of partial variance reduction in federated learning with heterogeneous data. arXiv. http://arxiv.org/abs/2212.02191v2
Chicago
Li, B., M. N. Schmidt, T. S. Alstrøm, and S. U. Stich. 2022. “On the Effectiveness of Partial Variance Reduction in Federated Learning with Heterogeneous Data”. arXiv. http://arxiv.org/abs/2212.02191v2.
Harvard
Li, B. et al. (2022) “On the effectiveness of partial variance reduction in federated learning with heterogeneous data”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.02191v2.
Vancouver
1. Li B, Schmidt MN, Alstrøm TS, Stich SU (2022) On the effectiveness of partial variance reduction in federated learning with heterogeneous data. arXiv

BibTeX

@article{li2022the,
  title = {On the effectiveness of partial variance reduction in federated learning with heterogeneous data},
  author = {Li, Bo and Schmidt, Mikkel N. and Alstrøm, Tommy S. and Stich, Sebastian U.},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.02191v2},
  eprint = {2212.02191}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE