Communication-Efficient Learning of Deep Networks from Decentralized Data

H. B. McMahanEider MooreDaniel RamageS. HampsonB. A. Y. Arcas

article2016AISTATS27,176 citations

Introduces federated learning and the FederatedAveraging algorithm to train deep networks directly on decentralized, privacy-sensitive devices, reducing communication rounds by up to two orders of magnitude while successfully handling unbalanced and non-IID data.

Listen

Modern mobile devices generate vast amounts of valuable data that can power intelligent applications such as next-word text prediction and automated photo sorting. However, centralizing this data in traditional data centers creates severe privacy risks, regulatory concerns, and bandwidth bottlenecks. The article investigates federated learning, a decentralized machine learning framework where user data remains stored locally on individual devices. Instead of uploading raw data, devices perform local model training and transmit only minimal model updates to a central coordinating server for aggregation.

The article set out to evaluate whether deep neural networks can be trained effectively across decentralized, highly variable client devices while dramatically reducing the network communication rounds required to reach target model performance.

To demonstrate feasibility, the authors introduced the FederatedAveraging algorithm and conducted comprehensive empirical evaluations using more than 2,000 individual model runs. The evaluation covered five distinct neural network architectures across four benchmark and real-world datasets, including image recognition tasks (MNIST and CIFAR-10) and language modeling tasks (Shakespeare plays and a large-scale social network dataset comprising over 500,000 users). The experiments explicitly accounted for key federated optimization challenges: highly unbalanced sample sizes across devices, non-identical data distributions per user, and severe communication bandwidth limitations.

The key findings demonstrate that FederatedAveraging drastically reduces network traffic and handles real-world data skew effectively. First, the algorithm reduces required communication rounds by 10x to 100x compared to baseline synchronized stochastic gradient descent by performing multiple training passes locally on devices between server syncs. Second, the method proved highly robust on non-identical and unbalanced data partitions, achieving a 95x communication speedup on real-world text role distributions and reaching target accuracy on highly fragmented image datasets without diverging. Third, on a large-scale real-world next-word prediction task across 500,000 clients, FederatedAveraging achieved the target accuracy in only 35 rounds compared to 820 rounds for the baseline (a 23x reduction) while exhibiting lower test variance. Fourth, client parallelism exhibited diminishing returns; selecting a modest subset of available devices (such as 10% per round) provided an optimal balance of convergence speed and network efficiency.

These findings have strong practical implications for technology leaders. By keeping raw training data on client devices, organizations can drastically reduce cloud storage costs, limit data privacy attack surfaces, and comply with data minimization principles. Furthermore, by trading readily available on-device compute power for expensive network communication, organizations can deploy sophisticated artificial intelligence models over mobile networks without overwhelming user bandwidth or server infrastructure.

Organizations developing mobile-centric machine learning applications should prioritize the adoption of federated averaging techniques. Teams should initially tune local minibatch sizes and modest local training epochs to maximize on-device compute before scaling network communication. However, engineering teams should exercise caution: training for excessive local epochs without server synchronization can cause model divergence in later training stages, making adaptive local epoch schedules advisable. While federated learning provides structural privacy benefits, future deployments requiring strict formal privacy guarantees should combine this framework with differential privacy and secure multiparty aggregation protocols.

arXiv: 1602.05629
Cover for Communication-Efficient Learning of Deep Networks from Decentralized Data

Abstract

Modern mobile devices have access to a wealth of data suitable for learning models, which in turn can greatly improve the user experience on the device. For example, language models can improve speech recognition and text entry, and image models can automatically select good photos. However, this rich data is often privacy sensitive, large in quantity, or both, which may preclude logging to the data center and training there using conventional approaches. We advocate an alternative that leaves the training data distributed on the mobile devices, and learns a shared model by aggregating locally-computed updates. We term this decentralized approach Federated Learning.

We present a practical method for the federated learning of deep networks based on iterative model averaging, and conduct an extensive empirical evaluation, considering five different model architectures and four datasets. These experiments demonstrate the approach is robust to the unbalanced and non-IID data distributions that are a defining characteristic of this setting. Communication costs are the principal constraint, and we show a reduction in required communication rounds by 10-100x as compared to synchronized stochastic gradient descent.

Table of Contents

  • 1 Introduction
  • 2 The FederatedAveraging Algorithm
  • 3 Experimental Results
  • 4 Conclusions and Future Work
  • References
  • A Supplemental Figures and Tables

Knowls

  1. Knowl 1 — Federated Optimization Problem Formulation and Characteristics

    definition

    Federated optimization considers learning a model with parameter vector w∈Rdw \in \mathbb{R}^d from data distributed across a loose federation of KK decentralized clients. The global objective function is formulated as a finite-sum problem:

    min⁡w∈Rdf(w)wheref(w)=∑k=1KnknFk(w)\min_{w \in \mathbb{R}^d} f(w) \quad \text{where} \quad f(w) = \sum_{k=1}^K \frac{n_k}{n} F_k(w)

    where n=∑k=1Knkn = \sum_{k=1}^K n_k is the total number of training data points across all clients, nk=∣Pk∣n_k = |\mathcal{P}_k| is the number of data samples on client kk, Pk\mathcal{P}_k is the set of sample indices on client kk, and the local objective Fk(w)F_k(w) on client kk is defined by:

    Fk(w)=1nk∑i∈Pkfi(w)F_k(w) = \frac{1}{n_k} \sum_{i \in \mathcal{P}_k} f_i(w)

    Here, fi(w)=ℓ(xi,yi;w)f_i(w) = \ell(x_i, y_i; w) is the loss computed on example (xi,yi)(x_i, y_i) using model parameters ww.

    Federated optimization differs from typical data center distributed optimization along four key properties:

    1. Non-IID data: The local dataset on a given client reflects that specific user's behavior, meaning EPk[Fk(w)]≠f(w)\mathbb{E}_{\mathcal{P}_k}[F_k(w)] \neq f(w); local distributions do not represent the overall population distribution.
    2. Unbalanced data: Different clients generate and store widely varying amounts of data nkn_k.
    3. Massively distributed: The total number of participating clients KK is typically much larger than the average number of examples per client n/Kn/K.
    4. Limited communication: Clients communicate over constrained, intermittent network connections (such as mobile upload speeds of 1 MB/s or less), making communication rounds significantly more expensive than local on-device computation.
  2. Knowl 2 — The FederatedAveraging Algorithm

    algorithm

    The FederatedAveraging (FedAvg\text{FedAvg}) algorithm iteratively computes local stochastic gradient descent updates on client devices and aggregates them via weighted parameter averaging at a central server. It is governed by three primary hyperparameters: C∈(0,1]C \in (0, 1], the fraction of clients randomly sampled per round; EE, the number of training passes (epochs) each selected client performs over its local data per round; and BB, the local minibatch size used for client gradient updates (where B=∞B = \infty indicates processing the entire local dataset as a single batch). When E=1E = 1 and B=∞B = \infty, the algorithm reduces to Federated Stochastic Gradient Descent (FedSGD\text{FedSGD}).

    Algorithm: FederatedAveraging (FedAvg)
    Server executes:
        initialize w0w_0
        for each round t=1,2,…t = 1, 2, \dots do
            m←max⁡(C⋅K,1)m \leftarrow \max(C \cdot K, 1)
            St←S_t \leftarrow (random set of mm clients chosen without replacement)
            for each client k∈Stk \in S_t in parallel do
                wt+1k←ClientUpdate(k,wt)w_{t+1}^k \leftarrow \text{ClientUpdate}(k, w_t)
            mt←∑k∈Stnkm_t \leftarrow \sum_{k \in S_t} n_k
            wt+1←∑k∈Stnkmtwt+1kw_{t+1} \leftarrow \sum_{k \in S_t} \frac{n_k}{m_t} w_{t+1}^k
    ClientUpdate(k,wk, w): // Run on client kk
        B←\mathcal{B} \leftarrow (split local dataset Pk\mathcal{P}_k into minibatches of size BB)
        for each local epoch ii from 1 to EE do
            for batch b∈Bb \in \mathcal{B} do
                w←w−η∇ℓ(w;b)w \leftarrow w - \eta \nabla \ell(w; b)
        return ww to server

    For a client kk with nkn_k local samples, the number of local gradient updates performed per round is uk=E⌈nk/B⌉u_k = E \lceil n_k / B \rceil.

  3. Knowl 3 — Model Parameter Averaging under Shared vs Independent Initialization

    empirical result

    In non-convex neural network optimization, naive averaging of weights from independently trained models behaves fundamentally differently depending on initialization:

    1. Independent initializations: When two identical neural network architectures are trained via SGD on distinct, non-overlapping data partitions starting from different random initial weight vectors ww and w′w', linearly interpolating their parameters as θw+(1−θ)w′\theta w + (1-\theta) w' for θ∈[0,1]\theta \in [0, 1] encounters a severe loss barrier, producing models with significantly higher total training loss than either parent model.
    2. Shared initialization: When the two models start from an identical initial parameter configuration wtw_t before training locally on separate data partitions, the simple linear average 12(w+w′)\frac{1}{2}(w + w') achieves a dramatically lower training loss on the combined dataset than the loss achieved by either model trained independently.

    This behavior indicates that synchronized iterations of local training starting from a shared global model state wtw_t remain within a shared optimization basin amenable to coordinate averaging.

  4. Knowl 4 — Experimental Suite for Decentralized Vision and Language Benchmarks

    experimental setup

    The empirical evaluation of federated optimization comprises four learning tasks across five neural architectures:

    1. MNIST 2NN: A 2-hidden-layer multilayer perceptron with 200 ReLU units per layer (199,210 total parameters) for MNIST digit classification.
    2. MNIST CNN: A convolutional network consisting of two 5×55 \times 5 convolution layers (32 channels then 64 channels, each followed by 2×22 \times 2 max pooling), a 512-unit fully connected ReLU layer, and a softmax output layer (1,663,370 parameters).
      • MNIST IID partition: 60,000 examples shuffled and uniformly assigned across 100 clients (600 examples per client).
      • MNIST Pathological Non-IID partition: 60,000 examples sorted by digit label, grouped into 200 shards of 300 examples, and distributed such that each of the 100 clients receives exactly 2 shards (at most 2 distinct digits per client).
    3. Shakespeare Character-Level LSTM: A dataset created from The Complete Works of William Shakespeare, where each client corresponds to a speaking role with at least two lines (1,146 clients; 3,564,579 train characters and 870,014 test characters partitioned temporally per role). The model embeds characters into an 8-dimensional space, followed by 2 LSTM layers of 256 units each and a character softmax (866,578 parameters), unrolled for 80 characters.
    4. CIFAR-10 CNN: 50,000 training and 10,000 test images partitioned uniformly across 100 clients (500 train, 100 test per client). The architecture uses two convolutional layers, two fully connected layers, and a linear output layer (approx. 10610^6 parameters).
    5. Large-Scale Word-Level Language Model: 10 million public social media posts grouped by author (>500,000 clients; capped at 5,000 words per client; test set of 10510^5 posts from held-out authors). The architecture is a 256-node LSTM over a 10,000-word vocabulary with 192-dimensional co-trained input/output embeddings (4,950,544 total parameters), unrolled for 10 words.
  5. Knowl 5 — Effect of Client Parallelism Fraction on Convergence Speed

    data/table

    Varying the client fraction CC controls the number of clients participating in parallel in each communication round. The table below presents the number of communication rounds required to reach a target test accuracy (97% for MNIST 2NN with E=1E=1; 99% for MNIST CNN with E=5E=5) across different values of CC and local batch size BB, along with the speedup relative to the single-client baseline (C=0.0C=0.0, selecting m=1m=1 client per round):

    2NN IID Non-IID
    CC B=∞B = \infty B=10B = 10 B=∞B = \infty B=10B = 10
    0.0 1455 316 4278 3275
    0.1 1474 (1.0)× 87 (3.6)× 1796 (2.4)× 664 (4.9)×
    0.2 1658 (0.9)× 77 (4.1)× 1528 (2.8)× 619 (5.3)×
    0.5 — 75 (4.2)× — 443 (7.4)×
    1.0 — 70 (4.5)× — 380 (8.6)×
    CNN, E=5E=5 IID Non-IID
    0.0 387 50 1181 956
    0.1 339 (1.1)× 18 (2.8)× 1100 (1.1)× 206 (4.6)×
    0.2 337 (1.1)× 18 (2.8)× 978 (1.2)× 200 (4.8)×
    0.5 164 (2.4)× 18 (2.8)× 1067 (1.1)× 261 (3.7)×
    1.0 246 (1.6)× 16 (3.1)× — 97 (9.9)×

    When full batch gradient steps are used (B=∞B = \infty), scaling CC yields negligible speedup. When combined with local minibatches (B=10B = 10), setting C≥0.1C \ge 0.1 provides substantial communication speedups, but increasing CC further beyond 0.10.1 yields diminishing returns.

  6. Knowl 6 — Communication Reduction via Increased Local Client Computation

    data/table

    Increasing local computation on each client—by increasing local epochs EE or reducing local batch size BB—increases the expected number of local updates per round u=nEKBu = \frac{n E}{K B} and dramatically reduces the total number of communication rounds needed to achieve target accuracy compared to FedSGD\text{FedSGD} (E=1,B=∞,u=1E=1, B=\infty, u=1).

    Model EE BB uu IID Rounds (Speedup) Non-IID Rounds (Speedup)
    MNIST CNN (99% Acc) 1 ∞\infty 1 626 483
    5 ∞\infty 5 179 (3.5)× 1000 (0.5)×
    1 50 12 65 (9.6)× 600 (0.8)×
    20 ∞\infty 20 234 (2.7)× 672 (0.7)×
    1 10 60 34 (18.4)× 350 (1.4)×
    5 50 60 29 (21.6)× 334 (1.4)×
    20 50 240 32 (19.6)× 426 (1.1)×
    5 10 300 20 (31.3)× 229 (2.1)×
    20 10 1200 18 (34.8)× 173 (2.8)×
    Shakespeare LSTM (54% Acc) 1 ∞\infty 1.0 2488 3906
    1 50 1.5 1635 (1.5)× 549 (7.1)×
    5 ∞\infty 5.0 613 (4.1)× 597 (6.5)×
    1 10 7.4 460 (5.4)× 164 (23.8)×
    5 50 7.4 401 (6.2)× 152 (25.7)×
    5 10 37.1 192 (13.0)× 41 (95.3)×

    For MNIST CNN, increasing local computation yields a 35×35\times speedup on IID data and a 2.8×2.8\times speedup on pathological non-IID data. For Shakespeare LSTM, increased computation achieves a 13×13\times speedup on balanced IID data and a 95.3×95.3\times speedup on the unbalanced, non-IID data by role.

  7. Knowl 7 — Communication Efficiency on CIFAR-10 Benchmark

    data/table

    On the CIFAR-10 image classification task (100 clients, 500 training examples per client, balanced IID partition), FedAvg\text{FedAvg} (C=0.1,E=5,B=50C=0.1, E=5, B=50) was evaluated against baseline centralized SGD\text{SGD} (minibatch size 100 on the full dataset, where each minibatch update is treated as equivalent to one communication round) and FedSGD\text{FedSGD} (C=0.1C=0.1).

    Target Accuracy 80% 82% 85%
    SGD\text{SGD} (Minibatch 100) 18,000 (—) 31,000 (—) 99,000 (—)
    FedSGD\text{FedSGD} (C=0.1C=0.1) 3,750 (4.8)× 6,600 (4.7)× N/A (—)
    FedAvg\text{FedAvg} (C=0.1,E=5,B=50C=0.1, E=5, B=50) 280 (64.3)× 630 (49.2)× 2,000 (49.5)×

    Centralized SGD achieves 86% test accuracy after 197,500 minibatch updates. FedAvg\text{FedAvg} reaches 85% accuracy in only 2,000 communication rounds, achieving a 49.5×49.5\times reduction in communication rounds compared to SGD, while FedSGD\text{FedSGD} fails to reach the 85% accuracy threshold within the evaluated rounds.

  8. Knowl 8 — Communication Speedup on Large-Scale Word Language Modeling

    empirical result

    On a large-scale real-world next-word prediction task across >500,000>500,000 client authors (10 million posts total) using a 4.95-million parameter LSTM with a 10,000-word vocabulary, FedAvg\text{FedAvg} was evaluated with 200 clients per round, B=8B=8, and E=1E=1 against baseline FedSGD\text{FedSGD} (B=∞,E=1B=\infty, E=1).

    • Baseline FedSGD\text{FedSGD} with optimal learning rate η=18.0\eta = 18.0 required 820 communication rounds to reach 10.5% top-1 accuracy on a test set of 10510^5 posts from held-out authors.
    • FedAvg\text{FedAvg} (E=1,B=8E=1, B=8) with optimal learning rate η=9.0\eta = 9.0 reached the same 10.5% top-1 test accuracy in only 35 communication rounds, representing a 23×23\times reduction in communication rounds.
    • FedAvg\text{FedAvg} exhibited noticeably lower variance in test accuracy across evaluation rounds compared to FedSGD\text{FedSGD}.
  9. Knowl 9 — Local Client Over-Optimization and Divergence at Extreme Epoch Counts

    limitation

    When clients perform an excessively large number of local epochs EE between averaging rounds (such as E≥50E \ge 50 on the Shakespeare character LSTM), FedAvg\text{FedAvg} can plateau or diverge rather than continue improving.

    Because the global parameters wtw_t only influence local optimization via initialization, letting E→∞E \to \infty causes local client models to move towards client-specific local minima that diverge from one another, rendering naive parameter averaging ineffective. Consequently, in the later stages of convergence or for complex recurrent models, decaying the amount of local computation per round (reducing EE or increasing BB) or decaying the learning rate η\eta is necessary to prevent divergence.

  10. Knowl 10 — Asymptotic Generalization Advantage and Regularization from Model Averaging

    empirical result

    Across multiple model architectures, FedAvg\text{FedAvg} with local iterations (B<∞,E>1B < \infty, E > 1) consistently reaches higher asymptotic test accuracy than FedSGD\text{FedSGD} (B=∞,E=1B = \infty, E = 1), even when FedSGD\text{FedSGD} is trained for thousands of additional rounds:

    • For the MNIST CNN, FedSGD\text{FedSGD} asymptotically plateaued at 99.22% test accuracy after 1,200 rounds (showing no further improvement through 6,000 rounds).
    • In contrast, FedAvg\text{FedAvg} with B=10,E=20B = 10, E = 20 achieved a higher test accuracy of 99.44% in only 300 rounds.

    This behavior indicates that periodic parameter averaging of models trained on decentralized, heterogeneous data partitions provides an implicit regularization benefit similar to dropout.

Coverage note — None was omitted; all primary methodological definitions, the FedAvg algorithm, core empirical benchmarks (MNIST, Shakespeare, CIFAR-10, large-scale LSTM), ablation data, and identified optimization limitations were captured.

References

  1. 1.Martin Abadi, Andy Chu, Ian Goodfellow, Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. Deep learning with differential privacy. In 23rd ACM Conference on Computer and Communications Security (ACM CCS), 2016.
  2. 2.Monica Anderson. Technology device ownership: 2015. http://www.pewinternet.org/2015/10/29/technology-device-ownership-2015/, 2015.
  3. 3.Yossi Arjevani and Ohad Shamir. Communication complexity of distributed convex learning and optimization. In Advances in Neural Information Processing Systems 28. 2015.
  4. 4.Maria-Florina Balcan, Avrim Blum, Shai Fine, and Yishay Mansour. Distributed learning, communication complexity and privacy. arXiv preprint arXiv:1204.3514, 2012.
  5. 5.Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. A neural probabilistic language model. J. Mach. Learn. Res., 2003.
  6. 6.Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H. Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth. Practical secure aggregation for federated learning on user-held data. In NIPS Workshop on Private Multi-Party Machine Learning, 2016.
  7. 7.David L. Chaum. Untraceable electronic mail, return addresses, and digital pseudonyms. Commun. ACM, 24(2), 1981.
  8. 8.Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz. Revisiting distributed synchronous sgd. In ICLR Workshop Track, 2016.
  9. 9.Anna Choromanska, Mikael Henaff, Michaèl Mathieu, Gérard Ben Arous, and Yann LeCun. The loss surfaces of multilayer networks. In AISTATS, 2015.
  10. 10.Greg Corrado. Computer, respond to this email. http://googleresearch.blogspot.com/2015/11/computer-respond-to-this-email.html, November 2015.
  11. 11.Yann N. Dauphin, Razvan Pascanu, Çaglar G÷lcçehre, KyungHyun Cho, Surya Ganguli, and Yoshua Bengio. Identifying and attacking the saddle point problem in high-dimensional non-convex optimization. In NIPS, 2014.
  12. 12.Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc'Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng. Large scale distributed deep networks. In NIPS, 2012.
  13. 13.John Duchi, Michael I. Jordan, and Martin J. Wainwright. Privacy aware learning. Journal of the Association for Computing Machinery, 2014.
  14. 14.Cynthia Dwork and Aaron Roth. The Algorithmic Foundations of Differential Privacy. Foundations and Trends in Theoretical Computer Science. Now Publishers, 2014.
  15. 15.Olivier Fercoq, Zheng Qu, Peter Richtárik, and Martin Takác. Fast distributed coordinate descent for non-strongly convex losses. In Machine Learning for Signal Processing (MLSP), 2014 IEEE International Workshop on, 2014.
  16. 16.Ian Goodfellow, Yoshua Bengio, and Aaron Courville. Deep learning. Book in preparation for MIT Press, 2016.
  17. 17.Ian J. Goodfellow, Oriol Vinyals, and Andrew M. Saxe. Qualitatively characterizing neural network optimization problems. In ICLR, 2015.
  18. 18.Slawomir Goryczka, Li Xiong, and Vaidy Sunderam. Secure multiparty aggregation with differential privacy: A comparative study. In Proceedings of the Joint EDBT/ICDT 2013 Workshops, 2013.
  19. 19.Benjamin Graham. Fractional max-pooling. CoRR, abs/1412.6071, 2014. URL http://arxiv.org/abs/1412.6071.
  20. 20.Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9(8), November 1997.
  21. 21.Sergey Ioffe and Christian Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML, 2015.
  22. 22.Yoon Kim, Yacine Jernite, David Sontag, and Alexander M. Rush. Character-aware neural language models. CoRR, abs/1508.06615, 2015.
  23. 23.Jakub Konečný, H. Brendan McMahan, Felix X. Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon. Federated learning: Strategies for improving communication efficiency. In NIPS Workshop on Private Multi-Party Machine Learning, 2016.
  24. 24.Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, 2009.
  25. 25.Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In NIPS. 2012.
  26. 26.Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11), 1998.
  27. 27.Chenxin Ma, Virginia Smith, Martin Jaggi, Michael I Jordan, Peter Richtárik, and Martin Takáč. Adding vs. averaging in distributed primal-dual optimization. In ICML, 2015.
  28. 28.Ryan McDonald, Keith Hall, and Gideon Mann. Distributed training strategies for the structured perceptron. In NAACL HLT, 2010.
  29. 29.Natalia Neverova, Christian Wolf, Griffin Lacey, Lex Fridman, Deepak Chandra, Brandon Barbello, and Graham W. Taylor. Learning human identity from motion patterns. IEEE Access, 4:1810–1820, 2016.
  30. 30.Jacob Poushter. Smartphone ownership and internet usage continues to climb in emerging economies. Pew Research Center Report, 2016.
  31. 31.Daniel Povey, Xiaohui Zhang, and Sanjeev Khudanpur. Parallel training of deep neural networks with natural gradient and parameter averaging. In ICLR Workshop Track, 2015.
  32. 32.William Shakespeare. The Complete Works of William Shakespeare. Publically available at https://www.gutenberg.org/ebooks/100.
  33. 33.Ohad Shamir and Nathan Srebro. Distributed stochastic optimization and learning. In Communication, Control, and Computing (Allerton), 2014.
  34. 34.Ohad Shamir, Nathan Srebro, and Tong Zhang. Communication efficient distributed optimization using an approximate newton-type method. arXiv preprint arXiv:1312.7853, 2013.
  35. 35.Reza Shokri and Vitaly Shmatikov. Privacy-preserving deep learning. In Proceedings of the 22Nd ACM SIGSAC Conference on Computer and Communications Security, CCS '15, 2015.
  36. 36.Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: A simple way to prevent neural networks from overfitting. 15, 2014.
  37. 37.Latanya Sweeney. Simple demographics often identify people uniquely. 2000.
  38. 38.TensorFlow team. Tensorflow convolutional neural networks tutorial, 2016. http://www.tensorflow.org/tutorials/deep_cnn.
  39. 39.White House Report. Consumer data privacy in a networked world: A framework for protecting privacy and promoting innovation in the global digital economy. Journal of Privacy and Confidentiality, 2013.
  40. 40.Tianbao Yang. Trading computation for communication: Distributed stochastic dual coordinate ascent. In Advances in Neural Information Processing Systems, 2013.
  41. 41.Ruiliang Zhang and James Kwok. Asynchronous distributed admm for consensus optimization. In ICML. JMLR Workshop and Conference Proceedings, 2014.
  42. 42.Sixin Zhang, Anna E Choromanska, and Yann LeCun. Deep learning with elastic averaging sgd. In NIPS. 2015.
  43. 43.Yuchen Zhang and Lin Xiao. Communication-efficient distributed optimization of self-concordant empirical loss. arXiv preprint arXiv:1501.00263, 2015.
  44. 44.Yuchen Zhang, Martin J Wainwright, and John C Duchi. Communication-efficient algorithms for statistical optimization. In NIPS, 2012.
  45. 45.Yuchen Zhang, John Duchi, Michael I Jordan, and Martin J Wainwright. Information-theoretic lower bounds for distributed statistical estimation with communication constraints. In Advances in Neural Information Processing Systems, 2013.
  46. 46.Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J. Smola. Parallelized stochastic gradient descent. In NIPS. 2010.

Citation

MLA
McMahan, H. B., et al. “Communication-Efficient Learning of Deep Networks from Decentralized Data”. Proceedings of the 20 Th International Conference on Artificial Intelligence and Statistics (AISTATS) 2017. JMLR: W&CP Volume 54, 2016, http://arxiv.org/abs/1602.05629v4.
APA
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., & Arcas, B. A. y . (2016). Communication-Efficient Learning of Deep Networks from Decentralized Data. Proceedings of the 20 Th International Conference on Artificial Intelligence and Statistics (AISTATS) 2017. JMLR: W&CP Volume 54. http://arxiv.org/abs/1602.05629v4
Chicago
McMahan, H. B., E. Moore, D. Ramage, S. Hampson, and B. A. y . Arcas. 2016. “Communication-Efficient Learning of Deep Networks from Decentralized Data”. Proceedings of the 20 Th International Conference on Artificial Intelligence and Statistics (AISTATS) 2017. JMLR: W&CP Volume 54. http://arxiv.org/abs/1602.05629v4.
Harvard
McMahan, H.B. et al. (2016) “Communication-Efficient Learning of Deep Networks from Decentralized Data”, Proceedings of the 20 th International Conference on Artificial Intelligence and Statistics (AISTATS) 2017. JMLR: W&CP volume 54 [Preprint]. Available at: http://arxiv.org/abs/1602.05629v4.
Vancouver
1. McMahan HB, Moore E, Ramage D, Hampson S, Arcas BA y (2016) Communication-Efficient Learning of Deep Networks from Decentralized Data. Proceedings of the 20 th International Conference on Artificial Intelligence and Statistics (AISTATS) 2017. JMLR: W&CP volume 54

BibTeX

@article{mcmahan2016communication,
  title = {Communication-Efficient Learning of Deep Networks from Decentralized Data},
  author = {McMahan, H. Brendan and Moore, Eider and Ramage, Daniel and Hampson, Seth and Arcas, Blaise Agüera y},
  year = {2016},
  journal = {Proceedings of the 20 th International Conference on Artificial Intelligence and Statistics (AISTATS) 2017. JMLR: W&CP volume 54},
  url = {http://arxiv.org/abs/1602.05629v4},
  eprint = {1602.05629}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Published with permission