LEAF: A Benchmark for Federated Settings

Sebastian CaldasPeter WuTian LiJakub KonecnýH. B. McMahanVirginia SmithAmeet Talwalkar

article2018arXiv1,821 citations

Proposes LEAF, an open-source benchmarking framework providing realistic datasets, standardized evaluation metrics, and reference implementations to enable reproducible research across federated, multi-task, and meta-learning environments.

Listen

Modern edge devices such as mobile phones, wearables, and autonomous vehicles generate massive quantities of data. Training machine learning models across these decentralized networks can improve user experiences, but researchers face major hurdles due to statistical variations across users, tight hardware and network constraints, and privacy requirements. Progress in federated and decentralized learning has been hindered because existing research relies either on unrealistic synthetic partitions, unreleased proprietary data, or public sources that are difficult to reproduce.

The article introduces and evaluates LEAF, an open-source benchmarking framework designed to provide realistic, reproducible datasets, system metrics, and baseline implementations for learning in distributed edge environments.

The authors curated six diverse datasets across text, image, and synthetic domains—including handwriting (FEMNIST), Twitter sentiment (Sent140), literature (Shakespeare), facial attributes (CelebA), and discussion forums (Reddit)—spanning thousands to millions of devices with naturally skewed data distributions. They developed evaluation metrics capturing percentile-based statistical performance as well as computational and network budgets (such as floating-point operations and network bytes transferred). The framework was demonstrated by replicating known federated learning behaviors and comparing alternative paradigms like meta-learning, local models, and centralized baselines across standardized test splits.

The analysis yielded several key findings regarding federated evaluation. First, LEAF successfully reproduced published algorithmic anomalies, confirming that the standard Federated Averaging algorithm diverges on text prediction tasks when local training epochs increase. Second, standard mean and median metrics obscure major performance disparities; in sentiment analysis, while median user accuracy degraded only slightly when low-data users were included, bottom-quartile (25th percentile) accuracy dropped dramatically. Third, communication-computation trade-offs vary substantially: Federated Averaging dramatically reduces network transfer compared to standard distributed stochastic gradient descent when targeting a 75% accuracy threshold on character classification. Finally, modular testing showed that optimal training paradigms depend heavily on data structure, with meta-learning achieving higher accuracy (80.24%) than Federated Averaging (74.72%) on handwriting data, whereas purely local models outperformed Federated Averaging on synthetic tasks (87.34% versus 71.89%) but failed on image classification (65.29% versus 89.46%).

These findings show that evaluating decentralized algorithms solely on aggregate accuracy or artificial benchmarks risks deploying models that fail for significant segments of users or overwhelm device resources. Because real-world performance depends heavily on user data skew and network limitations, researchers and technology leaders must evaluate both statistical tail distributions and systems costs before deployment.

Decision-makers and engineering teams should adopt standardized, multi-metric benchmarking suites to evaluate edge algorithms against both hardware budgets and user-level fairness metrics. Future work should expand the benchmark to cover richer media types, such as audio and video, and implement reference code for broader machine learning paradigms.

The source findings are highly credible for the evaluated image and text datasets. However, stakeholders should note that the current reference implementations are limited to a few specific learning paradigms, and performance characteristics may shift as new modalities and network constraints are introduced.

Cover for LEAF: A Benchmark for Federated Settings

Abstract

Modern federated networks, such as those comprised of wearable devices, mobile phones, or autonomous vehicles, generate massive amounts of data each day. This wealth of data can help to learn models that can improve the user experience on each device. However, the scale and heterogeneity of federated data presents new challenges in research areas such as federated learning, meta-learning, and multi-task learning. As the machine learning community begins to tackle these challenges, we are at a critical time to ensure that developments made in these areas are grounded with realistic benchmarks. To this end, we propose LEAF, a modular benchmarking framework for learning in federated settings. LEAF includes a suite of open-source federated datasets, a rigorous evaluation framework, and a set of reference implementations, all geared towards capturing the obstacles and intricacies of practical federated environments.

Table of Contents

  • 1 Introduction
  • 2 LEAF
  • 3 LEAF in action
  • 4 Conclusions
  • References
  • A Synthetic Dataset
  • B Experiment Details

Knowls

  1. Knowl 1 — LEAF Federated Benchmarking Framework Architecture

    model/method

    LEAF is a modular benchmarking framework designed for evaluating learning algorithms in massive, heterogeneous federated networks. The framework consists of three decoupled modules:

    1. Datasets Module: Preprocesses raw data and converts it into a standardized JSON format partitioned by natural device/user keys (capturing non-identically distributed sample sizes and statistical distributions across devices). It provides small and full dataset versions for rapid prototyping and benchmark testing.
    2. Reference Implementations Module: Contains implementations of baseline federated algorithms (including Minibatch Stochastic Gradient Descent, Federated Averaging (extFedAvg ext{FedAvg}), MOCHA, and meta-learning algorithms like Reptile). Each implementation outputs execution logs recording statistical accuracy and systems-level resource consumption.
    3. Metrics Module: Parses and aggregates execution logs to evaluate performance across device distributions and systems dimensions.
  2. Knowl 2 — LEAF Federated Dataset Suite and Summary Statistics

    data/table

    The LEAF suite provides five primary empirical datasets exhibiting natural keyed partitions and skew in the number of samples per device:

    Dataset Number of devices Total samples Samples/device (mean) Samples/device (stdev)
    FEMNIST 3,550 805,263 226.83 88.94
    Sent140 660,120 1,600,498 2.42 4.71
    Shakespeare 1,129 4,226,158 3,743.28 6,212.26
    CelebA 9,343 200,288 21.44 7.63
    Reddit 1,660,820 56,587,343 34.07 62.95
    • FEMNIST: Extended MNIST digit and character classification partitioned by the individual writer.
    • Sentiment140 (Sent140): Twitter sentiment classification dataset partitioned by individual Twitter user accounts.
    • Shakespeare: Next-character language modeling dataset built from The Complete Works of William Shakespeare, partitioned by individual speaking roles in each play.
    • CelebA: Face attribute classification dataset partitioned by the celebrity identity in the image.
    • Reddit: Next-token language prediction dataset built from December 2017 Reddit comments, partitioned by individual user accounts.
  3. Knowl 3 — Synthetic Clustered Federated Dataset Generation Procedure

    algorithm

    To generate controlled non-IID federated tasks with ground-truth models clustered around multiple centers, LEAF uses the following procedure:

    Input: Number of devices/tasks T≥1T \ge 1, cluster mixing probabilities p=(p1,…,pk)p = (p_1, \dots, p_k) with pj>0p_j > 0 and ∑j=1kpj=1\sum_{j=1}^k p_j = 1, feature dimension dd, latent dimension ss.
    Output: Federated task datasets {(Xt,Yt)}t=1T\{ (X_t, Y_t) \}_{t=1}^T.
    for j=1j = 1 to kk do
        Sample Bj∼N(0,Is)B_j \sim \mathcal{N}(0, I_s)
        Sample cluster mean μj∼N(Bj,Is)\mu_j \sim \mathcal{N}(B_j, I_s)
    Sample projection matrix Q∼N(0,I(d+1)×s)Q \sim \mathcal{N}(0, I_{(d+1) \times s})
    Construct diagonal covariance matrix Σ∈Rd×d\Sigma \in \mathbb{R}^{d \times d} with Σi,i=i−1.2\Sigma_{i,i} = i^{-1.2}
    for t=1t = 1 to TT do
        Sample cluster assignment j∈{1,…,k}j \in \{1, \dots, k\} with probability pjp_j, and set μt=μj\mu_t = \mu_j
        Sample latent task parameter ut∼N(μt,Is)u_t \sim \mathcal{N}(\mu_t, I_s)
        Compute task model weight vector wt=Qut∈Rd+1w_t = Q u_t \in \mathbb{R}^{d+1}
        Sample sample size generator mt∼Lognormal(μ=3,σ=2)m_t \sim \text{Lognormal}(\mu=3, \sigma=2)
        Set number of task samples nt=min⁡(⌊mt⌋+5,1000)n_t = \min(\lfloor m_t \rfloor + 5, 1000)
        Sample task feature bias Ct∼N(0,Id)C_t \sim \mathcal{N}(0, I_d)
        Sample task feature mean vt∼N(Ct,Id)v_t \sim \mathcal{N}(C_t, I_d)
        for i=1i = 1 to ntn_t do
            Sample feature vector xti∼N(vt,Σ)∈Rdx_t^i \sim \mathcal{N}(v_t, \Sigma) \in \mathbb{R}^d
            Augment feature vector with intercept: x~ti=[xti;1]∈Rd+1\tilde{x}_t^i = [x_t^i; 1] \in \mathbb{R}^{d+1}
            Sample observation noise ϵti∼N(0,0.1⋅I)\epsilon_t^i \sim \mathcal{N}(0, 0.1 \cdot I)
            Compute label yti=arg⁡max⁡(sigmoid(wt⊤x~ti+ϵti))y_t^i = \arg\max(\text{sigmoid}(w_t^\top \tilde{x}_t^i + \epsilon_t^i))
        Set task dataset (Xt,Yt)=({xti}i=1nt,{yti}i=1nt)(X_t, Y_t) = (\{x_t^i\}_{i=1}^{n_t}, \{y_t^i\}_{i=1}^{n_t})
    return {(Xt,Yt)}t=1T\{ (X_t, Y_t) \}_{t=1}^T
  4. Knowl 4 — Statistical and Systems Evaluation Metrics for Federated Benchmarks

    model/method

    LEAF establishes a multi-metric evaluation protocol separating statistical quality and systems-level costs across heterogeneous devices:

    1. Statistical Metrics:
      • Percentile-based accuracy: Reports performance at the 10th, 25th, 50th (median), 75th, and 90th percentiles across client devices to capture accuracy disparities across data-rich and data-poor clients.
      • Hierarchical stratification: Measures performance grouped by intrinsic domain structures (e.g., performance grouped by play in Shakespeare or by subreddit in Reddit).
      • Accuracy weighting: Distinguishes between uniform device weighting (equal weight per client device) and sample weighting (equal weight per data point, which biases toward heavy users).
    2. Systems Metrics:
      • Computation cost: Total floating-point operations (FLOPs) accumulated across all participating client devices.
      • Communication cost: Total network footprint measured in cumulative bytes uploaded and downloaded across all devices.
  5. Knowl 5 — Performance Comparison Across Federated and Alternative Learning Paradigms

    data/table

    LEAF evaluates baseline Federated Averaging (FedAvg) against alternative learning paradigms (purely local training, pooled global IID training, and meta-learning via Reptile) across different dataset structures:

    Dataset FedAvg (baseline) Additional Pipeline Accuracy
    CelebA 89.46% Local Models 65.29%
    Synthetic 71.89% Local Models 87.34%
    Reddit 13.35% Global IID model 12.60%
    FEMNIST 74.72% Reptile 80.24%
    • On CelebA (face attributes with high task alignment), FedAvg outperforms individual local models by 24.17%24.17\% absolute test accuracy due to collaborative feature learning.
    • On Synthetic (where true task models are distinct and task-dependent), training independent local models outperforms a single shared FedAvg model by 15.45%15.45\%.
    • On Reddit, retaining natural client partitioning via FedAvg outperforms pooling all data into an IID stream (13.35%13.35\% vs 12.60%12.60\%).
    • On FEMNIST, the meta-learning method Reptile with client fine-tuning achieves 80.24%80.24\% test accuracy, outperforming standard FedAvg (74.72%74.72\%) on unseen test devices.
  6. Knowl 6 — FedAvg Loss Divergence Under High Local Epochs on the Shakespeare Dataset

    empirical result

    When running Federated Averaging (FedAvg) on next-character prediction for the Shakespeare dataset using a 2-layer LSTM (256 units per layer, embedding dimension 8, sequence length 80, learning rate 0.8, and 10 devices sampled per round):

    • Setting local epochs to E=1E = 1 leads to smooth, stable convergence in training loss and monotonic increases in Top-1 test accuracy over 40 communication rounds.
    • Increasing local epochs to E=20E = 20 causes the training loss to diverge sharply, exhibiting large oscillations between rounds and degrading final model quality. This replicates the known instability of FedAvg when local client updates drift excessively between aggregation steps on non-IID text data.
  7. Knowl 7 — Tail Performance Degradation Under Device Sample Scarcity in Sentiment140

    empirical result

    On the Sentiment140 Twitter sentiment classification task (using a bag-of-words logistic regression model with learning rate 3×10−43 \times 10^{-4}), filtering users by minimum sample threshold k∈{3,10,30,100}k \in \{3, 10, 30, 100\} reveals significant disparities between central tendency and tail percentiles:

    • Median device accuracy remains relatively stable as kk decreases from 100 down to 3.
    • In contrast, the 10th and 25th percentile accuracies degrade sharply at lower thresholds (e.g., k=3k = 3), showing that overall average or median performance metrics conceal severe failure modes on data-deficient tail devices in non-IID federated distributions.
  8. Knowl 8 — Systems Computation-Communication Trade-Off of FedAvg versus Minibatch SGD on FEMNIST

    empirical result

    On a 5%5\% subsample of the FEMNIST dataset (using a convolutional network with two conv-pooling layers and a 2048-unit dense layer), reaching a fixed test sample accuracy threshold of 0.750.75 imposes contrasting computation and communication budgets between algorithms:

    • Minibatch SGD (learning rate 6×10−26 \times 10^{-2}): Incurs low computational cost per device (fewer local FLOPs) but requires roughly 101110^{11} bytes uploaded to the network due to frequent parameter synchronization.
    • FedAvg (learning rate 4×10−34 \times 10^{-3}, local epochs E∈{1,100}E \in \{1, 100\}, local batch size 5): Shifts the burden toward local computation (higher total FLOPs per round) while reducing total uploaded bytes by approximately three orders of magnitude ( ≈108\,\approx 10^8 bytes written).

Coverage note — None was omitted; all core benchmarking components, dataset summaries, synthetic generation equations, evaluation metrics, and empirical findings have been extracted.

References

  1. 1.Arvind Agarwal, Samuel Gerber, and Hal Daume. Learning multiple tasks using manifold regularization. In Advances in neural information processing systems, 2010.
  2. 2.Andreas Argyriou, Theodoros Evgeniou, and Massimiliano Pontil. Convex multi-task feature learning. Machine Learning, 73(3):243–272, 2008.
  3. 3.Eugene Bagdasaryan, Andreas Veit, Yiqing Hua, Deborah Estrin, and Vitaly Shmatikov. How to backdoor federated learning. arXiv preprint arXiv:1807.00459, 2018.
  4. 4.Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth. Practical secure aggregation for privacy-preserving machine learning. In ACM SIGSAC Conference on Computer and Communications Security, 2017.
  5. 5.Fei Chen, Zhenhua Dong, Zhenguo Li, and Xiuqiang He. Federated meta-learning for recommendation. arXiv preprint arXiv:1802.07876, 2018.
  6. 6.Gregory Cohen, Saeed Afshar, Jonathan Tapson, and André van Schaik. EMNIST: an extension of MNIST to handwritten letters. arXiv preprint arXiv:1702.05373, 2017.
  7. 7.Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning, 2017.
  8. 8.Robin C Geyer, Tassilo Klein, and Moin Nabi. Differentially private federated learning: A client level perspective. arXiv preprint arXiv:1712.07557, 2017.
  9. 9.Alec Go, Richa Bhayani, and Lei Huang. Twitter sentiment classification using distant supervision. Project Report, Stanford, 2009.
  10. 10.Yihan Jiang, Jakub Konecnˇ y, Keith Rush, and Sreeram Kannan. Improving federated learning ` personalization via model agnostic meta learning. arXiv preprint arXiv:1909.12488, 2019.
  11. 11.Michael Kamp, Linara Adilova, Joachim Sicking, Fabian Hüger, Peter Schlicht, Tim Wirtz, and Stefan Wrobel. Efficient decentralized deep learning by dynamic model averaging. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 2018.
  12. 12.Mikhail Khodak, Maria Florina-Balcan, and Ameet Talwalkar. Adaptive gradient-based meta-learning methods. Advances in Neural Information Processing Systems, 2019.
  13. 13.Jakub Konecnˇ y, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, ` and Dave Bacon. Federated learning: Strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492, 2016.
  14. 14.Abhishek Kumar and Hal Daume III. Learning task grouping and overlap in multi-task learning. arXiv preprint arXiv:1206.6417, 2012.
  15. 15.Brenden Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua Tenenbaum. One shot learning of simple visual concepts. In Annual Meeting of the Cognitive Science Society, 2011.
  16. 16.Yann LeCun. The MNIST database of handwritten digits. http: // yann. lecun. com/ exdb/ mnist/ , 1998.
  17. 17.Giwoong Lee, Eunho Yang, and Sung Hwang. Asymmetric multi-task learning based on task relatedness and loss. In International Conference on Machine Learning, 2016.
  18. 18.David Leroy, Alice Coucke, Thibaut Lavril, Thibault Gisselbrecht, and Joseph Dureau. Federated learning for keyword spotting. In IEEE International Conference on Acoustics, Speech and Signal Processing, 2019.
  19. 19.Tian Li, Anit Kumar Sahu, Ameet Talwalkar, and Virginia Smith. Federated learning: Challenges, methods, and future directions. arXiv preprint arXiv:1908.07873, 2019.
  20. 20.Tian Li, Maziar Sanjabi, and Virginia Smith. Fair resource allocation in federated learning. arXiv preprint arXiv:1905.10497, 2019.
  21. 21.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In International Conference on Computer Vision, 2015.
  22. 22.H Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas. Communication-efficient learning of deep networks from decentralized data. In Artificial Intelligence and Statistics, 2017.
  23. 23.H Brendan McMahan and Daniel Ramage. Federated Learning: Collaborative Machine Learning without Centralized Training Data. googblogs.com, 2017.
  24. 24.H Brendan McMahan, Daniel Ramage, Kunal Talwar, and Li Zhang. Learning differentially private recurrent language models. In International Conference on Learning Representations, 2018.
  25. 25.Luca Melis, Congzheng Song, Emiliano De Cristofaro, and Vitaly Shmatikov. Inference attacks against collaborative learning. arXiv preprint arXiv:1805.04049, 2018.
  26. 26.Keerthiram Murugesan and Jaime Carbonell. Multi-task multiple kernel relationship learning. In SIAM International Conference on Data Mining, 2017.
  27. 27.Alex Nichol, Joshua Achiam, and John Schulman. On first-order meta-learning algorithms. arXiv preprint arXiv:1803.02999, 2018.
  28. 28.Vasyl Pihur, Aleksandra Korolova, Frederick Liu, Subhash Sankuratripati, Moti Yung, Dachuan Huang, and Ruogu Zeng. Differentially-private "draw and discard" machine learning. arXiv preprint arXiv:1807.04369, 2018.
  29. 29.Sachin Ravi and Hugo Larochelle. Optimization as a model for few-shot learning. https: // openreview. net/ pdf? id= rJY0-Kcll , 2016.
  30. 30.Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, and Ameet S Talwalkar. Federated multi-task learning. In Advances in Neural Information Processing Systems, 2017.
  31. 31.Jake Snell, Kevin Swersky, and Richard Zemel. Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems, 2017.
  32. 32.Gregor Ulm, Emil Gustavsson, and Mats Jirstrand. Functional federated learning in erlang (ffl-erl). In International Workshop on Functional and Constraint Logic Programming, 2018.
  33. 33.Oriol Vinyals, Charles Blundell, Tim Lillicrap, Daan Wierstra, et al. Matching networks for one shot learning. In Advances in Neural Information Processing Systems, 2016.
  34. 34.Shiqiang Wang, Tiffany Tuor, Theodoros Salonidis, Kin K Leung, Christian Makaya, Ting He, and Kevin Chan. Adaptive federated learning in resource constrained edge computing systems. IEEE Journal on Selected Areas in Communications, 2019.
  35. 35.William Shakespeare. The Complete Works of William Shakespeare. Publicly available at //www.gutenberg.org/ebooks/100.
  36. 36.Ya Xue, Xuejun Liao, Lawrence Carin, and Balaji Krishnapuram. Multi-task learning for classification with dirichlet process priors. Journal of Machine Learning Research, 8(Jan):35–63, 2007.
  37. 37.Qiang Yang, Yang Liu, Tianjian Chen, and Yongxin Tong. Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology, 2019.
  38. 38.Yi Zhang and Jeff G Schneider. Learning multiple tasks with a sparse matrix-normal penalty. In Advances in Neural Information Processing Systems, 2010.

Citation

MLA
Caldas, S., et al. “LEAF: A Benchmark for Federated Settings”. arXiv, 2018, http://arxiv.org/abs/1812.01097v3.
APA
Caldas, S., Duddu, S. M. K., Wu, P., Li, T., Konečný, J., McMahan, H. B., Smith, V., & Talwalkar, A. (2018). LEAF: A Benchmark for Federated Settings. arXiv. http://arxiv.org/abs/1812.01097v3
Chicago
Caldas, S., S. M. K. Duddu, P. Wu, et al. 2018. “LEAF: A Benchmark for Federated Settings”. arXiv. http://arxiv.org/abs/1812.01097v3.
Harvard
Caldas, S. et al. (2018) “LEAF: A Benchmark for Federated Settings”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1812.01097v3.
Vancouver
1. Caldas S, Duddu SMK, Wu P, Li T, Konečný J, McMahan HB, Smith V, Talwalkar A (2018) LEAF: A Benchmark for Federated Settings. arXiv

BibTeX

@article{caldas2018leaf,
  title = {LEAF: A Benchmark for Federated Settings},
  author = {Caldas, Sebastian and Duddu, Sai Meher Karthik and Wu, Peter and Li, Tian and Konečný, Jakub and McMahan, H. Brendan and Smith, Virginia and Talwalkar, Ameet},
  year = {2018},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1812.01097v3},
  eprint = {1812.01097}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF

License: Published with permission