LEAF: A Benchmark for Federated Settings
Sebastian CaldasPeter WuTian LiJakub KonecnýH. B. McMahanVirginia SmithAmeet Talwalkar
Proposes LEAF, an open-source benchmarking framework providing realistic datasets, standardized evaluation metrics, and reference implementations to enable reproducible research across federated, multi-task, and meta-learning environments.
Modern edge devices such as mobile phones, wearables, and autonomous vehicles generate massive quantities of data. Training machine learning models across these decentralized networks can improve user experiences, but researchers face major hurdles due to statistical variations across users, tight hardware and network constraints, and privacy requirements. Progress in federated and decentralized learning has been hindered because existing research relies either on unrealistic synthetic partitions, unreleased proprietary data, or public sources that are difficult to reproduce.
The article introduces and evaluates LEAF, an open-source benchmarking framework designed to provide realistic, reproducible datasets, system metrics, and baseline implementations for learning in distributed edge environments.
The authors curated six diverse datasets across text, image, and synthetic domains—including handwriting (FEMNIST), Twitter sentiment (Sent140), literature (Shakespeare), facial attributes (CelebA), and discussion forums (Reddit)—spanning thousands to millions of devices with naturally skewed data distributions. They developed evaluation metrics capturing percentile-based statistical performance as well as computational and network budgets (such as floating-point operations and network bytes transferred). The framework was demonstrated by replicating known federated learning behaviors and comparing alternative paradigms like meta-learning, local models, and centralized baselines across standardized test splits.
The analysis yielded several key findings regarding federated evaluation. First, LEAF successfully reproduced published algorithmic anomalies, confirming that the standard Federated Averaging algorithm diverges on text prediction tasks when local training epochs increase. Second, standard mean and median metrics obscure major performance disparities; in sentiment analysis, while median user accuracy degraded only slightly when low-data users were included, bottom-quartile (25th percentile) accuracy dropped dramatically. Third, communication-computation trade-offs vary substantially: Federated Averaging dramatically reduces network transfer compared to standard distributed stochastic gradient descent when targeting a 75% accuracy threshold on character classification. Finally, modular testing showed that optimal training paradigms depend heavily on data structure, with meta-learning achieving higher accuracy (80.24%) than Federated Averaging (74.72%) on handwriting data, whereas purely local models outperformed Federated Averaging on synthetic tasks (87.34% versus 71.89%) but failed on image classification (65.29% versus 89.46%).
These findings show that evaluating decentralized algorithms solely on aggregate accuracy or artificial benchmarks risks deploying models that fail for significant segments of users or overwhelm device resources. Because real-world performance depends heavily on user data skew and network limitations, researchers and technology leaders must evaluate both statistical tail distributions and systems costs before deployment.
Decision-makers and engineering teams should adopt standardized, multi-metric benchmarking suites to evaluate edge algorithms against both hardware budgets and user-level fairness metrics. Future work should expand the benchmark to cover richer media types, such as audio and video, and implement reference code for broader machine learning paradigms.
The source findings are highly credible for the evaluated image and text datasets. However, stakeholders should note that the current reference implementations are limited to a few specific learning paradigms, and performance characteristics may shift as new modalities and network constraints are introduced.
- Paper: Federated Optimization: Distributed Machine Learning for On-Device Intelligence, Jakub Konečný et al. (2016). This seminal paper defines the core federated optimization problem and foundational distributed optimization setting that LEAF benchmarks.
- Paper: Federated Multi-Task Learning, Virginia Smith et al. (2017). This work establishes the connection between federated learning and multi-task learning, motivating LEAF's multi-task benchmark suites and shared authorship foundations.
- Paper: Federated Learning: Strategies for Improving Communication Efficiency, Jakub Konečný et al. (2016). This paper establishes key communication-efficiency strategies in federated learning that inform the baseline evaluations implemented in LEAF.
- Paper: Federated Optimization in Heterogeneous Networks, Tian Li et al. (2018). This paper introduces FedProx and directly uses LEAF datasets (such as FEMNIST, Shakespeare, and Sent140) to evaluate optimization across heterogeneous networks.
- Paper: Federated Learning: Challenges, Methods, and Future Directions, Tian Li et al. (2019). This comprehensive survey synthesizes the algorithmic challenges and realistic benchmarking paradigms established by the LEAF framework.
- Paper: Advances and Open Problems in Federated Learning, P. Kairouz et al. (2019). This broad field survey synthesizes open problems and evaluation methodologies across federated learning, building directly upon modular benchmarking efforts like LEAF.
- Paper: SCAFFOLD: Stochastic Controlled Averaging for Federated Learning, Sai Praneeth Karimireddy et al. (2019). This work develops the SCAFFOLD algorithm to resolve client drift on heterogeneous non-IID data distributions standardly evaluated via LEAF.
- Paper: Adaptive Federated Optimization, Sashank Reddi et al. (2020). This paper proposes server-side adaptive optimizers for federated learning, evaluating them extensively across realistic federated benchmarks.
- Paper: Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization, Jianyu Wang et al. (2020). This work builds on heterogeneous federated evaluation principles to solve objective inconsistency during client averaging with the FedNova algorithm.
- Paper: Model-Contrastive Federated Learning, Qinbin Li et al. (2021). This paper proposes model-contrastive learning to address non-IID client drift in federated benchmarking setups.
- Paper: On the Convergence of FedAvg on Non-IID Data, Xiang Li et al. (2019). This study establishes theoretical convergence rates of federated averaging under realistic non-IID and partial-device settings evaluated in LEAF.
- Paper: Towards Federated Learning at Scale: System Design, Keith Bonawitz et al. (2019). This paper details the systems-level deployment of production federated learning, complementing LEAF's simulated benchmarking framework.
