Built independently by an author, for readers. Read the story and support ChapterPal

keyword

gradient compression

Gradient compression is a collection of techniques in distributed and federated machine learning designed to reduce the volume of gradient data transmitted between computing nodes or client devices and a central server during model training. By employing strategies such as quantization to decrease the numerical bit precision of gradient values and sparsification to transmit only the most significant coordinates, these methods significantly diminish communication overhead and network bandwidth constraints. In addition to accelerating distributed optimization across communication-limited environments, gradient compression can also degrade shared information to help mitigate privacy leakage risks, often incorporating error-compensation mechanisms to preserve the convergence rate and predictive accuracy of the trained model.

10 items

Auditing Privacy Defenses in Federated Learning via Generative Gradient Leakage

Auditing Privacy Defenses in Federated Learning via Generative Gradient Leakage

Zhuohang Li, Jiaxin Zhang, Luyang Liu, Jian Liu

OrganizationsGoogleOak Ridge National LaboratoryUniversity of Tennessee

Why you should read this

Presents Generative Gradient Leakage, a framework combining deep generative priors and gradient-free optimization to reconstruct high-fidelity private images from defended gradients in federated learning, exposing vulnerabilities in existing gradient degradation defenses.

Federated Learning (FL) framework brings privacy benefits to distributed learning systems by allowing multiple clients to participate in a learning task under the coordination of a central server without exchanging their private data. However, recent studies have revealed that private information can still be leaked through shared gradient information. To further protect user's privacy, several defense mechanisms have been proposed to prevent privacy leakage via gradient information degradation methods, such as using additive noise or gradient compression before sharing it with the server. In this work, we validate that the private training data can still be leaked under certain defense settings with a new type of leakage, i.e., Generative Gradient Leakage (GGL). Unlike existing methods that only rely on gradient information to reconstruct data, our method leverages the latent space of generative adversarial networks (GAN) learned from public image datasets as a prior to compensate for the informational loss during gradient degradation. To address the nonlinearity caused by the gradient operator and the GAN model, we explore various gradient-free optimization methods (e.g., evolution strategies and Bayesian optimization) and empirically show their superiority in reconstructing high-quality images from gradients compared to gradient-based optimizers. We hope the proposed method can serve as a tool for empirically measuring the amount of privacy leakage to facilitate the design of more robust defense mechanisms.

Added

2026-10-06

Communication-Efficient Adaptive Federated Learning

Communication-Efficient Adaptive Federated Learning

Yujia Wang, Lu Lin, Jinghui Chen

OrganizationsPennsylvania State UniversityUniversity of Virginia

Why you should read this

Proposes FedCAMS, a communication-compressed adaptive federated learning algorithm that combines error-feedback compression with adaptive optimization to significantly reduce bandwidth overhead while maintaining the theoretical convergence rate of uncompressed methods.

Federated learning is a machine learning training paradigm that enables clients to jointly train models without sharing their own localized data. However, the implementation of federated learning in practice still faces numerous challenges, such as the large communication overhead due to the repetitive server-client synchronization and the lack of adaptivity by SGD-based model updates. Despite that various methods have been proposed for reducing the communication cost by gradient compression or quantization, and the federated versions of adaptive optimizers such as FedAdam are proposed to add more adaptivity, the current federated learning framework still cannot solve the aforementioned challenges all at once. In this paper, we propose a novel communication-efficient adaptive federated learning method (FedCAMS) with theoretical convergence guarantees. We show that in the nonconvex stochastic optimization setting, our proposed FedCAMS achieves the same convergence rate of O(1/(√(T K m))) as its non-compressed counterparts. Extensive experiments on various benchmarks verify our theoretical analysis.

Added

2026-10-01

On Biased Compression for Distributed Learning

On Biased Compression for Distributed Learning

Aleksandr Beznosikov, Samuel Horváth, Peter Richtárik, Mher Safaryan

OrganizationsInstitute of Science and Technology AustriaKing Abdullah University of Science and TechnologyMohamed bin Zayed University of Artificial IntelligenceMoscow Institute of Physics and TechnologySkolkovo Institute of Science and Technology

Why you should read this

Establishes the first theoretical framework proving linear convergence for biased gradient compression operators in single-node and distributed optimization, showing both theoretically and empirically why biased methods like Top-kk sparsification consistently outperform unbiased alternatives when combined with error feedback.

In the last few years, various communication compression techniques have emerged as an indispensable tool helping to alleviate the communication bottleneck in distributed learning. However, despite the fact biased compressors often show superior performance in practice when compared to the much more studied and understood unbiased compressors, very little is known about them. In this work we study three classes of biased compression operators, two of which are new, and their performance when applied to (stochastic) gradient descent and distributed (stochastic) gradient descent. We show for the first time that biased compressors can lead to linear convergence rates both in the single node and distributed settings. We prove that distributed compressed SGD method, employed with error feedback mechanism, enjoys the ergodic rate O (δL exp [− μK / δL] + (C+δD) / (Kμ)), where δ ≥ 1 is a compression parameter which grows when more compression is applied, L and μ are the smoothness and strong convexity constants, C captures stochastic gradient noise (C = 0 if full gradients are

Added

2026-09-28

Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities

Compute-Efficient Deep Learning: Algorithmic Trends and Opportunities

Brian R. Bartoldson, Bhavya Kailkhura, Davis W. Blalock

OrganizationsLawrence Livermore National LaboratoryMosaicML

Why you should read this

Presents a unified taxonomy and rigorous evaluation framework for algorithmically efficient deep learning training techniques, equipping practitioners to systematically identify bottlenecks and combine methods to cut computational costs.

Although deep learning has made great progress in recent years, the exploding economic and environmental costs of training neural networks are becoming unsustainable. To address this problem, there has been a great deal of research on algorithmically-efficient deep learning, which seeks to reduce training costs not at the hardware or implementation level, but through changes in the semantics of the training program. In this paper, we present a structured and comprehensive overview of the research in this field. First, we formalize the algorithmic speedup problem, then we use fundamental building blocks of algorithmically efficient training to develop a taxonomy. Our taxonomy highlights commonalities of seemingly disparate methods and reveals current research gaps. Next, we present evaluation best practices to enable comprehensive, fair, and reliable comparisons of speedup techniques. To further aid research and applications, we discuss common bottlenecks in the training pipeline (illustrated via experiments) and offer taxonomic mitigation strategies for them. Finally, we highlight some unsolved research challenges and present promising future directions.

Added

2026-09-26

FedNL: Making Newton-Type Methods Applicable to Federated Learning

FedNL: Making Newton-Type Methods Applicable to Federated Learning

Mher Safaryan, Rustem Islamov, Xun Qian, Peter Richtárik

OrganizationsInstitut polytechnique de ParisKing Abdullah University of Science and TechnologyMoscow Institute of Physics and Technology

Why you should read this

Proposes FedNL, a family of communication-efficient second-order optimization methods for federated learning that uses contractive Hessian compression and privacy-preserving local updates to achieve condition-number-independent local convergence.

Inspired by recent work of Islamov et al (2021), we propose a family of Federated Newton Learn (FedNL) methods, which we believe is a marked step in the direction of making second-order methods applicable to FL. In contrast to the aforementioned work, FedNL employs a different Hessian learning technique which i) enhances privacy as it does not rely on the training data to be revealed to the coordinating server, ii) makes it applicable beyond generalized linear models, and iii) provably works with general contractive compression operators for compressing the local Hessians, such as Top-K or Rank-R, which are vastly superior in practice. Notably, we do not need to rely on error feedback for our methods to work with contractive compressors. Moreover, we develop FedNL-PP, FedNL-CR and FedNL-LS, which are variants of FedNL that support partial participation, and globalization via cubic regularization and line search, respectively, and FedNL-BC, which is a variant that can further benefit from bidirectional compression of gradients and models, i.e., smart uplink gradient and smart downlink model compression. We prove local convergence rates that are independent of the condition number, the number of training data points, and compression variance. Our communication efficient Hessian learning technique provably learns the Hessian at the optimum. Finally, we perform a variety of numerical experiments that show that our FedNL methods have state-of-the-art communication complexity when compared to key baselines.

Added

2026-09-26

Robust and Communication-Efficient Federated Learning From Non-i.i.d. Data

Robust and Communication-Efficient Federated Learning From Non-i.i.d. Data

Felix Sattler, Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek

OrganizationsFraunhofer Heinrich Hertz InstituteKorea UniversityMax Planck Institute for InformaticsTechnische Universität Berlin

Why you should read this

Introduces Sparse Ternary Compression, a bidirectional compression framework that significantly reduces federated learning bandwidth requirements while outperforming Federated Averaging on heterogeneous, non-IID client data.

Federated Learning allows multiple parties to jointly train a deep learning model on their combined data, without any of the participants having to reveal their local data to a centralized server. This form of privacy-preserving collaborative learning however comes at the cost of a significant communication overhead during training. To address this problem, several compression methods have been proposed in the distributed training literature that can reduce the amount of required communication by up to three orders of magnitude. These existing methods however are only of limited utility in the Federated Learning setting, as they either only compress the upstream communication from the clients to the server (leaving the downstream communication uncompressed) or only perform well under idealized conditions such as iid distribution of the client data, which typically can not be found in Federated Learning. In this work, we propose Sparse Ternary Compression (STC), a new compression framework that is specifically designed to meet the requirements of the Federated Learning environment. Our experiments on four different learning tasks demonstrate that STC distinctively outperforms Federated Averaging in common Federated Learning scenarios where clients either a) hold non-iid data, b) use small batch sizes during training, or where c) the number of clients is large and the participation rate in every communication round is low. We furthermore show that even if the clients hold iid data and use medium sized batches for training, STC still behaves pareto-superior to Federated Averaging in the sense that it achieves fixed target accuracies on our benchmarks within both fewer training iterations and a smaller communication budget.

Added

2026-09-24

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, Shen Li

OrganizationsMeta

Why you should read this

Translates the zero-redundancy theories into a robust PyTorch implementation, relying heavily on flat-parameter abstractions and collective asynchronous prefetching.

It is widely acknowledged that large models have the potential to deliver superior performance across a broad range of domains. Despite the remarkable progress made in the field of machine learning systems research, which has enabled the development and exploration of large models, such abilities remain confined to a small group of advanced users and industry leaders, resulting in an implicit technical barrier for the wider community to access and leverage these technologies. In this paper, we introduce PyTorch Fully Sharded Data Parallel (FSDP) as an industry-grade solution for large model training. FSDP has been closely co-designed with several key PyTorch core components including Tensor implementation, dispatcher system, and CUDA memory caching allocator, to provide non-intrusive user experiences and high training efficiency. Additionally, FSDP natively incorporates a range of techniques and settings to optimize resource utilization across a variety of hardware configurations. The experimental results demonstrate that FSDP is capable of achieving comparable performance to Distributed Data Parallel while providing support for significantly larger models with near-linear scalability in terms of TFLOPS.

Added

2026-03-13

Zero-Shot Text-to-Image Generation

Zero-Shot Text-to-Image Generation

Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, Ilya Sutskever

OrganizationsOpenAI

Why you should read this

Demonstrates that a simple, large-scale autoregressive transformer can achieve competitive zero-shot text-to-image generation, challenging the reliance on complex architectures and auxiliary losses.

Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset. These assumptions might involve complex architectures, auxiliary losses, or side information such as object part labels or segmentation masks supplied during training. We describe a simple approach for this task based on a transformer that autoregressively models the text and image tokens as a single stream of data. With sufficient data and scale, our approach is competitive with previous domain-specific models when evaluated in a zero-shot fashion.

Added

2026-01-28

Creative Commons License