keyword
privacy-preserving learning
Privacy-preserving learning is a machine learning paradigm focused on developing and deploying statistical models while safeguarding sensitive or confidential training data from exposure, inference, and unauthorized reconstruction. This approach encompasses techniques that allow algorithms to extract actionable insights from data without directly accessing, centralizing, or revealing private information. Common strategies include decentralized frameworks such as federated learning that keep raw datasets localized on remote devices, representation learning methods that enforce statistical independence from protected attributes, and mathematical safeguards like differential privacy and cryptographic computation. By balancing predictive performance against formal privacy guarantees, privacy-preserving learning enables secure collaboration across institutional and personal boundaries without compromising individual confidentiality.
2 items

Fundamental Limits and Tradeoffs in Invariant Representation Learning
Han Zhao, Chen Dan, Bryon Aragam, Tommi S. Jaakkola, Geoffrey J. Gordon, Pradeep Ravikumar
Why you should read this
Establishes an information-theoretic framework that bounds the achievable tradeoffs between predictive accuracy and feature invariance across classification and regression tasks, providing a method to certify the suboptimality of representation learning algorithms.
A wide range of machine learning applications such as privacy-preserving learning, algorithmic fairness, and domain adaptation/generalization among others, involve learning invariant representations of the data that aim to achieve two competing goals: (a) maximize information or accuracy with respect to a target response, and (b) maximize invariance or independence with respect to a set of protected features (e.g. for fairness, privacy, etc). Despite their wide applicability, theoretical understanding of the optimal tradeoffs — with respect to accuracy, and invariance — achievable by invariant representations is still severely lacking. In this paper, we provide an information theoretic analysis of such tradeoffs under both classification and regression settings. More precisely, we provide a geometric characterization of the accuracy and invariance achievable by any representation of the data; we term this feasible region the information plane. We provide an inner bound for this feasible region for the classification case, and an exact characterization for the regression case, which allows us to either bound or exactly characterize the Pareto optimal frontier between accuracy and invariance. Although our contributions are mainly theoretical, a key practical application of our results is in certifying the potential sub-optimality of any given representation learning algorithm for either classification or regression tasks. Our results shed new light on the fundamental interplay between accuracy and invariance, and may be useful in guiding the design of future representation learning algorithms.
Added
2026-10-03

Federated Learning: Challenges, Methods, and Future Directions
Tian Li, Anit Kumar Sahu, Ameet Talwalkar, Virginia Smith
Why you should read this
Presents a systematic overview of federated learning by categorizing critical challenges in communication efficiency, systems and statistical heterogeneity, and privacy alongside the primary algorithms used to train models across decentralized networks.
Federated learning involves training statistical models over remote devices or siloed data centers, such as mobile phones or hospitals, while keeping data localized. Training in heterogeneous and potentially massive networks introduces novel challenges that require a fundamental departure from standard approaches for large-scale machine learning, distributed optimization, and privacy-preserving data analysis. In this article, we discuss the unique characteristics and challenges of federated learning, provide a broad overview of current approaches, and outline several directions of future work that are relevant to a wide range of research communities.
Added
2026-09-11
