Built independently by an author, for readers. Read the story and support ChapterPal

keyword

maximal mutual information

Maximal mutual information refers to the maximum amount of shared information, quantified through mutual information, that a learned representation or feature transformation of data can preserve about a target variable, often evaluated subject to constraints such as statistical independence from protected or nuisance attributes. In information theory and machine learning, mutual information measures how much knowing one random variable reduces uncertainty about another, so maximizing it ensures that representations retain optimal predictive utility for a downstream task. Under invariance or fairness constraints, maximal mutual information defines the fundamental theoretical limit on task-relevant information that can be preserved while completely or partially eliminating information related to sensitive attributes, characterizing an extremal point on the Pareto-optimal trade-off between predictive accuracy and attribute invariance.

1 item

Fundamental Limits and Tradeoffs in Invariant Representation Learning

Fundamental Limits and Tradeoffs in Invariant Representation Learning

Han Zhao, Chen Dan, Bryon Aragam, Tommi S. Jaakkola, Geoffrey J. Gordon, Pradeep Ravikumar

OrganizationsCarnegie Mellon UniversityMassachusetts Institute of TechnologyUniversity of ChicagoUniversity of Illinois Urbana-Champaign

Why you should read this

Establishes an information-theoretic framework that bounds the achievable tradeoffs between predictive accuracy and feature invariance across classification and regression tasks, providing a method to certify the suboptimality of representation learning algorithms.

A wide range of machine learning applications such as privacy-preserving learning, algorithmic fairness, and domain adaptation/generalization among others, involve learning invariant representations of the data that aim to achieve two competing goals: (a) maximize information or accuracy with respect to a target response, and (b) maximize invariance or independence with respect to a set of protected features (e.g. for fairness, privacy, etc). Despite their wide applicability, theoretical understanding of the optimal tradeoffs — with respect to accuracy, and invariance — achievable by invariant representations is still severely lacking. In this paper, we provide an information theoretic analysis of such tradeoffs under both classification and regression settings. More precisely, we provide a geometric characterization of the accuracy and invariance achievable by any representation of the data; we term this feasible region the information plane. We provide an inner bound for this feasible region for the classification case, and an exact characterization for the regression case, which allows us to either bound or exactly characterize the Pareto optimal frontier between accuracy and invariance. Although our contributions are mainly theoretical, a key practical application of our results is in certifying the potential sub-optimality of any given representation learning algorithm for either classification or regression tasks. Our results shed new light on the fundamental interplay between accuracy and invariance, and may be useful in guiding the design of future representation learning algorithms.

Added

2026-10-03