Built independently by an author, for readers. Read the story and support ChapterPal

keyword

conditional mutual information

Conditional mutual information is an information-theoretic measure that quantifies the expected amount of shared information between two random variables given the knowledge of a third conditioning variable. In practical terms, it measures how much uncertainty about one variable is reduced by observing a second variable, over and above the information already provided by the third variable. Mathematically, it can be expressed as the difference between conditional entropies or as the expected relative entropy between the joint conditional probability distribution of the two variables and the product of their individual conditional distributions given the third. The quantity is always non-negative and equals zero if and only if the two variables are conditionally independent given the third, making it a foundational tool for evaluating conditional dependence, feature relevance, and information flow across statistical and machine learning applications.

3 items

A Rigorous Information-Theoretic Definition of Redundancy and Relevancy in Feature Selection Based on (Partial) Information Decomposition

A Rigorous Information-Theoretic Definition of Redundancy and Relevancy in Feature Selection Based on (Partial) Information Decomposition

Patricia Wollstadt, Sebastian Schmitt, Michael Wibral

OrganizationsCampus Institute for Dynamics of Biological NetworksHonda Research InstituteUniversity of Göttingen

Why you should read this

Establishes a rigorous definition of feature relevancy and redundancy using partial information decomposition and introduces an iterative conditional mutual information algorithm that isolates unique, redundant, and synergistic feature contributions for machine learning.

Selecting a minimal feature set that is maximally informative about a target variable is a central task in machine learning and statistics. Information theory provides a powerful framework for formulating feature selection algorithms—yet, a rigorous, information-theoretic definition of feature relevancy, which accounts for feature interactions such as redundant and synergistic contributions, is still missing. We argue that this lack is inherent to classical information theory which does not provide measures to decompose the information a set of variables provides about a target into unique, redundant, and synergistic contributions. Such a decomposition has been introduced only recently by the partial information decomposition (PID) framework. Using PID, we clarify why feature selection is a conceptually difficult problem when approached using information theory and provide a novel definition of feature relevancy and redundancy in PID terms. From this definition, we show that the conditional mutual information (CMI) maximizes relevancy while minimizing redundancy and propose an iterative, CMI-based algorithm for practical feature selection. We demonstrate the power of our CMI-based algorithm in comparison to the unconditional mutual information on benchmark examples and provide corresponding PID estimates to highlight how PID allows to quantify information contribution of features and their interactions in feature-selection problems.

Added

2026-10-05

Deep Unlearning via Randomized Conditionally Independent Hessians

Deep Unlearning via Randomized Conditionally Independent Hessians

Ronak Mehta, Sourav Pal, Vikas Singh, Sathya N. Ravi

OrganizationsUniversity of Illinois ChicagoUniversity of Wisconsin Madison

Why you should read this

Proposes a scalable approximate machine unlearning method that avoids full Hessian inversion by using a conditional independence coefficient to identify and update only the most relevant parameter subsets.

Recent legislation has led to interest in machine unlearning, i.e., removing specific training samples from a predictive model as if they never existed in the training dataset. Unlearning may also be required due to corrupted/adversarial data or simply a user’s updated privacy requirement. For models which require no training (k-NN), simply deleting the closest original sample can be effective. But this idea is inapplicable to models which learn richer representations. Recent ideas leveraging optimization-based updates scale poorly with the model dimension d, due to inverting the Hessian of the loss function. We use a variant of a new conditional independence coefficient, L-CODEC, to identify a subset of the model parameters with the most semantic overlap on an individual sample level. Our approach completely avoids the need to invert a (possibly) huge matrix. By utilizing a Markov blanket selection, we premise that L-CODEC is also suitable for deep unlearning, as well as other applications in vision. Compared to alternatives, L-CODEC makes approximate unlearning possible in settings that would otherwise be infeasible, including vision models used for face recognition, person re-identification and NLP models that may require unlearning samples identified for exclusion. Code is available at https://github.com/vsingh-group/LCODEC-deep-unlearning

Added

2026-09-26

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Paul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou, Louis-Philippe Morency, Ruslan Salakhutdinov

OrganizationsCarnegie Mellon UniversityStanford UniversityUniversity of Pennsylvania

Why you should read this

Proposes FACTORCL, a multimodal contrastive learning framework that overcomes standard multi-view redundancy limitations by factorizing representations into task-relevant shared and unique information via conditional mutual information bounds and multimodal augmentations.

In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumption of multi-view redundancy - that shared information between modalities is necessary and sufficient for downstream tasks. However, in many real-world settings, task-relevant information is also contained in modality-unique regions: information that is only present in one modality but still relevant to the task. How can we learn self-supervised multimodal representations to capture both shared and unique information relevant to downstream tasks? This paper proposes FACTORCL, a new multimodal representation learning method to go beyond multi-view redundancy. FACTORCL is built from three new contributions: (1) factorizing task-relevant information into shared and unique representations, (2) capturing task-relevant information via maximizing MI lower bounds and removing task-irrelevant information via minimizing MI upper bounds, and (3) multimodal data augmentations to approximate task relevance without labels. On large-scale real-world datasets, FACTORCL captures both shared and unique information and achieves state-of-the-art results on six benchmarks.

Added

2026-09-26