keyword
mutual information
Mutual information is a fundamental concept in information theory and probability that measures the amount of information shared between two random variables, or the extent to which knowledge of one variable reduces uncertainty about the other. Unlike simple linear correlation, mutual information captures all forms of statistical dependency, including complex non-linear relationships. Mathematically, it is defined as the Kullback-Leibler divergence between the joint probability distribution of the variables and the product of their marginal distributions, which is equivalent to the difference between the marginal Shannon entropy of a variable and its conditional entropy given the other variable. Mutual information is always non-negative, symmetric between the two variables, and equals zero if and only if the variables are completely statistically independent. Consequently, it serves as a central metric for optimizing representation learning, feature selection, channel capacity, and model evaluation across machine learning, statistics, and communication theory.
14 items

Learning to Maximize Mutual Information for Dynamic Feature Selection
Ian Connick Covert, Wei Qiu, Mingyu Lu, Nayoon Kim, Nathan J. White, Su-In Lee
Why you should read this
Proposes an amortized optimization framework for dynamic feature selection that directly learns greedy conditional mutual information policies from standard labeled data, bypassing the instability of reinforcement learning and the cost of generative modeling.
Feature selection helps reduce data acquisition costs in ML, but the standard approach is to train models with static feature subsets. Here, we consider the dynamic feature selection (DFS) problem where a model sequentially queries features based on the presently available information. DFS is often addressed with reinforcement learning, but we explore a simpler approach of greedily selecting features based on their conditional mutual information. This method is theoretically appealing but requires oracle access to the data distribution, so we develop a learning approach based on amortized optimization. The proposed method is shown to recover the greedy policy when trained to optimality, and it outperforms numerous existing feature selection methods in our experiments, thus validating it as a simple but powerful approach for this problem.
Added
2026-10-02

How Does Information Bottleneck Help Deep Learning?
Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang Huang
Why you should read this
Establishes the first rigorous theoretical foundation linking the information bottleneck principle to generalization in deep neural networks by deriving novel sample complexity bounds that depend directly on intermediate layer compression rather than parameter counts.
Numerous deep learning algorithms have been inspired by and understood via the notion of information bottleneck, where unnecessary information is (often implicitly) minimized while task-relevant information is maximized. However, a rigorous argument for justifying why it is desirable to control information bottlenecks has been elusive. In this paper, we provide the first rigorous learning theory for justifying the benefit of information bottleneck in deep learning by mathematically relating information bottleneck to generalization errors. Our theory proves that controlling information bottleneck is one way to control generalization errors in deep learning, although it is not the only or necessary way. We investigate the merit of our new mathematical findings with experiments across a range of architectures and learning settings. In many cases, generalization errors are shown to correlate with the degree of information bottleneck: i.e., the amount of the unnecessary information at hidden layers. This paper provides a theoretical foundation for current and future methods through the lens of information bottleneck. Our new generalization bounds scale with the degree of information bottleneck, unlike the previous bounds that scale with the number of parameters, VC dimension, Rademacher complexity, stability or robustness. Our code is publicly available at: https://github.com/xu-ji/information-bottleneck
Added
2026-10-01

An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels
Taylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw, Kyle Jeffrey Rogers, Alexia Pauline Delorey, Mahmoud Khalil, Nancy Fulda, David Wingate
Why you should read this
Proposes an unsupervised, black-box prompt selection method that optimizes mutual information between inputs and language model outputs to identify top-performing prompt templates without requiring labeled data or access to model weights.
Pre-trained language models derive substantial linguistic and factual knowledge from the massive corpora on which they are trained, and prompt engineering seeks to align these models to specific tasks. Unfortunately, existing prompt engineering methods require significant amounts of labeled data, access to model parameters, or both. We introduce a new method for selecting prompt templates without labeled examples and without direct access to the model. Specifically, over a set of candidate templates, we choose the template that maximizes the mutual information between the input and the corresponding model output. Across 8 datasets representing 7 distinct NLP tasks, we show that when a template has high mutual information, it also has high accuracy on the task. On the largest model, selecting prompts with our method gets 90% of the way from the average prompt accuracy to the best prompt accuracy and requires no ground truth labels.
Added
2026-10-01

InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models
Yingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan, Fei Wang, Christopher De Sa, Volodymyr Kuleshov
Why you should read this
Proposes InfoDiffusion, a framework that integrates mutual information regularization into diffusion models to extract semantically meaningful, disentangled low-dimensional latent representations without sacrificing generative sample quality.
While diffusion models excel at generating high-quality samples, their latent variables typically lack semantic meaning and are not suitable for representation learning. Here, we propose InfoDiffusion, an algorithm that augments diffusion models with low-dimensional latent variables that capture high-level factors of variation in the data. InfoDiffusion relies on a learning objective regularized with the mutual information between observed and hidden variables, which improves latent space quality and prevents the latents from being ignored by expressive diffusion-based decoders. Empirically, we find that InfoDiffusion learns disentangled and human-interpretable latent representations that are competitive with state-of-the-art generative and contrastive methods, while retaining the high sample quality of diffusion models. Our method enables manipulating the attributes of generated images and has the potential to assist tasks that require exploring a learned latent space to generate quality samples, e.g., generative design.
Added
2026-09-26

Disentangling by Factorising
Hyunjik Kim, Andriy Mnih
Why you should read this
Proposes FactorVAE, an unsupervised representation learning method that encourages latent factor independence to achieve a superior trade-off between disentanglement and reconstruction quality compared to -VAE, while introducing a more reliable evaluation metric for disentangled representations.
We define and address the problem of unsupervised learning of disentangled representations on data generated from independent factors of variation. We propose FactorVAE, a method that disentangles by encouraging the distribution of representations to be factorial and hence independent across the dimensions. We show that it improves upon -VAE by providing a better trade-off between disentanglement and reconstruction quality. Moreover, we highlight the problems of a commonly used disentanglement metric and introduce a new metric that does not suffer from them.
Added
2026-09-25


Rényi Divergence and Kullback-Leibler Divergence
Tim van Erven, Peter Harremoës
Why you should read this
Develops a rigorous foundation for Rényi divergence by establishing its essential analytic properties, generalizing the Pythagorean inequality to arbitrary orders, and extending channel capacity and minimax redundancy equivalences to continuous inputs.
Rényi divergence is related to Rényi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by Rényi as a measure of information that satisfies almost the same axioms as Kullback-Leibler divergence, and depends on a parameter that is called its order. In particular, the Rényi divergence of order 1 equals the Kullback-Leibler divergence. We review and extend the most important properties of Rényi divergence and Kullback-Leibler divergence, including convexity, continuity, limits of -algebras and the relation of the special order 0 to the Gaussian dichotomy and contiguity. We also show how to generalize the Pythagorean inequality to orders different from 1, and we extend the known equivalence between channel capacity and minimax redundancy to continuous channel inputs (for all orders) and present several other minimax results.
Added
2026-09-24

Opening the Black Box of Deep Neural Networks via Information
Ravid Shwartz-Ziv, Naftali Tishby
Why you should read this
Demonstrates through the Information Bottleneck principle that deep neural network training consists of distinct label-fitting and representation-compression phases, explaining theoretically why hidden layers accelerate generalization.
Despite their great success, there is still no comprehensive theoretical understanding of learning with Deep Neural Networks (DNNs) or their inner organization. Previous work proposed to analyze DNNs in the \textit{Information Plane}; i.e., the plane of the Mutual Information values that each layer preserves on the input and output variables. They suggested that the goal of the network is to optimize the Information Bottleneck (IB) tradeoff between compression and prediction, successively, for each layer. In this work we follow up on this idea and demonstrate the effectiveness of the Information-Plane visualization of DNNs. Our main results are: (i) most of the training epochs in standard DL are spent on {\emph compression} of the input to efficient representation and not on fitting the training labels. (ii) The representation compression phase begins when the training errors becomes small and the Stochastic Gradient Decent (SGD) epochs change from a fast drift to smaller training error into a stochastic relaxation, or random diffusion, constrained by the training error value. (iii) The converged layers lie on or very close to the Information Bottleneck (IB) theoretical bound, and the maps from the input to any hidden layer and from this hidden layer to the output satisfy the IB self-consistent equations. This generalization through noise mechanism is unique to Deep Neural Networks and absent in one layer networks. (iv) The training time is dramatically reduced when adding more hidden layers. Thus the main advantage of the hidden layers is computational. This can be explained by the reduced relaxation time, as this it scales super-linearly (exponentially for simple diffusion) with the information compression from the previous layer.
Added
2026-09-24

Automatic Retrieval and Clustering of Similar Words
Dekang Lin
Why you should read this
Proposes an information-theoretic distributional similarity measure based on dependency parse triples to automatically construct broad-coverage thesauri from unannotated text corpora.
Bootstrapping semantics from text is one of the greatest challenges in natural language learning. We first define a word similarity measure based on the distributional pattern of words. The similarity measure allows us to construct a thesaurus using a parsed corpus. We then present a new evaluation methodology for the automatically constructed thesaurus. The evaluation results show that the thesaurus is significantly closer to WordNet than Roget Thesaurus is.
Added
2026-09-18

Deep learning and the information bottleneck principle
Naftali Tishby, Noga Zaslavsky
Why you should read this
Establishes an information-theoretic framework for deep learning by applying the Information Bottleneck principle to explain how successive layers compress input data while preserving target information to achieve generalization.
Deep Neural Networks (DNNs) are analyzed via the theoretical framework of the information bottleneck (IB) principle. We first show that any DNN can be quantified by the mutual information between the layers and the input and output variables. Using this representation we can calculate the optimal information theoretic limits of the DNN and obtain finite sample generalization bounds. The advantage of getting closer to the theoretical limit is quantifiable both by the generalization bound and by the network's simplicity. We argue that both the optimal architecture, number of layers and features/connections at each layer, are related to the bifurcation points of the information bottleneck tradeoff, namely, relevant compression of the input layer with respect to the output layer. The hierarchical representations at the layered network naturally correspond to the structural phase transitions along the information curve. We believe that this new insight can lead to new optimality bounds and deep learning algorithms.
Added
2026-09-18

Deep Variational Information Bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon
Why you should read this
Presents a variational framework for parameterizing the Information Bottleneck principle with deep neural networks, providing an effective regularization method that improves generalization and resilience against adversarial attacks.
We present a variational approximation to the information bottleneck of Tishby et al. (1999). This variational approach allows us to parameterize the information bottleneck model using a neural network and leverage the reparameterization trick for efficient training. We call this method "Deep Variational Information Bottleneck", or Deep VIB. We show that models trained with the VIB objective outperform those that are trained with other forms of regularization, in terms of generalization performance and robustness to adversarial attack.
Added
2026-09-16

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance
X. Nguyen, Julien Epps, James Bailey
Why you should read this
Establishes which information-theoretic clustering comparison measures satisfy metric, normalization, and chance-correction properties, motivating normalized information distance as a principled default.
Information theoretic measures form a fundamental class of measures for comparing clusterings, and have recently received increasing interest. Nevertheless, a number of questions concerning their properties and inter-relationships remain unresolved. In this paper, we perform an organized study of information theoretic measures for clustering comparison, including several existing popular measures in the literature, as well as some newly proposed ones. We discuss and prove their important properties, such as the metric property and the normalization property. We then highlight to the clustering community the importance of correcting information theoretic measures for chance, especially when the data size is small compared to the number of clusters present therein. Of the available information theoretic based measures, we advocate the normalized information distance (NID) as a general measure of choice, for it possesses concurrently several important properties, such as being both a metric and a normalized measure, admitting an exact analytical adjusted-for-chance form, and using the nominal [0, 1] range better than other normalized variants.
Added
2026-09-14

The information bottleneck method
Naftali Tishby, Fernando C. Pereira, William Bialek
Why you should read this
Proposes a fundamental information-theoretic framework and algorithm to compress input signals while preserving maximal information about a target variable, generalizing rate-distortion theory without requiring predefined distortion measures.
We define the relevant information in a signal as being the information that this signal provides about another signal . Examples include the information that face images provide about the names of the people portrayed, or the information that speech sounds provide about the words spoken. Understanding the signal requires more than just predicting , it also requires specifying which features of play a role in the prediction. We formalize this problem as that of finding a short code for that preserves the maximum information about . That is, we squeeze the information that provides about through a `bottleneck' formed by a limited set of codewords . This constrained optimization problem can be seen as a generalization of rate distortion theory in which the distortion measure emerges from the joint statistics of and . This approach yields an exact set of self consistent equations for the coding rules and . Solutions to these equations can be found by a convergent re-estimation method that generalizes the Blahut-Arimoto algorithm. Our variational principle provides a surprisingly rich framework for discussing a variety of problems in signal processing and learning, as will be described in detail elsewhere.
Added
2026-09-10

Word Association Norms, Mutual Information, and Lexicography
Kenneth Ward Church, Patrick Hanks
Why you should read this
Proposes a corpus-based mutual information metric that automates the extraction of semantic and syntactic word associations at scale, replacing expensive human subject testing for computational linguists and lexicographers.
The term word association is used in a very particular sense in the psycholinguistic literature. (Generally speaking, subjects respond quicker than normal to the word “nurse” if it follows a highly associated word such as “doctor.”) We will extend the term to provide the basis for a statistical description of a variety of interesting linguistic phenomena, ranging from semantic relations of the doctor/nurse type (content word/content word) to lexico-syntactic co-occurrence constraints between verbs and prepositions (content word/function word). This paper will propose a new objective measure based on the information theoretic notion of mutual information, for estimating word association norms from computer readable corpora. (The standard method of obtaining word association norms, testing a few thousand subjects on a few hundred words, is both costly and unreliable.) The proposed measure, the association ratio, estimates word association norms directly from computer readable corpora, making it possible to estimate norms for tens of thousands of words.
Added
2026-09-10

An exact information theory of generalization phase transitions in Bayesian diffusion models
Henry Hunt, Mason Kamb, Surya Ganguli
Why you should read this
Establishes an exact information-theoretic phase boundary between memorization and generalization in Bayesian diffusion models, explaining how spatial information restriction enables generative AI to overcome the curse of dimensionality.
How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models \textit{early in training}, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.
Added
2026-09-02

