Built independently by an author, for readers. Read the story and support ChapterPal

keyword

mutual information

Mutual information is a fundamental concept in information theory and probability that measures the amount of information shared between two random variables, or the extent to which knowledge of one variable reduces uncertainty about the other. Unlike simple linear correlation, mutual information captures all forms of statistical dependency, including complex non-linear relationships. Mathematically, it is defined as the Kullback-Leibler divergence between the joint probability distribution of the variables and the product of their marginal distributions, which is equivalent to the difference between the marginal Shannon entropy of a variable and its conditional entropy given the other variable. Mutual information is always non-negative, symmetric between the two variables, and equals zero if and only if the variables are completely statistically independent. Consequently, it serves as a central metric for optimizing representation learning, feature selection, channel capacity, and model evaluation across machine learning, statistics, and communication theory.

14 items

How Does Information Bottleneck Help Deep Learning?

How Does Information Bottleneck Help Deep Learning?

Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang Huang

OrganizationsColumbia UniversityMilaNational University of SingaporeUniversity of Pennsylvania

Why you should read this

Establishes the first rigorous theoretical foundation linking the information bottleneck principle to generalization in deep neural networks by deriving novel sample complexity bounds that depend directly on intermediate layer compression rather than parameter counts.

Numerous deep learning algorithms have been inspired by and understood via the notion of information bottleneck, where unnecessary information is (often implicitly) minimized while task-relevant information is maximized. However, a rigorous argument for justifying why it is desirable to control information bottlenecks has been elusive. In this paper, we provide the first rigorous learning theory for justifying the benefit of information bottleneck in deep learning by mathematically relating information bottleneck to generalization errors. Our theory proves that controlling information bottleneck is one way to control generalization errors in deep learning, although it is not the only or necessary way. We investigate the merit of our new mathematical findings with experiments across a range of architectures and learning settings. In many cases, generalization errors are shown to correlate with the degree of information bottleneck: i.e., the amount of the unnecessary information at hidden layers. This paper provides a theoretical foundation for current and future methods through the lens of information bottleneck. Our new generalization bounds scale with the degree of information bottleneck, unlike the previous bounds that scale with the number of parameters, VC dimension, Rademacher complexity, stability or robustness. Our code is publicly available at: https://github.com/xu-ji/information-bottleneck

Added

2026-10-01

An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels

An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels

Taylor Sorensen, Joshua Robinson, Christopher Michael Rytting, Alexander Glenn Shaw, Kyle Jeffrey Rogers, Alexia Pauline Delorey, Mahmoud Khalil, Nancy Fulda, David Wingate

OrganizationsBrigham Young University

Why you should read this

Proposes an unsupervised, black-box prompt selection method that optimizes mutual information between inputs and language model outputs to identify top-performing prompt templates without requiring labeled data or access to model weights.

Pre-trained language models derive substantial linguistic and factual knowledge from the massive corpora on which they are trained, and prompt engineering seeks to align these models to specific tasks. Unfortunately, existing prompt engineering methods require significant amounts of labeled data, access to model parameters, or both. We introduce a new method for selecting prompt templates without labeled examples and without direct access to the model. Specifically, over a set of candidate templates, we choose the template that maximizes the mutual information between the input and the corresponding model output. Across 8 datasets representing 7 distinct NLP tasks, we show that when a template has high mutual information, it also has high accuracy on the task. On the largest model, selecting prompts with our method gets 90% of the way from the average prompt accuracy to the best prompt accuracy and requires no ground truth labels.

Added

2026-10-01

InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models

InfoDiffusion: Representation Learning Using Information Maximizing Diffusion Models

Yingheng Wang, Yair Schiff, Aaron Gokaslan, Weishen Pan, Fei Wang, Christopher De Sa, Volodymyr Kuleshov

OrganizationsCornell University

Why you should read this

Proposes InfoDiffusion, a framework that integrates mutual information regularization into diffusion models to extract semantically meaningful, disentangled low-dimensional latent representations without sacrificing generative sample quality.

While diffusion models excel at generating high-quality samples, their latent variables typically lack semantic meaning and are not suitable for representation learning. Here, we propose InfoDiffusion, an algorithm that augments diffusion models with low-dimensional latent variables that capture high-level factors of variation in the data. InfoDiffusion relies on a learning objective regularized with the mutual information between observed and hidden variables, which improves latent space quality and prevents the latents from being ignored by expressive diffusion-based decoders. Empirically, we find that InfoDiffusion learns disentangled and human-interpretable latent representations that are competitive with state-of-the-art generative and contrastive methods, while retaining the high sample quality of diffusion models. Our method enables manipulating the attributes of generated images and has the potential to assist tasks that require exploring a learned latent space to generate quality samples, e.g., generative design.

Added

2026-09-26

Opening the Black Box of Deep Neural Networks via Information

Opening the Black Box of Deep Neural Networks via Information

Ravid Shwartz-Ziv, Naftali Tishby

OrganizationsSchool of Engineering and Computer ScienceThe Hebrew University of Jerusalem

Why you should read this

Demonstrates through the Information Bottleneck principle that deep neural network training consists of distinct label-fitting and representation-compression phases, explaining theoretically why hidden layers accelerate generalization.

Despite their great success, there is still no comprehensive theoretical understanding of learning with Deep Neural Networks (DNNs) or their inner organization. Previous work proposed to analyze DNNs in the \textit{Information Plane}; i.e., the plane of the Mutual Information values that each layer preserves on the input and output variables. They suggested that the goal of the network is to optimize the Information Bottleneck (IB) tradeoff between compression and prediction, successively, for each layer. In this work we follow up on this idea and demonstrate the effectiveness of the Information-Plane visualization of DNNs. Our main results are: (i) most of the training epochs in standard DL are spent on {\emph compression} of the input to efficient representation and not on fitting the training labels. (ii) The representation compression phase begins when the training errors becomes small and the Stochastic Gradient Decent (SGD) epochs change from a fast drift to smaller training error into a stochastic relaxation, or random diffusion, constrained by the training error value. (iii) The converged layers lie on or very close to the Information Bottleneck (IB) theoretical bound, and the maps from the input to any hidden layer and from this hidden layer to the output satisfy the IB self-consistent equations. This generalization through noise mechanism is unique to Deep Neural Networks and absent in one layer networks. (iv) The training time is dramatically reduced when adding more hidden layers. Thus the main advantage of the hidden layers is computational. This can be explained by the reduced relaxation time, as this it scales super-linearly (exponentially for simple diffusion) with the information compression from the previous layer.

Added

2026-09-24

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance

Information Theoretic Measures for Clusterings Comparison: Variants, Properties, Normalization and Correction for Chance

X. Nguyen, Julien Epps, James Bailey

OrganizationsCSIRO’s Data61University of MelbourneUniversity of New South Wales

Why you should read this

Establishes which information-theoretic clustering comparison measures satisfy metric, normalization, and chance-correction properties, motivating normalized information distance as a principled default.

Information theoretic measures form a fundamental class of measures for comparing clusterings, and have recently received increasing interest. Nevertheless, a number of questions concerning their properties and inter-relationships remain unresolved. In this paper, we perform an organized study of information theoretic measures for clustering comparison, including several existing popular measures in the literature, as well as some newly proposed ones. We discuss and prove their important properties, such as the metric property and the normalization property. We then highlight to the clustering community the importance of correcting information theoretic measures for chance, especially when the data size is small compared to the number of clusters present therein. Of the available information theoretic based measures, we advocate the normalized information distance (NID) as a general measure of choice, for it possesses concurrently several important properties, such as being both a metric and a normalized measure, admitting an exact analytical adjusted-for-chance form, and using the nominal [0, 1] range better than other normalized variants.

Added

2026-09-14

The information bottleneck method

The information bottleneck method

Naftali Tishby, Fernando C. Pereira, William Bialek

OrganizationsAT&T Labs—ResearchCenter for Neural ComputationInstitute for Computer ScienceNEC Laboratories America, Inc.The Hebrew University of Jerusalem

Why you should read this

Proposes a fundamental information-theoretic framework and algorithm to compress input signals while preserving maximal information about a target variable, generalizing rate-distortion theory without requiring predefined distortion measures.

We define the relevant information in a signal x∈Xx\in X as being the information that this signal provides about another signal y∈\Yy\in \Y. Examples include the information that face images provide about the names of the people portrayed, or the information that speech sounds provide about the words spoken. Understanding the signal xx requires more than just predicting yy, it also requires specifying which features of \X\X play a role in the prediction. We formalize this problem as that of finding a short code for \X\X that preserves the maximum information about \Y\Y. That is, we squeeze the information that \X\X provides about \Y\Y through a `bottleneck' formed by a limited set of codewords \tX\tX. This constrained optimization problem can be seen as a generalization of rate distortion theory in which the distortion measure d(x,\x)d(x,\x) emerges from the joint statistics of \X\X and \Y\Y. This approach yields an exact set of self consistent equations for the coding rules X→\tXX \to \tX and \tX→\Y\tX \to \Y. Solutions to these equations can be found by a convergent re-estimation method that generalizes the Blahut-Arimoto algorithm. Our variational principle provides a surprisingly rich framework for discussing a variety of problems in signal processing and learning, as will be described in detail elsewhere.

Added

2026-09-10

Word Association Norms, Mutual Information, and Lexicography

Word Association Norms, Mutual Information, and Lexicography

Kenneth Ward Church, Patrick Hanks

OrganizationsBell LaboratoriesCollins Publishers

Why you should read this

Proposes a corpus-based mutual information metric that automates the extraction of semantic and syntactic word associations at scale, replacing expensive human subject testing for computational linguists and lexicographers.

The term word association is used in a very particular sense in the psycholinguistic literature. (Generally speaking, subjects respond quicker than normal to the word “nurse” if it follows a highly associated word such as “doctor.”) We will extend the term to provide the basis for a statistical description of a variety of interesting linguistic phenomena, ranging from semantic relations of the doctor/nurse type (content word/content word) to lexico-syntactic co-occurrence constraints between verbs and prepositions (content word/function word). This paper will propose a new objective measure based on the information theoretic notion of mutual information, for estimating word association norms from computer readable corpora. (The standard method of obtaining word association norms, testing a few thousand subjects on a few hundred words, is both costly and unreliable.) The proposed measure, the association ratio, estimates word association norms directly from computer readable corpora, making it possible to estimate norms for tens of thousands of words.

Added

2026-09-10

An exact information theory of generalization phase transitions in Bayesian diffusion models

An exact information theory of generalization phase transitions in Bayesian diffusion models

Henry Hunt, Mason Kamb, Surya Ganguli

OrganizationsStanford University

Why you should read this

Establishes an exact information-theoretic phase boundary between memorization and generalization in Bayesian diffusion models, explaining how spatial information restriction enables generative AI to overcome the curse of dimensionality.

How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models \textit{early in training}, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.

Added

2026-09-02

Creative Commons License