keyword
predictive coding
Predictive coding is a computational and neuroscientific framework in which a system learns representations by generating predictions about incoming sensory inputs or internal latent states and updating its internal model based on the discrepancy between expected and observed signals. Originating as a theory of how biological brains process perception through top-down predictions and bottom-up error signals, the concept has become foundational in artificial intelligence and self-supervised learning. In neural network architectures, predictive coding techniques train models to anticipate future time steps, infer masked regions of data, or predict their own latent representations across different views, enabling the extraction of structured and semantically rich features without requiring manual annotations.
5 items

Contrastive Multiview Coding
Yonglong Tian, Dilip Krishnan, Phillip Isola
Why you should read this
Introduces Contrastive Multiview Coding, a scalable self-supervised framework that maximizes mutual information across multiple sensory views, demonstrating that contrastive objectives outperform predictive reconstruction and that representation quality consistently improves as more views are integrated.
Humans view the world through many sensory channels, e.g., the long-wavelength light channel, viewed by the left eye, or the high-frequency vibrations channel, heard by the right ear. Each view is noisy and incomplete, but important factors, such as physics, geometry, and semantics, tend to be shared between all views (e.g., a “dog” can be seen, heard, and felt). We investigate the classic hypothesis that a powerful representation is one that models view-invariant factors. We study this hypothesis under the framework of multiview contrastive learning, where we learn a representation that aims to maximize mutual information between different views of the same scene but is otherwise compact. Our approach scales to any number of views, and is view-agnostic. We analyze key properties of the approach that make it work, finding that the contrastive loss outperforms a popular alternative based on cross-view prediction, and that the more views we learn from, the better the resulting representation captures underlying scene semantics. Our approach achieves state-of-the-art results on image and video unsupervised learning benchmarks. Code is released at: http://github.com/HobbitLong/CMC/.
Added
2026-09-14

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, Abdelrahman Mohamed
Why you should read this
Introduces HuBERT, a self-supervised speech representation framework that applies masked prediction over offline k-means cluster targets to learn combined acoustic and language models from continuous audio, outperforming wav2vec 2.0 on speech recognition benchmarks.
Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase, and (3) sound units have variable lengths with no explicit segmentation. To deal with these three problems, we propose the Hidden-Unit BERT (HuBERT) approach for self-supervised speech representation learning, which utilizes an offline clustering step to provide aligned target labels for a BERT-like prediction loss. A key ingredient of our approach is applying the prediction loss over the masked regions only, which forces the model to learn a combined acoustic and language model over the continuous inputs. HuBERT relies primarily on the consistency of the unsupervised clustering step rather than the intrinsic quality of the assigned cluster labels. Starting with a simple k-means teacher of 100 clusters, and using two iterations of clustering, the HuBERT model either matches or improves upon the state-of-the-art wav2vec 2.0 performance on the Librispeech (960h) and Libri-light (60,000h) benchmarks with 10min, 1h, 10h, 100h, and 960h fine-tuning subsets. Using a 1B parameter model, HuBERT shows up to 19% and 13% relative WER reduction on the more challenging dev-other and test-other evaluation subsets.
Added
2026-09-13

Learning deep representations by mutual information estimation and maximization
R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Adam Trischler, Yoshua Bengio
Why you should read this
Introduces Deep InfoMax (DIM), an unsupervised representation learning method that maximizes mutual information between local input features and global encoder representations while using adversarial prior matching to rival supervised classification performance.
In this work, we perform unsupervised learning of representations by maximizing mutual information between an input and the output of a deep neural network encoder. Importantly, we show that structure matters: incorporating knowledge about locality of the input to the objective can greatly influence a representation's suitability for downstream tasks. We further control characteristics of the representation by matching to a prior distribution adversarially. Our method, which we call Deep InfoMax (DIM), outperforms a number of popular unsupervised learning methods and competes with fully-supervised learning on several classification tasks. DIM opens new avenues for unsupervised learning of representations and is an important step towards flexible formulations of representation-learning objectives for specific end-goals.
Added
2026-09-13

Learn from your own latents and not from tokens: A sample-complexity theory
Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
Why you should read this
Proves that latent prediction in generative models dramatically improves data efficiency, reducing sample complexity from exponential to constant, compared to token-level SSL, by effectively recovering hierarchical latent structures.
Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological learners require. An alternative paradigm has emerged in which networks are trained to predict their \emph{own} latent representations of related views or masked regions, as in data2vec and JEPA -- an idea related to predictive-coding accounts of the cortex. Despite strong empirical results, the theoretical understanding of these methods remains limited. Central questions include: by how much does latent prediction actually improve data efficiency? Is there a benefit to stacking such methods into multi-scale hierarchies? We answer both using as data a tractable probabilistic context-free grammar that captures the compositional structure of natural language and images. Such a grammar generates strings of visible tokens by recursively applying production rules along a tree of hidden symbols of depth . For such data, supervised or token-level SSL require a number of samples \emph{exponential} in to recover the latent tree; we prove that latent prediction achieves this with a number of samples \emph{constant} in , up to logarithmic factors. We confirm this bound with (i) a hierarchical clustering algorithm, (ii) an end-to-end neural network whose predictor-clusterer modules predict their own latents at each level via gradient descent, and (iii) the first sample-complexity analysis of data2vec, which we show implicitly performs hierarchical latent prediction. This suggests that explicit stacking such as H-JEPA is largely redundant.
Added
2026-06-08


Self-Supervised Learning of Pretext-Invariant Representations
Ishan Misra, Laurens van der Maaten
Why you should read this
Demonstrates that learning image representations to be invariant to transformations rather than predictive of them substantially improves the quality of self-supervised visual features, achieving state-of-the-art results and even surpassing supervised pre-training for object detection.
The goal of self-supervised learning from images is to construct image representations that are semantically meaningful via pretext tasks that do not require semantic annotations for a large training set of images. Many pretext tasks lead to representations that are covariant with image transformations. We argue that, instead, semantic representations ought to be invariant under such transformations. Specifically, we develop Pretext-Invariant Representation Learning (PIRL, pronounced as "pearl") that learns invariant representations based on pretext tasks. We use PIRL with a commonly used pretext task that involves solving jigsaw puzzles. We find that PIRL substantially improves the semantic quality of the learned image representations. Our approach sets a new state-of-the-art in self-supervised learning from images on several popular benchmarks for self-supervised learning. Despite being unsupervised, PIRL outperforms supervised pre-training in learning image representations for object detection. Altogether, our results demonstrate the potential of self-supervised learning of image representations with good invariance properties.
Added
2026-02-21
