Built independently by an author, for readers. Read the story and support ChapterPal

keyword

cross-view prediction

Cross-view prediction is a self-supervised representation learning technique in which a machine learning model is trained to predict, reconstruct, or infer one view, channel, or modality of a data sample given another view of the same underlying entity or scene. In this framework, different perspectives of an input—such as separate color channels, audio and visual streams, distinct sensor readings, or spatial viewpoints—serve mutually as inputs and supervision targets. By learning to map between complementary yet incomplete observations without human annotations, the model captures shared structural, geometric, and semantic factors that remain consistent across varying observations. Unlike contrastive methods that optimize representation distances through instance discrimination in an embedding space, cross-view prediction typically relies on generative or predictive reconstruction objectives across different perceptual domains.

1 item

Contrastive Multiview Coding

Contrastive Multiview Coding

Yonglong Tian, Dilip Krishnan, Phillip Isola

OrganizationsGoogleMassachusetts Institute of Technology

Why you should read this

Introduces Contrastive Multiview Coding, a scalable self-supervised framework that maximizes mutual information across multiple sensory views, demonstrating that contrastive objectives outperform predictive reconstruction and that representation quality consistently improves as more views are integrated.

Humans view the world through many sensory channels, e.g., the long-wavelength light channel, viewed by the left eye, or the high-frequency vibrations channel, heard by the right ear. Each view is noisy and incomplete, but important factors, such as physics, geometry, and semantics, tend to be shared between all views (e.g., a “dog” can be seen, heard, and felt). We investigate the classic hypothesis that a powerful representation is one that models view-invariant factors. We study this hypothesis under the framework of multiview contrastive learning, where we learn a representation that aims to maximize mutual information between different views of the same scene but is otherwise compact. Our approach scales to any number of views, and is view-agnostic. We analyze key properties of the approach that make it work, finding that the contrastive loss outperforms a popular alternative based on cross-view prediction, and that the more views we learn from, the better the resulting representation captures underlying scene semantics. Our approach achieves state-of-the-art results on image and video unsupervised learning benchmarks. Code is released at: http://github.com/HobbitLong/CMC/.

Added

2026-09-14