Built independently by an author, for readers. Read the story and support ChapterPal

keyword

SplitBrain autoencoders

Split-brain autoencoders are neural network architectures designed for self-supervised representation learning by performing cross-channel prediction instead of full input reconstruction. In a standard autoencoder, the model learns by reconstructing the complete input signal, which can risk learning trivial identity mappings. A split-brain autoencoder instead divides the input data across its channels or modalities into disjoint subsets, such as separating luminance from color channels or visual appearance from depth. The network is divided into separate subnetworks, where each subnetwork receives one subset of the channels as input and is trained to predict the complementary subset. By compelling the subnetworks to solve these cross-channel estimation tasks, the model learns high-level semantic and invariant feature representations across the full input signal, which can be aggregated and transferred effectively to downstream tasks without requiring manual supervision.

1 item

Contrastive Multiview Coding

Contrastive Multiview Coding

Yonglong Tian, Dilip Krishnan, Phillip Isola

OrganizationsGoogleMassachusetts Institute of Technology

Why you should read this

Introduces Contrastive Multiview Coding, a scalable self-supervised framework that maximizes mutual information across multiple sensory views, demonstrating that contrastive objectives outperform predictive reconstruction and that representation quality consistently improves as more views are integrated.

Humans view the world through many sensory channels, e.g., the long-wavelength light channel, viewed by the left eye, or the high-frequency vibrations channel, heard by the right ear. Each view is noisy and incomplete, but important factors, such as physics, geometry, and semantics, tend to be shared between all views (e.g., a “dog” can be seen, heard, and felt). We investigate the classic hypothesis that a powerful representation is one that models view-invariant factors. We study this hypothesis under the framework of multiview contrastive learning, where we learn a representation that aims to maximize mutual information between different views of the same scene but is otherwise compact. Our approach scales to any number of views, and is view-agnostic. We analyze key properties of the approach that make it work, finding that the contrastive loss outperforms a popular alternative based on cross-view prediction, and that the more views we learn from, the better the resulting representation captures underlying scene semantics. Our approach achieves state-of-the-art results on image and video unsupervised learning benchmarks. Code is released at: http://github.com/HobbitLong/CMC/.

Added

2026-09-14