keyword
contrastive multiview coding
Contrastive multiview coding is a self-supervised representation learning framework designed to extract view-invariant features by maximizing the mutual information between multiple distinct views or sensory channels of the same scene or data instance [0, 1]. Built on the principle that essential semantic and structural properties are shared across diverse sensory perspectives while view-specific noise is discarded, the framework maps different views—such as separate color channels, sensory modalities, or data augmentations—into latent representations [0, 1]. It trains neural network encoders using a contrastive loss objective that pulls representations of matching views from the same instance closer together while pushing apart representations from different instances [0, 1]. This view-agnostic approach scales across multiple modalities to produce compact, generalizable data representations without relying on human annotations [0, 1].
4 items

CROMA: Remote Sensing Representations with Contrastive Radar-Optical Masked Autoencoders
Anthony Fuller, Koreen Millard, James R. Green
Why you should read this
Introduces a self-supervised foundation model that combines cross-modal contrastive learning and masked autoencoding across spatially aligned optical and SAR satellite imagery, enabling flexible unimodal or multimodal inference and zero-shot spatial extrapolation up to 17.6× training resolution.
Added
2026-10-05

Robust Cross-Modal Representation Learning with Progressive Self-Distillation
Alex Andonian, Shixing Chen, Raffay Hamid
Why you should read this
Proposes a progressive self-distillation framework that replaces rigid one-to-one pairings in vision-language pretraining with dynamic soft-alignment targets, consistently outperforming CLIP across zero-shot, transfer, and retrieval benchmarks without extra computational overhead.
The learning objective of vision-language approach of CLIP [63] does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets, which contributes to its compute and data inefficiency. To address this challenge, we introduce a novel training framework based on cross-modal contrastive learning that uses progressive self-distillation and soft image-text alignments to more efficiently learn robust representations from noisy data. Our model distills its own knowledge to dynamically generate soft-alignment targets for a subset of images and captions in every minibatch, which are then used to update its parameters. Extensive evaluation across 14 benchmark datasets shows that our method consistently outperforms its CLIP counterpart in multiple settings, including: (a) zero-shot classification, (b) linear probe transfer, and (c) image-text retrieval, without incurring extra computational cost. Analysis using an ImageNet-based robustness test-bed [70] reveals that our method offers better effective robustness to natural distribution shifts compared to both ImageNet-trained models and CLIP itself. Lastly, pretraining with datasets spanning two orders of magnitude in size shows that our improvements over CLIP tend to scale with number of training examples.
Added
2026-09-26

FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space
Shengzhong Liu, Tomoyoshi Kimura, Dongxin Liu, Ruijie Wang, Jinyang Li, Suhas N. Diggavi, Mani B. Srivastava, Tarek F. Abdelzaher
Why you should read this
Proposes a self-supervised contrastive learning framework that factorizes multimodal time-series signals into orthogonal shared and private latent spaces alongside statistical temporal constraints to achieve state-of-the-art representation quality across diverse sensing datasets.
This paper proposes a novel contrastive learning framework, called FOCAL, for extracting comprehensive features from multimodal time-series sensing signals through self-supervised training. Existing multimodal contrastive frameworks mostly rely on the shared information between sensory modalities, but do not explicitly consider the exclusive modality information that could be critical to understanding the underlying sensing physics. Besides, contrastive frameworks for time series have not handled the temporal information locality appropriately. FOCAL solves these challenges by making the following contributions: First, given multimodal time series, it encodes each modality into a factorized latent space consisting of shared features and private features that are orthogonal to each other. The shared space emphasizes feature patterns consistent across sensory modalities through a modal-matching objective. In contrast, the private space extracts modality-exclusive information through a transformation-invariant objective. Second, we propose a temporal structural constraint for modality features, such that the average distance between temporally neighboring samples is no larger than that of temporally distant samples. Extensive evaluations are performed on four multimodal sensing datasets with two backbone encoders and two classifiers to demonstrate the superiority of FOCAL. It consistently outperforms the state-of-the-art baselines in downstream tasks with a clear margin, under different ratios of available labels. The code and self-collected dataset are available at https://github.com/tomoyoshki/focal.
Added
2026-09-26

Contrastive Multiview Coding
Yonglong Tian, Dilip Krishnan, Phillip Isola
Why you should read this
Introduces Contrastive Multiview Coding, a scalable self-supervised framework that maximizes mutual information across multiple sensory views, demonstrating that contrastive objectives outperform predictive reconstruction and that representation quality consistently improves as more views are integrated.
Humans view the world through many sensory channels, e.g., the long-wavelength light channel, viewed by the left eye, or the high-frequency vibrations channel, heard by the right ear. Each view is noisy and incomplete, but important factors, such as physics, geometry, and semantics, tend to be shared between all views (e.g., a “dog” can be seen, heard, and felt). We investigate the classic hypothesis that a powerful representation is one that models view-invariant factors. We study this hypothesis under the framework of multiview contrastive learning, where we learn a representation that aims to maximize mutual information between different views of the same scene but is otherwise compact. Our approach scales to any number of views, and is view-agnostic. We analyze key properties of the approach that make it work, finding that the contrastive loss outperforms a popular alternative based on cross-view prediction, and that the more views we learn from, the better the resulting representation captures underlying scene semantics. Our approach achieves state-of-the-art results on image and video unsupervised learning benchmarks. Code is released at: http://github.com/HobbitLong/CMC/.
Added
2026-09-14
