Built independently by an author, for readers. Read the story and support ChapterPal

keyword

factorized contrastive learning

Factorized contrastive learning is a self-supervised representation learning framework that decomposes multimodal data into distinct shared and modality-unique representations. Unlike standard multimodal contrastive methods that capture only the redundant information overlapping across paired modalities, factorized contrastive learning explicitly isolates features common to all modalities from features unique to individual modalities. By optimizing information-theoretic objectives to preserve task-relevant shared and modality-specific signals while discarding irrelevant noise, this approach enables machine learning models to capture both cross-modal correlations and complementary, single-modality information for downstream tasks.

1 item

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Paul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou, Louis-Philippe Morency, Ruslan Salakhutdinov

OrganizationsCarnegie Mellon UniversityStanford UniversityUniversity of Pennsylvania

Why you should read this

Proposes FACTORCL, a multimodal contrastive learning framework that overcomes standard multi-view redundancy limitations by factorizing representations into task-relevant shared and unique information via conditional mutual information bounds and multimodal augmentations.

In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumption of multi-view redundancy - that shared information between modalities is necessary and sufficient for downstream tasks. However, in many real-world settings, task-relevant information is also contained in modality-unique regions: information that is only present in one modality but still relevant to the task. How can we learn self-supervised multimodal representations to capture both shared and unique information relevant to downstream tasks? This paper proposes FACTORCL, a new multimodal representation learning method to go beyond multi-view redundancy. FACTORCL is built from three new contributions: (1) factorizing task-relevant information into shared and unique representations, (2) capturing task-relevant information via maximizing MI lower bounds and removing task-irrelevant information via minimizing MI upper bounds, and (3) multimodal data augmentations to approximate task relevance without labels. On large-scale real-world datasets, FACTORCL captures both shared and unique information and achieves state-of-the-art results on six benchmarks.

Added

2026-09-26