Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multi-view redundancy

Multi-view redundancy is the presence of overlapping, shared information across multiple distinct perspectives, modalities, or sensory inputs representing the same underlying entity or event. In multimodal machine learning and representation learning, the term frequently describes the foundational assumption that this shared mutual information between views is both necessary and sufficient for downstream tasks. Under this paradigm, learning algorithms such as contrastive models maximize the similarity between representations from different modalities to capture invariant core semantics, treating view-unique variations as secondary or non-essential features.

1 item

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Paul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou, Louis-Philippe Morency, Ruslan Salakhutdinov

OrganizationsCarnegie Mellon UniversityStanford UniversityUniversity of Pennsylvania

Why you should read this

Proposes FACTORCL, a multimodal contrastive learning framework that overcomes standard multi-view redundancy limitations by factorizing representations into task-relevant shared and unique information via conditional mutual information bounds and multimodal augmentations.

In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumption of multi-view redundancy - that shared information between modalities is necessary and sufficient for downstream tasks. However, in many real-world settings, task-relevant information is also contained in modality-unique regions: information that is only present in one modality but still relevant to the task. How can we learn self-supervised multimodal representations to capture both shared and unique information relevant to downstream tasks? This paper proposes FACTORCL, a new multimodal representation learning method to go beyond multi-view redundancy. FACTORCL is built from three new contributions: (1) factorizing task-relevant information into shared and unique representations, (2) capturing task-relevant information via maximizing MI lower bounds and removing task-irrelevant information via minimizing MI upper bounds, and (3) multimodal data augmentations to approximate task relevance without labels. On large-scale real-world datasets, FACTORCL captures both shared and unique information and achieves state-of-the-art results on six benchmarks.

Added

2026-09-26