Built independently by an author, for readers. Read the story and support ChapterPal

keyword

optimal multimodal augmentation

An optimal multimodal augmentation refers to a transformed version of multimodal data where the only information preserved between the original inputs and the augmented counterparts is strictly relevant to a downstream task, free from task-irrelevant noise. Formulated within the framework of self-supervised representation learning and information theory, this condition occurs when the mutual information between the original multimodal inputs and their augmented views exactly equals the mutual information between the original inputs and the target task label. In practice, such augmentations perfectly preserve all essential task signals, capturing both redundant features shared across modalities and distinct features unique to individual modalities, while altering or removing superfluous background details and noise. By maintaining task-relevant properties invariant across views, optimal multimodal augmentations provide a theoretical benchmark that allows models to isolate sufficient, task-specific representations from unlabeled multimodal data.

1 item

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Paul Pu Liang, Zihao Deng, Martin Q. Ma, James Y. Zou, Louis-Philippe Morency, Ruslan Salakhutdinov

OrganizationsCarnegie Mellon UniversityStanford UniversityUniversity of Pennsylvania

Why you should read this

Proposes FACTORCL, a multimodal contrastive learning framework that overcomes standard multi-view redundancy limitations by factorizing representations into task-relevant shared and unique information via conditional mutual information bounds and multimodal augmentations.

In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumption of multi-view redundancy - that shared information between modalities is necessary and sufficient for downstream tasks. However, in many real-world settings, task-relevant information is also contained in modality-unique regions: information that is only present in one modality but still relevant to the task. How can we learn self-supervised multimodal representations to capture both shared and unique information relevant to downstream tasks? This paper proposes FACTORCL, a new multimodal representation learning method to go beyond multi-view redundancy. FACTORCL is built from three new contributions: (1) factorizing task-relevant information into shared and unique representations, (2) capturing task-relevant information via maximizing MI lower bounds and removing task-irrelevant information via minimizing MI upper bounds, and (3) multimodal data augmentations to approximate task relevance without labels. On large-scale real-world datasets, FACTORCL captures both shared and unique information and achieves state-of-the-art results on six benchmarks.

Added

2026-09-26