Factorized Contrastive Learning: Going Beyond Multi-view Redundancy
Paul Pu LiangZihao DengMartin Q. MaJames Y. ZouLouis-Philippe MorencyRuslan Salakhutdinov
Proposes FACTORCL, a multimodal contrastive learning framework that overcomes standard multi-view redundancy limitations by factorizing representations into task-relevant shared and unique information via conditional mutual information bounds and multimodal augmentations.
Modern multimodal machine learning relies heavily on contrastive self-supervised pre-training to learn useful representations from paired data, such as images matched with text or video paired with audio. However, standard methods rest on the assumption of multi-view redundancy: the premise that the only information needed for downstream tasks is the data shared between modalities. In real-world domains such as healthcare diagnostics, robotics, and figurative language understanding, critical task-relevant information often exists uniquely in one modality (e.g., specific medical sensor readings or facial expressions) or modalities share very little overlap. When applied to these non-redundant settings, conventional contrastive learning discards modality-unique signals, degrading predictive performance.
The article introduces and evaluates Factorized Contrastive Learning (FACTORCL), a framework designed to overcome multi-view redundancy. The main objective is to establish a method that explicitly separates and captures both shared and modality-unique task-relevant information while filtering out irrelevant noise, enabling effective self-supervised multimodal representation learning across diverse data scenarios.
The authors develop an information-theoretic approach that factorizes multimodal representations into four components: shared and unique features for each modality. To train these representations without relying on ground-truth task labels, FACTORCL introduces multimodal data augmentations that approximate task relevance and optimizes mutual information using paired lower-bound (to capture useful signal) and upper-bound (to remove task-irrelevant noise) estimators. The researchers evaluated FACTORCL through synthetic experiments with controllable shared-to-unique information ratios and across six large-scale real-world multimodal benchmarks encompassing over 84,000 data points in healthcare (MIMIC-III), affective computing (MOSI, MOSEI, UR-FUNNY, MUSTARD), and figurative language recognition (IRFL).
The evaluation yielded several key findings. First, FACTORCL achieved state-of-the-art results across all six real-world benchmarks, outperforming standard self-supervised and supervised contrastive baselines, particularly where unique information is vital—such as ICU mortality prediction and sarcasm detection. Second, probing experiments confirmed that FACTORCL captured roughly 75% more unique information (7 bits versus 4 bits) compared to standard methods while retaining strong shared information. Third, ablation analyses showed that explicit representation factorization improved performance by an average of 6.1% (up to 8.6% on healthcare data), and removing irrelevant noise via the upper-bound objective improved accuracy by an average of 13.6% (and up to 23.5% on sentiment tasks). Fourth, domain-aware multimodal augmentations outperformed independent unimodal augmentations (e.g., reaching 95.18% versus 92.77% on figurative language tasks). Finally, the framework maintained parity with standard contrastive baselines in traditional settings where information was predominantly redundant.
These findings demonstrate that discarding modality-unique information introduces a substantial performance gap in practical multimodal AI deployments. Incorporating factorized representations and noise removal mitigates this risk without significant computational overhead, as the upper-bound estimates reuse existing network components. Organizations deploying AI in sensor-rich, medical, or complex interaction environments can achieve higher accuracy and robustness by moving beyond pure cross-modal alignment.
Practitioners developing multimodal pipelines should adopt factorized contrastive objectives and design data augmentations that preserve cross-modal context rather than augmenting each data stream in isolation. Future work should explore automating the search for optimal multimodal augmentations, dynamically weighting shared versus unique loss terms based on task structure, and extending these principles to non-contrastive and masked pre-training frameworks.
While the empirical results across synthetic and public benchmarks provide high confidence in FACTORCL's effectiveness, the framework's self-supervised variant depends on the quality of domain-specific data augmentations. Readers should exercise caution when applying the method to novel modalities where designing augmentations that cleanly isolate task relevance is non-trivial.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). It formalizes contrastive multiview coding under the assumption that shared cross-view information is sufficient, providing the baseline multi-view redundancy framework that FACTORCL explicitly critiques and generalizes.
- Paper: What makes for good views for contrastive learning, Yonglong Tian et al. (2020). It establishes the InfoMin principle for view generation in contrastive learning, supplying core information-theoretic foundations on shared mutual information that FACTORCL extends to capture non-redundant, modality-unique information.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). It introduces the concept of factorizing multimodal inputs into shared invariant and modality-specific subspaces for affective benchmarks like MOSI, establishing representation factorization techniques preceding FACTORCL.
- Paper: Representation Learning with Contrastive Predictive Coding, Aäron van den Oord et al. (2018). It introduces the foundational contrastive mutual information lower-bound estimator (InfoNCE), which serves as a core mathematical tool adapted in FACTORCL's optimization objective.
- Paper: A Simple Framework for Contrastive Learning of Visual Representations, Ting Chen et al. (2020). It defines the foundational self-supervised contrastive learning framework using data augmentations and projection heads that FACTORCL builds upon and factorizes.
- Paper: Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere, Tongzhou Wang et al. (2020). It provides fundamental theoretical insights into the alignment and uniformity properties of contrastive loss on representation spaces that underpin FACTORCL's optimization mechanics.
- Paper: Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks, Nan Wu et al. (2022). It analyzes the failure modes of joint multimodal learning and modality under-utilization, motivating FACTORCL's goal of retaining modality-unique signals.
- Paper: FOCAL: Contrastive Learning for Multimodal Time-Series Sensing Signals in Factorized Orthogonal Latent Space, Shengzhong Liu et al. (2023). It applies the principle of factorizing representations into shared and private orthogonal latent spaces specifically to multimodal time-series and sensor perception tasks.
- Paper: Multi-View Causal Representation Learning with Partial Observability, Dingling Yao et al. (2024). It provides theoretical causal identifiability guarantees for disentangling shared and view-specific latent components under partial observability in multi-view contrastive learning.
- Paper: Decoupled Contrastive Multi-View Clustering with High-Order Random Walks, Yiding Lu et al. (2024). It extends decoupled contrastive representation learning by separating intra-view and inter-view objectives to preserve view-specific information in multi-view clustering.
- Paper: ReconBoost: Boosting Can Achieve Modality Reconcilement, Cong Hua et al. (2024). It introduces an alternating boosting scheme to reconcile dominant modality competition with single-modality feature retention in multimodal representation learning.
- Paper: UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All, Yuanhuiyi Lyu et al. (2024). It addresses multimodal alignment imbalance across numerous sensory modalities by leveraging language models to establish unified and balanced representation spaces.
