Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Paul Pu LiangZihao DengMartin Q. MaJames Y. ZouLouis-Philippe MorencyRuslan Salakhutdinov

article2023NeurIPS135 citations

Proposes FACTORCL, a multimodal contrastive learning framework that overcomes standard multi-view redundancy limitations by factorizing representations into task-relevant shared and unique information via conditional mutual information bounds and multimodal augmentations.

Listen

Modern multimodal machine learning relies heavily on contrastive self-supervised pre-training to learn useful representations from paired data, such as images matched with text or video paired with audio. However, standard methods rest on the assumption of multi-view redundancy: the premise that the only information needed for downstream tasks is the data shared between modalities. In real-world domains such as healthcare diagnostics, robotics, and figurative language understanding, critical task-relevant information often exists uniquely in one modality (e.g., specific medical sensor readings or facial expressions) or modalities share very little overlap. When applied to these non-redundant settings, conventional contrastive learning discards modality-unique signals, degrading predictive performance.

The article introduces and evaluates Factorized Contrastive Learning (FACTORCL), a framework designed to overcome multi-view redundancy. The main objective is to establish a method that explicitly separates and captures both shared and modality-unique task-relevant information while filtering out irrelevant noise, enabling effective self-supervised multimodal representation learning across diverse data scenarios.

The authors develop an information-theoretic approach that factorizes multimodal representations into four components: shared and unique features for each modality. To train these representations without relying on ground-truth task labels, FACTORCL introduces multimodal data augmentations that approximate task relevance and optimizes mutual information using paired lower-bound (to capture useful signal) and upper-bound (to remove task-irrelevant noise) estimators. The researchers evaluated FACTORCL through synthetic experiments with controllable shared-to-unique information ratios and across six large-scale real-world multimodal benchmarks encompassing over 84,000 data points in healthcare (MIMIC-III), affective computing (MOSI, MOSEI, UR-FUNNY, MUSTARD), and figurative language recognition (IRFL).

The evaluation yielded several key findings. First, FACTORCL achieved state-of-the-art results across all six real-world benchmarks, outperforming standard self-supervised and supervised contrastive baselines, particularly where unique information is vital—such as ICU mortality prediction and sarcasm detection. Second, probing experiments confirmed that FACTORCL captured roughly 75% more unique information (7 bits versus 4 bits) compared to standard methods while retaining strong shared information. Third, ablation analyses showed that explicit representation factorization improved performance by an average of 6.1% (up to 8.6% on healthcare data), and removing irrelevant noise via the upper-bound objective improved accuracy by an average of 13.6% (and up to 23.5% on sentiment tasks). Fourth, domain-aware multimodal augmentations outperformed independent unimodal augmentations (e.g., reaching 95.18% versus 92.77% on figurative language tasks). Finally, the framework maintained parity with standard contrastive baselines in traditional settings where information was predominantly redundant.

These findings demonstrate that discarding modality-unique information introduces a substantial performance gap in practical multimodal AI deployments. Incorporating factorized representations and noise removal mitigates this risk without significant computational overhead, as the upper-bound estimates reuse existing network components. Organizations deploying AI in sensor-rich, medical, or complex interaction environments can achieve higher accuracy and robustness by moving beyond pure cross-modal alignment.

Practitioners developing multimodal pipelines should adopt factorized contrastive objectives and design data augmentations that preserve cross-modal context rather than augmenting each data stream in isolation. Future work should explore automating the search for optimal multimodal augmentations, dynamically weighting shared versus unique loss terms based on task structure, and extending these principles to non-contrastive and masked pre-training frameworks.

While the empirical results across synthetic and public benchmarks provide high confidence in FACTORCL's effectiveness, the framework's self-supervised variant depends on the quality of domain-specific data augmentations. Readers should exercise caution when applying the method to novel modalities where designing augmentations that cleanly isolate task relevance is non-trivial.

Cover for Factorized Contrastive Learning: Going Beyond Multi-view Redundancy

Abstract

In a wide range of multimodal tasks, contrastive learning has become a particularly appealing approach since it can successfully learn representations from abundant unlabeled data with only pairing information (e.g., image-caption or video-audio pairs). Underpinning these approaches is the assumption of multi-view redundancy - that shared information between modalities is necessary and sufficient for downstream tasks. However, in many real-world settings, task-relevant information is also contained in modality-unique regions: information that is only present in one modality but still relevant to the task. How can we learn self-supervised multimodal representations to capture both shared and unique information relevant to downstream tasks? This paper proposes FACTORCL, a new multimodal representation learning method to go beyond multi-view redundancy. FACTORCL is built from three new contributions: (1) factorizing task-relevant information into shared and unique representations, (2) capturing task-relevant information via maximizing MI lower bounds and removing task-irrelevant information via minimizing MI upper bounds, and (3) multimodal data augmentations to approximate task relevance without labels. On large-scale real-world datasets, FACTORCL captures both shared and unique information and achieves state-of-the-art results on six benchmarks.

Table of Contents

  • 1 Introduction
  • 2 Analysis of Multi-view Contrastive Learning
  • 3 FACTORIZED CONTRASTIVE LEARNING
  • 3.1 Supervised FACTORCL with shared and unique information
  • 3.2 Self-supervised FACTORCL via multimodal augmentations
  • 3.3 Overall method and implementation
  • 4 Experiments
  • 4.1 Controlled experiments on synthetic datasets
  • 4.2 Self-supervised multimodal learning with low redundancy and high uniqueness
  • 5 Related Work
  • 6 Conclusion
  • Acknowledgements
  • References
  • Appendix
  • A Broader Impact
  • B Analysis of Multi-view Contrastive Learning
  • C FACTORIZED CONTRASTIVE LEARNING
  • C.1 Contrastive estimators
  • C.2 Unimodal and multimodal augmentations
  • C.2.1 Implementing conditional CL via kernel
  • C.3 Final estimators in FACTORCL
  • C.4 Extensions to masking and non-contrastive learning
  • D Experimental Details
  • D.1 Implementation details
  • D.2 Datasets
  • D.3 Additional analysis and results

Knowls

  1. Knowl 1 — Decomposition of Multimodal Task-Relevant Information into Shared and Unique Components

    definition

    In multimodal representation learning with two input modalities represented by random variables X1X_1 and X2X_2 and a downstream task label YY, the total multimodal task-relevant mutual information I(X1,X2;Y)I(X_1, X_2; Y) is decomposed into task-relevant shared information SS and modality-unique task-relevant information U1U_1 and U2U_2:

    I(X1,X2;Y)=S+U1+U2I(X_1, X_2; Y) = S + U_1 + U_2

    where:

    • S=I(X1;X2;Y)=I(X1;X2)−I(X1;X2∣Y)S = I(X_1; X_2; Y) = I(X_1; X_2) - I(X_1; X_2 \mid Y) represents the task-relevant shared mutual information between X1X_1 and X2X_2, with I(X1;X2∣Y)I(X_1; X_2 \mid Y) denoting the task-irrelevant shared information.
    • U1=I(X1;Y∣X2)U_1 = I(X_1; Y \mid X_2) denotes the task-relevant unique information present exclusively in modality X1X_1.
    • U2=I(X2;Y∣X1)U_2 = I(X_2; Y \mid X_1) denotes the task-relevant unique information present exclusively in modality X2X_2.

    Multi-view redundancy is defined by the existence of ϵ>0\epsilon > 0 such that I(X1;Y∣X2)≤ϵI(X_1; Y \mid X_2) \le \epsilon and I(X2;Y∣X1)≤ϵI(X_2; Y \mid X_1) \le \epsilon. Conversely, multi-view non-redundancy holds when I(X1;Y∣X2)>ϵI(X_1; Y \mid X_2) > \epsilon or I(X2;Y∣X1)>ϵI(X_2; Y \mid X_1) > \epsilon.

  2. Knowl 2 — Suboptimality and Uniqueness Gap of Standard Contrastive Learning

    theoretical result

    When multimodal data exhibits multi-view non-redundancy (task-relevant unique information I(X1;Y∣X2)>ϵI(X_1; Y \mid X_2) > \epsilon or I(X2;Y∣X1)>ϵI(X_2; Y \mid X_1) > \epsilon), optimal representations Z1=arg⁡max⁡Z1=fθ(X1)I(Z1;X2)Z_1 = \arg\max_{Z_1=f_\theta(X_1)} I(Z_1; X_2) and Z2=arg⁡max⁡Z2=fθ(X2)I(X1;Z2)Z_2 = \arg\max_{Z_2=f_\theta(X_2)} I(X_1; Z_2) that satisfy the minimal sufficient representation assumption (I(Z1;Y∣X2)=I(Z2;Y∣X1)=0I(Z_1; Y \mid X_2) = I(Z_2; Y \mid X_1) = 0) fail to capture the full task-relevant information content:

    I(Z1,Z2;Y)=I(X1,X2;Y)−I(X1;Y∣X2)−I(X2;Y∣X1)=I(X1;X2)−I(X1;X2∣Y)<I(X1,X2;Y)I(Z_1, Z_2; Y) = I(X_1, X_2; Y) - I(X_1; Y \mid X_2) - I(X_2; Y \mid X_1) = I(X_1; X_2) - I(X_1; X_2 \mid Y) < I(X_1, X_2; Y)

    The difference I(X1,X2;Y)−I(Z1,Z2;Y)=I(X1;Y∣X2)+I(X2;Y∣X1)I(X_1, X_2; Y) - I(Z_1, Z_2; Y) = I(X_1; Y \mid X_2) + I(X_2; Y \mid X_1) is termed the uniqueness gap, which quantifies the loss in task-relevant information between input data and contrastively encoded representations.

    Furthermore, the Bayes error rate Pe(Z1,Z2):=1−Ep(z1,z2)[max⁡y∈YP(Y^=y∣z1,z2)]P_e(Z_1, Z_2) := 1 - \mathbb{E}_{p(z_1, z_2)}[\max_{y \in \mathcal{Y}} P(\hat{Y} = y \mid z_1, z_2)] of downstream linear prediction using the learned representations satisfies:

    Pe(Z1,Z2)≤1−exp⁡(I(X1;X2;Y)−H(Y))=1−exp⁡(I(X1,X2;Y)−I(X1;Y∣X2)−I(X2;Y∣X1)−H(Y))P_e(Z_1, Z_2) \le 1 - \exp\left( I(X_1; X_2; Y) - H(Y) \right) = 1 - \exp\left( I(X_1, X_2; Y) - I(X_1; Y \mid X_2) - I(X_2; Y \mid X_1) - H(Y) \right)

    where H(Y)H(Y) is the task entropy. As unique task-relevant information increases or shared information decreases, the Bayes error bound degrades.

  3. Knowl 3 — Factorized Representations and Contrastive Estimation of Shared and Unique Information

    model/method

    Factorized Contrastive Learning (FACTORCL) learns four factorized representations to isolate shared and unique task-relevant features across modalities X1X_1 and X2X_2:

    • ZS1=arg⁡max⁡Z1=fθ(X1)I(Z1;X2;Y)Z_{S_1} = \arg\max_{Z_1=f_\theta(X_1)} I(Z_1; X_2; Y) (information in X1X_1 shared with X2X_2 relevant to YY)
    • ZS2=arg⁡max⁡Z2=fθ(X2)I(Z2;X1;Y)Z_{S_2} = \arg\max_{Z_2=f_\theta(X_2)} I(Z_2; X_1; Y) (information in X2X_2 shared with X1X_1 relevant to YY)
    • ZU1=arg⁡max⁡Z1=fθ(X1)I(Z1;Y∣X2)Z_{U_1} = \arg\max_{Z_1=f_\theta(X_1)} I(Z_1; Y \mid X_2) (information in X1X_1 unique from X2X_2 relevant to YY)
    • ZU2=arg⁡max⁡Z2=fθ(X2)I(Z2;Y∣X1)Z_{U_2} = \arg\max_{Z_2=f_\theta(X_2)} I(Z_2; Y \mid X_1) (information in X2X_2 unique from X1X_1 relevant to YY)

    To estimate intractable mutual information terms, lower bounds are estimated via InfoNCE and upper bounds via the plug-in Contrastive Log-ratio Upper Bound (NCE-CLUB):

    INCE(X1;X2)=Ex1,x2+∼p(x1,x2)x2−∼p(x2)[log⁡exp⁡f(x1,x2+)∑kexp⁡f(x1,x2−)]I_{\text{NCE}}(X_1; X_2) = \mathbb{E}_{\substack{x_1, x_2^+ \sim p(x_1, x_2) \\ x_2^- \sim p(x_2)}} \left[ \log \frac{\exp f(x_1, x_2^+)}{\sum_k \exp f(x_1, x_2^-)} \right]

    INCE-CLUB(X1;X2)=Ex1,x2+∼p(x1,x2)[f∗(x1,x2+)]−Ex1∼p(x1)x2−∼p(x2)[f∗(x1,x2−)]I_{\text{NCE-CLUB}}(X_1; X_2) = \mathbb{E}_{x_1, x_2^+ \sim p(x_1, x_2)} [f^*(x_1, x_2^+)] - \mathbb{E}_{\substack{x_1 \sim p(x_1) \\ x_2^- \sim p(x_2)}} [f^*(x_1, x_2^-)]

    where f∗f^* is the optimal critic trained by INCE(X1;X2)I_{\text{NCE}}(X_1; X_2), satisfying INCE(X1;X2)≤I(X1;X2)≤INCE-CLUB(X1;X2)I_{\text{NCE}}(X_1; X_2) \le I(X_1; X_2) \le I_{\text{NCE-CLUB}}(X_1; X_2).

    Task-relevant shared information SS and unique information UiU_i (i∈{1,2}i \in \{1, 2\}, with complementary modality X−iX_{-i}) are bounded and optimized via:

    S=I(X1;X2;Y)≥INCE(X1;X2)−INCE-CLUB(X1;X2∣Y)S = I(X_1; X_2; Y) \ge I_{\text{NCE}}(X_1; X_2) - I_{\text{NCE-CLUB}}(X_1; X_2 \mid Y)

    Ui=I(Xi;Y∣X−i)≥INCE(Xi;Y)−INCE-CLUB(X1;X2)+INCE(X1;X2∣Y)U_i = I(X_i; Y \mid X_{-i}) \ge I_{\text{NCE}}(X_i; Y) - I_{\text{NCE-CLUB}}(X_1; X_2) + I_{\text{NCE}}(X_1; X_2 \mid Y)

  4. Knowl 4 — Self-Supervised FACTORCL via Optimal Multimodal Augmentations

    model/method

    In self-supervised settings without access to labels YY, FACTORCL approximates task relevance by substituting YY with semantic multimodal data augmentations (X1′,X2′)(X_1', X_2').

    Optimal unimodal augmentations satisfy I(Xi;Xi′)=I(Xi;Y)I(X_i; X_i') = I(X_i; Y). To satisfy the joint condition I(X1,X2;X1′,X2′)=I(X1,X2;Y)I(X_1, X_2; X_1', X_2') = I(X_1, X_2; Y), FACTORCL creates augmentations sequentially:

    1. Unimodal Augmentation: X1′X_1' such that I(X1;X1′)=I(X1;Y)I(X_1; X_1') = I(X_1; Y).
    2. Unique Augmentation: X2′X_2' conditioned on X1X_1 such that I(X2;X2′∣X1)=I(X2;Y∣X1)I(X_2; X_2' \mid X_1) = I(X_2; Y \mid X_1).

    Unique augmentation ensures that perturbations applied to X2X_2 preserve any information shared with X1X_1 (e.g., in image-captioning, image cropping is avoided to prevent removing objects mentioned in the caption, using flipping or color jittering instead).

    Substituting (X1′,X2′)(X_1', X_2') into the MI bounds gives the self-supervised objectives:

    S=I(X1;X2;Y)≥INCE(X1;X2)−INCE-CLUB(X1;X2∣X1′,X2′)S = I(X_1; X_2; Y) \ge I_{\text{NCE}}(X_1; X_2) - I_{\text{NCE-CLUB}}(X_1; X_2 \mid X_1', X_2')

    Ui=I(Xi;Y∣X−i)≥INCE(Xi;Xi′)−INCE-CLUB(X1;X2)+INCE(X1;X2∣X1′,X2′)U_i = I(X_i; Y \mid X_{-i}) \ge I_{\text{NCE}}(X_i; X_i') - I_{\text{NCE-CLUB}}(X_1; X_2) + I_{\text{NCE}}(X_1; X_2 \mid X_1', X_2')

    When representations ZS1,ZS2,ZU1,ZU2Z_{S_1}, Z_{S_2}, Z_{U_1}, Z_{U_2} maximize these objectives and bounds are tight, FACTORCL achieves complete task-relevant coverage: I(X1,X2;Y)=I(ZS1;ZS2;Y)+I(ZU1;Y∣ZS2)+I(ZU2;Y∣ZS1)I(X_1, X_2; Y) = I(Z_{S_1}; Z_{S_2}; Y) + I(Z_{U_1}; Y \mid Z_{S_2}) + I(Z_{U_2}; Y \mid Z_{S_1}).

  5. Knowl 5 — FACTORCL Training Algorithm with Sampling for Inner Critic Optimization

    algorithm

    FACTORCL optimizes encoder/projection parameters θ\theta and critic parameters ϕ\phi using a two-loop training strategy where the optimal critic f∗f^* needed for NCE-CLUB upper bounds is updated in an inner step.

    Input: Multimodal dataset {X1,X2}\{X_1, X_2\}, inner critic update steps kk (default k=1k = 1)
    Output: Trained multimodal feature extractors fθf_\theta
    Initialize encoder/projection parameters θ\theta and critic parameters ϕ\phi
    while not converged do
        Sample mini-batch {x1,x2}∼{X1,X2}\{x_1, x_2\} \sim \{X_1, X_2\}
        Generate unimodal augmentation x1′←Augment(x1)x_1' \leftarrow \text{Augment}(x_1)
        Generate unique augmentation x2′←Unique-Augment(x2∣x1)x_2' \leftarrow \text{Unique-Augment}(x_2 \mid x_1)
        Compute shared and unique estimators:
            S←INCE(x1;x2)−INCE-CLUB(x1;x2∣x1′,x2′;ϕ)S \leftarrow I_{\text{NCE}}(x_1; x_2) - I_{\text{NCE-CLUB}}(x_1; x_2 \mid x_1', x_2'; \phi)
            U1←INCE(x1;x1′)−INCE-CLUB(x1;x2;ϕ)+INCE(x1;x2∣x1′,x2′)U_1 \leftarrow I_{\text{NCE}}(x_1; x_1') - I_{\text{NCE-CLUB}}(x_1; x_2; \phi) + I_{\text{NCE}}(x_1; x_2 \mid x_1', x_2')
            U2←INCE(x2;x2′)−INCE-CLUB(x1;x2;ϕ)+INCE(x1;x2∣x1′,x2′)U_2 \leftarrow I_{\text{NCE}}(x_2; x_2') - I_{\text{NCE-CLUB}}(x_1; x_2; \phi) + I_{\text{NCE}}(x_1; x_2 \mid x_1', x_2')
        Compute overall loss LFACTORCL=−(S+U1+U2)\mathcal{L}_{\text{FACTORCL}} = -(S + U_1 + U_2)
        Update parameters θ\theta by descending along ∇θLFACTORCL\nabla_\theta \mathcal{L}_{\text{FACTORCL}}
        
        for i=1i = 1 to kk do
            Sample mini-batch {x1′,x2′}∼{X1,X2}\{x_1', x_2'\} \sim \{X_1, X_2\}
            Compute surrogate critic loss LNCE=−[INCE(x1;x2)+INCE(x1;x2∣x1′,x2′)]\mathcal{L}_{\text{NCE}} = -[I_{\text{NCE}}(x_1; x_2) + I_{\text{NCE}}(x_1; x_2 \mid x_1', x_2')]
            Update critic parameters ϕ\phi by descending along ∇ϕLNCE\nabla_\phi \mathcal{L}_{\text{NCE}}
        end for
    end while
    return fθf_\theta

    Conditioning on augmentations is implemented by concatenating representations into the critic: f(x1,x2,x1′,x2′)=g([z1,z1′])Th([z2,z2′])f(x_1, x_2, x_1', x_2') = g([z_1, z_1'])^T h([z_2, z_2']).

  6. Knowl 6 — Information Decomposition Probing on Synthetic Latent Variables

    empirical result

    On synthetic datasets generated from latent variables w1,w2,ws∼N(0d,Σd2)w_1, w_2, w_s \sim \mathcal{N}(0_d, \Sigma_d^2) (d=50d=50), where wsw_s is shared across X1,X2X_1, X_2 and w1,w2w_1, w_2 are unique to X1X_1 and X2X_2, representation probing evaluates mutual information captured between representations ZZ and ground-truth latent components:

    Model SimCLR Cross+self SupCon FACTORCL
    Representations Z1Z2Z_1 \quad Z_2 Z1Z2Z_1 \quad Z_2 Z1Z2Z_1 \quad Z_2 ZU1Z_{U_1} ZU2Z_{U_2} ZS1Z_{S_1} ZS2Z_{S_2}
    I(Z;w1)I(Z; w_1) 4.450.164.45 \quad 0.16 4.390.144.39 \quad 0.14 5.170.195.17 \quad 0.19 7.83 0.03 6.25 0.04
    I(Z;w2)I(Z; w_2) 0.173.920.17 \quad 3.92 0.134.260.13 \quad 4.26 0.235.170.23 \quad 5.17 0.06 7.17 0.05 5.79
    I(Z;ws)I(Z; w_s) 12.6112.0612.61 \quad 12.06 11.3011.4711.30 \quad 11.47 7.487.177.48 \quad 7.17 9.47 9.89 10.13 9.40

    Standard contrastive learning (SimCLR) and Cross+self focus heavily on shared information wsw_s (∼11–12\sim 11\text{--}12 bits) but capture little unique information (∼4\sim 4 bits). FACTORCL successfully factorizes representations: ZU1Z_{U_1} and ZU2Z_{U_2} capture significantly higher unique information (7.837.83 and 7.177.17 bits, respectively) while ZS1Z_{S_1} and ZS2Z_{S_2} capture shared information wsw_s (∼10\sim 10 bits).

  7. Knowl 7 — Multimodal Fusion Classification Performance Across MultiBench Datasets

    empirical result

    Self-supervised (SSL) and supervised (SUP) FACTORCL models were evaluated against contrastive baselines across five MultiBench datasets spanning healthcare (MIMIC-III), sentiment (CMU-MOSEI, CMU-MOSI), humor (UR-FUNNY), and sarcasm (MUSTARD). Downstream linear evaluation accuracy (mean ±\pm std across 5 runs) is reported:

    Model MIMIC MOSEI MOSI UR-FUNNY MUSTARD
    SimCLR 66.7±0.0%66.7 \pm 0.0\% 71.9±0.3%71.9 \pm 0.3\% 47.8±1.8%47.8 \pm 1.8\% 50.1±1.9%50.1 \pm 1.9\% 53.5±2.9%53.5 \pm 2.9\%
    Cross+Self 65.2±0.0%65.2 \pm 0.0\% 71.1±0.2%71.1 \pm 0.2\% 48.6±1.2%48.6 \pm 1.2\% 56.5±0.7%56.5 \pm 0.7\% 53.9±4.5%53.9 \pm 4.5\%
    Cross+Self+Fact 65.5±0.0%65.5 \pm 0.0\% 71.9±0.2%71.9 \pm 0.2\% 49.0±1.1%49.0 \pm 1.1\% 59.9±0.9%59.9 \pm 0.9\% 53.9±4.0%53.9 \pm 4.0\%
    OurCL-SSL (no fact.) 65.2±0.0%65.2 \pm 0.0\% 71.2±0.2%71.2 \pm 0.2\% 49.0±0.8%49.0 \pm 0.8\% 58.8±1.3%58.8 \pm 1.3\% 54.0±2.5%54.0 \pm 2.5\%
    FACTORCL-SSL 67.3±0.0%\mathbf{67.3 \pm 0.0\%} 74.5±0.1%\mathbf{74.5 \pm 0.1\%} 51.2±1.6%\mathbf{51.2 \pm 1.6\%} 60.5±0.8%\mathbf{60.5 \pm 0.8\%} 55.8±0.9%\mathbf{55.8 \pm 0.9\%}
    SupCon 67.4±0.0%67.4 \pm 0.0\% 71.0±0.1%71.0 \pm 0.1\% 47.2±1.2%47.2 \pm 1.2\% 50.1±2.0%50.1 \pm 2.0\% 52.7±2.2%52.7 \pm 2.2\%
    OurCL-SUP (no fact.) 68.2±0.0%68.2 \pm 0.0\% 71.1±0.2%71.1 \pm 0.2\% 65.3±0.8%65.3 \pm 0.8\% 58.3±1.1%58.3 \pm 1.1\% 65.1±1.8%65.1 \pm 1.8\%
    FACTORCL-SUP 76.8±0.0%\mathbf{76.8 \pm 0.0\%} 77.8±0.3%\mathbf{77.8 \pm 0.3\%} 69.1±0.6%\mathbf{69.1 \pm 0.6\%} 63.5±0.8%\mathbf{63.5 \pm 0.8\%} 69.9±1.9%\mathbf{69.9 \pm 1.9\%}

    FACTORCL achieves state-of-the-art results across all benchmarks in both self-supervised and supervised configurations. Performance gains are most pronounced on datasets with high unique information: MUSTARD (+16.4% in SUP over SimCLR) and MIMIC-III (+10.1% in SUP over SimCLR).

  8. Knowl 8 — Figurative Language Image-Text Classification on IRFL

    empirical result

    Continued pre-training using FACTORCL objectives on top of a pre-trained CLIP (CLIP-ViT-B/32) backbone was evaluated on the IRFL dataset (6,697 matching image-caption pairs across idioms, similes, and metaphors where captions are non-literal):

    Method IRFL Accuracy
    Zero-shot CLIP 89.2±0.0%89.2 \pm 0.0\%
    Fine-tuned CLIP 96.4±0.0%96.4 \pm 0.0\%
    SimCLR 91.6±0.0%91.6 \pm 0.0\%
    Cross+Self 91.1±1.2%91.1 \pm 1.2\%
    FACTORCL-IndAug (independent augmentations) 91.6±1.3%91.6 \pm 1.3\%
    FACTORCL-SSL (unique multimodal augmentations) 93.8±1.4%\mathbf{93.8 \pm 1.4\%}
    SupCon 87.7±4.7%87.7 \pm 4.7\%
    FACTORCL-SUP 98.3±1.2%\mathbf{98.3 \pm 1.2\%}

    FACTORCL-SSL outperforms standard continued pre-training baselines (SimCLR at 91.6%91.6\% and Cross+Self at 91.1%91.1\%), and unique multimodal augmentations provide a +2.2%+2.2\% gain over independent augmentations. Supervised FACTORCL-SUP achieves 98.3%98.3\%, outperforming fine-tuned CLIP (96.4%96.4\%) and SupCon (87.7%87.7\%).

  9. Knowl 9 — Ablations on Representation Factorization, Upper-Bound Removal, and Sub-Representations

    empirical result

    Ablation experiments across benchmark datasets quantify the impact of key FACTORCL components:

    1. Representation Factorization: Removing factorization (collapsing shared and unique representations into a single feature vector per modality, as in OurCL-SSL and OurCL-SUP) causes an average performance degradation of 6.1%6.1\%, with drops reaching 8.6%8.6\% on MIMIC-III.
    2. Task-Irrelevant Information Removal via Upper Bound: Removing the INCE-CLUBI_{\text{NCE-CLUB}} upper-bound penalty (which strips task-irrelevant shared information) results in an average accuracy decline of 13.6%13.6\% across benchmarks, with drops up to 23.5%23.5\% on CMU-MOSI.
    3. Sub-representation Isolation: Evaluating linear classifiers on individual factorized subspaces shows that predicting with both shared and unique representations is required for optimal performance:
      • On CMU-MOSEI: Full model (77.8%77.8\%) vs ZSZ_S only (77.2%77.2\%), ZU1Z_{U_1} only (77.1%77.1\%), ZU2Z_{U_2} only (71.0%71.0\%).
      • On MUSTARD: Full model (69.9%69.9\%) vs ZU1Z_{U_1} (video/audio sarcasm features) (59.4%59.4\%), ZSZ_S (57.3%57.3\%), ZU2Z_{U_2} (53.6%53.6\%).
  10. Knowl 10 — Limitations of FACTORCL

    limitation

    FACTORCL has three primary limitations:

    1. Manual Design of Multimodal Augmentations: The self-supervised formulation requires domain-specific data augmentations that approximately satisfy optimal multimodal augmentation conditions (I(X1,X2;X1′,X2′)=I(X1,X2;Y)I(X_1, X_2; X_1', X_2') = I(X_1, X_2; Y)) without automated discovery.
    2. High-Dimensional MI Bound Estimation: In high-dimensional feature spaces and complex continuous distributions, estimating conditional mutual information bounds remains challenging and may suffer from approximation error.
    3. Equal Weighting of Information Objectives: FACTORCL weights shared and unique terms uniformly in the joint objective L=−(S+U1+U2)\mathcal{L} = -(S + U_1 + U_2), rather than adaptively reweighting the components based on the task-specific balance of shared versus unique information.

Coverage note — No substantial contributed material was omitted. High shared information settings (CIFAR-10/MNIST in Appendix D.3) and kernel-based CMI formulations in Appendix C.2 were excluded as secondary extensions to prioritize core theory, algorithms, and primary benchmark evaluations.

References

  1. 1.Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. Advances in Neural Information Processing Systems, 34:24206–24221, 2021.
  2. 2.Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelovic, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. Self-supervised multimodal versatile networks. Advances in Neural Information Processing Systems, 33:25–37, 2020.
  3. 3.Relja Arandjelovic and Andrew Zisserman. Look, listen and learn. In Proceedings of the IEEE international conference on computer vision, pages 609–617, 2017.
  4. 4.Philip Bachman, R Devon Hjelm, and William Buchwalter. Learning representations by maximizing mutual information across views. Advances in neural information processing systems, 32, 2019.
  5. 5.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems, 33: 12449–12460, 2020.
  6. 6.Adrien Bardes, Jean Ponce, and Yann LeCun. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. In International Conference on Learning Representations, 2021.
  7. 7.Anthony J Bell. The co-information lattice. In Proceedings of the fifth international workshop on independent component analysis and blind signal separation: ICA, volume 2003, 2003.
  8. 8.Yoshua Bengio, Aaron Courville, and Pascal Vincent. Representation learning: A review and new perspectives. TPAMI, 35(8), August 2013.
  9. 9.Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. Man is to computer programmer as woman is to homemaker? debiasing word embeddings. Advances in neural information processing systems, 29, 2016.
  10. 10.Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott. Multimodal pretraining unmasked: A meta-analysis and a unified framework of vision-and-language berts. Transactions of the Association for Computational Linguistics, 9:978–994, 2021.
  11. 11.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. Advances in neural information processing systems, 33:9912–9924, 2020.
  12. 12.Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria. Towards multimodal sarcasm detection (an obviously perfect paper). arXiv preprint arXiv:1906.01815, 2019.
  13. 13.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020.
  14. 14.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15750–15758, 2021.
  15. 15.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9640–9649, 2021.
  16. 16.Pengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu, Zhe Gan, and Lawrence Carin. Club: A contrastive log-ratio upper bound of mutual information. In International conference on machine learning, pages 1779–1788. PMLR, 2020.
  17. 17.Jianfeng Chi, William Shand, Yaodong Yu, Kai-Wei Chang, Han Zhao, and Yuan Tian. Conditional supervised contrastive learning for fair text classification. arXiv preprint arXiv:2205.11485, 2022.
  18. 18.Thomas M Cover and Joy A Thomas. Information theory and statistics. Elements of information theory, 1 (1):279–335, 1991.
  19. 19.Li Deng. The mnist database of handwritten digit images for machine learning research. IEEE Signal Processing Magazine, 29(6):141–142, 2012.
  20. 20.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  21. 21.Benjamin Eysenbach, Tianjun Zhang, Sergey Levine, and Russ R Salakhutdinov. Contrastive learning as goal-conditioned reinforcement learning. Advances in Neural Information Processing Systems, 35: 35603–35620, 2022.
  22. 22.Meir Feder and Neri Merhav. Relations between entropy and error probability. IEEE Transactions on Information theory, 40(1):259–266, 1994.
  23. 23.Marco Federici, Anjan Dutta, Patrick Forré, Nate Kushman, and Zeynep Akata. Learning robust representations via multi-view information bottleneck. arXiv preprint arXiv:2002.07017, 2020.
  24. 24.Tianyu Gao, Xingcheng Yao, and Danqi Chen. Simcse: Simple contrastive learning of sentence embeddings. arXiv preprint arXiv:2104.08821, 2021.
  25. 25.Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neural information processing systems, 33:21271–21284, 2020.
  26. 26.Qing Guo, Junya Chen, Dong Wang, Yuewei Yang, Xinwei Deng, Jing Huang, Larry Carin, Fan Li, and Chenyang Tao. Tight mutual information estimation with contrastive fenchel-legendre optimization. Advances in Neural Information Processing Systems, 35:28319–28334, 2022.
  27. 27.Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, and Mohammed Ehsan Hoque. Ur-funny: A multimodal language dataset for understanding humor. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2046–2056, 2019.
  28. 28.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9729–9738, 2020.
  29. 29.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022.
  30. 30.Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-vae: Learning basic visual concepts with a constrained variational framework. 2016.
  31. 31.R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio. Learning deep representations by mutual information estimation and maximization. In International Conference on Learning Representations, 2018.
  32. 32.Wei-Ning Hsu and James Glass. Disentangling by partitioning: A representation learning framework for multimodal sensory data. arXiv preprint arXiv:1805.11264, 2018.
  33. 33.Po-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, and Alexander Hauptmann. Multilingual multimodal pre-training for zero-shot cross-lingual transfer of vision-language models. arXiv preprint arXiv:2103.08849, 2021.
  34. 34.Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, and Longbo Huang. What makes multi-modal learning better than single (provably). Advances in Neural Information Processing Systems, 34:10944–10956, 2021.
  35. 35.HyeongJoo Hwang, Geon-Hyeong Kim, Seunghoon Hong, and Kee-Eung Kim. Variational interaction information maximization for cross-domain disentanglement. Advances in Neural Information Processing Systems, 33:22479–22491, 2020.
  36. 36.Aashi Jain, Mandy Guo, Krishna Srinivasan, Ting Chen, Sneha Kudugunta, Chao Jia, Yinfei Yang, and Jason Baldridge. Mural: multimodal, multitask retrieval across languages. arXiv preprint arXiv:2109.05125, 2021.
  37. 37.Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International Conference on Machine Learning, pages 4904–4916. PMLR, 2021.
  38. 38.Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9, 2016.
  39. 39.Jonathan Kahana and Yedid Hoshen. A contrastive objective for learning disentangled representations. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXVI, pages 579–595. Springer, 2022.
  40. 40.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186, 2019.
  41. 41.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. Advances in neural information processing systems, 33:18661–18673, 2020.
  42. 42.Byoungjip Kim, Sungik Choi, Dasol Hwang, Moontae Lee, and Honglak Lee. Transferring pre-trained multimodal representations with cross-modal similarity matching. Advances in Neural Information Processing Systems, 35:30826–30839, 2022.
  43. 43.Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. Cifar-10 (canadian institute for advanced research). URL http://www.cs.toronto.edu/~kriz/cifar.html.
  44. 44.Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, and Yale Song. Parameter efficient multimodal transformers for video representation learning. arXiv preprint arXiv:2012.04124, 2020.
  45. 45.Paul Pu Liang, Yiwei Lyu, Xiang Fan, Zetian Wu, Yun Cheng, Jason Wu, Leslie Chen, Peter Wu, Michelle A Lee, Yuke Zhu, Ruslan Salakhutdinov, and Louis-Philippe Morency. Multibench: Multiscale benchmarks for multimodal representation learning. NeurIPS Datasets and Benchmarks Track, 2021.
  46. 46.Paul Pu Liang, Chiyu Wu, Louis-Philippe Morency, and Ruslan Salakhutdinov. Towards understanding and mitigating social biases in language models. In International Conference on Machine Learning, pages 6565–6576. PMLR, 2021.
  47. 47.Paul Pu Liang, Yiwei Lyu, Xiang Fan, Shengtong Mo, Dani Yogatama, et al. Highmmt: Towards modality and task generalization for high-modality representation learning. arXiv preprint arXiv:2203.01311, 2022.
  48. 48.Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. Foundations and recent trends in multimodal machine learning: Principles, challenges, and open questions. arXiv preprint arXiv:2209.03430, 2022.
  49. 49.Paul Pu Liang, Yun Cheng, Xiang Fan, Chun Kai Ling, Suzanne Nie, Richard Chen, Zihao Deng, Faisal Mahmood, Ruslan Salakhutdinov, and Louis-Philippe Morency. Quantifying & modeling feature interactions: An information decomposition framework. arXiv preprint arXiv:2302.12247, 2023.
  50. 50.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, pages 13–23, 2019.
  51. 51.Martin Q Ma, Yao-Hung Hubert Tsai, Paul Pu Liang, Han Zhao, Kun Zhang, Ruslan Salakhutdinov, and Louis-Philippe Morency. Conditional contrastive learning for improving fairness in self-supervised learning. arXiv preprint arXiv:2106.02866, 2021.
  52. 52.Emily E Marsh and Marilyn Domas White. A taxonomy of relationships between images and text. Journal of documentation, 2003.
  53. 53.William McGill. Multivariate information transmission. Transactions of the IRE Professional Group on Information Theory, 4(4):93–111, 1954.
  54. 54.Yu Meng, Chenyan Xiong, Payal Bajaj, Paul Bennett, Jiawei Han, Xia Song, et al. Coco-lm: Correcting and contrasting text sequences for language model pretraining. Advances in Neural Information Processing Systems, 34:23102–23114, 2021.
  55. 55.Sudipto Mukherjee, Himanshu Asnani, and Sreeram Kannan. Ccmi: Classifier based conditional mutual information estimation. In Uncertainty in artificial intelligence, pages 1083–1093. PMLR, 2020.
  56. 56.Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al. Text and code embeddings by contrastive pre-training. arXiv preprint arXiv:2201.10005, 2022.
  57. 57.XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan. Estimating divergence functionals and the likelihood ratio by convex risk minimization. IEEE Transactions on Information Theory, 56(11): 5847–5861, 2010.
  58. 58.Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  59. 59.Sherjil Ozair, Corey Lynch, Yoshua Bengio, Aaron Van den Oord, Sergey Levine, and Pierre Sermanet. Wasserstein dependency measure for representation learning. Advances in Neural Information Processing Systems, 32, 2019.
  60. 60.Ben Poole, Sherjil Ozair, Aaron Van Den Oord, Alex Alemi, and George Tucker. On variational bounds of mutual information. In International Conference on Machine Learning, pages 5171–5180. PMLR, 2019.
  61. 61.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021.
  62. 62.Anil Rahate, Rahee Walambe, Sheela Ramanna, and Ketan Kotecha. Multimodal co-learning: challenges, applications with datasets, recent advances and future directions. Information Fusion, 81:203–239, 2022.
  63. 63.Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli. wav2vec: Unsupervised pre-training for speech recognition. arXiv preprint arXiv:1904.05862, 2019.
  64. 64.Bin Shan, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. Ernie-vil 2.0: Multi-view contrastive learning for image-text pre-training, 2022.
  65. 65.Claude Elwood Shannon. A mathematical theory of communication. The Bell system technical journal, 27 (3):379–423, 1948.
  66. 66.Yuge Shi, Brooks Paige, Philip Torr, et al. Variational mixture-of-experts autoencoders for multi-modal deep generative models. Advances in Neural Information Processing Systems, 32, 2019.
  67. 67.Ravid Shwartz-Ziv and Yann LeCun. To compress or not to compress–self-supervised learning and information theory: A review. arXiv preprint arXiv:2304.09355, 2023.
  68. 68.Jiaming Song and Stefano Ermon. Understanding the limitations of variational mutual information estimators. CoRR, abs/1910.06222, 2019. URL http://arxiv.org/abs/1910.06222.
  69. 69.Alessandro Sordoni, Nouha Dziri, Hannes Schulz, Geoff Gordon, Philip Bachman, and Remi Tachet Des Combes. Decomposed mutual information estimation for contrastive representation learning. In International Conference on Machine Learning, pages 9859–9869. PMLR, 2021.
  70. 70.Karthik Sridharan and Sham M Kakade. An information theoretic framework for multi-view learning. In Conference on Learning Theory, 2008.
  71. 71.Yonglong Tian, Dilip Krishnan, and Phillip Isola. Contrastive multiview coding. ECCV, 2020.
  72. 72.Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. What makes for good views for contrastive learning? Advances in Neural Information Processing Systems, 33:6827–6839, 2020.
  73. 73.Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu. Contrastive learning, multi-view redundancy, and linear models. In Algorithmic Learning Theory, pages 1179–1206. PMLR, 2021.
  74. 74.Yao-Hung Hubert Tsai, Tianqin Li, Weixin Liu, Peiyuan Liao, Ruslan Salakhutdinov, and Louis-Philippe Morency. Learning weakly-supervised contrastive representations. In International Conference on Learning Representations.
  75. 75.Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. Learning factorized multimodal representations. ICLR, 2019.
  76. 76.Yao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, and Louis-Philippe Morency. Self-supervised learning from a multi-view perspective. In International Conference on Learning Representations, 2020.
  77. 77.Yao-Hung Hubert Tsai, Han Zhao, Makoto Yamada, Louis-Philippe Morency, and Russ R Salakhutdinov. Neural methods for point-wise dependency estimation. Advances in Neural Information Processing Systems, 33:62–72, 2020.
  78. 78.Yao-Hung Hubert Tsai, Tianqin Li, Martin Q Ma, Han Zhao, Kun Zhang, Louis-Philippe Morency, and Ruslan Salakhutdinov. Conditional contrastive learning with kernel. arXiv preprint arXiv:2202.05458, 2022.
  79. 79.Michael Tschannen, Josip Djolonga, Paul K Rubenstein, Sylvain Gelly, and Mario Lucic. On mutual information maximization for representation learning. In International Conference on Learning Representations, 2019.
  80. 80.Jorge R Vergara and Pablo A Estévez. A review of feature selection methods based on mutual information. Neural computing and applications, 24:175–186, 2014.
  81. 81.Haoqing Wang, Xun Guo, Zhi-Hong Deng, and Yan Lu. Rethinking minimal sufficient representation in contrastive learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16041–16050, 2022.
  82. 82.Paul L Williams and Randall D Beer. Nonnegative decomposition of multivariate information. arXiv preprint arXiv:1004.2515, 2010.
  83. 83.Mike Wu and Noah Goodman. Multimodal generative models for scalable weakly-supervised learning. Advances in Neural Information Processing Systems, 31, 2018.
  84. 84.Mike Wu, Chengxu Zhuang, Milan Mosse, Daniel Yamins, and Noah Goodman. On mutual information in contrastive learning for visual representations. arXiv preprint arXiv:2005.13149, 2020.
  85. 85.Jianwei Yang, Chunyuan Li, Pengchuan Zhang, Bin Xiao, Ce Liu, Lu Yuan, and Jianfeng Gao. Unified contrastive learning in image-text-label space, 2022.
  86. 86.Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang. Vision-language pre-training with triple contrastive learning, 2022.
  87. 87.Zesheng Ye and Lina Yao. Contrastive conditional neural processes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9687–9696, 2022.
  88. 88.Ron Yosef, Yonatan Bitton, and Dafna Shahaf. Irfl: Image recognition of figurative language. arXiv preprint arXiv:2303.15445, 2023.
  89. 89.Xin Yuan, Zhe Lin, Jason Kuen, Jianming Zhang, Yilin Wang, Michael Maire, Ajinkya Kale, and Baldo Faieta. Multimodal contrastive training for visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6995–7004, 2021.
  90. 90.Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6023–6032, 2019.
  91. 91.Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos. arXiv preprint arXiv:1606.06259, 2016.
  92. 92.Amir Zadeh, Paul Pu Liang, and Louis-Philippe Morency. Foundations of multimodal co-learning. Information Fusion, 64:188–193, 2020.
  93. 93.AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236–2246, 2018.
  94. 94.Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning, pages 12310–12320. PMLR, 2021.

Citation

MLA
Liang, P. P., et al. “Factorized Contrastive Learning: Going Beyond Multi-view Redundancy”. Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 32971–98, https://proceedings.neurips.cc/paper_files/paper/2023/file/6818dcc65fdf3cbd4b05770fb957803e-Paper-Conference.pdf.
APA
Liang, P. P., Deng, Z., Ma, M. Q., Zou, J., Morency, L.-P., & Salakhutdinov, R. (2023). Factorized Contrastive Learning: Going Beyond Multi-view Redundancy. Advances in Neural Information Processing Systems, 36, 32971–32998. https://proceedings.neurips.cc/paper_files/paper/2023/file/6818dcc65fdf3cbd4b05770fb957803e-Paper-Conference.pdf
Chicago
Liang, P. P., Z. Deng, M. Q. Ma, J. Zou, L.-P. Morency, and R. Salakhutdinov. 2023. “Factorized Contrastive Learning: Going Beyond Multi-view Redundancy”. Advances in Neural Information Processing Systems 36: 32971–98. https://proceedings.neurips.cc/paper_files/paper/2023/file/6818dcc65fdf3cbd4b05770fb957803e-Paper-Conference.pdf.
Harvard
Liang, P.P. et al. (2023) “Factorized Contrastive Learning: Going Beyond Multi-view Redundancy”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 32971–32998. Available at: https://proceedings.neurips.cc/paper_files/paper/2023/file/6818dcc65fdf3cbd4b05770fb957803e-Paper-Conference.pdf.
Vancouver
1. Liang PP, Deng Z, Ma MQ, Zou J, Morency L-P, Salakhutdinov R (2023) Factorized Contrastive Learning: Going Beyond Multi-view Redundancy. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 32971–32998

BibTeX

@inproceedings{liang2023factorized,
  title = {Factorized Contrastive Learning: Going Beyond Multi-view Redundancy},
  author = {Liang, Paul Pu and Deng, Zihao and Ma, Martin Q. and Zou, James and Morency, Louis-Philippe and Salakhutdinov, Ruslan},
  year = {2023},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {36},
  pages = {32971-32998},
  url = {https://proceedings.neurips.cc/paper_files/paper/2023/file/6818dcc65fdf3cbd4b05770fb957803e-Paper-Conference.pdf}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors