Geometric Multimodal Contrastive Representation Learning

Petra PoklukarMiguel VascoHang YinFrancisco S. MeloAna PaivaDanica Kragic

article2022ICML67 citations

Proposes a geometric multimodal contrastive learning framework that aligns modality-specific encoders with complete observations in a shared latent space, maintaining high task performance even when individual modalities are missing during evaluation.

Listen

Modern artificial intelligence systems frequently rely on multiple data sources, such as video, audio, text, and physical sensors, to make decisions and control automated tasks. However, in practical deployment, systems often experience sensor failures, communication dropouts, or missing data streams. Existing machine learning models often struggle under these conditions because the internal representations derived from a single data source become geometrically misaligned with the representations learned when all data streams are present simultaneously.

The article introduces the Geometric Multimodal Contrastive (GMC) framework to solve this vulnerability. The primary objective is to demonstrate that GMC can create semantically rich, geometrically aligned data representations that maintain state-of-the-art performance across classification and control tasks, even when one or more data streams are entirely absent during testing.

To evaluate this capability, the authors conducted experiments across three distinct learning environments: unsupervised classification using handwriting digits across four channels (images, sound, motion trajectories, and labels), supervised sentiment classification on video datasets (CMU-MOSEI and CMU-MOSI), and reinforcement learning control on a simulated multimodal inverted pendulum. The approach uses a two-level neural network architecture combining modality-specific base encoders with a shared projection head. It applies a contrastive learning objective that directly aligns single-source representations with full multimodal representations in a shared space, scaling linearly with the number of input sources.

Across all benchmarks, GMC demonstrated significant performance and efficiency advantages. In unsupervised digit classification with incomplete input data, GMC achieved between 93.04% and 99.96% accuracy across individual modalities where competing baselines dropped as low as 10% to 79%. In supervised video sentiment analysis missing visual or audio streams, GMC raised binary accuracy to roughly 65% compared to baseline drops below 54%. In reinforcement learning control with missing image or sound data, agents using GMC experienced virtually zero performance degradation, maintaining an average return of approximately -0.94 to -0.96 compared to sharp declines to -6.64 in baseline methods. In addition, GMC achieved these results with high computational efficiency, requiring 50% to 68% fewer parameters than alternative multimodal baseline models.

These findings indicate that explicit geometric alignment prevents downstream task failure caused by sudden data source loss, directly improving system reliability, safety, and resilience in autonomous operations. By operating with significantly fewer parameters and offering straightforward integration with existing network architectures, the method reduces computing costs and deployment overhead without sacrificing baseline accuracy.

Organizations deploying multimodal models in real-world environments should consider adopting contrastive alignment methods to mitigate sensor failure risks. Next steps should include exploring modality-specific data augmentations and conducting pilot tests in real-world physical and industrial environments. Decision-makers should note that GMC's performance gains are slightly moderated when trained on smaller datasets, as demonstrated by the smaller CMU-MOSI evaluation, though confidence in the framework remains high given consistent improvements across diverse test domains.

Poklukar et al (2022).pdf
Cover for Geometric Multimodal Contrastive Representation Learning

Abstract

Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we present a novel Geometric Multimodal Contrastive (GMC) representation learning method consisting of two main components: i) a two-level architecture consisting of modality-specific base encoders, allowing to process an arbitrary number of modalities to an intermediate representation of fixed dimensionality, and a shared projection head, mapping the intermediate representations to a latent representation space; ii) a multimodal contrastive loss function that encourages the geometric alignment of the learned representations. We experimentally demonstrate that GMC representations are semantically rich and achieve state-of-the-art performance with missing modality information on three different learning problems including prediction and reinforcement learning tasks.

Table of Contents

  • 1. Introduction
  • 2. The Problem of Geometric Misalignment in Multimodal Representation Learning
  • 3. Geometric Multimodal Contrastive Learning
  • 4. Related Work
  • 5. Experiments
  • 5.1. Experiment 1: Unsupervised Learning
  • 5.2. Experiment 2: Supervised Learning
  • 5.3. Experiment 3: Reinforcement Learning
  • 5.4. Ablation studies
  • 6. Conclusion
  • Acknowledgements
  • References
  • A. Delaunay Component Analysis
  • B. Ablation Study on GMC
  • C. Experiment 2: Supervised Learning with the CMU-MOSI dataset
  • D. Model Architecture
  • E. Training Hyperparameters
  • F. Additional Visualizations of the Alignment of Complete and Modality-Specific Representations

Knowls

  1. Knowl 1 — Two-level architecture for aligned multimodal representations

    model/method

    Geometric Multimodal Contrastive (GMC) represents observations from MM modalities using modality-specific base encoders fmf_m and a complete-observation encoder f1:Mf_{1:M}. Each encoder maps its input to an intermediate vector in Rd\mathbb{R}^d; a single shared projection head gg maps these vectors into a common latent space Z⊆Rs\mathcal{Z}\subseteq\mathbb{R}^s. Thus, for a complete observation x1:M=(x1,…,xM)x_{1:M}=(x_1,\ldots,x_M) and its modality-specific observations xmx_m, the corresponding representations are z1:M=g(f1:M(x1:M))z_{1:M}=g(f_{1:M}(x_{1:M})) and zm=g(fm(xm))z_m=g(f_m(x_m)). Training aligns each zmz_m with the complete representation z1:Mz_{1:M} from the same sample while contrasting against representations from other samples. The modality encoders can be tailored to different data types, while the shared projection head and per-modality alignment objective allow the framework to accommodate an arbitrary number of modalities. The intended outcome is a complete representation and modality-specific representations that retain task-relevant information and can be used when some modalities are unavailable.

  2. Knowl 2 — GMC multimodal contrastive objective

    equation

    For a minibatch of BB paired observations, let zmi∈Rsz_m^i\in\mathbb{R}^s be the representation of modality mm for sample ii, and let z1:Mi∈Rsz_{1:M}^i\in\mathbb{R}^s be the complete representation of that sample. Here m∈{1,…,M}m\in\{1,\ldots,M\}, i,j∈{1,…,B}i,j\in\{1,\ldots,B\}, and the temperature τ\tau is positive. Define the exponentiated cosine similarity between representations of types aa and bb from samples ii and jj by

    sa,b(i,j)=exp⁡ ⁣(cos⁡(zai,zbj)τ),s_{a,b}(i,j)=\exp\!\left(\frac{\operatorname{cos}(z_a^i,z_b^j)}{\tau}\right),

    where aa or bb may denote an individual modality or the complete observation. For each modality mm and anchor sample ii, the positive pair is (zmi,z1:Mi)(z_m^i,z_{1:M}^i). The negative-pair similarity sum and per-pair loss are

    Ωm(i)=∑j≠i[sm,1:M(i,j)+sm,m(i,j)+s1:M,1:M(i,j)],ℓm(i)=−log⁡ ⁣(sm,1:M(i,i)Ωm(i)).\Omega_m(i)=\sum_{j\ne i}\left[s_{m,1:M}(i,j)+s_{m,m}(i,j)+s_{1:M,1:M}(i,j)\right], \qquad \ell_m(i)=-\log\!\left(\frac{s_{m,1:M}(i,i)}{\Omega_m(i)}\right).

    The total minibatch objective sums over all modalities and samples:

    LGMC(B)=∑m=1M∑i=1Bℓm(i).\mathcal{L}_{\mathrm{GMC}}(B)=\sum_{m=1}^{M}\sum_{i=1}^{B}\ell_m(i).

    The negative terms compare modality-specific-to-complete, same-modality-to-same-modality, and complete-to-complete representations from different samples. Contrasting each individual modality against the complete observation separately makes the objective scale linearly with the number of modalities.

  3. Knowl 3 — GMC improves MHD classification with partial observations

    empirical result

    On the Multimodal Handwritten Digits (MHD) dataset, GMC representations were evaluated by training a 10-class classifier on complete training representations and testing it with either complete or single-modality representations. MHD has image, sound, trajectory, and label modalities, with 50,000 training and 10,000 test samples. The GMC configuration used 64-dimensional intermediate and latent representations, temperature τ=0.1\tau=0.1, 100 representation-training epochs, and a learning rate of 10−310^{-3}. Accuracy below is test accuracy in percent, averaged over five runs, except MVAE, which used three runs because of training divergence. GMC achieved near-perfect performance from each individual modality and matched perfect accuracy on complete inputs; the strongest baselines were substantially less reliable for sound and trajectory inputs.

    Input MVAE MMVAE Nexus MUSE MFM GMC
    Complete (x1:4x_{1:4}) 100.0±0.00100.0\pm0.00 99.81±0.2199.81\pm0.21 99.98±0.0599.98\pm0.05 99.99±4e−599.99\pm4\mathrm{e}{-5} 100.0±0.00100.0\pm0.00 100.0±0.00100.0\pm0.00
    Image (x1x_1) 77.94±3.1677.94\pm3.16 94.63±2.6194.63\pm2.61 95.89±0.3495.89\pm0.34 79.37±2.7579.37\pm2.75 34.66±6.4834.66\pm6.48 99.75±0.0399.75\pm0.03
    Sound (x2x_2) 61.75±4.5961.75\pm4.59 69.43±26.4369.43\pm26.43 39.07±5.8239.07\pm5.82 41.39±0.1841.39\pm0.18 10.07±0.2010.07\pm0.20 93.04±0.4593.04\pm0.45
    Trajectory (x3x_3) 10.03±0.0610.03\pm0.06 95.33±2.5695.33\pm2.56 98.55±0.3498.55\pm0.34 89.49±2.4489.49\pm2.44 25.61±5.4125.61\pm5.41 99.96±0.0299.96\pm0.02
    Label (x4x_4) 100.0±0.00100.0\pm0.00 87.99±7.4987.99\pm7.49 100.0±0.00100.0\pm0.00 100.0±0.00100.0\pm0.00 100.0±0.00100.0\pm0.00 100.0±0.00100.0\pm0.00
  4. Knowl 4 — MHD representations are better aligned with fewer parameters

    empirical result

    The paper measured geometric alignment on held-out MHD representations, using complete representations z1:4z_{1:4} as the Delaunay Component Analysis (DCA) reference set and each modality-specific representation as the evaluation set. The reported scalar is the harmonic mean of DCA precision PP, recall RR, and network quality qq, namely 3/(1/P+1/R+1/q)3/(1/P+1/R+1/q) when all three are positive, and zero otherwise. A higher score indicates that more points from the two sets lie in geometrically well-aligned components. GMC aligned all four modalities strongly with the complete representation, whereas most baselines had zero or near-zero scores for sound and trajectory. The scores are means over five runs.

    Evaluation modality MVAE MMVAE Nexus MUSE MFM GMC
    Image (z1z_1) 0.01±0.010.01\pm0.01 0.21±0.290.21\pm0.29 0.00±0.000.00\pm0.00 0.54±0.440.54\pm0.44 0.00±0.000.00\pm0.00 0.96±0.020.96\pm0.02
    Sound (z2z_2) 0.00±0.000.00\pm0.00 0.00±0.000.00\pm0.00 0.00±0.000.00\pm0.00 0.00±0.000.00\pm0.00 0.00±0.000.00\pm0.00 0.87±0.160.87\pm0.16
    Trajectory (z3z_3) 0.00±0.000.00\pm0.00 0.01±0.010.01\pm0.01 0.08±0.020.08\pm0.02 0.00±0.000.00\pm0.00 0.00±0.000.00\pm0.00 0.86±0.050.86\pm0.05
    Label (z4z_4) 0.99±0.010.99\pm0.01 0.74±0.220.74\pm0.22 0.43±0.050.43\pm0.05 0.93±0.050.93\pm0.05 0.85±0.060.85\pm0.06 1.00±0.001.00\pm0.00

    The representation models used 9.3 million parameters for MVAE, 9.0 million for MMVAE, 12.9 million for Nexus, 9.9 million for MUSE, 9.4 million for MFM, and 2.9 million for GMC. Thus, GMC used about 68% fewer parameters than the smallest baseline, MMVAE, in this experiment.

  5. Knowl 5 — GMC preserves CMU-MOSEI performance and improves single-modality prediction

    empirical result

    On CMU-MOSEI sentiment prediction, GMC was added to a Multimodal Transformer-based complete-observation encoder and evaluated against the baseline using the same prediction metrics. Results are means over five runs. With complete text, audio, and video inputs, GMC was competitive with the baseline: it slightly improved mean absolute error (MAE), while the baseline had higher correlation, F1, and binary accuracy. With any one modality alone, GMC improved every reported metric, with especially large gains in F1 and binary accuracy. The table reports MAE (lower is better), correlation (Cor), F1, and binary accuracy (Acc, percent; higher is better), as baseline / GMC.

    Input Metric Baseline GMC
    Complete MAE 0.643±0.0190.643\pm0.019 0.634±0.0080.634\pm0.008
    Complete Cor 0.664±0.0040.664\pm0.004 0.653±0.0040.653\pm0.004
    Complete F1 0.809±0.0030.809\pm0.003 0.798±0.0080.798\pm0.008
    Complete Acc (%) 80.75±00.2880.75\pm00.28 79.73±00.6979.73\pm00.69
    Text only MAE 0.805±0.0280.805\pm0.028 0.712±0.0150.712\pm0.015
    Text only Cor 0.427±0.0610.427\pm0.061 0.590±0.0130.590\pm0.013
    Text only F1 0.713±0.0860.713\pm0.086 0.779±0.0050.779\pm0.005
    Text only Acc (%) 66.53±09.8666.53\pm09.86 77.85±00.3677.85\pm00.36
    Audio only MAE 0.873±0.0650.873\pm0.065 0.837±0.0080.837\pm0.008
    Audio only Cor 0.090±0.0620.090\pm0.062 0.256±0.0070.256\pm0.007
    Audio only F1 0.622±0.1220.622\pm0.122 0.676±0.0150.676\pm0.015
    Audio only Acc (%) 53.17±09.4753.17\pm09.47 65.59±00.6265.59\pm00.62
    Video only MAE 1.025±0.1641.025\pm0.164 0.845±0.0100.845\pm0.010
    Video only Cor 0.110±0.0600.110\pm0.060 0.278±0.0110.278\pm0.011
    Video only F1 0.574±0.0950.574\pm0.095 0.655±0.0030.655\pm0.003
    Video only Acc (%) 44.33±09.4044.33\pm09.40 65.02±00.2865.02\pm00.28
  6. Knowl 6 — CMU-MOSEI modality-specific embeddings align with complete embeddings

    empirical result

    For CMU-MOSEI test samples, DCA compared the complete representation set z1:3z_{1:3} against each single-modality set. GMC produced higher alignment scores than the baseline for text, audio, and video, consistent with its improved single-modality prediction results. Scores are means over five runs; larger is better. GMC used 1.4 million parameters, compared with 1.1 million for the baseline, an additional 300,000 parameters.

    Evaluation modality Baseline DCA GMC DCA
    Text (z1z_1) 0.50±0.050.50\pm0.05 0.95±0.010.95\pm0.01
    Audio (z2z_2) 0.41±0.140.41\pm0.14 0.86±0.040.86\pm0.04
    Vision (z3z_3) 0.50±0.140.50\pm0.14 0.92±0.020.92\pm0.02
  7. Knowl 7 — GMC enables zero-shot control from incomplete Pendulum observations

    empirical result

    In the multimodal inverted Pendulum task, image and sound observations were used to train representation models from 20,000 randomly collected training samples; 2,000 samples were held out for testing. A DDPG controller was trained using complete-observation representations and then evaluated without additional training when given complete observations, images alone, or sound alone. Returns are total reward per episode, averaged over 100 episodes and 10 random seeds; higher (less negative) is better. GMC maintained similar returns across all three observation conditions, whereas MVAE and MUSE lost substantial performance for at least one missing-modality condition. The DCA scores compare each modality-specific test representation against the complete representation; parameter counts refer to representation models.

    Observation MVAE + DDPG MUSE + DDPG GMC + DDPG
    Complete (x1:2x_{1:2}) −1.114±0.110-1.114\pm0.110 −1.005±0.117-1.005\pm0.117 −0.935±0.057-0.935\pm0.057
    Image (x1x_1) −1.116±0.121-1.116\pm0.121 −4.752±0.994-4.752\pm0.994 −0.940±0.056-0.940\pm0.056
    Sound (x2x_2) −6.642±0.106-6.642\pm0.106 −3.459±0.519-3.459\pm0.519 −0.956±0.075-0.956\pm0.075
    Alignment comparison MVAE + DDPG MUSE + DDPG GMC + DDPG
    Complete vs. image 0.79±0.010.79\pm0.01 0.20±0.090.20\pm0.09 0.87±0.010.87\pm0.01
    Complete vs. sound 0.00±0.000.00\pm0.00 0.01±0.010.01\pm0.01 0.88±0.020.88\pm0.02

    The representation models had 3.8 million parameters for MVAE, 4.3 million for MUSE, and 1.9 million for GMC. MUSE results used nine seeds because training diverged on the remaining seed.

  8. Knowl 8 — CMU-MOSI results show improved alignment but mixed prediction gains

    empirical result

    On CMU-MOSI, which had 1,513 training samples, GMC was evaluated using the same supervised setup as on CMU-MOSEI. The table gives test metrics as baseline / GMC, averaged over five runs. GMC was competitive on complete inputs and improved all four metrics for text-only input. For audio-only and video-only input, it improved correlation and binary accuracy but not every metric: audio MAE and F1 and video MAE and F1 were worse than the baseline. Despite these mixed prediction results, the DCA alignment score increased for all three modalities. The authors hypothesize that the smaller training set makes it harder to form useful contrastive pairs; this is offered as an explanation, not established as a cause.

    Input Metric Baseline GMC
    Complete MAE 1.033±0.0371.033\pm0.037 1.010±0.0701.010\pm0.070
    Complete Cor 0.642±0.0080.642\pm0.008 0.649±0.0190.649\pm0.019
    Complete F1 0.770±0.0170.770\pm0.017 0.776±0.0230.776\pm0.023
    Complete Acc (%) 77.07±01.6777.07\pm01.67 77.59±02.2077.59\pm02.20
    Text only MAE 1.244±0.1001.244\pm0.100 1.119±0.0331.119\pm0.033
    Text only Cor 0.431±0.2080.431\pm0.208 0.573±0.0160.573\pm0.016
    Text only F1 0.698±0.0530.698\pm0.053 0.727±0.0130.727\pm0.013
    Text only Acc (%) 66.28±07.7466.28\pm07.74 72.32±0.01372.32\pm0.013
    Audio only MAE 1.431±0.0251.431\pm0.025 1.434±0.0171.434\pm0.017
    Audio only Cor 0.056±0.0710.056\pm0.071 0.211±0.0100.211\pm0.010
    Audio only F1 0.588±0.0760.588\pm0.076 0.570±0.0060.570\pm0.006
    Audio only Acc (%) 47.20±05.6747.20\pm05.67 55.91±01.1155.91\pm01.11
    Video only MAE 1.406±0.0411.406\pm0.041 1.452±0.0351.452\pm0.035
    Video only Cor 0.021±0.0280.021\pm0.028 0.176±0.0280.176\pm0.028
    Video only F1 0.659±0.0490.659\pm0.049 0.550±0.0150.550\pm0.015
    Video only Acc (%) 53.87±05.7753.87\pm05.77 54.30±01.9654.30\pm01.96

    For text, audio, and video, respectively, the baseline DCA scores were 0.54±0.070.54\pm0.07, 0.14±0.060.14\pm0.06, and 0.36±0.090.36\pm0.09; GMC scores were 0.93±0.020.93\pm0.02, 0.75±0.050.75\pm0.05, and 0.85±0.040.85\pm0.04. MAE is lower-is-better; correlation, F1, and binary accuracy (Acc) are higher-is-better.

  9. Knowl 9 — Ablations identify the role of modality-specific negatives

    empirical result

    On MHD, GMC classification performance was broadly stable across temperatures τ∈{0.05,0.1,0.2,0.3,0.5}\tau\in\{0.05,0.1,0.2,0.3,0.5\} and intermediate and latent dimensions in {32,64,128}\{32,64,128\}. For example, sound accuracy ranged from 91.87±0.58%91.87\pm0.58\% to 95.01±0.38%95.01\pm0.38\% over the tested temperatures; with intermediate dimensions 32, 64, and 128 it was 93.31±0.41%93.31\pm0.41\%, 93.04±0.45%93.04\pm0.45\%, and 93.34±0.51%93.34\pm0.51\%, respectively. Alignment could vary more than classification: at τ=0.5\tau=0.5, trajectory-to-complete DCA fell to 0.64±0.110.64\pm0.11, compared with 0.86±0.050.86\pm0.05 at the default τ=0.1\tau=0.1.

    The loss ablation used a modified objective that retained only complete-observation representations as negative pairs, removing the modality-specific negative comparisons. Classification accuracy changed little, but alignment deteriorated sharply for some modalities. In image, sound, trajectory, and label order, default versus modified-loss DCA scores were 0.96±0.020.96\pm0.02 vs. 0.80±0.020.80\pm0.02, 0.87±0.160.87\pm0.16 vs. 0.27±0.140.27\pm0.14, 0.86±0.050.86\pm0.05 vs. 0.86±0.030.86\pm0.03, and 1.00±0.001.00\pm0.00 vs. 0.24±0.100.24\pm0.10. The corresponding accuracies were 99.75±0.03%99.75\pm0.03\% vs. 99.87±0.01%99.87\pm0.01\%, 93.04±0.45%93.04\pm0.45\% vs. 92.79±0.24%92.79\pm0.24\%, 99.96±0.02%99.96\pm0.02\% vs. 99.98±0.01%99.98\pm0.01\%, and 100.00±0.00%100.00\pm0.00\% vs. 100.00±0.00%100.00\pm0.00\%. These results indicate that modality-specific negative comparisons are important for geometric alignment, even when downstream classification accuracy changes minimally.

Coverage note — No substantial contributed material was omitted. The detailed layer-by-layer network diagrams and additional UMAP visualizations are omitted because they implement or illustrate the architecture and alignment results captured above, rather than adding a separate finding.

References

  1. 1.Bagher Zadeh, A., Liang, P. P., Poria, S., Cambria, E., and Morency, L.-P. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2236–2246, 2018.
  2. 2.Baltrusaitis, T., Ahuja, C., and Morency, L.-P. Multimodal machine learning: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 41 (2):423–443, 2018.
  3. 3.Cao, Y.-H. and Wu, J. Rethinking self-supervised learning: Small is beautiful. arXiv preprint arXiv:2103.13559, 2021.
  4. 4.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. In Proceedings of the 37th International Conference on Machine Learning (ICML), pp. 1597–1607, 2020.
  5. 5.Guo, W., Wang, J., and Wang, S. Deep multimodal representation learning: A survey. IEEE Access, 7:63373–63394, 2019.
  6. 6.Higgins, I., Pal, A., Rusu, A., Matthey, L., Burgess, C., Pritzel, A., Botvinick, M., Blundell, C., and Lerchner, A. Darla: Improving zero-shot transfer in reinforcement learning. In Procedings of the 34th International Conference on Machine Learning (ICML), pp. 1480–1490, 2017.
  7. 7.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In Proceedings of the 2nd International Conference on Learning Representations (ICLR), 2014.
  8. 8.Liang, P., Lyu, Y., Fan, X., Wu, Z., Cheng, Y., Wu, J., Chen, L., Wu, P., Lee, M., Zhu, Y., et al. Multibench: Multiscale benchmarks for multimodal representation learning. In Proceedings of the 35th International Conference on Neural Information Processing Systems (Neurips), 2021.
  9. 9.Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015.
  10. 10.McInnes, L., Healy, J., Saul, N., and Großberger, L. Umap: Uniform manifold approximation and projection. Journal of Open Source Software, 3(29):861, 2018.
  11. 11.Meo, C. and Lanillos, P. Multimodal vae active inference controller. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 2693–2699. IEEE, 2021.
  12. 12.Poklukar, P., Varava, A., and Kragic, D. Geomca: Geometric evaluation of data representations. In Proceedings of the 38th International Conference on Machine Learning (ICML), pp. 8588–8598, 2021.
  13. 13.Poklukar, P., Polianskii, V., Varava, A., Pokorny, F., and Kragic, D. Delaunay component analysis for evaluation of data representations. In Procedings of the 10th International Conference on Learning Representations (ICLR), 2022.
  14. 14.Shi, Y., Siddharth, N., Paige, B., and Torr, P. H. Variational mixture-of-experts autoencoders for multi-modal deep generative models. In Proceedings of the 33rd International Conference on Neural Information Processing Systems (Neurips), pp. 15718–15729, 2019.
  15. 15.Silva, R., Vasco, M., Melo, F. S., Paiva, A., and Veloso, M. Playing games in the dark: An approach for cross-modality transfer in reinforcement learning. In Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS), pp. 1260–1268, 2020.
  16. 16.Suzuki, M., Nakayama, K., and Matsuo, Y. Joint multimodal learning with deep generative models. arXiv preprint arXiv:1611.01891, 2016.
  17. 17.Tremblay, J.-F., Manderson, T., Noca, A., Dudek, G., and Meger, D. Multimodal dynamics modeling for off-road autonomous vehicles. In Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), pp. 1796–1802. IEEE, 2021.
  18. 18.Tsai, Y.-H. H., Bai, S., Liang, P. P., Kolter, J. Z., Morency, L.-P., and Salakhutdinov, R. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 6558–6569, 2019a.
  19. 19.Tsai, Y.-H. H., Liang, P. P., Zadeh, A., Morency, L.-P., and Salakhutdinov, R. Learning factorized multimodal representations. In Proceedings of the 7th International Conference on Learning Representations (ICLR), 2019b.
  20. 20.Vasco, M., Yin, H., Melo, F. S., and Paiva, A. How to sense the world: Leveraging hierarchy in multimodal perception for robust reinforcement learning agents. In Proceedings of the 21st International Conference on Autonomous Agents and MultiAgent Systems (AAMAS), pp. 1301–1309, 2022a.
  21. 21.Vasco, M., Yin, H., Melo, F. S., and Paiva, A. Leveraging hierarchy in multimodal generative models for effective cross-modality inference. Neural Networks, 146:238–255, 2022b.
  22. 22.Wu, M. and Goodman, N. Multimodal generative models for scalable weakly-supervised learning. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Neurips), pp. 5580–5590, 2018.
  23. 23.Yin, H., Melo, F. S., Billard, A., and Paiva, A. Associate latent encodings in learning from demonstrations. In Procedings of the 31st AAAI Conference on Artificial Intelligence (AAAI), 2017.
  24. 24.Zadeh, A., Zellers, R., Pincus, E., and Morency, L.-P. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages. IEEE Intelligent Systems, 31(6):82–88, 2016.
  25. 25.Zambelli, M., Cully, A., and Demiris, Y. Multimodal representation models for prediction and control from partial information. Robotics and Autonomous Systems, 123:103312, 2020.

Citation

MLA
Poklukar, P., et al. “Geometric Multimodal Contrastive Representation Learning”. International Conference on Machine Learning, vol. 162, 2022, pp. 17782–800, https://proceedings.mlr.press/v162/poklukar22a.html.
APA
Poklukar, P., Vasco, M., Yin, H., Melo, F. S., Paiva, A., & Kragic, D. (2022). Geometric Multimodal Contrastive Representation Learning. International Conference on Machine Learning, 162, 17782–17800. https://proceedings.mlr.press/v162/poklukar22a.html
Chicago
Poklukar, P., M. Vasco, H. Yin, F. S. Melo, A. Paiva, and D. Kragic. 2022. “Geometric Multimodal Contrastive Representation Learning”. International Conference on Machine Learning 162: 17782–800. https://proceedings.mlr.press/v162/poklukar22a.html.
Harvard
Poklukar, P. et al. (2022) “Geometric Multimodal Contrastive Representation Learning”, International Conference on Machine Learning. PMLR, pp. 17782–17800. Available at: https://proceedings.mlr.press/v162/poklukar22a.html.
Vancouver
1. Poklukar P, Vasco M, Yin H, Melo FS, Paiva A, Kragic D (2022) Geometric Multimodal Contrastive Representation Learning. In: International Conference on Machine Learning. PMLR, pp 17782–17800

BibTeX

@InProceedings{pmlr-v162-poklukar22a,
  title = 	 {Geometric Multimodal Contrastive Representation Learning},
  author =       {Poklukar, Petra and Vasco, Miguel and Yin, Hang and Melo, Francisco S. and Paiva, Ana and Kragic, Danica},
  booktitle = 	 {Proceedings of the 39th International Conference on Machine Learning},
  pages = 	 {17782--17800},
  year = 	 {2022},
  editor = 	 {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan},
  volume = 	 {162},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {17--23 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v162/poklukar22a/poklukar22a.pdf},
  url = 	 {https://proceedings.mlr.press/v162/poklukar22a.html},
  abstract = 	 {Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we present a novel Geometric Multimodal Contrastive (GMC) representation learning method consisting of two main components: i) a two-level architecture consisting of modality-specific base encoders, allowing to process an arbitrary number of modalities to an intermediate representation of fixed dimensionality, and a shared projection head, mapping the intermediate representations to a latent representation space; ii) a multimodal contrastive loss function that encourages the geometric alignment of the learned representations. We experimentally demonstrate that GMC representations are semantically rich and achieve state-of-the-art performance with missing modality information on three different learning problems including prediction and reinforcement learning tasks.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/