Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Geometric multimodal contrastive representation learning

Geometric multimodal contrastive representation learning is a machine learning paradigm that trains models to map data from diverse modalities into a shared representation space by aligning their geometric structures through contrastive objectives. In this approach, modality-specific encoders process heterogeneous inputs into intermediate features before a shared projection maps them into a common latent embedding space. By using contrastive loss functions that enforce geometric alignment among corresponding multimodal inputs, the framework preserves shared semantic relationships across different sensory channels. This structured alignment produces semantically rich representations that allow the system to effectively integrate information and maintain robust performance even when certain modalities are missing during evaluation.

1 item

Geometric Multimodal Contrastive Representation Learning

Geometric Multimodal Contrastive Representation Learning

Petra Poklukar, Miguel Vasco, Hang Yin, Francisco S. Melo, Ana Paiva, Danica Kragic

Why you should read this

Proposes a geometric multimodal contrastive learning framework that aligns modality-specific encoders with complete observations in a shared latent space, maintaining high task performance even when individual modalities are missing during evaluation.

Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we present a novel Geometric Multimodal Contrastive (GMC) representation learning method consisting of two main components: i) a two-level architecture consisting of modality-specific base encoders, allowing to process an arbitrary number of modalities to an intermediate representation of fixed dimensionality, and a shared projection head, mapping the intermediate representations to a latent representation space; ii) a multimodal contrastive loss function that encourages the geometric alignment of the learned representations. We experimentally demonstrate that GMC representations are semantically rich and achieve state-of-the-art performance with missing modality information on three different learning problems including prediction and reinforcement learning tasks.

Added

2026-10-02