keyword
modality-specific representations
Modality-specific representations are learned feature encodings that capture the unique characteristics, structures, and information exclusive to a single type of data input, such as text, audio, or images. In multimodal machine learning systems, these representations are generated by dedicated encoders tailored to the distinct statistical properties and formats of an individual data channel. Unlike modality-invariant or shared representations, which capture common semantic elements across multiple channels, modality-specific representations preserve the specialized, non-redundant nuances inherent to each medium. Retaining these channel-specific features allows computational models to effectively integrate complementary information from diverse sensory streams, better analyze complex multimodal relationships, and maintain functional robustness even when certain data sources are incomplete or noisy.
2 items

Geometric Multimodal Contrastive Representation Learning
Petra Poklukar, Miguel Vasco, Hang Yin, Francisco S. Melo, Ana Paiva, Danica Kragic
Why you should read this
Proposes a geometric multimodal contrastive learning framework that aligns modality-specific encoders with complete observations in a shared latent space, maintaining high task performance even when individual modalities are missing during evaluation.
Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we present a novel Geometric Multimodal Contrastive (GMC) representation learning method consisting of two main components: i) a two-level architecture consisting of modality-specific base encoders, allowing to process an arbitrary number of modalities to an intermediate representation of fixed dimensionality, and a shared projection head, mapping the intermediate representations to a latent representation space; ii) a multimodal contrastive loss function that encourages the geometric alignment of the learned representations. We experimentally demonstrate that GMC representations are semantically rich and achieve state-of-the-art performance with missing modality information on three different learning problems including prediction and reinforcement learning tasks.
Added
2026-10-02

ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment Analysis
Jiuding Yang, Yakun Yu, Di Niu, Weidong Guo, Yu Xu
Why you should read this
Proposes a multimodal sentiment analysis framework that combines inter-sample contrastive learning with text-centered feature decomposition to isolate shared and modality-specific information, achieving state-of-the-art results across standard video benchmarks.
Multimodal Sentiment Analysis aims to predict the sentiment of video content. Recent research suggests that multimodal sentiment analysis critically depends on learning a good representation of multimodal information, which should contain both modality-invariant representations that are consistent across modalities as well as modality-specific representations. In this paper, we propose ConFEDE, a unified learning framework that jointly performs contrastive representation learning and contrastive feature decomposition to enhance representation of multimodal information. It decomposes each of the three modalities of a video sample, including text, video frames, and audio, into a similarity feature and a dissimilarity feature, which are learned by a contrastive relation centered around text. We conducted extensive experiments on CH-SIMS, MOSI and MOSEI to evaluate various state-of-the-art multimodal sentiment analysis methods. Experimental results show that ConFEDE outperforms all baselines on these datasets on a range of metrics.
Added
2026-10-01
