Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Cross-modal interactions

Cross-modal interactions refer to the mutual influences, relationships, and exchanges of information that occur between distinct data modalities, such as text, vision, and audio, when they are combined and processed within a computational model. In multimodal machine learning, these interactions capture complementary, redundant, and contextual dependencies across heterogeneous input streams rather than treating each modality in isolation. Modeling cross-modal dynamics allows features from one data channel to guide, reinforce, or disambiguate the interpretation of another, which is critical for effective multimodal fusion, joint representation learning, and holistic data understanding.

2 items

Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks

Characterizing and Overcoming the Greedy Nature of Learning in Multi-modal Deep Neural Networks

Nan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, Krzysztof J. Geras

OrganizationsCIFARGenentechNew York University

Why you should read this

Explains why multi-modal neural networks often over-rely on a single modality and introduces a training algorithm that balances learning speeds across modalities to improve overall generalization.

We hypothesize that due to the greedy nature of learning in multi-modal deep neural networks, these models tend to rely on just one modality while under-fitting the other modalities. Such behavior is counter-intuitive and hurts the models’ generalization, as we observe empirically. To estimate the model’s dependence on each modality, we compute the gain on the accuracy when the model has access to it in addition to another modality. We refer to this gain as the conditional utilization rate. In the experiments, we consistently observe an imbalance in conditional utilization rates between modalities, across multiple tasks and architectures. Since conditional utilization rate cannot be computed efficiently during training, we introduce a proxy for it based on the pace at which the model learns from each modality, which we refer to as the conditional learning speed. We propose an algorithm to balance the conditional learning speeds between modalities during training and demonstrate that it indeed addresses the issue of greedy learning.1 The proposed algorithm improves the model’s generalization on three datasets: Colored MNIST, ModelNet40, and NVIDIA Dynamic Hand Gesture.

Added

2026-09-26

Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

Amir Zadeh, P. Liang, Soujanya Poria, E. Cambria, Louis-philippe Morency

OrganizationsAgency for Science, Technology and ResearchCarnegie Mellon UniversityNanyang Technological University

Why you should read this

Introduces the CMU-MOSEI benchmark alongside the Dynamic Fusion Graph model to advance large-scale multimodal sentiment analysis and emotion recognition through interpretable cross-modal dynamics.

Analyzing human multimodal language is an emerging area of research in NLP. Intrinsically human communication is multimodal (heterogeneous), temporal and asynchronous; it consists of the language (words), visual (expressions), and acoustic (paralinguistic) modalities all in the form of asynchronous coordinated sequences. From a resource perspective, there is a genuine need for large scale datasets that allow for in-depth studies of multimodal language. In this paper we introduce CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI), the largest dataset of sentiment analysis and emotion recognition to date. Using data from CMU-MOSEI and a novel multimodal fusion technique called the Dynamic Fusion Graph (DFG), we conduct experimentation to investigate how modalities interact with each other in human multimodal language. Unlike previously proposed fusion techniques, DFG is highly interpretable and achieves competitive performance compared to the current state of the art.

Added

2026-09-24