keyword
multimodal sentiment analysis
Multimodal sentiment analysis is an artificial intelligence research field and computational task that determines the sentiment, attitude, or affective state expressed in content by integrating information from multiple communication modalities, primarily text, visual signals, and acoustic features. While traditional sentiment analysis focuses strictly on written language, multimodal approaches combine spoken words with nonverbal cues such as facial expressions, vocal inflections, and gestures, typically extracted from video and audio recordings. The process involves learning representations from heterogeneous data streams, addressing alignment and sequence variations across modalities, and applying fusion techniques to capture both modality-specific nuances and cross-modal interactions for a more comprehensive understanding of human communication.
12 items

Decoupled Multimodal Distilling for Emotion Recognition
Yong Li, Yuanzhi Wang, Zhen Cui
Why you should read this
Proposes a decoupled multimodal distillation framework that separates representations into shared and modality-exclusive spaces and applies dynamic graph distillation to adaptively transfer knowledge across language, visual, and acoustic streams for more accurate emotion recognition.
Human multimodal emotion recognition (MER) aims to perceive human emotions via language, visual and acoustic modalities. Despite the impressive performance of previous MER approaches, the inherent multimodal heterogeneities still haunt and the contribution of different modalities varies significantly. In this work, we mitigate this issue by proposing a decoupled multimodal distillation (DMD) approach that facilitates flexible and adaptive crossmodal knowledge distillation, aiming to enhance the discriminative features of each modality. Specially, the representation of each modality is decoupled into two parts, i.e., modality-irrelevant/-exclusive spaces, in a self-regression manner. DMD utilizes a graph distillation unit (GD-Unit) for each decoupled part so that each GD can be performed in a more specialized and effective manner. A GD-Unit consists of a dynamic graph where each vertex represents a modality and each edge indicates a dynamic knowledge distillation. Such GD paradigm provides a flexible knowledge transfer manner where the distillation weights can be automatically learned, thus enabling diverse crossmodal knowledge transfer patterns. Experimental results show DMD consistently obtains superior performance than state-of-the-art MER methods. Visualization results show the graph edges in DMD exhibit meaningful distributional patterns w.r.t. the modality-irrelevant/-exclusive feature spaces. Codes are released at https://github.com/mdszwyz/DMD.
Added
2026-10-05

AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
Zheng Lian, Haoyu Chen, Lan Chen, Haiyang Sun, Licai Sun, Yong Ren, Zebang Cheng, Bin Liu, Rui Liu, Xiaojiang Peng, Jiangyan Yi, Jianhua Tao
Why you should read this
Presents AffectGPT alongside MER-Caption, a 115K-sample descriptive video emotion dataset spanning over 2,000 fine-grained categories, and the MER-UniBench benchmark to improve complex, open-ended emotion recognition using multimodal large language models.
The emergence of multimodal large language models (MLLMs) advances multimodal emotion recognition (MER) to the next level—from naive discriminative tasks to complex emotion understanding with advanced video understanding abilities and natural language description. However, the current community suffers from a lack of large-scale datasets with intensive, descriptive emotion annotations, as well as a multimodal-centric framework to maximize the potential of MLLMs for emotion understanding. To address this, we establish a new benchmark for MLLM-based emotion understanding with a novel dataset (MER-Caption) and a new model (AffectGPT). Utilizing our model-based crowd-sourcing data collection strategy, we construct the largest descriptive emotion dataset to date (by far), featuring over 2K fine-grained emotion categories across 115K samples. We also introduce the AffectGPT model, designed with pre-fusion operations to enhance multimodal integration. Finally, we present MER-UniBench, a unified benchmark with evaluation metrics tailored for typical MER tasks and the free-form, natural language output style of MLLMs. Extensive experimental results show AffectGPT’s robust performance across various MER tasks. We have released both the code and the dataset to advance research and development in emotion understanding: https://github.com/zeroQiaoba/AffectGPT.
Added
2026-10-04

Multimodal Multi-loss Fusion Network for Sentiment Analysis
Zehui Wu, Ziwei Gong, Jaywon Koo, Julia Hirschberg
Why you should read this
Proposes an end-to-end multimodal fusion framework that combines cross-modal attention, self-attention, and auxiliary single-modality training losses to achieve state-of-the-art sentiment analysis performance on CMU-MOSI, CMU-MOSEI, and CH-SIMS benchmarks.
Sentiment analysis has become increasingly important due to the rapid growth of social media data. Recent research has shown that multimodal approaches which combine text, acoustic, and visual modalities can outperform unimodal methods. In particular, fusion strategies such as early fusion, late fusion, and model-level fusion have been explored extensively. However, effectively leveraging multiple losses to guide the training process remains under-explored. To address this gap, we propose a Multimodal Multi-loss Fusion Network (MMLFN). Our approach incorporates three different loss functions—focal loss, contrastive loss, and center loss—into a unified framework. We evaluate MMLFN on several benchmark datasets, achieving state-of-the-art performance. Extensive experiments demonstrate that our method significantly improves upon existing techniques. Moreover, comprehensive ablation studies verify the effectiveness of each component in our design. With joint optimization using complementary information across modalities as well as diverse objectives via distinct supervision signals during learning phase yields robust representations leading toward superior results compared prior art hence benefit...
Added
2026-10-02

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis
Haoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu, Yuanyuan Liu, Tianshu Yu
Why you should read this
Proposes an adaptive language-guided transformer that suppresses conflicting and irrelevant visual and acoustic signals using multi-scale language cues, achieving state-of-the-art multimodal sentiment analysis performance across standard benchmarks like MOSI and MOSEI.
Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.
Added
2026-10-01

ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment Analysis
Jiuding Yang, Yakun Yu, Di Niu, Weidong Guo, Yu Xu
Why you should read this
Proposes a multimodal sentiment analysis framework that combines inter-sample contrastive learning with text-centered feature decomposition to isolate shared and modality-specific information, achieving state-of-the-art results across standard video benchmarks.
Multimodal Sentiment Analysis aims to predict the sentiment of video content. Recent research suggests that multimodal sentiment analysis critically depends on learning a good representation of multimodal information, which should contain both modality-invariant representations that are consistent across modalities as well as modality-specific representations. In this paper, we propose ConFEDE, a unified learning framework that jointly performs contrastive representation learning and contrastive feature decomposition to enhance representation of multimodal information. It decomposes each of the three modalities of a video sample, including text, video frames, and audio, into a similarity feature and a dissimilarity feature, which are learned by a contrastive relation centered around text. We conducted extensive experiments on CH-SIMS, MOSI and MOSEI to evaluate various state-of-the-art multimodal sentiment analysis methods. Experimental results show that ConFEDE outperforms all baselines on these datasets on a range of metrics.
Added
2026-10-01

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition
Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, Yongbin Li
Why you should read this
Proposes UniMSE, a generative framework that unifies multimodal sentiment analysis and emotion recognition in conversation by combining label spaces, integrating acoustic and visual signals directly into a T5 backbone, and applying inter-modality contrastive learning.
Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period. However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models. We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions. Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMOCAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.
Added
2026-09-26

COGMEN: COntextualized GNN based Multimodal Emotion recognitioN
Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Vikram Singh, Ashutosh Modi
Why you should read this
Proposes a contextualized graph neural network architecture that models both global conversational context and speaker dependencies across multiple modalities to achieve state-of-the-art multimodal emotion recognition on IEMOCAP and MOSEI.
Emotions are an inherent part of human interactions, and consequently, it is imperative to develop AI systems that understand and recognize human emotions. During a conversation involving various people, a person’s emotions are influenced by the other speaker’s utterances and their own emotional state over the utterances. In this paper, we propose COntextualized Graph Neural Network based Multimodal Emotion recognitioN (COGMEN) system that leverages local information (i.e., inter/intra dependency between speakers) and global information (context). The proposed model uses Graph Neural Network (GNN) based architecture to model the complex dependencies (local and global information) in a conversation. Our model gives state-of-the-art (SOTA) results on IEMOCAP and MOSI datasets, and detailed ablation experiments show the importance of modeling information at both levels.
Added
2026-09-26

MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis
Devamanyu Hazarika, Roger Zimmermann, Soujanya Poria
Why you should read this
Proposes a multimodal representation framework, MISA, that factorizes signals into invariant and modality-specific subspaces, effectively resolving heterogeneous modality gaps to achieve state-of-the-art performance in sentiment analysis and humor detection.
Multimodal Sentiment Analysis is an active area of research that leverages multimodal signals for affective understanding of user-generated videos. The predominant approach, addressing this task, has been to develop sophisticated fusion techniques. However, the heterogeneous nature of the signals creates distributional modality gaps that pose significant challenges. In this paper, we aim to learn effective modality representations to aid the process of fusion. We propose a novel framework, MISA, which projects each modality to two distinct subspaces. The first subspace is modality-invariant, where the representations across modalities learn their commonalities and reduce the modality gap. The second subspace is modality-specific, which is private to each modality and captures their characteristic features. These representations provide a holistic view of the multimodal data, which is used for fusion that leads to task predictions. Our experiments on popular sentiment analysis benchmarks, MOSI and MOSEI, demonstrate significant gains over state-of-the-art models. We also consider the task of Multimodal Humor Detection and experiment on the recently proposed UR_FUNNY dataset. Here too, our model fares better than strong baselines, establishing MISA as a useful multimodal framework.
Added
2026-09-25

MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, Rada Mihalcea
Why you should read this
Introduces MELD, a large-scale benchmark containing over 13,000 text, audio, and visual utterances designed to advance emotion and sentiment recognition in multi-party conversations.
Emotion recognition in conversations is a challenging task that has recently gained popularity due to its potential applications. Until now, however, a large-scale multimodal multi-party emotional conversational database containing more than two speakers per dialogue was missing. Thus, we propose the Multimodal EmotionLines Dataset (MELD), an extension and enhancement of EmotionLines. MELD contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends. Each utterance is annotated with emotion and sentiment labels, and encompasses audio, visual and textual modalities. We propose several strong multimodal baselines and show the importance of contextual and multimodal information for emotion recognition in conversations. The full dataset is available for use at http:// this http URL.
Added
2026-09-24

Tensor Fusion Network for Multimodal Sentiment Analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, Louis-Philippe Morency
Why you should read this
Introduces the Tensor Fusion Network, an end-to-end model that integrates intra-modality and inter-modality dynamics across language, acoustic, and visual streams to advance multimodal sentiment analysis.
Multimodal sentiment analysis is an increasingly popular research area, which extends the conventional language-based definition of sentiment analysis to a multimodal setup where other relevant modalities accompany language. In this paper, we pose the problem of multimodal sentiment analysis as modeling intra-modality and inter-modality dynamics. We introduce a novel model, termed Tensor Fusion Network, which learns both such dynamics end-to-end. The proposed approach is tailored for the volatile nature of spoken language in online videos as well as accompanying gestures and voice. In the experiments, our model outperforms state-of-the-art approaches for both multimodal and unimodal sentiment analysis.
Added
2026-09-18

Multimodal Transformer for Unaligned Multimodal Language Sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, Ruslan Salakhutdinov
Why you should read this
Introduces the Multimodal Transformer to model asynchronous language, audio, and visual streams end-to-end via directional crossmodal attention without requiring explicit word-level alignment preprocessing.
Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors. However, two major challenges in modeling such multimodal human language time-series data exist: 1) inherent data non-alignment due to variable sampling rates for the sequences from each modality; and 2) long-range dependencies between elements across modalities. In this paper, we introduce the Multimodal Transformer (MulT) to generically address the above issues in an end-to-end manner without explicitly aligning the data. At the heart of our model is the directional pairwise crossmodal attention, which attends to interactions between multimodal sequences across distinct time steps and latently adapt streams from one modality to another. Comprehensive experiments on both aligned and non-aligned multimodal time-series show that our model outperforms state-of-the-art methods by a large margin. In addition, empirical analysis suggests that correlated crossmodal signals are able to be captured by the proposed crossmodal attention mechanism in MulT.
Added
2026-09-15

Multimodal Machine Learning: A Survey and Taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, Louis-Philippe Morency
Why you should read this
Establishes a comprehensive taxonomy for multimodal machine learning by structuring the field around five fundamental technical challenges: representation, translation, alignment, fusion, and co-learning.
Our experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors. Modality refers to the way in which something happens or is experienced and a research problem is characterized as multimodal when it includes multiple such modalities. In order for Artificial Intelligence to make progress in understanding the world around us, it needs to be able to interpret such multimodal signals together. Multimodal machine learning aims to build models that can process and relate information from multiple modalities. It is a vibrant multi-disciplinary field of increasing importance and with extraordinary potential. Instead of focusing on specific multimodal applications, this paper surveys the recent advances in multimodal machine learning itself and presents them in a common taxonomy. We go beyond the typical early and late fusion categorization and identify broader challenges that are faced by multimodal machine learning, namely: representation, translation, alignment, fusion, and co-learning. This new taxonomy will enable researchers to better understand the state of the field and identify directions for future research.
Added
2026-09-10
