Built independently by an author, for readers. Read the story and support ChapterPal

keyword

multimodal emotion recognition

Multimodal emotion recognition is a subfield of artificial intelligence and affective computing that aims to identify, interpret, and categorize human emotional states by processing and fusing information from multiple sensory and communication channels. Instead of relying on a single data source, these systems integrate complementary modalities—predominantly spoken or written language, acoustic features such as vocal tone and pitch, and visual cues including facial expressions and gestures. By modeling the interactions, correlations, and contextual dependencies across these diverse inputs, multimodal emotion recognition improves the accuracy and robustness of affective prediction, enabling more natural human-computer interaction, empathetic conversational systems, and nuanced analysis of interpersonal communication.

7 items

Decoupled Multimodal Distilling for Emotion Recognition

Decoupled Multimodal Distilling for Emotion Recognition

Yong Li, Yuanzhi Wang, Zhen Cui

Why you should read this

Proposes a decoupled multimodal distillation framework that separates representations into shared and modality-exclusive spaces and applies dynamic graph distillation to adaptively transfer knowledge across language, visual, and acoustic streams for more accurate emotion recognition.

Human multimodal emotion recognition (MER) aims to perceive human emotions via language, visual and acoustic modalities. Despite the impressive performance of previous MER approaches, the inherent multimodal heterogeneities still haunt and the contribution of different modalities varies significantly. In this work, we mitigate this issue by proposing a decoupled multimodal distillation (DMD) approach that facilitates flexible and adaptive crossmodal knowledge distillation, aiming to enhance the discriminative features of each modality. Specially, the representation of each modality is decoupled into two parts, i.e., modality-irrelevant/-exclusive spaces, in a self-regression manner. DMD utilizes a graph distillation unit (GD-Unit) for each decoupled part so that each GD can be performed in a more specialized and effective manner. A GD-Unit consists of a dynamic graph where each vertex represents a modality and each edge indicates a dynamic knowledge distillation. Such GD paradigm provides a flexible knowledge transfer manner where the distillation weights can be automatically learned, thus enabling diverse crossmodal knowledge transfer patterns. Experimental results show DMD consistently obtains superior performance than state-of-the-art MER methods. Visualization results show the graph edges in DMD exhibit meaningful distributional patterns w.r.t. the modality-irrelevant/-exclusive feature spaces. Codes are released at https://github.com/mdszwyz/DMD.

Added

2026-10-05

A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation

A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation

Xiaoheng Zhang, Yang Li

OrganizationsBeihang University

Why you should read this

Proposes CMCF-SRNet, a framework that integrates audio and text through locality-constrained cross-modal attention and graph-based semantic refinement to capture conversational context and speaker emotional inertia for emotion recognition in conversation.

Emotion recognition in conversation (ERC) has attracted enormous attention for its applications in empathetic dialogue systems. However, most previous researches simply concatenate multi-modal representations, leading to an accumulation of redundant information and a limited context interaction between modalities. Furthermore, they only consider simple contextual features ignoring semantic clues, resulting in an insufficient capture of the semantic coherence and consistency in conversations. To address these limitations, we propose a cross-modality context fusion and semantic refinement network (CMCF-SRNet). Specifically, we first design a cross-modal locality-constrained transformer to explore the multimodal interaction. Second, we investigate a graph-based semantic refinement transformer, which solves the limitation of insufficient semantic relationship information between utterances. Extensive experiments on two public benchmark datasets show the effectiveness of our proposed method compared with other state-of-the-art methods, indicating its potential application in emotion recognition. Our model is available at https://github.com/zxiaohen/CMCF-SRNet.

Added

2026-10-04

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, Yongbin Li

OrganizationsHarbin Institute of Technology

Why you should read this

Proposes UniMSE, a generative framework that unifies multimodal sentiment analysis and emotion recognition in conversation by combining label spaces, integrating acoustic and visual signals directly into a T5 backbone, and applying inter-modality contrastive learning.

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period. However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models. We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions. Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMOCAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.

Added

2026-09-26

A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

Wenjie Zheng, Jianfei Yu, Rui Xia, Shijin Wang

OrganizationsiFLYTEKNanjing University of Science and TechnologyState Key Laboratory of Cognitive Intelligence

Why you should read this

Proposes a two-stage multimodal framework that isolates the true speaker's face sequence from complex multi-party video scenes to accurately guide conversational emotion recognition via multi-task learning.

Multimodal Emotion Recognition in Multi-party Conversations (MERMC) has recently attracted considerable attention. Due to the complexity of visual scenes in multi-party conversations, most previous MERMC studies mainly focus on text and audio modalities while ignoring visual information. Recently, several works proposed to extract face sequences as visual features and have shown the importance of visual information in MERMC. However, given an utterance, the face sequence extracted by previous methods may contain multiple people’s faces, which will inevitably introduce noise to the emotion prediction of the real speaker. To tackle this issue, we propose a two-stage framework named Facial expression-aware Multimodal Multi-Task learning (FacialMMT). Specifically, a pipeline method is first designed to extract the face sequence of the real speaker of each utterance, which consists of multimodal face recognition, unsupervised face clustering, and face matching. With the extracted face sequences, we propose a multimodal facial expression-aware emotion recognition model, which leverages the frame-level facial emotion distributions to help improve utterance-level emotion recognition based on multi-task learning. Experiments demonstrate the effectiveness of the proposed FacialMMT framework on the benchmark MELD dataset. The source code is publicly released at https://github.com/NUSTM/FacialMMT.

Added

2026-09-26

Multimodal Machine Learning: A Survey and Taxonomy

Multimodal Machine Learning: A Survey and Taxonomy

Tadas Baltrušaitis, Chaitanya Ahuja, Louis-Philippe Morency

OrganizationsCarnegie Mellon University

Why you should read this

Establishes a comprehensive taxonomy for multimodal machine learning by structuring the field around five fundamental technical challenges: representation, translation, alignment, fusion, and co-learning.

Our experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors. Modality refers to the way in which something happens or is experienced and a research problem is characterized as multimodal when it includes multiple such modalities. In order for Artificial Intelligence to make progress in understanding the world around us, it needs to be able to interpret such multimodal signals together. Multimodal machine learning aims to build models that can process and relate information from multiple modalities. It is a vibrant multi-disciplinary field of increasing importance and with extraordinary potential. Instead of focusing on specific multimodal applications, this paper surveys the recent advances in multimodal machine learning itself and presents them in a common taxonomy. We go beyond the typical early and late fusion categorization and identify broader challenges that are faced by multimodal machine learning, namely: representation, translation, alignment, fusion, and co-learning. This new taxonomy will enable researchers to better understand the state of the field and identify directions for future research.

Added

2026-09-10