UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition
Guimin HuTing-En LinYi ZhaoGuangming LuYuchuan WuYongbin Li
Proposes UniMSE, a generative framework that unifies multimodal sentiment analysis and emotion recognition in conversation by combining label spaces, integrating acoustic and visual signals directly into a T5 backbone, and applying inter-modality contrastive learning.
Understanding human intent and affective state through automated systems is vital for advanced conversational interfaces and automated customer service. Currently, machine learning approaches divide this challenge into two distinct tasks: evaluating general sentiment over longer periods and identifying specific, short-term emotions during conversations. Most research treats these problems in isolation, which prevents models from sharing complementary insights across text, audio, and video modalities.
The article establishes a unified framework that combines multimodal sentiment analysis and conversational emotion recognition into a single generative architecture. The researchers designed this system to evaluate whether sharing knowledge between sentiment and emotion tasks improves predictive accuracy across multiple standard benchmarks.
To achieve this, the approach transforms the two separate tasks into a shared sequence-generation problem. Audio and video features were standardized across datasets, while a universal label scheme mapped sentiment polarities, intensities, and emotion categories into a shared format using sentence-level semantic matching. The architecture embeds audio and visual signals directly into intermediate layers of a standard language model and applies contrastive learning to pull corresponding modalities of the same sample closer while pushing unrelated samples apart. The framework was evaluated across four widely used multimodal benchmarks comprising thousands of conversational and video review segments.
The experimental findings show that the unified architecture consistently outperforms existing specialized models across all tested datasets. On sentiment tasks, the framework achieved an accuracy of 85.85% to 86.9% on one benchmark and 85.86% to 87.5% on another, reflecting gains of roughly 1.1% to 1.7% over previous state-of-the-art systems. For conversational emotion recognition, classification accuracy reached 65.09% and 70.56% across the two datasets, surpassing prior methods by approximately 2.3% to 2.6%. Ablation analyses confirmed that removing non-verbal modalities or omitting the multi-task training sets systematically degraded performance, highlighting the value of acoustic cues and cross-dataset knowledge sharing.
These results demonstrate that sentiment and emotion share a functional embedding space that enhances machine learning performance when modeled together. Consolidating separate analytical pipelines into a single model can reduce architectural complexity and training overhead. However, while the system establishes new benchmark performance, overall emotion recognition accuracy remains around 65% to 71%, which is insufficient for fully autonomous, high-stakes operational deployment.
Organizations developing affective computing systems should consider adopting unified architectures over isolated models to capture cross-modal efficiencies. Before deploying these models into production, practitioners should run targeted pilot programs and expand conversational context modeling across longer dialogues to improve baseline reliability.
The study notes limitations in its label completion process, which relied primarily on text similarity rather than acoustic or visual cues, and observed that broader context was not integrated across all evaluated datasets. Consequently, while the framework reliably advances foundational research, stakeholders should maintain human-in-the-loop oversight when deploying emotion detection systems in sensitive environments.
- Paper: MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, Soujanya Poria et al. (2018). Introduces the MELD benchmark for multimodal conversational emotion and sentiment recognition, which serves as one of the primary evaluation datasets unified by UniMSE.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). Introduces the CMU-MOSEI dataset and foundational multimodal fusion paradigms that UniMSE benchmarks and seeks to unify with conversational emotion recognition.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). Establishes modality-invariant and modality-specific representation learning with contrastive and orthogonal losses for multimodal sentiment analysis, directly informing UniMSE's contrastive embedding space.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). Develops cross-modal attention mechanisms for unaligned multimodal sequences, establishing the core technical basis for fusing acoustic and visual features into language models.
- Paper: Tensor Fusion Network for Multimodal Sentiment Analysis, Amir Zadeh et al. (2017). Provides the seminal Tensor Fusion Network architecture for modeling inter-modality dynamics in multimodal sentiment analysis evaluated on benchmark video datasets.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). Establishes the core taxonomy of representation, alignment, and fusion challenges in multimodal machine learning that UniMSE addresses through a unified generative model.
- Paper: VL-BERT: Pre-training of Generic Visual-Linguistic Representations, Weijie Su et al. (2019). Demonstrates how Transformer backbones can be pre-trained to embed cross-modal signals into intermediate layers, underpinning UniMSE's LM-based architecture.
- Paper: A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations, Wenjie Zheng et al. (2023). Extends multimodal conversational emotion recognition by addressing visual noise from non-speaking participants through active speaker facial expression extraction and multi-task learning.
- Paper: PMR: Prototypical Modal Rebalance for Multimodal Learning, Yunfeng Fan et al. (2023). Investigates and resolves the modality competition and imbalance issues inherent in joint multimodal training frameworks like UniMSE using prototypical rebalancing.
- Paper: ReconBoost: Boosting Can Achieve Modality Reconcilement, Cong Hua et al. (2024). Proposes an alternating boosting framework to prevent dominant modalities from suppressing weaker acoustic or visual representations during joint multimodal optimization.
- Paper: SECap: Speech Emotion Captioning with Large Language Model, Yaoxun Xu et al. (2024). Extends generative affective computing by using large language models to generate rich, descriptive natural language captions of vocal emotions rather than discrete label mappings.
- Paper: When did you become so smart, oh wise one?! Sarcasm Explanation in Multi-modal Multi-party Dialogues, Shivani Kumar et al. (2022). Applies multimodal sequence generation to complex conversational reasoning by generating natural language explanations for sarcasm across multi-party dialogues.
- Paper: SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words, Junyi Ao et al. (2024). Introduces a standardized benchmark to evaluate spoken dialogue comprehension and response generation conditioned on acoustic emotion and non-verbal cues beyond transcribed text.
