COGMEN: COntextualized GNN based Multimodal Emotion recognitioN
Abhinav JoshiAshwani BhatAyush JainAtin Vikram SinghAshutosh Modi
Proposes a contextualized graph neural network architecture that models both global conversational context and speaker dependencies across multiple modalities to achieve state-of-the-art multimodal emotion recognition on IEMOCAP and MOSEI.
Accurately identifying human emotions during multi-person conversations is essential for developing intuitive artificial intelligence systems, such as virtual digital assistants. However, conversational emotion recognition remains difficult because a speaker's emotional state fluctuates based on both the overarching conversation topic and the immediate back-and-forth interactions between participants. Additionally, human emotion is inherently multimodal, requiring systems to interpret complementary cues from text, speech audio, and facial video simultaneously.
The article introduces and evaluates a novel artificial intelligence architecture named COGMEN, which combines contextual language models and graph-based network processing. The main objective of the article is to demonstrate how simultaneously capturing full conversational context alongside immediate speaker-to-speaker interactions enhances multimodal emotion recognition.
To evaluate the system, the authors conducted empirical experiments using two benchmark conversational datasets: IEMOCAP, which contains multi-speaker dialogue videos categorized into emotional states, and CMU-MOSEI, a large-scale multimodal sentiment and emotion dataset. The approach combines a transformer network to extract global contextual features from concatenated audio, visual, and textual inputs with a graph neural network framework that explicitly maps internal speaker continuity and inter-speaker reactions across surrounding utterances.
The findings show that COGMEN establishes new performance benchmarks. On the IEMOCAP four-emotion classification task, the model achieved an 84.5% weighted F1-score, representing a 7.7 percentage point improvement over previous state-of-the-art approaches. On the IEMOCAP six-emotion benchmark, the system reached 68.2% accuracy and a 67.6% F1-score, outperforming existing multimodal models across several challenging categories. On CMU-MOSEI, COGMEN achieved the highest binary sentiment classification accuracy at 85.0% and outperformed competitive baselines across multi-label emotion tasks. Furthermore, ablation experiments confirmed that removing the relational graph structure or reducing conversational context significantly decreased model performance.
These results demonstrate that combining broad dialogue context with localized speaker-relationship modeling improves the reliability of emotion recognition in complex interactions. While previous multimodal models often suffered performance degradation when adding noisy visual data, the proposed graph architecture successfully leveraged multi-modal features without requiring complex, computationally expensive fusion schemes. This improves classification accuracy while maintaining a streamlined model training pipeline.
Before deploying this technology in operational environments, organizations should focus on several next steps. Because the current architecture relies on offline processing that analyzes both past and future utterances within a conversation, further research and development are needed to adapt the system for real-time, online streaming environments, such as live customer support or telecommunications. Implementing dynamic context buffers represents a promising direction to balance latency and accuracy. Additionally, future model iterations should incorporate dedicated mechanisms to better handle abrupt emotional transitions and distinguish between closely related emotional classes, such as excitement versus happiness.
- Paper: MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, Soujanya Poria et al. (2018). This paper establishes the foundational multi-party conversational emotion recognition benchmark (MELD) and demonstrates the critical need for modeling conversational context and speaker dependencies.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). It introduces the crossmodal transformer paradigm for aligning and fusing unaligned multimodal sequences across standard affective datasets like IEMOCAP and MOSI.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). This work formulates the separation of multimodal inputs into modality-invariant and modality-specific representations, providing key representational techniques leveraged in conversational affect modeling.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). It provides foundational methodology and benchmark data for modeling multi-channel cross-modal dynamics using graph-based memory structures across sentiment and emotion tasks.
- Paper: Tensor Fusion Network for Multimodal Sentiment Analysis, Amir Zadeh et al. (2017). This paper introduces tensor-based multimodal fusion, establishing the core mechanics of capturing multi-way interdependencies across language, visual, and acoustic features.
- Paper: The Graph Neural Network Model, Franco Scarselli et al. (2009). It establishes the foundational Graph Neural Network model and message-passing theory necessary for understanding relational context modeling over conversational graphs.
- Paper: A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations, Wenjie Zheng et al. (2023). This work extends multimodal multi-party conversational emotion recognition by addressing visual noise from non-speaking participants through active speaker facial expression extraction and multi-task learning.
- Paper: UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition, Guimin Hu et al. (2022). It unifies conversational emotion recognition and multimodal sentiment analysis into a single generative framework using cross-modal contrastive learning.
- Paper: SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words, Junyi Ao et al. (2024). It broadens conversational emotion and non-verbal understanding beyond classification by establishing a benchmark for end-to-end spoken dialogue response generation.
