Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
Cam-Van Thi NguyenAnh-Tuan MaiThe-Son LeHai-Dang KieuDuc-Trong Le
Proposes a relational temporal graph neural network framework, CORECT, that jointly models utterance-level temporal dependencies and conversation-level cross-modal interactions while preserving modality-specific representations to achieve state-of-the-art multimodal emotion recognition on IEMOCAP and CMU-MOSEI.
Emotion recognition in spoken and text conversations is essential for developing intelligent, human-aware systems. Real-world human interaction relies on multiple signals—including text, vocal tone, and visual expressions—that evolve dynamically over time. Most existing systems struggle to accurately understand these emotions because they either merge different sensory streams too early into a single representation, discarding modality-specific nuances, or they fail to properly capture how past and future utterances influence the current emotional state.
The article demonstrates a novel framework called CORECT, designed to improve multimodal emotion recognition by jointly modeling local temporal dependencies and global cross-modality interactions without prematurely collapsing individual modality features.
To evaluate this framework, the authors conducted extensive experiments on two benchmark conversation datasets: IEMOCAP, which includes 151 dyadic dialogues across 7,433 utterances, and CMU-MOSEI, which contains over 22,000 utterances. The method extracts separate features for text, audio, and video, models their relationships and time order using a relational temporal graph network, and enriches these with a pairwise cross-modal interaction module. The system's performance was compared against leading baseline models across standard emotion and sentiment classification metrics.
The framework achieved new state-of-the-art results across both benchmarks. On the six-class IEMOCAP benchmark, CORECT improved overall accuracy by 2.89% and weighted F1-score by 2.75% over the previous best-performing model, reaching an accuracy of 69.93%. On the four-class version, it achieved an accuracy of 84.73%, representing a 2.44% improvement. Ablation analyses confirmed that removing the relational temporal graph module caused the largest performance decline (up to 4.10%), proving that local structural and temporal modeling is critical. Furthermore, contextual analysis showed an asymmetrical time effect: past utterances exert a significantly stronger influence on current emotional state than future utterances, with an optimal window of eleven past utterances versus nine future ones.
These findings demonstrate that retaining separate modality representations while explicitly modeling how conversation history unfolds yields substantial gains in conversational understanding. For organizations building conversational agents, customer service analytics, or affective monitoring tools, adopting architecture that explicitly tracks temporal flow and cross-modal dependencies improves classification reliability, particularly in identifying difficult or underrepresented emotions like fear and surprise that standard models fail to distinguish.
For future development, organizations and researchers should focus on implementing automated hyperparameter tuning to optimize past and future contextual window sizes dynamically. Additionally, developers should explore adaptive attention mechanisms that selectively weight earlier utterances rather than relying on fixed sliding windows, ensuring computational efficiency during deployment in real-time conversational systems.
Confidence in the findings is high based on consistent improvements across diverse datasets and rigorous module-by-module ablation. However, decision-makers should note that performance remains lower for ambiguous emotion pairs—such as distinguishing excitement from happiness or frustration from sadness—and visual input quality remains susceptible to real-world noise such as camera angle and lighting variations.
- Paper: COGMEN: COntextualized GNN based Multimodal Emotion recognitioN, Abhinav Joshi et al. (2022). COGMEN establishes a closely related contextual graph approach to multimodal conversational emotion recognition, making CORECT’s temporal and relational design easier to situate.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). This survey’s taxonomy of multimodal representation, alignment, and fusion provides the conceptual groundwork for understanding CORECT’s separate modality features and cross-modal interactions.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). The Multimodal Transformer introduces cross-modal attention for unaligned language, audio, and video streams, clarifying a key family of methods behind CORECT’s modality interactions.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). This paper introduces CMU-MOSEI, one of CORECT’s evaluation benchmarks, and its dynamic fusion graph offers useful context for modeling interactions across modalities.
- Paper: MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, Soujanya Poria et al. (2018). MELD establishes a foundational multimodal conversational emotion dataset and baseline context for interpreting the benchmark design used in this research area.
- Paper: Tensor Fusion Network for Multimodal Sentiment Analysis, Amir Zadeh et al. (2017). Tensor Fusion Network explains an influential approach to representing interactions among language, audio, and visual features, helping motivate CORECT’s explicit cross-modal modeling.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). MISA’s separation of shared and modality-specific representations provides relevant groundwork for understanding why CORECT preserves distinct modality features during fusion.
No sufficiently relevant recommendations were found.
