A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations
Wenjie ZhengJianfei YuRui XiaShijin Wang
Proposes a two-stage multimodal framework that isolates the true speaker's face sequence from complex multi-party video scenes to accurately guide conversational emotion recognition via multi-task learning.
Emotion recognition in multi-party conversations is an essential capability for conversational artificial intelligence, customer service analytics, and interactive video systems. While human communication naturally relies on spoken words, vocal tone, and visual cues, existing automated systems often struggle to leverage video effectively. In multi-person scenes, background environments and non-speaking individuals frequently introduce visual noise, misleading models when predicting the primary speaker's actual emotion.
The article addresses this challenge by developing and evaluating FacialMMT, a two-stage multimodal framework designed to isolate the true speaker's facial expressions and use them to enhance conversation-level emotion recognition across video, audio, and text.
The approach first extracts the exact facial sequence of the active speaker through active speaker detection, rule-based lip-movement analysis, unsupervised graph clustering, and face matching against a reference library. In the second stage, the framework employs a multi-task learning architecture that predicts frame-by-frame facial emotion distributions using an auxiliary dynamic facial expression dataset, filters out ambiguous frames via a gating mechanism, and fuses the clarified visual features with spoken and textual representations using cross-modal attention transformers.
Evaluations on the benchmark conversational dataset demonstrate that the framework achieves state-of-the-art performance. The system reached an overall conversation emotion recognition score of 66.58 percent, outperforming leading baseline systems. When evaluating visual data alone, isolating the real speaker's face yielded a recognition score of 36.48 percent, noticeably exceeding prior methods that relied on full video frames (31.26 to 32.34 percent) or mixed-speaker face sequences (33.27 percent). Ablation analysis confirmed that facial visual cues provided greater marginal value to overall system accuracy than audio features, and removing the auxiliary frame-level emotion learning caused performance to decline.
These findings indicate that the visual modality is far more informative for conversational emotion recognition than previously assumed, provided that non-speaking faces and scene distractions are cleanly filtered out. For organizations deploying conversational AI, improving speaker isolation and incorporating frame-level expression tracking offers a reliable path to higher accuracy without requiring completely new model architectures.
Organizations considering implementation should evaluate two-stage speaker isolation pipelines for multi-party video processing while planning for technical refinements. Before enterprise deployment, technical teams should explore unified, end-to-end architectures to reduce operational overhead, test across more diverse demographic datasets to mitigate demographic bias, and implement strong privacy safeguards around facial data processing.
Confidence in these findings is high for conversational video benchmarks structured similarly to the evaluated dataset. However, stakeholders should exercise caution because the framework relies on a multi-stage pipeline rather than an end-to-end system and depends on reference facial libraries that may require adaptation in open-world environments where speaker identities are entirely unconstrained.
- Paper: MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, Soujanya Poria et al. (2018). This paper establishes the foundational MELD benchmark and conversational emotion recognition setup that FacialMMT directly uses and builds upon.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). It introduces the crossmodal attention transformer mechanism that underpins FacialMMT's fusion of spoken, textual, and isolated visual streams.
- Paper: OpenFace: An open source facial behavior analysis toolkit, Tadas Baltrusaitis et al. (2016). It presents essential open-source tools and pipelines for real-time facial landmark detection, action unit analysis, and gaze tracking crucial for speaker face processing.
- Paper: FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos, Yan Wang et al. (2022). It provides a comprehensive benchmark for video facial expression recognition across dynamic scenes, establishing foundational challenges in video affect processing.
- Paper: Deep Facial Expression Recognition: A Survey, Shan Li et al. (2018). This survey provides essential background on deep learning techniques and spatio-temporal architectures for facial expression recognition in uncontrolled conditions.
- Paper: AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild, Ali Mollahosseini et al. (2017). It introduces large-scale in-the-wild facial affect datasets and models, establishing the frame-level emotion recognition paradigms utilized in multi-task frameworks.
- Paper: UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition, Guimin Hu et al. (2022). It formulates joint multi-task representation learning for conversational emotion recognition and sentiment analysis across multimodal streams.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). It establishes principles for disentangling modality-invariant and modality-specific representations in multimodal sentiment and emotion pipelines.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). This survey outlines the overarching taxonomy of multimodal representation, cross-modal alignment, and fusion strategies leveraged by conversational AI frameworks.
- Paper: ReconBoost: Boosting Can Achieve Modality Reconcilement, Cong Hua et al. (2024). This work introduces alternating boosting techniques to address modality competition and imbalance, advancing beyond standard cross-modal transformer fusion.
- Paper: SECap: Speech Emotion Captioning with Large Language Model, Yaoxun Xu et al. (2024). It extends discrete conversational emotion classification toward open-ended natural language descriptions and emotion captioning powered by large language models.
- Paper: Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation, Dongjin Kang et al. (2024). It builds upon conversational emotion understanding by deploying large language models as multi-stage empathetic conversational agents in emotional support dialogues.
- Paper: Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, Chaoyou Fu et al. (2025). It establishes a comprehensive evaluation benchmark for multimodal LLMs handling long-form dynamic audio-visual reasoning across full conversational and scene contexts.
- Paper: LLaVA-OneVision: Easy Visual Task Transfer, Bo Li et al. (2024). It generalizes multi-stage visual extraction pipelines into a unified multimodal foundation model capable of processing continuous conversational video and audio.
- Paper: Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling, Zhe Chen et al. (2024). It scales vision-language architectures and test-time reasoning to handle multi-image and video comprehension in complex conversational settings.
- Paper: Qwen3.5-Omni Technical Report, Qwen Team (2026). It synthesizes multimodal perception and conversation into an end-to-end omnimodal framework that unifies real-time audio-visual understanding and speech interaction.
