A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

Wenjie ZhengJianfei YuRui XiaShijin Wang

article2023ACL79 citations

Proposes a two-stage multimodal framework that isolates the true speaker's face sequence from complex multi-party video scenes to accurately guide conversational emotion recognition via multi-task learning.

Listen

Emotion recognition in multi-party conversations is an essential capability for conversational artificial intelligence, customer service analytics, and interactive video systems. While human communication naturally relies on spoken words, vocal tone, and visual cues, existing automated systems often struggle to leverage video effectively. In multi-person scenes, background environments and non-speaking individuals frequently introduce visual noise, misleading models when predicting the primary speaker's actual emotion.

The article addresses this challenge by developing and evaluating FacialMMT, a two-stage multimodal framework designed to isolate the true speaker's facial expressions and use them to enhance conversation-level emotion recognition across video, audio, and text.

The approach first extracts the exact facial sequence of the active speaker through active speaker detection, rule-based lip-movement analysis, unsupervised graph clustering, and face matching against a reference library. In the second stage, the framework employs a multi-task learning architecture that predicts frame-by-frame facial emotion distributions using an auxiliary dynamic facial expression dataset, filters out ambiguous frames via a gating mechanism, and fuses the clarified visual features with spoken and textual representations using cross-modal attention transformers.

Evaluations on the benchmark conversational dataset demonstrate that the framework achieves state-of-the-art performance. The system reached an overall conversation emotion recognition score of 66.58 percent, outperforming leading baseline systems. When evaluating visual data alone, isolating the real speaker's face yielded a recognition score of 36.48 percent, noticeably exceeding prior methods that relied on full video frames (31.26 to 32.34 percent) or mixed-speaker face sequences (33.27 percent). Ablation analysis confirmed that facial visual cues provided greater marginal value to overall system accuracy than audio features, and removing the auxiliary frame-level emotion learning caused performance to decline.

These findings indicate that the visual modality is far more informative for conversational emotion recognition than previously assumed, provided that non-speaking faces and scene distractions are cleanly filtered out. For organizations deploying conversational AI, improving speaker isolation and incorporating frame-level expression tracking offers a reliable path to higher accuracy without requiring completely new model architectures.

Organizations considering implementation should evaluate two-stage speaker isolation pipelines for multi-party video processing while planning for technical refinements. Before enterprise deployment, technical teams should explore unified, end-to-end architectures to reduce operational overhead, test across more diverse demographic datasets to mitigate demographic bias, and implement strong privacy safeguards around facial data processing.

Confidence in these findings is high for conversational video benchmarks structured similarly to the evaluated dataset. However, stakeholders should exercise caution because the framework relies on a multi-stage pipeline rather than an end-to-end system and depends on reference facial libraries that may require adaptation in open-world environments where speaker identities are entirely unconstrained.

Cover for A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

Abstract

Multimodal Emotion Recognition in Multi-party Conversations (MERMC) has recently attracted considerable attention. Due to the complexity of visual scenes in multi-party conversations, most previous MERMC studies mainly focus on text and audio modalities while ignoring visual information. Recently, several works proposed to extract face sequences as visual features and have shown the importance of visual information in MERMC. However, given an utterance, the face sequence extracted by previous methods may contain multiple people’s faces, which will inevitably introduce noise to the emotion prediction of the real speaker. To tackle this issue, we propose a two-stage framework named Facial expression-aware Multimodal Multi-Task learning (FacialMMT). Specifically, a pipeline method is first designed to extract the face sequence of the real speaker of each utterance, which consists of multimodal face recognition, unsupervised face clustering, and face matching. With the extracted face sequences, we propose a multimodal facial expression-aware emotion recognition model, which leverages the frame-level facial emotion distributions to help improve utterance-level emotion recognition based on multi-task learning. Experiments demonstrate the effectiveness of the proposed FacialMMT framework on the benchmark MELD dataset. The source code is publicly released at https://github.com/NUSTM/FacialMMT.

Table of Contents

  • 1 Introduction
  • 2 Method
  • 2.1 Task Formulation
  • 2.2 Framework Overview
  • 2.3 Face Sequence Extraction
  • 2.4 A Multimodal Facial Expression-Aware Multi-Task Learning Model
  • 2.4.1 Unimodal Feature Extraction
  • 2.4.2 Emotion-Aware Visual Representation
  • 2.4.3 Multimodal Fusion
  • 3 Experiments and Analysis
  • 3.1 Dataset
  • 3.2 Compared Systems
  • 3.3 Implementation
  • 3.4 Main Results on the MERMC task
  • 3.5 Results on the DFER task
  • 3.6 Ablation Study
  • 3.7 Case Study
  • 4 Related Work
  • 4.1 Emotion Recognition in Conversations
  • 4.2 Dynamic Facial Expression Recognition
  • 5 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Multimodal Rules
  • A.2 Pseudo-code of MARIO

Knowls

  1. Knowl 1 — FacialMMT Framework for Multimodal Emotion Recognition in Multi-party Conversations

    model/method

    The Facial Expression-Aware Multimodal Multi-Task learning (FacialMMT) framework is a two-stage approach for Multimodal Emotion Recognition in Multi-party Conversations (MERMC). In MERMC, given a dialogue d={u1,u2,…,un}d = \{u_1, u_2, \dots, u_n\} of nn utterances where each utterance ui={uil,uia,uiv}u_i = \{u_{il}, u_{ia}, u_{iv}\} comprises text ll, audio aa, and visual vv modalities, the objective is to classify each utterance uiu_i into one of CC emotion classes yi∈Yy_i \in \mathcal{Y} corresponding to the specific active speaker of that utterance.

    FacialMMT operates in two sequential stages:

    1. Stage 1 (Face Sequence Extraction): Isolate the face sequence corresponding exclusively to the annotated real speaker for each utterance uiu_i. This stage combines active speaker detection and rule-based multimodal heuristics, unsupervised graph clustering with the InfoMap algorithm, and reference-based face cosine matching to eliminate extraneous background faces and non-speaking conversational participants.
    2. Stage 2 (Multimodal Facial Expression-Aware Multi-Task Learning - MARIO): Enhances visual representations by integrating frame-level facial emotion distributions from an auxiliary Dynamic Facial Expression Recognition (DFER) task. These frame-level distributions are relaxed differentiably via Gumbel-Softmax, filtered via an emotion clarity gating mechanism, and fused with textual and acoustic features through intra-modal self-attention and Cross-Modal Transformers (CMT) to predict the utterance emotion sequence y={y1,y2,…,yn}y = \{y_1, y_2, \dots, y_n\}.
  2. Knowl 2 — Real-Speaker Face Sequence Extraction via Multimodal Recognition, Clustering, and Matching

    model/method

    To extract the visual face sequence belonging strictly to the real speaker of an utterance in multi-party conversation videos, a three-step pipeline is employed:

    1. Multimodal Face Recognition (MFR): A pre-trained active speaker detection (ASD) network (TalkNet) detects speaker candidates across video frames by evaluating audio-visual synchrony. For short or noisy video segments where TalkNet fails, heuristic multimodal rules are applied, measuring mouth open/close frequency, lip coordinate displacement across consecutive frames, and voice activity alignment.
    2. Unsupervised Graph Clustering (UC): Potential speaker faces are organized into a KK-nearest neighbor graph where edge weights correspond to normalized facial visual similarities. Random walks and the InfoMap community detection algorithm are run to partition faces into KK distinct facial sequence clusters by minimizing the description length of information flow.
    3. Face Matching (FM): A reference library consisting of 20 representative face images per recurring leading role is constructed and embedded using a ResNet-50 network pre-trained on MS-Celeb-1M. The average embedding of each extracted face cluster is compared against the reference embeddings using cosine similarity. If the annotated speaker label is a known leading character, the cluster exhibiting the highest cosine similarity is selected; if the annotated speaker is a non-recurring passerby, the cluster exhibiting the lowest similarity to the leading character library is retained.
  3. Knowl 3 — Objective Function for InfoMap Face Sequence Clustering

    equation

    In the unsupervised face clustering stage of FacialMMT, the InfoMap algorithm groups detected face crops into coherent individual face sequences by minimizing the two-level map equation describing the minimum average encoding length of a random walk across the facial similarity graph:

    arg⁡min⁡K,YL(P,K,Y)=q↷(−∑i=1Kqi↷q↷log⁡qi↷q↷)+∑i=1Kpi↺(−qi↷pi↺log⁡qi↷pi↺−∑α∈ipαpi↺log⁡pαpi↺)\arg\min_{K, Y} \mathcal{L}(P, K, Y) = q_{\curvearrowright} \left( - \sum_{i=1}^K \frac{q_{i\curvearrowright}}{q_{\curvearrowright}} \log \frac{q_{i\curvearrowright}}{q_{\curvearrowright}} \right) + \sum_{i=1}^K p_{i\circlearrowleft} \left( - \frac{q_{i\curvearrowright}}{p_{i\circlearrowleft}} \log \frac{q_{i\curvearrowright}}{p_{i\circlearrowleft}} - \sum_{\alpha \in i} \frac{p_\alpha}{p_{i\circlearrowleft}} \log \frac{p_\alpha}{p_{i\circlearrowleft}} \right)

    where:

    • K∈N+K \in \mathbb{N}^+ is the number of predicted face sequence clusters.
    • YY represents the predicted face sequence assignment.
    • qi↷∈[0,1]q_{i\curvearrowright} \in [0, 1] is the transition probability of exiting class/cluster ii.
    • q↷=∑i=1Kqi↷q_{\curvearrowright} = \sum_{i=1}^K q_{i\curvearrowright} is the total exit probability across all clusters.
    • pα∈[0,1]p_\alpha \in [0, 1] is the stationary ergodic probability of visiting face image node α\alpha.
    • pi↺=qi↷+∑α∈ipαp_{i\circlearrowleft} = q_{i\curvearrowright} + \sum_{\alpha \in i} p_\alpha is the total probability of movement within or exiting cluster ii.
  4. Knowl 4 — Emotion-Aware Visual Representation via Differentiable Gumbel-Softmax and Clarity Gating

    model/method

    To enrich frame-level facial features with emotional semantics without breaking end-to-end gradient propagation, the MARIO model leverages an auxiliary Dynamic Facial Expression Recognition (DFER) branch and an emotion clarity gating mechanism:

    1. Feature Extraction: For an extracted face sequence s={s1,…,sm}s = \{s_1, \dots, s_m\} of length mm, visual spatial representations Hs={h1s,…,hms}∈Rm×duH^s = \{h_1^s, \dots, h_m^s\} \in \mathbb{R}^{m \times d_u} are extracted using a Swin Transformer pre-trained on MS-Celeb-1M. Unimodal baseline visual features Ev∈RL×dvE_v \in \mathbb{R}^{L \times d_v} (dv=512d_v = 512) are extracted using InceptionResNetv1 pre-trained on CASIA-WebFace.
    2. Differentiable Expression Distribution: To avoid non-differentiable argmax\text{argmax} hard assignments when passing predicted frame emotion distributions to the MERMC task, the Gumbel-Softmax distribution parameterized by temperature τ\tau is applied: gi=softmax(g+hisτ)g_i = \text{softmax}\left( \frac{g + h_i^s}{\tau} \right) where gi∈RCg_i \in \mathbb{R}^C, g=−log⁡(−log⁡(u))g = -\log(-\log(u)), and u∼Uniform(0,1)u \sim \text{Uniform}(0, 1). As τ→0\tau \to 0, gig_i smoothly approximates an indicator one-hot emotion distribution.
    3. Clarity Gating and Filtering: Emotion clarity for frame ii is quantified by the self-dot product δi=gi⋅gi⊤∈[0,1]\delta_i = g_i \cdot g_i^\top \in [0, 1]. Frames with δi<0.2\delta_i < 0.2 (uniform, ambiguous emotion distributions) are filtered out, leaving m′m' clear frames with visual representations Ev′∈Rm′×dvE'_v \in \mathbb{R}^{m' \times d_v}.
    4. Emotion-Aware Representation: The filtered frame features Ev′E'_v and their corresponding frame emotion distributions Ee={g1,…,gm′}E_e = \{g_1, \dots, g_{m'}\} are concatenated along feature dimensions: E^v=Ev′⊕Ee∈Rm′×(dv+C)\hat{E}_v = E'_v \oplus E_e \in \mathbb{R}^{m' \times (d_v + C)}
  5. Knowl 5 — MARIO Multimodal Fusion and Multi-Task Optimization

    model/method

    The MARIO model integrates textual, acoustic, and emotion-aware visual features through hierarchical self-attention and cross-attention, optimized jointly with an auxiliary expression recognition loss:

    1. Unimodal Representations:
      • Text representation El∈R512E_l \in \mathbb{R}^{512} is obtained from the [CLS][\text{CLS}] token representation of a pre-trained language model (BERT or RoBERTa) fed with the target utterance and its dialogue history.
      • Audio representation Ea∈RdaE_a \in \mathbb{R}^{d_a} (da=768d_a = 768) is obtained from a Wav2vec 2.0 encoder pre-trained on LibriSpeech-960h.
      • Emotion-aware visual representation E^v∈Rm′×(512+C)\hat{E}_v \in \mathbb{R}^{m' \times (512+C)} is obtained via clarity-gated Gumbel-Softmax expression pooling.
    2. Intra-Modal Modeling: Audio and visual features are contextualized through separate 12-head Transformer self-attention layers: Ha=Transformer(Ea),Hv=Transformer(E^v)H_a = \text{Transformer}(E_a), \quad H_v = \text{Transformer}(\hat{E}_v)
    3. Inter-Modal Fusion: Cross-Modal Transformers (CMT) fuse the modalities hierarchically: Hl−a=CM-Transformer(El,Ha)H_{l-a} = \text{CM-Transformer}(E_l, H_a) Hl−a−v=CM-Transformer(Hl−a,Hv)H_{l-a-v} = \text{CM-Transformer}(H_{l-a}, H_v)
    4. Classification & Objectives: Utterance-level predictions are generated via q(y)=softmax(W⊤Hl−a−v+b)q(y) = \text{softmax}(W^\top H_{l-a-v} + b). Training alternates between the DFER auxiliary loss LDFER=−1M∑i=1M∑j=1mlog⁡p(zij)\mathcal{L}_{\text{DFER}} = -\frac{1}{M}\sum_{i=1}^M \sum_{j=1}^m \log p(z_{ij}) on DFER data and the MERMC cross-entropy loss LMERMC=−1N∑i=1Nlog⁡q(yi)\mathcal{L}_{\text{MERMC}} = -\frac{1}{N}\sum_{i=1}^N \log q(y_i) on conversation data.
  6. Knowl 6 — MARIO Multi-Task Training Algorithm

    algorithm

    The training procedure of the MARIO model alternates batch updates across the auxiliary Dynamic Facial Expression Recognition (DFER) dataset and the main Multimodal Emotion Recognition in Multi-party Conversations (MERMC) dataset.

    Algorithm: Multitask Training Procedure of MARIO
    Input: Auxiliary DFER dataset Ds\mathcal{D}^s, Main MERMC dataset D\mathbb{D}
    Output: Trained parameters θSwin\theta_{\text{Swin}} (Swin-Transformer), θT\theta_T (Text encoder), θself-attn\theta_{\text{self-attn}} (Intra-modal Transformers), θCMT\theta_{\text{CMT}} (Cross-Modal Transformers)
    repeat
        for each batch in Ds\mathcal{D}^s do
            Forward face sequences through Swin-Transformer
            Compute auxiliary expression loss LDFER=−1M∑i=1M∑j=1mlog⁡p(zij)\mathcal{L}_{\text{DFER}} = -\frac{1}{M}\sum_{i=1}^M \sum_{j=1}^m \log p(z_{ij})
            Update θSwin\theta_{\text{Swin}} using gradient ∇LDFER\nabla \mathcal{L}_{\text{DFER}}
        end for
        for each batch in D\mathbb{D} do
            Forward conversation context through text encoder to get ElE_l
            Forward speaker face sequences through Swin-Transformer to get frame expressions gig_i
            Filter frames below clarity threshold δi<0.2\delta_i < 0.2 and construct E^v=Ev′⊕Ee\hat{E}_v = E'_v \oplus E_e
            Process EaE_a and E^v\hat{E}_v via intra-modal self-attention to obtain HaH_a and HvH_v
            Fuse text and audio: Hl−a=CM-Transformer(El,Ha)H_{l-a} = \text{CM-Transformer}(E_l, H_a)
            Fuse text-audio and visual: Hl−a−v=CM-Transformer(Hl−a,Hv)H_{l-a-v} = \text{CM-Transformer}(H_{l-a}, H_v)
            Compute MERMC task loss LMERMC=−1N∑i=1Nlog⁡q(yi)\mathcal{L}_{\text{MERMC}} = -\frac{1}{N}\sum_{i=1}^N \log q(y_i)
            Update θself-attn\theta_{\text{self-attn}}, θCMT\theta_{\text{CMT}}, and fine-tune θSwin\theta_{\text{Swin}}, θT\theta_T using ∇LMERMC\nabla \mathcal{L}_{\text{MERMC}}
        end for
    until maximum training epochs reached
  7. Knowl 7 — Multimodal Active Speaker Detection Heuristic Rules

    algorithm

    When pre-trained active speaker detection (ASD) fails due to short video duration or acoustic noise, the following heuristic algorithm extracts candidate speaker face sequences using facial landmarks and audio-visual correlation:

    Algorithm: Rule-Based Multimodal Speaker Face Identification
    Input: Video clip VV, Audio track AA
    Output: Candidate speaker face sequences
    Sample video frames using FFmpeg
    Extract audio waveform from VV using FFmpeg
    Detect all faces across frames using OpenFace, outputting FaceIDs, 68 facial landmark coordinates, and aligned face crops
    for each unique FaceID candidate do
        Compute Mouth Open-Close Count: evaluate distance between upper and lower lip landmarks; record frame as mouth-open if vertical lip separation exceeds threshold
        Compute Mouth Movement: measure total inter-frame displacement by summing the absolute differences of inner corner mouth width and inner lip vertical separation between consecutive frames
        Compute Audio-Visual Synchrony: evaluate correlation between vertical lip motion velocity and audio signal envelope energy (Voice Activity Detection matching)
    end for
    Rank FaceID candidates by composite score of open-close frequency, movement magnitude, and audio synchrony to identify active speaker sequence
  8. Knowl 8 — Multimodal Emotion Recognition Performance on MELD Benchmark

    data/table

    Performance of FacialMMT compared with baseline models on the MELD conversation dataset (13,707 video clips, 7 emotion classes). Results are reported as class-wise F1 scores and overall weighted average F1 score (%). Baselines include text-only models (italics) and multimodal models utilizing BERT (⋆^\star), RoBERTa (∗^*), or T5 (▲^\blacktriangle).

    Models Neutral Surprise Fear Sadness Joy Disgust Anger Weighted F1
    DialogueRNN 73.50 49.40 1.20 23.80 50.70 1.70 41.50 57.03
    ConGCN 76.70 50.30 8.70 28.50 53.10 10.60 46.80 59.40
    MMGCN - - - - - - - 58.65
    DialogueTRM∗^* - - - - - - - 63.50
    DAG-ERC∗^* - - - - - - - 63.65
    MM-DFN 77.76 50.69 - 22.94 54.78 - 47.82 59.46
    EmoCaps⋆^\star 77.12 63.19 3.03 42.52 57.50 7.69 57.54 64.00
    UniMSE▲^\blacktriangle - - - - - - - 65.51
    GA2MIF 76.92 49.08 - 27.18 51.87 - 48.52 58.94
    FacialMMT-BERT 78.55 58.17 13.04 38.51 61.10 30.30 53.66 64.69
    FacialMMT-RoBERTa 80.13 59.63 19.18 41.99 64.88 18.18 56.00 66.58

    FacialMMT-RoBERTa achieves the state-of-the-art weighted F1 of 66.58%, outperforming the previous best multimodal system UniMSE (65.51% F1) and showing substantial improvements on low-resource minority emotion classes such as Fear (19.18% F1) and Disgust (18.18% F1).

  9. Knowl 9 — Visual Modality Emotion Recognition and Extraction Ablation on MELD

    data/table

    Comparison of isolated visual-modality emotion recognition weighted F1 scores (%) on MELD across different visual feature compositions and ablation stages of the face sequence extraction pipeline.

    Model Composition of Visual Information Weighted F1
    EmoCaps Raw video frames (3D-CNN) 31.26
    MM-DFN Raw video frames (3D-CNN) 32.34
    MMGCN All detected speakers' face sequences 33.27
    FacialMMT Real speaker's face sequence 36.48
    – w/o UC, FM Face sequence from MFR alone 34.36
    – w/o MFR, UC, FM Raw video frames 32.27

    Where MFR denotes Multimodal Face Recognition, UC denotes Unsupervised Clustering (InfoMap), and FM denotes Face Matching. Extracting solely the real speaker's face sequence yields 36.48% F1, surpassing generic video frame features (32.27%) and unfiltered multi-speaker faces (33.27%), demonstrating that removing background and non-speaking facial noise directly improves visual emotion classification.

  10. Knowl 10 — Modality and Component Ablation Analysis in FacialMMT

    data/table

    Ablation study evaluating the contribution of individual modalities and the auxiliary Dynamic Facial Expression Recognition (DFER) task within FacialMMT on the MELD dataset, evaluated by weighted average F1 score (%):

    Configuration Weighted F1 (%)
    FacialMMT (Full Model) 66.58
    – w/o Audio 66.20
    – w/o Vision 65.55
    – w/o Audio, Vision (Text-only) 63.98
    – w/o Text, Vision (Audio-only) 38.02
    – w/o Auxiliary DFER Module 66.08

    Removing the visual modality causes a larger performance degradation (−1.03%-1.03\%) than removing the acoustic modality (−0.38%-0.38\%), indicating that precise real-speaker facial representations are more informative than audio in MERMC. Furthermore, removing the auxiliary DFER supervision decreases F1 by 0.50%0.50\%, verifying that frame-level facial expression signals provide beneficial inductive bias for utterance-level conversation emotion recognition.

  11. Knowl 11 — Limitations of the FacialMMT Framework

    limitation

    The FacialMMT framework possesses two main limitations:

    1. Decoupled Two-Stage Architecture: The framework is not fully end-to-end. Face sequence extraction (active speaker detection, InfoMap clustering, and cosine face matching) is performed as a pre-processing pipeline independent of the multimodal emotion recognition model, preventing joint optimization of visual extraction and emotion classification.
    2. Visual-Centric Scope: The approach primarily focuses on enhancing visual modality representation through facial expression awareness and filtering, leaving the exploration of advanced cross-modal alignment mechanisms or enhanced acoustic/textual representation modules unaddressed.

Coverage note — None was omitted; all contributed models, equations, algorithms, experimental evaluations on MELD and Aff-Wild2, ablation studies, and stated limitations are fully represented.

References

  1. 1.Shahin Amiriparian, Lukas Christ, Andreas König, Eva-Maria Meßner, Alan Cowen, Erik Cambria, and Björn W Schuller. 2022. Muse 2022 challenge: Multimodal humour, emotional reactions, and stress. In Proceedings of ACM MM.
  2. 2.Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020. wav2vec 2.0: A framework for self-supervised learning of speech representations. Proceedings of NeurIPS.
  3. 3.Tadas Baltrušaitis, Peter Robinson, and Louis-Philippe Morency. 2016. Openface: an open source facial behavior analysis toolkit. In Proceedings of WACV.
  4. 4.Feiyu Chen, Zhengxiao Sun, Deqiang Ouyang, Xueliang Liu, and Jie Shao. 2021. Learning what and when to drop: Adaptive multimodal and contextual dynamics for emotion recognition in conversation. In Proceedings of ACM MM.
  5. 5.Wenliang Dai, Samuel Cahyawijaya, Zihan Liu, and Pascale Fung. 2021. Multimodal end-to-end sparse model for emotion recognition. In Proceedings of NAACL-HLT.
  6. 6.Kevin Delgado, Juan Manuel Origgi, Tania Hasanpoor, Hao Yu, Danielle Allessio, Ivon Arroyo, William Lee, Margrit Betke, Beverly Woolf, and Sarah Adel Bargal. 2021. Student engagement dataset. In Proceedings of ICCV.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL.
  8. 8.Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao. 2016. Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. In Proceedings of ECCV.
  9. 9.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of CVPR.
  10. 10.Dou Hu, Xiaolong Hou, Lingwei Wei, Lianxin Jiang, and Yang Mo. 2022a. Mm-dfn: Multimodal dynamic fusion network for emotion recognition in conversations. In Proceedings of ICASSP.
  11. 11.Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, and Yongbin Li. 2022b. Unimse: Towards unified multimodal sentiment analysis and emotion recognition. In Proceedings of EMNLP.
  12. 12.Jingwen Hu, Yuchen Liu, Jinming Zhao, and Qin Jin. 2021. Mmgcn: Multimodal fusion via deep graph convolution network for emotion recognition in conversation. In Proceedings of ACL.
  13. 13.Eric Jang, Shixiang Gu, and Ben Poole. 2017. Categorical reparametrization with gumble-softmax. In Proceedings of ICLR.
  14. 14.Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2010. 3d convolutional neural networks for human action recognition. Proceedings of ICML.
  15. 15.Xingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang, Wanchuang Xia, Cheng Lu, and Jiateng Liu. 2020. Dfew: A large-scale database for recognizing dynamic facial expressions in the wild. In Proceedings of ACM MM.
  16. 16.Xiao Jin, Jianfei Yu, Zixiang Ding, Rui Xia, Xiangsheng Zhou, and Yaofeng Tu. 2020. Hierarchical multimodal transformer with localness and speaker aware attention for emotion recognition in conversations. In Proceedings of NLPCC.
  17. 17.Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Vikram Singh, and Ashutosh Modi. 2022. Cogmen: Contextualized gnn based multimodal emotion recognition. In Proceedings of NAACL-HLT.
  18. 18.Dimitrios Kollias. 2022. Abaw: Valence-arousal estimation, expression recognition, action unit detection & multi-task learning challenges. In Proceedings of CVPR.
  19. 19.Dimitrios Kollias and Stefanos Zafeiriou. 2019. Expression, affect, action unit recognition: Aff-wild2, multi-task learning and arcface. arXiv preprint arXiv:1910.04855.
  20. 20.Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In Proceedings of ACL.
  21. 21.Jiang Li, Xiaoping Wang, Guoqing Lv, and Zhigang Zeng. 2023. Ga2mif: Graph and attention based two-stage multi-source information fusion for conversational emotion detection. IEEE Trans. Affect. Comput.
  22. 22.Jingye Li, Donghong Ji, Fei Li, Meishan Zhang, and Yijiang Liu. 2020. Hitrans: A transformer-based context-and speaker-sensitive model for emotion detection in conversations. In Proceedings of COLING.
  23. 23.Shan Li and Weihong Deng. 2020. Deep facial expression recognition: A survey. IEEE Trans. Affect. Comput.
  24. 24.Shimin Li, Hang Yan, and Xipeng Qiu. 2022a. Contrast and generation make bart a good dialogue emotion recognizer. In Proceedings of AAAI.
  25. 25.Zaijing Li, Fengxiao Tang, Ming Zhao, and Yusen Zhu. 2022b. Emocaps: Emotion capsule based model for conversational emotion recognition. In Proceedings of ACL (Findings).
  26. 26.Jingjun Liang, Ruichen Li, and Qin Jin. 2020. Semi-supervised multi-modal emotion recognition with cross-modal distribution matching. In Proceedings of ACM MM.
  27. 27.Yunlong Liang, Fandong Meng, Ying Zhang, Yufeng Chen, Jinan Xu, and Jie Zhou. 2021. Infusing multi-source knowledge with heterogeneous graph neural network for emotional conversation generation. In Proceedings of AAAI.
  28. 28.Feng Liu, Han-Yang Wang, Si-Yuan Shen, Xun Jia, Jing-Yi Hu, Jia-Hao Zhang, Xi-Yi Wang, Ying Lei, Ai-Min Zhou, Jia-Yin Qi, et al. 2022a. Opo-fcm: A computational affection based occ-pad-ocean federation cognitive modeling approach. IEEE Trans. Comput. Soc. Syst.
  29. 29.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692.
  30. 30.Yuanyuan Liu, Wei Dai, Chuanxu Feng, Wenbin Wang, Guanghao Yin, Jiabei Zeng, and Shiguang Shan. 2022b. Mafw: A large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild. In Proceedings of ACM MM.
  31. 31.Yuchen Liu, Jinming Zhao, Jingwen Hu, Ruichen Li, and Qin Jin. 2022c. Dialogueein: Emotion interaction network for dialogue affective analysis. In Proceedings of COLING.
  32. 32.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of ICCV.
  33. 33.Patrick Lucey, Jeffrey F Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews. 2010. The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression. In Proceedings of CVPR.
  34. 34.Fengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan, and Guosheng Lin. 2021. Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences. In Proceedings of CVPR.
  35. 35.Sijie Mai, Haifeng Hu, and Songlong Xing. 2019. Divide, conquer and combine: Hierarchical feature fusion network with local and global perspectives for multimodal affective computing. In Proceedings of ACL.
  36. 36.Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander Gelbukh, and Erik Cambria. 2019. Dialoguernn: An attentive rnn for emotion detection in conversations. In Proceedings of AAAI.
  37. 37.Yuzhao Mao, Guang Liu, Xiaojie Wang, Weiguo Gao, and Xuan Li. 2021. Dialoguetrm: Exploring multimodal emotional dynamics in a conversation. In Proceedings of EMNLP (Findings).
  38. 38.Naval Kishore Mehta, Shyam Sunder Prasad, Sumeet Saurav, Ravi Saini, and Sanjay Singh. 2022. Three-dimensional densenet self-attention neural network for automatic detection of student's engagement. Appl. Intell.
  39. 39.Donovan Ong, Jian Su, Bin Chen, Anh Tuan Luu, Ashok Narendranath, Yue Li, Shuqi Sun, Yingzhan Lin, and Haifeng Wang. 2022. Is discourse role important for emotion recognition in conversation? In Proceedings of AAAI.
  40. 40.Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In Proceedings of ICASSP.
  41. 41.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. Meld: A multimodal multi-party dataset for emotion recognition in conversations. In Proceedings of ACL.
  42. 42.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res.
  43. 43.Martin Rosvall and Carl T Bergstrom. 2008. Maps of random walks on complex networks reveal community structure. Proceedings of the national academy of sciences.
  44. 44.Andrey V Savchenko. 2022. Video-based frame-level facial analysis of affective behavior on mobile devices using efficientnets. In Proceedings of CVPR.
  45. 45.Weizhou Shen, Siyue Wu, Yunyi Yang, and Xiaojun Quan. 2021. Directed acyclic graph network for conversational emotion recognition. In Proceedings of ACL.
  46. 46.Luke Stark and Jesse Hoey. 2021. The ethics of emotion in artificial intelligence systems. In Proceedings of ACM FAccT.
  47. 47.Luke Stark and Jevan Hutson. 2021. Physiognomic artificial intelligence. Fordham Intell. Prop. Media & Ent. LJ.
  48. 48.Ömer Sümer, Patricia Goldberg, Sidney D'Mello, Peter Gerjets, Ulrich Trautwein, and Enkelejda Kasneci. 2021. Multimodal engagement analysis from facial videos in the classroom. IEEE Trans. Affect. Comput.
  49. 49.Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi. 2017. Inception-v4, inception-resnet and the impact of residual connections on learning. In Proceedings of AAAI.
  50. 50.Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li. 2021. Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection. In Proceedings of ACM MM.
  51. 51.Antoine Toisoul, Jean Kossaifi, Adrian Bulat, Georgios Tzimiropoulos, and Maja Pantic. 2021. Estimation of continuous valence and arousal levels from faces in naturalistic conditions. Nat. Mach. Intell.
  52. 52.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of ICCV.
  53. 53.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of ACL.
  54. 54.Michel Valstar, Maja Pantic, et al. 2010. Induced disgust, happiness and surprise: an addition to the mmi facial expression database. In Proceedings of LREC.
  55. 55.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Proceedings of NeurIPS.
  56. 56.Fanfan Wang, Zixiang Ding, Rui Xia, Zhaoyu Li, and Jianfei Yu. 2022a. Multimodal emotion-cause pair extraction in conversations. IEEE Trans. Affect. Comput.
  57. 57.Yan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu, Shuyong Gao, Wei Zhang, Weifeng Ge, and Wenqiang Zhang. 2022b. Ferv39k: A large-scale multiscene dataset for facial expression recognition in videos. In Proceedings of CVPR.
  58. 58.Dong Yi, Zhen Lei, Shengcai Liao, and Stan Z Li. 2014. Learning face representation from scratch. arXiv preprint arXiv:1411.7923.
  59. 59.Jeewoo Yoon, Chaewon Kang, Seungbae Kim, and Jinyoung Han. 2022. D-vlog: Multimodal vlog dataset for depression detection. In Proceedings of AAAI.
  60. 60.Dong Zhang, Liangqing Wu, Changlong Sun, Shoushan Li, Qiaoming Zhu, and Guodong Zhou. 2019. Modeling both context-and speaker-sensitive dependence for emotion detection in multi-speaker conversations. In Proceedings of IJCAI.
  61. 61.Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. 2016. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Process. Lett.
  62. 62.Guoying Zhao, Xiaohua Huang, Matti Taini, Stan Z Li, and Matti PietikäInen. 2011. Facial expression recognition from near-infrared videos. Image Vis. Comput.
  63. 63.Jinming Zhao, Ruichen Li, Qin Jin, Xinchao Wang, and Haizhou Li. 2022a. Memobert: Pre-training model with prompt-based learning for multimodal emotion recognition. In ICASSP.
  64. 64.Jinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, and Haizhou Li. 2022b. M3ed: Multi-modal multi-scene multi-label emotional dialogue database. In Proceedings of ACL.
  65. 65.Zengqun Zhao and Qingshan Liu. 2021. Former-dfer: Dynamic facial expression recognition transformer. In Proceedings of ACM MM.
  66. 66.ShiHao Zou, Xianying Huang, XuDong Shen, and Hankai Liu. 2022. Improving multimodal fusion with main modal transformer for emotion recognition in conversation. Knowl. Based Syst.

Citation

MLA
Zheng, W., et al. “A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 15445–59, https://doi.org/10.18653/v1/2023.acl-long.861.
APA
Zheng, W., Yu, J., Xia, R., & Wang, S. (2023). A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15445–15459. https://doi.org/10.18653/v1/2023.acl-long.861
Chicago
Zheng, W., J. Yu, R. Xia, and S. Wang. 2023. “A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 15445–59. https://doi.org/10.18653/v1/2023.acl-long.861.
Harvard
Zheng, W. et al. (2023) “A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 15445–15459. Available at: https://doi.org/10.18653/v1/2023.acl-long.861.
Vancouver
1. Zheng W, Yu J, Xia R, Wang S (2023) A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 15445–15459

BibTeX

@inproceedings{zheng-etal-2023-facial,
    title = "A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations",
    author = "Zheng, Wenjie  and
      Yu, Jianfei  and
      Xia, Rui  and
      Wang, Shijin",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.861/",
    doi = "10.18653/v1/2023.acl-long.861",
    pages = "15445--15459"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/