A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation

Xiaoheng ZhangYang Li

article2023ACL61 citations

Proposes CMCF-SRNet, a framework that integrates audio and text through locality-constrained cross-modal attention and graph-based semantic refinement to capture conversational context and speaker emotional inertia for emotion recognition in conversation.

Listen

Emotion recognition in conversation is essential for creating empathetic dialogue systems and automated customer support interfaces. However, existing conversational emotion recognition models often fail to capture subtle cross-modal interactions between spoken audio and text. Furthermore, conventional methods struggle to balance a speaker's short-term emotional inertia in local context with broader, dialogue-wide semantic relationships.

The article designs and demonstrates CMCF-SRNet, a unified neural network framework that combines cross-modality context fusion with semantic refinement. The primary objective is to improve the accuracy of conversational emotion recognition by capturing both immediate conversational dynamics and global semantic dependencies across speech and text.

The researchers evaluated their framework against multiple state-of-the-art baselines using two widely recognized conversational datasets: IEMOCAP (a dyadic scripted and improvised dialogue dataset) and MELD (a multi-party conversational dataset containing over 13,000 utterances from television scripts). Their approach integrates acoustic features extracted via standard audio tools with text embeddings generated by sentence-level language models. The architecture uses a cross-modal transformer constrained by local context and speaker awareness, coupled with a relational graph neural network and graph-transformer to model global semantic structures.

The evaluation yielded several key findings. First, CMCF-SRNet achieved new state-of-the-art performance, reaching an 86.5% weighted F1 score on the 4-way IEMOCAP benchmark (an absolute improvement of 2.0% over prior methods) and 69.6% on the 6-way benchmark (a 1.5% to 10.6% improvement over previous models). Second, the model scored 62.3% weighted F1 on the complex MELD benchmark, outperforming existing multimodal baselines. Third, ablation experiments revealed that the locality-constrained cross-modal attention module contributed a 3.4% boost to weighted F1, while the graph-based semantic refinement components provided an additional 2.7% performance increase over base configurations.

These results demonstrate that prioritizing local conversational context while simultaneously tracking global semantic coherence substantially enhances model performance without incurring excessive dimensional complexity. By strategically weighting spoken and textual cues, systems can achieve higher classification fidelity, reducing errors in downstream human-computer interaction applications.

Organizations developing affective or conversational artificial intelligence should consider adopting locality-aware multimodal fusion and semantic graph architectures to upgrade their conversational agents. When deploying these architectures, technical teams should dynamically calibrate the conversational context window based on the dialogue domain—using narrower windows for rapidly shifting multi-party chats and broader windows for sustained, topic-focused interactions.

Confidence in these findings is supported by rigorous benchmarking across repeated experimental runs. However, operational limitations remain. The model experiences confusion when distinguishing between closely related emotional states, such as anger versus frustration or happiness versus excitement. Additionally, in datasets dominated by neutral interactions, the model displays a tendency to misclassify minority emotional states as neutral. Further research is recommended to introduce fine-grained emotion discrimination before deployment in high-stakes environments.

No sufficiently relevant recommendations were found.

Cover for A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation

Abstract

Emotion recognition in conversation (ERC) has attracted enormous attention for its applications in empathetic dialogue systems. However, most previous researches simply concatenate multi-modal representations, leading to an accumulation of redundant information and a limited context interaction between modalities. Furthermore, they only consider simple contextual features ignoring semantic clues, resulting in an insufficient capture of the semantic coherence and consistency in conversations. To address these limitations, we propose a cross-modality context fusion and semantic refinement network (CMCF-SRNet). Specifically, we first design a cross-modal locality-constrained transformer to explore the multimodal interaction. Second, we investigate a graph-based semantic refinement transformer, which solves the limitation of insufficient semantic relationship information between utterances. Extensive experiments on two public benchmark datasets show the effectiveness of our proposed method compared with other state-of-the-art methods, indicating its potential application in emotion recognition. Our model is available at https://github.com/zxiaohen/CMCF-SRNet.

Table of Contents

  • 1 Introduction
  • 2 Methodology
  • 2.1 Cross-modality Context Fusion Module
  • 2.2 Semantic Refinement Module
  • 2.3 Emotion Classifier
  • 3 Experiments and Results
  • 3.1 Datasets
  • 3.2 Implementation Details and Metrics
  • 3.3 Overall Performance
  • 4 Discussion
  • 4.1 Effect of Cross-modality Context Fusion
  • 4.2 Effect of Semantic Refinement
  • 4.3 Visualization and Interpretability
  • 5 Conclusion
  • Limitation
  • Acknowledgements
  • References
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — CMCF-SRNet combines local multimodal interaction with semantic graph refinement

    model/method

    CMCF-SRNet predicts an emotion for each utterance in a dialogue using audio and text. It first forms modality-specific utterance representations and uses intra-modal transformers to capture temporal dependencies. A two-stream cross-modal transformer then exchanges information between audio and text, with locality-constrained attention (LCA) controlling which utterances receive weight. An attentive selection block combines the audio, text, and cross-modal representations. The resulting utterance features define a semantic graph, whose node features are refined first by a relational graph convolutional network (RGCN) and then by a semantic graph transformer. A multilayer perceptron and softmax layer produce the emotion prediction for each utterance; training uses categorical cross-entropy.

  2. Knowl 2 — Speaker-aware locality-constrained cross-modal attention

    model/method

    For a target utterance at position mm in modality ii, CMCF-SRNet attends to utterances at positions nn in the other modality jj, where i,j∈{a,t}i,j\in\{a,t\} and i≠ji\ne j (aa is audio and tt is text). Queries qm(i)q_m^{(i)}, keys kn(j)k_n^{(j)}, and values vn(j)v_n^{(j)} are learned projections of the corresponding modality features. The attention weights are modulated by a locality and speaker mask:

    Lmn=σ ⁣(M−C(n−m)2)1[sm=sn],omi→j=∑nAmni→jLmnvn(j),L_{mn}=\sigma\!\left(M-C(n-m)^2\right)\mathbf{1}[s_m=s_n],\qquad o_m^{i\to j}=\sum_n A_{mn}^{i\to j}L_{mn}v_n^{(j)},

    where sms_m and sns_n are the speakers of the two utterances, 1[sm=sn]\mathbf{1}[s_m=s_n] is one when they are the same speaker and zero otherwise, and σ\sigma is the sigmoid function. Amni→jA_{mn}^{i\to j} is the ordinary scaled dot-product attention weight, normalized across source positions nn; M=5M=5 and C=1.5C=1.5. The quadratic position term gives closer utterances greater locality weight, while the speaker mask emphasizes same-speaker emotional inertia. The resulting attention is used in the cross-modal transformer to update one modality using features from the other.

  3. Knowl 3 — Semantic graph construction and relation-aware local refinement

    model/method

    CMCF-SRNet represents every dialogue utterance as a graph node and connects utterances in a past-and-future context window. For utterances mm and nn, with nonzero fused feature vectors gmg_m and gng_n, the semantic edge score is simm,n=1−arccos⁡ ⁣(gm⊤gn∥gm∥∥gn∥)\mathrm{sim}_{m,n}=1-\arccos\!\left(\frac{g_m^\top g_n}{\lVert g_m\rVert\lVert g_n\rVert}\right). Edges are directed and encode past or future relations; their relation types also distinguish utterances spoken by the same speaker from those spoken by different speakers. The window sizes PP and FF specify how many past and future utterances are considered. A two-layer correlation-based RGCN uses these graph relations and edge weights to update utterance nodes: it aggregates neighboring features with learned relation-specific transformations, includes a self-node transformation, and applies ReLU activations. This stage captures local inter-utterance dependencies and produces features for the subsequent semantic graph transformer.

  4. Knowl 4 — Graph transformer incorporates semantic and topological relations

    model/method

    After RGCN refinement, CMCF-SRNet applies a graph transformer to the utterance-node features. For pairs of nodes, it uses both a topological relative-position encoding based on their shortest-path distance in the semantic graph and a semantic encoding based on their semantic edge score. These relation encodings affect the attention logits and are also added to the value features being aggregated. Thus, the attention computation uses node content together with graph position and semantic relation information, while the value aggregation carries relation information into the updated node representations. The refined node features are then passed to the emotion classifier.

  5. Knowl 5 — Attentive selection weights audio, text, and cross-modal features

    model/method

    For each utterance, CMCF-SRNet's attentive selection block (ASB) combines three equal-dimension representations: acoustic hm(a)h_m^{(a)}, textual hm(t)h_m^{(t)}, and cross-modal hm(c)h_m^{(c)}. A learned projection followed by ReLU assigns a score to each representation; a softmax across the three scores produces weights αa,αt,αc\alpha_a,\alpha_t,\alpha_c that sum to one. The output is their weighted concatenation, gm=[αahm(a);αthm(t);αchm(c)]g_m=[\alpha_a h_m^{(a)};\alpha_t h_m^{(t)};\alpha_c h_m^{(c)}]. This lets the model emphasize modality features relevant to the utterance rather than treating modalities equally.

  6. Knowl 6 — Benchmark performance on IEMOCAP and MELD

    empirical result

    In the reported IEMOCAP six-class evaluation, CMCF-SRNet obtains 70.5% weighted average accuracy (WAA) and 69.6% weighted F1 (WF1). Its class WF1 scores are 52.2% for happy, 80.9% for sad, 68.8% for neutral, 70.3% for angry, 76.7% for excited, and 61.6% for frustrated. On MELD, the reported WF1 is 62.3%. The authors report that the model outperforms the compared baselines on the reported benchmarks. For the IEMOCAP four-class audio-and-text setting, the ablation results report 86.8% WAA and 86.5% WF1. These results are averages over five runs, with standard deviations reported for the main six-class and MELD scores.

  7. Knowl 7 — Ablations identify contributions from cross-modal attention and semantic refinement

    empirical result

    The component and fusion ablations report WAA/WF1 on IEMOCAP four-class and MELD, in that order. The complete audio-and-text model scores 86.8/86.5 and 62.8/62.3. Removing LCA gives 83.6/83.2 and 60.5/59.3; removing ASB gives 84.5/84.1 and 61.1/60.3. Removing semantic edge weights (SEW) gives 84.2/83.6 and 59.8/57.9, while removing semantic-positional encoding (SPE) gives 83.6/83.8 and 60.8/59.6. The modality and fusion comparisons are: text only, 85.6/85.1 and 60.4/59.7; audio only, 60.6/59.2 and 55.5/53.2; concatenation, 85.6/84.2 and 60.2/59.62; addition, 84.3/83.9 and 59.8/58.5; and tensor fusion, 83.6/83.1 and 53.5/60.3. The results show that removing any of the tested components lowers performance relative to the full model, and that ASB outperforms the reported alternative fusion strategies.

  8. Knowl 8 — Training and evaluation configuration

    experimental setup

    The experiments use IEMOCAP and MELD. IEMOCAP contains 10,039 utterances; the authors evaluate four-class and six-class settings using the first four and first six emotion categories, respectively, and use utterances from the first eight speakers for training and validation and the remaining speakers for testing. MELD contains 13,708 utterances in 1,433 conversations, annotated with seven emotion labels. Acoustic features have dimension 100 and are extracted with OpenSmile; textual features have dimension 768 and are extracted with sBERT. Training uses Adam with an initial learning rate of 0.0001 and dropout of 0.5. The cross-modal transformer uses four attention heads and the graph transformer uses two. Each dataset is run five times and results are reported as means and standard deviations. Evaluation includes WAA, a class-accuracy average weighted by class frequency, and WF1, a class-F1 average weighted by class frequency.

  9. Knowl 9 — Attention and window-size analyses indicate both local and distant context matter

    empirical result

    The authors analyze attention distances for correctly classified MELD utterances and report that many rely on nearby context, while distant utterances also contribute; distant context is more evident among the second-highest-attended utterances. Contextual attention occurs toward both past and future utterances. Their window-size experiments indicate that larger windows tend to help when inter-speaker and intra-speaker dependencies persist over longer sequences, whereas smaller windows are more suitable when dialogue topics change frequently and speakers are less influenced by one another. The authors also report that feature visualizations show better-formed emotion clusters when semantic refinement is included.

  10. Knowl 10 — The model confuses similar emotions and can overpredict neutral

    limitation

    CMCF-SRNet does not reliably distinguish some similar emotion labels: frustration can be confused with anger, and happiness with excitement. On MELD, the model also tends to assign other emotions to neutral, which the authors attribute to the large proportion of neutral examples in the dataset. They identify fine-grained emotion modeling as a direction for addressing these errors.

Coverage note — No substantial contributed material was omitted; individual attention-map examples and confusion-matrix cell counts were not separately encoded because the paper discusses their qualitative implications in the context and error-analysis knowls.

References

  1. 1.Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008. Iemocap: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42(4):335–359.
  2. 2.Shizhe Chen and Qin Jin. 2016. Multi-modal conditional attention fusion for dimensional emotion prediction. In Proceedings of the 24th ACM international conference on Multimedia, pages 571–575.
  3. 3.Florian Eyben, Martin Wöllmer, and Björn Schuller. 2010. Opensmile: the munich versatile and fast open-source audio feature extractor. In Proceedings of the 18th ACM international conference on Multimedia, pages 1459–1462.
  4. 4.Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, and Alexander Gelbukh. 2019. DialogueGCN: A graph convolutional neural network for emotion recognition in conversation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 154–164, Hong Kong, China. Association for Computational Linguistics.
  5. 5.Devamanyu Hazarika, Soujanya Poria, Rada Mihalcea, Erik Cambria, and Roger Zimmermann. 2018a. ICON: Interactive conversational memory network for multimodal emotion detection. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2594–2604, Brussels, Belgium. Association for Computational Linguistics.
  6. 6.Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria, Louis-Philippe Morency, and Roger Zimmermann. 2018b. Conversational memory network for emotion recognition in dyadic dialogue videos. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2122–2132, New Orleans, Louisiana. Association for Computational Linguistics.
  7. 7.Dou Hu, Xiaolong Hou, Lingwei Wei, Lianxin Jiang, and Yang Mo. 2022. Mm-dfn: Multimodal dynamic fusion network for emotion recognition in conversations. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7037–7041.
  8. 8.Jingwen Hu, Yuchen Liu, Jinming Zhao, and Qin Jin. 2021. MMGCN: Multimodal fusion via deep graph convolution network for emotion recognition in conversation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 5666–5675, Online. Association for Computational Linguistics.
  9. 9.Elvin Isufi, Fernando Gama, and Alejandro Ribeiro. 2022. Edgenets: Edge varying graph neural networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(11):7457–7473.
  10. 10.Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Singh, and Ashutosh Modi. 2022. COGMEN: COntextualized GNN based multimodal emotion recognitioN. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4148–4164, Seattle, United States. Association for Computational Linguistics.
  11. 11.Zheng Lian, Bin Liu, and Jianhua Tao. 2021. Ct-net: Conversational transformer network for emotion recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:985–1000.
  12. 12.Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, and Louis-Philippe Morency. 2018. Efficient low-rank multimodal fusion with modality-specific factors. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2247–2256, Melbourne, Australia. Association for Computational Linguistics.
  13. 13.Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander Gelbukh, and Erik Cambria. 2019. Dialoguernn: An attentive rnn for emotion detection in conversations. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 6818–6825.
  14. 14.Weizhi Nie, Rihao Chang, Minjie Ren, Yuting Su, and Anan Liu. 2022. I-gcn: Incremental graph convolution network for conversation emotion detection. IEEE Transactions on Multimedia, 24:4471–4481.
  15. 15.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017a. Context-dependent sentiment analysis in user-generated videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 873–883, Vancouver, Canada. Association for Computational Linguistics.
  16. 16.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017b. Context-dependent sentiment analysis in user-generated videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 873–883, Vancouver, Canada. Association for Computational Linguistics.
  17. 17.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017c. Context-dependent sentiment analysis in user-generated videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 873–883, Vancouver, Canada. Association for Computational Linguistics.
  18. 18.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. MELD: A multimodal multi-party dataset for emotion recognition in conversations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 527–536, Florence, Italy. Association for Computational Linguistics.
  19. 19.Aravind Sesagiri Raamkumar and Yinping Yang. 2022. Empathetic conversational systems: A review of current advances, gaps, and opportunities. IEEE Transactions on Affective Computing, pages 1–20.
  20. 20.Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084.
  21. 21.Minjie Ren, Xiangdong Huang, Wenhui Li, Dan Song, and Weizhi Nie. 2022. Lr-gcn: Latent relation-aware graph convolutional network for conversational emotion recognition. IEEE Transactions on Multimedia, 24:4422–4432.
  22. 22.Weizhou Shen, Siyue Wu, Yunyi Yang, and Xiaojun Quan. 2021. Directed acyclic graph network for conversational emotion recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1551–1560, Online. Association for Computational Linguistics.
  23. 23.Yao-Hung Hubert Tsai, Shaojie Bai, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6558–6569, Florence, Italy. Association for Computational Linguistics.
  24. 24.Geng Tu, Bin Liang, Dazhi Jiang, and Ruifeng Xu. 2022. Sentiment- emotion- and context-guided knowledge selection framework for emotion recognition in conversations. IEEE Transactions on Affective Computing, pages 1–14.
  25. 25.Songlong Xing, Sijie Mai, and Haifeng Hu. 2022. Adapted dynamic memory network for emotion recognition in conversation. IEEE Transactions on Affective Computing, 13(3):1426–1439.
  26. 26.Hao Yuan, Haiyang Yu, Shurui Gui, and Shuiwang Ji. 2022. Explainability in graph neural networks: A taxonomic survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–19.
  27. 27.Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1103–1114, Copenhagen, Denmark. Association for Computational Linguistics.
  28. 28.Tong Zhu, Leida Li, Jufeng Yang, Sicheng Zhao, and Xiao Xiao. 2022. Multimodal emotion classification with multi-level semantic reasoning network. IEEE Transactions on Multimedia, pages 1–13.

Citation

MLA
Zhang, X., and Y. Li. “A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 13099–110, https://doi.org/10.18653/v1/2023.acl-long.732.
APA
Zhang, X., & Li, Y. (2023). A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13099–13110. https://doi.org/10.18653/v1/2023.acl-long.732
Chicago
Zhang, X., and Y. Li. 2023. “A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 13099–110. https://doi.org/10.18653/v1/2023.acl-long.732.
Harvard
Zhang, X. and Li, Y. (2023) “A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 13099–13110. Available at: https://doi.org/10.18653/v1/2023.acl-long.732.
Vancouver
1. Zhang X, Li Y (2023) A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 13099–13110

BibTeX

@inproceedings{zhang-li-2023-cross,
    title = "A Cross-Modality Context Fusion and Semantic Refinement Network for Emotion Recognition in Conversation",
    author = "Zhang, Xiaoheng  and
      Li, Yang",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.732/",
    doi = "10.18653/v1/2023.acl-long.732",
    pages = "13099--13110"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/