Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction

Cam-Van Thi NguyenAnh-Tuan MaiThe-Son LeHai-Dang KieuDuc-Trong Le

article2023EMNLP50 citations

Proposes a relational temporal graph neural network framework, CORECT, that jointly models utterance-level temporal dependencies and conversation-level cross-modal interactions while preserving modality-specific representations to achieve state-of-the-art multimodal emotion recognition on IEMOCAP and CMU-MOSEI.

Listen

Emotion recognition in spoken and text conversations is essential for developing intelligent, human-aware systems. Real-world human interaction relies on multiple signals—including text, vocal tone, and visual expressions—that evolve dynamically over time. Most existing systems struggle to accurately understand these emotions because they either merge different sensory streams too early into a single representation, discarding modality-specific nuances, or they fail to properly capture how past and future utterances influence the current emotional state.

The article demonstrates a novel framework called CORECT, designed to improve multimodal emotion recognition by jointly modeling local temporal dependencies and global cross-modality interactions without prematurely collapsing individual modality features.

To evaluate this framework, the authors conducted extensive experiments on two benchmark conversation datasets: IEMOCAP, which includes 151 dyadic dialogues across 7,433 utterances, and CMU-MOSEI, which contains over 22,000 utterances. The method extracts separate features for text, audio, and video, models their relationships and time order using a relational temporal graph network, and enriches these with a pairwise cross-modal interaction module. The system's performance was compared against leading baseline models across standard emotion and sentiment classification metrics.

The framework achieved new state-of-the-art results across both benchmarks. On the six-class IEMOCAP benchmark, CORECT improved overall accuracy by 2.89% and weighted F1-score by 2.75% over the previous best-performing model, reaching an accuracy of 69.93%. On the four-class version, it achieved an accuracy of 84.73%, representing a 2.44% improvement. Ablation analyses confirmed that removing the relational temporal graph module caused the largest performance decline (up to 4.10%), proving that local structural and temporal modeling is critical. Furthermore, contextual analysis showed an asymmetrical time effect: past utterances exert a significantly stronger influence on current emotional state than future utterances, with an optimal window of eleven past utterances versus nine future ones.

These findings demonstrate that retaining separate modality representations while explicitly modeling how conversation history unfolds yields substantial gains in conversational understanding. For organizations building conversational agents, customer service analytics, or affective monitoring tools, adopting architecture that explicitly tracks temporal flow and cross-modal dependencies improves classification reliability, particularly in identifying difficult or underrepresented emotions like fear and surprise that standard models fail to distinguish.

For future development, organizations and researchers should focus on implementing automated hyperparameter tuning to optimize past and future contextual window sizes dynamically. Additionally, developers should explore adaptive attention mechanisms that selectively weight earlier utterances rather than relying on fixed sliding windows, ensuring computational efficiency during deployment in real-time conversational systems.

Confidence in the findings is high based on consistent improvements across diverse datasets and rigorous module-by-module ablation. However, decision-makers should note that performance remains lower for ambiguous emotion pairs—such as distinguishing excitement from happiness or frustration from sadness—and visual input quality remains susceptible to real-world noise such as camera angle and lighting variations.

arXiv: 2311.04507leson502/CORECT_EMNLP2023

No sufficiently relevant recommendations were found.

Cover for Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction

Abstract

Emotion recognition is a crucial task for human conversation understanding. It becomes more challenging with the notion of multimodal data, e.g., language, voice, and facial expressions. As a typical solution, the global and local context information are exploited to predict the emotional label for every single sentence, i.e., utterance, in the dialogue. Specifically, the global representation could be captured via modeling of cross-modal interactions at the conversation level. The local one is often inferred using the temporal information of speakers or emotional shifts, which neglects vital factors at the utterance level. Additionally, most existing approaches take fused features of multiple modalities in a unified input without leveraging modality-specific representations. Motivating from these problems, we propose the Relational Temporal Graph Neural Network with Auxiliary Cross-Modality Interaction (CORECT), an novel neural network framework that effectively captures conversation-level cross-modality interactions and utterance-level temporal dependencies with the modality-specific manner for conversation understanding. Extensive experiments demonstrate the effectiveness of CORECT via its state-of-the-art results on the IEMOCAP and CMU-MOSEI datasets for the multimodal ERC task.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Multimodal Emotion Recognition in Conversation
  • 2.2 Graph Neural Networks
  • 3 Methodology
  • 3.1 Utterance-level Feature Extraction
  • 3.1.1 Unimodal Encoder
  • 3.1.2 Speaker Embedding
  • 3.2 Relational Temporal Graph Convolutional Network (RT-GCN)
  • 3.2.1 Multimodal Graph Construction
  • 3.2.2 Graph Learning
  • 3.3 Pairwise Cross-modal Feature Interaction
  • 3.4 Multimodal Emotion Classification
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.2 Comparison With Baselines
  • 4.3 Ablation study
  • 5 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Appendix
  • A.1 Implementation Details
  • A.1.1 Multimodal Raw Feature Extraction
  • A.1.2 Pairwise Cross-modal Feature Interaction
  • A.2 Additional Experiment Result
  • A.3 Reproducibility

Knowls

  1. Knowl 1 — CORECT combines utterance-level graph context with dialogue-level cross-modal context

    model/method

    CORECT is a multimodal emotion-recognition model for dialogue. It keeps audio, visual, and text representations modality-specific while using two complementary context mechanisms: RT-GCN builds local utterance representations from multimodal and temporal graph relations, and P-CM builds conversation-level representations through pairwise cross-modal attention. For each utterance, CORECT concatenates the graph representations for the three modalities with the pairwise cross-modal representations, then applies a feed-forward classifier and a softmax to predict the utterance's emotion.

  2. Knowl 2 — RT-GCN represents modality and temporal relations in a graph

    model/method

    For a dialogue of NN utterances, RT-GCN constructs three nodes per utterance, one each for audio, visual, and text features. It uses 15 directed relation types: nine within-utterance multimodal relations (the six directed connections between distinct modalities and one self-connection per modality), plus six temporal relations (past and future for each modality). Temporal edges connect nodes of the same modality, with separate past and future windows of widths PP and FF. A relation-specific graph convolution aggregates normalized neighbor messages using learned transformations, with a separate self-node transformation. A multi-head Graph Transformer then attends over connected nodes and produces one updated representation for each utterance and modality. The resulting representations provide RT-GCN's local-context features.

  3. Knowl 3 — P-CM transfers information between modality sequences using bidirectional cross-attention

    model/method

    Pairwise Cross-modal Feature Interaction (P-CM) operates on the full sequence of utterance features for each modality. For a target modality and a source modality, it forms queries from the target and keys and values from the source, then applies scaled dot-product cross-attention so each target position can draw information from source positions without requiring aligned sequences. Each cross-modal transformer layer applies layer normalization, attention with a residual connection, and a position-wise feed-forward network with another residual connection. P-CM stacks DD such layers and computes transfers in both directions for each of the three modality pairs: audio–visual, visual–text, and text–audio. The final pairwise representations are concatenated to provide dialogue-level cross-modal context.

  4. Knowl 4 — Input features and training configuration used for evaluation

    experimental setup

    CORECT was evaluated on IEMOCAP and CMU-MOSEI. IEMOCAP uses 108/12/31 train/validation/test dialogues, with 5,146/664/1,623 utterances for the six-class task and 3,200/400/943 for the four-class task. CMU-MOSEI uses 2,249/300/646 dialogues and 16,327/1,871/4,662 utterances for those splits. IEMOCAP features have dimensions 100 for audio (OpenSmile), 512 for visual (OpenFace), and 768 for text (sBERT); CMU-MOSEI features have dimensions 80 for audio (librosa filter banks), 35 for visual, and 768 for text (sBERT). Text is encoded with a Transformer and audio and visual features with fully connected encoders. A learned speaker embedding is added to each modality's utterance features, scaled by a contribution coefficient. Training uses Adam and dropout 0.5. The Graph Transformer uses seven attention heads and P-CM uses two. Learning rates are 0.0003 for IEMOCAP and 0.0006 for CMU-MOSEI. The selected temporal windows are [P,F]=[11,9][P,F]=[11,9] for IEMOCAP and [5,4][5,4] for CMU-MOSEI.

  5. Knowl 5 — CORECT improves overall results on both IEMOCAP emotion-label settings

    empirical result

    With audio, visual, and text inputs, CORECT obtains 69.93% accuracy and 70.02% weighted F1 on IEMOCAP six-class emotion recognition. The strongest prior baseline reported, COGMEN, obtains 67.04% accuracy and 67.27% weighted F1, so CORECT's gains are 2.89 and 2.75 percentage points. CORECT's F1 scores by class are 59.30% happy, 80.53% sad, 66.94% neutral, 69.59% angry, 72.69% excited, and 68.50% frustrated; it does not have the highest reported F1 for sad or excited. On the four-class setting, CORECT reaches 84.73% accuracy and 84.64% weighted F1, compared with COGMEN's 82.29% and 82.15%, gains of 2.44 and 2.49 percentage points.

  6. Knowl 6 — CORECT achieves the best reported CMU-MOSEI scores across sentiment and emotion tasks

    empirical result

    Using all three modalities on CMU-MOSEI, CORECT records 83.66% accuracy on two-class sentiment and 46.31% on seven-class sentiment. Its weighted F1 scores for the one-versus-all emotion tasks are 71.35% happiness, 72.86% sadness, 76.77% anger, 87.90% fear, 84.26% disgust, and 86.48% surprise. These exceed the corresponding COGMEN results of 82.95%, 45.22%, 70.88%, 70.91%, 74.20%, 87.79%, 81.83%, and 86.05%, respectively. The authors report that reproduced baseline classifiers did not distinguish any samples for fear or surprise; CORECT's scores on those labels are slightly higher.

  7. Knowl 7 — Ablations show contributions from both context modules and both graph relation groups

    empirical result

    Removing either major CORECT module lowers performance on IEMOCAP. On the six-class task, the full model scores 69.93% accuracy and 70.02% weighted F1; without RT-GCN it scores 66.61% and 66.55%, and without P-CM it scores 66.54% and 66.64%. On the four-class task, the full model scores 84.73% and 84.64%; without RT-GCN it scores 80.69% and 80.54%, and without P-CM it scores 82.18% and 82.16%. Removing either relation group also reduces performance: without multimodal relations, scores are 66.54%/66.82% for six-class and 82.61%/82.53% for four-class; without temporal relations, they are 67.04%/67.34% and 82.08%/82.07%, respectively. These results support the use of both local graph context and dialogue-level cross-modal interaction, as well as both relation types.

  8. Knowl 8 — Text is the strongest single IEMOCAP modality, while all three modalities perform best together

    empirical result

    In CORECT's IEMOCAP modality comparison, text is the strongest unimodal input and visual features are the weakest. For the six-class task, text alone gives 67.22% accuracy and 67.26% weighted F1, compared with 52.31%/51.49% for audio and 38.63%/37.67% for vision. Audio plus text is the strongest pair, at 68.27%/68.36%; text plus vision gives 65.50%/65.61%, and vision plus audio gives 54.16%/53.82%. All three modalities give the highest scores, 69.93%/70.02%. The same ordering of single modalities and modality pairs holds on the four-class task: text scores 82.82%/82.65%, audio 67.02%/65.48%, vision 49.73%/47.97%; audio plus text scores 83.14%/83.13%, text plus vision 81.76%/81.75%, vision plus audio 69.03%/68.21%, and all three modalities 84.73%/84.64%.

  9. Knowl 9 — The best tested IEMOCAP temporal window favors more past than future context

    empirical result

    For six-class IEMOCAP, the authors varied the past and future temporal-window widths over values from 1 to 15. The best tested setting was [P,F]=[11,9][P,F]=[11,9]. Their analysis reports that past context has a stronger influence on performance than future context, while also showing that the two window widths do not have identical effects.

  10. Knowl 10 — Hyperparameter exploration was limited

    limitation

    The authors state that time and computational-resource constraints prevented exhaustive tuning of CORECT's hyperparameters, including the number of P-CM attention heads and the past and future window sizes. They note that this limited exploration could lead to convergence at local minima. They suggest automated hyperparameter optimization and mechanisms that learn the relative importance of past and future utterances as possible ways to address the limitation.

Coverage note — The detailed confusion matrices and the CMU-MOSEI modality-combination ablation are omitted; the principal benchmark comparisons, IEMOCAP modality trends, and the stated label-imbalance observation are retained.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  2. 2.AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236–2246, Melbourne, Australia. Association for Computational Linguistics.
  3. 3.Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443.
  4. 4.Tadas Baltrusaitis, Amir Zadeh, Yao Chong Lim, and Louis-Philippe Morency. 2018. Openface 2.0: Facial behavior analysis toolkit. In 2018 13th IEEE international conference on automatic face & gesture recognition (FG 2018), pages 59–66. IEEE.
  5. 5.Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008. Iemocap: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42(4):335–359.
  6. 6.Feiyu Chen, Jie Shao, Shuyuan Zhu, and Heng Tao Shen. 2023. Multivariate, multi-frequency and multimodal: Rethinking graph neural networks for emotion recognition in conversation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10761–10770.
  7. 7.Jean-Benoit Delbrouck, Noé Tits, Mathilde Brousmiche, and Stéphane Dupont. 2020. A transformer-based joint-encoding for emotion recognition and sentiment analysis. In Second Grand-Challenge and Workshop on Multimodal Language (Challenge-HML), pages 1–7, Seattle, USA. Association for Computational Linguistics.
  8. 8.Florian Eyben, Martin Wöllmer, and Björn Schuller. 2010. Opensmile: the munich versatile and fast open-source audio feature extractor. In Proceedings of the 18th ACM international conference on Multimedia, pages 1459–1462.
  9. 9.Deepanway Ghosal, Navonil Majumder, Alexander Gelbukh, Rada Mihalcea, and Soujanya Poria. 2020. COSMIC: COmmonSense knowledge for eMotion identification in conversations. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 2470–2481, Online. Association for Computational Linguistics.
  10. 10.Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, and Alexander Gelbukh. 2019. DialogueGCN: A graph convolutional neural network for emotion recognition in conversation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 154–164, Hong Kong, China. Association for Computational Linguistics.
  11. 11.Marco Gori, Gabriele Monfardini, and Franco Scarselli. 2005. A new model for learning in graph domains. In Proceedings. 2005 IEEE International Joint Conference on Neural Networks, 2005., volume 2, pages 729–734. IEEE.
  12. 12.Devamanyu Hazarika, Soujanya Poria, Rada Mihalcea, Erik Cambria, and Roger Zimmermann. 2018a. ICON: Interactive conversational memory network for multimodal emotion detection. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2594–2604, Brussels, Belgium. Association for Computational Linguistics.
  13. 13.Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria, Louis-Philippe Morency, and Roger Zimmermann. 2018b. Conversational memory network for emotion recognition in dyadic dialogue videos. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2122–2132, New Orleans, Louisiana. Association for Computational Linguistics.
  14. 14.Dou Hu, Lingwei Wei, and Xiaoyong Huai. 2021. DialogueCRN: Contextual reasoning networks for emotion recognition in conversations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 7042–7052, Online. Association for Computational Linguistics.
  15. 15.Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Singh, and Ashutosh Modi. 2022. COGMEN: COntextualized GNN based multimodal emotion recognitioN. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4148–4164, Seattle, United States. Association for Computational Linguistics.
  16. 16.Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In International Conference on Learning Representations (ICLR).
  17. 17.Jiang Li, Xiaoping Wang, Guoqing Lv, and Zhigang Zeng. 2023. Graphmft: A graph network based multimodal fusion technique for emotion recognition in conversation. Neurocomputing, page 126427.
  18. 18.Zheng Lian, Bin Liu, and Jianhua Tao. 2022. Smin: Semi-supervised multi-modal interaction network for conversational emotion recognition. IEEE Transactions on Affective Computing.
  19. 19.Zheng Lian, Jianhua Tao, Bin Liu, Jian Huang, Zhanlei Yang, and Rongjun Li. 2020. Conversational emotion recognition using self-attention mechanisms and graph neural networks. In INTERSPEECH, pages 2347–2351.
  20. 20.Navonil Majumder, Devamanyu Hazarika, Alexander Gelbukh, Erik Cambria, and Soujanya Poria. 2018. Multimodal sentiment analysis using hierarchical fusion with context modeling. Knowledge-based systems, 161:124–133.
  21. 21.Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander Gelbukh, and Erik Cambria. 2019. Dialoguernn: An attentive rnn for emotion detection in conversations. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pages 6818–6825.
  22. 22.Brian McFee, Colin Raffel, Dawen Liang, Daniel P Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015. librosa: Audio and music signal analysis in python. In Proceedings of the 14th python in science conference, volume 8, pages 18–25.
  23. 23.Andrei Nicolicioiu, Iulia Duta, and Marius Leordeanu. 2019. Recurrent space-time graph neural networks. Advances in neural information processing systems, 32.
  24. 24.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Context-dependent sentiment analysis in user-generated videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 873–883, Vancouver, Canada. Association for Computational Linguistics.
  25. 25.Soujanya Poria, Navonil Majumder, Rada Mihalcea, and Eduard Hovy. 2019. Emotion recognition in conversation: Research challenges, datasets, and recent advances. IEEE Access, 7:100943–100953.
  26. 26.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  27. 27.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. 2008. The graph neural network model. IEEE transactions on neural networks, 20(1):61–80.
  28. 28.Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne Van Den Berg, Ivan Titov, and Max Welling. 2018. Modeling relational data with graph convolutional networks. In The Semantic Web: 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3–7, 2018, Proceedings 15, pages 593–607. Springer.
  29. 29.Garima Sharma and Abhinav Dhall. 2021. A survey on automatic multimodal emotion recognition in the wild. Advances in data science: Methodologies and applications, pages 35–64.
  30. 30.Weizhou Shen, Junqing Chen, Xiaojun Quan, and Zhixian Xie. 2021a. Dialogxl: All-in-one xlnet for multi-party conversation emotion recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13789–13797.
  31. 31.Weizhou Shen, Siyue Wu, Yunyi Yang, and Xiaojun Quan. 2021b. Directed acyclic graph network for conversational emotion recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1551–1560, Online. Association for Computational Linguistics.
  32. 32.Aman Shenoy and Ashish Sardana. 2020. Multilogue-net: A context-aware RNN for multi-modal emotion detection and sentiment analysis in conversation. In Second Grand-Challenge and Workshop on Multimodal Language (Challenge-HML), pages 19–28, Seattle, USA. Association for Computational Linguistics.
  33. 33.Tao Shi and Shao-Lun Huang. 2023. MultiEMO: An attention-based correlation-aware multimodal fusion framework for emotion recognition in conversations. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14752–14766, Toronto, Canada. Association for Computational Linguistics.
  34. 34.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6558–6569, Florence, Italy. Association for Computational Linguistics.
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems, 30.
  36. 36.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018. Graph attention networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  37. 37.Yan Wang, Wei Song, Wei Tao, Antonio Liotta, Dawei Yang, Xinlei Li, Shuyong Gao, Yixuan Sun, Weifeng Ge, Wei Zhang, et al. 2022. A systematic review on affective computing: Emotion models, databases, and recent advances. Information Fusion.
  38. 38.Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. Mmgcn: Multi-modal graph convolution network for personalized recommendation of micro-video. In Proceedings of the 27th ACM international conference on multimedia, pages 1437–1445.
  39. 39.Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, and Louis-Philippe Morency. 2021. MTAG: Modal-temporal attention graph for unaligned human multimodal language sequences. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1009–1021, Online. Association for Computational Linguistics.
  40. 40.Seongjun Yun, Minbyul Jeong, Raehyun Kim, Jaewoo Kang, and Hyunwoo J. Kim. 2019. Graph transformer networks. CoRR, abs/1911.06455.
  41. 41.Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Memory fusion network for multi-view sequential learning. In Proceedings of the AAAI conference on artificial intelligence, volume 32.
  42. 42.Dong Zhang, Liangqing Wu, Changlong Sun, Shoushan Li, Qiaoming Zhu, and Guodong Zhou. 2019. Modeling both context-and speaker-sensitive dependence for emotion detection in multi-speaker conversations. In IJCAI, pages 5415–5421.

Citation

MLA
Nguyen, C.-V. T., et al. “Conversation Understanding Using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 15154–67, https://doi.org/10.18653/v1/2023.emnlp-main.937.
APA
Nguyen, C.-V. T., Mai, A.-T., Le, T.-S., Kieu, H.-D., & Le, D.-T. (2023). Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 15154–15167. https://doi.org/10.18653/v1/2023.emnlp-main.937
Chicago
Nguyen, C.-V. T., A.-T. Mai, T.-S. Le, H.-D. Kieu, and D.-T. Le. 2023. “Conversation Understanding Using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 15154–67. https://doi.org/10.18653/v1/2023.emnlp-main.937.
Harvard
Nguyen, C.-V.T. et al. (2023) “Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 15154–15167. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.937.
Vancouver
1. Nguyen C-VT, Mai A-T, Le T-S, Kieu H-D, Le D-T (2023) Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 15154–15167

BibTeX

@inproceedings{nguyen-etal-2023-conversation,
    title = "Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction",
    author = "Nguyen, Cam-Van Thi  and
      Mai, Anh-Tuan  and
      Le, The-Son  and
      Kieu, Hai-Dang  and
      Le, Duc-Trong",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.937/",
    doi = "10.18653/v1/2023.emnlp-main.937",
    pages = "15154--15167"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/