COGMEN: COntextualized GNN based Multimodal Emotion recognitioN

Abhinav JoshiAshwani BhatAyush JainAtin Vikram SinghAshutosh Modi

article2022NAACL150 citations

Proposes a contextualized graph neural network architecture that models both global conversational context and speaker dependencies across multiple modalities to achieve state-of-the-art multimodal emotion recognition on IEMOCAP and MOSEI.

Listen

Accurately identifying human emotions during multi-person conversations is essential for developing intuitive artificial intelligence systems, such as virtual digital assistants. However, conversational emotion recognition remains difficult because a speaker's emotional state fluctuates based on both the overarching conversation topic and the immediate back-and-forth interactions between participants. Additionally, human emotion is inherently multimodal, requiring systems to interpret complementary cues from text, speech audio, and facial video simultaneously.

The article introduces and evaluates a novel artificial intelligence architecture named COGMEN, which combines contextual language models and graph-based network processing. The main objective of the article is to demonstrate how simultaneously capturing full conversational context alongside immediate speaker-to-speaker interactions enhances multimodal emotion recognition.

To evaluate the system, the authors conducted empirical experiments using two benchmark conversational datasets: IEMOCAP, which contains multi-speaker dialogue videos categorized into emotional states, and CMU-MOSEI, a large-scale multimodal sentiment and emotion dataset. The approach combines a transformer network to extract global contextual features from concatenated audio, visual, and textual inputs with a graph neural network framework that explicitly maps internal speaker continuity and inter-speaker reactions across surrounding utterances.

The findings show that COGMEN establishes new performance benchmarks. On the IEMOCAP four-emotion classification task, the model achieved an 84.5% weighted F1-score, representing a 7.7 percentage point improvement over previous state-of-the-art approaches. On the IEMOCAP six-emotion benchmark, the system reached 68.2% accuracy and a 67.6% F1-score, outperforming existing multimodal models across several challenging categories. On CMU-MOSEI, COGMEN achieved the highest binary sentiment classification accuracy at 85.0% and outperformed competitive baselines across multi-label emotion tasks. Furthermore, ablation experiments confirmed that removing the relational graph structure or reducing conversational context significantly decreased model performance.

These results demonstrate that combining broad dialogue context with localized speaker-relationship modeling improves the reliability of emotion recognition in complex interactions. While previous multimodal models often suffered performance degradation when adding noisy visual data, the proposed graph architecture successfully leveraged multi-modal features without requiring complex, computationally expensive fusion schemes. This improves classification accuracy while maintaining a streamlined model training pipeline.

Before deploying this technology in operational environments, organizations should focus on several next steps. Because the current architecture relies on offline processing that analyzes both past and future utterances within a conversation, further research and development are needed to adapt the system for real-time, online streaming environments, such as live customer support or telecommunications. Implementing dynamic context buffers represents a promising direction to balance latency and accuracy. Additionally, future model iterations should incorporate dedicated mechanisms to better handle abrupt emotional transitions and distinguish between closely related emotional classes, such as excitement versus happiness.

arXiv: 2205.02455
Cover for COGMEN: COntextualized GNN based Multimodal Emotion recognitioN

Abstract

Emotions are an inherent part of human interactions, and consequently, it is imperative to develop AI systems that understand and recognize human emotions. During a conversation involving various people, a person’s emotions are influenced by the other speaker’s utterances and their own emotional state over the utterances. In this paper, we propose COntextualized Graph Neural Network based Multimodal Emotion recognitioN (COGMEN) system that leverages local information (i.e., inter/intra dependency between speakers) and global information (context). The proposed model uses Graph Neural Network (GNN) based architecture to model the complex dependencies (local and global information) in a conversation. Our model gives state-of-the-art (SOTA) results on IEMOCAP and MOSI datasets, and detailed ablation experiments show the importance of modeling information at both levels.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Proposed Model
  • 3.1 Overall Architecture
  • 4 Experiments
  • 5 Results and Analysis
  • 6 Discussion
  • 7 Conclusion and Future Work
  • 8 Acknowledgements
  • References
  • Appendix
  • A Hyperparameter Setting
  • B Dataset Analysis
  • C Evaluation Metrics
  • D Results on Modality Combinations
  • E Additional Analysis
  • F Graph Formation
  • G Discussion
  • G.1 Modality Fusing Mechanisms
  • G.2 Effect of window size in Graph Formation

Knowls

  1. Knowl 1 — COGMEN architecture for multimodal conversational emotion recognition

    model/method

    COGMEN predicts one emotion label for every utterance in a conversation by combining three modality-specific utterance representations—audio, text, and video—with both global and local conversational information. For each utterance, the modality features are concatenated and passed through a Transformer encoder without positional encodings; this produces a context-aware representation in which every utterance can attend to all utterances in the dialogue. The resulting representations are used as nodes in a directed, relation-labeled utterance graph. A Relational Graph Convolutional Network (RGCN) followed by a GraphTransformer models speaker-specific local dependencies, including both inter-speaker and intra-speaker effects. A shared feed-forward emotion classifier then predicts the label for each graph node. The architecture diagram on page 4 depicts this sequence: modality concatenation, global context extraction, graph formation, relational graph processing, and per-utterance classification.

  2. Knowl 2 — Global context extraction with a positional-free Transformer

    equation

    For a dialogue containing nn utterances, let ui(a)∈Rdau_i^{(a)}\in\mathbb{R}^{d_a}, ui(t)∈Rdtu_i^{(t)}\in\mathbb{R}^{d_t}, and ui(v)∈Rdvu_i^{(v)}\in\mathbb{R}^{d_v} denote the audio, text, and video feature vectors for utterance ii. COGMEN concatenates the available modalities as

    xi=[ui(a)⊕ui(t)⊕ui(v)]∈Rd,d=da+dt+dv,x_i=[u_i^{(a)}\oplus u_i^{(t)}\oplus u_i^{(v)}]\in\mathbb{R}^{d},\qquad d=d_a+d_t+d_v,

    and forms X=[x1,…,xn]T∈Rn×dX=[x_1,\ldots,x_n]^\mathsf{T}\in\mathbb{R}^{n\times d}. For attention head hh, with head dimension kk, the query, key, and value matrices are

    Q(h)=XWh,q,K(h)=XWh,k,V(h)=XWh,v,Q^{(h)}=XW_{h,q},\qquad K^{(h)}=XW_{h,k},\qquad V^{(h)}=XW_{h,v},

    where Wh,q,Wh,k,Wh,v∈Rd×kW_{h,q},W_{h,k},W_{h,v}\in\mathbb{R}^{d\times k}. Row-wise scaled dot-product attention is

    α(h)=softmax⁡row(Q(h)(K(h))Tk),head⁡(h)=α(h)V(h).\alpha^{(h)}=\operatorname{softmax}_{\mathrm{row}}\left(\frac{Q^{(h)}(K^{(h)})^\mathsf{T}}{\sqrt{k}}\right),\qquad \operatorname{head}^{(h)}=\alpha^{(h)}V^{(h)}.

    With HH heads and output matrix Wo∈RHk×dW_o\in\mathbb{R}^{Hk\times d}, the multi-head output is U′=[head⁡(1)⊕⋯⊕head⁡(H)]Wo∈Rn×dU'= [\operatorname{head}^{(1)}\oplus\cdots\oplus\operatorname{head}^{(H)}]W_o\in\mathbb{R}^{n\times d}. COGMEN adds a residual connection and applies layer normalization, followed by a two-layer feed-forward block:

    U=LayerNorm⁡(X+U′;γ1,β1),U=\operatorname{LayerNorm}(X+U';\gamma_1,\beta_1), Z′=ReLU⁡(UW1)W2,Z=LayerNorm⁡(U+Z′;γ2,β2),Z'=\operatorname{ReLU}(UW_1)W_2,\qquad Z=\operatorname{LayerNorm}(U+Z';\gamma_2,\beta_2),

    where W1∈Rd×mW_1\in\mathbb{R}^{d\times m}, W2∈Rm×dW_2\in\mathbb{R}^{m\times d}, mm is the feed-forward hidden size, and γ1,β1,γ2,β2∈Rd\gamma_1,\beta_1,\gamma_2,\beta_2\in\mathbb{R}^{d} are layer-normalization parameters. The iith row ziz_i of ZZ is the global context representation of utterance ii. No positional encoding is added, so the Transformer is used to expose each utterance to the entire dialogue context rather than to impose a sequential positional representation.

  3. Knowl 3 — Relation-labeled local graph over utterances

    model/method

    COGMEN represents every utterance as a node and connects it to nearby utterances using directed, typed edges. For each central utterance, the graph includes up to PP past and FF future utterances from the same speaker and from every other speaker. Same-speaker edges are intra-speaker relations, while different-speaker edges are inter-speaker relations. Each edge is additionally labeled by temporal direction: past or future. Thus, the graph distinguishes, for example, a past utterance by the same speaker from a past utterance by another speaker, and distinguishes both from their future counterparts. A self-edge is included among the intra-speaker relations. If a conversation has MM speakers, the relation inventory contains 2M22M^2 possible speaker-pair and temporal-direction types, including the two temporal types for every ordered speaker pair. The hyperparameters PP and FF control how much local context is propagated; they are independent of the Transformer’s global context.

  4. Knowl 4 — Relational GCN and GraphTransformer for local dependency modeling

    equation

    Let ziz_i be the global Transformer representation of utterance node ii, let RR be the set of directed relation types, and let Nr(i)N_r(i) be the neighbors of node ii connected by relation rr. The RGCN first computes a relation-specific aggregation:

    z~i=Θrootzi+∑r∈R∑j∈Nr(i)1∣Nr(i)∣Θrzj,\widetilde z_i=\Theta_{\mathrm{root}}z_i+\sum_{r\in R}\sum_{j\in N_r(i)}\frac{1}{|N_r(i)|}\Theta_r z_j,

    where Θroot\Theta_{\mathrm{root}} is the learnable transformation of the central node, Θr\Theta_r is the learnable transformation for relation rr, and ∣Nr(i)∣|N_r(i)| normalizes the contribution from relation rr. This layer preserves the distinction between intra-speaker, inter-speaker, past, and future dependencies.

    The resulting node features are refined by a GraphTransformer that attends only over graph neighbors. For graph hidden dimension dgd_g, one attention head can be written as

    hi=Az~i+∑j∈N(i)αijBz~j,h_i=A\widetilde z_i+\sum_{j\in N(i)}\alpha_{ij}B\widetilde z_j, αij=softmax⁡j∈N(i)((Cz~i)T(Dz~j)dg),\alpha_{ij}=\operatorname{softmax}_{j\in N(i)}\left(\frac{(C\widetilde z_i)^\mathsf{T}(D\widetilde z_j)}{\sqrt{d_g}}\right),

    where N(i)=⋃r∈RNr(i)N(i)=\bigcup_{r\in R}N_r(i) is the full neighbor set and AA, BB, CC, and DD are learnable projection matrices. COGMEN uses the multi-head extension of this dot-product attention. The RGCN supplies relation-specific local transformations, while the GraphTransformer selectively weights the connected utterances when constructing the final local representation hih_i.

  5. Knowl 5 — Shared per-utterance emotion classifier

    equation

    For each utterance node ii, COGMEN applies the same classifier to the GraphTransformer representation hih_i. Let WhW_h, WpW_p, bhb_h, and bpb_p be learnable parameters and let CC be the number of emotion classes. The hidden representation, class probabilities, and predicted class are

    ri=ReLU⁡(Whhi+bh),r_i=\operatorname{ReLU}(W_hh_i+b_h), pi=softmax⁡(Wpri+bp)∈RC,p_i=\operatorname{softmax}(W_pr_i+b_p)\in\mathbb{R}^{C}, y^i=arg⁡max⁡c∈{1,…,C}pi,c.\hat y_i=\arg\max_{c\in\{1,\ldots,C\}}p_{i,c}.

    The classifier is shared across all utterance positions and speakers, so the graph encoder—not a position-specific output layer—provides the conversational and speaker-dependent information used for each prediction.

  6. Knowl 6 — Datasets, modality features, and training configuration

    experimental setup

    COGMEN was evaluated on IEMOCAP and CMU-MOSEI using weighted F1 and accuracy. IEMOCAP was tested in both a four-class setting—anger, sadness, happiness, and neutral—and a six-class setting—anger, excited, sadness, happiness, frustrated, and neutral. The reported split contains 120 training dialogues with 5,810 utterances, listed as 5,146+6645,146+664, and 31 test dialogues with 1,623 utterances. MOSEI contains 2,249 training dialogues with 16,327 utterances, 300 validation dialogues with 1,871 utterances, and 646 test dialogues with 4,662 utterances. MOSEI includes seven sentiment classes from −3-3 to +3+3 and six emotion labels: happiness, sadness, disgust, fear, surprise, and anger.

    For IEMOCAP, COGMEN used OpenSmile audio features of dimension 100, OpenFace video features of dimension 512, and sBERT text features of dimension 768. For MOSEI, audio features had dimension 80, video features had dimension 35, and sBERT text features had dimension 768. Audio and video token-level features were averaged to the utterance level, and available modalities were fused by concatenation. Concatenation was selected because the authors’ experiments found no significant gain from more complex pairwise or cross-modal attention mechanisms.

    The IEMOCAP configuration used dropout 0.10.1, seven GNN heads, a sequence-context setting of 4, and initial learning rate 10−410^{-4}. For MOSEI, the searched configurations were: text only—dropout 0.3990.399, three GNN heads, sequence context 5, learning rate 3.3×10−33.3\times10^{-3}; audio plus text—dropout 0.1030.103, one GNN head, sequence context 2, learning rate 6.9×10−36.9\times10^{-3}; and audio plus text plus video—dropout 0.3370.337, two GNN heads, sequence context 1, learning rate 1.1×10−31.1\times10^{-3}. The IEMOCAP implementation had 55,932,052 parameters and required approximately 7 minutes for 50 epochs on an NVIDIA Tesla K80.

  7. Knowl 7 — IEMOCAP multimodal performance

    data/table

    The reported IEMOCAP comparisons on page 6 show that COGMEN achieved the best weighted F1 and accuracy among the listed six-class multimodal systems. With audio, text, and video, COGMEN obtained class-wise F1 scores of 51.9% for happiness, 81.7% for sadness, 68.6% for neutral, 66.0% for anger, 75.3% for excited, and 58.2% for frustrated, together with 68.2% accuracy and 67.6% weighted average F1. The strongest listed earlier system, Af-CAN, obtained 64.6% accuracy and 63.7% weighted F1, so COGMEN improved the weighted F1 by 3.9 percentage points in the six-class comparison.

    In the four-class IEMOCAP setting, COGMEN achieved 84.50% weighted F1, compared with 76.80% for CHFusion and 75.13% for bc-LSTM. Thus, the reported improvement over the strongest listed prior result in this setting was 7.7 percentage points. These results were obtained with the multimodal audio-plus-text-plus-video input.

  8. Knowl 8 — MOSEI sentiment and emotion results

    data/table

    The MOSEI results reported on page 7 evaluate two sentiment tasks and two emotion formulations. Two-class sentiment is binary negative-versus-positive classification measured by accuracy; seven-class sentiment uses the −3-3 through +3+3 labels and is also measured by accuracy. The six emotion labels are evaluated either as separate binary classifiers or as one multi-label classifier, both using weighted F1.

    COGMEN’s results were:

    • Text only: two-class sentiment accuracy 84.42%, seven-class sentiment accuracy 43.50%; separate binary-emotion F1 scores were happiness 69.28%, sadness 70.49%, anger 73.04%, fear 87.80%, disgust 83.69%, and surprise 85.83%; multi-label emotion F1 scores were happiness 69.92%, sadness 72.16%, anger 77.34%, fear 86.39%, disgust 86.00%, and surprise 88.27%.
    • Audio plus text: two-class sentiment accuracy 85.00%, seven-class sentiment accuracy 44.31%; separate binary-emotion F1 scores were 68.39%, 73.28%, 74.98%, 88.08%, 83.90%, and 85.35% in the order happiness, sadness, anger, fear, disgust, and surprise; multi-label emotion F1 scores were 69.62%, 72.67%, 76.93%, 86.39%, 85.35%, and 88.21% in the same order.
    • Audio plus text plus video: two-class sentiment accuracy 84.34%, seven-class sentiment accuracy 43.90%; separate binary-emotion F1 scores were 70.42%, 72.31%, 76.20%, 88.17%, 83.69%, and 85.28%; multi-label emotion F1 scores were 72.74%, 73.90%, 78.04%, 86.71%, 85.48%, and 88.37%.

    COGMEN reached the best reported two-class sentiment accuracy, 85.00%, with audio plus text. Its seven-class sentiment performance was comparable to the baselines, while it outperformed the compared systems in most individual emotion settings. Adding video did not improve binary sentiment accuracy over audio plus text, but it generally improved the multi-label emotion results.

  9. Knowl 9 — Ablations establish the value of global context, graph propagation, and relation types

    empirical result

    The ablations reported on page 7 show that both global context and local relational processing contribute to COGMEN. On IEMOCAP four-class classification with audio, text, and video, using every utterance in a dialogue gave 84.50% weighted F1. Restricting the input to 10 utterances reduced F1 to 77.43%, a 7.07-point drop, and restricting it to 3 utterances reduced F1 to 75.39%, a 9.11-point drop.

    Removing the GNN while retaining the Transformer context reduced four-class F1 from 84.50% to 80.28% with all three modalities, and reducing the graph to a single undifferentiated relation type reduced it to 79.61%. In the six-class setting with all three modalities, the corresponding values were 67.63% for the complete model, 62.96% without the GNN, and 62.13% without relation distinctions. The same pattern held for text-only and audio-plus-text inputs: for six classes, complete-model F1 was 66.00% and 65.42%, compared with 64.34% and 61.69% without the GNN; for four classes, complete-model F1 was 81.55% and 81.59%, compared with 81.18% and 80.16% without the GNN. These results support the paper’s claim that the Transformer captures global context whereas the relational graph layers provide useful local inter-speaker and intra-speaker information.

  10. Knowl 10 — Offline-context limitation and emotion-shift failure modes

    limitation

    COGMEN is an offline emotion-recognition system: its global Transformer and graph construction can use future utterances when predicting the emotion of an earlier utterance. Consequently, it cannot be directly deployed for real-time conversation processing without replacing the full dialogue with a bounded buffer. The authors’ bounded-context experiment illustrates the cost of this restriction: on IEMOCAP four-class classification with all modalities, F1 fell from 84.50% with the full dialogue to 77.43% with 10 utterances and 75.39% with 3 utterances.

    The reported error analysis also found that COGMEN often confuses similar emotions, especially happiness with excited and anger with frustrated, and tends to predict neutral for other classes because neutral examples are more frequent. Accuracy was 53.6% on utterances involving an emotion shift, compared with 74.2% when the emotion remained the same. The paper therefore identifies online inference and explicit modeling of emotional shifts as unresolved limitations and future directions.

Coverage note — Omitted the detailed confusion matrices, UMAP visualizations, utterance-masking plots, transition-count analyses, modality-combination appendix, and full window-size sweep because they are supporting diagnostics rather than additional load-bearing model components or headline results.

References

  1. 1.Harsh Agarwal, Keshav Bansal, Abhinav Joshi, and Ashutosh Modi. 2021. Shapes of emotions: Multimodal emotion recognition in conversations via emotion shifts. CoRR, abs/2112.01938.
  2. 2.AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236–2246, Melbourne, Australia. Association for Computational Linguistics.
  3. 3.Tadas Baltrusaitis, Amir Zadeh, Yao Chong Lim, and Louis-Philippe Morency. 2018. Openface 2.0: Facial behavior analysis toolkit. In 2018 13th IEEE International Conference on Automatic Face Gesture Recognition (FG 2018), pages 59–66.
  4. 4.Etienne Becht, Leland McInnes, John Healy, Charles-Antoine Dutertre, Immanuel WH Kwok, Lai Guan Ng, Florent Ginhoux, and Evan W Newell. 2019. Dimensionality reduction for visualizing single-cell data using umap. Nature biotechnology, 37(1):38–44.
  5. 5.Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jean nette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008. Iemocap: Interactive emotional dyadic motion capture database. Language resources and evaluation, 42(4):335–359.
  6. 6.Pierre Colombo, Wojciech Witon, Ashutosh Modi, James Kennedy, and Mubbasir Kapadia. 2019. Affect-driven dialog generation. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3734–3743, Minneapolis, Minnesota. Association for Computational Linguistics.
  7. 7.Comet.ML. 2021. Comet.ML home page.
  8. 8.Dragos Datcu and Leon JM Rothkrantz. 2014. Semantic audio-visual data fusion for automatic emotion recognition. Emotion recognition: a pattern analysis approach, pages 411–435.
  9. 9.Jean-Benoit Delbrouck, Noé Tits, Mathilde Brousmiche, and Stéphane Dupont. 2020. A transformer-based joint-encoding for emotion recognition and sentiment analysis. In Second Grand-Challenge and Workshop on Multimodal Language (Challenge-HML), pages 1–7, Seattle, USA. Association for Computational Linguistics.
  10. 10.Marwan Dhuheir, Abdullatif Albaseer, Emna Baccour, Aiman Erbad, Mohamed Abdallah, and Mounir Hamdi. 2021. Emotion recognition for healthcare surveillance systems using neural networks: A survey.
  11. 11.Paul Ekman. 1993. Facial expression and emotion. American psychologist, 48(4):384.
  12. 12.Florian Eyben, Martin Wöllmer, and Björn Schuller. 2010. Opensmile: The munich versatile and fast open-source audio feature extractor. In Proceedings of the 18th ACM International Conference on Multimedia, MM ’10, page 1459–1462, New York, NY, USA. Association for Computing Machinery.
  13. 13.Matthias Fey and Jan Eric Lenssen. 2019. Fast graph representation learning with pytorch geometric. CoRR, abs/1903.02428.
  14. 14.Marc Franzen, Michael Stephan Gresser, Tobias Müller, and Prof. Dr. Sebastian Mauser. 2021. Developing emotion recognition for video conference software to support people with autism.
  15. 15.Yahui Fu, Shogo Okada, Longbiao Wang, Lili Guo, Yaodong Song, Jiaxing Liu, and Jianwu Dang. 2021. Consk-gcn: Conversational semantic- and knowledge-oriented graph convolutional network for multimodal emotion recognition. In 2021 IEEE International Conference on Multimedia and Expo (ICME), pages 1–6.
  16. 16.Deepanway Ghosal, Md Shad Akhtar, Dushyant Chauhan, Soujanya Poria, Asif Ekbal, and Pushpak Bhattacharyya. 2018. Contextual inter-modal attention for multi-modal sentiment analysis. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3454–3466, Brussels, Belgium. Association for Computational Linguistics.
  17. 17.Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, and Alexander Gelbukh. 2019. DialogueGCN: A graph convolutional neural network for emotion recognition in conversation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 154–164, Hong Kong, China. Association for Computational Linguistics.
  18. 18.Tushar Goswamy, Ishika Singh, Ahsan Barkati, and Ashutosh Modi. 2020. Adapting a language model for controlled affective text generation. In Proceedings of the 28th International Conference on Computational Linguistics, pages 2787–2801, Barcelona, Spain (Online). International Committee on Computational Linguistics.
  19. 19.M. Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh, Rada Mihalcea, Louis-Philippe Morency, and Ehsan Hoque. 2021. Humor knowledge enriched transformer for understanding multimodal humor. In AAAI.
  20. 20.Devamanyu Hazarika, Soujanya Poria, Rada Mihalcea, Erik Cambria, and Roger Zimmermann. 2018a. ICON: Interactive conversational memory network for multimodal emotion detection. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2594–2604, Brussels, Belgium. Association for Computational Linguistics.
  21. 21.Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria, Louis-Philippe Morency, and Roger Zimmermann. 2018b. Conversational memory network for emotion recognition in dyadic dialogue videos. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 2122–2132, New Orleans, Louisiana. Association for Computational Linguistics.
  22. 22.Dou Hu, Lingwei Wei, and Xiaoyong Huai. 2021. Dialoguecrn: Contextual reasoning networks for emotion recognition in conversations. In ACL/IJCNLP (1), pages 7042–7052. Association for Computational Linguistics.
  23. 23.Taichi Ishiwatari, Yuki Yasuda, Taro Miyazaki, and Jun Goto. 2020. Relation-aware graph attention networks with relational position encodings for emotion recognition in conversations. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 7360–7370.
  24. 24.Sepehr Janghorbani, Ashutosh Modi, Jakob Buhmann, and Mubbasir Kapadia. 2019. Domain authoring assistant for intelligent virtual agent. AAMAS ’19, page 104–112, Richland, SC. International Foundation for Autonomous Agents and Multiagent Systems.
  25. 25.Taewoon Kim and Piek Vossen. 2021. EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa. arXiv e-prints, page arXiv:2108.12009.
  26. 26.Agata Kołakowska, Agnieszka Landowska, Mariusz Szwoch, Wioleta Szwoch, and Michal R Wrobel. 2014. Emotion recognition and its applications. In Human-Computer Systems Interaction: Backgrounds and Applications 3, pages 51–62. Springer.
  27. 27.Ayush Kumar and Jithendra Vepa. 2020. Gated mechanism for attention based multi modal sentiment analysis. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4477–4481. IEEE.
  28. 28.John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of the Eighteenth International Conference on Machine Learning, ICML ’01, page 282–289, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
  29. 29.Zheng Lian, Jianhua Tao, Bin Liu, Jian Huang, Zhanlei Yang, and Rongjun Li. 2020. Conversational emotion recognition using self-attention mechanisms and graph neural networks. In INTERSPEECH, pages 2347–2351.
  30. 30.Navonil Majumder, Devamanyu Hazarika, Alexander Gelbukh, Erik Cambria, and Soujanya Poria. 2018. Multimodal sentiment analysis using hierarchical fusion with context modeling. Knowledge-based systems, 161:124–133.
  31. 31.Navonil Majumder, Soujanya Poria, Devamanyu Hazarika, Rada Mihalcea, Alexander Gelbukh, and Erik Cambria. 2019. Dialoguernn: An attentive rnn for emotion detection in conversations. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6818–6825.
  32. 32.Brian McFee, Colin Raffel, Dawen Liang, Daniel P. W. Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015. librosa: Audio and music signal analysis in python.
  33. 33.Marvin Minsky. 2007. The Emotion Machine: Commonsense Thinking, Artificial Intelligence, and the Future of the Human Mind. SIMON & SCHUSTER.
  34. 34.Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. Pytorch: An imperative style, high-performance deep learning library. In H. Wallach, H. Larochelle, A. Beygelzimer, A. d'Alché-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Information Processing Systems 32, pages 8024–8035. Curran Associates, Inc.
  35. 35.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  36. 36.Sally Planalp, Julie Fitness, and Beverley A. Fehr. 2018. The Roles of Emotion in Relationships, 2 edition, Cambridge Handbooks in Psychology, page 256–268. Cambridge University Press.
  37. 37.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Context-dependent sentiment analysis in user-generated videos. In Proceedings of the 55th annual meeting of the association for computational linguistics (volume 1: Long papers), pages 873–883.
  38. 38.Soujanya Poria, Navonil Majumder, Rada Mihalcea, and Eduard Hovy. 2019. Emotion recognition in conversation: Research challenges, datasets, and recent advances. IEEE Access, 7:100943–100953.
  39. 39.Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 3982–3992, Hong Kong, China. Association for Computational Linguistics.
  40. 40.Michael Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. 2018. Modeling relational data with graph convolutional networks. In The Semantic Web, pages 593–607, Cham. Springer International Publishing.
  41. 41.Nicu Sebe, Ira Cohen, Theo Gevers, and Thomas S Huang. 2005. Multimodal approaches for emotion recognition: A survey. Proceedings of SPIE - The International Society for Optical Engineering, 5670:56–67. Proceedings of SPIE-IS and T Electronic Imaging - Internet Imaging VI ; Conference date: 18-01-2005 Through 20-01-2005.
  42. 42.Garima Sharma and Abhinav Dhall. 2021. A Survey on Automatic Multimodal Emotion Recognition in the Wild, pages 35–64.
  43. 43.Weizhou Shen, Junqing Chen, Xiaojun Quan, and Zhixian Xie. 2021a. Dialogxl: All-in-one xlnet for multiparty conversation emotion recognition. Proceedings of the AAAI Conference on Artificial Intelligence, 35(15):13789–13797.
  44. 44.Weizhou Shen, Siyue Wu, Yunyi Yang, and Xiaojun Quan. 2021b. Directed acyclic graph network for conversational emotion recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1551–1560, Online. Association for Computational Linguistics.
  45. 45.Dongming Sheng, Dong Wang, Ying Shen, Haitao Zheng, and Haozhuang Liu. 2020. Summarize before aggregate: A global-to-local heterogeneous graph inference network for conversational emotion recognition. In Proceedings of the 28th International Conference on Computational Linguistics, pages 4153–4163.
  46. 46.Aman Shenoy and Ashish Sardana. 2020. Multilogue-net: A context-aware RNN for multi-modal emotion detection and sentiment analysis in conversation. In Second Grand-Challenge and Workshop on Multimodal Language (Challenge-HML), pages 19–28, Seattle, USA. Association for Computational Linguistics.
  47. 47.Yunsheng Shi, Zhengjie Huang, Wenjin Wang, Hui Zhong, Shikun Feng, and Yu Sun. 2021. Masked label prediction: Unified massage passing model for semi-supervised classification. In IJCAI.
  48. 48.Aaditya Singh, Shreeshail Hingane, Saim Wani, and Ashutosh Modi. 2021a. An end-to-end network for emotion-cause pair extraction. In Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 84–91, Online. Association for Computational Linguistics.
  49. 49.Gargi Singh, Dhanajit Brahma, Piyush Rai, and Ashutosh Modi. 2021b. Fine-Grained Emotion Prediction by Modeling Emotion Definitions. In 2021 9th International Conference on Affective Computing and Intelligent Interaction (ACII), pages 1–8.
  50. 50.Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, and Rob Fergus. 2015. End-to-end memory networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 2, NIPS’15, page 2440–2448, Cambridge, MA, USA. MIT Press.
  51. 51.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
  52. 52.C Vinola and K Vimaladevi. 2015. A survey on human emotion recognition approaches, databases and applications. ELCVIA Electronic Letters on Computer Vision and Image Analysis, 14(2):24–44.
  53. 53.Tana Wang, Yaqing Hou, Dongsheng Zhou, and Qiang Zhang. 2021a. A contextual attention network for multimodal emotion recognition in conversation. In 2021 International Joint Conference on Neural Networks (IJCNN), pages 1–7.
  54. 54.Tana Wang, Yaqing Hou, Dongsheng Zhou, and Qiang Zhang. 2021b. A contextual attention network for multimodal emotion recognition in conversation. In 2021 International Joint Conference on Neural Networks (IJCNN), pages 1–7. IEEE.
  55. 55.Yan Wang, Jiayu Zhang, Jun Ma, Shaojun Wang, and Jing Xiao. 2020. Contextualized emotion recognition in conversation as sequence tagging. In Proceedings of the 21th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 186–195.
  56. 56.Martin Wollmer, Angeliki Metallinou, Florian Eyben, Bjorn Schuller, and Shrikanth Narayanan. 2010. Context-sensitive multimodal emotion recognition from speech and facial expression using bidirectional lstm modeling. In Proc. INTERSPEECH 2010, Makuhari, Japan, pages 2362–2365.
  57. 57.Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R. Salakhutdinov, and Quoc V. Le. 2019. Xlnet: Generalized autoregressive pretraining for language understanding. In Advances in Neural Information Processing Systems, volume 32.
  58. 58.Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019. Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6281–6290.
  59. 59.Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1103–1114, Copenhagen, Denmark. Association for Computational Linguistics.
  60. 60.Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018a. Memory fusion network for multiview sequential learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’18/IAAI’18/EAAI’18. AAAI Press.
  61. 61.AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018b. Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236–2246.
  62. 62.Dong Zhang, Liangqing Wu, Changlong Sun, Shoushan Li, Qiaoming Zhu, and Guodong Zhou. 2019. Modeling both context- and speaker-sensitive dependence for emotion detection in multi-speaker conversations. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019, pages 5415–5421. ijcai.org.

Citation

MLA
Joshi, A., et al. “COGMEN: COntextualized GNN Based Multimodal Emotion recognitioN”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 4148–64, https://doi.org/10.18653/v1/2022.naacl-main.306.
APA
Joshi, A., Bhat, A., Jain, A., Singh, A., & Modi, A. (2022). COGMEN: COntextualized GNN based Multimodal Emotion recognitioN. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4148–4164. https://doi.org/10.18653/v1/2022.naacl-main.306
Chicago
Joshi, A., A. Bhat, A. Jain, A. Singh, and A. Modi. 2022. “COGMEN: COntextualized GNN Based Multimodal Emotion recognitioN”. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 4148–64. https://doi.org/10.18653/v1/2022.naacl-main.306.
Harvard
Joshi, A. et al. (2022) “COGMEN: COntextualized GNN based Multimodal Emotion recognitioN”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp. 4148–4164. Available at: https://doi.org/10.18653/v1/2022.naacl-main.306.
Vancouver
1. Joshi A, Bhat A, Jain A, Singh A, Modi A (2022) COGMEN: COntextualized GNN based Multimodal Emotion recognitioN. In: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, pp 4148–4164

BibTeX

@inproceedings{joshi-etal-2022-cogmen,
    title = "{COGMEN}: {CO}ntextualized {GNN} based Multimodal Emotion recognitio{N}",
    author = "Joshi, Abhinav  and
      Bhat, Ashwani  and
      Jain, Ayush  and
      Singh, Atin  and
      Modi, Ashutosh",
    editor = "Carpuat, Marine  and
      de Marneffe, Marie-Catherine  and
      Meza Ruiz, Ivan Vladimir",
    booktitle = "Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
    month = jul,
    year = "2022",
    address = "Seattle, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.naacl-main.306/",
    doi = "10.18653/v1/2022.naacl-main.306",
    pages = "4148--4164"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/