UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

Guimin HuTing-En LinYi ZhaoGuangming LuYuchuan WuYongbin Li

article2022EMNLP270 citations

Proposes UniMSE, a generative framework that unifies multimodal sentiment analysis and emotion recognition in conversation by combining label spaces, integrating acoustic and visual signals directly into a T5 backbone, and applying inter-modality contrastive learning.

Listen

Understanding human intent and affective state through automated systems is vital for advanced conversational interfaces and automated customer service. Currently, machine learning approaches divide this challenge into two distinct tasks: evaluating general sentiment over longer periods and identifying specific, short-term emotions during conversations. Most research treats these problems in isolation, which prevents models from sharing complementary insights across text, audio, and video modalities.

The article establishes a unified framework that combines multimodal sentiment analysis and conversational emotion recognition into a single generative architecture. The researchers designed this system to evaluate whether sharing knowledge between sentiment and emotion tasks improves predictive accuracy across multiple standard benchmarks.

To achieve this, the approach transforms the two separate tasks into a shared sequence-generation problem. Audio and video features were standardized across datasets, while a universal label scheme mapped sentiment polarities, intensities, and emotion categories into a shared format using sentence-level semantic matching. The architecture embeds audio and visual signals directly into intermediate layers of a standard language model and applies contrastive learning to pull corresponding modalities of the same sample closer while pushing unrelated samples apart. The framework was evaluated across four widely used multimodal benchmarks comprising thousands of conversational and video review segments.

The experimental findings show that the unified architecture consistently outperforms existing specialized models across all tested datasets. On sentiment tasks, the framework achieved an accuracy of 85.85% to 86.9% on one benchmark and 85.86% to 87.5% on another, reflecting gains of roughly 1.1% to 1.7% over previous state-of-the-art systems. For conversational emotion recognition, classification accuracy reached 65.09% and 70.56% across the two datasets, surpassing prior methods by approximately 2.3% to 2.6%. Ablation analyses confirmed that removing non-verbal modalities or omitting the multi-task training sets systematically degraded performance, highlighting the value of acoustic cues and cross-dataset knowledge sharing.

These results demonstrate that sentiment and emotion share a functional embedding space that enhances machine learning performance when modeled together. Consolidating separate analytical pipelines into a single model can reduce architectural complexity and training overhead. However, while the system establishes new benchmark performance, overall emotion recognition accuracy remains around 65% to 71%, which is insufficient for fully autonomous, high-stakes operational deployment.

Organizations developing affective computing systems should consider adopting unified architectures over isolated models to capture cross-modal efficiencies. Before deploying these models into production, practitioners should run targeted pilot programs and expand conversational context modeling across longer dialogues to improve baseline reliability.

The study notes limitations in its label completion process, which relied primarily on text similarity rather than acoustic or visual cues, and observed that broader context was not integrated across all evaluated datasets. Consequently, while the framework reliably advances foundational research, stakeholders should maintain human-in-the-loop oversight when deploying emotion detection systems in sensitive environments.

arXiv: 2211.11256
Cover for UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

Abstract

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period. However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models. We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions. Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMOCAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Overall Architecture
  • 3.2 Task Formalization
  • 3.2.1 Input Formalization
  • 3.2.2 Label Formalization
  • 3.3 Pre-trained Modality Fusion (PMF)
  • 3.4 Inter-modality Contrastive Learning
  • 3.5 Grounding UL to MSA and ERC
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Evaluation metrics
  • 4.3 Baselines
  • 4.4 Experimental Settings
  • 4.5 Results
  • 4.6 Ablation Study
  • 4.7 Visualization
  • 5 Conclusion
  • Limitations
  • Ethics Statement
  • Acknowledgement
  • References
  • A Appendix
  • A.1 Datasets
  • A.2 Decoding Algorithm for MSA and ERC tasks
  • A.3 Experimental Environment

Knowls

  1. Knowl 1 — Universal Label Scheme and Task Formalization for Unified MSA and ERC

    model/method

    The UniMSE framework unifies Multimodal Sentiment Analysis (MSA) and Emotion Recognition in Conversations (ERC) into a single sequence-to-sequence generative format by defining a Universal Label (UL) target sequence:

    yi={yip,yir,yic}y_i = \{y_i^p, y_i^r, y_i^c\}

    where yip∈{positive,negative,neutral}y_i^p \in \{\text{positive}, \text{negative}, \text{neutral}\} denotes sentiment polarity, yir∈[−3,+3]y_i^r \in [-3, +3] is a real-valued sentiment intensity score (the primary supervisory signal of MSA), and yicy_i^c is a categorical emotion label (the primary supervisory signal of ERC, such as joy, sadness, anger, fear, surprise, or disgust).

    Because standard MSA datasets lack emotion category annotations and standard ERC datasets lack continuous sentiment intensity scores, missing labels are completed offline using sentence semantic similarity:

    1. All MSA and ERC samples are split into positive, neutral, and negative subsets based on sentiment polarity.
    2. For an MSA sample mm lacking an emotion label, its textual embedding is compared against the textual embeddings of ERC samples in the same polarity subset using SimCSE cosine similarity. The categorical emotion ycy^c of the most similar ERC sample is assigned to mm.
    3. Conversely, for an ERC sample ee lacking a sentiment intensity score, the score yry^r of the most similar MSA sample in the corresponding polarity subset is assigned to ee.

    For conversational inputs, textual modality context is incorporated by concatenating the current utterance uiu_i with its preceding two utterances and following two utterances into Iit=[ui−2,ui−1,ui,ui+1,ui+2]I_i^t = [u_{i-2}, u_{i-1}, u_i, u_{i+1}, u_{i+2}], paired with a segment mask Sit=[0,…,0,1,…,1,0,…,0]S_i^t = [0, \dots, 0, 1, \dots, 1, 0, \dots, 0] where only tokens corresponding to uiu_i are assigned the value 11.

  2. Knowl 2 — Pre-trained Modality Fusion (PMF) Layer Formulation

    model/method

    To fuse non-verbal acoustic and visual representations with multi-level textual representations inside a pre-trained sequence-to-sequence Transformer (T5), UniMSE integrates adapter-style Pre-trained Modality Fusion (PMF) layers following the feed-forward sublayers of the encoder.

    Let Xit∈Rlt×dtX_i^t \in \mathbb{R}^{l_t \times d_t} be the textual representation from the first Transformer layer (Fi(0)=XitF_i^{(0)} = X_i^t). Acoustic features (Mel-spectrograms) and visual features (facial representations) are encoded by two separate LSTM networks to yield sequence representations Xia∈Rla×daX_i^a \in \mathbb{R}^{l_a \times d_a} and Xiv∈Rlv×dvX_i^v \in \mathbb{R}^{l_v \times d_v}, with their final time-step hidden vectors denoted as Xia,la∈R1×daX_i^{a, l_a} \in \mathbb{R}^{1 \times d_a} and Xiv,lv∈R1×dvX_i^{v, l_v} \in \mathbb{R}^{1 \times d_v}.

    For the jj-th PMF layer, multimodal fusion is computed as:

    Fi=[Fi(j−1)⋅Xia,la⋅Xiv,lv]F_i = [F_i^{(j-1)} \cdot X_i^{a, l_a} \cdot X_i^{v, l_v}]

    Fid=σ(WdFi+bd)F_i^d = \sigma(W^d F_i + b^d)

    Fiu=WuFid+buF_i^u = W^u F_i^d + b^u

    Fi(j)=W(Fiu⊙Fi(j−1))F_i^{(j)} = W(F_i^u \odot F_i^{(j-1)})

    where [⋅][\cdot] denotes concatenation along the feature dimension, σ(⋅)\sigma(\cdot) is the Sigmoid activation function, ⊙\odot represents element-wise addition, and {Wd,Wu,W,bd,bu}\{W^d, W^u, W, b^d, b^u\} are learnable projection weights and biases. The resulting fusion representation Fi(j)F_i^{(j)} is passed to layer normalization.

    To avoid disrupting early linguistic text representations and prevent parameter overfitting, the first N−3N - 3 Transformer layers of the T5 encoder process text unimodally, while PMF units are inserted only into the final 3 Transformer layers of the encoder.

  3. Knowl 3 — Inter-Modality Contrastive Learning Objective

    equation

    UniMSE applies self-supervised inter-modality contrastive learning on the representations of the last 3 encoder Transformer layers to narrow the representation distance between modalities of the same instance while pushing representations of different instances apart.

    To align temporal sequence lengths across modalities, acoustic representations XiaX_i^a, visual representations XivX_i^v, and fused representations Fi(j)F_i^{(j)} are projected through 1D temporal convolutional layers:

    X^iu=Conv1D(Xiu,ku),u∈{a,v}\hat{X}_i^u = \text{Conv1D}(X_i^u, k^u), \quad u \in \{a, v\}

    F^i(j)=Conv1D(Fi(j),kf)\hat{F}_i^{(j)} = \text{Conv1D}(F_i^{(j)}, k^f)

    where kuk^u and kfk^f are the convolutional kernel sizes for modality uu and the fusion modality, respectively.

    In a mini-batch of KK samples, text modality acts as the anchor. For sample ii, the positive pairs are (F^i(j),X^ia)(\hat{F}_i^{(j)}, \hat{X}_i^a) and (F^i(j),X^iv)(\hat{F}_i^{(j)}, \hat{X}_i^v), while negative pairs are formed using acoustic and visual representations from the other K−1K-1 samples in the batch. The text-acoustic contrastive loss Lta,jL^{ta,j} and text-visual contrastive loss Ltv,jL^{tv,j} at the jj-th Transformer layer are defined as:

    Lta,j=−log⁡exp⁡(F^i(j)X^ia)exp⁡(F^i(j)X^ia)+∑k=1Kexp⁡(F^i(j)X^ka)L^{ta,j} = -\log \frac{\exp(\hat{F}_i^{(j)} \hat{X}_i^a)}{\exp(\hat{F}_i^{(j)} \hat{X}_i^a) + \sum_{k=1}^K \exp(\hat{F}_i^{(j)} \hat{X}_k^a)}

    Ltv,j=−log⁡exp⁡(F^i(j)X^iv)exp⁡(F^i(j)X^iv)+∑k=1Kexp⁡(F^i(j)X^kv)L^{tv,j} = -\log \frac{\exp(\hat{F}_i^{(j)} \hat{X}_i^v)}{\exp(\hat{F}_i^{(j)} \hat{X}_i^v) + \sum_{k=1}^K \exp(\hat{F}_i^{(j)} \hat{X}_k^v)}

  4. Knowl 4 — UniMSE Joint Multi-Task Training Loss

    equation

    The training objective of UniMSE optimizes the generative sequence-to-sequence loss alongside inter-modal contrastive learning losses across the selected Transformer layers of the encoder:

    L=Ltask+α∑jLta,j+β∑jLtv,jL = L^{\text{task}} + \alpha \sum_j L^{ta,j} + \beta \sum_j L^{tv,j}

    where:

    • LtaskL^{\text{task}} is the negative log-likelihood loss for autoregressively generating the target universal label sequence yi={yip,yir,yic}y_i = \{y_i^p, y_i^r, y_i^c\}.
    • jj indexes the encoder Transformer layers equipped with modality fusion (specifically the last 3 layers of the T5 encoder).
    • Lta,jL^{ta,j} and Ltv,jL^{tv,j} are the text-acoustic and text-visual contrastive losses at layer jj.
    • α\alpha and β\beta are weighting hyperparameters balancing the contrastive objectives, set to α=0.5\alpha = 0.5 and β=0.5\beta = 0.5.
  5. Knowl 5 — Decoding Universal Label Sequences into Task Predictions

    algorithm

    During inference, the generated Universal Label sequence yi=(yip,yir,yic)y_i = (y_i^p, y_i^r, y_i^c) output by UniMSE is parsed to produce predictions tailored to either Multimodal Sentiment Analysis (MSA) or Emotion Recognition in Conversations (ERC).

    Input: Target task t∈{MSA,ERC}t \in \{\text{MSA}, \text{ERC}\}, predicted universal label sequences Y={y1,y2,…,yN}Y = \{y_1, y_2, \dots, y_N\} where each yi=(yip,yir,yic)y_i = (y_i^p, y_i^r, y_i^c)
    Output: Task-specific predictions $Y^t = \{y_1^t, y_2^t, \dots, y_N^t\}
    Yt={}Y^t = \{\}
    for each yiy_i in YY do
        yir=yi[1]y_i^r = y_i[1]
        yic=yi[2]y_i^c = y_i[2]
        if tt is MSA then
            Yt.append(yir)Y^t.\text{append}(y_i^r)
        end
        if tt is ERC then
            Yt.append(yic)Y^t.\text{append}(y_i^c)
        end
    end
    return YtY^t

    For MSA, the predicted real-valued token yiry_i^r is parsed as the continuous sentiment intensity score. For ERC, the predicted emotion category token yicy_i^c is extracted as the categorical classification.

  6. Knowl 6 — Performance Comparison on MSA and ERC Benchmark Datasets

    data/table

    UniMSE was evaluated on four multimodal benchmarks: MOSI and MOSEI for Multimodal Sentiment Analysis (MSA), and MELD and IEMOCAP for Emotion Recognition in Conversations (ERC). Metrics include Mean Absolute Error (MAE ↓\downarrow), Pearson Correlation (Corr ↑\uparrow), seven-class accuracy (ACC-7 ↑\uparrow), binary accuracy (ACC-2 ↑\uparrow, reported as non-negative/negative and positive/negative), F1 score (↑\uparrow), standard classification accuracy (ACC ↑\uparrow), and weighted F1 (WF1 ↑\uparrow).

    MOSI MOSEI MELD IEMOCAP
    Method MAE Corr ACC-7 ACC-2 F1 MAE Corr ACC-7 ACC-2 F1 ACC WF1 ACC WF1
    LMF 0.917 0.695 33.20 -/82.5 -/82.4 0.623 0.700 48.00 -/82.0 -/82.1 61.15 58.30 56.50 56.49
    TFN 0.901 0.698 34.90 -/80.8 -/80.7 0.593 0.677 50.20 -/82.5 -/82.1 60.70 57.74 55.02 55.13
    MFM 0.877 0.706 35.40 -/81.7 -/81.6 0.568 0.703 51.30 -/84.4 -/84.3 60.80 57.80 61.24 61.60
    MTAG 0.866 0.722 38.90 -/82.3 -/82.1 - - - - - - - - -
    MulT 0.861 0.711 - 81.50/84.10 80.60/83.90 0.580 0.713 - -/82.5 -/82.3 - - - -
    MISA 0.804 0.764 - 80.79/82.10 80.77/82.03 0.568 0.717 - 82.59/84.23 82.67/83.97 - - - -
    COGMEN - - 43.90 -/84.34 - - - - - - - - 68.20 67.63
    Self-MM 0.713 0.798 - 84.00/85.98 84.42/85.95 0.530 0.765 - 82.81/85.17 82.53/85.30 - - - -
    MAG-BERT 0.712 0.796 - 84.20/86.10 84.10/86.00 - - - 84.70/- 84.50/- - - - -
    MMIM 0.700 0.800 46.65 84.14/86.06 84.00/85.98 0.526 0.772 54.24 82.24/85.97 82.66/85.94 - - - -
    DialogueGCN - - - - - - - - - - 59.46 58.10 65.25 64.18
    DAG-ERC - - - - - - - - - - - 63.65 - 68.03
    COSMIC - - - - - - - - - - - 65.21 - 65.28
    MM-DFN - - - - - - - - - - 62.49 59.46 68.21 68.18
    UniMSE 0.691 0.809 48.68 85.85/86.90 85.83/86.42 0.523 0.773 54.39 85.86/87.50 85.79/87.46 65.09 65.51 70.56 70.66

    UniMSE outperforms previous state-of-the-art models on all evaluated benchmarks, achieving improvements of +1.65% ACC-2 on MOSI, +1.16% ACC-2 on MOSEI, +2.60% ACC on MELD, and +2.35% ACC on IEMOCAP.

  7. Knowl 7 — Ablation Analysis of Modalities, Modules, and Datasets in UniMSE

    data/table

    An ablation study evaluated on the MOSI test set demonstrates the contribution of each modality (Acoustic AA, Visual VV), architectural components (Pre-trained Modality Fusion PMF, Contrastive Learning CL), and cross-dataset multi-task training data (IEMOCAP, MELD, MOSEI).

    Configuration MAE ↓\downarrow Corr ↑\uparrow ACC-2 ↑\uparrow F1 ↑\uparrow
    UniMSE (Full Model) 0.691 0.809 85.85 / 86.90 85.83 / 86.42
    – w/o A 0.719 0.794 83.82 / 85.20 83.86 / 85.69
    – w/o V 0.714 0.798 84.37 / 85.37 84.71 / 85.78
    – w/o A, V 0.721 0.780 83.72 / 85.11 83.52 / 85.11
    – w/o PMF 0.722 0.785 85.13 / 86.59 85.03 / 86.37
    – w/o CL 0.713 0.795 85.28 / 86.59 85.27 / 86.55
    – w/o IEMOCAP 0.718 0.784 84.11 / 85.88 84.75 / 85.47
    – w/o MELD 0.722 0.776 84.05 / 84.96 84.50 / 84.64
    – w/o MOSEI 0.775 0.727 80.68 / 81.22 81.35 / 81.83

    Key takeaways:

    • Removing non-verbal modalities impairs performance, with acoustic information showing greater relative importance than visual information on MOSI (0.7190.719 vs. 0.7140.714 MAE).
    • Removing PMF or CL worsens both MAE and correlation, confirming their role in multimodal fusion and discriminative representation.
    • Removing external datasets during unified training degrades performance on MOSI, showing that knowledge transfer from ERC datasets (MELD, IEMOCAP) and large-scale MSA data (MOSEI) benefits MSA performance.
  8. Knowl 8 — Video Segment Duration Disparity Between MSA and ERC

    empirical result

    Empirical measurement of utterance video segment durations supports psychological theories defining emotions as short-term affective reactions and sentiments as longer-term dispositions:

    Task Dataset Dataset Average Length (s) Task Average Length (s)
    MSA MOSI 4.2 7.3
    MOSEI 7.6
    ERC MELD 3.2 3.7
    IEMOCAP 4.6

    Across the benchmark datasets, the average duration of video segments in MSA tasks (7.3 s7.3\text{ s}) is nearly double that of ERC tasks (3.7 s3.7\text{ s}). The longer duration of MOSEI segments (7.6 s7.6\text{ s}) aligns with its primary utility for sentiment intensity modeling rather than transient emotion recognition.

  9. Knowl 9 — Multimodal Feature Extraction and Implementation Configuration

    experimental setup

    The UniMSE framework employs a standardized feature extraction and training pipeline:

    • Text Backbone: Pre-trained T5-Base (220M220\text{M} parameters, 12 layers, hidden size 768768, 12 attention heads).
    • Audio Features: Extracted as Mel-spectrogram power spectrum representations using Librosa and encoded via an LSTM into hidden dimension 6464.
    • Visual Features: A fixed number of TT frames per segment are extracted, encoded with EfficientNet pre-trained on VGGFace and AFEW datasets, and passed through an LSTM into hidden dimension 6464.
    • Fusion Layer Size: Fused vector dimension is 768768.
    • Optimization: Training batches are constructed by combining the training sets of MOSI, MOSEI, MELD, and IEMOCAP. Batch size is 9696. The learning rate for T5 fine-tuning is 3×10−43 \times 10^{-4}, and the learning rate for the main network and PMF adapters is 1×10−41 \times 10^{-4}.
  10. Knowl 10 — Limitations of UniMSE

    limitation

    The UniMSE architecture and evaluation exhibit three primary limitations identified by the authors:

    1. Conversational Context Restriction: Conversational context windowing (IitI_i^t using former and latter 2 turns) was integrated only for the dialogue datasets (MELD and IEMOCAP), leaving utterance-level context integration for MOSI and MOSEI unaddressed.
    2. Text-Only Universal Label Completion: Offline pseudo-label completion relies exclusively on textual semantic similarity computed via SimCSE sentence embeddings, without incorporating acoustic or visual modal similarities into the label alignment step.
    3. Practical Accuracy Barrier: Even with unified multi-task training, six-class emotion recognition accuracy on MELD remains at 65.09%65.09\%, indicating that model precision is still limited for downstream practical deployments such as emotional companion robots or automated customer service.

Coverage note — None was omitted; all key contributions including framework design, formalizations, equations, empirical findings, ablations, setup, and limitations are fully covered.

References

  1. 1.Md. Shad Akhtar, Dushyant Singh Chauhan, Deepanway Ghosal, Soujanya Poria, Asif Ekbal, and Pushpak Bhattacharyya. 2019. Multi-task learning for multi-modal emotion recognition and sentiment analysis. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 370–379. Association for Computational Linguistics.
  2. 2.Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer normalization. CoRR, abs/1607.06450.
  3. 3.Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2018. Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443.
  4. 4.C Daniel Batson, Laura L Shaw, and Kathryn C Oleson. 1992. Differentiating affect, mood, and emotion: Toward functionally based conceptual distinctions.
  5. 5.Aaron Ben-Ze’ev. 2001. The subtlety of emotions. MIT press.
  6. 6.Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N. Chang, Sungbok Lee, and Shrikanth S. Narayanan. 2008. IEMOCAP: interactive emotional dyadic motion capture database. Lang. Resour. Evaluation, 42(4):335–359.
  7. 7.Dushyant Singh Chauhan, Md. Shad Akhtar, Asif Ekbal, and Pushpak Bhattacharyya. 2019. Context-aware interactive attention for multi-modal sentiment and emotion analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 5646–5656. Association for Computational Linguistics.
  8. 8.Zhi Chen, Lu Chen, Bei Chen, Libo Qin, Yuncong Liu, Su Zhu, Jian-Guang Lou, and Kai Yu. 2022. Unidu: Towards A unified generative dialogue understanding framework. CoRR, abs/2204.04637.
  9. 9.Junyan Cheng, Iordanis Fostiropoulos, Barry W. Boehm, and Mohammad Soleymani. 2021a. Multimodal phased transformer for sentiment analysis. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 2447–2458. Association for Computational Linguistics.
  10. 10.Kewei Cheng, Ziqing Yang, Ming Zhang, and Yizhou Sun. 2021b. Uniker: A unified framework for combining embedding and definite horn rule reasoning for knowledge graph inference. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 9753–9771. Association for Computational Linguistics.
  11. 11.Yinpei Dai, Hangyu Li, Yongbin Li, Jian Sun, Fei Huang, Luo Si, and Xiaodan Zhu. 2021. Preview, attend and review: Schema-aware curriculum learning for multi-domain dialogue state tracking. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 879–885.
  12. 12.Yinpei Dai, Hangyu Li, Chengguang Tang, Yongbin Li, Jian Sun, and Xiaodan Zhu. 2020a. Learning low-resource end-to-end goal-oriented dialog for fast and reliable system deployment. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 609–618.
  13. 13.Yinpei Dai, Huihua Yu, Yixuan Jiang, Chengguang Tang, Yongbin Li, and Jian Sun. 2020b. A survey on dialog management: Recent advances and challenges. arXiv preprint arXiv:2005.02233.
  14. 14.Richard J Davidson, Klaus R Sherer, and H Hill Goldsmith. 2009. Handbook of affective sciences. Oxford University Press.
  15. 15.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6894–6910. Association for Computational Linguistics.
  16. 16.Deepanway Ghosal, Md. Shad Akhtar, Dushyant Singh Chauhan, Soujanya Poria, Asif Ekbal, and Pushpak Bhattacharyya. 2018. Contextual inter-modal attention for multi-modal sentiment analysis. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, October 31 - November 4, 2018, pages 3454–3466. Association for Computational Linguistics.
  17. 17.Deepanway Ghosal, Navonil Majumder, Alexander F. Gelbukh, Rada Mihalcea, and Soujanya Poria. 2020. COSMIC: commonsense knowledge for emotion identification in conversations. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, pages 2470–2481.
  18. 18.Deepanway Ghosal, Navonil Majumder, Soujanya Poria, Niyati Chhaya, and Alexander F. Gelbukh. 2019. Dialoguegcn: A graph convolutional neural network for emotion recognition in conversation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 154–164. Association for Computational Linguistics.
  19. 19.Michael Gutmann and Aapo Hyvarinen. 2010. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2010, Chia Laguna Resort, Sardinia, Italy, May 13-15, 2010, pages 297–304.
  20. 20.Wei Han, Hui Chen, and Soujanya Poria. 2021. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 9180–9192. Association for Computational Linguistics.
  21. 21.Devamanyu Hazarika, Soujanya Poria, Roger Zimmermann, and Rada Mihalcea. 2019. Emotion recognition in conversations with transfer learning from generative conversation modeling. CoRR, abs/1910.04980.
  22. 22.Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria. 2020. MISA: modality-invariant and -specific representations for multimodal sentiment analysis. In MM ’20: The 28th ACM International Conference on Multimedia, Virtual Event / Seattle, WA, USA, October 12-16, 2020, pages 1122–1131. ACM.
  23. 23.Wanwei He, Yinpei Dai, Binyuan Hui, Min Yang, Zheng Cao, Jianbo Dong, Fei Huang, Luo Si, and Yongbin Li. 2022a. Space-2: Tree-structured semi-supervised contrastive pre-training for task-oriented dialog understanding. In Proceedings of the 29th International Conference on Computational Linguistics.
  24. 24.Wanwei He, Yinpei Dai, Min Yang, Jian Sun, Fei Huang, Luo Si, and Yongbin Li. 2022b. Space-3: Unified dialog model pre-training for task-oriented dialog understanding and generation. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 187–200.
  25. 25.Wanwei He, Yinpei Dai, Yinhe Zheng, Yuchuan Wu, Zheng Cao, Dermot Liu, Peng Jiang, Min Yang, Fei Huang, Luo Si, et al. 2022c. Space: A generative pre-trained model for task-oriented dialog with semi-supervised learning and explicit policy injection. Proceedings of the AAAI Conference on Artificial Intelligence.
  26. 26.Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, pages 2790–2799.
  27. 27.Dou Hu, Xiaolong Hou, Lingwei Wei, Lian-Xin Jiang, and Yang Mo. 2022. MM-DFN: multimodal dynamic fusion network for emotion recognition in conversations. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2022, Virtual and Singapore, 23-27 May 2022, pages 7037–7041.
  28. 28.Guimin Hu, Guangming Lu, and Yi Zhao. 2021a. Bidirectional hierarchical attention networks based on document-level context for emotion cause extraction. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 558–568.
  29. 29.Guimin Hu, Guangming Lu, and Yi Zhao. 2021b. FSS-GCN: A graph convolutional networks with fusion of semantic and structure for emotion cause analysis. Knowl. Based Syst., 212:106584.
  30. 30.Jingwen Hu, Yuchen Liu, Jinming Zhao, and Qin Jin. 2021c. MMGCN: multimodal fusion via deep graph convolution network for emotion recognition in conversation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 5666–5675. Association for Computational Linguistics.
  31. 31.Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Vikram Singh, and Ashutosh Modi. 2022. COGMEN: contextualized GNN based multimodal emotion recognition. CoRR, abs/2205.02455.
  32. 32.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised contrastive learning. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  33. 33.Joosung Lee and Wooin Lee. 2021. Compm: Context modeling with speaker’s pre-trained memory tracking for emotion recognition in conversation. CoRR, abs/2108.11626.
  34. 34.Jiangnan Li, Zheng Lin, Peng Fu, and Weiping Wang. 2021a. Past, present, and future: Conversational emotion recognition through structural modeling of psychological knowledge. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 1204–1214. Association for Computational Linguistics.
  35. 35.Shimin Li, Hang Yan, and Xipeng Qiu. 2021b. Contrast and generation make BART a good dialogue emotion recognizer. CoRR, abs/2112.11202.
  36. 36.Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2022. Foundations and recent trends in multimodal machine learning: Principles, challenges, and open questions. arXiv preprint arXiv:2209.03430.
  37. 37.Ting-En Lin, Yuchuan Wu, Fei Huang, Luo Si, Jian Sun, and Yongbin Li. 2022. Duplex conversation: Towards human-like interaction in spoken dialogue system. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery & Data Mining.
  38. 38.Ting-En Lin and Hua Xu. 2019a. Deep unknown intent detection with margin loss. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5491–5496.
  39. 39.Ting-En Lin and Hua Xu. 2019b. A post-processing method for detecting unknown intent of dialogue system via pre-trained deep neural network classifier. Knowledge-Based Systems, 186:104979.
  40. 40.Ting-En Lin, Hua Xu, and Hanlei Zhang. 2020. Discovering new intents via constrained deep adaptive clustering with cluster refinement. In Proceedings of AAAI, pages 8360–8367.
  41. 41.Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2018. Efficient low-rank multimodal fusion with modality-specific factors. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 2247–2256. Association for Computational Linguistics.
  42. 42.Huaishao Luo, Lei Ji, Yanyong Huang, Bin Wang, Shenggong Ji, and Tianrui Li. 2021. Scalevlad: Improving multimodal sentiment analysis via multi-scale fusion of locally descriptors. CoRR, abs/2112.01368.
  43. 43.Sijie Mai, Haifeng Hu, and Songlong Xing. 2020. Modality to modality translation: An adversarial representation learning and graph fusion network for multimodal fusion. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 164–172. AAAI Press.
  44. 44.Huisheng Mao, Ziqi Yuan, Hua Xu, Wenmeng Yu, Yihe Liu, and Kai Gao. 2022. M-sena: An integrated platform for multimodal sentiment analysis. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 204–213.
  45. 45.Yuzhao Mao, Guang Liu, Xiaojie Wang, Weiguo Gao, and Xuan Li. 2021. Dialoguetrm: Exploring multimodal emotional dynamics in a conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 2694–2704.
  46. 46.Louis-Philippe Morency, Rada Mihalcea, and Payal Doshi. 2011. Towards multimodal sentiment analysis: harvesting opinions from the web. In Proceedings of the 13th International Conference on Multimodal Interfaces, ICMI 2011, Alicante, Spain, November 14-18, 2011, pages 169–176.
  47. 47.Myriam Munezero, Calkin Suero Montero, Erkki Sutinen, and John Pajunen. 2014. Are they different? affect, feeling, emotion, sentiment, and opinion detection in text. IEEE Trans. Affect. Comput., 5(2):101–111.
  48. 48.Henry Alexander Murray and Christiana D Morgan. 1945. A clinical study of sentiments (i & ii). Genetic Psychology Monographs.
  49. 49.Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng. 2011. Multimodal deep learning. In Proceedings of the 28th International Conference on Machine Learning, ICML 2011, Bellevue, Washington, USA, June 28 - July 2, 2011, pages 689–696. Omnipress.
  50. 50.Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018. Deep contextualized word representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2018, New Orleans, Louisiana, USA, June 1-6, 2018, Volume 1 (Long Papers), pages 2227–2237. Association for Computational Linguistics.
  51. 51.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Multi-level multiple attentions for contextual multimodal sentiment analysis. In 2017 IEEE International Conference on Data Mining, ICDM 2017, New Orleans, LA, USA, November 18-21, 2017, pages 1033–1038.
  52. 52.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. MELD: A multimodal multi-party dataset for emotion recognition in conversations. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 527–536. Association for Computational Linguistics.
  53. 53.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67.
  54. 54.Wasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh, Chengfeng Mao, Louis-Philippe Morency, and Mohammed E. Hoque. 2020. Integrating multimodal information in large pre-trained transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 2359–2369.
  55. 55.Robert K Shelly. 2004. Emotions, sentiments, and performance expectations. In Theory and research on human emotions. Emerald Group Publishing Limited.
  56. 56.Weizhou Shen, Siyue Wu, Yunyi Yang, and Xiaojun Quan. 2021. Directed acyclic graph network for conversational emotion recognition. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 1551–1560. Association for Computational Linguistics.
  57. 57.Jan E Stets. 2006. Emotions and sentiments. In Handbook of social psychology, pages 309–335. Springer.
  58. 58.Yang Sun, Nan Yu, and Guohong Fu. 2021. A discourse-aware graph neural network for emotion recognition in multi-party conversation. In Findings of the Association for Computational Linguistics: EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 16-20 November, 2021, pages 2949–2958. Association for Computational Linguistics.
  59. 59.Zhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, and Yingyu Liang. 2020. Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 8992–8999.
  60. 60.Mingxing Tan and Quoc V. Le. 2019. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 6105–6114. PMLR.
  61. 61.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019a. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 6558–6569. Association for Computational Linguistics.
  62. 62.Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019b. Learning factorized multimodal representations. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  63. 63.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  64. 64.Wenhui Wang, Hangbo Bao, Li Dong, and Furu Wei. 2021a. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. CoRR, abs/2111.02358.
  65. 65.Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2019. Words can shift: Dynamically adjusting word representations using nonverbal behaviors. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pages 7216–7223.
  66. 66.Yijun Wang, Changzhi Sun, Yuanbin Wu, Hao Zhou, Lei Li, and Junchi Yan. 2021b. Unire: A unified label space for entity relation extraction. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 220–231.
  67. 67.Tianbao Xie, Chen Henry Wu, Peng Shi, Ruiqi Zhong, Torsten Scholak, Michihiro Yasunaga, Chien-Sheng Wu, Ming Zhong, Pengcheng Yin, Sida I. Wang, Victor Zhong, Bailin Wang, Chengzu Li, Connor Boyle, Ansong Ni, Ziyu Yao, Dragomir R. Radev, Caiming Xiong, Lingpeng Kong, Rui Zhang, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2022. Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models. CoRR, abs/2201.05966.
  68. 68.Hang Yan, Junqi Dai, Tuo Ji, Xipeng Qiu, and Zheng Zhang. 2021a. A unified generative framework for aspect-based sentiment analysis. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 2416–2429.
  69. 69.Hang Yan, Tao Gui, Junqi Dai, Qipeng Guo, Zheng Zhang, and Xipeng Qiu. 2021b. A unified generative framework for various NER subtasks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 5808–5822. Association for Computational Linguistics.
  70. 70.Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, and Louis-Philippe Morency. 2021. MTAG: modal-temporal attention graph for unaligned human multimodal language sequences. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 1009–1021.
  71. 71.Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu. 2021a. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021, pages 10790–10797. AAAI Press.
  72. 72.Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu. 2021b. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10790–10797.
  73. 73.Ziqi Yuan, Wei Li, Hua Xu, and Wenmeng Yu. 2021. Transformer-based feature reconstruction network for robust multimodal sentiment analysis. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4400–4407.
  74. 74.Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, EMNLP 2017, Copenhagen, Denmark, September 9-11, 2017, pages 1103–1114. Association for Computational Linguistics.
  75. 75.Amir Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2018, Volume 1: Long Papers, pages 2236–2246. Association for Computational Linguistics.
  76. 76.Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages. IEEE Intell. Syst., 31(6):82–88.
  77. 77.Hanlei Zhang, Hua Xu, and Ting-En Lin. 2021a. Deep open intent classification with adaptive decision boundary. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14374–1438.
  78. 78.Hanlei Zhang, Hua Xu, Ting-En Lin, and Rui Lyu. 2021b. Discovering new intents with deep aligned clustering. Proceedings of the AAAI Conference on Artificial Intelligence, 35(16):14365–14373.
  79. 79.Hanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou, Shaojie Zhao, and Jiayan Teng. 2022a. Mintrec: A new dataset for multimodal intent recognition. arXiv preprint arXiv:2209.04355.
  80. 80.Sai Zhang, Yuwei Hu, Yuchuan Wu, Jiaman Wu, Yongbin Li, Jian Sun, Caixia Yuan, and Xiaojie Wang. 2022b. A slot is not built in one utterance: Spoken language dialogs with sub-slots. In Findings of the Association for Computational Linguistics: ACL 2022, pages 309–321, Dublin, Ireland. Association for Computational Linguistics.
  81. 81.Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang, Qun Liu, and Zhenglu Yang. 2021c. Unims: A unified framework for multimodal summarization with knowledge distillation. CoRR, abs/2109.05812.
  82. 82.Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang, Qun Liu, and Zhenglu Yang. 2022c. Unims: A unified framework for multimodal summarization with knowledge distillation. In Thirty-Sixth AAAI Conference on Artificial Intelligence, AAAI 2022, Thirty-Fourth Conference on Innovative Applications of Artificial Intelligence, IAAI 2022, The Twelveth Symposium on Educational Advances in Artificial Intelligence, EAAI 2022 Virtual Event, February 22 - March 1, 2022, pages 11757–11764.
  83. 83.Lixing Zhu, Gabriele Pergola, Lin Gui, Deyu Zhou, and Yulan He. 2021. Topic-driven and knowledge-aware transformer for dialogue emotion detection. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 1571–1582.

Citation

MLA
Hu, G., et al. “UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 7837–51, https://doi.org/10.18653/v1/2022.emnlp-main.534.
APA
Hu, G., Lin, T.-E., Zhao, Y., Lu, G., Wu, Y., & Li, Y. (2022). UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 7837–7851. https://doi.org/10.18653/v1/2022.emnlp-main.534
Chicago
Hu, G., T.-E. Lin, Y. Zhao, G. Lu, Y. Wu, and Y. Li. 2022. “UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 7837–51. https://doi.org/10.18653/v1/2022.emnlp-main.534.
Harvard
Hu, G. et al. (2022) “UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 7837–7851. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.534.
Vancouver
1. Hu G, Lin T-E, Zhao Y, Lu G, Wu Y, Li Y (2022) UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 7837–7851

BibTeX

@inproceedings{hu-etal-2022-unimse,
    title = "{U}ni{MSE}: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition",
    author = "Hu, Guimin  and
      Lin, Ting-En  and
      Zhao, Yi  and
      Lu, Guangming  and
      Wu, Yuchuan  and
      Li, Yongbin",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.534/",
    doi = "10.18653/v1/2022.emnlp-main.534",
    pages = "7837--7851"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/