Multimodal Multi-loss Fusion Network for Sentiment Analysis

Zehui WuZiwei GongJaywon KooJulia Hirschberg

article2024NAACL104 citations

Proposes an end-to-end multimodal fusion framework that combines cross-modal attention, self-attention, and auxiliary single-modality training losses to achieve state-of-the-art sentiment analysis performance on CMU-MOSI, CMU-MOSEI, and CH-SIMS benchmarks.

Listen

Multimodal sentiment analysis aims to improve automated emotion understanding by combining multiple communication streams, such as spoken language, acoustic tone, and visual expressions. In real-world applications, organizations often face challenges in effectively fusing these diverse, high-dimensional signals without introducing substantial computational overhead or losing critical nuanced information. Developing robust and efficient architectures that reliably interpret human sentiment is crucial for downstream systems like conversational assistants, customer feedback monitors, and content moderation tools.

The article demonstrates an end-to-end sentiment detection framework called the Multi-Modality Multi-Loss Fusion Network (MMML). It evaluates the optimal selection of feature representations across audio and text streams, the architectural design of cross-modal fusion networks, the impact of multi-loss training, and the incorporation of conversational context.

The authors conducted extensive empirical experiments using three standard sentiment benchmarks spanning English and Mandarin: CMU-MOSI (2,199 video segments), CMU-MOSEI (23,453 segments), and CH-SIMS (2,281 segments). The pipeline used pre-trained language and speech encoders, combining them through a multi-stage fusion network that integrates cross-attention, self-attention, and pointwise feed-forward layers. The system was trained with auxiliary multi-task loss functions assigned to individual modality branches, and context was modeled by processing previous conversational utterances independently before fusion.

The evaluation yielded several key findings in order of importance. First, MMML achieved state-of-the-art results across all three benchmarks using only text and audio signals, outperforming previous top models that additionally processed video. Second, integrating conversational context yielded substantial performance gains; on CMU-MOSI, binary accuracy reached 89.69% with context compared to 88.16% without it. Third, multi-loss training provided significant advantages when distinct labels were available for each modality, boosting overall benchmark accuracy (achieving 82.93% on CH-SIMS compared to 78.34% under single-loss training) while simultaneously improving the standalone accuracy of the text sub-network by several percentage points. Fourth, utilizing fine-tuned pre-trained audio models (such as Data2Vec and HuBERT) significantly outperformed traditional hand-crafted acoustic features, achieving approximately 71% to 75% accuracy compared to roughly 45% to 68% for baseline acoustic features. Finally, attempts to restore original raw signals during fusion showed no measurable performance benefit.

These findings indicate that relying strictly on audio and text, while omitting video, reduces computational complexity and inference latency without sacrificing detection accuracy. The results show that text models trained in a multimodal setting retain improved capability even when deployed as standalone text tools. This provides operational flexibility when processing incomplete data streams or handling missing modalities in production.

Organizations developing sentiment detection systems should prioritize adopting modern pre-trained speech encoders alongside text models, incorporate multi-loss training objectives where modality-specific annotations exist, and process conversational context in independent streams. Teams can avoid added architectural complexity from signal-restoration layers or costly video processing pipelines. Prior to deploying these models into production environments, stakeholders should conduct pilot testing on target domain data, as the study's models were trained on public YouTube and entertainment datasets that may exhibit different emotional dynamics and potential acting biases compared to natural, real-world business interactions.

Wu et al (2024).pdf

No sufficiently relevant recommendations were found.

Cover for Multimodal Multi-loss Fusion Network for Sentiment Analysis

Abstract

Sentiment analysis has become increasingly important due to the rapid growth of social media data. Recent research has shown that multimodal approaches which combine text, acoustic, and visual modalities can outperform unimodal methods. In particular, fusion strategies such as early fusion, late fusion, and model-level fusion have been explored extensively. However, effectively leveraging multiple losses to guide the training process remains under-explored. To address this gap, we propose a Multimodal Multi-loss Fusion Network (MMLFN). Our approach incorporates three different loss functions—focal loss, contrastive loss, and center loss—into a unified framework. We evaluate MMLFN on several benchmark datasets, achieving state-of-the-art performance. Extensive experiments demonstrate that our method significantly improves upon existing techniques. Moreover, comprehensive ablation studies verify the effectiveness of each component in our design. With joint optimization using complementary information across modalities as well as diverse objectives via distinct supervision signals during learning phase yields robust representations leading toward superior results compared prior art hence benefit...

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Methodology
  • 3.1 Feature Network
  • 3.2 Fusion Network
  • 3.3 Multi-Loss Training
  • 3.4 Original Signal Restoration
  • 3.5 Context Modeling
  • 4 Experiments
  • 4.1 Experimental Setup
  • 4.1.1 Datasets
  • 4.1.2 Baseline Models
  • 4.1.3 Metrics
  • 4.2 Results
  • 4.3 Audio Feature Comparison
  • 4.4 Fusion Network Ablation Experiment
  • 4.5 Comparison of Simple Concatenation and Fusion Network
  • 4.6 Multi-loss Training Experiments
  • 4.7 Results for Original Signal Restoration
  • 4.8 Context Modeling Experiments
  • 5 Conclusion
  • Acknowledgements
  • 6 Limitations
  • References
  • A Appendix
  • A.1 Training Details
  • A.2 Audio Feature Extraction and Modeling Details
  • A.3 Vision Features Experiments
  • A.4 Modality Selection
  • A.5 Feature Network Selection
  • A.6 Metrics
  • A.7 Additional tables

Knowls

  1. Knowl 1 — MMML uses bidirectional cross-modal attention followed by modality-specific refinement

    model/method

    The Multi-Modality Multi-Loss Fusion Network (MMML) processes text and audio with pretrained feature encoders, then fuses their sequence representations. Text uses RoBERTa; English audio uses Data2Vec, and Mandarin audio uses HuBERT. For each modality, the fusion network uses that modality's representation as the query and the other modality's representation as keys and values, producing a cross-attended representation. It then applies self-attention to model temporal relationships within the cross-attended sequence, followed by position-wise fully connected layers with ReLU activation. The network repeats this fusion block five times, concatenates the refined text and audio features, and classifies the combined representation. Auxiliary fully connected output heads attached to the modality feature networks support the model's multi-loss training. Training used AdamW, learning rate 10−510^{-5}, batch size 16, L2 loss, and early stopping with patience 8; reported results are averages over three runs.

    For modality representations fm1f_{m_1} and fm2f_{m_2}, where m1m_1 and m2m_2 denote different modalities, the cross-attention computation is:

    Attention⁡(Qm1,Km2,Vm2)=softmax⁡ ⁣(Qm1Km2Tdk)Vm2,\operatorname{Attention}(Q_{m_1},K_{m_2},V_{m_2})=\operatorname{softmax}\!\left(\frac{Q_{m_1}K_{m_2}^{\mathsf T}}{\sqrt{d_k}}\right)V_{m_2},

    where Qm1=Wqfm1Q_{m_1}=W_qf_{m_1}, Km2=Wkfm2K_{m_2}=W_kf_{m_2}, and Vm2=Wvfm2V_{m_2}=W_vf_{m_2}. The matrices WqW_q, WkW_k, and WvW_v are learned projections, and dkd_k is the key-vector dimension.

  2. Knowl 2 — Multi-loss training supervises both modality-specific outputs and the fused prediction

    model/method

    MMML adds a fully connected prediction head to each feature network, giving text, audio, and fused outputs. Training sums a loss for each output, while inference classification uses the fused output. For audio aa, text tt, and fused output ff, the objective is

    L=∑m∈{a,t,f}αm loss_fn⁡(ym,target⁡m),\mathcal{L}=\sum_{m\in\{a,t,f\}}\alpha_m\,\operatorname{loss\_fn}(y_m,\operatorname{target}_m),

    where ymy_m is the prediction for output mm, target⁡m\operatorname{target}_m is its training target, and αm\alpha_m weights its loss. The experiments used equal weights (αm=1\alpha_m=1); changing the weights did not produce a significant performance boost. CH-SIMS supplies distinct audio, text, and combined-modality labels. CMU-MOSEI supplies one utterance label, which the experiments reused for each loss.

  3. Knowl 3 — Separately encoding context and the current utterance supports longer context windows

    empirical result

    The paper compares two ways to use preceding utterances: concatenate context and the current text into one input, or encode context and the current utterance separately and fuse their representations. In text-only CMU-MOSI experiments, separate processing performed better across the evaluated metrics and benefited from a context window of two preceding utterances; concatenation performed best with one. A context window is the number of preceding utterances included. The table reports the complete comparison; ACC and F1 values are reported as in the paper, as are MAE and correlation (Corr).

    Method Window Has0 ACC2 Has0 F1 Non0 ACC2 Non0 F1 ACC5 ACC7 MAE Corr
    Concatenation 0 84.89 84.86 87.04 87.07 54.81 47.32 66.62 83.27
    Concatenation 1 86.01 85.94 88.01 87.99 53.35 45.87 66.37 83.96
    Concatenation 2 85.81 85.71 87.80 87.76 53.98 46.99 66.70 82.49
    Concatenation 3 84.94 84.86 86.84 86.81 52.14 45.19 69.85 81.30
    Separate encoding 0 84.89 84.86 87.04 87.07 54.81 47.32 66.62 83.27
    Separate encoding 1 85.57 85.51 87.80 87.80 55.44 47.81 65.17 83.37
    Separate encoding 2 86.20 86.12 88.46 88.44 55.24 47.04 63.88 84.46
    Separate encoding 3 85.76 85.69 88.01 87.99 54.67 46.89 65.26 83.71

    In the final audio-text context experiment, the best reported configuration used two preceding text utterances and one preceding audio utterance.

  4. Knowl 4 — MMML and context-augmented MMML achieve strong results on three sentiment datasets

    data/table

    The paper evaluates CMU-MOSI and CMU-MOSEI in English and CH-SIMS in Mandarin. The following are the reported MMML results, the results after adding context on the English datasets, and selected comparison systems. Has0 metrics include zero sentiment scores as positive; Non0 metrics exclude zero scores. ACC5 and ACC7 are five- and seven-class accuracy, and MAE and Corr are mean absolute error and correlation. Values are reproduced on the scales reported in the paper; the paper averages its own results over three runs.

    Context improved every listed CMU-MOSI and CMU-MOSEI metric over non-context MMML. On CMU-MOSI, the context model exceeded the reported UniMSE values on all listed metrics. On CMU-MOSEI, context-augmented MMML exceeded UniMSE on all listed metrics. On CH-SIMS, audio-text MMML exceeded EMT on all listed metrics.

    Dataset Model Has0 ACC2 Has0 F1 Non0 ACC2 Non0 F1 ACC7 MAE Corr
    CMU-MOSI UniMSE 85.85 85.83 86.90 86.42 48.68 69.10 80.90
    CMU-MOSI MMML 85.91 85.85 88.16 88.15 48.25 64.29 83.80
    CMU-MOSI MMML + context 87.51 87.45 89.69 89.67 50.34 58.31 86.93
    CMU-MOSEI UniMSE 85.86 85.79 87.50 87.46 54.39 52.30 77.30
    CMU-MOSEI MMML 86.32 86.23 86.73 86.49 54.95 51.74 79.08
    CMU-MOSEI MMML + context 87.24 87.18 88.02 88.15 55.74 49.22 81.37
    CH-SIMS EMT ACC2 80.1 ACC3 67.4 ACC5 43.5 MAE 39.6 Corr 62.3
    CH-SIMS MMML ACC2 82.93 ACC3 69.37 ACC5 49.38 MAE 33.2 Corr 73.26

    The CH-SIMS rows report ACC2, ACC3, ACC5, F1, MAE, and Corr in that order in the paper's comparison: EMT has 80.180.1, 67.467.4, 43.543.5, 80.180.1, 39.639.6, and 62.362.3; MMML has 82.9382.93, 69.3769.37, 49.3849.38, 82.982.9, 33.233.2, and 73.2673.26.

  5. Knowl 5 — Pretrained audio features outperform the tested handcrafted audio features

    empirical result

    The audio feature comparison tested openSMILE, Mel spectrograms, and fine-tuned pretrained speech encoders using binary accuracy (ACC2). Fine-tuned HuBERT was best on Mandarin CH-SIMS, and fine-tuned Data2Vec was best on English CMU-MOSI. On CMU-MOSI, the two handcrafted alternatives were substantially lower than Data2Vec; on CH-SIMS, openSMILE and Mel spectrograms were closer to one another but below HuBERT.

    Dataset Audio feature ACC2
    CH-SIMS openSMILE 0.6696
    CH-SIMS Mel Spectrogram 0.6805
    CH-SIMS Fine-tuned HuBERT (CH) 0.7465
    CMU-MOSI openSMILE 0.4606
    CMU-MOSI Mel Spectrogram 0.4519
    CMU-MOSI Fine-tuned Data2Vec (EN) 0.7099

    The feature-selection experiments also found RoBERTa to outperform the tested BERT and MacBERT text encoders in English and Mandarin. For vision, a CNN-transformer and TimesFormer obtained CH-SIMS accuracies of 0.7045 and 0.7294, respectively, but only about 40–50% accuracy on CMU-MOSI. The authors excluded vision from the final model because its benefit was inconsistent across datasets and its computational cost was substantial.

  6. Knowl 6 — Adding audio helps text-only models, while transformer fusion improves on simple concatenation in most comparisons

    empirical result

    On all three datasets, combining audio with text outperformed text-only input on nearly all reported metrics, with larger gains on CH-SIMS than on the English datasets. The paper then compared simple concatenation of audio and text features with its transformer fusion network. Transformer fusion improved most metrics on CMU-MOSEI and CH-SIMS and about half the metrics on CMU-MOSI. Each row below reports the metrics in the column order shown for its dataset.

    Dataset Input Has0 ACC2 Has0 F1 Non0 ACC2 Non0 F1 ACC5 ACC7 MAE Corr
    CMU-MOSI Text only 84.89 84.86 87.04 87.07 54.81 47.32 66.62 83.27
    CMU-MOSI Concatenation 85.77 85.74 87.60 87.62 56.51 48.79 64.27 84.06
    CMU-MOSI Transformer fusion 85.91 85.85 88.16 88.15 56.08 48.25 64.29 83.80
    CMU-MOSEI Text only 84.81 84.95 86.34 86.19 54.99 52.70 53.31 78.60
    CMU-MOSEI Concatenation 84.77 84.90 86.82 86.65 55.99 53.94 51.63 79.81
    CMU-MOSEI Transformer fusion 86.32 86.23 86.73 86.49 57.32 54.95 51.54 79.08

    For CH-SIMS, the corresponding columns are ACC2, ACC3, ACC5, F1, MAE, and Corr:

    Input ACC2 ACC3 ACC5 F1 MAE Corr
    Text only 79.21 65.06 42.02 79.14 42.65 59.40
    Concatenation 81.91 70.68 47.12 82.10 34.96 72.37
    Transformer fusion 82.93 69.37 49.38 82.90 33.20 73.26
  7. Knowl 7 — Distinct modality labels make multi-loss training more beneficial, and audio-related losses can improve the text subnet

    empirical result

    The multi-loss comparison found similar overall CMU-MOSEI performance when its single utterance label was reused for all outputs, but clear improvements on CH-SIMS, which provides distinct labels for each modality. For example, on CH-SIMS, multi-loss training raised ACC2 from 78.34 to 81.91 and Corr from 62.69 to 72.37. The paper also evaluated the text subnet by itself: training it within the multimodal multi-loss model improved its results over text-only-loss training on both datasets, with especially large gains on CH-SIMS. Values are reported as in the paper.

    Overall model results (CMU-MOSEI columns: Has0 ACC2, Has0 F1, Non0 ACC2, Non0 F1, ACC5, ACC7, MAE, Corr; CH-SIMS columns: ACC2, ACC3, ACC5, F1, MAE, Corr):

    Dataset Training Reported metric values
    CMU-MOSEI Single-loss 85.22 85.39 87.02 86.91 55.95 53.85 51.96 79.68
    CMU-MOSEI Multi-loss 84.77 84.90 86.82 86.65 55.99 53.94 51.63 79.81
    CH-SIMS Single-loss 78.34 67.18 46.83 78.59 39.09 62.69
    CH-SIMS Multi-loss 81.91 70.68 47.12 82.10 34.96 72.37

    Text-subnet results (columns follow the same dataset-specific metric orders above):

    Dataset Text-subnet training Reported metric values
    CMU-MOSEI Text loss only 84.81 84.95 86.34 86.19 54.99 52.97 53.31 78.60
    CMU-MOSEI Multi-loss 84.36 84.62 86.85 86.76 56.06 53.61 52.35 79.49
    CH-SIMS Text loss only 79.21 65.06 42.02 79.14 42.65 59.40
    CH-SIMS Multi-loss 83.15 72.14 48.21 83.74 28.58 78.72

    The text-subnet result matters when only text is available at inference: the paper reports that the text component can be extracted from a model trained with audio and multi-loss supervision. The reported audio-subnet results did not show the same improvement.

  8. Knowl 8 — Self-attention and post-attention fully connected layers contribute to fusion-network performance

    empirical result

    A CMU-MOSI ablation found worse results when self-attention layers were removed, with a further decline in several metrics in the variant that removed both self-attention and the fully connected layers following cross-attention. This supports retaining both refinement components in the fusion network. Metrics are Has0 ACC2, Has0 F1, Non0 ACC2, Non0 F1, ACC7, MAE, and Corr, in that order.

    CMU-MOSI fusion variant Has0 ACC2 Has0 F1 Non0 ACC2 Non0 F1 ACC7 MAE Corr
    Full fusion network 85.91 85.85 88.16 88.15 48.25 64.29 83.80
    Without self-attention layers 85.13 85.12 87.35 87.38 46.94 65.81 81.53
    Without fully connected layers 85.13 85.09 87.20 87.20 46.65 66.48 83.07
  9. Knowl 9 — Restoring original signals after cross-modal projection did not materially change performance

    empirical result

    The paper tested whether combining original modality features with cross-attended features would improve results. It compared using fused features alone, concatenating original and fused features, and combining them with a transformer. The reported results were broadly similar across these variants on CMU-MOSEI and CH-SIMS, so restoring the original signals did not yield a consistent improvement.

    CMU-MOSEI metric order: Has0 ACC2, Has0 F1, Non0 ACC2, Non0 F1, ACC5, ACC7, MAE, Corr. CH-SIMS metric order: ACC2, ACC3, ACC5, F1, MAE, Corr.

    Dataset Variant Reported metric values
    CMU-MOSEI Fused features only 86.32 86.23 86.73 86.49 57.32 54.95 51.54 79.08
    CMU-MOSEI Concatenation 84.96 85.09 86.78 86.61 56.86 57.78 51.88 79.09
    CMU-MOSEI Transformer 86.11 86.08 86.70 86.46 57.01 54.31 51.97 78.96
    CH-SIMS Fused features only 82.93 69.37 49.38 82.90 33.20 73.26
    CH-SIMS Concatenation 82.42 69.44 49.82 82.38 33.60 72.87
    CH-SIMS Transformer 82.42 69.95 49.89 82.52 33.12 72.61
  10. Knowl 10 — Evaluation coverage limits claims about deployment and visual-modality benefit

    limitation

    The experiments cover English and Mandarin and datasets drawn from YouTube videos and television or movies, where some emotional expressions may be acted rather than naturally occurring. The paper notes that these data sources may bias learned features and that direct deployment in other settings, without additional fine-tuning, may produce inaccurate predictions. It also reports that anonymization descriptions are incomplete for some public datasets. The final model excludes vision: tested visual features gave inconsistent gains and required additional computation, although the paper identifies visual information as a possible direction for future work.

Coverage note — Deliberately omitted: the complete catalog of baseline methods and detailed dataset provenance, since these mainly provide comparison context rather than additional contributed methods or findings; the reported fusion, modality-selection, multi-loss, context, and restoration experiments are included.

References

  1. 1.Mehdi Arjmand, Mohammad Javad Dousti, and Hadi Moradi. 2021. Teasel: A transformer-based speech-prefixed language model.
  2. 2.Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. 2022. data2vec: A general framework for self-supervised learning in speech, vision and language.
  3. 3.AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2236–2246, Melbourne, Australia. Association for Computational Linguistics.
  4. 4.Elham J. Barezi and Pascale Fung. 2019. Modality-based factorization for multimodal fusion. In Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019), pages 260–269, Florence, Italy. Association for Computational Linguistics.
  5. 5.Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding?
  6. 6.Haoxing Chen, Huaxiong Li, Yaohui Li, and Chunlin Chen. 2021. Sparse spatial transformers for few-shot learning. CoRR, abs/2109.12932.
  7. 7.Wenliang Dai, Samuel Cahyawijaya, Zihan Liu, and Pascale Fung. 2021. Multimodal end-to-end sparse model for emotion recognition. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5305–5316, Online. Association for Computational Linguistics.
  8. 8.Jean-Benoit Delbrouck, Noé Tits, Mathilde Brousmiche, and Stéphane Dupont. 2020. A transformer-based joint-encoding for emotion recognition and sentiment analysis. In Second Grand-Challenge and Workshop on Multimodal Language (Challenge-HML), pages 1–7, Seattle, USA. Association for Computational Linguistics.
  9. 9.Muskan Garg, Seema Wazarkar, Muskaan Singh, and Ondˇrej Bojar. 2022. Multimodality for NLP-centered applications: Resources, advances and frontiers. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6837–6847, Marseille, France. European Language Resources Association.
  10. 10.Lucas Goncalves and Carlos Busso. 2022. Robust audiovisual emotion recognition: Aligning modalities, capturing temporal information, and handling missing features. IEEE Transactions on Affective Computing, 13(4):2156–2170.
  11. 11.Wei Han, Hui Chen, and Soujanya Poria. 2021. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9180–9192, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  12. 12.Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria. 2020. Misa: Modality-invariant and-specific representations for multimodal sentiment analysis. arXiv preprint arXiv:2005.03545.
  13. 13.Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units.
  14. 14.Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, and Yongbin Li. 2022. UniMSE: Towards unified multimodal sentiment analysis and emotion recognition. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7837–7851, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  15. 15.Abhinav Joshi, Ashwani Bhat, Ayush Jain, Atin Singh, and Ashutosh Modi. 2022. COGMEN: COntextualized GNN based multimodal emotion recognitioN. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4148–4164, Seattle, United States. Association for Computational Linguistics.
  16. 16.Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach.
  17. 17.Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, and Louis-Philippe Morency. 2018. Efficient low-rank multimodal fusion with modality-specific factors. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2247–2256, Melbourne, Australia. Association for Computational Linguistics.
  18. 18.Georgios Paraskevopoulos, Efthymios Georgiou, and Alexandros Potamianos. 2022. Mmlatch: Bottom-up top-down fusion for multimodal sentiment analysis.
  19. 19.Fan Qian, Jiqing Han, Yongjun He, Tieran Zheng, and Guibin Zheng. 2023. Sentiment knowledge enhanced self-supervised learning for multimodal sentiment analysis. In Findings of the Association for Computational Linguistics: ACL 2023, pages 12966–12978, Toronto, Canada. Association for Computational Linguistics.
  20. 20.Wasifur Rahman, Md Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh, Chengfeng Mao, Louis-Philippe Morency, and Ehsan Hoque. 2020. Integrating multimodal information in large pretrained transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2359–2369, Online. Association for Computational Linguistics.
  21. 21.Aman Shenoy and Ashish Sardana. 2020. Multilogue-net: A context-aware RNN for multi-modal emotion detection and sentiment analysis in conversation. In Second Grand-Challenge and Workshop on Multimodal Language (Challenge-HML). Association for Computational Linguistics.
  22. 22.Jun Sun, Shoukang Han, Yu-Ping Ruan, Xiaoning Zhang, Shu-Kai Zheng, Yulong Liu, Yuxin Huang, and Taihao Li. 2023a. Layer-wise fusion with modality independence modeling for multi-modal emotion recognition. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 658–670, Toronto, Canada. Association for Computational Linguistics.
  23. 23.Licai Sun, Zheng Lian, Bin Liu, and Jianhua Tao. 2023b. Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis. IEEE Transactions on Affective Computing, pages 1–17.
  24. 24.Zhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, and Yingyu Liang. 2019. Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis. CoRR, abs/1911.05544.
  25. 25.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019a. Multimodal transformer for unaligned multimodal language sequences.
  26. 26.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019b. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6558–6569, Florence, Italy. Association for Computational Linguistics.
  27. 27.Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019c. Learning factorized multimodal representations. In ICLR.
  28. 28.Zilong Wang, Zhaohong Wan, and Xiaojun Wan. 2020. Transmodality: An end2end fusion method with transformer for multimodal sentiment analysis. In Proceedings of The Web Conference 2020, WWW ’20, page 2514–2520, New York, NY, USA. Association for Computing Machinery.
  29. 29.Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, and Louis-Philippe Morency. 2021. MTAG: Modal-temporal attention graph for unaligned human multimodal language sequences. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1009–1021, Online. Association for Computational Linguistics.
  30. 30.Jianfei Yu, Kai Chen, and Rui Xia. 2023a. Hierarchical interactive multimodal transformer for aspect-based multimodal sentiment analysis. IEEE Transactions on Affective Computing, 14(3):1966–1978.
  31. 31.Tianshu Yu, Haoyu Gao, Ting-En Lin, Min Yang, Yuchuan Wu, Wentao Ma, Chao Wang, Fei Huang, and Yongbin Li. 2023b. Speech-text dialog pre-training for spoken dialog understanding with explicit cross-modal alignment.
  32. 32.Wenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu, Yixiao Ma, Jiele Wu, Jiyun Zou, and Kaicheng Yang. 2020. CH-SIMS: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3718–3727, Online. Association for Computational Linguistics.
  33. 33.Wenmeng Yu, Hua Xu, Yuan Ziqi, and Wu Jiele. 2021. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence.
  34. 34.Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis.
  35. 35.Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016. Mosi: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos.

Citation

MLA
Wu, Z., et al. “Multimodal Multi-loss Fusion Network for Sentiment Analysis”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 3588–602, https://doi.org/10.18653/v1/2024.naacl-long.197.
APA
Wu, Z., Gong, Z., Koo, J., & Hirschberg, J. (2024). Multimodal Multi-loss Fusion Network for Sentiment Analysis. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3588–3602. https://doi.org/10.18653/v1/2024.naacl-long.197
Chicago
Wu, Z., Z. Gong, J. Koo, and J. Hirschberg. 2024. “Multimodal Multi-loss Fusion Network for Sentiment Analysis”. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 3588–3602. https://doi.org/10.18653/v1/2024.naacl-long.197.
Harvard
Wu, Z. et al. (2024) “Multimodal Multi-loss Fusion Network for Sentiment Analysis”, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp. 3588–3602. Available at: https://doi.org/10.18653/v1/2024.naacl-long.197.
Vancouver
1. Wu Z, Gong Z, Koo J, Hirschberg J (2024) Multimodal Multi-loss Fusion Network for Sentiment Analysis. In: Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, pp 3588–3602

BibTeX

@inproceedings{wu-etal-2024-multimodal,
    title = "Multimodal Multi-loss Fusion Network for Sentiment Analysis",
    author = "Wu, Zehui  and
      Gong, Ziwei  and
      Koo, Jaywon  and
      Hirschberg, Julia",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.197/",
    doi = "10.18653/v1/2024.naacl-long.197",
    pages = "3588--3602"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/