Multimodal Multi-loss Fusion Network for Sentiment Analysis
Zehui WuZiwei GongJaywon KooJulia Hirschberg
Proposes an end-to-end multimodal fusion framework that combines cross-modal attention, self-attention, and auxiliary single-modality training losses to achieve state-of-the-art sentiment analysis performance on CMU-MOSI, CMU-MOSEI, and CH-SIMS benchmarks.
Multimodal sentiment analysis aims to improve automated emotion understanding by combining multiple communication streams, such as spoken language, acoustic tone, and visual expressions. In real-world applications, organizations often face challenges in effectively fusing these diverse, high-dimensional signals without introducing substantial computational overhead or losing critical nuanced information. Developing robust and efficient architectures that reliably interpret human sentiment is crucial for downstream systems like conversational assistants, customer feedback monitors, and content moderation tools.
The article demonstrates an end-to-end sentiment detection framework called the Multi-Modality Multi-Loss Fusion Network (MMML). It evaluates the optimal selection of feature representations across audio and text streams, the architectural design of cross-modal fusion networks, the impact of multi-loss training, and the incorporation of conversational context.
The authors conducted extensive empirical experiments using three standard sentiment benchmarks spanning English and Mandarin: CMU-MOSI (2,199 video segments), CMU-MOSEI (23,453 segments), and CH-SIMS (2,281 segments). The pipeline used pre-trained language and speech encoders, combining them through a multi-stage fusion network that integrates cross-attention, self-attention, and pointwise feed-forward layers. The system was trained with auxiliary multi-task loss functions assigned to individual modality branches, and context was modeled by processing previous conversational utterances independently before fusion.
The evaluation yielded several key findings in order of importance. First, MMML achieved state-of-the-art results across all three benchmarks using only text and audio signals, outperforming previous top models that additionally processed video. Second, integrating conversational context yielded substantial performance gains; on CMU-MOSI, binary accuracy reached 89.69% with context compared to 88.16% without it. Third, multi-loss training provided significant advantages when distinct labels were available for each modality, boosting overall benchmark accuracy (achieving 82.93% on CH-SIMS compared to 78.34% under single-loss training) while simultaneously improving the standalone accuracy of the text sub-network by several percentage points. Fourth, utilizing fine-tuned pre-trained audio models (such as Data2Vec and HuBERT) significantly outperformed traditional hand-crafted acoustic features, achieving approximately 71% to 75% accuracy compared to roughly 45% to 68% for baseline acoustic features. Finally, attempts to restore original raw signals during fusion showed no measurable performance benefit.
These findings indicate that relying strictly on audio and text, while omitting video, reduces computational complexity and inference latency without sacrificing detection accuracy. The results show that text models trained in a multimodal setting retain improved capability even when deployed as standalone text tools. This provides operational flexibility when processing incomplete data streams or handling missing modalities in production.
Organizations developing sentiment detection systems should prioritize adopting modern pre-trained speech encoders alongside text models, incorporate multi-loss training objectives where modality-specific annotations exist, and process conversational context in independent streams. Teams can avoid added architectural complexity from signal-restoration layers or costly video processing pipelines. Prior to deploying these models into production environments, stakeholders should conduct pilot testing on target domain data, as the study's models were trained on public YouTube and entertainment datasets that may exhibit different emotional dynamics and potential acting biases compared to natural, real-world business interactions.
- Paper: Tensor Fusion Network for Multimodal Sentiment Analysis, Amir Zadeh et al. (2017). Read this foundational MOSI study first to understand the multimodal sentiment fusion problem and benchmark that MMML revisits.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). Its cross-modal attention approach for unaligned language, audio, and video sequences provides essential context for MMML’s attention-based fusion design.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). This paper introduced MOSEI and an interpretable fusion architecture, grounding MMML’s use of that benchmark and its multimodal comparisons.
- Paper: MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis, Devamanyu Hazarika et al. (2020). MISA’s shared-versus-modality-specific representations and multi-component objectives clarify earlier strategies for multimodal sentiment fusion and training.
- Paper: MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations, Soujanya Poria et al. (2018). Read this MELD dataset paper to understand the conversational context and multimodal dialogue setting that informs MMML’s context experiments.
No sufficiently relevant recommendations were found.
