MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis
Devamanyu HazarikaRoger ZimmermannSoujanya Poria
Proposes a multimodal representation framework, MISA, that factorizes signals into invariant and modality-specific subspaces, effectively resolving heterogeneous modality gaps to achieve state-of-the-art performance in sentiment analysis and humor detection.
Understanding human sentiment and humor from user-generated online videos requires analyzing multiple communication streams simultaneously, including spoken language, vocal tone, and visual facial expressions. However, combining these diverse data streams is difficult because each modality possesses fundamentally different characteristics and statistical distributions, creating significant modality gaps. While prior research has focused primarily on designing complex fusion architectures to bridge these gaps, such methods often place an excessive burden on the fusion mechanism to resolve discrepancies while failing to adequately separate shared information from unique, modality-specific nuances.
The main objective of the article is to demonstrate that learning factorized, disentangled modality representations before applying fusion provides a more effective and comprehensive foundation for multimodal sentiment analysis and humor detection. The authors evaluate a novel framework, named MISA, which explicitly separates each modality into modality-invariant and modality-specific representations to improve downstream predictive accuracy.
To achieve this, the authors designed a framework that maps initial utterance features into two distinct subspaces per modality: a shared modality-invariant space that aligns common information across modalities and a private modality-specific space that captures unique stylistic traits. The system is trained using a multi-component loss function that enforces distributional alignment using moment matching, imposes orthogonality constraints to eliminate redundancy, applies reconstruction loss to retain essential content, and optimizes task-specific prediction. The approach was evaluated on two widely used sentiment analysis benchmarks—the CMU-MOSI dataset of 2,198 video segments and the CMU-MOSEI dataset of 23,453 video segments—as well as the UR_FUNNY dataset for multimodal humor detection.
The article demonstrates five key findings. First, the framework outperformed existing state-of-the-art models across all sentiment benchmarks, achieving a mean absolute error reduction of 0.077 on MOSI and 0.010 on MOSEI, along with higher correlation and classification accuracy. Second, it surpassed state-of-the-art baselines in multimodal humor detection on UR_FUNNY, reaching an accuracy of 70.61% and improving performance by over two percentage points. Third, ablation analyses confirmed that combining both invariant and specific representations yields superior performance compared to using either representation alone or relying on unfactorized baselines. Fourth, language features proved to be the most critical contributor to overall accuracy, though integrating audio and visual signals consistently improved outcomes. Finally, the framework achieved these gains using only single-utterance information, outperforming complex models that incorporate surrounding contextual dialogue.
These findings indicate that prior investments in intricate fusion mechanisms may be suboptimal compared to improving initial representation learning. In practical applications, resolving cross-modal distribution gaps beforehand simplifies model complexity, reduces reliance on expensive contextual tracking architectures, and enhances the system's ability to detect nuanced affective states like sarcasm or humor. The results establish that proper representation factorization lowers overall engineering overhead while delivering robust, state-of-the-art accuracy.
Engineering and research teams developing affective computing systems should transition from designing complex fusion networks to adopting factorized subspace representation learning. For near-term implementation, practitioners should prioritize high-quality language encoders while using similarity and orthogonality constraints to integrate audio and visual streams. Future efforts should focus on piloting this representation approach across broader affective dimensions, such as fine-grained emotion recognition, and exploring alternative distance metrics for shared subspace alignment.
Confidence in these findings is high, supported by statistically significant improvements on multiple established benchmarks and thorough ablation testing. However, decision-makers should note certain limitations: the performance heavily relies on pre-trained language models, and the findings reflect dataset-specific feature distributions derived from structured monologues and presentation videos, which may require validation when applied to unaligned or highly noisy conversational data in production environments.
- Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). This paper introduces the CMU-MOSEI benchmark dataset and foundational cross-modal dynamic fusion concepts directly utilized and evaluated in MISA.
- Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). This work establishes key cross-modal attention baselines on the MOSI and MOSEI benchmarks that MISA directly compares against and builds upon.
- Paper: Tensor Fusion Network for Multimodal Sentiment Analysis, Amir Zadeh et al. (2017). This foundational paper develops the Tensor Fusion Network for multimodal sentiment analysis on MOSI, representing the primary fusion paradigm that MISA seeks to improve by learning disentangled representations.
- Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). This comprehensive survey outlines the core taxonomy of multimodal machine learning—including coordinated representation spaces and fusion—which underpins MISA's modality-invariant and specific design.
- Paper: Deep Canonical Correlation Analysis, Galen Andrew et al. (2013). This paper provides the theoretical basis for learning shared, correlated subspaces across disparate modalities using deep networks, which directly inspires MISA's invariant representation objective.
- Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). This work introduces contrastive multi-view learning principles that motivate isolating shared inter-modal signals while separating modality-specific components.
- Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). This seminal paper introduces deep cross-modal autoencoders for learning shared representations across audio and video channels, establishing the background for subspace factorization.
- Paper: Align before Fuse: Vision and Language Representation Learning with Momentum Distillation, Junnan Li et al. (2021). This work extends the paradigm of decoupling alignment and fusion by explicitly contrasting representations before multimodal fusion via transformers.
- Paper: ImageBind One Embedding Space to Bind Them All, Rohit Girdhar et al. (2023). This paper generalizes the concept of learning a unified modality-invariant embedding space across six distinct sensory modalities using contrastive objectives.
- Paper: BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation, Junnan Li et al. (2022). This work broadens the pre-fusion representation alignment strategy to unified vision-language understanding and generation architectures.
- Paper: Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, Bin Lin et al. (2023). This study adopts the principle of aligning multi-modal representations prior to feature projection and fusion within large language model backbones.
- Paper: A survey on multimodal large language models, Shukang Yin et al. (2023). This survey examines the broader trajectory of multimodal alignment and interface connectors in contemporary multimodal foundation models.
