MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis

Devamanyu HazarikaRoger ZimmermannSoujanya Poria

article2020ACM Multimedia1,325 citations

Proposes a multimodal representation framework, MISA, that factorizes signals into invariant and modality-specific subspaces, effectively resolving heterogeneous modality gaps to achieve state-of-the-art performance in sentiment analysis and humor detection.

Listen

Understanding human sentiment and humor from user-generated online videos requires analyzing multiple communication streams simultaneously, including spoken language, vocal tone, and visual facial expressions. However, combining these diverse data streams is difficult because each modality possesses fundamentally different characteristics and statistical distributions, creating significant modality gaps. While prior research has focused primarily on designing complex fusion architectures to bridge these gaps, such methods often place an excessive burden on the fusion mechanism to resolve discrepancies while failing to adequately separate shared information from unique, modality-specific nuances.

The main objective of the article is to demonstrate that learning factorized, disentangled modality representations before applying fusion provides a more effective and comprehensive foundation for multimodal sentiment analysis and humor detection. The authors evaluate a novel framework, named MISA, which explicitly separates each modality into modality-invariant and modality-specific representations to improve downstream predictive accuracy.

To achieve this, the authors designed a framework that maps initial utterance features into two distinct subspaces per modality: a shared modality-invariant space that aligns common information across modalities and a private modality-specific space that captures unique stylistic traits. The system is trained using a multi-component loss function that enforces distributional alignment using moment matching, imposes orthogonality constraints to eliminate redundancy, applies reconstruction loss to retain essential content, and optimizes task-specific prediction. The approach was evaluated on two widely used sentiment analysis benchmarks—the CMU-MOSI dataset of 2,198 video segments and the CMU-MOSEI dataset of 23,453 video segments—as well as the UR_FUNNY dataset for multimodal humor detection.

The article demonstrates five key findings. First, the framework outperformed existing state-of-the-art models across all sentiment benchmarks, achieving a mean absolute error reduction of 0.077 on MOSI and 0.010 on MOSEI, along with higher correlation and classification accuracy. Second, it surpassed state-of-the-art baselines in multimodal humor detection on UR_FUNNY, reaching an accuracy of 70.61% and improving performance by over two percentage points. Third, ablation analyses confirmed that combining both invariant and specific representations yields superior performance compared to using either representation alone or relying on unfactorized baselines. Fourth, language features proved to be the most critical contributor to overall accuracy, though integrating audio and visual signals consistently improved outcomes. Finally, the framework achieved these gains using only single-utterance information, outperforming complex models that incorporate surrounding contextual dialogue.

These findings indicate that prior investments in intricate fusion mechanisms may be suboptimal compared to improving initial representation learning. In practical applications, resolving cross-modal distribution gaps beforehand simplifies model complexity, reduces reliance on expensive contextual tracking architectures, and enhances the system's ability to detect nuanced affective states like sarcasm or humor. The results establish that proper representation factorization lowers overall engineering overhead while delivering robust, state-of-the-art accuracy.

Engineering and research teams developing affective computing systems should transition from designing complex fusion networks to adopting factorized subspace representation learning. For near-term implementation, practitioners should prioritize high-quality language encoders while using similarity and orthogonality constraints to integrate audio and visual streams. Future efforts should focus on piloting this representation approach across broader affective dimensions, such as fine-grained emotion recognition, and exploring alternative distance metrics for shared subspace alignment.

Confidence in these findings is high, supported by statistically significant improvements on multiple established benchmarks and thorough ablation testing. However, decision-makers should note certain limitations: the performance heavily relies on pre-trained language models, and the findings reflect dataset-specific feature distributions derived from structured monologues and presentation videos, which may require validation when applied to unaligned or highly noisy conversational data in production environments.

  • Paper: Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph, Amir Zadeh et al. (2018). This paper introduces the CMU-MOSEI benchmark dataset and foundational cross-modal dynamic fusion concepts directly utilized and evaluated in MISA.
  • Paper: Multimodal Transformer for Unaligned Multimodal Language Sequences, Yao-Hung Hubert Tsai et al. (2019). This work establishes key cross-modal attention baselines on the MOSI and MOSEI benchmarks that MISA directly compares against and builds upon.
  • Paper: Tensor Fusion Network for Multimodal Sentiment Analysis, Amir Zadeh et al. (2017). This foundational paper develops the Tensor Fusion Network for multimodal sentiment analysis on MOSI, representing the primary fusion paradigm that MISA seeks to improve by learning disentangled representations.
  • Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). This comprehensive survey outlines the core taxonomy of multimodal machine learning—including coordinated representation spaces and fusion—which underpins MISA's modality-invariant and specific design.
  • Paper: Deep Canonical Correlation Analysis, Galen Andrew et al. (2013). This paper provides the theoretical basis for learning shared, correlated subspaces across disparate modalities using deep networks, which directly inspires MISA's invariant representation objective.
  • Paper: Contrastive Multiview Coding, Yonglong Tian et al. (2019). This work introduces contrastive multi-view learning principles that motivate isolating shared inter-modal signals while separating modality-specific components.
  • Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). This seminal paper introduces deep cross-modal autoencoders for learning shared representations across audio and video channels, establishing the background for subspace factorization.
Cover for MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis

Abstract

Multimodal Sentiment Analysis is an active area of research that leverages multimodal signals for affective understanding of user-generated videos. The predominant approach, addressing this task, has been to develop sophisticated fusion techniques. However, the heterogeneous nature of the signals creates distributional modality gaps that pose significant challenges. In this paper, we aim to learn effective modality representations to aid the process of fusion. We propose a novel framework, MISA, which projects each modality to two distinct subspaces. The first subspace is modality-invariant, where the representations across modalities learn their commonalities and reduce the modality gap. The second subspace is modality-specific, which is private to each modality and captures their characteristic features. These representations provide a holistic view of the multimodal data, which is used for fusion that leads to task predictions. Our experiments on popular sentiment analysis benchmarks, MOSI and MOSEI, demonstrate significant gains over state-of-the-art models. We also consider the task of Multimodal Humor Detection and experiment on the recently proposed UR_FUNNY dataset. Here too, our model fares better than strong baselines, establishing MISA as a useful multimodal framework.

Table of Contents

  • 1 Introduction
  • 2 Related Works
  • 2.1 Multimodal Sentiment Analysis.
  • 2.2 Multimodal Representation Learning.
  • 3 Approach
  • 3.1 Task Setup
  • 3.2 MISA
  • 3.3 Modality Representation Learning
  • 3.4 Modality Fusion
  • 3.5 Learning
  • 3.5.1 ℒsim\mathcal{L}_{\text{sim}} – Similarity Loss
  • 3.5.2 ℒdiff\mathcal{L}_{\text{diff}} – Difference Loss
  • 3.5.3 ℒrecon\mathcal{L}_{\text{recon}} – Reconstruction Loss
  • 3.5.4 ℒtask\mathcal{L}_{\text{task}} – Task Loss
  • 4 Experiments
  • 4.1 Datasets
  • 4.1.1 CMU-MOSI
  • 4.1.2 CMU-MOSEI
  • 4.1.3 UR_\_FUNNY
  • 4.2 Evaluation Criteria
  • 4.3 Feature Extraction
  • 4.3.1 Language Features
  • 4.3.2 Visual Features
  • 4.3.3 Acoustic Features
  • 4.4 Baselines
  • 4.4.1 Previous Models.
  • 4.4.2 State of the Art.
  • 5 Results and Analysis
  • 5.1 Quantitative Results
  • 5.1.1 Multimodal Sentiment Analysis
  • 5.1.2 Multimodal Humor Detection
  • 5.1.3 BERT vs. GloVe.
  • 5.2 Ablation Study
  • 5.2.1 Role of Modalities.
  • 5.2.2 Role of Regularization.
  • 5.2.3 Role of subspaces.
  • 5.2.4 Visualizing Attention
  • 6 Conclusion
  • References
  • A Baseline Models
  • B Dataset Sizes
  • C Hyper-parameter Selection
  • D Network Topology

Knowls

  1. Knowl 1 — Modality-Invariant and -Specific Representation Learning in MISA

    model/method

    The MISA framework factorizes multimodal inputs into modality-invariant and modality-specific latent subspaces to aid downstream multimodal fusion. An input utterance UU comprises sequences of low-level features across three modalities: language (ll), visual (vv), and acoustic (aa), denoted as Ul∈RTl×dlU_l \in \mathbb{R}^{T_l \times d_l}, Uv∈RTv×dvU_v \in \mathbb{R}^{T_v \times d_v}, and Ua∈RTa×daU_a \in \mathbb{R}^{T_a \times d_a}, where TmT_m is sequence length and dmd_m is feature dimension for modality m∈{l,v,a}m \in \{l, v, a\}.

    First, each modality sequence is mapped to a fixed-dimensional utterance representation vector um∈Rdhu_m \in \mathbb{R}^{d_h}. For visual and acoustic modalities (and language when using GloVe), a stacked bidirectional Long Short-Term Memory (sLSTM) network with a dense layer produces: um=sLSTM(Um;θmlstm)u_m = \text{sLSTM}(U_m; \theta_m^{lstm}) When using pre-trained BERT for language, the utterance vector ul=BERT(Ul;θbert)u_l = \text{BERT}(U_l; \theta^{bert}) is computed as the mean of the final 768-dimensional token representations.

    Next, each utterance vector umu_m is projected into two separate latent subspaces:

    1. A modality-invariant representation hmc∈Rdhh_m^c \in \mathbb{R}^{d_h}, capturing shared cross-modal commonalities.
    2. A modality-specific representation hmp∈Rdhh_m^p \in \mathbb{R}^{d_h}, capturing idiosyncratic, modality-private characteristics.

    These representations are generated via feed-forward encoders: hmc=Ec(um;θc),hmp=Ep(um;θmp)h_m^c = E_c(u_m; \theta^c), \quad h_m^p = E_p(u_m; \theta_m^p) where EcE_c shares parameters θc\theta^c across all modalities m∈{l,v,a}m \in \{l, v, a\}, whereas EpE_p assigns modality-specific parameter sets θmp\theta_m^p.

  2. Knowl 2 — Transformer-Based Fusion and Prediction in MISA

    model/method

    After mapping the modalities (l,v,al, v, a) into six hidden vectors {hlc,hvc,hac,hlp,hvp,hap}⊂Rdh\{h_l^c, h_v^c, h_a^c, h_l^p, h_v^p, h_a^p\} \subset \mathbb{R}^{d_h}, MISA applies multi-head self-attention to facilitate cross-modal and cross-subspace interactions before prediction.

    The six vectors are stacked into a matrix M∈R6×dhM \in \mathbb{R}^{6 \times d_h}: M=[hlc,hvc,hac,hlp,hvp,hap]M = [h_l^c, h_v^c, h_a^c, h_l^p, h_v^p, h_a^p]

    A Transformer multi-head self-attention module with query, key, and value matrices Q=K=V=MQ = K = V = M computes transformed representations: Mˉ=MultiHead(M;θatt)=(head1⊕⋯⊕headn)Wo\bar{M} = \text{MultiHead}(M; \theta^{att}) = (\text{head}_1 \oplus \dots \oplus \text{head}_n) W^o where ⊕\oplus denotes concatenation, θatt={Wq,Wk,Wv,Wo}\theta^{att} = \{W^q, W^k, W^v, W^o\}, and the ii-th attention head is calculated as: headi=softmax(QWiq(KWik)⊤dh)VWiv\text{head}_i = \text{softmax}\left(\frac{Q W_i^q (K W_i^k)^\top}{\sqrt{d_h}}\right) V W_i^v with projection parameter matrices Wiq,Wik,Wiv∈Rdh×dhW_i^q, W_i^k, W_i^v \in \mathbb{R}^{d_h \times d_h} and output projection WoW^o.

    The transformed vectors [hˉlc,hˉvc,hˉac,hˉlp,hˉvp,hˉap][\bar{h}_l^c, \bar{h}_v^c, \bar{h}_a^c, \bar{h}_l^p, \bar{h}_v^p, \bar{h}_a^p] from Mˉ\bar{M} are concatenated into a joint multimodal vector: hout=[hˉlc⊕hˉvc⊕hˉac⊕hˉlp⊕hˉvp⊕hˉap]∈R6dhh^{out} = [\bar{h}_l^c \oplus \bar{h}_v^c \oplus \bar{h}_a^c \oplus \bar{h}_l^p \oplus \bar{h}_v^p \oplus \bar{h}_a^p] \in \mathbb{R}^{6 d_h}

    The final prediction y^\hat{y} is generated by a dense feed-forward network GG: y^=G(hout;θout)\hat{y} = G(h^{out}; \theta^{out})

  3. Knowl 3 — Central Moment Discrepancy Similarity Loss

    equation

    To minimize the distributional discrepancy between modality-invariant representations in the shared subspace, MISA optimizes the Central Moment Discrepancy (CMD) metric across pairs of modalities. For bounded random samples XX and YY with probability distributions on [a,b]N[a, b]^N, the empirical CMD metric of order KK is defined as: CMDK(X,Y)=1∣b−a∣∥E(X)−E(Y)∥2+∑k=2K1∣b−a∣k∥Ck(X)−Ck(Y)∥2\text{CMD}_K(X, Y) = \frac{1}{|b - a|} \| \mathbb{E}(X) - \mathbb{E}(Y) \|_2 + \sum_{k=2}^K \frac{1}{|b - a|^k} \| C_k(X) - C_k(Y) \|_2 where E(X)=1∣X∣∑x∈Xx\mathbb{E}(X) = \frac{1}{|X|} \sum_{x \in X} x is the empirical expectation vector and Ck(X)=E((x−E(X))k)C_k(X) = \mathbb{E}((x - \mathbb{E}(X))^k) is the vector of all kk-th order sample central moments of the coordinates of XX.

    The similarity loss Lsim\mathcal{L}_{sim} minimizes CMD across all modality pairs among the invariant representations hlc,hvc,hach_l^c, h_v^c, h_a^c: Lsim=13∑(m1,m2)∈{(l,a),(l,v),(a,v)}CMDK(hm1c,hm2c)\mathcal{L}_{sim} = \frac{1}{3} \sum_{(m_1, m_2) \in \{(l, a), (l, v), (a, v)\}} \text{CMD}_K(h_{m_1}^c, h_{m_2}^c) CMD enables explicit matching of higher-order probability moments without requiring distance matrix or kernel computations, and avoids the training instability and extra parameters associated with adversarial discriminators.

  4. Knowl 4 — Difference and Reconstruction Losses for Subspace Disentanglement

    equation

    To prevent redundancy between invariant and specific representations as well as across modality-specific representations, MISA enforces a soft orthogonality difference loss Ldiff\mathcal{L}_{diff}. Let HmcH_m^c and HmpH_m^p be batch matrices (normalized to zero mean and unit ℓ2\ell_2 row norm) of invariant and specific hidden vectors for modality m∈{l,v,a}m \in \{l, v, a\}. The difference loss is defined as: Ldiff=∑m∈{l,v,a}∥Hmc⊤Hmp∥F2+∑(m1,m2)∈{(l,a),(l,v),(a,v)}∥Hm1p⊤Hm2p∥F2\mathcal{L}_{diff} = \sum_{m \in \{l, v, a\}} \| {H_m^c}^\top H_m^p \|_F^2 + \sum_{(m_1, m_2) \in \{(l, a), (l, v), (a, v)\}} \| {H_{m_1}^p}^\top H_{m_2}^p \|_F^2 where ∥⋅∥F2\| \cdot \|_F^2 is the squared Frobenius norm.

    To ensure that the modality-specific encoders do not learn unrepresentative or trivial orthogonal vectors, a reconstruction loss Lrecon\mathcal{L}_{recon} is applied. The combined representations hmc+hmph_m^c + h_m^p are decoded back to the original utterance vectors um∈Rdhu_m \in \mathbb{R}^{d_h} using a decoder network D(⋅;θd)D(\cdot; \theta^d): u^m=D(hmc+hmp;θd)\hat{u}_m = D(h_m^c + h_m^p; \theta^d) Lrecon=13∑m∈{l,v,a}∥um−u^m∥22dh\mathcal{L}_{recon} = \frac{1}{3} \sum_{m \in \{l, v, a\}} \frac{\| u_m - \hat{u}_m \|_2^2}{d_h} where ∥⋅∥22\| \cdot \|_2^2 is the squared L2L_2-norm.

  5. Knowl 5 — Overall Optimization Objective for MISA

    equation

    The complete MISA architecture is trained end-to-end by minimizing the joint loss function: L=Ltask+αLsim+βLdiff+γLrecon\mathcal{L} = \mathcal{L}_{task} + \alpha \mathcal{L}_{sim} + \beta \mathcal{L}_{diff} + \gamma \mathcal{L}_{recon} where α,β,γ≥0\alpha, \beta, \gamma \ge 0 are hyperparameter interaction weights controlling the contributions of the similarity loss Lsim\mathcal{L}_{sim}, difference loss Ldiff\mathcal{L}_{diff}, and reconstruction loss Lrecon\mathcal{L}_{recon}.

    For a training batch of NbN_b utterances with ground-truth targets yiy_i and model predictions y^i\hat{y}_i, the task-specific loss Ltask\mathcal{L}_{task} is computed as: Ltask=−1Nb∑i=1Nbyi⋅log⁡y^ifor classification tasks\mathcal{L}_{task} = -\frac{1}{N_b} \sum_{i=1}^{N_b} y_i \cdot \log \hat{y}_i \quad \text{for classification tasks} Ltask=1Nb∑i=1Nb∥yi−y^i∥22for regression tasks\mathcal{L}_{task} = \frac{1}{N_b} \sum_{i=1}^{N_b} \| y_i - \hat{y}_i \|_2^2 \quad \text{for regression tasks}

  6. Knowl 6 — Multimodal Sentiment Analysis Performance on CMU-MOSI and CMU-MOSEI

    data/table

    MISA was evaluated on the CMU-MOSI and CMU-MOSEI multimodal sentiment analysis benchmarks across regression metrics (Mean Absolute Error, MAE ↓\downarrow; Pearson correlation, Corr ↑\uparrow) and classification metrics (Binary Accuracy, Acc-2 ↑\uparrow, formatted as negative/non-negative / negative/positive; F-Score ↑\uparrow; Seven-class Accuracy, Acc-7 ↑\uparrow). Models using BERT-based language features are designated with (B).

    Model (MOSI) MAE (↓\downarrow) Corr (↑\uparrow) Acc-2 (↑\uparrow) F-Score (↑\uparrow) Acc-7 (↑\uparrow)
    BC-LSTM 1.079 0.581 73.9 / - 73.9 / - 28.7
    MV-LSTM 1.019 0.601 73.9 / - 74.0 / - 33.2
    TFN 0.970 0.633 73.9 / - 73.4 / - 32.1
    MARN 0.968 0.625 77.1 / - 77.0 / - 34.7
    MFN 0.965 0.632 77.4 / - 77.3 / - 34.1
    LMF 0.912 0.668 76.4 / - 75.7 / - 32.8
    CH-Fusion - - 80.0 / - - -
    MFM 0.951 0.662 78.1 / - 78.1 / - 36.2
    RAVEN 0.915 0.691 78.0 / - 76.6 / - 33.2
    RMFN 0.922 0.681 78.4 / - 78.0 / - 38.3
    MCTN 0.909 0.676 79.3 / - 79.1 / - 35.6
    CIA 0.914 0.689 79.8 / - - / 79.5 38.9
    HFFN - - - / 80.2 - / 80.3 -
    LMFN - - - / 80.9 - / 80.9 -
    DFF-ATMF (B) - - - / 80.9 - / 81.2 -
    ARGF - - - / 81.4 - / 81.5 -
    MulT 0.871 0.698 - / 83.0 - / 82.8 40.0
    TFN (B) 0.901 0.698 - / 80.8 - / 80.7 34.9
    LMF (B) 0.917 0.695 - / 82.5 - / 82.4 33.2
    MFM (B) 0.877 0.706 - / 81.7 - / 81.6 35.4
    ICCN (B) 0.860 0.710 - / 83.0 - / 83.0 39.0
    MISA (B) 0.783 0.761 81.8 / 83.4 81.7 / 83.6 42.3
    Model (MOSEI) MAE (↓\downarrow) Corr (↑\uparrow) Acc-2 (↑\uparrow) F-Score (↑\uparrow) Acc-7 (↑\uparrow)
    MFN - - 76.0 / - 76.0 / - -
    MV-LSTM - - 76.4 / - 76.4 / - -
    Graph-MFN 0.710 0.540 76.9 / - 77.0 / - 45.0
    RAVEN 0.614 0.662 79.1 / - 79.5 / - 50.0
    MCTN 0.609 0.670 79.8 / - 80.6 / - 49.6
    CIA 0.680 0.590 80.4 / - 78.2 / - 50.1
    CIM-MTL - - 80.5 / - 78.8 / - -
    DFF-ATMF (B) - - - / 77.1 - / 78.3 -
    MulT 0.580 0.703 - / 82.5 - / 82.3 51.8
    TFN (B) 0.593 0.700 - / 82.5 - / 82.1 50.2
    LMF (B) 0.623 0.677 - / 82.0 - / 82.1 48.0
    MFM (B) 0.568 0.717 - / 84.4 - / 84.3 51.3
    ICCN (B) 0.565 0.713 - / 84.2 - / 84.2 51.6
    MISA (B) 0.555 0.756 83.6 / 85.5 83.8 / 85.3 52.2

    MISA outperforms previous state-of-the-art models on all evaluated metrics on both datasets. On MOSI, MISA improves MAE by 0.077, Corr by 0.051, binary accuracy by 2.0% (neg/non-neg) and 0.4% (neg/pos), F-Score by 2.6% / 0.6%, and Acc-7 by 3.3% over the best baseline. On MOSEI, MISA improves MAE by 0.010, Corr by 0.043, binary accuracy by 3.1% / 1.3%, F-Score by 5.0% / 1.1%, and Acc-7 by 0.6%. The binary classification gains are statistically significant (p<0.05p < 0.05 under McNemar's test).

  7. Knowl 7 — Multimodal Humor Detection Performance on UR_FUNNY

    data/table

    MISA was evaluated on the UR_FUNNY multimodal humor detection dataset for binary classification accuracy (Acc-2). The comparison evaluates context-dependent models (which utilize preceding utterances) versus target-only utterance models, under both GloVe and BERT text representations.

    Model Context Target Acc-2 (%) (↑\uparrow)
    C-MFN ✓ 58.45
    C-MFN ✓ 64.47
    TFN ✓ 64.71
    LMF ✓ 65.16
    C-MFN ✓ ✓ 65.23
    LMF (BERT) ✓ 67.53
    TFN (BERT) ✓ 68.57
    MISA (GloVe) ✓ 68.60
    MISA (BERT) ✓ 70.61

    MISA (BERT) achieves an accuracy of 70.61%, surpassing the previous contextual state-of-the-art model C-MFN (65.23%) by 2.07% absolute improvement (with p<0.05p < 0.05 under McNemar's test compared to TFN and LMF). Furthermore, MISA using GloVe features (68.60%) performs comparably to or better than BERT-based baselines TFN (BERT) (68.57%) and LMF (BERT) (67.53%), even without using inter-utterance contextual signals.

  8. Knowl 8 — Ablation Analysis of Modalities, Loss Terms, and Subspace Configurations

    data/table

    An ablation study evaluated the influence of modality inputs, individual loss components, and subspace architectures on CMU-MOSI, CMU-MOSEI, and UR_FUNNY.

    Model Configuration MOSI MOSEI UR_FUNNY
    MAE (↓\downarrow) Corr (↑\uparrow) MAE (↓\downarrow) Corr (↑\uparrow) Acc-2 (%) (↑\uparrow)
    (1) Full MISA 0.783 0.761 0.555 0.756 70.6
    Modality Ablations:
    (2) (-) Language ll 1.450 0.041 0.801 0.090 55.5
    (3) (-) Visual vv 0.798 0.756 0.558 0.753 69.7
    (4) (-) Audio aa 0.849 0.732 0.562 0.753 70.2
    Regularization Ablations:
    (5) (-) Lsim\mathcal{L}_{sim} (α=0\alpha=0) 0.807 0.740 0.566 0.751 69.3
    (6) (-) Ldiff\mathcal{L}_{diff} (β=0\beta=0) 0.824 0.749 0.565 0.742 69.3
    (7) (-) Lrecon\mathcal{L}_{recon} (γ=0\gamma=0) 0.794 0.757 0.559 0.754 69.7
    Subspace Variants:
    (8) MISA-base (no factorized subspaces) 0.810 0.750 0.568 0.752 69.2
    (9) MISA-inv (only invariant space) 0.811 0.737 0.561 0.743 68.8
    (10) MISA-sFusion (fuse specific hph^p only) 0.858 0.716 0.563 0.752 70.1
    (11) MISA-iFusion (fuse invariant hch^c only) 0.850 0.735 0.555 0.750 69.8

    Key takeaways:

    1. Multimodal fusion achieves the highest performance; removing language causes severe performance drops across all datasets (e.g., Pearson Corr drops from 0.761 to 0.041 on MOSI).
    2. Removing either similarity loss Lsim\mathcal{L}_{sim} or difference loss Ldiff\mathcal{L}_{diff} degrades performance more than removing reconstruction loss Lrecon\mathcal{L}_{recon}, indicating that distributional alignment and subspace orthogonality are critical.
    3. Training only an invariant subspace (MISA-inv) performs slightly worse than the un-factorized baseline (MISA-base). Optimal performance requires both factorized representation learning and subsequent fusion of both invariant and specific subspaces (Full MISA).
  9. Knowl 9 — Subspace Alignment and Self-Attention Dynamics in MISA

    empirical result

    Analysis of the learned representations and self-attention weights reveals two distinct empirical properties of MISA:

    1. Invariant Alignment and Specific Separation: In t-SNE projections of test representations, enforcing the regularization losses (α>0,β>0\alpha > 0, \beta > 0) causes the modality-invariant vectors (hlc,hvc,hach_l^c, h_v^c, h_a^c) to overlap and align in a shared distribution, while the modality-specific vectors (hlp,hvp,haph_l^p, h_v^p, h_a^p) form distinct, isolated clusters. When α=0,β=0\alpha = 0, \beta = 0, invariant vectors fail to align across modalities.

    2. Attention Weight Distribution in Transformer Fusion: In the Transformer self-attention layer, the invariant representations (hlc,hvc,hach_l^c, h_v^c, h_a^c) receive virtually equal average attention weights across all modalities, confirming that the modality gap is effectively reduced. Conversely, modality-specific representations exhibit asymmetric importance: language-specific features hlph_l^p contribute the highest attention weight across all queries, while acoustic and visual specific features contribute complementary, modality-distinct attention signals.

Coverage note — Omitted fine-grained implementation hyperparameters (e.g., specific batch sizes, learning rates, optimizer schedulers, and detailed dense layer dimensionalities described in the appendix) as standard reproducibility parameters rather than fundamental architectural concepts.

References

  1. 1.Roee Aharoni and Yoav Goldberg. 2020. Unsupervised Domain Clusters in Pretrained Language Models. CoRR abs/2004.02105 (2020). arXiv:2004.02105 https://arxiv.org/abs/2004.02105
  2. 2.Md. Shad Akhtar, Dushyant Singh Chauhan, Deepanway Ghosal, Soujanya Poria, Asif Ekbal, and Pushpak Bhattacharyya. 2019. Multi-task Learning for Multi-modal Emotion Recognition and Sentiment Analysis. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, MN, USA, 370–379. https://doi.org/10.18653/v1/n19-1034
  3. 3.Galen Andrew, Raman Arora, Jeff A. Bilmes, and Karen Livescu. 2013. Deep Canonical Correlation Analysis. In Proceedings of the 30th International Conference on Machine Learning, ICML 2013 (JMLR Workshop and Conference Proceedings, Vol. 28). JMLR.org, Atlanta, GA, USA, 1247–1255. http://proceedings.mlr.press/v28/andrew13.html
  4. 4.Tadas Baltrusaitis, Peter Robinson, and Louis-Philippe Morency. 2016. OpenFace: An open source facial behavior analysis toolkit. In 2016 IEEE Winter Conference on Applications of Computer Vision, WACV 2016. IEEE Computer Society, Lake Placid, NY, USA, 1–10. https://doi.org/10.1109/WACV.2016.7477553
  5. 5.Konstantinos Bousmalis, George Trigeorgis, Nathan Silberman, Dilip Krishnan, and Dumitru Erhan. 2016. Domain Separation Networks. In Advances in Neural Information Processing Systems 29. Curran Associates, Inc., Barcelona, Spain, 343–351. http://papers.nips.cc/paper/6254-domain-separation-networks
  6. 6.Dushyant Singh Chauhan, Md Shad Akhtar, Asif Ekbal, and Pushpak Bhattacharyya. 2019. Context-aware Interactive Attention for Multi-modal Sentiment and Emotion Analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, 5647–5657. https://doi.org/10.18653/v1/D19-1566
  7. 7.Feiyang Chen, Ziqian Luo, Yanyan Xu, and Dengfeng Ke. 2019. Complementary Fusion of Multi-Features and Multi-Modalities in Sentiment Analysis. Technical Report. EasyChair.
  8. 8.Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, and Louis-Philippe Morency. 2017. Multimodal Sentiment Analysis with Word-Level Fusion and Reinforcement Learning. In Proceedings of the 19th ACM International Conference on Multimodal Interaction (Glasgow, UK) (ICMI ’17). Association for Computing Machinery, New York, NY, USA, 163–171. https://doi.org/10.1145/3136755.3136801
  9. 9.Randall Davis. 2012. Multi-View Latent Variable Discriminative Models for Action Recognition. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (CVPR ’12). IEEE Computer Society, USA, 2120–2127.
  10. 10.Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer. 2014. COVAREP - A collaborative voice analysis repository for speech technologies. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2014. IEEE, Florence, Italy, 960–964. https://doi.org/10.1109/ICASSP.2014.6853739
  11. 11.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186. https://doi.org/10.18653/v1/N19-1423
  12. 12.Thomas Drugman and Abeer Alwan. 2011. Joint Robust Voicing Detection and Pitch Estimation Based on Residual Harmonics. In INTERSPEECH 2011, 12th Annual Conference of the International Speech Communication Association. ISCA, Florence, Italy, 1973–1976. http://www.isca-speech.org/archive/interspeech_2011/i11_1973.html
  13. 13.Thomas Drugman, Mark R. P. Thomas, Jón Guðnason, Patrick A. Naylor, and Thierry Dutoit. 2012. Detection of Glottal Closure Instants From Speech Signals: A Quantitative Review. IEEE Trans. Audio, Speech & Language Processing 20, 3 (2012), 994–1006. https://doi.org/10.1109/TASL.2011.2170835
  14. 14.Rosenberg Ekman. 1997. What the face reveals: Basic and applied studies of spontaneous expression using the Facial Action Coding System (FACS). Oxford University Press, USA.
  15. 15.Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016. Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Austin, Texas, 457–468. https://doi.org/10.18653/v1/D16-1044
  16. 16.Deepanway Ghosal, Md Shad Akhtar, Dushyant Chauhan, Soujanya Poria, Asif Ekbal, and Pushpak Bhattacharyya. 2018. Contextual Inter-modal Attention for Multi-modal Sentiment Analysis. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Brussels, Belgium, 3454–3466. https://doi.org/10.18653/v1/D18-1382
  17. 17.Yue Gu, Xinyu Li, Kaixiang Huang, Shiyu Fu, Kangning Yang, Shuhong Chen, Moliang Zhou, and Ivan Marsic. 2018. Human Conversation Analysis Using Attentive Multimodal Networks with Hierarchical Encoder-Decoder. In Proceedings of the 26th ACM International Conference on Multimedia (Seoul, Republic of Korea) (MM ’18). Association for Computing Machinery, New York, NY, USA, 537–545. https://doi.org/10.1145/3240508.3240714
  18. 18.Wenzhong Guo, Jianwen Wang, and Shiping Wang. 2019. Deep Multimodal Representation Learning: A Survey. IEEE Access 7 (2019), 63373–63394. https://doi.org/10.1109/ACCESS.2019.2916887
  19. 19.Md Kamrul Hasan, Wasifur Rahman, AmirAli Bagher Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, and Mohammed (Ehsan) Hoque. 2019. UR-FUNNY: A Multimodal Language Dataset for Understanding Humor. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Association for Computational Linguistics, Hong Kong, China, 2046–2056. https://doi.org/10.18653/v1/D19-1211
  20. 20.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long Short-Term Memory. Neural Computation 9, 8 (1997), 1735–1780. https://doi.org/10.1162/neco.1997.9.8.1735
  21. 21.Guosheng Hu, Yang Hua, Yang Yuan, Zhihong Zhang, Zheng Lu, Sankha S. Mukherjee, Timothy M. Hospedales, Neil Martin Robertson, and Yongxin Yang. 2017. Attribute-Enhanced Face Recognition with Neural Tensor Fusion Networks. In IEEE International Conference on Computer Vision, ICCV 2017. IEEE Computer Society, Venice, Italy, 3764–3773. https://doi.org/10.1109/ICCV.2017.404
  22. 22.Douwe Kiela, Suvrat Bhooshan, Hamed Firooz, and Davide Testuggine. 2019. Supervised Multimodal Bitransformers for Classifying Images and Text. In Visually Grounded Interaction and Language (ViGIL), NeurIPS 2019 Workshop. Curran Associates, Inc., Vancouver, Canada. https://vigilworkshop.github.io/static/papers/40.pdf
  23. 23.Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019. VisualBERT: A Simple and Performant Baseline for Vision and Language. CoRR abs/1908.03557 (2019). arXiv:1908.03557 http://arxiv.org/abs/1908.03557
  24. 24.Paul Pu Liang, Ziyin Liu, Amir Zadeh, and Louis-Philippe Morency. 2018. Multimodal Language Analysis with Recurrent Multistage Fusion. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Brussels, Belgium, 150–161. https://doi.org/10.18653/v1/d18-1014
  25. 25.Pengfei Liu, Xipeng Qiu, and Xuanjing Huang. 2017. Adversarial Multi-task Learning for Text Classification. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Volume 1: Long Papers. Association for Computational Linguistics, Vancouver, Canada, 1–10. https://doi.org/10.18653/v1/P17-1001
  26. 26.Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2018. Efficient Low-rank Multimodal Fusion With Modality-Specific Factors. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Volume 1: Long Papers. Association for Computational Linguistics, Melbourne, Australia, 2247–2256. https://doi.org/10.18653/v1/P18-1209
  27. 27.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. In Advances in Neural Information Processing Systems 32. Curran Associates, Inc., Vancouver, BC, Canada, 13–23. http://papers.nips.cc/paper/8297-vilbert-pretraining-task-agnostic-visiolinguistic-representations-for-vision-and-language-tasks
  28. 28.Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research 9, Nov (2008), 2579–2605.
  29. 29.Sijie Mai, Haifeng Hu, and Songlong Xing. 2019. Divide, Conquer and Combine: Hierarchical Feature Fusion Network with Local and Global Perspectives for Multimodal Affective Computing. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 481–492. https://doi.org/10.18653/v1/P19-1046
  30. 30.Sijie Mai, Haifeng Hu, and Songlong Xing. 2020. Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal Fusion. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020. AAAI Press, New York, NY, USA, 164–172. https://aaai.org/ojs/index.php/AAAI/article/view/5347
  31. 31.Sijie Mai, Songlong Xing, and Haifeng Hu. 2020. Locally Confined Modality Fusion Network With a Global Perspective for Multimodal Human Affective Computing. IEEE Trans. Multimedia 22, 1 (2020), 122–137. https://doi.org/10.1109/TMM.2019.2925966
  32. 32.Navonil Majumder, Devamanyu Hazarika, Alexander F. Gelbukh, Erik Cambria, and Soujanya Poria. 2018. Multimodal sentiment analysis using hierarchical fusion with context modeling. Knowl. Based Syst. 161 (2018), 124–133. https://doi.org/10.1016/j.knosys.2018.07.041
  33. 33.Rada Mihalcea. 2012. Multimodal Sentiment Analysis. In Proceedings of the 3rd Workshop in Computational Approaches to Subjectivity and Sentiment Analysis (Jeju, Republic of Korea) (WASSA ’12). Association for Computational Linguistics, USA, 1.
  34. 34.David Olson. 1977. From utterance to text: The bias of language in speech and writing. Harvard educational review 47, 3 (1977), 257–281.
  35. 35.Gwangbeen Park and Woobin Im. 2016. Image-Text Multi-Modal Representation Learning by Adversarial Backpropagation. CoRR abs/1612.08354 (2016). arXiv:1612.08354 http://arxiv.org/abs/1612.08354
  36. 36.Minlong Peng, Qi Zhang, Yu-Gang Jiang, and Xuanjing Huang. 2018. Cross-Domain Sentiment Classification with Target Domain Specific Information. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Volume 1: Long Papers. Association for Computational Linguistics, Melbourne, Australia, 2505–2513. https://doi.org/10.18653/v1/P18-1233
  37. 37.Yuxin Peng and Jinwei Qi. 2019. CM-GANs: Cross-Modal Generative Adversarial Networks for Common Representation Learning. ACM Trans. Multimedia Comput. Commun. Appl. 15, 1, Article 22 (Feb. 2019), 24 pages. https://doi.org/10.1145/3284750
  38. 38.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global Vectors for Word Representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Doha, Qatar, 1532–1543. https://doi.org/10.3115/v1/D14-1162
  39. 39.Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabás Póczos. 2019. Found in Translation: Learning Robust Joint Representations by Cyclic Translations between Modalities. In The Thirty-Third AAAI Conference on Artificial Intelligence. AAAI Press, Honolulu, Hawaii, 6892–6899. https://doi.org/10.1609/aaai.v33i01.33016892
  40. 40.Hai Pham, Thomas Manzini, Paul Pu Liang, and Barnabás Poczós. 2018. Seq2Seq2Sentiment: Multimodal Sequence to Sequence Models for Sentiment Analysis. In Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML). Association for Computational Linguistics, Melbourne, Australia, 53–63. https://doi.org/10.18653/v1/W18-3308
  41. 41.Soujanya Poria, Erik Cambria, Rajiv Bajpai, and Amir Hussain. 2017. A review of affective computing: From unimodal analysis to multimodal fusion. Information Fusion 37 (2017), 98 – 125. https://doi.org/10.1016/j.inffus.2017.02.003
  42. 42.Soujanya Poria, Erik Cambria, and Alexander Gelbukh. 2015. Deep Convolutional Neural Network Textual Features and Multiple Kernel Learning for Utterance-level Multimodal Sentiment Analysis. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Lisbon, Portugal, 2539–2544. https://doi.org/10.18653/v1/D15-1303
  43. 43.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Multi-level Multiple Attentions for Contextual Multimodal Sentiment Analysis. In 2017 IEEE International Conference on Data Mining, ICDM 2017. IEEE Computer Society, New Orleans, LA, USA, 1033–1038. https://doi.org/10.1109/ICDM.2017.134
  44. 44.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017. Context-Dependent Sentiment Analysis in User-Generated Videos. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Vancouver, Canada, 873–883. https://doi.org/10.18653/v1/P17-1081
  45. 45.Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, and Rada Mihalcea. 2020. Beneath the Tip of the Iceberg: Current Challenges and New Directions in Sentiment Analysis Research. CoRR abs/2005.00357 (2020). arXiv:2005.00357 https://arxiv.org/abs/2005.00357
  46. 46.Shyam Sundar Rajagopalan, Louis-Philippe Morency, Tadas Baltrusaitis, and Roland Goecke. 2016. Extending Long Short-Term Memory for Multi-View Structured Learning. In Computer Vision - ECCV 2016 - 14th European Conference, Proceedings, Part VII (Lecture Notes in Computer Science, Vol. 9911). Springer, Amsterdam, The Netherlands, 338–353. https://doi.org/10.1007/978-3-319-46478-7_21
  47. 47.Sebastian Ruder and Barbara Plank. 2018. Strong Baselines for Neural Semi-Supervised Learning under Domain Shift. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Volume 1: Long Papers. Association for Computational Linguistics, Melbourne, Australia, 1044–1054. https://doi.org/10.18653/v1/P18-1096
  48. 48.Mathieu Salzmann, Carl Henrik Ek, Raquel Urtasun, and Trevor Darrell. 2010. Factorized Orthogonal Latent Spaces. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2010 (JMLR Proceedings, Vol. 9). JMLR.org, Sardinia, Italy, 701–708. http://proceedings.mlr.press/v9/salzmann10a.html
  49. 49.Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019. VL-BERT: Pre-training of Generic Visual-Linguistic Representations. CoRR abs/1908.08530 (2019). arXiv:1908.08530 http://arxiv.org/abs/1908.08530
  50. 50.Zhongkai Sun, Prathusha K. Sarma, William A. Sethares, and Yingyu Liang. 2019. Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language Analysis. CoRR abs/1911.05544 (2019). arXiv:1911.05544 http://arxiv.org/abs/1911.05544
  51. 51.Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Learning Factorized Multimodal Representations. In 7th International Conference on Learning Representations, ICLR 2019. OpenReview.net, New Orleans, LA, USA. https://openreview.net/forum?id=rygqqsA9KX
  52. 52.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019. Multimodal Transformer for Unaligned Multimodal Language Sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 6558–6569. https://doi.org/10.18653/v1/P19-1656
  53. 53.Eric Tzeng, Judy Hoffman, Trevor Darrell, and Kate Saenko. 2015. Simultaneous Deep Transfer Across Domains and Tasks. In Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV) (ICCV ’15). IEEE Computer Society, USA, 4068–4076. https://doi.org/10.1109/ICCV.2015.463
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems 30. Curran Associates, Inc., Long Beach, CA, USA, 5998–6008. http://papers.nips.cc/paper/7181-attention-is-all-you-need
  55. 55.Weiran Wang, Honglak Lee, and Karen Livescu. 2016. Deep Variational Canonical Correlation Analysis. CoRR abs/1610.03454 (2016). arXiv:1610.03454 http://arxiv.org/abs/1610.03454
  56. 56.Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2019. Words Can Shift: Dynamically Adjusting Word Representations Using Nonverbal Behaviors. In The Thirty-Third AAAI Conference on Artificial Intelligence. AAAI Press, Honolulu, Hawaii, 7216–7223. https://doi.org/10.1609/aaai.v33i01.33017216
  57. 57.Chen Xi, Guanming Lu, and Jingjie Yan. 2020. Multimodal Sentiment Analysis Based on Multi-Head Attention Mechanism. In Proceedings of the 4th International Conference on Machine Learning and Soft Computing (Haiphong City, Viet Nam) (ICMLSC 2020). Association for Computing Machinery, New York, NY, USA, 34–39. https://doi.org/10.1145/3380688.3380693
  58. 58.Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor Fusion Network for Multimodal Sentiment Analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Copenhagen, Denmark, 1103–1114. https://doi.org/10.18653/v1/D17-1115
  59. 59.Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Memory Fusion Network for Multi-view Sequential Learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18). AAAI Press, New Orleans, Louisiana, USA, 5634–5641. https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/17341
  60. 60.Amir Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Volume 1: Long Papers. Association for Computational Linguistics, Melbourne, Australia, 2236–2246. https://doi.org/10.18653/v1/P18-1208
  61. 61.Amir Zadeh, Paul Pu Liang, Soujanya Poria, Prateek Vij, Erik Cambria, and Louis-Philippe Morency. 2018. Multi-attention Recurrent Network for Human Communication Comprehension. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI-18). AAAI Press, New Orleans, Louisiana, USA, 5642–5649. https://www.aaai.org/ocs/index.php/AAAI/AAAI18/paper/view/17390
  62. 62.Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016. Multimodal Sentiment Intensity Analysis in Videos: Facial Gestures and Verbal Messages. IEEE Intelligent Systems 31, 6 (2016), 82–88. https://doi.org/10.1109/MIS.2016.94
  63. 63.Werner Zellinger, Thomas Grubinger, Edwin Lughofer, Thomas Natschläger, and Susanne Saminger-Platz. 2017. Central Moment Discrepancy (CMD) for Domain-Invariant Representation Learning. In 5th International Conference on Learning Representations, ICLR 2017. OpenReview.net, Toulon, France. https://openreview.net/forum?id=SkB-_mcel

Citation

MLA
Hazarika, D., et al. “MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis”. arXiv, 2020, http://arxiv.org/abs/2005.03545v3.
APA
Hazarika, D., Zimmermann, R., & Poria, S. (2020). MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis. arXiv. http://arxiv.org/abs/2005.03545v3
Chicago
Hazarika, D., R. Zimmermann, and S. Poria. 2020. “MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis”. arXiv. http://arxiv.org/abs/2005.03545v3.
Harvard
Hazarika, D., Zimmermann, R. and Poria, S. (2020) “MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2005.03545v3.
Vancouver
1. Hazarika D, Zimmermann R, Poria S (2020) MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis. arXiv

BibTeX

@article{hazarika2020misa,
  title = {MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment Analysis},
  author = {Hazarika, Devamanyu and Zimmermann, Roger and Poria, Soujanya},
  year = {2020},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2005.03545v3},
  eprint = {2005.03545}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by-sa/4.0/