Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

Amir ZadehP. LiangSoujanya PoriaE. CambriaLouis-philippe Morency

article2018ACL1,791 citations

Introduces the CMU-MOSEI benchmark alongside the Dynamic Fusion Graph model to advance large-scale multimodal sentiment analysis and emotion recognition through interpretable cross-modal dynamics.

Listen

Human communication is intrinsically multimodal, combining spoken words, visual facial expressions, and vocal acoustics. However, computational models designed to analyze human sentiment and emotion have historically been constrained by small datasets that lack diversity across speakers, topics, and expressive behaviors. Furthermore, existing machine learning fusion techniques often function as black boxes, providing limited visibility into how different communication channels interact. Addressing these limitations is essential for developing reliable, interpretable artificial intelligence systems capable of processing real-world human interactions.

The article introduces CMU-MOSEI, the largest multimodal dataset of sentiment and emotion recognition to date, and demonstrates a novel, interpretable machine learning architecture called the Dynamic Fusion Graph to evaluate cross-modal dynamics.

To establish a robust foundation, the researchers curated 23,453 video segment annotations from 3,228 online monologue videos spanning 1,000 distinct speakers and 250 diverse topics, totaling nearly 66 hours of content. The dataset incorporates gender balancing, phoneme-level audio-to-text alignment, and rigorous multi-annotator crowdsourced labeling for sentiment intensity and six basic emotions. In parallel, the researchers engineered the Dynamic Fusion Graph and integrated it into a sequential architecture called the Graph Memory Fusion Network. This framework models unimodal, bimodal, and trimodal interactions across language, visual, and acoustic inputs while using dynamically calculated connection weights, termed efficacies, to trace how information is combined over time.

The evaluation produced several significant findings. First, the Graph Memory Fusion Network achieved superior performance in multimodal sentiment analysis, reaching 76.9% binary accuracy and an F1 score of 77.0%, while maintaining competitive performance across six emotion categories. Second, the dynamic graph structure confirmed that multimodal fusion is highly volatile, continuously altering its pathways depending on which modalities are informative at any given moment. Third, the system learned consistent communication priors: language and acoustic channels consistently fused together first, whereas unimodal connections directly to the final decision state were suppressed. Finally, visual signals were found to act conditionally, engaging heavily in the fusion process only when delivering meaningful and non-contradictory information.

These findings indicate that artificial intelligence systems perform best when designed to model the nuanced interdependencies of human expression rather than analyzing text, voice, or video in isolation. The interpretability provided by the Dynamic Fusion Graph mitigates operational risks by enabling stakeholders to understand the internal decision pathways of affective computing models. Because the model achieves state-of-the-art accuracy with efficient parameter usage, it provides a scalable, explainable framework for deploying automated sentiment and emotion recognition in customer experience, media analysis, and behavioral monitoring.

Organizations developing or deploying multimodal language systems should adopt diverse, large-scale benchmarks like CMU-MOSEI to prevent models from overfitting to specific speaker identities or narrow domains. Technical teams should implement dynamic, graph-based fusion mechanisms when explainability and cross-modal interpretability are required. Future initiatives should leverage the openly available dataset to expand multi-task learning investigations and explore dynamic fusion architectures across conversational and multi-speaker environments.

Confidence in these results is supported by the scale of the dataset, comprehensive feature extraction pipelines, and rigorous benchmarking against established baseline models. Nevertheless, users should consider certain boundary conditions: the dataset primarily features English-language monologues recorded directly in front of stationary cameras, and crowdsourced annotations exhibit a natural real-world skew toward positive sentiment and happiness over less frequent emotions such as fear.

  • Paper: Tensor Fusion Network for Multimodal Sentiment Analysis, Amir Zadeh et al. (2017). This paper establishes the foundational Tensor Fusion Network on the CMU-MOSI dataset, providing the prior multimodal fusion paradigm and dataset lineage that CMU-MOSEI and the Dynamic Fusion Graph directly expand upon.
  • Paper: Multimodal Machine Learning: A Survey and Taxonomy, Tadas Baltrušaitis et al. (2017). This survey provides the foundational taxonomy of representation, alignment, and fusion challenges that structure the multimodal language problem tackled by CMU-MOSEI.
  • Paper: AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild, Ali Mollahosseini et al. (2017). This work introduces in-the-wild continuous valence and arousal affect modeling, establishing key principles for unconstrained affective computing that underpin MOSEI's emotion annotations.
  • Paper: Multimodal Deep Learning, Jiquan Ngiam et al. (2011). This early foundational paper demonstrates cross-modal feature learning across audio and visual streams, setting early precedents for joint multimodal representation.
  • Paper: Multimodal learning with deep Boltzmann machines, Nitish Srivastava et al. (2012). This study introduces generative joint multimodal representation learning across disparate input distributions, providing conceptual roots for multimodal fusion modeling.
  • Paper: CROWDSOURCING A WORD–EMOTION ASSOCIATION LEXICON, Saif M. Mohammad et al. (2013). This paper pioneers crowdsourced annotation protocols for discrete emotions, informing the crowdsourcing methodology used to label CMU-MOSEI.
  • Paper: Automatic Analysis of Facial Expressions: The State of the Art, Maja Pantic et al. (2000). This benchmark review details the limitations of laboratory-constrained facial expression analysis, highlighting the exact shortcomings that CMU-MOSEI was curated to overcome.
Cover for Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph

Abstract

Analyzing human multimodal language is an emerging area of research in NLP. Intrinsically human communication is multimodal (heterogeneous), temporal and asynchronous; it consists of the language (words), visual (expressions), and acoustic (paralinguistic) modalities all in the form of asynchronous coordinated sequences. From a resource perspective, there is a genuine need for large scale datasets that allow for in-depth studies of multimodal language. In this paper we introduce CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI), the largest dataset of sentiment analysis and emotion recognition to date. Using data from CMU-MOSEI and a novel multimodal fusion technique called the Dynamic Fusion Graph (DFG), we conduct experimentation to investigate how modalities interact with each other in human multimodal language. Unlike previously proposed fusion techniques, DFG is highly interpretable and achieves competitive performance compared to the current state of the art.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Comparison to other Datasets
  • 2.1.1 Multimodal Datasets
  • 2.1.2 Language Datasets
  • 2.1.3 Visual and Acoustic Datasets
  • 2.2 Baseline Models
  • 3 CMU-MOSEI Dataset
  • 3.1 Data Acquisition
  • 3.2 Annotation
  • 3.3 Extracted Features
  • 4 Multimodal Fusion Study
  • 4.1 Dynamic Fusion Graph
  • 4.2 Graph-MFN
  • 5 Experiments and Discussion
  • 6 Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — CMU-MOSEI Dataset Specification and Annotation

    definition

    The CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) dataset is a large-scale corpus for multimodal sentiment analysis and emotion recognition in monologue video format. It comprises 23,45323{,}453 video segments (sentences) drawn from 3,2283{,}228 monologue videos across 1,0001{,}000 distinct speakers (channels) and 250250 topics, totaling 6565 hours, 5353 minutes, and 3636 seconds of video. Each speaker contributes between 1010 and 5050 sentence segments. The dataset maintains a balanced gender distribution (57%57\% male to 43%43\% female), with an average sentence duration of 7.287.28 seconds and a total vocabulary of 23,02623{,}026 unique words (447,143447{,}143 total word tokens).

    Annotations were gathered on Amazon Mechanical Turk from master workers (requiring a >98%>98\% approval rating), where each sentence was annotated by three independent judges:

    1. Sentiment Intensity: Evaluated on a 7-point Likert scale ranging from −3-3 to +3+3: {−3:highly negative,−2:negative,−1:weakly negative,0:neutral,+1:weakly positive,+2:positive,+3:highly positive}\{-3: \text{highly negative}, -2: \text{negative}, -1: \text{weakly negative}, 0: \text{neutral}, +1: \text{weakly positive}, +2: \text{positive}, +3: \text{highly positive}\}

    2. Ekman Emotion Intensity: Evaluated for six basic emotions {happiness,sadness,anger,fear,disgust,surprise}\{\text{happiness}, \text{sadness}, \text{anger}, \text{fear}, \text{disgust}, \text{surprise}\} on a 4-point Likert presence scale from 00 to 33: {0:no evidence,1:weakly present,2:present,3:highly present}\{0: \text{no evidence}, 1: \text{weakly present}, 2: \text{present}, 3: \text{highly present}\}

  2. Knowl 2 — Dynamic Fusion Graph (DFG) Architecture

    model/method

    The Dynamic Fusion Graph (DFG) is an interpretable neural fusion model that explicitly captures hierarchical nn-modal interactions among a set of modalities M={l,v,a}M = \{l, v, a\} (language, vision, acoustic) with dynamic structural routing during inference.

    DFG is formulated as a directed acyclic graph G=(V,E)G = (V, E) containing 8 vertices and 19 directed edges:

    • Vertices (VV): Represents latent modal dynamic representations. It includes 3 unimodal vertices ({l},{v},{a})(\{l\}, \{v\}, \{a\}), 3 bimodal vertices ({l,v},{v,a},{l,a})(\{l, v\}, \{v, a\}, \{l, a\}), 1 trimodal vertex ({l,v,a})(\{l, v, a\}), and 1 output terminal vertex T\mathcal{T}.
    • Edges (EE): A directed neural connection eij=(vi,vj)e_{ij} = (v_i, v_j) exists between two non-terminal vertices if and only if vi⊂vjv_i \subset v_j (e.g., {l}⊂{l,v}\{l\} \subset \{l, v\}, {l,v}⊂{l,v,a}\{l, v\} \subset \{l, v, a\}). In addition, directed edges connect every non-terminal vertex vk∈V∖{T}v_k \in V \setminus \{\mathcal{T}\} to the terminal output vertex T\mathcal{T}.

    Each edge eije_{ij} is weighted by a scalar dynamic efficacy αij∈[0,1]\alpha_{ij} \in [0, 1]. Vertex viv_i is multiplied element-wise by αij\alpha_{ij} before serving as input to the neural network Dj\mathcal{D}_j associated with vertex vjv_j.

  3. Knowl 3 — Mathematical Formulation of Dynamic Fusion Graph Computations and Efficacies

    equation

    Let l,v,a∈Rdl, v, a \in \mathbb{R}^d denote the input singleton latent representations for language, vision, and audio modalities at a given time-step. The Dynamic Fusion Graph (DFG) evaluates dynamic edge efficacies and higher-order modal vertices as follows:

    1. Dynamic Efficacy Vector Calculation: All 19 edge efficacies α={αij}eij∈E\boldsymbol{\alpha} = \{\alpha_{ij}\}_{e_{ij} \in E} are computed concurrently via an efficacy neural network Dα\mathcal{D}_\alpha using only singleton inputs: α=σ(Dα([l;v;a]))\boldsymbol{\alpha} = \sigma\left(\mathcal{D}_\alpha([l; v; a])\right) where [⋅;⋅][\cdot; \cdot] denotes vector concatenation and σ(⋅)\sigma(\cdot) is the element-wise sigmoid activation function ensuring αij∈[0,1]\alpha_{ij} \in [0, 1].

    2. Vertex Representation Propagation: For any intermediate composite vertex vjv_j (bimodal or trimodal), its latent state is computed by a deep neural network Dj\mathcal{D}_j taking all incoming efficacy-scaled parent vertices: vj=Dj([αijvi]vi⊂vj)v_j = \mathcal{D}_j\left( \left[ \alpha_{ij} v_i \right]_{v_i \subset v_j} \right)

    3. Terminal Output Assembly: The output representation Tt\mathcal{T}_t combines all unimodal, bimodal, and trimodal dynamic representations scaled by their respective direct connections to T\mathcal{T}: Tt=∑vk∈V∖{T}αkTvk\mathcal{T}_t = \sum_{v_k \in V \setminus \{\mathcal{T}\}} \alpha_{k\mathcal{T}} v_k

  4. Knowl 4 — Graph Memory Fusion Network (Graph-MFN)

    model/method

    The Graph Memory Fusion Network (Graph-MFN) integrates the Dynamic Fusion Graph (DFG) into a recurrent sequential architecture for multimodal time-series modeling:

    1. System of Unimodal LSTMs: Three parallel Long Short-Term Memory networks model intra-modal temporal dynamics for language (ll), vision (vv), and acoustic (aa) streams. For each modality m∈{l,v,a}m \in \{l, v, a\}, hidden states htmh^m_t and memory cells ctmc^m_t are maintained.

    2. Singleton Vertex Generation: At time-step tt, a dedicated feedforward network Dm\mathcal{D}_m processes consecutive hidden states to capture local temporal variation: vtm=Dm([ht−1m;htm])v^m_t = \mathcal{D}_m\left( [h^m_{t-1}; h^m_t] \right) yielding singleton vertex inputs lt=vtll_t = v^l_t, vt=vtvv_t = v^v_t, and at=vtaa_t = v^a_t for the DFG.

    3. Cross-Modal Fusion and Memory Update: DFG processes (lt,vt,at)(l_t, v_t, a_t) and outputs multimodal fusion representation Tt\mathcal{T}_t. A Multi-view Gated Memory state utu_t stores cross-modal representations through time: u^t=Du(Tt)\hat{u}_t = \mathcal{D}_u(\mathcal{T}_t) γ1=σ(Dγ1(Tt)),γ2=σ(Dγ2(Tt))\gamma_1 = \sigma\left(\mathcal{D}_{\gamma_1}(\mathcal{T}_t)\right), \quad \gamma_2 = \sigma\left(\mathcal{D}_{\gamma_2}(\mathcal{T}_t)\right) ut=γ1⊙ut−1+γ2⊙u^tu_t = \gamma_1 \odot u_{t-1} + \gamma_2 \odot \hat{u}_t where γ1\gamma_1 and γ2\gamma_2 denote the retain and update gates, respectively.

    4. Feedback and Task Output: A feedback representation zt=Dz(Tt)z_t = \mathcal{D}_z(\mathcal{T}_t) updates the unimodal LSTMs. At the final sequence step TT, prediction heads operate on the concatenated state [hTl;hTv;hTa;uT][h^l_T; h^v_T; h^a_T; u_T].

  5. Knowl 5 — Multimodal Feature Extraction and Alignment Pipeline for CMU-MOSEI

    experimental setup

    Features across language, visual, and acoustic modalities in CMU-MOSEI are extracted and aligned at the word level:

    • Language: Manual transcripts are converted into word vectors using GloVe pre-trained word embeddings. Words and speech audio are aligned at the phoneme level using the P2FA forced alignment model.
    • Visual: Video frames are extracted at 30 Hz30\text{ Hz}. Faces are detected using Multi-task Cascaded Convolutional Networks (MTCNN). Facial Action Units (FAUs) are extracted using the Facial Action Coding System (FACS) via OpenFace. OpenFace also extracts 6868 facial landmarks, 2020 facial shape parameters, HoG features, head pose, head orientation, and eye gaze. Six basic static facial emotions are extracted using Emotient FACET. Facial identity embeddings are extracted using DeepFace, FaceNet, and SphereFace. Visual features are aligned to word intervals by interpolation.
    • Acoustic: Extracted using the COVAREP framework at speech intervals and interpolated to word intervals. Extracted acoustic features include 1212 Mel-frequency cepstral coefficients (MFCCs), fundamental frequency (pitch), voiced/unvoiced segmenting features, glottal source parameters, peak slope parameters, and maxima dispersion quotients.
  6. Knowl 6 — CMU-MOSEI Benchmark Evaluation Results

    data/table

    Performance of Graph-MFN compared against the previous best (SOTA1) and second-best (SOTA2) unimodal and multimodal models on CMU-MOSEI. Metrics include binary classification accuracy (A2A^2), 5-class accuracy (A5A^5), 7-class accuracy (A7A^7), binary F1 score, Mean Absolute Error (MAE), Pearson correlation (rr), and Weighted Accuracy (WA) alongside F1 score for six Ekman emotions:

    Dataset MOSEI Sentiment MOSEI Emotions
    Task Sentiment Anger Disgust Fear Happy Sad Surprise
    Metric A2A^2 F1 A5A^5 A7A^7 MAE rr WA F1 WA F1 WA F1 WA F1 WA F1 WA F1
    Language SOTA2 74.1 74.1 43.1 42.9 0.75 0.46 56.0 71.0 59.0 67.1 56.2 79.7 53.0 44.1 53.8 49.9 53.2 70.0
    Language SOTA1 74.3 74.1 43.2 43.2 0.74 0.47 56.6 71.8 64.0 72.6 58.8 89.8 54.0 47.0 54.0 61.2 54.3 85.3
    Visual SOTA2 73.8 73.5 42.5 42.5 0.78 0.41 54.4 64.6 54.4 71.5 51.3 78.4 53.4 40.8 54.3 60.8 51.3 84.2
    Visual SOTA1 73.9 73.7 42.7 42.7 0.78 0.43 60.0 71.0 60.3 72.4 64.2 89.8 57.4 49.3 57.7 61.5 51.8 85.4
    Acoustic SOTA2 74.2 73.8 42.1 42.1 0.78 0.43 55.5 51.8 58.9 72.4 58.5 89.8 57.2 55.5 58.9 65.9 52.2 83.6
    Acoustic SOTA1 74.2 73.9 42.4 42.4 0.74 0.43 56.4 71.9 60.9 72.4 62.7 89.8 61.5 61.4 62.0 69.2 54.3 85.4
    Multimodal SOTA2 76.0 76.0 44.7 44.6 0.72 0.52 56.0 71.4 65.2 71.4 56.7 89.9 57.8 66.6 58.9 60.8 52.2 85.4
    Multimodal SOTA1 76.4 76.4 44.8 44.7 0.72 0.52 60.5 72.0 67.0 73.2 60.0 89.9 66.5 71.0 59.2 61.8 53.3 85.4
    Graph-MFN 76.9 77.0 45.1 45.0 0.71 0.54 62.6 72.8 69.1 76.6 62.0 89.9 66.3 66.3 60.4 66.9 53.7 85.5

    Graph-MFN outperforms all previous state-of-the-art multimodal baselines across all sentiment metrics (A2=76.9%A^2 = 76.9\%, MAE=0.71\text{MAE} = 0.71, r=0.54r = 0.54) and achieves top or competitive results across emotion categories.

  7. Knowl 7 — Empirical Dynamics and Priors Discovered by Dynamic Fusion Graph

    empirical result

    Analysis of the learned edge efficacies αij\alpha_{ij} across time and test cases reveals key properties of human multimodal communication:

    1. Dynamic Volatility: Graph connection efficacies change dynamically over time steps and input samples, selectively routing information through relevant nn-modal paths when modalities are informative and attenuating paths when inputs are noisy or uninformative.
    2. Language-Audio Coupling Prior: The model learns an invariant prior that strongly favors fusing language and audio first: efficacies for l→{l,a}l \to \{l, a\} and a→{l,a}a \to \{l, a\} are consistently high across almost all samples and time steps. Concurrently, direct unilateral connections from language or audio to the terminal vertex (l→Tl \to \mathcal{T} and a→Ta \to \mathcal{T}) maintain consistently low efficacy, indicating that the network avoids relying on uncombined speech modalities alone.
    3. Conditional Visual Modulation: Visual modality interactions (v→{l,v}v \to \{l, v\}, v→{l,a,v}v \to \{l, a, v\}, and {l,a}→{l,a,v}\{l, a\} \to \{l, a, v\}) remain active when facial gestures convey clear emotional signals (e.g., gaze aversion, surprise), but are deactivated (efficacy approaches zero) when visual expressions are neutral, uninformative, or contradictory.

Coverage note — None omitted; all core contributions—the CMU-MOSEI dataset design and feature extraction, the Dynamic Fusion Graph architecture and mathematical formulation, the Graph-MFN recurrent model, full benchmark empirical results, and the interpretability analysis of multimodal fusion dynamics—are included.

References

  1. 1.Paavo Alku, Tom Backstr¨om, and Erkki Vilkman. 2002. Normalized amplitude quotient for parametrization of the glottal flow. the Journal of the Acoustical Society of America 112(2):701–710.
  2. 2.Paavo Alku, Helmer Strik, and Erkki Vilkman. 1997. Parabolic spectral parameter—a new method for quantification of the glottal flow. Speech Communication 22(1):67–79.
  3. 3.Tadas Baltrušaitis, Chaitanya Ahuja, and LouisPhilippe Morency. 2017. Multimodal machine learning: A survey and taxonomy. arXiv preprint arXiv:1705.09406 .
  4. 4.Tadas Baltrušaitis, Peter Robinson, and Louis-Philippe Morency. 2016. Openface: an open source facial behavior analysis toolkit. In Applications of Computer Vision (WACV), 2016 IEEE Winter Conference on. IEEE, pages 1–10.
  5. 5.Sanjay Bilakhia, Stavros Petridis, Anton Nijholt, and Maja Pantic. 2015. The mahnob mimicry database: A database of naturalistic human interactions. Pattern Recognition Letters 66(Supplement C):52 – 61. Pattern Recognition in Human Computer Interaction. https://doi.org/https://doi.org/10.1016/j.patrec.2015.03.005.
  6. 6.Leo Breiman. 2001. Random forests. Mach. Learn. 45(1):5–32. https://doi.org/10.1023/A:1010933404324.
  7. 7.Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette Chang, Sungbok Lee, and Shrikanth S. Narayanan. 2008. Iemocap: Interactive emotional dyadic motion capture database. Journal of Language Resources and Evaluation 42(4):335–359. https://doi.org/10.1007/s10579-008-9076-6.
  8. 8.Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, and Louis-Philippe Morency. 2017. Multimodal sentiment analysis with wordlevel fusion and reinforcement learning. In Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, New York, NY, USA, ICMI 2017, pages 163–171. https://doi.org/10.1145/3136755.3136801.
  9. 9.Glen Coppersmith and Erin Kelly. 2014. Dynamic wordclouds and vennclouds for exploratory data analysis. In Proceedings of the Workshop on Interactive Language Learning, Visualization, and Interfaces. Association for Computational Linguistics, Baltimore, Maryland, USA, pages 22–29.
  10. 10.Corinna Cortes and Vladimir Vapnik. 1995. Supportvector networks. Mach. Learn. 20(3):273–297. https://doi.org/10.1023/A:1022627411411.
  11. 11.Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer. 2014. Covarep—a collaborative voice analysis repository for speech technologies. In Acoustics, Speech and Signal Processing (ICASSP), 2014 IEEE International Conference on. IEEE, pages 960–964.
  12. 12.A. Dhall, R. Goecke, S. Lucey, and T. Gedeon. 2012. Collecting large, richly annotated facial-expression databases from movies. IEEE MultiMedia 19(3):34–41. https://doi.org/10.1109/MMUL.2012.26.
  13. 13.Abhinav Dhall, O.V. Ramana Murthy, Roland Goecke, Jyoti Joshi, and Tom Gedeon. 2015. Video and image based emotion recognition challenges in the wild: Emotiw 2015. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, New York, NY, USA, ICMI '15, pages 423–426. https://doi.org/10.1145/2818346.2829994.
  14. 14.Thomas Drugman and Abeer Alwan. 2011. Joint robust voicing detection and pitch estimation based on residual harmonics. In Interspeech. pages 1973–1976.
  15. 15.Thomas Drugman, Mark Thomas, Jon Gudnason, Patrick Naylor, and Thierry Dutoit. 2012. Detection of glottal closure instants from speech signals: A quantitative review. IEEE Transactions on Audio, Speech, and Language Processing 20(3):994–1006.
  16. 16.Paul Ekman, Wallace V Freisen, and Sonia Ancoli. 1980. Facial signs of emotional experience. Journal of personality and social psychology 39(6):1125.
  17. 17.A. Graves, A. r. Mohamed, and G. Hinton. 2013. Speech recognition with deep recurrent neural networks. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing. pages 6645–6649. https://doi.org/10.1109/ICASSP.2013.6638947.
  18. 18.Michael Grimm, Kristian Kroschel, and Shrikanth Narayanan. 2008. The vera am mittag german audiovisual emotional speech database. In ICME. IEEE, pages 865–868.
  19. 19.Devamanyu Hazarika, Soujanya Poria, Amir Zadeh, Erik Cambria, Louis-Philippe Morency, and Roger Zimmerman. 2018. Memn: Multimodal emotional memory network for emotion recognition in dyadic conversational videos. In NAACL.
  20. 20.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation 9(8):1735–1780.
  21. 21.iMotions. 2017. Facial expression analysis. goo.gl/1rh1JN.
  22. 22.Mohit Iyyer, Varun Manjunatha, Jordan L BoydGraber, and Hal Daumé III. 2015. Deep unordered composition rivals syntactic methods for text classification. In ACL (1). pages 1681–1691.
  23. 23.Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2013. 3d convolutional neural networks for human action recognition. IEEE Trans. Pattern Anal. Mach. Intell. 35(1):221–231. https://doi.org/10.1109/TPAMI.2012.59.
  24. 24.Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. arXiv preprint arXiv:1404.2188 .
  25. 25.John Kane and Christer Gobl. 2013. Wavelet maxima dispersion for breathy to tense voice discrimination. IEEE Transactions on Audio, Speech, and Language Processing 21(6):1170–1179.
  26. 26.Wootaek Lim, Daeyoung Jang, and Taejin Lee. 2016. Speech emotion recognition using convolutional and recurrent neural networks. In Signal and Information Processing Association Annual Summit and Conference (APSIPA), 2016 Asia-Pacific. IEEE, pages 1–4.
  27. 27.Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song. 2017. Sphereface: Deep hypersphere embedding for face recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition.
  28. 28.Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, Portland, Oregon, USA, pages 142–150. http://www.aclweb.org/anthology/P11-1015.
  29. 29.Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In Association for Computational Linguistics (ACL) System Demonstrations. pages 55–60. http://www.aclweb.org/anthology/P/P14/P14-5010.
  30. 30.Louis-Philippe Morency, Rada Mihalcea, and Payal Doshi. 2011. Towards multimodal sentiment analysis: Harvesting opinions from the web. In Proceedings of the 13th International Conference on Multimodal Interactions. ACM, pages 169–176.
  31. 31.Friedrich Max Müller. 1866. Lectures on the science of language: Delivered at the Royal Institution of Great Britain in April, May, & June 1861, volume 1. Longmans, Green.
  32. 32.Behnaz Nojavanasghari, Deepak Gopinath, Jayanth Koushik, Tadas Baltrušaitis, and Louis-Philippe Morency. 2016. Deep multimodal fusion for persuasiveness prediction. In Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, New York, NY, USA, ICMI 2016, pages 284–288. https://doi.org/10.1145/2993148.2993176.
  33. 33.Bo Pang, Lillian Lee, and Shivakumar Vaithyanathan. 2002. Thumbs up? sentiment classification using machine learning techniques. In Proceedings of EMNLP. pages 79–86.
  34. 34.Sunghyun Park, Han Suk Shim, Moitreya Chatterjee, Kenji Sagae, and Louis-Philippe Morency. 2014. Computational analysis of persuasiveness in social multimedia: A novel dataset and multimodal prediction approach. In Proceedings of the 16th International Conference on Multimodal Interaction. ACM, New York, NY, USA, ICMI '14, pages 50–57. https://doi.org/10.1145/2663204.2663260.
  35. 35.Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In EMNLP. volume 14, pages 1532–1543.
  36. 36.Veronica Perez-Rosas, Rada Mihalcea, and LouisPhilippe Morency. 2013. Utterance-Level Multimodal Sentiment Analysis. In Association for Computational Linguistics (ACL). Sofia, Bulgaria.
  37. 37.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Mazumder, Amir Zadeh, and LouisPhilippe Morency. 2017a. Context dependent sentiment analysis in user generated videos. In Association for Computational Linguistics.
  38. 38.Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Mazumder, Amir Zadeh, and LouisPhilippe Morency. 2017b. Context-dependent sentiment analysis in user-generated videos. In Association for Computational Linguistics.
  39. 39.Soujanya Poria, Iti Chaturvedi, Erik Cambria, and Amir Hussain. 2016. Convolutional mkl based multimodal emotion recognition and sentiment analysis. In Data Mining (ICDM), 2016 IEEE 16th International Conference on. IEEE, pages 439–448.
  40. 40.Shyam Sundar Rajagopalan, Louis-Philippe Morency, Tadas Baltrušaitis, and Roland Goecke. 2016. Extending long short-term memory for multi-view structured learning. In European Conference on Computer Vision.
  41. 41.Fabien Ringeval, Andreas Sonderegger, Jürgen S. Sauer, and Denis Lalanne. 2013. Introducing the recola multimodal corpus of remote collaborative and affective interactions. In FG. IEEE Computer Society, pages 1–8.
  42. 42.Florian Schroff, Dmitry Kalenichenko, and James Philbin. 2015. Facenet: A unified embedding for face recognition and clustering. In CVPR. IEEE Computer Society, pages 815–823.
  43. 43.M. Schuster and K.K. Paliwal. 1997. Bidirectional recurrent neural networks. Trans. Sig. Proc. 45(11):2673–2681. https://doi.org/10.1109/78.650093.
  44. 44.Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, Christopher Potts, et al. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the conference on empirical methods in natural language processing (EMNLP). Citeseer, volume 1631, page 1642.
  45. 45.Rupesh K Srivastava, Klaus Greff, and Juergen Schmidhuber. 2015. Training very deep networks. In C. Cortes, N. D. Lawrence, D. D. Lee, M. Sugiyama, and R. Garnett, editors, Advances in Neural Information Processing Systems 28, Curran Associates, Inc., pages 2377–2385. http://papers.nips.cc/paper/5850-training-very-deep-networks.pdf.
  46. 46.Yaniv Taigman, Ming Yang, Marc'Aurelio Ranzato, and Lior Wolf. 2014. Deepface: Closing the gap to human-level performance in face verification. In Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition. IEEE Computer Society, Washington, DC, USA, CVPR '14, pages 1701–1708. https://doi.org/10.1109/CVPR.2014.220.
  47. 47.Edmund Tong, Amir Zadeh, Cara Jones, and LouisPhilippe Morency. 2017. Combating human trafficking with multimodal deep models. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). volume 1, pages 1547–1556.
  48. 48.George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A Nicolaou, Björn Schuller, and Stefanos Zafeiriou. 2016. Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network. In Acoustics, Speech and Signal Processing (ICASSP), 2016 IEEE International Conference on. IEEE, pages 5200–5204.
  49. 49.Haohan Wang, Aaksha Meghawat, Louis-Philippe Morency, and Eric P Xing. 2016. Select-additive learning: Improving cross-individual generalization in multimodal sentiment analysis. arXiv preprint arXiv:1609.05244 .
  50. 50.Martin Wöllmer, Felix Weninger, Tobias Knaup, Björn Schuller, Congkai Sun, Kenji Sagae, and LouisPhilippe Morency. 2013. Youtube movie reviews: Sentiment analysis in an audio-visual context. IEEE Intelligent Systems 28(3):46–53.
  51. 51.Jiahong Yuan and Mark Liberman. 2008. Speaker identification on the scotus corpus. Journal of the Acoustical Society of America 123(5):3878.
  52. 52.Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In Empirical Methods in Natural Language Processing, EMNLP.
  53. 53.Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018a. Memory fusion network for multi-view sequential learning. arXiv preprint arXiv:1802.00927 .
  54. 54.Amir Zadeh, Paul Pu Liang, Soujanya Poria, Prateek Vij, Erik Cambria, and Louis-Philippe Morency. 2018b. Multi-attention recurrent network for human communication comprehension. arXiv preprint arXiv:1802.00923 .
  55. 55.Amir Zadeh, Rowan Zellers, Eli Pincus, and LouisPhilippe Morency. 2016a. Mosi: Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos. arXiv preprint arXiv:1606.06259 .
  56. 56.Amir Zadeh, Rowan Zellers, Eli Pincus, and LouisPhilippe Morency. 2016b. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages. IEEE Intelligent Systems 31(6):82–88.
  57. 57.Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. 2016. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters 23(10):1499–1503.
  58. 58.Chunting Zhou, Chonglin Sun, Zhiyuan Liu, and Francis C. M. Lau. 2015. A c-lstm neural network for text classification. CoRR abs/1511.08630.
  59. 59.Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber. 2016. Recurrent Highway Networks. arXiv preprint arXiv:1607.03474 .

Citation

MLA
Zadeh, A. B., et al. “Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph”. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 2236–46, https://doi.org/10.18653/v1/P18-1208.
APA
Zadeh, A. B., Liang, P. P., Poria, S., Cambria, E., & Morency, L.-P. (2018). Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2236–2246. https://doi.org/10.18653/v1/P18-1208
Chicago
Zadeh, A. B., P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency. 2018. “Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph”. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2236–46. https://doi.org/10.18653/v1/P18-1208.
Harvard
Zadeh, A.B. et al. (2018) “Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 2236–2246. Available at: https://doi.org/10.18653/v1/P18-1208.
Vancouver
1. Zadeh AB, Liang PP, Poria S, Cambria E, Morency L-P (2018) Multimodal Language Analysis in the Wild: CMU-MOSEI Dataset and Interpretable Dynamic Fusion Graph. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 2236–2246

BibTeX

@inproceedings{bagher-zadeh-etal-2018-multimodal,
    title = "Multimodal Language Analysis in the Wild: {CMU}-{MOSEI} Dataset and Interpretable Dynamic Fusion Graph",
    author = "Bagher Zadeh, AmirAli  and
      Liang, Paul Pu  and
      Poria, Soujanya  and
      Cambria, Erik  and
      Morency, Louis-Philippe",
    editor = "Gurevych, Iryna  and
      Miyao, Yusuke",
    booktitle = "Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2018",
    address = "Melbourne, Australia",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/P18-1208/",
    doi = "10.18653/v1/P18-1208",
    pages = "2236--2246"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/