Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

Haoyu ZhangYu WangGuanghao YinKejun LiuYuanyuan LiuTianshu Yu

article2023EMNLP221 citations

Proposes an adaptive language-guided transformer that suppresses conflicting and irrelevant visual and acoustic signals using multi-scale language cues, achieving state-of-the-art multimodal sentiment analysis performance across standard benchmarks like MOSI and MOSEI.

Listen

Multimodal sentiment analysis aims to automatically understand human attitudes and emotions by evaluating spoken language, facial expressions in video, and acoustic cues in audio. In practical deployments such as healthcare monitoring and automated human-computer interaction, non-verbal modalities frequently introduce sentiment-irrelevant or conflicting information—such as sudden background noise, shifts in lighting, and head poses. Because conventional multimodal methods fuse all signals directly without accounting for these discrepancies, extraneous noise degrades the overall reliability and accuracy of automated decision-making systems.

The article demonstrates that using text as a primary anchor to guide and filter visual and acoustic inputs suppresses distracting, non-verbal noise and yields more accurate sentiment predictions. To achieve this, the authors introduce the Adaptive Language-guided Multimodal Transformer framework, which extracts low-dimensional representations across video, audio, and language before combining them.

To evaluate the system, the authors conducted extensive experiments across three standard benchmark datasets: MOSI, MOSEI, and CH-SIMS, which encompass diverse video clips annotated with sentiment scores. The architecture first compresses raw modality inputs into unified, low-dimensional tokens. An Adaptive Hyper-modality Learning module then employs multi-scale language features to dynamically guide and fuse the visual and acoustic features into a single complementary representation, which is merged with language features in a cross-modality fusion step to produce the final sentiment prediction.

The evaluation produced four key findings. First, the proposed framework achieves top-tier results across standard benchmarks, reaching seven-class sentiment accuracy of 49.42% on MOSI and five-class accuracy of 45.73% on CH-SIMS, outperforming prior baselines. Second, removing the language-guided hyper-modality learning component causes seven-class accuracy on MOSI to plunge from 49.42% to 34.40%, confirming that filtering auxiliary modalities is vital for robust performance. Third, visual inputs were found to contribute more complementary sentiment value than audio inputs, as evidenced by higher learned attention weights. Fourth, the architecture maintains computational efficiency, operating with 2.50 million parameters and training via a single standard loss function without the fragile multi-task optimization tuning required by competing models.

These findings indicate that multimodal systems in production can achieve higher accuracy and operational stability by prioritizing text as a structural guide rather than treating all sensory inputs as equal peers. This design choice reduces deployment risk and computational overhead, avoiding complex hyperparameter tuning while yielding consistent cross-scenario performance. For organizations operating automated customer service or clinical sentiment tools, adopting a language-guided filtering approach can prevent false alerts triggered by environmental background noise.

Organizations developing multimodal sentiment solutions should transition from symmetric fusion pipelines to language-anchored fusion architectures to reduce error rates. In terms of limitations, the framework relies on Transformer blocks that require substantial training data, meaning current improvements on continuous regression metrics remain modest due to the relatively small size of academic benchmark datasets. Decision-makers should validate the architecture against larger, domain-specific proprietary datasets before rolling it out to production-scale operations.

arXiv: 2310.05804Haoyu-ha/ALMT
Cover for Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

Abstract

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Multimodal Sentiment Analysis
  • 2.2 Transformer
  • 3 Method
  • 3.1 Overview
  • 3.2 Multimodal Input
  • 3.3 Modality Embedding
  • 3.4 Adaptive Hyper-modality Learning
  • 3.4.1 Construction of Two-scale Language Features
  • 3.4.2 Adaptive Hyper-modality Learning Layer
  • 3.5 Multimodal Fusion and Output
  • 3.6 Overall Learning Objectives
  • 4 Experiments
  • 4.1 Datasets
  • 4.2 Evaluation Criteria
  • 4.3 Baselines
  • 4.4 Performance Comparison
  • 4.5 Ablation Study and Analysis
  • 4.5.1 Effects of Different Modalities
  • 4.5.2 Effects of Different Components
  • 4.5.3 Effects of Different Query, Key, and Value Settings in Fusion Transformer
  • 4.5.4 Effects of the Guidance of Different Language Features in AHL
  • 4.5.5 Effects of Different Fusion Techniques
  • 4.5.6 Analysis on Model Complexity
  • 4.5.7 Visualization of Attention in AHL
  • 4.5.8 Visualization of Robustness of AHL
  • 4.5.9 Visualization of Different Representations
  • 4.5.10 Visualization of Convergence Performance
  • 5 Conclusion
  • Limitations
  • References
  • A Hyper-parameters
  • A.1 Overview
  • A.2 Effects of Length Settings of Modality Feature
  • A.3 Effects of Depth Settings of AHL
  • A.4 Effects of Depth Settings of Fusion Transformer

Knowls

  1. Knowl 1 — ALMT compresses each modality into a short token sequence

    model/method

    Adaptive Language-guided Multimodal Transformer (ALMT) processes language, audio, and video features separately before cross-modal learning. It uses precomputed BERT language features, Librosa audio features, and OpenFace visual features. For modality m∈{l,a,v}m\in\{l,a,v\}, let Um∈RTm×dmU_m\in\mathbb{R}^{T_m\times d_m} denote its input sequence, where TmT_m is its input length and dmd_m its feature dimension. ALMT concatenates each sequence with a modality-specific initialized token sequence, then applies a one-layer Transformer to produce Hm1∈RT×dH_m^1\in\mathbb{R}^{T\times d}. The learned token sequence retains modality information while reducing the input to a common length and dimension. The experiments use T=8T=8 and d=128d=128.

  2. Knowl 2 — Adaptive Hyper-modality Learning uses language to weight audio and visual features

    model/method

    ALMT’s Adaptive Hyper-modality Learning (AHL) module forms a hyper-modality sequence from audio and visual features, guided by language features at three scales. The low-scale language sequence L1L_1 is the embedded language representation; two successive Transformer layers produce the middle- and high-scale sequences L2L_2 and L3L_3. At AHL layer j∈{1,2,3}j\in\{1,2,3\}, let Lj∈RT×dL_j\in\mathbb{R}^{T\times d} be its language guidance, and let A,V∈RT×dA,V\in\mathbb{R}^{T\times d} be the embedded audio and visual sequences. Language-derived queries are compared separately with audio and visual keys: α=softmax(QjKa⊤/dk)\alpha=softmax(Q_jK_a^\top/\sqrt{d_k}) and β=softmax(QjKv⊤/dk)\beta=softmax(Q_jK_v^\top/\sqrt{d_k}), where Qj=LjWQQ_j=L_jW_Q, Ka=AWKaK_a=AW_{Ka}, and Kv=VWKvK_v=VW_{Kv}. The matrices α,β∈RT×T\alpha,\beta\in\mathbb{R}^{T\times T} are row-normalized attention weights; the projection matrices map dimension dd to per-head dimension dkd_k. With audio and visual values Va,VvV_a,V_v formed by value projections, head concatenation, and output projection to dimension dd, the hyper-modality sequence is updated as Xj=Xj−1+αVa+βVvX_j=X_{j-1}+\alpha V_a+\beta V_v, starting from an initialized sequence X0∈RT×dX_0\in\mathbb{R}^{T\times d}. The model uses eight attention heads with dk=16d_k=16. The intended effect is to reduce irrelevant or conflicting auxiliary-modality information while retaining information complementary to language.

  3. Knowl 3 — Language-anchored fusion produces the sentiment representation

    model/method

    After three AHL layers, ALMT prepends an initialized token to the final language sequence Hl3H_l^3 and final hyper-modality sequence Hhyper3H_{hyper}^3. A cross-modality fusion Transformer uses the language sequence as query and the hyper-modality sequence as key and value, producing a joint representation H∈R1×dH\in\mathbb{R}^{1\times d} from the prepended token. A classifier maps HH to sentiment prediction y^\hat y. Training uses only mean squared error between the scalar sentiment labels yny_n and predictions y^n\hat y_n: L=(1/N)∑n=1N(yn−y^n)2L=(1/N)\sum_{n=1}^{N}(y_n-\hat y_n)^2, where NN is the number of training examples. The fusion Transformer uses eight attention heads; its depth is dataset-dependent.

  4. Knowl 4 — Datasets and common ALMT training configuration

    experimental setup

    ALMT was evaluated on three trimodal sentiment datasets. MOSI contains 2,199 samples, split into 1,284 training, 229 validation, and 686 test samples, with sentiment scores from −3 to 3. MOSEI contains 22,856 samples, split into 16,326 training, 1,871 validation, and 4,659 test samples, also scored from −3 to 3. Chinese CH-SIMS contains 2,281 samples, split into 1,368 training, 456 validation, and 457 test samples, scored from −1 to 1. Across datasets, the modality-token length is 8, embedding dimension is 128, modality-embedding Transformer depth is 1, and AHL depth is 3. Fusion-Transformer depth is 2 for MOSI and 4 for MOSEI and CH-SIMS. Training uses AdamW, batch size 64, initial learning rate 10−410^{-4}, 200 epochs, warm-up, and cosine annealing. Reported metrics include classification accuracy at different sentiment granularities (Acc-2, Acc-3, Acc-5, Acc-7), F1, mean absolute error (MAE), and prediction–human correlation (Corr).

  5. Knowl 5 — ALMT benchmark results on MOSI, MOSEI, and CH-SIMS

    empirical result

    On MOSI, ALMT reports Acc-7 49.42, Acc-5 56.41, Acc-2 84.55/86.43, F1 84.57/86.47, MAE 0.683, and Corr 0.805. On MOSEI, it reports Acc-7 54.28, Acc-5 55.96, Acc-2 84.78/86.79, F1 85.19/86.86, MAE 0.526, and Corr 0.779. The paired Acc-2 and F1 values are the two binary evaluation settings reported by the paper. On CH-SIMS, ALMT reports Acc-5 45.73, Acc-3 68.93, Acc-2 81.19, F1 81.57, MAE 0.404, and Corr 0.619. The paper reports state-of-the-art performance on nearly all evaluated metrics. It highlights a 1.69% relative Acc-7 improvement over CHFN on MOSI, and relative improvements over Self-MM on CH-SIMS of 1.44% in Acc-2 and 1.40% in F1.

  6. Knowl 6 — AHL and the other ALMT components improve ablation performance

    empirical result

    Ablations on MOSI and CH-SIMS show that AHL contributes substantially to ALMT performance. The full model scores MOSI Acc-7 49.42 and MAE 0.683, and CH-SIMS Acc-5 45.73 and MAE 0.404. Replacing AHL with feature concatenation (w/o AHL) lowers these to 34.40 and 0.952 on MOSI, and 38.29 and 0.444 on CH-SIMS. Removing the fusion Transformer gives 48.69/0.703 on MOSI and 43.76/0.410 on CH-SIMS; removing modality embedding gives 47.96/0.701 and 43.11/0.429, respectively. Modality-removal tests also reduce performance: removing audio gives 48.69/0.705 on MOSI and 45.08/0.416 on CH-SIMS; removing video gives 47.96/0.704 and 44.64/0.403. Removing audio or video together with AHL gives, respectively, 46.91/0.724 and 47.08/0.726 on MOSI, and 43.54/0.407 and 43.76/0.406 on CH-SIMS. Removing both audio and video gives 46.79/0.752 on MOSI and 40.26/0.405 on CH-SIMS. The results support the contribution of AHL and indicate that auxiliary modalities can be useful when their information is handled by the model.

  7. Knowl 7 — Guidance from all three language scales yields the best AHL result

    empirical result

    On MOSI and CH-SIMS, the best tested AHL configuration guides hyper-modality learning with all three language representations L1,L2,L3L_1,L_2,L_3. It achieves MOSI Acc-7 49.42 and MAE 0.683, and CH-SIMS Acc-5 45.73 and MAE 0.404. With no language-scale guidance, the corresponding results are 34.40/0.952 and 38.29/0.444. Single-scale guidance by L1L_1, L2L_2, or L3L_3 yields, respectively, MOSI 47.38/0.704, 48.10/0.709, or 48.54/0.711, and CH-SIMS 43.54/0.412, 43.11/0.415, or 43.98/0.412. Using pairs of scales yields MOSI 46.36/0.736 for L1,L2L_1,L_2, 48.10/0.707 for L1,L3L_1,L_3, and 47.81/0.729 for L2,L3L_2,L_3; the corresponding CH-SIMS results are 45.51/0.417, 44.20/0.409, and 43.76/0.416. Each pair is accuracy/MAE. These ablations show that using every tested language scale together outperforms the tested single-scale and two-scale alternatives on both datasets.

  8. Knowl 8 — Language queries and Transformer fusion outperform tested alternatives

    empirical result

    The fusion ablations on MOSI and CH-SIMS favor using language features as query and hyper-modality features as key and value. This setting achieves MOSI Acc-7 49.42 and MAE 0.683, and CH-SIMS Acc-5 45.73 and MAE 0.404. Reversing query and key/value roles gives 48.10/0.707 and 44.64/0.410. The paper also compares fusion techniques; each pair below is accuracy/MAE. Concatenation gives 48.69/0.703 on MOSI and 43.76/0.410 on CH-SIMS; addition gives 46.36/0.706 and 42.45/0.411; GRU gives 47.81/0.710 and 44.86/0.414; tensor fusion gives 47.23/0.710 and 44.20/0.403; low-rank fusion gives 46.65/0.715 and 45.08/0.408. The cross-modality Transformer used by ALMT gives 49.42/0.683 and 45.73/0.404. Thus the proposed Transformer fusion is strongest on MOSI and has the highest classification accuracy among these methods on CH-SIMS, though tensor fusion has lower CH-SIMS MAE (0.403).

  9. Knowl 9 — Attention visualizations provide qualitative evidence of AHL filtering

    empirical result

    On CH-SIMS, visualizations of average attention in the final AHL layer show greater attention to visual than audio features, which the authors interpret as visual features providing more complementary information in that analysis. In a test on one randomly selected sample, adding random noise to a peak visual frame substantially decreases the attention linking language features to that frame. This is qualitative evidence that AHL can reduce the influence of a perturbed visual feature; it is not a quantified robustness result across the dataset. A t-SNE visualization also shows audio and visual feature distributions separated before AHL, while the learned hyper-modality representations occupy a more convergent distribution, which the authors interpret as making fusion easier.

  10. Knowl 10 — The authors identify limited training data as an ALMT limitation

    limitation

    ALMT is a Transformer-based model that requires substantial training, and its performance may therefore be constrained by the small size of current sentiment datasets. The paper specifically notes that fine-grained regression measures such as MAE and Corr may need more training data than classification metrics and consequently show relatively small improvements over other advanced methods.

Coverage note — The parameter-count and convergence comparisons, and the token-length and Transformer-depth tuning sweeps, are omitted as secondary efficiency and hyperparameter analyses rather than load-bearing parts of the method or its main findings.

References

  1. 1.Tadas Baltrusaitis, Amir Zadeh, Yao Chong Lim, and Louis-Philippe Morency. 2018. Openface 2.0: Facial behavior analysis toolkit. In 13th IEEE International Conference on Automatic Face & Gesture Recognition, FG 2018, Xi’an, China, May 15-19, 2018, pages 59–66. IEEE Computer Society.
  2. 2.Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In Proceedings of the 16th European Conference on Computer Vision, volume 12346, pages 213–229.
  3. 3.Weidong Chen, Xiaofen Xing, Xiangmin Xu, Jianxin Pang, and Lan Du. 2022. Speechformer: A hierarchical efficient framework incorporating the characteristics of speech. In Proceedings of the 23rd Annual Conference of the International Speech Communication Association, pages 346–350.
  4. 4.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR.
  5. 5.Jiwei Guo, Jiajia Tang, Weichen Dai, Yu Ding, and Wanzeng Kong. 2022. Dynamically adjust word representations using unaligned multimodal information. In Proceedings of the 30th ACM International Conference on Multimedia, MM ’22, page 3394–3402. Association for Computing Machinery.
  6. 6.Wei Han, Hui Chen, and Soujanya Poria. 2021. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 9180–9192, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  7. 7.Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria. 2020. MISA: modality-invariant and -specific representations for multimodal sentiment analysis. In Proceedings of the 28th ACM international conference on multimedia, pages 1122–1131. ACM.
  8. 8.Jian Huang, Jianhua Tao, Bin Liu, Zheng Lian, and Mingyue Niu. 2020. Multimodal transformer fusion for continuous emotion recognition. In 2020 IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP), pages 3507–3511. IEEE.
  9. 9.Yingying Jiang, Wei Li, M. Shamim Hossain, Min Chen, Abdulhameed Alelaiwi, and Muneer Al-Hammadi. 2020. A snapshot research and implementation of multimodal information fusion for data-driven emotion recognition. Information Fusion, 53:209–221.
  10. 10.Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT, pages 4171–4186.
  11. 11.Yuanyuan Liu, Wenbin Wang, Chuanxu Feng, Haoyu Zhang, Zhe Chen, and Yibing Zhan. 2023a. Expression snippet transformer for robust video-based facial expression recognition. Pattern Recognition, 138:109368.
  12. 12.Yuanyuan Liu, Haoyu Zhang, Yibing Zhan, Zijing Chen, Guanghao Yin, Lin Wei, and Zhe Chen. 2023b. Noise-resistant multimodal transformer for emotion recognition. arXiv preprint arXiv:2305.02814.
  13. 13.Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, AmirAli Bagher Zadeh, and Louis-Philippe Morency. 2018. Efficient low-rank multimodal fusion with modality-specific factors. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2247–2256, Melbourne, Australia. Association for Computational Linguistics.
  14. 14.Fengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan, and Guosheng Lin. 2021. Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2554–2562. Computer Vision Foundation / IEEE.
  15. 15.Huisheng Mao, Ziqi Yuan, Hua Xu, Wenmeng Yu, Yihe Liu, and Kai Gao. 2022. M-sena: An integrated platform for multimodal sentiment analysis. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 204–213.
  16. 16.Brian McFee, Colin Raffel, Dawen Liang, Daniel P. W. Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. 2015. librosa: Audio and music signal analysis in python. In Proceedings of the 14th Python in Science Conference 2015 (SciPy 2015), Austin, Texas, July 6 - 12, 2015, pages 18–24. scipy.org.
  17. 17.Yongfeng Qian, Yin Zhang, Xiao Ma, Han Yu, and Limei Peng. 2019. EARS: emotion-aware recommender system based on hybrid information fusion. Information Fusion, 46:141–146.
  18. 18.Wasifur Rahman, Md Kamrul Hasan, Sangwu Lee, Amir Zadeh, Chengfeng Mao, Louis-Philippe Morency, and Ehsan Hoque. 2020. Integrating multimodal information in large pretrained transformers. In Proceedings of the conference. Association for Computational Linguistics. Meeting, volume 2020, page 2359. NIH Public Access.
  19. 19.Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019a. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6558–6569, Florence, Italy. Association for Computational Linguistics.
  20. 20.Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019b. Learning factorized multimodal representations. In ICLR.
  21. 21.Laurens Van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-sne. Journal of machine learning research, 9(11).
  22. 22.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, pages 5998–6008.
  23. 23.Dingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du, and Lihua Zhang. 2022. Disentangled representation learning for multimodal emotion recognition. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1642–1651. ACM.
  24. 24.Wenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu, Yixiao Ma, Jiele Wu, Jiyun Zou, and Kaicheng Yang. 2020. CH-SIMS: A chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 3718–3727. Association for Computational Linguistics.
  25. 25.Wenmeng Yu, Hua Xu, Ziqi Yuan, and Jiele Wu. 2021. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10790–10797.
  26. 26.Ziqi Yuan, Wei Li, Hua Xu, and Wenmeng Yu. 2021. Transformer-based feature reconstruction network for robust multimodal sentiment analysis. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4400–4407. ACM.
  27. 27.Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2017. Tensor fusion network for multimodal sentiment analysis. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1103–1114, Copenhagen, Denmark. Association for Computational Linguistics.
  28. 28.Amir Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency. 2018. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, pages 2236–2246.
  29. 29.Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency. 2016. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages. IEEE Intelligent Systems, 31(6):82–88.

Citation

MLA
Zhang, H., et al. “Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 756–67, https://doi.org/10.18653/v1/2023.emnlp-main.49.
APA
Zhang, H., Wang, Y., Yin, G., Liu, K., Liu, Y., & Yu, T. (2023). Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 756–767. https://doi.org/10.18653/v1/2023.emnlp-main.49
Chicago
Zhang, H., Y. Wang, G. Yin, K. Liu, Y. Liu, and T. Yu. 2023. “Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis”. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 756–67. https://doi.org/10.18653/v1/2023.emnlp-main.49.
Harvard
Zhang, H. et al. (2023) “Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 756–767. Available at: https://doi.org/10.18653/v1/2023.emnlp-main.49.
Vancouver
1. Zhang H, Wang Y, Yin G, Liu K, Liu Y, Yu T (2023) Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 756–767

BibTeX

@inproceedings{zhang-etal-2023-learning-language,
    title = "Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis",
    author = "Zhang, Haoyu  and
      Wang, Yu  and
      Yin, Guanghao  and
      Liu, Kejun  and
      Liu, Yuanyuan  and
      Yu, Tianshu",
    editor = "Bouamor, Houda  and
      Pino, Juan  and
      Bali, Kalika",
    booktitle = "Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2023",
    address = "Singapore",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.emnlp-main.49/",
    doi = "10.18653/v1/2023.emnlp-main.49",
    pages = "756--767"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/