Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis

Hao YangYanyan ZhaoBing Qin

article2022EMNLP70 citations

Proposes a face-sensitive cross-modal translation framework that extracts and textualizes visual facial emotions to align them with specific target aspects, setting state-of-the-art results in multimodal aspect-based sentiment analysis on Twitter benchmarks.

Listen

Organizations analyzing public sentiment from social media face a persistent challenge: social posts increasingly pair text with imagery, yet automated systems struggle to accurately judge the emotional tone directed at specific subjects or entities. Human faces appear in over half of social media images and represent potent visual emotional cues. However, conventional sentiment models either focus solely on neutral scene descriptions or attempt complex image-text fusion that often introduces misleading noise and misaligns conflicting visual emotions.

The article demonstrates a novel framework called Face-Sensitive Image-to-Emotional-Text Translation (FITE). The primary objective is to evaluate how explicitly translating facial emotions into descriptive text and aligning those expressions to specific target aspects improves fine-grained sentiment classification across multimodal data.

The researchers designed an automated pipeline evaluated across two benchmark collections (Twitter-2015 and Twitter-2017) as well as a specialized facial subset (Twitter-face). The approach identifies faces and predicts key attributes such as age, gender, race, and expression (e.g., happy, sad, angry). It converts these attributes into structured text descriptions using predefined templates, eliminating the need to train specialized emotional image captioning models. In multi-face images, the system uses a pre-trained vision-language model (CLIP) to match the target subject to the correct face. Finally, it uses a gating mechanism to denoise and fuse textual content, scene captions, and facial emotion descriptions before classifying sentiment.

The evaluation yielded several key findings. First, the proposed method achieved state-of-the-art performance, with the large-scale configuration reaching a macro-F1 score of 74.79% on Twitter-2015 and 73.03% on Twitter-2017, outperforming leading baseline approaches. Second, on the dedicated Twitter-face dataset, the method improved accuracy by up to 6.68 percentage points and macro-F1 by 7.79 percentage points over the CapBERT baseline. Third, ablation analyses confirmed that facial emotion is the single most critical visual cue, with its removal causing the largest performance decline of over 2.1 percentage points. Additionally, removing the gating noise-reduction mechanism or the fine-grained alignment module consistently decreased accuracy by 1 to 2 percentage points.

These findings indicate that explicit cross-modal translation of human emotions offers a higher-performing, lower-complexity alternative to raw feature-fusion approaches. For decision-makers, this framework enhances the reliability and interpretability of automated social media monitoring, customer feedback analysis, and brand sentiment tracking while mitigating the computational cost of developing complex visual-language architectures from scratch.

Organizations handling multimodal analytics should consider integrating facial emotional parsing into their sentiment pipelines, especially when evaluating public-facing figures or customer imagery. Prior to production deployment, technical teams should conduct pilot testing to determine the proportion of facial content in their target data streams. Next steps should focus on extending visual sentiment extraction to non-human imagery using common-sense knowledge and exploring generative language models to produce more flexible emotional descriptions.

The primary limitation of this approach is its reliance on detectable human faces; it cannot identify visual sentiment in images that only feature landscapes, text graphics, logos, or buildings. Furthermore, system accuracy remains constrained by the quality of upstream facial recognition tools, which can degrade with blurry images, inaccurate attribute predictions, or sarcastic content where textual context contradicts facial expressions. Within datasets containing human imagery, confidence in the reported performance gains remains high.

No sufficiently relevant recommendations were found.

Cover for Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis

Abstract

Aspect-level multimodal sentiment analysis, which aims to identify the sentiment of the target aspect from multimodal data, recently has attracted extensive attention in the community of multimedia and natural language processing. Despite the recent success in textual aspect-based sentiment analysis, existing models mainly focused on utilizing the object-level semantic information in the image but ignore explicitly using the visual emotional cues, especially the facial emotions. How to distill visual emotional cues and align them with the textual content remains a key challenge to solve the problem. In this work, we introduce a face-sensitive image-to-emotional-text translation (FITE) method, which focuses on capturing visual sentiment cues through facial expressions and selectively matching and fusing with the target aspect in textual modality. To the best of our knowledge, we are the first that explicitly utilize the emotional information from images in the multimodal aspect-based sentiment analysis task. Experiment results show that our method achieves state-of-the-art results on the Twitter-2015 and Twitter-2017 datasets. The improvement demonstrates the superiority of our model in capturing aspect-level sentiment in multimodal data with facial expressions^1.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Task Formulation
  • 3.2 Overview
  • 3.3 Face-Sensitive Image-to-Emotional-Text Translation
  • 3.4 Aspect-Sensitive Alignment and Modification
  • 3.5 Gate-Based Multimodal Fusion Module
  • 4 Experiment
  • 4.1 Experimental Setup
  • 4.2 Compared Baselines
  • 4.3 Experimental Results and Analysis
  • 4.4 Ablation Study
  • 4.5 Case Study
  • 4.6 Visualization of Embedding Space
  • 5 Conclusion
  • Acknowledgments
  • Limitations
  • Ethics Statement
  • References
  • A Appendix
  • A.1 Face Description Rewriting
  • A.2 Analysis of Cases with Multiple Faces
  • A.3 Error Analysis

Knowls

  1. Knowl 1 — FITE translates facial emotion into text before aspect-level fusion

    model/method

    Face-sensitive image-to-emotional-text translation (FITE) predicts a sentiment label—positive, negative, or neutral—for a target aspect in a text–image sample. Its central design is to express facial attributes as text, align that text with the target aspect, and combine it with a separate textual description of the image scene. The pipeline has three stages: (1) detect faces and generate attribute-based face descriptions; (2) use CLIP to select and rewrite the face description relevant to the target aspect, while a caption transformer describes the overall scene; and (3) encode the aligned face description and scene caption with separate language models and fuse their representations using a learned gate before classification. The method uses pre-trained face-analysis tools rather than training a new emotional image-captioning model.

  2. Knowl 2 — Face attributes are converted into templated emotional descriptions

    model/method

    For each detected face, FITE applies pre-trained classifiers to predict age, race, gender, and facial expression. The expression classes are angry, fear, neutral, sad, disgust, happy, and surprise; the race classes are Asian, Black, White, Middle Eastern, Indian, and Latino. An attribute prediction whose confidence is below 0.40.4 is filtered out. The remaining attributes are placed in a manually designed natural-language pattern, exemplified by “A [Age]-year-old [Race] [Gender] with a [Emotion] expression.” This produces text descriptions that make facial emotion available to downstream language models without training a separate emotional image-captioning system.

  3. Knowl 3 — CLIP selects the face description most relevant to the target aspect

    model/method

    When an image contains multiple facial expressions, FITE forms a candidate text for each face by combining the target aspect with that face’s generated description. CLIP’s text encoder represents each candidate, and CLIP’s image encoder represents the image. Learnable projections map the two kinds of representations into a shared space; after L2L_2 normalization, FITE scores each candidate against the image using temperature-scaled cosine similarity and selects the highest-scoring description. The rewritten description retains the target aspect and the predicted facial expression, providing the emotional cue used for that aspect. This matching step is intended to reduce noise from faces whose expressions are irrelevant to the target. For a single-face image, FITE directly modifies the face description with the aspect and adds the sentence “There are 1 people in the image.” For an image with no detected face, it supplies “There are 0 people in the image.”

  4. Knowl 4 — Scene captions and face descriptions are fused with a learned gate

    equation

    FITE creates two text inputs, each containing the tweet text TT, target aspect AA, and either the aligned face description or the overall scene caption. Separate fine-tuned pre-trained language models encode these inputs, yielding pooled [CLS] vectors HD[CLS]H_D^{[CLS]} and HC[CLS]H_C^{[CLS]} in R768\mathbb{R}^{768}. Here, DD denotes the aligned face description and CC the scene caption. A learned elementwise gate combines the vectors to reduce noise in generated text, and a linear classifier predicts the sentiment class:

    gt=tanh⁡(WDHD[CLS]+WCHC[CLS]+bg),H=gt⊙HD[CLS]+(1−gt)⊙HC[CLS],g_t = \tanh(W_D H_D^{[CLS]} + W_C H_C^{[CLS]} + b_g), \qquad H = g_t \odot H_D^{[CLS]} + (1-g_t) \odot H_C^{[CLS]}, p(y∣H)=softmax⁡(WH+b).p(y\mid H)=\operatorname{softmax}(WH+b).

    The matrices WD,WC∈R768×768W_D,W_C\in\mathbb{R}^{768\times768}, classifier matrix W∈R768×3W\in\mathbb{R}^{768\times3}, and biases bg∈R768b_g\in\mathbb{R}^{768} and b∈R3b\in\mathbb{R}^{3} are learned. The gate uses the hyperbolic tangent transformation specified by the paper; ⊙\odot denotes elementwise multiplication. The three output classes are positive, negative, and neutral. The model parameters are optimized with standard cross-entropy loss.

  5. Knowl 5 — FITE improves benchmark performance, especially with stronger language encoders

    data/table

    The authors evaluated on Twitter-2015 and Twitter-2017, reporting accuracy and macro-F1 averaged over five independent training runs. Training used learning rate 5×10−55\times10^{-5}, 12 attention heads, dropout 0.1, batch size 16, 8 fine-tuning epochs, and maximum text length 256. The models were implemented in PyTorch and run on two NVIDIA Tesla V100 GPUs. FITE uses BERT; FITE-DE replaces BERT with BERTweet-base, and FITE-DE-Large uses BERTweet-Large. The results show that FITE-DE-Large achieved the highest scores among the listed systems on both datasets, while the BERT-based FITE was competitive but did not lead every metric.

    Method T-2015 Acc T-2015 Macro-F1 T-2017 Acc T-2017 Macro-F1
    Res-Target 59.88 46.48 58.59 53.98
    LSTM 70.30 63.43 61.67 57.97
    MGAN 71.17 64.21 64.75 61.46
    Bert 74.25 70.04 68.88 66.12
    MIMN 71.84 65.69 65.88 62.99
    ESAFN 73.38 67.37 67.83 64.22
    VilBERT 73.69 69.53 67.86 64.93
    TomBERT 77.15 71.15 70.34 68.03
    CapBERT 78.01 73.25 69.77 68.42
    CapBERT-DE 77.92 73.90 72.30 70.20
    VLP-MABSA 78.60 73.80 73.80 71.80
    FITE 78.49 73.90 70.90 68.70
    FITE-DE 78.64 74.30 72.98 71.97
    FITE-DE-Large 78.76 74.79 73.87 73.03

    FITE-DE-Large’s scores were 78.76 accuracy and 74.79 macro-F1 on Twitter-2015, and 73.87 accuracy and 73.03 macro-F1 on Twitter-2017. The paper reports that FITE-DE and FITE-DE-Large exceeded CapBERT-DE in macro-F1 by about 0.4 and 0.9 points on Twitter-2015, and by about 1.8 and 2.8 points on Twitter-2017, respectively. The authors also observed that the stronger language encoder improved FITE, indicating the importance of language-model context representations in the fusion stage.

  6. Knowl 6 — Face-containing examples show gains over text-only and caption-based models

    data/table

    The authors formed Twitter-face by selecting examples containing faces from Twitter-2015 and Twitter-2017, then evaluated sentiment classification on this subset. The split statistics and results below show higher scores for FITE than for text-only BERT and CapBERT, with further gains for FITE-DE. The parenthetical gains for the FITE variants are those reported in the paper relative to CapBERT.

    Split POS Neutral NEG Total Mean aspects Mean length
    Train 1285 1531 408 3204 1.37 16.52
    Valid. 449 514 137 1100 1.37 16.54
    Test 442 494 156 1092 1.39 16.53
    Method Accuracy Macro-F1
    BERT 67.02 63.32
    CapBERT 67.52 64.33
    FITE 69.50 (+1.98) 66.89 (+2.56)
    FITE-DE 74.20 (+6.68) 72.12 (+7.79)

    The paper also reports performance by face count. In Twitter-2015, 23.2% of images had one face and 14.7% had multiple faces; in Twitter-2017, the corresponding proportions were 28.9% and 22.4%. On Twitter-face, FITE-base scored 66.97 accuracy / 63.88 F1 for one-face examples and 73.21 / 71.86 for multiple-face examples. CapBERT scored 65.13 / 62.82 and 67.50 / 65.72 on those subsets, respectively. The larger difference on the multiple-face subset is consistent with the value of aspect-to-face alignment, though these results alone do not isolate that component.

  7. Knowl 7 — Ablations support alignment, scene captions, and gated fusion

    data/table

    The BERT-base FITE model was ablated on Twitter-2015 and Twitter-2017. Removing the gate replaces gated fusion with concatenation of the language-model pooled outputs before classification. Removing fine-grained alignment eliminates aspect-to-face matching, and removing the scene caption drops that image-to-text input. Each change lowers the reported scores, supporting the use of all three components; the largest accuracy drop in this comparison occurs when the scene caption is removed on Twitter-2015.

    Setting T-2015 Acc T-2015 Macro-F1 T-2017 Acc T-2017 Macro-F1
    FITE 78.49 73.90 70.90 68.70
    Without gate mechanism 77.47 72.70 68.93 67.12
    Without fine-grained alignment 77.89 72.90 70.03 68.32
    Without image caption of scene 76.56 72.28 70.01 67.95

    The authors interpret the loss without gating as evidence that the gate helps manage noise in generated descriptions, and the loss without alignment as evidence that matching visual emotional cues to the target aspect matters. The scene-caption ablation also indicates that scene semantics complement the facial-expression description.

  8. Knowl 8 — Structured facial descriptions outperform simpler descriptions and raw visual features

    data/table

    Two sets of experiments test how FITE represents visual information. In the description experiment, “single face” uses only the face with the highest emotion-prediction confidence in a multi-face image; “No Pattern” uses predicted attributes directly rather than the manually designed sentence pattern. Removing each facial attribute tests its contribution. In the raw-feature comparison, Res-BERT+BL combines ResNet image features with language-model text features, while Res-BERT_face+BL uses face-region features; FITE+image feature and FITE+face feature add those visual features to FITE. Across the comparisons, the complete templated description is strongest, omitting emotion causes the largest drop among the facial-attribute removals, and the tested raw-feature additions do not match complete FITE.

    Description setting T-2015 Acc T-2015 Macro-F1 T-2017 Acc T-2017 Macro-F1
    Single face 76.60 71.19 68.53 66.92
    No Pattern 77.33 72.43 69.36 67.54
    FITE 78.49 73.90 70.90 68.70
    Without age 78.20 73.28 70.62 68.44
    Without race 77.82 73.44 70.25 67.97
    Without gender 77.82 73.07 69.93 67.57
    Without emotion 76.22 71.44 68.80 66.35
    Visual-feature setting T-2015 Acc T-2015 Macro-F1 T-2017 Acc T-2017 Macro-F1
    Bert 74.25 70.04 68.88 66.12
    Res-BERT+BL 75.02 69.21 69.20 66.48
    Res-BERT_face+BL 76.37 71.34 69.25 67.42
    FITE 78.49 73.90 70.90 68.70
    FITE+face feature 77.37 73.01 69.35 67.32
    FITE+image feature 76.54 71.83 69.25 67.24

    The results show that removing any one of age, race, gender, or emotion reduces performance, with the largest decline when emotion is omitted. The authors suggest that adding raw visual features can introduce noise when those features lack cross-modal alignment, whereas translating facial cues into text supplies more directly usable emotional information.

  9. Knowl 9 — Qualitative cases illustrate aspect matching and non-human targets

    empirical result

    In the authors’ illustrative comparisons with BERT and CapBERT, FITE’s generated descriptions expose facial emotion that neutral scene captions may miss. A happy-expression description enabled a correct positive prediction in one example where the other systems predicted neutral. In a multi-face example, CLIP-based matching linked the target person to the relevant happy expression rather than an expression from another face. Another example concerned “FBI,” a non-human aspect: FITE used emotional information from people in the image to help classify the aspect. A further example illustrated the gate suppressing an emotional cue associated with a non-matching aspect. These cases show how the proposed components can be useful beyond aspects that directly name a person, but they are illustrative examples rather than a controlled quantitative analysis.

  10. Knowl 10 — FITE depends on detectable and informative facial expressions

    limitation

    The authors identify the method’s reliance on faces as a major limitation: it is not well suited to examples without facial emotion, and its fallback description that nobody appears in the image does not recover emotional information from other objects or scenes. They argue that interpreting emotion from arbitrary visual objects is difficult and can vary between observers. Errors also arise when no useful face is detected or facial-attribute classifiers predict incorrectly, especially for blurry images. Facial expressions may fail to represent the complete visual sentiment when image and text sentiment conflict or when sarcasm is involved. Finally, face descriptions rely on hand-designed templates; the authors report that generative alternatives were constrained by model performance and the lack of public emotional image-to-text datasets.

Coverage note — The t-SNE embedding visualization and full Twitter-2015/Twitter-2017 split-level class-count tables are omitted because they are supplementary descriptive analyses rather than independent method claims; background, related work, and acknowledgements are outside scope.

References

  1. 1.Navneet Dalal and Bill Triggs. 2005. Histograms of oriented gradients for human detection. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), volume 1, pages 886–893. Ieee.
  2. 2.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  3. 3.Feifan Fan, Yansong Feng, and Dongyan Zhao. 2018a. Multi-grained attention network for aspect-level sentiment classification. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 3433–3442.
  4. 4.Shaojing Fan, Zhiqi Shen, Ming Jiang, Bryan L Koenig, Juan Xu, Mohan S Kankanhalli, and Qi Zhao. 2018b. Emotional attention: A study of image sentiment and visual attention. In Proceedings of the IEEE Conference on computer vision and pattern recognition, pages 7521–7531.
  5. 5.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778.
  6. 6.Mengting Hu, Shiwan Zhao, Li Zhang, Keke Cai, Zhong Su, Renhong Cheng, and Xiaowei Shen. 2019. CAN: Constrained attention networks for multi-aspect sentiment analysis. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4601–4610, Hong Kong, China. Association for Computational Linguistics.
  7. 7.Xincheng Ju, Dong Zhang, Rong Xiao, Junhui Li, Shoushan Li, Min Zhang, and Guodong Zhou. 2021. Joint multi-modal aspect-sentiment analysis with auxiliary cross-modal relation detection. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 4395–4405.
  8. 8.Zaid Khan and Yun Fu. 2021. Exploiting bert for multimodal target sentiment classification through input space translation. In Proceedings of the 29th ACM International Conference on Multimedia, pages 3034–3042.
  9. 9.Shan Li, Weihong Deng, and JunPing Du. 2017. Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2852–2861.
  10. 10.Yuncong Li, Cunxiang Yin, and Sheng-hua Zhong. 2020. Sentence constituent-aware aspect-category sentiment analysis with graph attention networks. In CCF International Conference on Natural Language Processing and Chinese Computing, pages 815–827. Springer.
  11. 11.Yan Ling, Jianfei Yu, and Rui Xia. 2022. Vision-language pre-training for multimodal aspect-based sentiment analysis. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 2149–2159. Association for Computational Linguistics.
  12. 12.Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems, 32.
  13. 13.Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. 2020. BERTweet: A pre-trained language model for English Tweets. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 9–14.
  14. 14.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR.
  15. 15.Sefik Ilkin Serengil and Alper Ozpinar. 2020. Lightface: A hybrid deep face recognition framework. In 2020 Innovations in Intelligent Systems and Applications Conference (ASYU), pages 23–27. IEEE.
  16. 16.Sefik Ilkin Serengil and Alper Ozpinar. 2021. Hyper-extended lightface: A facial attribute analysis framework. In 2021 International Conference on Engineering and Emerging Technologies (ICEET), pages 1–4. IEEE.
  17. 17.Jiawei Wang, Zhe Liu, Victor Sheng, Yuqing Song, and Chenjian Qiu. 2021. Saliencybert: Recurrent attention network for target-oriented multimodal sentiment classification. In Chinese Conference on Pattern Recognition and Computer Vision (PRCV), pages 3–15. Springer.
  18. 18.Yequan Wang, Minlie Huang, Xiaoyan Zhu, and Li Zhao. 2016. Attention-based lstm for aspect-level sentiment classification. In Proceedings of the 2016 conference on empirical methods in natural language processing, pages 606–615.
  19. 19.Nan Xu, Wenji Mao, and Guandan Chen. 2019. Multi-interactive memory network for aspect based multimodal sentiment analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 371–378.
  20. 20.Wei Xue and Tao Li. 2018. Aspect based sentiment analysis with gated convolutional networks. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2514–2523, Melbourne, Australia. Association for Computational Linguistics.
  21. 21.Jianfei Yu and Jing Jiang. 2019. Adapting bert for target-oriented multimodal sentiment classification. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19, pages 5408–5414. International Joint Conferences on Artificial Intelligence Organization.
  22. 22.Jianfei Yu, Jing Jiang, and Rui Xia. 2019. Entity-sensitive attention and fusion network for entity-level multimodal sentiment classification. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:429–439.
  23. 23.Peisong Zhu, Zhuang Chen, Haojie Zheng, and Tieyun Qian. 2019. Aspect aware learning for aspect category sentiment analysis. ACM Transactions on Knowledge Discovery from Data (TKDD), 13(6):1–21.

Citation

MLA
Yang, H., et al. “Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3324–35, https://doi.org/10.18653/V1/2022.EMNLP-MAIN.219.
APA
Yang, H., Zhao, Y., & Qin, B. (2022). Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3324–3335. https://doi.org/10.18653/V1/2022.EMNLP-MAIN.219
Chicago
Yang, H., Y. Zhao, and B. Qin. 2022. “Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3324–35. https://doi.org/10.18653/V1/2022.EMNLP-MAIN.219.
Harvard
Yang, H., Zhao, Y. and Qin, B. (2022) “Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3324–3335. Available at: https://doi.org/10.18653/V1/2022.EMNLP-MAIN.219.
Vancouver
1. Yang H, Zhao Y, Qin B (2022) Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3324–3335

BibTeX

@inproceedings{Yang_2022, title={Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis}, url={http://dx.doi.org/10.18653/V1/2022.EMNLP-MAIN.219}, DOI={10.18653/v1/2022.emnlp-main.219}, booktitle={Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing}, publisher={Association for Computational Linguistics}, author={Yang, Hao and Zhao, Yanyan and Qin, Bing}, year={2022}, pages={3324–3335} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/