Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis
Hao YangYanyan ZhaoBing Qin
Proposes a face-sensitive cross-modal translation framework that extracts and textualizes visual facial emotions to align them with specific target aspects, setting state-of-the-art results in multimodal aspect-based sentiment analysis on Twitter benchmarks.
Organizations analyzing public sentiment from social media face a persistent challenge: social posts increasingly pair text with imagery, yet automated systems struggle to accurately judge the emotional tone directed at specific subjects or entities. Human faces appear in over half of social media images and represent potent visual emotional cues. However, conventional sentiment models either focus solely on neutral scene descriptions or attempt complex image-text fusion that often introduces misleading noise and misaligns conflicting visual emotions.
The article demonstrates a novel framework called Face-Sensitive Image-to-Emotional-Text Translation (FITE). The primary objective is to evaluate how explicitly translating facial emotions into descriptive text and aligning those expressions to specific target aspects improves fine-grained sentiment classification across multimodal data.
The researchers designed an automated pipeline evaluated across two benchmark collections (Twitter-2015 and Twitter-2017) as well as a specialized facial subset (Twitter-face). The approach identifies faces and predicts key attributes such as age, gender, race, and expression (e.g., happy, sad, angry). It converts these attributes into structured text descriptions using predefined templates, eliminating the need to train specialized emotional image captioning models. In multi-face images, the system uses a pre-trained vision-language model (CLIP) to match the target subject to the correct face. Finally, it uses a gating mechanism to denoise and fuse textual content, scene captions, and facial emotion descriptions before classifying sentiment.
The evaluation yielded several key findings. First, the proposed method achieved state-of-the-art performance, with the large-scale configuration reaching a macro-F1 score of 74.79% on Twitter-2015 and 73.03% on Twitter-2017, outperforming leading baseline approaches. Second, on the dedicated Twitter-face dataset, the method improved accuracy by up to 6.68 percentage points and macro-F1 by 7.79 percentage points over the CapBERT baseline. Third, ablation analyses confirmed that facial emotion is the single most critical visual cue, with its removal causing the largest performance decline of over 2.1 percentage points. Additionally, removing the gating noise-reduction mechanism or the fine-grained alignment module consistently decreased accuracy by 1 to 2 percentage points.
These findings indicate that explicit cross-modal translation of human emotions offers a higher-performing, lower-complexity alternative to raw feature-fusion approaches. For decision-makers, this framework enhances the reliability and interpretability of automated social media monitoring, customer feedback analysis, and brand sentiment tracking while mitigating the computational cost of developing complex visual-language architectures from scratch.
Organizations handling multimodal analytics should consider integrating facial emotional parsing into their sentiment pipelines, especially when evaluating public-facing figures or customer imagery. Prior to production deployment, technical teams should conduct pilot testing to determine the proportion of facial content in their target data streams. Next steps should focus on extending visual sentiment extraction to non-human imagery using common-sense knowledge and exploring generative language models to produce more flexible emotional descriptions.
The primary limitation of this approach is its reliance on detectable human faces; it cannot identify visual sentiment in images that only feature landscapes, text graphics, logos, or buildings. Furthermore, system accuracy remains constrained by the quality of upstream facial recognition tools, which can degrade with blurry images, inaccurate attribute predictions, or sarcastic content where textual context contradicts facial expressions. Within datasets containing human imagery, confidence in the reported performance gains remains high.
- Paper: Vision-Language Pre-Training for Multimodal Aspect-Based Sentiment Analysis, Yan Ling et al. (2022). Read this task-specific vision-language pretraining work first to understand the multimodal aspect-sentiment framework that FITE’s face-focused translation and alignment approach builds upon.
No sufficiently relevant recommendations were found.
