Built independently by an author, for readers. Read the story and support ChapterPal

keyword

Twitter-face dataset

The Twitter-face dataset is a multimodal dataset composed of social media posts containing paired text and images with detectable human faces and facial expressions. Derived from standard multimodal Twitter sentiment benchmarks, it is designed to support multimodal aspect-based sentiment analysis. The dataset is used to evaluate how effectively computational models can extract visual emotional cues from facial expressions and align them with textual entities to classify the sentiment polarity of specific target aspects.

1 item

Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis

Face-Sensitive Image-to-Emotional-Text Cross-modal Translation for Multimodal Aspect-based Sentiment Analysis

Hao Yang, Yanyan Zhao, Bing Qin

OrganizationsHarbin Institute of Technology

Why you should read this

Proposes a face-sensitive cross-modal translation framework that extracts and textualizes visual facial emotions to align them with specific target aspects, setting state-of-the-art results in multimodal aspect-based sentiment analysis on Twitter benchmarks.

Aspect-level multimodal sentiment analysis, which aims to identify the sentiment of the target aspect from multimodal data, recently has attracted extensive attention in the community of multimedia and natural language processing. Despite the recent success in textual aspect-based sentiment analysis, existing models mainly focused on utilizing the object-level semantic information in the image but ignore explicitly using the visual emotional cues, especially the facial emotions. How to distill visual emotional cues and align them with the textual content remains a key challenge to solve the problem. In this work, we introduce a face-sensitive image-to-emotional-text translation (FITE) method, which focuses on capturing visual sentiment cues through facial expressions and selectively matching and fusing with the target aspect in textual modality. To the best of our knowledge, we are the first that explicitly utilize the emotional information from images in the multimodal aspect-based sentiment analysis task. Experiment results show that our method achieves state-of-the-art results on the Twitter-2015 and Twitter-2017 datasets. The improvement demonstrates the superiority of our model in capturing aspect-level sentiment in multimodal data with facial expressions^1.

Added

2026-10-03