Built independently by an author, for readers. Read the story and support ChapterPal

keyword

facial expression

A facial expression is the visible movement or configuration of the facial muscles beneath the skin that communicates an individual emotional state, intention, or social signal. These muscular adjustments alter the appearance of key features such as the eyebrows, eyes, cheeks, and mouth to reflect discrete emotion categories, including happiness, sadness, anger, fear, surprise, and disgust, as well as continuous affective dimensions such as valence and arousal. Operating as a fundamental form of nonverbal communication, facial expressions can occur voluntarily or involuntarily, serving as critical visual cues in interpersonal interaction, psychological analysis, and computer vision applications focused on automated emotion recognition and facial generation.

3 items

CelebV-Text: A Large-Scale Facial Text-Video Dataset

CelebV-Text: A Large-Scale Facial Text-Video Dataset

Jianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy, Weidong Cai, Wayne Wu

Why you should read this

Presents a large-scale dataset of 70,000 in-the-wild facial video clips paired with 1.4 million detailed static and dynamic text descriptions to advance and standardize face-centric text-to-video generation.

Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents CelebV-Text, a large-scale, diverse, and high-quality dataset of facial text-video pairs, to facilitate research on facial text-to-video generation tasks. CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy. The provided texts are of high quality, describing both static and dynamic attributes precisely. The superiority of CelebV-Text over other datasets is demonstrated via comprehensive statistical analysis of the videos, texts, and text-video relevance. The effectiveness and potential of CelebV-Text are further shown through extensive self-evaluation. A benchmark is constructed with representative methods to standardize the evaluation of the facial text-to-video generation task. All data and models are publicly available^1.

Added

2026-09-26

AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild

AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild

Ali Mollahosseini, Behzad Hasani, Mohammad H. Mahoor

OrganizationsUniversity of Denver

Why you should read this

Presents AffectNet, a large-scale in-the-wild facial expression database annotated for both categorical emotions and continuous valence-arousal dimensions, establishing deep learning baselines to advance automated affective computing.

Automated affective computing in the wild setting is a challenging problem in computer vision. Existing annotated databases of facial expressions in the wild are small and mostly cover discrete emotions (aka the categorical model). There are very limited annotated facial databases for affective computing in the continuous dimensional model (e.g., valence and arousal). To meet this need, we collected, annotated, and prepared for public distribution a new database of facial emotions in the wild (called AffectNet). AffectNet contains more than 1,000,000 facial images from the Internet by querying three major search engines using 1250 emotion related keywords in six different languages. About half of the retrieved images were manually annotated for the presence of seven discrete facial expressions and the intensity of valence and arousal. AffectNet is by far the largest database of facial expression, valence, and arousal in the wild enabling research in automated facial expression recognition in two different emotion models. Two baseline deep neural networks are used to classify images in the categorical model and predict the intensity of valence and arousal. Various evaluation metrics show that our deep neural network baselines can perform better than conventional machine learning methods and off-the-shelf facial expression recognition systems.

Added

2026-09-16