keyword
facial expression
A facial expression is the visible movement or configuration of the facial muscles beneath the skin that communicates an individual emotional state, intention, or social signal. These muscular adjustments alter the appearance of key features such as the eyebrows, eyes, cheeks, and mouth to reflect discrete emotion categories, including happiness, sadness, anger, fear, surprise, and disgust, as well as continuous affective dimensions such as valence and arousal. Operating as a fundamental form of nonverbal communication, facial expressions can occur voluntarily or involuntarily, serving as critical visual cues in interpersonal interaction, psychological analysis, and computer vision applications focused on automated emotion recognition and facial generation.
3 items

CelebV-Text: A Large-Scale Facial Text-Video Dataset
Jianhui Yu, Hao Zhu, Liming Jiang, Chen Change Loy, Weidong Cai, Wayne Wu
Why you should read this
Presents a large-scale dataset of 70,000 in-the-wild facial video clips paired with 1.4 million detailed static and dynamic text descriptions to advance and standardize face-centric text-to-video generation.
Text-driven generation models are flourishing in video generation and editing. However, face-centric text-to-video generation remains a challenge due to the lack of a suitable dataset containing high-quality videos and highly relevant texts. This paper presents CelebV-Text, a large-scale, diverse, and high-quality dataset of facial text-video pairs, to facilitate research on facial text-to-video generation tasks. CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy. The provided texts are of high quality, describing both static and dynamic attributes precisely. The superiority of CelebV-Text over other datasets is demonstrated via comprehensive statistical analysis of the videos, texts, and text-video relevance. The effectiveness and potential of CelebV-Text are further shown through extensive self-evaluation. A benchmark is constructed with representative methods to standardize the evaluation of the facial text-to-video generation task. All data and models are publicly available^1.
Added
2026-09-26

The CMU Pose, Illumination, and Expression Database
Terence Sim, Simon Baker, Maan Bsat
Why you should read this
Presents the CMU Pose, Illumination, and Expression database, providing an extensive benchmark of over 40,000 systematically varied facial images across 68 subjects to advance face recognition and 3D modeling research.
In the Fall of 2000 we collected a database of over 40,000 facial images of 68 people. Using the CMU 3D Room we imaged each person across 13 different poses, under 43 different illumination conditions, and with 4 different expressions. We call this the CMU Pose, Illumination, and Expression (PIE) database. We describe the imaging hardware, the collection procedure, the organization of the images, several possible uses, and how to obtain the database.
Added
2026-09-18

AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild
Ali Mollahosseini, Behzad Hasani, Mohammad H. Mahoor
Why you should read this
Presents AffectNet, a large-scale in-the-wild facial expression database annotated for both categorical emotions and continuous valence-arousal dimensions, establishing deep learning baselines to advance automated affective computing.
Automated affective computing in the wild setting is a challenging problem in computer vision. Existing annotated databases of facial expressions in the wild are small and mostly cover discrete emotions (aka the categorical model). There are very limited annotated facial databases for affective computing in the continuous dimensional model (e.g., valence and arousal). To meet this need, we collected, annotated, and prepared for public distribution a new database of facial emotions in the wild (called AffectNet). AffectNet contains more than 1,000,000 facial images from the Internet by querying three major search engines using 1250 emotion related keywords in six different languages. About half of the retrieved images were manually annotated for the presence of seven discrete facial expressions and the intensity of valence and arousal. AffectNet is by far the largest database of facial expression, valence, and arousal in the wild enabling research in automated facial expression recognition in two different emotion models. Two baseline deep neural networks are used to classify images in the categorical model and predict the intensity of valence and arousal. Various evaluation metrics show that our deep neural network baselines can perform better than conventional machine learning methods and off-the-shelf facial expression recognition systems.
Added
2026-09-16
