FERV39k is a large-scale, multi-scene dataset designed for dynamic facial expression recognition (DFER) in the wild. It comprises 38,935 video clips with an average duration of 1.5 seconds (durations range from 0.5 to 4.0 seconds), corresponding to approximately 1 million video frames and cropped face images. The dataset provides cropped face images at 224×224 resolution and full scene context frames at 336×504 resolution.
The dataset is structured into a two-level hierarchy of 4 isolated scenarios subdivided into 22 fine-grained scenes:
- Daily Life (DL11k): Composed of 6 scenes: Argue, Social, School, Medicine, Conflict, and Daily-Life.
- Weak-Interactive Shows (WIS9k): Composed of 6 scenes: Action, Scholar-Reports, Speech, Elegant-Art, Live-Show, and Talk-Show.
- Strong-Interactive Activities (SIA10k): Composed of 6 scenes: Business, Experiment, Official-Event, Crime, Interview, and Contest.
- Anomaly Issues (AI9k): Composed of 4 scenes: History, Terror, War, and Crisis.
The video clips are annotated with 7 basic emotion classes: Angry, Disgust, Fear, Happy, Sad, Surprise, and Neutral. To resolve semantic ambiguity during labeling, annotations are collected across 26 fine-grained emotion words mapped onto the 7 basic classes (for example, Angry maps to Furious, Wrath, Outraged, and Sore; Happy maps to Pride, Cheerful, and Thrill).