FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos

Yan WangYixuan SunYiwen HuangZhongying LiuShuyong GaoWei ZhangWeifeng GeWenqiang Zhang

article2022CVPR138 citations

Presents FERV39k, a large-scale video benchmark of nearly 39,000 clips spanning 22 distinct scenes to advance dynamic facial expression recognition across diverse real-world contexts.

Abstract

Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the “Happy” expression with high intensity in Talk-Show is more discriminating than the same expression with low intensity in Official-Event. To fill this gap, we build a large-scale multi-scene dataset, coined as FERV39k. We analyze the important ingredients of constructing such a novel dataset in three aspects: (1) multi-scene hierarchy and expression class, (2) generation of candidate video clips, (3) trusted manual labelling process. Based on these guidelines, we select 4 scenarios subdivided into 22 scenes, annotate 86k samples automatically obtained from 4k videos based on the well-designed workflow, and finally build 38,935 video clips labeled with 7 classic expressions. Experiment benchmarks on four kinds of baseline frameworks were also provided and further analysis on their performance across different scenes and some challenges for future research were given. Besides, we systematically investigate key components of DFER by ablation studies.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Video-based Datasets for DFER
  • 2.2. Dynamic FER Approaches
  • 3. The FERV39k Dataset
  • 3.1. Key Challenges
  • 3.2. Dataset Construction Procedure
  • 3.3. Dataset Statistics
  • 3.4. Dataset Characteristics
  • 4. Benchmark Performance
  • 4.1. Experiment Setup
  • 4.2. Baseline Network
  • 4.3. Baseline Evaluation
  • 4.4. Ablation Studies
  • 5. Conclusion
  • 6. Acknowledgments
  • References

Knowls

  1. Knowl 1 — FERV39k Multi-Scene Dynamic Facial Expression Dataset Structure and Taxonomy

    definition

    FERV39k is a large-scale, multi-scene dataset designed for dynamic facial expression recognition (DFER) in the wild. It comprises 38,935 video clips with an average duration of 1.5 seconds (durations range from 0.5 to 4.0 seconds), corresponding to approximately 1 million video frames and cropped face images. The dataset provides cropped face images at 224×224224 \times 224 resolution and full scene context frames at 336×504336 \times 504 resolution.

    The dataset is structured into a two-level hierarchy of 4 isolated scenarios subdivided into 22 fine-grained scenes:

    1. Daily Life (DL11k): Composed of 6 scenes: Argue, Social, School, Medicine, Conflict, and Daily-Life.
    2. Weak-Interactive Shows (WIS9k): Composed of 6 scenes: Action, Scholar-Reports, Speech, Elegant-Art, Live-Show, and Talk-Show.
    3. Strong-Interactive Activities (SIA10k): Composed of 6 scenes: Business, Experiment, Official-Event, Crime, Interview, and Contest.
    4. Anomaly Issues (AI9k): Composed of 4 scenes: History, Terror, War, and Crisis.

    The video clips are annotated with 7 basic emotion classes: Angry, Disgust, Fear, Happy, Sad, Surprise, and Neutral. To resolve semantic ambiguity during labeling, annotations are collected across 26 fine-grained emotion words mapped onto the 7 basic classes (for example, Angry maps to Furious, Wrath, Outraged, and Sore; Happy maps to Pride, Cheerful, and Thrill).

  2. Knowl 2 — Four-Stage Automated Candidate Video Clip Generation Pipeline

    model/method

    To build a candidate pool of short, single-expression video clips from diverse in-the-wild video sources at low manual cost, a four-stage video collection and segmentation pipeline is used:

    1. Keyword-Based Metadata Crawling: Over 6,000 video metadata records spanning Asian, African, and European/American platforms are downloaded from 8 worldwide open-source video and search engines using scene-specific keyword lists.
    2. Temporal Balancing and Segmentation: Raw videos are sorted, pruned, and balanced across scenes down to 4,000 video records, which are then segmented into candidate clips of durations between 0.5 and 4.0 seconds.
    3. Heuristic Rule Selection: An automated rule list filters the segments to select a candidate set containing approximately twenty times the target dataset size.
    4. FER Detector Refinement and Distribution Alignment: A pre-trained lightweight ResNet-50 facial expression recognition detector evaluates the candidate clips to generate preliminary expression predictions. The candidate pool is pruned to approximately twice the target dataset size (~86,000 candidate video clips), sub-sampled to ensure the latency and frequency distributions of the predicted expressions match natural real-world distributions.
  3. Knowl 3 — Two-Stage Quality-Controlled Crowd-Professional Annotation Workflow with WWTA Voting

    model/method

    To ensure annotation accuracy at an affordable cost, dataset annotation combines crowdsourcing annotators (CAs, 20 workers) and professional researchers (PRs, 10 workers) via a custom web platform providing face bounding boxes and 26 fine-grained emotion choices alongside a PASS option.

    The workflow operates as follows:

    1. Grouping and PR Ground-Truth Seeding: Video clips are partitioned into batches where 5% of clips within each batch are pre-annotated by PRs as hidden verification samples. Each batch is replicated and independently distributed to 3 distinct CAs.
    2. Flag-Recaptured Statistical Screening: CA annotations are evaluated against the 5% PR ground-truth samples using the Flag-Recaptured Statistic method:
      • Accuracy ≥80%\ge 80\%: Marked as Accept (AC).
      • 40%≤Accuracy<80%40\% \le \text{Accuracy} < 80\%: Marked as Improper (IP).
      • Accuracy <40%< 40\%: Marked as Unacceptable (UA).
    3. Feedback and Professional Adjudication: Batches marked as UA are rejected and returned to CAs with corrective feedback. Batches marked as IP and AC proceed to PRs for review; persistent IP ambiguities are manually re-annotated by PRs.
    4. Weighted-Winner-Take-All (WWTA) Label Generation: Final basic expression labels are computed by applying WWTA voting to the CA and PR selections and automatically converting the 26 fine-grained emotion categories into the 7 primary expression classes.
  4. Knowl 4 — DFER Benchmark Protocol and Baseline Architectures

    experimental setup

    The FERV39k benchmark partitions the dataset clips into an 80% training set (including validation) and a 20% testing set without overlap across clips. The evaluation protocol includes 27 evaluation configurations: 22 per-scene setups, 4 isolated scenario setups, and 1 full-dataset setup. Performance is measured using Weighted Average Recall (WAR\text{WAR}, overall classification accuracy) and Unweighted Average Recall (UAR\text{UAR}, average recall across the 7 expression classes).

    Four categories of deep baseline architectures are evaluated:

    1. 2D ConvNets: Single-frame backbones (ResNet-18, ResNet-50, VGG-13, VGG-16) extract frame embeddings across all clip frames; embeddings are concatenated and fed to a linear classifier.
    2. 2D ConvNet-LSTM: A recurrent LSTM layer with 1024 hidden units and batch normalization is appended after the global average pooling layer of the 2D ConvNet backbone, followed by a fully connected classifier.
    3. 3D ConvNets: Spatio-temporal 3D convolutional networks (C3D, I3D, 3D-ResNet-18) process clip volumes directly using 3D spatio-temporal filters.
    4. Two-Stream Networks: Dual-branch networks process cropped facial image sequences in one branch and whole scene context frames in a parallel branch, fusing the representations before final classification (e.g., Two-Stream C3D, Two-Stream I3D, Two-Stream 3D-ResNet-18, Two-Stream R18-LSTM, Two-Stream VGG13-LSTM).

    Models are trained from scratch for 60 epochs using stochastic gradient descent (SGD) with momentum 0.9, batch size 32, weight decay 1×10−41\times 10^{-4}, uniform frame sampling interval of 8, and initial learning rate in [1×10−3,1×10−2][1\times 10^{-3}, 1\times 10^{-2}] decayed by a factor of 0.95 per epoch. Cropped face images are resized to 112×112112 \times 112 and full scene context frames to 112×168112 \times 168 with random cropping, illumination changes, and horizontal flipping.

  5. Knowl 5 — Benchmark Performance of Deep Architectures on FERV39k

    data/table

    The table presents Weighted Average Recall and Unweighted Average Recall (WAR/UAR\text{WAR}/\text{UAR}, reported in %) for 15 baseline architectures trained from scratch on FERV39k across the overall dataset, 4 scenarios, and 9 representative scenes.

    Method All DL11k WIS9k SIA10k AI9k
    R18 39.33/30.30 39.75/31.36 40.50/28.67 42.31/30.02 33.90/27.20
    R50 30.57/22.47 30.46/21.52 32.52/23.50 30.56/22.68 30.14/19.94
    VGG13 41.02/31.19 40.40/31.59 43.04/30.23 43.44/29.99 38.86/29.94
    VGG16 41.66/32.01 41.81/32.59 42.93/30.77 42.31/29.58 39.60/31.46
    R18-LSTM 42.59/30.92 43.34/32.24 44.12/29.59 42.85/28.78 39.66/30.40
    R50-LSTM 40.75/32.12 40.93/32.91 41.74/30.70 42.16/30.39 38.01/31.16
    VGG13-LSTM 43.37/32.41 42.29/32.46 44.23/30.81 45.00/31.45 41.20/31.49
    VGG16-LSTM 41.70/30.93 42.99/32.32 41.63/28.42 43.83/29.83 37.04/29.39
    C3D 31.69/22.68 26.95/21.02 30.15/19.94 42.70/29.22 27.29/19.80
    I3D 38.78/30.17 38.56/29.25 38.52/29.11 40.55/31.07 37.44/28.15
    3D-R18 37.57/26.67 37.69/27.47 38.40/24.85 40.40/26.08 33.45/25.40
    Two C3D 41.77/30.72 41.45/31.37 43.44/29.77 44.71/30.15 37.89/28.09
    Two I3D 41.30/31.01 41.02/31.55 42.31/30.14 43.63/31.20 38.75/28.53
    Two 3D-R18 42.28/30.55 42.77/32.72 44.12/29.63 42.95/27.83 38.46/28.54
    Two R18-LSTM 43.20/31.28 42.20/31.66 44.91/30.37 46.33/31.09 40.40/30.04
    Two VGG13-LSTM 44.54/32.79 44.65/32.96 45.25/31.45 46.57/31.88 40.63/30.96
    Average 39.58/29.34 39.27/29.80 40.61/28.11 42.04/28.94 36.55/27.61

    Two-stream 2D ConvNet-LSTM methods outperform other categories, with Two-Stream VGG13-LSTM achieving the highest overall WAR/UAR\text{WAR}/\text{UAR} of 44.54%/32.79%44.54\%/32.79\%. Models consistently perform best on Strong-Interactive Activities (SIA10k, average 42.04%42.04\% WAR; Experiment scene reaching 63.72%63.72\% WAR for Two C3D) and worst on Anomaly Issues (AI9k, average 36.55%36.55\% WAR; Terror scene averaging 33.75%33.75\% WAR), due to differences in expression intensity, duration, and contextual clarity.

  6. Knowl 6 — Cross-Scenario Transfer Generalization Discrepancy

    data/table

    Cross-scenario generalization of dynamic facial expression models is evaluated by training a ResNet50-LSTM model on one source scenario and testing on all 4 isolated target scenarios. The matrix below reports performance in WAR (%) / UAR (%)\text{WAR}\,(\%)\,/\,\text{UAR}\,(\%):

    Source Target DL11k Target WIS9k Target SIA10k Target AI9k
    DL11k 37.69 / 27.21 29.98 / 19.93 31.15 / 21.87 24.27 / 18.54
    WIS9k 27.04 / 19.95 40.50 / 26.60 31.78 / 19.90 24.62 / 19.24
    SIA10k 28.57 / 21.92 31.39 / 19.95 39.72 / 24.90 27.75 / 20.28
    AI9k 26.29 / 20.21 23.30 / 18.29 23.85 / 17.93 31.62 / 24.16

    Cross-domain evaluation demonstrates an average accuracy decline of nearly 8% compared to within-scenario training and testing. Transferring models trained on weak-interactive scenarios (WIS9k) to stronger-interaction or anomaly scenarios yields substantial performance drops, illustrating significant distribution shifts in facial dynamics and expression manifestations across different social contexts.

  7. Knowl 7 — Ablation Analysis of Pre-training, Frame Sampling, and Contextual Stream Fusion

    empirical result

    Ablation experiments on key architectural components and training protocols reveal three empirical properties of dynamic facial expression recognition (DFER):

    1. Pre-training Transferability: Pre-training ResNet18-LSTM and ResNet50-LSTM on large-scale face/FER datasets (MS-Celeb-1M or DFEW) does not consistently outperform training from scratch on FERV39k. This occurs because the feature distribution, contextual interactions, and multi-scene diversity in FERV39k differ substantially from existing datasets.
    2. Frame Sampling Capacity: Scaling the number of uniformly sampled input frames from 2 to 16 in steps of 2 across 2D ConvNet-LSTM models shows that accuracy does not steadily increase with frame count. Beyond a modest frame budget, performance plateaus or slightly decreases, indicating that naive sparse sampling is insufficient and dynamic keyframe selection is required.
    3. Context Stream Contribution: Comparing single-stream I3D (face-only) to two-stream I3D (face + scene context) shows that incorporating scene context improves accuracy across representative scenes from all scenarios (e.g., Medicine, Argue, TalkShow, Action, Business, Experiment, Terror, History), demonstrating that environmental and interaction context provides complementary signals for resolving ambiguous facial expressions.
  8. Knowl 8 — Inherent Challenges in Multi-Scene Video Facial Expression Recognition

    limitation

    Existing video FER architectures encounter performance degradation on FERV39k due to four core challenges:

    1. Sparse Expression Frames: In dynamic clips, emotionally relevant facial muscle activations often occur in only a few frames while the rest are neutral or transitioning, causing standard temporal aggregation methods to dilute key cues.
    2. Context-Dependent Expression Intensity: Identical facial expressions exhibit disparate intensity and feature distributions across scenes (for example, Happy displays high activation in Talk-Show but subtle activation in Official-Event).
    3. Complex Spatio-Temporal Variations: In-the-wild facial clips contain unconstrained head rotations, varying motion speeds, camera movement, and duration variations (ranging from 0.5s to 4.0s) that challenge fixed-kernel 3D convolutions and recurrent models.
    4. Severe Scene-Level Long-Tailed Imbalance: Expression occurrences follow strong long-tailed distributions conditioned on the scene (for example, Fear represents 18% of clips in Terror but is rare in other scenes; Happy represents 33% of Live-Show clips), inducing substantial model prediction biases.

Coverage note — None was omitted; all primary contributions—including the FERV39k dataset taxonomy, the four-stage clip generation and two-stage annotation workflows, baseline benchmarks, cross-scenario experiments, ablation studies, and dataset challenge analyses—are fully represented.

References

  1. 1.Dawood Adel Al Chanti and Alice Caplier. Deep learning for spatio-temporal modeling of dynamic spontaneous emotions. IEEE Transactions on Affective Computing, 2018. 3
  2. 2.Amal Azazi, Syaheerah Lebai Lutfi, Ibrahim Venkat, and Fernando Fern'andez-Mart'ınez. Towards a robust affect recognition: Automatic facial expression recognition in 3d faces. Expert Systems with Applications, 42(6):3056–3066, 2015. 1
  3. 3.Asit Barman and Paramartha Dutta. Facial expression recognition using distance and texture signature relevant features. Applied Soft Computing, 77:88–105, 2019. 1
  4. 4.Graham Bell. Population estimates from recapture studies in which no recaptures have been made. Nature, 248(5449):616–616, 1974. 4
  5. 5.Xianye Ben, Yi Ren, Junping Zhang, Su-Jing Wang, Kidiyo Kpalma, Weixiao Meng, and Yong-Jin Liu. Video-based facial micro-expression analysis: A survey of datasets, features and algorithms. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021. 4
  6. 6.Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. 8
  7. 7.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017. 2, 6, 7
  8. 8.Chun-Fu Richard Chen, Rameswar Panda, Kandan Ramakrishnan, Rogerio Feris, John Cohn, Aude Oliva, and Quanfu Fan. Deep analysis of cnn-based spatio-temporal representations for action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6165–6175, 2021. 6
  9. 9.Alan S Cowen, Dacher Keltner, Florian Schroff, Brendan Jou, Hartwig Adam, and Gautam Prasad. Sixteen facial expressions occur in similar contexts worldwide. Nature, 589(7841):251–257, 2021. 2, 3, 4
  10. 10.Arnaud Dapogny, Kevin Bailly, and S'everine Dubuisson. Dynamic facial expression recognition by joint static and multi-time gap transition classification. In 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), volume 1, pages 1–6. IEEE, 2015. 6
  11. 11.Abhinav Dhall, Amanjot Kaur, Roland Goecke, and Tom Gedeon. Emotiw 2018: Audio-video, student engagement and group-level affect prediction. In Proceedings of the 20th ACM International Conference on Multimodal Interaction, pages 653–656, 2018. 2, 3, 4
  12. 12.Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. Long-term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2625–2634, 2015. 6
  13. 13.Yin Fan, Xiangju Lu, Dian Li, and Yuanliu Liu. Video-based emotion recognition using cnn-rnn and c3d hybrid networks. In Proceedings of the 18th ACM international conference on multimodal interaction, pages 445–450, 2016. 3
  14. 14.Amir Hossein Farzaneh and Xiaojun Qi. Facial expression recognition in the wild via deep attentive center loss. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2402–2411, 2021. 1
  15. 15.Yongjian Fu, Xintian Wu, Xi Li, Zhijie Pan, and Daxin Luo. Semantic neighborhood-aware deep facial expression recognition. IEEE Transactions on Image Processing, 29:6535–6548, 2020. 1
  16. 16.Amir Ghodrati, Babak Ehteshami Bejnordi, and Amirhossein Habibian. Frameexit: Conditional early exiting for efficient video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15608–15618, 2021. 8
  17. 17.Shreyank N Gowda, Marcus Rohrbach, and Laura Sevilla-Lara. Smart frame selection for action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1451–1459, 2021. 8
  18. 18.Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao. Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. In European conference on computer vision, pages 87–102. Springer, 2016. 8
  19. 19.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 6
  20. 20.Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 3d convolutional neural networks for human action recognition. IEEE transactions on pattern analysis and machine intelligence, 35(1):221–231, 2012. 6
  21. 21.Xingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang, Wanchuang Xia, Cheng Lu, and Jiateng Liu. Dfew: A large-scale database for recognizing dynamic facial expressions in the wild. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2881–2889, 2020. 2, 3, 4, 6, 8
  22. 22.Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. 2, 6
  23. 23.Dimitrios Kollias, Panagiotis Tzirakis, Mihalis A Nicolaou, Athanasios Papaioannou, Guoying Zhao, Bj"orn Schuller, Irene Kotsia, and Stefanos Zafeiriou. Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond. International Journal of Computer Vision, 127(6):907–929, 2019. 2, 3, 4
  24. 24.Jean Kossaifi, Georgios Tzimiropoulos, Sinisa Todorovic, and Maja Pantic. Afew-va database for valence and arousal estimation in-the-wild. Image and Vision Computing, 65:23–36, 2017. 2, 3
  25. 25.Jiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park, and Kwanghoon Sohn. Context-aware emotion recognition networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10143–10152, 2019. 2, 3, 6
  26. 26.Shan Li and Weihong Deng. Deep facial expression recognition: A survey. IEEE transactions on affective computing, 2020. 1, 3, 5
  27. 27.Shan Li, Weihong Deng, and JunPing Du. Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2852–2861, 2017. 1, 4
  28. 28.Liqian Liang, Congyan Lang, Yidong Li, Songhe Feng, and Jian Zhao. Fine-grained facial expression recognition in the wild. IEEE Transactions on Information Forensics and Security, 16:482–494, 2020. 4
  29. 29.Daizong Liu, Xi Ouyang, Shuangjie Xu, Pan Zhou, Kun He, and Shiping Wen. Saanet: Siamese action-units attention network for improving dynamic facial expression recognition. Neurocomputing, 413:145–157, 2020. 1, 3
  30. 30.Xin Liu, Silvia L Pintea, Fatemeh Karimi Nejadasl, Olaf Booij, and Jan C van Gemert. No frame left behind: Full video action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14892–14901, 2021. 2
  31. 31.Ling Lo, Hong-Xia Xie, Hong-Han Shuai, and Wen-Huang Cheng. Mer-gcn: Micro-expression recognition based on relation modeling with graph convolutional networks. In 2020 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), pages 79–84. IEEE, 2020. 3
  32. 32.Patrick Lucey, Jeffrey F Cohn, Takeo Kanade, Jason Saragih, Zara Ambadar, and Iain Matthews. The extended cohnkanade dataset (ck+): A complete dataset for action unit and emotion-specified expression. In 2010 ieee computer society conference on computer vision and pattern recognition-workshops, pages 94–101. IEEE, 2010. 1, 2, 3
  33. 33.Jiaxin Ma, Hao Tang, Wei-Long Zheng, and Bao-Liang Lu. Emotion recognition using multimodal residual lstm network. In Proceedings of the 27th ACM international conference on multimedia, pages 176–183, 2019. 3
  34. 34.Ali Mollahosseini, Behzad Hasani, and Mohammad H Mahoor. Affectnet: A database for facial expression, valence, and arousal computing in the wild. IEEE Transactions on Affective Computing, 10(1):18–31, 2017. 1, 3, 4, 6
  35. 35.Naima Otberdout, Mohammed Daoudi, Anis Kacem, Lahoucine Ballihi, and Stefano Berretti. Dynamic facial expression generation on hilbert hypersphere with conditional wasserstein generative adversarial nets. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020. 3
  36. 36.Waseem Rawat and Zenghui Wang. Deep convolutional neural networks for image classification: A comprehensive review. Neural computation, 29(9):2352–2449, 2017. 6
  37. 37.Bj"orn Schuller, Bogdan Vlasenko, Florian Eyben, Martin W"ollmer, Andre Stuhlsatz, Andreas Wendemuth, and Gerhard Rigoll. Cross-corpus acoustic emotion recognition: Variances and strategies. IEEE Transactions on Affective Computing, 1(2):119–131, 2010. 6
  38. 38.Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. Finegym: A hierarchical video dataset for fine-grained action understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2616–2625, 2020. 3, 5, 8
  39. 39.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 6
  40. 40.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 4489–4497, 2015. 3, 6, 7
  41. 41.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 6450–6459, 2018. 7
  42. 42.Ekaterina P Volkova, Betty J Mohler, Trevor J Dodds, Joachim Tesch, and Heinrich H B"ulthoff. Emotion categorization of body expressions in narrative scenarios. Frontiers in psychology, 5:623, 2014. 4
  43. 43.Kai Wang, Xiaojiang Peng, Jianfei Yang, Debin Meng, and Yu Qiao. Region attention networks for pose and occlusion robust facial expression recognition. IEEE Transactions on Image Processing, 29:4057–4069, 2020. 1
  44. 44.Yan Wang, Wei Song, Wei Tao, Antonio Liotta, Dawei Yang, Xinlei Li, Shuyong Gao, Yixuan Sun, Weifeng Ge, Wei Zhang, and Wenqiang Zhang. A systematic review on affective computing: Emotion models, databases, and recent advances. arXiv preprint arXiv:2203.06935, 2022. 1
  45. 45.Zhenbo Yu, Guangcan Liu, Qingshan Liu, and Jiankang Deng. Spatio-temporal convolutional features with nested lstm for facial expression recognition. Neurocomputing, 317:50–57, 2018. 1
  46. 46.Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4694–4702, 2015. 6
  47. 47.Kaihao Zhang, Yongzhen Huang, Yong Du, and Liang Wang. Facial expression recognition based on deep evolutional spatial-temporal networks. IEEE Transactions on Image Processing, 26(9):4193–4203, 2017. 3
  48. 48.Guoying Zhao, Xiaohua Huang, Matti Taini, Stan Z Li, and Matti Pietik"ainen. Facial expression recognition from near-infrared videos. Image and Vision Computing, 29(9):607–619, 2011. 1, 2, 3
  49. 49.Xiangyun Zhao, Xiaodan Liang, Luoqi Liu, Teng Li, Yugang Han, Nuno Vasconcelos, and Shuicheng Yan. Peak-piloted deep network for facial expression recognition. In European conference on computer vision, pages 425–442. Springer, 2016. 1
  50. 50.Yin-Dong Zheng, Zhaoyang Liu, Tong Lu, and Limin Wang. Dynamic sampling networks for efficient action recognition in videos. IEEE Transactions on Image Processing, 29:7970–7983, 2020. 8

Citation

MLA
Wang, Y., et al. “FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos”. arXiv, 2022, http://arxiv.org/abs/2203.09463v2.
APA
Wang, Y., Sun, Y., Huang, Y., Liu, Z., Gao, S., Zhang, W., Ge, W., & Zhang, W. (2022). FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos. arXiv. http://arxiv.org/abs/2203.09463v2
Chicago
Wang, Y., Y. Sun, Y. Huang, et al. 2022. “FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos”. arXiv. http://arxiv.org/abs/2203.09463v2.
Harvard
Wang, Y. et al. (2022) “FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.09463v2.
Vancouver
1. Wang Y, Sun Y, Huang Y, Liu Z, Gao S, Zhang W, Ge W, Zhang W (2022) FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos. arXiv

BibTeX

@article{wang2022ferv39k,
  title = {FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos},
  author = {Wang, Yan and Sun, Yixuan and Huang, Yiwen and Liu, Zhongying and Gao, Shuyong and Zhang, Wei and Ge, Weifeng and Zhang, Wenqiang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.09463v2},
  eprint = {2203.09463}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE