keyword
video-based emotion recognition
Video-based emotion recognition is an affective computing and computer vision task that automatically identifies and classifies human emotional states from dynamic video recordings. Unlike static image analysis, which evaluates isolated snapshots, video-based approaches analyze continuous image sequences to capture the temporal evolution, onset, peak, and offset of facial expressions and related behavioral cues. Computational models in this domain leverage spatial and temporal feature extraction techniques, such as multidimensional convolutional networks, recurrent architectures, and attention mechanisms, to track muscle movements and changes in expression over time. By incorporating both frame-level visual information and time-dependent dynamics, video-based emotion recognition aims to deliver accurate, context-sensitive interpretations of emotional behavior across changing conditions and naturalistic settings.
2 items

Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition
Hanyang Wang, Bo Li, Shuang Wu, Siyuan Shen, Feng Liu, Shouhong Ding, Aimin Zhou
Why you should read this
Proposes a multi-instance learning framework for dynamic facial expression recognition that treats non-target video frames as weakly supervised data and balances short- and long-term temporal dependencies to achieve state-of-the-art accuracy using a standard 3D CNN backbone.
Dynamic Facial Expression Recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video format. Previous research has considered non-target frames as noisy frames, but we propose that it should be treated as a weakly supervised problem. We also identify the imbalance of short- and long-term temporal relationships in DFER. Therefore, we introduce the Multi-3D Dynamic Facial Expression Learning (M3DFEL) framework, which utilizes Multi-Instance Learning (MIL) to handle inexact labels. M3DFEL generates 3D-instances to model the strong short-term temporal relationship and utilizes 3DCNNs for feature extraction. The Dynamic Long-term Instance Aggregation Module (DLIAM) is then utilized to learn the long-term temporal relationships and dynamically aggregate the instances. Our experiments on DFEW and FERV39K datasets show that M3DFEL outperforms existing state-of-the-art approaches with a vanilla R3D18 backbone. The source code is available at https://github.com/faceeyes/M3DFEL.
Added
2026-09-26

Deep Facial Expression Recognition: A Survey
Shan Li, Weihong Deng
Why you should read this
Presents a systematic guide to deep facial expression recognition, categorizing architectures for static and dynamic data while outlining practical strategies to handle real-world obstacles such as identity bias, head pose, and illumination variations.
With the transition of facial expression recognition (FER) from laboratory-controlled to challenging in-the-wild conditions and the recent success of deep learning techniques in various fields, deep neural networks have increasingly been leveraged to learn discriminative representations for automatic FER. Recent deep FER systems generally focus on two important issues: overfitting caused by a lack of sufficient training data and expression-unrelated variations, such as illumination, head pose and identity bias. In this paper, we provide a comprehensive survey on deep FER, including datasets and algorithms that provide insights into these intrinsic problems. First, we describe the standard pipeline of a deep FER system with the related background knowledge and suggestions of applicable implementations for each stage. We then introduce the available datasets that are widely used in the literature and provide accepted data selection and evaluation principles for these datasets. For the state of the art in deep FER, we review existing novel deep neural networks and related training strategies that are designed for FER based on both static images and dynamic image sequences, and discuss their advantages and limitations. Competitive performances on widely used benchmarks are also summarized in this section. We then extend our survey to additional related issues and application scenarios. Finally, we review the remaining challenges and corresponding opportunities in this field as well as future directions for the design of robust deep FER systems.
Added
2026-09-24
