Recognizing Action Units for Facial Expression Analysis
Ying-li TianT. KanadeJeffrey F. Cohn
Develops an automated facial expression analysis system that accurately recognizes subtle and combined Facial Action Coding System action units by combining multi-state tracking of permanent features and transient wrinkles with neural network classification.
Most automated facial expression systems focus on classifying a small set of basic emotions, such as anger or happiness. In daily life, however, these exaggerated expressions occur rarely. Human communication relies far more on subtle, fine-grained movements in isolated facial features, such as raised eyebrows or tightened lips. Manual identification of these muscle movements using the Facial Action Coding System requires extensive human training and labor. Automating this process has proven difficult because over 7,000 combinations of facial actions can occur, many of which alter the visual appearance of individual features when combined.
The article evaluated a computer vision system designed to automatically track facial features and accurately classify discrete facial action units in video sequences. It demonstrated the capability to identify these movements whether they occur alone or in complex combinations across both the upper and lower face.
The approach uses multi-state geometric models to track permanent features, such as the eyes, brows, and lips, alongside transient features like forehead wrinkles and cheek furrows in near-frontal video sequences. The system calculates 15 normalized parameters for the upper face and 9 parameters for the lower face relative to a neutral baseline frame. These parameters feed into artificial neural networks to classify 7 upper face and 11 lower face action units. The framework was evaluated across two datasets: a 24-subject dataset comprising 236 video sequences for the upper face, and a diverse 122-subject dataset comprising 463 sequences for the lower face.
The key findings demonstrate high classification performance across facial regions. The system achieved a 95% recognition accuracy on upper face action units and combinations, with a false alarm rate of 6.4%. For lower face movements, the system achieved a 96.71% recognition accuracy when modeling non-additive combinations and 96.3% when modeling basic units alone. Testing on novel subjects showed a 92.9% recognition rate with zero false alarms on single movements, confirming that the normalized geometric ratios generalize well to unseen individuals. Furthermore, separately modeling complex, non-additive upper face combinations provided no accuracy gain over standard models, whereas doing so for lower face movements yielded a slight performance improvement.
These results show that geometric feature tracking and normalized parameter ratios provide an effective, computationally efficient alternative to full-image holistic analysis or dense optical flow methods. Rather than requiring distinct models for thousands of possible action unit combinations, a unified neural network can reliably interpret isolated and combined expressions. This capability enables objective, high-throughput facial analysis that matches the reliability of certified human coders, reducing labor costs and measurement subjectivity in behavioral research, psychological assessment, and human-computer interaction.
Moving forward, developers should prioritize gathering larger training datasets for rare, non-additive combinations, such as combined nose wrinkling and chin raising, which caused the majority of classification errors. The primary recommendation is to integrate jaw and mandible tracking into lower face models to better distinguish between open-lip variations, as well as to expand the system's multi-state framework to handle significant out-of-plane head rotations.
Readers should interpret the findings within the article's operational boundary conditions. The current system relies on an initial neutral-expression frame for baseline comparison and assumes near-frontal head views with only minor rotations. While confidence in the reported accuracies is high for controlled, frontal video sequences across diverse adult demographics, performance in unconstrained settings with dynamic head poses and spontaneous baseline expressions remains subject to further validation.
- Paper: Automatic Analysis of Facial Expressions: The State of the Art, M. Pantic et al. (2000). Read this survey first to see the late-1990s expression-analysis systems and limitations that frame the source’s move from basic emotion labels to action-unit recognition.
- Paper: OpenFace: An open source facial behavior analysis toolkit, Tadas Baltrusaitis et al. (2016). OpenFace carries automated action-unit analysis forward into an integrated, real-time toolkit, adding landmark tracking, pose estimation, and gaze analysis to the source’s recognition focus.
