Recognizing Action Units for Facial Expression Analysis

Ying-li TianT. KanadeJeffrey F. Cohn

article2001TPAMI1,982 citations

Develops an automated facial expression analysis system that accurately recognizes subtle and combined Facial Action Coding System action units by combining multi-state tracking of permanent features and transient wrinkles with neural network classification.

Listen

Most automated facial expression systems focus on classifying a small set of basic emotions, such as anger or happiness. In daily life, however, these exaggerated expressions occur rarely. Human communication relies far more on subtle, fine-grained movements in isolated facial features, such as raised eyebrows or tightened lips. Manual identification of these muscle movements using the Facial Action Coding System requires extensive human training and labor. Automating this process has proven difficult because over 7,000 combinations of facial actions can occur, many of which alter the visual appearance of individual features when combined.

The article evaluated a computer vision system designed to automatically track facial features and accurately classify discrete facial action units in video sequences. It demonstrated the capability to identify these movements whether they occur alone or in complex combinations across both the upper and lower face.

The approach uses multi-state geometric models to track permanent features, such as the eyes, brows, and lips, alongside transient features like forehead wrinkles and cheek furrows in near-frontal video sequences. The system calculates 15 normalized parameters for the upper face and 9 parameters for the lower face relative to a neutral baseline frame. These parameters feed into artificial neural networks to classify 7 upper face and 11 lower face action units. The framework was evaluated across two datasets: a 24-subject dataset comprising 236 video sequences for the upper face, and a diverse 122-subject dataset comprising 463 sequences for the lower face.

The key findings demonstrate high classification performance across facial regions. The system achieved a 95% recognition accuracy on upper face action units and combinations, with a false alarm rate of 6.4%. For lower face movements, the system achieved a 96.71% recognition accuracy when modeling non-additive combinations and 96.3% when modeling basic units alone. Testing on novel subjects showed a 92.9% recognition rate with zero false alarms on single movements, confirming that the normalized geometric ratios generalize well to unseen individuals. Furthermore, separately modeling complex, non-additive upper face combinations provided no accuracy gain over standard models, whereas doing so for lower face movements yielded a slight performance improvement.

These results show that geometric feature tracking and normalized parameter ratios provide an effective, computationally efficient alternative to full-image holistic analysis or dense optical flow methods. Rather than requiring distinct models for thousands of possible action unit combinations, a unified neural network can reliably interpret isolated and combined expressions. This capability enables objective, high-throughput facial analysis that matches the reliability of certified human coders, reducing labor costs and measurement subjectivity in behavioral research, psychological assessment, and human-computer interaction.

Moving forward, developers should prioritize gathering larger training datasets for rare, non-additive combinations, such as combined nose wrinkling and chin raising, which caused the majority of classification errors. The primary recommendation is to integrate jaw and mandible tracking into lower face models to better distinguish between open-lip variations, as well as to expand the system's multi-state framework to handle significant out-of-plane head rotations.

Readers should interpret the findings within the article's operational boundary conditions. The current system relies on an initial neutral-expression frame for baseline comparison and assumes near-frontal head views with only minor rotations. While confidence in the reported accuracies is high for controlled, frontal video sequences across diverse adult demographics, performance in unconstrained settings with dynamic head poses and spontaneous baseline expressions remains subject to further validation.

  • Paper: OpenFace: An open source facial behavior analysis toolkit, Tadas Baltrusaitis et al. (2016). OpenFace carries automated action-unit analysis forward into an integrated, real-time toolkit, adding landmark tracking, pose estimation, and gaze analysis to the source’s recognition focus.
Cover for Recognizing Action Units for Facial Expression Analysis

Abstract

Most automatic expression analysis systems attempt to recognize a small set of prototypic expressions (e.g. happiness and anger). Such prototypic expressions, however, occur infrequently. Human emotions and intentions are communicated more often by changes in one or two discrete facial features. We develop an automatic system to analyze subtle changes in facial expressions based on both permanent facial features (brows, eyes, mouth) and transient facial features (deepening of facial furrows) in a nearly frontal image sequence. Unlike most existing systems, our system attempts to recognize fine-grained changes in facial expression based on Facial Action Coding System (FACS) action units (AUs), instead of six basic expressions (e.g. happiness and anger). Multi-state face and facial component models are proposed for tracking and modeling different facial features, including lips, eyes, brows, cheeks, and their related wrinkles and facial furrows. Then we convert the results of tracking to detailed parametric descriptions of the facial features. With these features as the inputs, 11 lower face action units (AUs) and 7 upper face AUs are recognized by a neural network algorithm. A recognition rate of 96.7% for lower face AUs and 95% for upper face AUs is obtained respectively. The recognition results indicate that our system can identify action units regardless of whether they occurred singly or in combinations.

Table of Contents

  • 1. Introduction
  • 2. Multi-State Models for Face and Facial Components
  • 2.1. Multi-state face model
  • 2.2. Multi-state face component models
  • 3. Facial Feature Extraction
  • 3.1. Permanent features
  • 3.2. Transient features
  • 4. Facial Feature Representation
  • 4.1. Upper Face Feature Representation
  • 4.2. Lower Face Feature Representation
  • 5. Facial Action Unit Definitions
  • 5.1. Upper Face Action Units
  • 5.2. Lower face action units
  • 6. Image Database
  • 6.1. Image Database for Upper Face AU Recognition
  • 6.2. Image Database for Lower Face AU Recognition
  • 7. Face Action Units Recognition
  • 7.1. Upper Face Action Units Recognition
  • 7.1.1 Upper Face Individual AU Recognition
  • 7.1.2 Upper Face AU Combination Recognition When Modeling 7 Individual AUs
  • 7.1.3 Upper Face AU Combination Recognition When Modeling Nonadditive Combinations
  • 7.2. Lower Face Action Units Recognition
  • 8. Conclusion and Discussion
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Multi-State Parametric Facial Feature Tracking

    model/method

    To track permanent facial features across varying expressions and head orientations in nearly frontal video sequences, multi-state geometric templates are employed:

    • Lips (Three-State Model): State is classified as open, closed, or tightly closed. For open and closed states, lip contours are modeled using two parabolic arcs parameterized by center position (xc,yc)(x_c, y_c), height parameters h1h_1 (upper lip height) and h2h_2 (lower lip height), lip width ww, and inclination angle θ\theta. For tightly closed lips, the dark boundary line connecting the two lip corners is detected directly. Initial lip color is modeled in the first neutral frame using a Gaussian mixture model, and feature points are tracked across subsequent frames.

    • Eyes (Two-State Model): State is classified as open or closed. For open eyes, the iris is modeled as a circle with radius rr and center (x0,y0)(x_0, y_0) detected via a half-circle edge mask, while eyelid contours are modeled by two parabolic arcs parameterized by (xc,yc,h1,h2,w,θ)(x_c, y_c, h_1, h_2, w, \theta). For closed eyes, the eye is parameterized by the coordinates of the two corners, (x1,y1)(x_1, y_1) and (x2,y2)(x_2, y_2), connected by a linear boundary segment without eyelid contour tracking.

    • Brows and Cheeks (Single-State Models): Brows and cheeks are each tracked using a three-point triangular template (x1,y1)(x_1, y_1), (x2,y2)(x_2, y_2), and (x3,y3)(x_3, y_3) propagated across frames using a gradient-based Lucas-Kanade feature tracker.

  2. Knowl 2 — Normalized Coordinate System and Parametric Upper Face Representation

    model/method

    To achieve invariance to facial scale and subject conformation, upper facial features are mapped into a canonical coordinate system where the xx-axis connects the left and right inner eye corners and the yy-axis is orthogonal. All continuous motion parameters are computed as ratio scores relative to measurements from the neutral initial frame (denoted with subscript 00):

    • Inner brow motion:

      rbinner=bi−bi0bi0r_{binner} = \frac{b_i - b_{i0}}{b_{i0}}

      (positive values indicate upward displacement).

    • Outer brow motion:

      rbouter=bo−bo0bo0r_{bouter} = \frac{b_o - b_{o0}}{b_{o0}}

      (positive values indicate upward displacement).

    • Eye height change:

      reheight=(h1+h2)−(h10+h20)h10+h20r_{eheight} = \frac{(h_1 + h_2) - (h_{10} + h_{20})}{h_{10} + h_{20}}

      where h1h_1 and h2h_2 are upper and lower eyelid heights.

    • Upper eyelid motion:

      rtop=h1−h10h10r_{top} = \frac{h_1 - h_{10}}{h_{10}}
    • Lower eyelid motion:

      rbtm=h2−h20h20r_{btm} = \frac{h_2 - h_{20}}{h_{20}}
    • Cheek motion:

      rcheek=c−c0c0r_{cheek} = \frac{c - c_0}{c_0}

      where cc is cheek elevation.

    • Inter-brow distance change:

      Dbrow=D−D0D0D_{brow} = \frac{D - D_0}{D_0}

      where DD is the horizontal distance between the inner brows.

    • Crows-feet wrinkles: Binary indicators Wleft,Wright∈{0,1}W_{left}, W_{right} \in \{0, 1\} indicating absence (00) or presence (11).

    These produce a 15-dimensional input vector consisting of bilateral pairs for inner brow, outer brow, eye height, top lid, bottom lid, cheek elevation, plus DbrowD_{brow}, WleftW_{left}, and WrightW_{right}.

  3. Knowl 3 — Normalized Parametric Lower Face Representation

    model/method

    Lower facial actions are described using a set of normalized parameters aligned with the inner eye corner baseline:

    • Lip height ratio:

      rheight=(h1+h2)−(h10+h20)h10+h20r_{height} = \frac{(h_1 + h_2) - (h_{10} + h_{20})}{h_{10} + h_{20}}

      where h1h_1 and h2h_2 denote upper and lower lip heights relative to lip center, and h10,h20h_{10}, h_{20} denote baseline values.

    • Lip width ratio:

      rwidth=w−w0w0r_{width} = \frac{w - w_0}{w_0}

      where ww is current mouth width.

    • Lip corner vertical displacements:

      rleft=Dleft−Dleft0Dleft0,rright=Dright−Dright0Dright0r_{left} = \frac{D_{left} - D_{left0}}{D_{left0}}, \quad r_{right} = \frac{D_{right} - D_{right0}}{D_{right0}}

      where DleftD_{left} and DrightD_{right} denote vertical distances from the left and right lip corners to the inner eye corner line.

    • Top and bottom lip vertical motion:

      rtop=Dtop−Dtop0Dtop0,rbtm=Dbtm−Dbtm0Dbtm0r_{top} = \frac{D_{top} - D_{top0}}{D_{top0}}, \quad r_{btm} = \frac{D_{btm} - D_{btm0}}{D_{btm0}}

      where DtopD_{top} and DbtmD_{btm} denote vertical distances from the upper and lower lip margins to the eye baseline.

    • Nose wrinkle state: Binary flag Snosew∈{0,1}S_{nosew} \in \{0, 1\} indicating presence (11) or absence (00).

    • Nasolabial furrow angles: Continuous angles AngleleftAngle_{left} and AnglerightAngle_{right} relative to the eye baseline. Due to high inter-subject anatomical variation, these angles are omitted from cross-subject classification networks, yielding a 7-dimensional cross-subject lower face feature vector.

  4. Knowl 4 — Detection of Transient Furrows and Facial Wrinkles

    model/method

    Transient facial features are classified into binary states (present vs. absent) based on Canny edge detection within regions determined by permanent feature positions:

    • Nose Wrinkles: Evaluated inside a square bounding box situated between the two inner eye corners.
    • Crows-Feet Wrinkles: Evaluated inside rectangular regions located lateral to each outer eye corner.

    For nose wrinkles and crows-feet wrinkles, the count of detected edge pixels in the current frame EE is compared against the baseline edge pixel count E0E_0 of the neutral initial frame. Furrows are marked as present if:

    EE0>T\frac{E}{E_0} > T

    where TT is a detection threshold; otherwise they are absent.

    • Nasolabial Furrows: Evaluated in the facial zone between the inner eye corners horizontal line and the lip corners horizontal line. Presence is determined by extracting continuous diagonal edges, and the angle of the furrow relative to the eye corner baseline is computed.
  5. Knowl 5 — Neural Network Architecture for Action Unit Classification

    model/method

    Action unit recognition is implemented using three-layer feedforward neural networks (multi-layer perceptrons) with one hidden layer:

    • Upper Face Network: Receives the 15 normalized upper face parameters. For single AU classification, the network uses 6 hidden units and 7 output nodes corresponding to AU0 (neutral), AU1 (inner brow raiser), AU2 (outer brow raiser), AU4 (brow lowerer), AU5 (upper lid raiser), AU6 (cheek raiser), and AU7 (lid tightener). For recognizing combinations, the hidden layer is increased to 12 hidden units to allow multiple output nodes to fire concurrently.

    • Lower Face Network: Receives the 7 normalized lower face parameters (rheight,rwidth,rleft,rright,rtop,rbtm,Snosewr_{height}, r_{width}, r_{left}, r_{right}, r_{top}, r_{btm}, S_{nosew}). The output layer contains 11 nodes corresponding to AU0 (neutral), AU9 (nose wrinkler), AU10 (upper lip raiser), AU12 (lip corner puller), AU15 (lip corner depressor), AU17 (chin raiser), AU20 (lip stretcher), AU25 (lips part), AU26 (jaw drop), AU27 (mouth stretch), and the combined unit AU23+24 (lip tightener and presser).

  6. Knowl 6 — Upper Face Action Unit Recognition Performance

    empirical result

    The upper face AU recognition system was evaluated on image sequences from 24 Caucasian subjects (236 total sequences: 99 containing single AUs and 137 containing AU combinations):

    • On single AU sequences with novel subjects absent from the training set (TrainS3/TestS3TrainS3 / TestS3), the classifier achieved an average recognition rate of 92.9%92.9\% with zero false alarms.

    • When tested on complex AU combinations using a 7-output network modeling individual AUs (TrainC1/TestC1TrainC1 / TestC1), the system achieved an average recognition rate of 95.0%95.0\% with a false alarm rate of 6.4%6.4\%.

    • When nonadditive combinations (AU1+2, AU1+4, AU4+5, AU1+2+4) were explicitly assigned dedicated output units (11 outputs total), average recognition accuracy was 93.7%93.7\% with a 4.5%4.5\% false alarm rate, indicating that separate combinatorial outputs do not improve recognition accuracy over multi-label basic AU modeling.

  7. Knowl 7 — Lower Face Action Unit Recognition Performance

    empirical result

    Lower face action unit recognition was evaluated on the Pitt-CMU database across 463 image sequences from 122 subjects (65% female, 35% male; 85% European-American, 15% African-American or Asian) using 400 sequences (1220 frames) for training and 63 sequences (243 frames) for testing:

    • When modeling 11 basic lower face AUs, the system achieved an average recognition rate of 96.30%96.30\% across 243 test tokens.

    • When explicitly modeling nonadditive combination classes (such as AU9+17 and AU10+17), the recognition rate increased slightly to 96.71%96.71\%.

    • Perfect recognition (100%100\%) was achieved on AU0, AU9, AU12, AU15, AU20, AU25, AU27, and AU23+24.

  8. Knowl 8 — Upper Face Action Unit Combination Performance Breakdown

    data/table

    The table below displays classification performance for upper face AU combination sequences on the TestC1TestC1 test partition (94 total target action unit occurrences) using the 7-output individual AU neural network architecture:

    AU No. Correct False Missed Confused Recognition Rate
    0 22 22 - - - 100%
    1 18 18 - - - 100%
    2 12 12 2 - - 100%
    4 10 10 4 - - 100%
    5 10 7 - 1 2 70.0%
    6 14 12 - 2 - 85.7%
    7 8 8 - - - 100%
    Total 94 89 6 3 2 95.0%

    The overall average recognition accuracy is 95.0%95.0\%, with an overall false alarm rate of 6.4%6.4\%. The majority of false alarms arise when single AU tokens trigger secondary outputs during nonadditive combination presentations (such as AU2 triggering an AU1 output).

  9. Knowl 9 — Lower Face Action Unit Performance Breakdown and Error Modes

    data/table

    The table below shows the lower face action unit recognition results on the 243-token test set when nonadditive combinations are modeled:

    AU No. Correct False Missed Confused Recognition Rate
    0 63 63 - - - 100%
    9 16 16 - - - 100%
    10 12 11 - 1 - 91.67%
    12 14 14 2 - - 100%
    15 12 12 - - - 100%
    17 36 34 - - 2 (AU12) 94.44%
    20 12 12 - - - 100%
    25 50 50 5 - - 100%
    26 14 9 - - 5 (AU25) 64.29%
    27 8 8 - - - 100%
    23+24 6 6 - - - 100%
    Total 243 235 7 1 7 96.71%

    Specific failure modes include:

    • AU26 (jaw drop): Recognized at 64.29%64.29\%; all misidentifications are confusions with AU25 (lips part) because both involve parted lips, and the parametric representation does not explicitly measure mandible tracking.
    • AU10 and AU17: Confusions occur exclusively during AU10+17 co-occurrence, where combined muscle contraction alters local feature morphology and few training examples (10 out of 1220 frames) were present.

Coverage note — No substantial contributed material was omitted from the extraction.

References

  1. 1.M. Bartlett, J. Hager, P.Ekman, and T. Sejnowski. Measuring facial expressions by computer image analysis. Psychophysiology, 36:253–264, 1999.
  2. 2.J. M. Carroll and J. Russell. Facial expression in hollywood’s portrayal of emotion. Journal of Personality and Social Psychology., 72:164–176, 1997.
  3. 3.G. Donato, M. S. Bartlett, J. C. Hager, P. Ekman, and T. J. Sejnowski. Classifying facial actions. International Journal of Pattern Analysis and Machine Intelligence, 21(10):974–989, October 1999.
  4. 4.P. Ekman and W. V. Friesen. The Facial Action Coding System: A Technique For The Measurement of Facial Movement. Consulting Psychologists Press Inc., San Francisco, CA, 1978.
  5. 5.I. A. Essa and A. P. Pentland. Coding, analysis, interpretation, and recognition of facial expressions. IEEE Transc. On Pattern Analysis and Machine Intelligence, 19(7):757–763, JULY 1997.
  6. 6.K. Fukui and O. Yamaguchi. Facial feature point extraction method based on combination of shape extraction and pattern matching. Systems and Computers in Japan, 29(6):49–58, 1998.
  7. 7.M. Kirby and L. Sirovich. Application of the k-l procedure for the characterization of human faces. IEEE Transc. On Pattern Analysis and Machine Intelligence, 12(1):103–108, Jan. 1990.
  8. 8.Y. Kwon and N. Lobo. Age classification from facial images. In Proc. IEEE Conf. Computer Vision and Pattern Recognition, pages 762–767, 1994.
  9. 9.J.-J. J. Lien, T. Kanade, J. F. Chon, and C. C. Li. Detection, tracking, and classification of action units in facial expression. Journal of Robotics and Autonomous System, in press.
  10. 10.B. Lucas and T. Kanade. An interative image registration technique with an application in stereo vision. In The 7th International Joint Conference on Artificial Intelligence, pages 674–679, 1981.
  11. 11.K. Mase. Recognition of facial expression from optical flow. IEICE Transc., E. 74(10):3474–3483, 0ctober 1991.
  12. 12.K. Scherer and P. Ekman. Handbook of methods in nonverbal behavior research. Cambridge University Press, Cambridge, UK, 1982.
  13. 13.D. Terzopoulos and K. Waters. Analysis of facial images using physical and anatomical models. In IEEE International Conference on Computer Vision, pages 727–732, 1990.
  14. 14.Y. Tian, T. Kanade, and J. Cohn. Dual-state parametric eye tracking. In Submit to International Conference on Face and Gesture Recognition, 1999.
  15. 15.Y. Tian, T. Kanade, and J. Cohn. Robust lip tracking by combining shape, color and motion. In Proc. Of ACCV'2000, 2000.
  16. 16.M. Turk and A. Pentland. face recognition using eigenfaces. In Proc. IEEE Conf. Computer Vision and Pattern Recognition, pages 586–591, 1991.
  17. 17.Y. Yacoob and L. S. Davis. Recognizing human facial expression from long image sequences using optical flow. IEEE Trans. On Pattern Analysis and machine Intelligence, 18(6):636–642, June 1996.
  18. 18.A. Yuille, P. Haallinan, and D. S. Cohen. Feature extraction from faces using deformable templates. International Journal of Computer Vision,, 8(2):99–111, 1992.
  19. 19.Z. Zhang. Feature-based facial expression recognition: Sensitivity analysis and experiments with a multilayer perceptron. International Journal of Pattern Recognition and Artificial Intelligence, 13(6):893–911, 1999.

Citation

MLA
Tian, Y.-I., et al. “Recognizing Action Units for Facial Expression Analysis”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 23, no. 2, 2001, pp. 97–115, https://doi.org/10.1109/34.908962.
APA
Tian, Y.-I., Kanade, T., & Cohn, J. F. (2001). Recognizing action units for facial expression analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 23(2), 97–115. https://doi.org/10.1109/34.908962
Chicago
Tian, Y.-I., T. Kanade, and J. F. Cohn. 2001. “Recognizing Action Units for Facial Expression Analysis”. IEEE Transactions on Pattern Analysis and Machine Intelligence 23 (2): 97–115. https://doi.org/10.1109/34.908962.
Harvard
Tian, Y.-I., Kanade, T. and Cohn, J.F. (2001) “Recognizing action units for facial expression analysis”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 23(2), pp. 97–115. Available at: https://doi.org/10.1109/34.908962.
Vancouver
1. Tian Y-I, Kanade T, Cohn JF (2001) Recognizing action units for facial expression analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence 23:97–115

BibTeX

@article{Tian_2001, title={Recognizing action units for facial expression analysis}, volume={23}, ISSN={0162-8828}, url={http://dx.doi.org/10.1109/34.908962}, DOI={10.1109/34.908962}, number={2}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Tian, Y.-I. and Kanade, T. and Cohn, J.F.}, year={2001}, pages={97–115} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF