Deep Facial Expression Recognition: A Survey

Shan LiWeihong Deng

article2018IEEE Transactions on Affective Computing1,750 citations

Presents a systematic guide to deep facial expression recognition, categorizing architectures for static and dynamic data while outlining practical strategies to handle real-world obstacles such as identity bias, head pose, and illumination variations.

Listen

Automatic facial expression recognition is increasingly important for applications such as sociable robotics, medical assessment, driver fatigue monitoring, and interactive user interfaces. While earlier systems relied on laboratory-controlled environments, recent real-world applications require systems that can handle unconstrained conditions, including varying lighting, head poses, occlusions, and diverse personal identities. The article provides a comprehensive survey of deep facial expression recognition, evaluating how modern deep learning architectures and training strategies overcome the core bottlenecks of data scarcity, overfitting, and expression-unrelated variations.

The article systematically assesses state-of-the-art methods across static images and dynamic video sequences. It reviews public benchmark datasets—spanning legacy lab-controlled sets like CK+ and MMI to modern unconstrained sets such as FER2013, RAF-DB, and AffectNet—along with pre-processing pipelines, network architectures, and classification techniques. The analysis covers convolutional neural networks, recurrent networks, deep belief networks, autoencoders, and generative adversarial networks, detailing how specific adaptations target expression classification.

Several key findings emerge from the review. First, multi-stage transfer learning and pre-training on large-scale face recognition datasets substantially reduce overfitting caused by limited emotional training data. Second, deep spatio-temporal networks combining convolutional networks with recurrent units or three-dimensional convolutions consistently outperform static models on video data by capturing dynamic motion and subtle expression intensities. Third, task-specific loss functions—such as island loss and locality-preserving loss—effectively tighten within-class clustering while widening margins between different emotions. Fourth, network ensembles and multi-task learning frameworks successfully isolate expression-specific information from confounding factors like subject identity and head angle.

These findings indicate that transitioning from isolated image models to integrated spatio-temporal and multi-task architectures is essential for practical, high-accuracy emotion recognition. Incorporating identity-disentangling techniques reduces risk in production deployments where personal appearance variations might otherwise degrade performance. However, deploying multi-network ensembles increases computational overhead and memory requirements, presenting clear trade-offs between accuracy and system latency in resource-constrained environments.

To build robust systems, organizations should adopt multi-stage fine-tuning pipelines and incorporate temporal modeling for video streams. Future efforts must focus on expanding large-scale datasets with detailed annotations for pose, occlusion, and demographic attributes, while developing cost-sensitive training methods to address severe class imbalances, such as the scarcity of disgust or fear samples relative to happiness. Furthermore, combining facial analysis with audio, infrared, or three-dimensional depth data is strongly recommended to enhance real-world reliability.

While the article demonstrates high confidence in the evaluated architectures across standard benchmarks, caution is warranted regarding real-world generalization. Many current models exhibit reduced accuracy in cross-database evaluations due to annotation inconsistencies and environmental bias across source datasets.

Cover for Deep Facial Expression Recognition: A Survey

Abstract

With the transition of facial expression recognition (FER) from laboratory-controlled to challenging in-the-wild conditions and the recent success of deep learning techniques in various fields, deep neural networks have increasingly been leveraged to learn discriminative representations for automatic FER. Recent deep FER systems generally focus on two important issues: overfitting caused by a lack of sufficient training data and expression-unrelated variations, such as illumination, head pose and identity bias. In this paper, we provide a comprehensive survey on deep FER, including datasets and algorithms that provide insights into these intrinsic problems. First, we describe the standard pipeline of a deep FER system with the related background knowledge and suggestions of applicable implementations for each stage. We then introduce the available datasets that are widely used in the literature and provide accepted data selection and evaluation principles for these datasets. For the state of the art in deep FER, we review existing novel deep neural networks and related training strategies that are designed for FER based on both static images and dynamic image sequences, and discuss their advantages and limitations. Competitive performances on widely used benchmarks are also summarized in this section. We then extend our survey to additional related issues and application scenarios. Finally, we review the remaining challenges and corresponding opportunities in this field as well as future directions for the design of robust deep FER systems.

Table of Contents

  • I Introduction
  • II Facial expression databases
  • III Deep facial expression recognition
  • III-A Pre-processing
  • III-A1 Face alignment
  • III-A2 Data augmentation
  • III-A3 Face normalization
  • III-B Deep networks for feature learning
  • III-B1 Convolutional neural network (CNN)
  • III-B2 Deep belief network (DBN)
  • III-B3 Deep autoencoder (DAE)
  • III-B4 Recurrent neural network (RNN)
  • III-B5 Generative Adversarial Network (GAN)
  • III-C Facial expression classification
  • IV The state of the art
  • IV-A Deep FER networks for static images
  • IV-A1 Pre-training and fine-tuning
  • IV-A2 Diverse network input
  • IV-A3 Auxiliary blocks & layers
  • IV-A4 Network ensemble
  • IV-A5 Multitask networks
  • IV-A6 Cascaded networks
  • IV-A7 Generative adversarial networks (GANs)
  • IV-A8 Discussion
  • IV-B Deep FER networks for dynamic image sequences
  • IV-B1 Frame aggregation
  • IV-B2 Expression Intensity network
  • IV-B3 Deep spatio-temporal FER network
  • IV-B4 Discussion
  • V Additional Related Issues
  • V-A Occlusion and non-frontal head pose
  • V-B FER on infrared data
  • V-C FER on 3D static and dynamic data
  • V-D Facial expression synthesis
  • V-E Visualization techniques
  • V-F Other special issues
  • VI Challenges and Opportunities
  • VI-A Facial expression datasets
  • VI-B Incorporating other affective models
  • VI-C Dataset bias and imbalanced distribution
  • VI-D Multimodal affect recognition
  • References

Knowls

  1. Knowl 1 — Standard Pipeline and Functional Stages for Deep Facial Expression Recognition

    model/method

    Facial Expression Recognition (FER) systems using deep learning architectures follow a three-stage processing pipeline:

    1. Pre-processing:

      • Face Detection and Landmark Alignment: Face bounding boxes are extracted (e.g., using the Viola-Jones detector or MTCNN) and normalized using localized facial landmarks (e.g., via Active Appearance Models, Supervised Descent Method, or Multi-Task Cascaded CNNs) to remove scale discrepancies and in-plane rotations.
      • Data Augmentation: On-the-fly transformations (e.g., random five-point cropping and horizontal flips yielding 10×10\times data expansion) and offline perturbations (affine geometric transformations, rotation, scaling, Gaussian/salt-and-pepper noise, HSV color jittering, 3D model synthesis, or generative adversarial generation) mitigate overfitting.
      • Face Normalization: Illumination normalization (isotropic diffusion, discrete cosine transform normalization, difference of Gaussians, homomorphic filtering, or histogram equalization) and pose normalization/frontalization (3D facial reconstruction or GAN-based synthesis) reduce non-expression appearance variations.
    2. Deep Feature Learning:

      • Deep neural networks extract high-level discriminative representations. Architectural families include Convolutional Neural Networks (CNNs), Deep Belief Networks (DBNs) trained with restricted Boltzmann machines via greedy layerwise contrastive divergence, Deep Autoencoders (DAEs, sparse autoencoders, and contractive autoencoders), Recurrent Neural Networks (RNNs, LSTMs, and Bidirectional RNNs), and Generative Adversarial Networks (GANs).
    3. Facial Expression Classification:

      • End-to-end networks apply classification loss layers directly to output probabilities (e.g., Softmax cross-entropy loss, linear Support Vector Machines, or Deep Neural Decision Forests). Alternatively, deep networks act as feature extractors whose outputs feed independent classifiers such as SVMs, Random Forests, or covariance pooling on Symmetric Positive Definite (SPD) Riemannian manifolds.
  2. Knowl 2 — Taxonomy and Characteristics of Benchmark Facial Expression Datasets

    data/table

    Facial expression recognition benchmarks span laboratory-controlled conditions and unconstrained in-the-wild settings across static image and dynamic video modalities:

    Dataset Samples Subjects Condition Elicitation Classes Modality
    CK+ 593 seq. 123 Lab Posed Spontaneous 6 basic + contempt + neutral Dynamic
    MMI 740 img, 2900 vid 25 Lab Posed 6 basic + neutral Dynamic
    JAFFE 213 images 10 Lab Posed 6 basic + neutral Static
    TFD 112,234 images N/A Lab Posed 6 basic + neutral Static
    FER-2013 35,887 images N/A Web Posed Spontaneous 6 basic + neutral Static
    AFEW 7.0 1,809 videos N/A Movie Posed Spontaneous 6 basic + neutral Dynamic
    SFEW 2.0 1,766 images N/A Movie Posed Spontaneous 6 basic + neutral Static
    Multi-PIE 755,370 images 337 Lab Posed 6 expressions (multi-view) Static
    BU-3DFE 2,500 images 100 Lab Posed 6 basic + neutral (3D) Static
    Oulu-CASIA 2,880 seq. 80 Lab Posed 6 basic (VIS/NIR) Dynamic
    RaFD 1,608 images 67 Lab Posed 6 basic + contempt + neutral Static
    KDEF 4,900 images 70 Lab Posed 6 basic + neutral Static
    EmotioNet 1,000,000 images N/A Web Posed Spontaneous 23 basic / compound Static
    RAF-DB 29,672 images N/A Web Posed Spontaneous 7 basic + 12 compound Static
    AffectNet 450,000 labeled N/A Web Posed Spontaneous 8 basic / Valence-Arousal Static
    ExpW 91,793 images N/A Web Posed Spontaneous 6 basic + neutral Static

    Standard evaluation setups for CK+ and MMI employ person-independent nn-fold cross-validation (typically n∈{5,8,10}n \in \{5, 8, 10\} or leave-one-subject-out) on peak frames (static) or sequences (dynamic). Web-collected benchmarks (FER2013, SFEW 2.0, RAF-DB, AffectNet, AFEW) define official disjoint training, validation, and test splits.

  3. Knowl 3 — Specialized Loss Functions for Intra-Class Compactness and Inter-Class Discrimination in Deep FER

    equation

    Standard softmax cross-entropy loss forces features of distinct expression classes to diverge, but does not sufficiently penalize large intra-class variations caused by identity and pose differences. Specialized loss layers enforce intra-class compactness and enlarged inter-class margins:

    1. Island Loss: Combines the standard softmax loss with a feature-level penalty that pulls deep features toward their respective class centroids while maximizing pairwise Euclidean distances among all class centroids: LIsland=Lsoftmax+λ1∑i=1N∥xi−cyi∥22+λ2∑j=1C∑k=1k≠jC1∥cj−ck∥22+ϵL_{\text{Island}} = L_{\text{softmax}} + \lambda_1 \sum_{i=1}^{N} \| x_i - c_{y_i} \|_2^2 + \lambda_2 \sum_{j=1}^{C} \sum_{\substack{k=1 \\ k \neq j}}^{C} \frac{1}{\| c_j - c_k \|_2^2 + \epsilon} where xi∈Rdx_i \in \mathbb{R}^d is the extracted deep feature vector for sample ii, yi∈{1,…,C}y_i \in \{1, \dots, C\} is its ground-truth expression label, cj∈Rdc_j \in \mathbb{R}^d represents the learned center of class jj, NN is the batch size, and λ1,λ2,ϵ>0\lambda_1, \lambda_2, \epsilon > 0 are balancing hyperparameters.

    2. Locality-Preserving (LP) Loss: Preserves local neighborhood graphs to ensure intra-class local feature clusters remain compact on the feature manifold: LLP=12∑i=1N∑j:xj∈N(xi,yi)Wij∥xi−xj∥22L_{\text{LP}} = \frac{1}{2} \sum_{i=1}^{N} \sum_{j: x_j \in \mathcal{N}(x_i, y_i)} W_{ij} \| x_i - x_j \|_2^2 where N(xi,yi)\mathcal{N}(x_i, y_i) is the set of kk-nearest neighbors of sample xix_i sharing the same expression label yiy_i, and WijW_{ij} is the affinity weight measuring local similarity between xix_i and xjx_j.

    3. (N+M)-Tuples Cluster Loss: Extends triplet metric learning by mining NN identity-aware hard negative samples and MM online positive samples simultaneously, pulling samples of the same expression from different subjects together while pushing different expressions apart without requiring manual anchor selection.

  4. Knowl 4 — Peak-Piloted Deep Network and Peak Gradient Suppression for Expression Intensity Invariance

    model/method

    The Peak-Piloted Deep Network (PPDN) recognizes subtle, non-peak facial expressions by transferring high-intensity expressive feature knowledge to low-intensity samples from the same subject:

    • Input Pairs: Training requires sample pairs (xp,xnp)(x_p, x_{np}) consisting of a peak-intensity expression image xpx_p and a non-peak (or neutral) expression image xnpx_{np} from the same individual with the same underlying expression label yy.

    • Optimization Objective: LPPDN=LCE(f(xp),y)+LCE(f(xnp),y)+λ2∥f(xp)−f(xnp)∥22L_{\text{PPDN}} = L_{\text{CE}}(f(x_p), y) + L_{\text{CE}}(f(x_{np}), y) + \frac{\lambda}{2} \| f(x_p) - f(x_{np}) \|_2^2 where f(⋅)f(\cdot) is the deep feature embedding, LCEL_{\text{CE}} is the cross-entropy classification loss, and λ>0\lambda > 0 is a scalar hyperparameter balancing the feature distance penalty.

    • Peak Gradient Suppression (PGS): Standard backpropagation on ∥f(xp)−f(xnp)∥22\| f(x_p) - f(x_{np}) \|_2^2 would backpropagate gradients into both f(xp)f(x_p) and f(xnp)f(x_{np}), pulling peak features toward non-peak features and corrupting the discriminative peak representation. PGS sets ∂LL2∂f(xp)=0\frac{\partial L_{L2}}{\partial f(x_p)} = 0 during backpropagation, enforcing that only non-peak features f(xnp)f(x_{np}) are updated toward peak features f(xp)f(x_p).

    • Inference: At test time, PPDN operates on individual still images without requiring paired images or intensity annotations.

  5. Knowl 5 — Comparative Taxonomy and Trade-offs of Static Deep FER Paradigms

    data/table

    Deep FER methods for static images utilize diverse architectures and learning formulations to address data scarcity and expression-unrelated variations:

    Network Type Data Size Req. Environmental Variations Identity Bias Efficiency Accuracy Training Difficulty
    Pre-train Fine-tune Low Fair Vulnerable High Fair Easy
    Diverse Input Low Good Vulnerable Low Fair Easy
    Auxiliary Layers Varies Good Varies Varies Good Varies
    Network Ensemble Low Good Fair Low Good Medium
    Multitask Network High Varies Good Fair Varies Hard
    Cascaded Network Fair Good Fair Fair Fair Medium
    GAN Fair Good Good Fair Good Hard
    • Pre-training and fine-tuning transfers representations from face recognition datasets, but risk identity-biased feature dominance.
    • Diverse inputs (mapped LBP, SIFT, part-based crops) offer robustness against lighting and misregistration, but increase input dimensionality.
    • Auxiliary layers and custom losses (Island loss, LP loss, Supervised Scoring Ensembles) directly optimize discriminative margins.
    • Network ensembles fuse diverse models at feature or decision levels (majority voting, simple/weighted averaging) to boost performance at the expense of computational and memory overhead.
    • Multitask networks jointly optimize FER alongside secondary tasks (facial landmark detection, Action Unit detection, face verification) to disentangle identity and head pose.
    • GAN-based methods explicitly frontalize faces, synthesize multi-view appearances, or regenerate neutral/expressive counterpart pairs to cancel identity bias.
  6. Knowl 6 — Spatio-Temporal Deep Architectures for Dynamic Facial Expression Recognition

    model/method

    Dynamic Facial Expression Recognition encodes temporal dependencies across contiguous video frames using several deep learning paradigms:

    1. Frame Aggregation:

      • Decision-level Aggregation: Class probability vectors across a sequence are unified via fixed-length frame averaging (averaging probabilities across k=10k=10 temporal bins for long sequences) or frame expansion (repeating frames uniformly for short sequences), or via statistical summary pooling (mean, max, square-average, maximum suppression).
      • Feature-level Aggregation: Deep feature representations of individual frames are combined through statistical pooling (concatenating mean, variance, min, max vectors), covariance matrices on Riemannian manifolds, or multi-instance dictionary learning (NetVLAD).
    2. 3D Convolutional Neural Networks (C3D):

      • 3D convolutions with shared kernels along the temporal dimension simultaneously capture spatial appearance and temporal dynamics. Architectural variants include Deep Temporal Appearance Networks (DTAN, using unshared temporal filters allowing time-varying kernel importance) and 3DCNN-DAP (incorporating deformable facial action part constraints).
    3. Facial Landmark Trajectory Modeling:

      • Tracks coordinate displacements or inter-landmark Euclidean distance variations over consecutive frames. Part-based Hierarchical RNNs (PHRNN) divide facial landmarks into anatomical subsets (eyebrows, eyes, nose, mouth) and feed them sequentially into hierarchical recurrent sub-networks.
    4. Cascaded CNN-RNN/LSTM Networks:

      • Spatially deep CNN modules extract per-frame visual embeddings, which feed into Recurrent Neural Networks (LSTM, Bidirectional RNNs, IRNNs with ReLU and identity initialization, or Nested LSTMs) or Conditional Random Fields (CRFs) to model sequential transitions.
    5. Multi-Stream Spatial-Temporal Ensembles:

      • Parallel deep streams separately model appearance (still RGB frames), motion dynamics (dense optical flow), and structural geometry (landmark trajectories), fusing representations via score averaging, learned neural classifiers, or joint fine-tuning (e.g., Deep Temporal Appearance-Geometry Networks, DTAGN).
  7. Knowl 7 — Two-Stage Knowledge Transfer via FaceNet2ExpNet

    model/method

    Directly training deep networks on small expression datasets causes overfitting, whereas fine-tuning an off-the-shelf face recognition (FR) model leads to identity-related features dominating the representation. The FaceNet2ExpNet algorithm eliminates face-identity distraction while preserving representational capacity via a two-stage training scheme:

    Input: Pre-trained face recognition network Net_{face}, Training dataset of facial expression images with emotion labels \mathcal{D} = \{(x_i, y_i)\}_{i=1}^N
    Output: Trained facial expression recognition network Net_{exp}
    Initialize convolutional layers of Net_{exp} with random weights;
    Freeze all parameters of Net_{face};
    Stage 1: Feature-Level Regularization
    for each mini-batch of images \{x_b\} in \mathcal{D} do
        Forward propagate \{x_b\} through Net_{face} to obtain reference convolutional features \Phi_{face}(x_b);
        Forward propagate \{x_b\} through convolutional layers of Net_{exp} to obtain \Phi_{exp}(x_b);
        Compute distribution divergence loss L_{dist}(\Phi_{exp}(x_b), \Phi_{face}(x_b));
        Update convolutional parameters of Net_{exp} via backpropagation of L_{dist};
    end
    Stage 2: Joint Supervised Expression Fine-Tuning
    Append randomly initialized fully connected layers to the trained convolutional layers of Net_{exp};
    Unfreeze all layers of Net_{exp};
    for each mini-batch \{(x_b, y_b)\} in \mathcal{D} do
        Forward propagate \{x_b\} through the entire Net_{exp} to compute expression logits \hat{y}_b;
        Compute classification cross-entropy loss L_{CE}(\hat{y}_b, y_b);
        Update all parameters of Net_{exp} via backpropagation of L_{CE};
    end
    return Net_{exp}
  8. Knowl 8 — Comparative Taxonomy and Trade-offs of Dynamic Deep FER Paradigms

    data/table

    Dynamic deep FER methods process sequence data with differing capabilities in spatio-temporal modeling, sequence length handling, and computational overhead:

    Network Type Data Req. Spatial Rep. Temporal Rep. Frame Length Accuracy Efficiency
    Frame Aggregation Low Good No Variable/Fixed Fair High
    Expression Intensity Net Fair Good Low Fixed Fair Varies
    Spatio-Temporal: RNN Low Low Good Variable Low Fair
    Spatio-Temporal: C3D High Good Fair Fixed Low Fair
    Spatio-Temporal: Landmark Trajectory Fair Fair Fair Fixed Low High
    Spatio-Temporal: Cascaded CNN+LSTM High Good Good Variable Good Fair
    Spatio-Temporal: Network Ensemble Low Good Good Fixed Good Low
    • Frame aggregation is computationally efficient but ignores contiguous temporal transitions.
    • RNNs alone lack spatial feature extraction strength, whereas C3Ds require large training sets and apply only to short fixed temporal windows.
    • Landmark trajectories provide illumination-invariant motion cues but degrade under face registration errors.
    • Cascaded CNN-LSTM architectures and multi-stream ensembles (combining spatial appearance, optical flow, and landmark geometry) achieve state-of-the-art recognition performance on video benchmarks.
  9. Knowl 9 — Representative Benchmark Performance of Static and Dynamic Deep FER Systems

    empirical result

    Representative state-of-the-art deep FER methods evaluated under person-independent protocols on major static and dynamic datasets demonstrate the impact of specialized architectures and temporal modeling:

    • Static Image Benchmarks:

      • CK+ (7 or 8 classes): FaceNet2ExpNet achieves 98.6% (6-class) and 96.8% (8-class); Island Loss CNN achieves 94.39% (7-class); (N+M)-tuple clusters loss achieves 97.1% (7-class); cGAN (DeRL) achieves 97.30% (7-class).
      • MMI (6 or 7 classes): FaceNet2ExpNet reaches 78.46% (6-class); (N+M)-tuple cluster loss reaches 78.53% (6-class); Identity-Aware CNN (IACNN) reaches 75.85% (7-class).
      • FER2013 (7 classes): CNN with linear SVM loss achieves 71.20%; Supervised Scoring Ensemble (SSE) reaches 73.73%; Multi-network ensemble of deep CNNs reaches 75.20%.
      • SFEW 2.0 (In-the-wild, 7 classes): Baseline fine-tuned AlexNet validation accuracy is 48.50% (test 55.60%); Island Loss CNN achieves 59.41% test accuracy; Hierarchical committee ensemble of CNNs reaches 61.60% (test).
    • Dynamic Video Benchmarks:

      • CK+: Peak-Piloted Deep Network (PPDN) achieves 99.3% (6-class); Deeper Cascaded Peak-Piloted Network (DCPN) achieves 99.6%; Spatio-temporal PHRNN + MSCNN ensemble achieves 98.50% (7-class); Joint fine-tuned DTAGN achieves 97.25% (7-class).
      • MMI: Multi-channel ensemble network achieves 91.46% (6-class); Spatio-temporal PHRNN + MSCNN achieves 81.18% (6-class); Multi-objective intensity CNN achieves 78.61%.
      • Oulu-CASIA (6 classes): PPDN achieves 84.59%; DCPN achieves 86.23%; Spatio-temporal PHRNN + MSCNN achieves 86.25%; Joint fine-tuned DTAGN achieves 91.67%.
      • AFEW 7.0 (In-the-wild challenge): Single-modality VGG16-LSTM achieves 48.60% validation accuracy; Spatio-temporal audio-visual multimodal fusion systems achieve 58.81% to 59.02% test accuracy.
  10. Knowl 10 — Open Challenges in Deep Facial Expression Recognition

    limitation

    Deep FER systems operate under four foundational constraints and open research challenges:

    1. Data Scarcity and Annotation Quality: Deep neural networks require substantial training data to capture subtle facial deformations without overfitting. Merging heterogeneous datasets introduces inconsistent annotations that degrade learning without latent truth modeling (e.g., LTNet / IPA2LT). Furthermore, head-pose and occlusion annotations remain sparse.

    2. Severe Class Imbalance: Naturalistic datasets suffer from severe class skew (e.g., happy expressions are abundant and easily captured, while disgust and fear are rare). This produces a pronounced drop in mean per-class accuracy compared to global accuracy, particularly in wild benchmarks like SFEW 2.0 and AFEW.

    3. Unconstrained Pose and Occlusion Variations: Non-frontal head orientations and facial occlusions non-linearly distort visual appearance. While 3D face models, GAN-based frontalization, and multi-channel pose-aware networks offer partial solutions, end-to-end joint modeling under extreme poses remains an open challenge.

    4. Limitations of the Categorical Emotion Model: The prototypical 6-to-7 discrete emotion taxonomy cannot represent the continuous and subtle spectrum of real-world affect. Bridging categorical FER with the Facial Action Coding System (FACS Action Units), continuous dimensional models (Valence-Arousal), and compound emotion categories is essential for comprehensive affective computing.

Coverage note — None was omitted. All principal contributions of the survey—including deep FER pipelines, dataset categorization, static and dynamic deep architectures, specialized loss functions, intensity-invariant modeling, transfer learning protocols, benchmark performance comparisons, and future challenges—are covered.

References

  1. 1.C. Darwin and P. Prodger, The expression of the emotions in man and animals. Oxford University Press, USA, 1998.
  2. 2.Y.-I. Tian, T. Kanade, and J. F. Cohn, "Recognizing action units for facial expression analysis," IEEE Transactions on pattern analysis and machine intelligence, vol. 23, no. 2, pp. 97–115, 2001.
  3. 3.P. Ekman and W. V. Friesen, "Constants across cultures in the face and emotion." Journal of personality and social psychology, vol. 17, no. 2, pp. 124–129, 1971.
  4. 4.P. Ekman, "Strong evidence for universals in facial expressions: a reply to russell's mistaken critique," Psychological bulletin, vol. 115, no. 2, pp. 268–287, 1994.
  5. 5.D. Matsumoto, "More evidence for the universality of a contempt expression," Motivation and Emotion, vol. 16, no. 4, pp. 363–368, 1992.
  6. 6.R. E. Jack, O. G. Garrod, H. Yu, R. Caldara, and P. G. Schyns, "Facial expressions of emotion are not culturally universal," Proceedings of the National Academy of Sciences, vol. 109, no. 19, pp. 7241–7244, 2012.
  7. 7.Z. Zeng, M. Pantic, G. I. Roisman, and T. S. Huang, "A survey of affect recognition methods: Audio, visual, and spontaneous expressions," IEEE transactions on pattern analysis and machine intelligence, vol. 31, no. 1, pp. 39–58, 2009.
  8. 8.E. Sariyanidi, H. Gunes, and A. Cavallaro, "Automatic analysis of facial affect: A survey of registration, representation, and recognition," IEEE transactions on pattern analysis and machine intelligence, vol. 37, no. 6, pp. 1113–1133, 2015.
  9. 9.B. Martinez and M. F. Valstar, "Advances, challenges, and opportunities in automatic facial expression recognition," in Advances in Face Detection and Facial Image Analysis. Springer, 2016, pp. 63–100.
  10. 10.P. Ekman, "Facial action coding system (facs)," A human face, 2002.
  11. 11.H. Gunes and B. Schuller, "Categorical and dimensional affect analysis in continuous input: Current trends and future directions," Image and Vision Computing, vol. 31, no. 2, pp. 120–136, 2013.
  12. 12.C. Shan, S. Gong, and P. W. McOwan, "Facial expression recognition based on local binary patterns: A comprehensive study," Image and Vision Computing, vol. 27, no. 6, pp. 803–816, 2009.
  13. 13.P. Liu, S. Han, Z. Meng, and Y. Tong, "Facial expression recognition via a boosted deep belief network," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 1805–1812.
  14. 14.A. Mollahosseini, D. Chan, and M. H. Mahoor, "Going deeper in facial expression recognition using deep neural networks," in Applications of Computer Vision (WACV), 2016 IEEE Winter Conference on. IEEE, 2016, pp. 1–10.
  15. 15.G. Zhao and M. Pietikainen, "Dynamic texture recognition using local binary patterns with an application to facial expressions," IEEE transactions on pattern analysis and machine intelligence, vol. 29, no. 6, pp. 915–928, 2007.
  16. 16.H. Jung, S. Lee, J. Yim, S. Park, and J. Kim, "Joint fine-tuning in deep neural networks for facial expression recognition," in Computer Vision (ICCV), 2015 IEEE International Conference on. IEEE, 2015, pp. 2983–2991.
  17. 17.X. Zhao, X. Liang, L. Liu, T. Li, Y. Han, N. Vasconcelos, and S. Yan, "Peak-piloted deep network for facial expression recognition," in European conference on computer vision. Springer, 2016, pp. 425–442.
  18. 18.C. A. Corneanu, M. O. Simón, J. F. Cohn, and S. E. Guerrero, "Survey on rgb, 3d, thermal, and multimodal approaches for facial expression recognition: History, trends, and affect-related applications," IEEE transactions on pattern analysis and machine intelligence, vol. 38, no. 8, pp. 1548–1568, 2016.
  19. 19.R. Zhi, M. Flierl, Q. Ruan, and W. B. Kleijn, "Graph-preserving sparse nonnegative matrix factorization with application to facial expression recognition," IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 41, no. 1, pp. 38–52, 2011.
  20. 20.L. Zhong, Q. Liu, P. Yang, B. Liu, J. Huang, and D. N. Metaxas, "Learning active facial patches for expression analysis," in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on. IEEE, 2012, pp. 2562–2569.
  21. 21.I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cukierski, Y. Tang, D. Thaler, D.-H. Lee et al., "Challenges in representation learning: A report on three machine learning contests," in International Conference on Neural Information Processing. Springer, 2013, pp. 117–124.
  22. 22.A. Dhall, O. Ramana Murthy, R. Goecke, J. Joshi, and T. Gedeon, "Video and image based emotion recognition challenges in the wild: Emotiw 2015," in Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, 2015, pp. 423–426.
  23. 23.A. Dhall, R. Goecke, J. Joshi, J. Hoey, and T. Gedeon, "Emotiw 2016: Video and group-level emotion recognition challenges," in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 427–432.
  24. 24.A. Dhall, R. Goecke, S. Ghosh, J. Joshi, J. Hoey, and T. Gedeon, "From individual to group-level emotion recognition: Emotiw 5.0," in Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, 2017, pp. 524–528.
  25. 25.A. Krizhevsky, I. Sutskever, and G. E. Hinton, "Imagenet classification with deep convolutional neural networks," in Advances in neural information processing systems, 2012, pp. 1097–1105.
  26. 26.K. Simonyan and A. Zisserman, "Very deep convolutional networks for large-scale image recognition," arXiv preprint arXiv:1409.1556, 2014.
  27. 27.C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, "Going deeper with convolutions," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 1–9.
  28. 28.K. He, X. Zhang, S. Ren, and J. Sun, "Deep residual learning for image recognition," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778.
  29. 29.M. Pantic and L. J. M. Rothkrantz, "Automatic analysis of facial expressions: The state of the art," IEEE Transactions on pattern analysis and machine intelligence, vol. 22, no. 12, pp. 1424–1445, 2000.
  30. 30.B. Fasel and J. Luettin, "Automatic facial expression analysis: a survey," Pattern recognition, vol. 36, no. 1, pp. 259–275, 2003.
  31. 31.T. Zhang, "Facial expression recognition based on deep learning: A survey," in International Conference on Intelligent and Interactive Systems and Applications. Springer, 2017, pp. 345–352.
  32. 32.M. F. Valstar, M. Mehu, B. Jiang, M. Pantic, and K. Scherer, "Metaanalysis of the first facial expression recognition challenge," IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 42, no. 4, pp. 966–979, 2012.
  33. 33.P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews, "The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression," in Computer Vision and Pattern Recognition Workshops (CVPRW), 2010 IEEE Computer Society Conference on. IEEE, 2010, pp. 94–101.
  34. 34.M. Pantic, M. Valstar, R. Rademaker, and L. Maat, "Web-based database for facial expression analysis," in Multimedia and Expo, 2005. ICME 2005. IEEE International Conference on. IEEE, 2005, pp. 5–pp.
  35. 35.M. Valstar and M. Pantic, "Induced disgust, happiness and surprise: an addition to the mmi facial expression database," in Proc. 3rd Intern. Workshop on EMOTION (satellite of LREC): Corpora for Research on Emotion and Affect, 2010, p. 65.
  36. 36.M. Lyons, S. Akamatsu, M. Kamachi, and J. Gyoba, "Coding facial expressions with gabor wavelets," in Automatic Face and Gesture Recognition, 1998. Proceedings. Third IEEE International Conference on. IEEE, 1998, pp. 200–205.
  37. 37.J. M. Susskind, A. K. Anderson, and G. E. Hinton, "The toronto face database," Department of Computer Science, University of Toronto, Toronto, ON, Canada, Tech. Rep, vol. 3, 2010.
  38. 38.R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker, "Multi-pie," Image and Vision Computing, vol. 28, no. 5, pp. 807–813, 2010.
  39. 39.L. Yin, X. Wei, Y. Sun, J. Wang, and M. J. Rosato, "A 3d facial expression database for facial behavior research," in Automatic face and gesture recognition, 2006. FGR 2006. 7th international conference on. IEEE, 2006, pp. 211–216.
  40. 40.G. Zhao, X. Huang, M. Taini, S. Z. Li, and M. PietikaInen, "Facial expression recognition from near-infrared videos," Image and Vision Computing, vol. 29, no. 9, pp. 607–619, 2011.
  41. 41.O. Langner, R. Dotsch, G. Bijlstra, D. H. Wigboldus, S. T. Hawk, and A. van Knippenberg, "Presentation and validation of the radboud faces database," Cognition and Emotion, vol. 24, no. 8, pp. 1377–1388, 2010.
  42. 42.D. Lundqvist, A. Flykt, and A. Öhman, "The karolinska directed emotional faces (kdef)," CD ROM from Department of Clinical Neuroscience, Psychology section, Karolinska Institutet, no. 1998, 1998.
  43. 43.C. F. Benitez-Quiroz, R. Srinivasan, and A. M. Martinez, "Emotionet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild," in Proceedings of IEEE International Conference on Computer Vision & Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016.
  44. 44.S. Li, W. Deng, and J. Du, "Reliable crowdsourcing and deep localitypreserving learning for expression recognition in the wild," in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2017, pp. 2584–2593.
  45. 45.S. Li and W. Deng, "Reliable crowdsourcing and deep localitypreserving learning for unconstrained facial expression recognition," IEEE Transactions on Image Processing, 2018.
  46. 46.A. Mollahosseini, B. Hasani, and M. H. Mahoor, "Affectnet: A database for facial expression, valence, and arousal computing in the wild," IEEE Transactions on Affective Computing, vol. PP, no. 99, pp. 1–1, 2017.
  47. 47.Z. Zhang, P. Luo, C. L. Chen, and X. Tang, "From facial expression recognition to interpersonal relation prediction," International Journal of Computer Vision, vol. 126, no. 5, pp. 1–20, 2018.
  48. 48.A. Dhall, R. Goecke, S. Lucey, T. Gedeon et al., "Collecting large, richly annotated facial-expression databases from movies," IEEE multimedia, vol. 19, no. 3, pp. 34–41, 2012.
  49. 49.A. Dhall, R. Goecke, S. Lucey, and T. Gedeon, "Acted facial expressions in the wild database," Australian National University, Canberra, Australia, Technical Report TR-CS-11, vol. 2, p. 1, 2011.
  50. 50.——, "Static facial expression analysis in tough conditions: Data, evaluation protocol and benchmark," in Computer Vision Workshops (ICCV Workshops), 2011 IEEE International Conference on. IEEE, 2011, pp. 2106–2112.
  51. 51.C. F. Benitez-Quiroz, R. Srinivasan, Q. Feng, Y. Wang, and A. M. Martinez, "Emotionet challenge: Recognition of facial expressions of emotion in the wild," arXiv preprint arXiv:1703.01210, 2017.
  52. 52.S. Du, Y. Tao, and A. M. Martinez, "Compound facial expressions of emotion," Proceedings of the National Academy of Sciences, vol. 111, no. 15, pp. E1454–E1462, 2014.
  53. 53.T. F. Cootes, G. J. Edwards, and C. J. Taylor, "Active appearance models," IEEE Transactions on Pattern Analysis & Machine Intelligence, no. 6, pp. 681–685, 2001.
  54. 54.N. Zeng, H. Zhang, B. Song, W. Liu, Y. Li, and A. M. Dobaie, "Facial expression recognition via learning deep sparse autoencoders," Neurocomputing, vol. 273, pp. 643–649, 2018.
  55. 55.B. Hasani and M. H. Mahoor, "Spatio-temporal facial expression recognition using convolutional neural networks and conditional random fields," in Automatic Face & Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on. IEEE, 2017, pp. 790–795.
  56. 56.X. Zhu and D. Ramanan, "Face detection, pose estimation, and landmark localization in the wild," in Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on. IEEE, 2012, pp. 2879–2886.
  57. 57.S. E. Kahou, C. Pal, X. Bouthillier, P. Froumenty, Ç. Gülçehre, R. Memisevic, P. Vincent, A. Courville, Y. Bengio, R. C. Ferrari et al., "Combining modality specific deep neural networks for emotion recognition in video," in Proceedings of the 15th ACM on International conference on multimodal interaction. ACM, 2013, pp. 543–550.
  58. 58.T. Devries, K. Biswaranjan, and G. W. Taylor, "Multi-task learning of facial landmarks and expression," in Computer and Robot Vision (CRV), 2014 Canadian Conference on. IEEE, 2014, pp. 98–103.
  59. 59.A. Asthana, S. Zafeiriou, S. Cheng, and M. Pantic, "Robust discriminative response map fitting with constrained local models," in Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on. IEEE, 2013, pp. 3444–3451.
  60. 60.M. Shin, M. Kim, and D.-S. Kwon, "Baseline cnn structure analysis for facial expression recognition," in Robot and Human Interactive Communication (RO-MAN), 2016 25th IEEE International Symposium on. IEEE, 2016, pp. 724–729.
  61. 61.Z. Meng, P. Liu, J. Cai, S. Han, and Y. Tong, "Identity-aware convolutional neural network for facial expression recognition," in Automatic Face & Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on. IEEE, 2017, pp. 558–565.
  62. 62.X. Xiong and F. De la Torre, "Supervised descent method and its applications to face alignment," in Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on. IEEE, 2013, pp. 532–539.
  63. 63.H.-W. Ng, V. D. Nguyen, V. Vonikakis, and S. Winkler, "Deep learning for emotion recognition on small datasets using transfer learning," in Proceedings of the 2015 ACM on international conference on multimodal interaction. ACM, 2015, pp. 443–449.
  64. 64.S. Ren, X. Cao, Y. Wei, and J. Sun, "Face alignment at 3000 fps via regressing local binary features," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 1685–1692.
  65. 65.A. Asthana, S. Zafeiriou, S. Cheng, and M. Pantic, "Incremental face alignment in the wild," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 1859–1866.
  66. 66.D. H. Kim, W. Baddar, J. Jang, and Y. M. Ro, "Multi-objective based spatio-temporal feature representation learning robust to expression intensity variations for facial expression recognition," IEEE Transactions on Affective Computing, 2017.
  67. 67.Y. Sun, X. Wang, and X. Tang, "Deep convolutional network cascade for facial point detection," in Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on. IEEE, 2013, pp. 3476–3483.
  68. 68.K. Zhang, Y. Huang, Y. Du, and L. Wang, "Facial expression recognition based on deep evolutional spatial-temporal networks," IEEE Transactions on Image Processing, vol. 26, no. 9, pp. 4193–4203, 2017.
  69. 69.K. Zhang, Z. Zhang, Z. Li, and Y. Qiao, "Joint face detection and alignment using multitask cascaded convolutional networks," IEEE Signal Processing Letters, vol. 23, no. 10, pp. 1499–1503, 2016.
  70. 70.Z. Yu, Q. Liu, and G. Liu, "Deeper cascaded peak-piloted network for weak expression recognition," The Visual Computer, pp. 1–9, 2017.
  71. 71.Z. Yu, G. Liu, Q. Liu, and J. Deng, "Spatio-temporal convolutional features with nested lstm for facial expression recognition," Neurocomputing, vol. 317, pp. 50–57, 2018.
  72. 72.P. Viola and M. Jones, "Rapid object detection using a boosted cascade of simple features," in Computer Vision and Pattern Recognition, 2001. CVPR 2001. Proceedings of the 2001 IEEE Computer Society Conference on, vol. 1. IEEE, 2001, pp. I–I.
  73. 73.F. De la Torre, W.-S. Chu, X. Xiong, F. Vicente, X. Ding, and J. F. Cohn, "Intraface," in IEEE International Conference on Automatic Face and Gesture Recognition (FG), 2015.
  74. 74.Z. Zhang, P. Luo, C. C. Loy, and X. Tang, "Facial landmark detection by deep multi-task learning," in European Conference on Computer Vision. Springer, 2014, pp. 94–108.
  75. 75.Z. Yu and C. Zhang, "Image based static facial expression recognition with multiple deep network learning," in Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, 2015, pp. 435–442.
  76. 76.B.-K. Kim, H. Lee, J. Roh, and S.-Y. Lee, "Hierarchical committee of deep cnns with exponentially-weighted decision fusion for static facial expression recognition," in Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, 2015, pp. 427–434.
  77. 77.X. Liu, B. Kumar, J. You, and P. Jia, "Adaptive deep metric learning for identity-aware facial expression recognition," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshops (CVPRW), 2017, pp. 522–531.
  78. 78.G. Levi and T. Hassner, "Emotion recognition in the wild via convolutional neural networks and mapped binary patterns," in Proceedings of the 2015 ACM on international conference on multimodal interaction. ACM, 2015, pp. 503–510.
  79. 79.D. A. Pitaloka, A. Wulandari, T. Basaruddin, and D. Y. Liliana, "Enhancing cnn with preprocessing stage in automatic emotion recognition," Procedia Computer Science, vol. 116, pp. 523–529, 2017.
  80. 80.A. T. Lopes, E. de Aguiar, A. F. De Souza, and T. Oliveira-Santos, "Facial expression recognition with convolutional neural networks: coping with few data and the training sample order," Pattern Recognition, vol. 61, pp. 610–628, 2017.
  81. 81.M. V. Zavarez, R. F. Berriel, and T. Oliveira-Santos, "Cross-database facial expression recognition based on fine-tuned deep convolutional network," in Graphics, Patterns and Images (SIBGRAPI), 2017 30th SIBGRAPI Conference on. IEEE, 2017, pp. 405–412.
  82. 82.W. Li, M. Li, Z. Su, and Z. Zhu, "A deep-learning approach to facial expression recognition with candid images," in Machine Vision Applications (MVA), 2015 14th IAPR International Conference on. IEEE, 2015, pp. 279–282.
  83. 83.I. Abbasnejad, S. Sridharan, D. Nguyen, S. Denman, C. Fookes, and S. Lucey, "Using synthetic data to improve facial expression analysis with 3d convolutional networks," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 1609–1618.
  84. 84.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, "Generative adversarial nets," in Advances in neural information processing systems, 2014, pp. 2672–2680.
  85. 85.W. Chen, M. J. Er, and S. Wu, "Illumination compensation and normalization for robust face recognition using discrete cosine transform in logarithm domain," IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 36, no. 2, pp. 458–466, 2006.
  86. 86.J. Li and E. Y. Lam, "Facial expression recognition using deep neural networks," in Imaging Systems and Techniques (IST), 2015 IEEE International Conference on. IEEE, 2015, pp. 1–6.
  87. 87.S. Ebrahimi Kahou, V. Michalski, K. Konda, R. Memisevic, and C. Pal, "Recurrent neural networks for emotion recognition in video," in Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, 2015, pp. 467–474.
  88. 88.S. A. Bargal, E. Barsoum, C. C. Ferrer, and C. Zhang, "Emotion recognition in the wild from videos using images," in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 433–436.
  89. 89.C.-M. Kuo, S.-H. Lai, and M. Sarkis, "A compact deep learning model for robust facial expression recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018, pp. 2121–2129.
  90. 90.A. Yao, D. Cai, P. Hu, S. Wang, L. Sha, and Y. Chen, "Holonet: towards robust emotion recognition in the wild," in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 472–478.
  91. 91.P. Hu, D. Cai, S. Wang, A. Yao, and Y. Chen, "Learning supervised scoring ensemble for emotion recognition in the wild," in Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, 2017, pp. 553–560.
  92. 92.T. Hassner, S. Harel, E. Paz, and R. Enbar, "Effective face frontalization in unconstrained images," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 4295–4304.
  93. 93.C. Sagonas, Y. Panagakis, S. Zafeiriou, and M. Pantic, "Robust statistical face frontalization," in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 3871–3879.
  94. 94.X. Yin, X. Yu, K. Sohn, X. Liu, and M. Chandraker, "Towards largepose face frontalization in the wild," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 3990–3999.
  95. 95.R. Huang, S. Zhang, T. Li, and R. He, "Beyond face rotation: Global and local perception gan for photorealistic and identity preserving frontal view synthesis," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 2439–2448.
  96. 96.L. Tran, X. Yin, and X. Liu, "Disentangled representation learning gan for pose-invariant face recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 1415–1424.
  97. 97.L. Deng, D. Yu et al., "Deep learning: methods and applications," Foundations and Trends® in Signal Processing, vol. 7, no. 3–4, pp. 197–387, 2014.
  98. 98.B. Fasel, "Robust face analysis using convolutional neural networks," in Pattern Recognition, 2002. Proceedings. 16th International Conference on, vol. 2. IEEE, 2002, pp. 40–43.
  99. 99.——, "Head-pose invariant facial expression recognition using convolutional neural networks," in Proceedings of the 4th IEEE International Conference on Multimodal Interfaces. IEEE Computer Society, 2002, p. 529.
  100. 100.M. Matsugu, K. Mori, Y. Mitari, and Y. Kaneda, "Subject independent facial expression recognition with robust face detection using a convolutional neural network," Neural Networks, vol. 16, no. 5-6, pp. 555–559, 2003.
  101. 101.B. Sun, L. Li, G. Zhou, X. Wu, J. He, L. Yu, D. Li, and Q. Wei, "Combining multimodal features within a fusion network for emotion recognition in the wild," in Proceedings of the 2015 ACM on International Conference on Multimodal Interaction. ACM, 2015, pp. 497–502.
  102. 102.B. Sun, L. Li, G. Zhou, and J. He, "Facial expression recognition in the wild based on multimodal texture features," Journal of Electronic Imaging, vol. 25, no. 6, p. 061407, 2016.
  103. 103.R. Girshick, J. Donahue, T. Darrell, and J. Malik, "Rich feature hierarchies for accurate object detection and semantic segmentation," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 580–587.
  104. 104.J. Li, D. Zhang, J. Zhang, J. Zhang, T. Li, Y. Xia, Q. Yan, and L. Xun, "Facial expression recognition with faster r-cnn," Procedia Computer Science, vol. 107, pp. 135–140, 2017.
  105. 105.S. Ren, K. He, R. Girshick, and J. Sun, "Faster r-cnn: Towards real-time object detection with region proposal networks," in Advances in neural information processing systems, 2015, pp. 91–99.
  106. 106.S. Ji, W. Xu, M. Yang, and K. Yu, "3d convolutional neural networks for human action recognition," IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 1, pp. 221–231, 2013.
  107. 107.D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, "Learning spatiotemporal features with 3d convolutional networks," in Computer Vision (ICCV), 2015 IEEE International Conference on. IEEE, 2015, pp. 4489–4497.
  108. 108.Y. Fan, X. Lu, D. Li, and Y. Liu, "Video-based emotion recognition using cnn-rnn and c3d hybrid networks," in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 445–450.
  109. 109.D. Nguyen, K. Nguyen, S. Sridharan, A. Ghasemi, D. Dean, and C. Fookes, "Deep spatio-temporal features for multimodal emotion recognition," in Applications of Computer Vision (WACV), 2017 IEEE Winter Conference on. IEEE, 2017, pp. 1215–1223.
  110. 110.S. Ouellet, "Real-time emotion recognition for gaming using deep convolutional network features," arXiv preprint arXiv:1408.3750, 2014.
  111. 111.H. Ding, S. K. Zhou, and R. Chellappa, "Facenet2expnet: Regularizing a deep face recognition net for expression recognition," in Automatic Face & Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on. IEEE, 2017, pp. 118–126.
  112. 112.B. Hasani and M. H. Mahoor, "Facial expression recognition using enhanced deep 3d convolutional neural networks," in Computer Vision and Pattern Recognition Workshops (CVPRW), 2017 IEEE Conference on. IEEE, 2017, pp. 2278–2288.
  113. 113.G. E. Hinton, S. Osindero, and Y.-W. Teh, "A fast learning algorithm for deep belief nets," Neural computation, vol. 18, no. 7, pp. 1527–1554, 2006.
  114. 114.G. E. Hinton and T. J. Sejnowski, "Learning and releaming in boltzmann machines," Parallel distributed processing: Explorations in the microstructure of cognition, vol. 1, no. 282-317, p. 2, 1986.
  115. 115.G. E. Hinton, "A practical guide to training restricted boltzmann machines," in Neural networks: Tricks of the trade. Springer, 2012, pp. 599–619.
  116. 116.Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle, "Greedy layerwise training of deep networks," in Advances in neural information processing systems, 2007, pp. 153–160.
  117. 117.G. E. Hinton, "Training products of experts by minimizing contrastive divergence," Neural computation, vol. 14, no. 8, pp. 1771–1800, 2002.
  118. 118.G. E. Hinton and R. R. Salakhutdinov, "Reducing the dimensionality of data with neural networks," science, vol. 313, no. 5786, pp. 504–507, 2006.
  119. 119.P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P.-A. Manzagol, "Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion," Journal of Machine Learning Research, vol. 11, no. Dec, pp. 3371–3408, 2010.
  120. 120.Q. V. Le, "Building high-level features using large scale unsupervised learning," in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on. IEEE, 2013, pp. 8595–8598.
  121. 121.S. Rifai, P. Vincent, X. Muller, X. Glorot, and Y. Bengio, "Contractive auto-encoders: Explicit invariance during feature extraction," in Proceedings of the 28th International Conference on International Conference on Machine Learning. Omnipress, 2011, pp. 833–840.
  122. 122.J. Masci, U. Meier, D. Cireşan, and J. Schmidhuber, "Stacked convolutional auto-encoders for hierarchical feature extraction," in International Conference on Artificial Neural Networks. Springer, 2011, pp. 52–59.
  123. 123.D. P. Kingma and M. Welling, "Auto-encoding variational bayes," arXiv preprint arXiv:1312.6114, 2013.
  124. 124.P. J. Werbos, "Backpropagation through time: what it does and how to do it," Proceedings of the IEEE, vol. 78, no. 10, pp. 1550–1560, 1990.
  125. 125.S. Hochreiter and J. Schmidhuber, "Long short-term memory," Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997.
  126. 126.M. Mirza and S. Osindero, "Conditional generative adversarial nets," arXiv preprint arXiv:1411.1784, 2014.
  127. 127.A. Radford, L. Metz, and S. Chintala, "Unsupervised representation learning with deep convolutional generative adversarial networks," arXiv preprint arXiv:1511.06434, 2015.
  128. 128.A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther, "Autoencoding beyond pixels using a learned similarity metric," arXiv preprint arXiv:1512.09300, 2015.
  129. 129.X. Chen, Y. Duan, R. Houthooft, J. Schulman, I. Sutskever, and P. Abbeel, "Infogan: Interpretable representation learning by information maximizing generative adversarial nets," in Advances in neural information processing systems, 2016, pp. 2172–2180.
  130. 130.Y. Tang, "Deep learning using linear support vector machines," arXiv preprint arXiv:1306.0239, 2013.
  131. 131.A. Dapogny and K. Bailly, "Investigating deep neural forests for facial expression recognition," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 629–633.
  132. 132.P. Kontschieder, M. Fiterau, A. Criminisi, and S. Rota Bulo, "Deep neural decision forests," in Proceedings of the IEEE international conference on computer vision, 2015, pp. 1467–1475.
  133. 133.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell, "Decaf: A deep convolutional activation feature for generic visual recognition," in International conference on machine learning, 2014, pp. 647–655.
  134. 134.A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson, "Cnn features off-the-shelf: an astounding baseline for recognition," in Computer Vision and Pattern Recognition Workshops (CVPRW), 2014 IEEE Conference on. IEEE, 2014, pp. 512–519.
  135. 135.N. Otberdout, A. Kacem, M. Daoudi, L. Ballihi, and S. Berretti, "Deep covariance descriptors for facial expression recognition," in BMVC, 2018.
  136. 136.D. Acharya, Z. Huang, D. Pani Paudel, and L. Van Gool, "Covariance pooling for facial expression recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018, pp. 367–374.
  137. 137.M. Liu, S. Li, S. Shan, and X. Chen, "Au-aware deep networks for facial expression recognition," in Automatic Face and Gesture Recognition (FG), 2013 10th IEEE International Conference and Workshops on. IEEE, 2013, pp. 1–6.
  138. 138.——, "Au-inspired deep networks for facial expression feature learning," Neurocomputing, vol. 159, pp. 126–136, 2015.
  139. 139.P. Khorrami, T. Paine, and T. Huang, "Do deep neural networks learn facial action units when doing expression recognition?" arXiv preprint arXiv:1510.02969v3, 2015.
  140. 140.J. Cai, Z. Meng, A. S. Khan, Z. Li, J. OReilly, and Y. Tong, "Island loss for learning discriminative features in facial expression recognition," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 302–309.
  141. 141.H. Yang, U. Ciftci, and L. Yin, "Facial expression recognition by deexpression residue learning," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 2168–2177.
  142. 142.D. Hamester, P. Barros, and S. Wermter, "Face expression recognition with a 2-channel convolutional neural network," in Neural Networks (IJCNN), 2015 International Joint Conference on. IEEE, 2015, pp. 1–8.
  143. 143.S. Reed, K. Sohn, Y. Zhang, and H. Lee, "Learning to disentangle factors of variation with manifold interaction," in International Conference on Machine Learning, 2014, pp. 1431–1439.
  144. 144.Z. Zhang, P. Luo, C.-C. Loy, and X. Tang, "Learning social relation traits from face images," in Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 3631–3639.
  145. 145.Y. Guo, D. Tao, J. Yu, H. Xiong, Y. Li, and D. Tao, "Deep neural networks with relativity learning for facial expression recognition," in Multimedia & Expo Workshops (ICMEW), 2016 IEEE International Conference on. IEEE, 2016, pp. 1–6.
  146. 146.B.-K. Kim, S.-Y. Dong, J. Roh, G. Kim, and S.-Y. Lee, "Fusing aligned and non-aligned face information for automatic affect recognition in the wild: A deep learning approach," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2016, pp. 48–57.
  147. 147.C. Pramerdorfer and M. Kampel, "Facial expression recognition using convolutional neural networks: State of the art," arXiv preprint arXiv:1612.02903, 2016.
  148. 148.O. M. Parkhi, A. Vedaldi, A. Zisserman et al., "Deep face recognition." in BMVC, vol. 1, no. 3, 2015, p. 6.
  149. 149.T. Kaneko, K. Hiramatsu, and K. Kashino, "Adaptive visual feedback generation for facial expression improvement with multi-task deep neural networks," in Proceedings of the 2016 ACM on Multimedia Conference. ACM, 2016, pp. 327–331.
  150. 150.D. Yi, Z. Lei, S. Liao, and S. Z. Li, "Learning face representation from scratch," arXiv preprint arXiv:1411.7923, 2014.
  151. 151.X. Zhang, L. Zhang, X.-J. Wang, and H.-Y. Shum, "Finding celebrities in billions of web images," IEEE Transactions on Multimedia, vol. 14, no. 4, pp. 995–1007, 2012.
  152. 152.H.-W. Ng and S. Winkler, "A data-driven approach to cleaning large face datasets," in Image Processing (ICIP), 2014 IEEE International Conference on. IEEE, 2014, pp. 343–347.
  153. 153.H. Kaya, F. Gürpınar, and A. A. Salah, "Video-based emotion recognition in the wild using deep transfer learning and score fusion," Image and Vision Computing, vol. 65, pp. 66–75, 2017.
  154. 154.B. Knyazev, R. Shvetsov, N. Efremova, and A. Kuharenko, "Convolutional neural networks pretrained on large face recognition datasets for emotion classification from video," arXiv preprint arXiv:1711.04598, 2017.
  155. 155.D. G. Lowe, "Object recognition from local scale-invariant features," in Computer vision, 1999. The proceedings of the seventh IEEE international conference on, vol. 2. Ieee, 1999, pp. 1150–1157.
  156. 156.T. Zhang, W. Zheng, Z. Cui, Y. Zong, J. Yan, and K. Yan, "A deep neural network-driven feature learning method for multi-view facial expression recognition," IEEE Transactions on Multimedia, vol. 18, no. 12, pp. 2528–2536, 2016.
  157. 157.Z. Luo, J. Chen, T. Takiguchi, and Y. Ariki, "Facial expression recognition with deep age," in Multimedia & Expo Workshops (ICMEW), 2017 IEEE International Conference on. IEEE, 2017, pp. 657–662.
  158. 158.L. Chen, M. Zhou, W. Su, M. Wu, J. She, and K. Hirota, "Softmax regression based deep sparse autoencoder network for facial emotion recognition in human-robot interaction," Information Sciences, vol. 428, pp. 49–61, 2018.
  159. 159.V. Mavani, S. Raman, and K. P. Miyapuram, "Facial expression recognition using visual saliency and deep learning," arXiv preprint arXiv:1708.08016, 2017.
  160. 160.M. Cornia, L. Baraldi, G. Serra, and R. Cucchiara, "A deep multi-level network for saliency prediction," in Pattern Recognition (ICPR), 2016 23rd International Conference on. IEEE, 2016, pp. 3488–3493.
  161. 161.B.-F. Wu and C.-H. Lin, "Adaptive feature mapping for customizing deep learning based facial expression recognition model," IEEE Access, 2018.
  162. 162.J. Lu, V. E. Liong, and J. Zhou, "Cost-sensitive local binary feature learning for facial age estimation," IEEE Transactions on Image Processing, vol. 24, no. 12, pp. 5356–5368, 2015.
  163. 163.W. Shang, K. Sohn, D. Almeida, and H. Lee, "Understanding and improving convolutional neural networks via concatenated rectified linear units," in International Conference on Machine Learning, 2016, pp. 2217–2225.
  164. 164.C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, "Rethinking the inception architecture for computer vision," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2818–2826.
  165. 165.C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, "Inception-v4, inception-resnet and the impact of residual connections on learning." in AAAI, vol. 4, 2017, p. 12.
  166. 166.S. Zhao, H. Cai, H. Liu, J. Zhang, and S. Chen, "Feature selection mechanism in cnns for facial expression recognition," in BMVC, 2018.
  167. 167.J. Zeng, S. Shan, and X. Chen, "Facial expression recognition with inconsistently annotated datasets," in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 222–237.
  168. 168.Y. Wen, K. Zhang, Z. Li, and Y. Qiao, "A discriminative feature learning approach for deep face recognition," in European Conference on Computer Vision. Springer, 2016, pp. 499–515.
  169. 169.F. Schroff, D. Kalenichenko, and J. Philbin, "Facenet: A unified embedding for face recognition and clustering," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 815–823.
  170. 170.G. Zeng, J. Zhou, X. Jia, W. Xie, and L. Shen, "Hand-crafted feature guided deep learning for facial expression recognition," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 423–430.
  171. 171.D. Ciregan, U. Meier, and J. Schmidhuber, "Multi-column deep neural networks for image classification," in Computer vision and pattern recognition (CVPR), 2012 IEEE conference on. IEEE, 2012, pp. 3642–3649.
  172. 172.G. Pons and D. Masip, "Supervised committee of convolutional neural networks in automated facial expression analysis," IEEE Transactions on Affective Computing, 2017.
  173. 173.B.-K. Kim, J. Roh, S.-Y. Dong, and S.-Y. Lee, "Hierarchical committee of deep convolutional neural networks for robust facial expression recognition," Journal on Multimodal User Interfaces, vol. 10, no. 2, pp. 173–189, 2016.
  174. 174.K. Liu, M. Zhang, and Z. Pan, "Facial expression recognition with cnn ensemble," in Cyberworlds (CW), 2016 International Conference on. IEEE, 2016, pp. 163–166.
  175. 175.G. Pons and D. Masip, "Multi-task, multi-label and multi-domain learning with residual convolutional networks for emotion recognition," arXiv preprint arXiv:1802.06664, 2018.
  176. 176.P. Ekman and E. L. Rosenberg, What the face reveals: Basic and applied studies of spontaneous expression using the Facial Action Coding System (FACS). Oxford University Press, USA, 1997.
  177. 177.R. Ranjan, S. Sankaranarayanan, C. D. Castillo, and R. Chellappa, "An all-in-one convolutional neural network for face analysis," in Automatic Face & Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on. IEEE, 2017, pp. 17–24.
  178. 178.Y. Lv, Z. Feng, and C. Xu, "Facial expression recognition via deep learning," in Smart Computing (SMARTCOMP), 2014 International Conference on. IEEE, 2014, pp. 303–308.
  179. 179.S. Rifai, Y. Bengio, A. Courville, P. Vincent, and M. Mirza, "Disentangling factors of variation for facial expression recognition," in European Conference on Computer Vision. Springer, 2012, pp. 808–822.
  180. 180.Y.-H. Lai and S.-H. Lai, "Emotion-preserving representation learning via generative adversarial network for multi-view facial expression recognition," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 263–270.
  181. 181.F. Zhang, T. Zhang, Q. Mao, and C. Xu, "Joint pose and expression modeling for facial expression recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 3359–3368.
  182. 182.H. Yang, Z. Zhang, and L. Yin, "Identity-adaptive facial expression recognition through expression regeneration using conditional generative adversarial networks," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 294–301.
  183. 183.J. Chen, J. Konrad, and P. Ishwar, "Vgan-based image representation learning for privacy-preserving facial expression recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2018, pp. 1570–1579.
  184. 184.Y. Kim, B. Yoo, Y. Kwak, C. Choi, and J. Kim, "Deep generativecontrastive networks for facial expression recognition," arXiv preprint arXiv:1703.07140, 2017.
  185. 185.N. Sun, Q. Li, R. Huan, J. Liu, and G. Han, "Deep spatial-temporal feature fusion for facial expression recognition in static images," Pattern Recognition Letters, 2017.
  186. 186.W. Ding, M. Xu, D. Huang, W. Lin, M. Dong, X. Yu, and H. Li, "Audio and face video emotion recognition in the wild using deep neural networks and small datasets," in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 506–513.
  187. 187.J. Yan, W. Zheng, Z. Cui, C. Tang, T. Zhang, Y. Zong, and N. Sun, "Multi-clue fusion for emotion recognition in the wild," in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 458–463.
  188. 188.Z. Cui, S. Xiao, Z. Niu, S. Yan, and W. Zheng, "Recurrent shape regression," IEEE Transactions on Pattern Analysis and Machine Intelligence, 2018.
  189. 189.X. Ouyang, S. Kawaai, E. G. H. Goh, S. Shen, W. Ding, H. Ming, and D.-Y. Huang, "Audio-visual emotion recognition using deep transfer learning and multiple temporal models," in Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, 2017, pp. 577–582.
  190. 190.V. Vielzeuf, S. Pateux, and F. Jurie, "Temporal multimodal fusion for video emotion classification in the wild," in Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, 2017, pp. 569–576.
  191. 191.S. E. Kahou, X. Bouthillier, P. Lamblin, C. Gulcehre, V. Michalski, K. Konda, S. Jean, P. Froumenty, Y. Dauphin, N. BoulangerLewandowski et al., "Emonets: Multimodal deep learning approaches for emotion recognition in video," Journal on Multimodal User Interfaces, vol. 10, no. 2, pp. 99–111, 2016.
  192. 192.M. Liu, R. Wang, S. Li, S. Shan, Z. Huang, and X. Chen, "Combining multiple kernel methods on riemannian manifold for emotion recognition in the wild," in Proceedings of the 16th International Conference on Multimodal Interaction. ACM, 2014, pp. 494–501.
  193. 193.B. Xu, Y. Fu, Y.-G. Jiang, B. Li, and L. Sigal, "Video emotion recognition with transferred deep feature encodings," in Proceedings of the 2016 ACM on International Conference on Multimedia Retrieval. ACM, 2016, pp. 15–22.
  194. 194.J. Chen, R. Xu, and L. Liu, "Deep peak-neutral difference feature for facial expression recognition," Multimedia Tools and Applications, pp. 1–17, 2018.
  195. 195.Q. V. Le, N. Jaitly, and G. E. Hinton, "A simple way to initialize recurrent networks of rectified linear units," arXiv preprint arXiv:1504.00941, 2015.
  196. 196.M. Schuster and K. K. Paliwal, "Bidirectional recurrent neural networks," IEEE Transactions on Signal Processing, vol. 45, no. 11, pp. 2673–2681, 1997.
  197. 197.P. Barros and S. Wermter, "Developing crossmodal expression recognition based on a deep neural model," Adaptive behavior, vol. 24, no. 5, pp. 373–396, 2016.
  198. 198.J. Zhao, X. Mao, and J. Zhang, "Learning deep facial expression features from image and optical flow sequences using 3d cnn," The Visual Computer, pp. 1–15, 2018.
  199. 199.M. Liu, S. Li, S. Shan, R. Wang, and X. Chen, "Deeply learning deformable facial action parts model for dynamic expression analysis," in Asian conference on computer vision. Springer, 2014, pp. 143–157.
  200. 200.P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan, "Object detection with discriminatively trained part-based models," IEEE transactions on pattern analysis and machine intelligence, vol. 32, no. 9, pp. 1627–1645, 2010.
  201. 201.S. Pini, O. B. Ahmed, M. Cornia, L. Baraldi, R. Cucchiara, and B. Huet, "Modeling multimodal cues in a deep learning-based framework for emotion recognition in the wild," in Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, 2017, pp. 536–543.
  202. 202.R. Arandjelovic, P. Gronat, A. Torii, T. Pajdla, and J. Sivic, "Netvlad: Cnn architecture for weakly supervised place recognition," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 5297–5307.
  203. 203.D. H. Kim, M. K. Lee, D. Y. Choi, and B. C. Song, "Multi-modal emotion recognition using semi-supervised learning and multiple neural networks in the wild," in Proceedings of the 19th ACM International Conference on Multimodal Interaction. ACM, 2017, pp. 529–535.
  204. 204.J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, "Long-term recurrent convolutional networks for visual recognition and description," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, pp. 2625–2634.
  205. 205.D. K. Jain, Z. Zhang, and K. Huang, "Multi angle optimal pattern-based deep learning for automatic facial expression recognition," Pattern Recognition Letters, 2017.
  206. 206.M. Baccouche, F. Mamalet, C. Wolf, C. Garcia, and A. Baskurt, "Spatio-temporal convolutional sparse auto-encoder for sequence classification." in BMVC, 2012, pp. 1–12.
  207. 207.S. Kankanamge, C. Fookes, and S. Sridharan, "Facial analysis in the wild with lstm networks," in Image Processing (ICIP), 2017 IEEE International Conference on. IEEE, 2017, pp. 1052–1056.
  208. 208.J. D. Lafferty, A. Mccallum, and F. C. N. Pereira, "Conditional random fields: Probabilistic models for segmenting and labeling sequence data," Proceedings of Icml, vol. 3, no. 2, pp. 282–289, 2001.
  209. 209.K. Simonyan and A. Zisserman, "Two-stream convolutional networks for action recognition in videos," in Advances in neural information processing systems, 2014, pp. 568–576.
  210. 210.J. Susskind, V. Mnih, G. Hinton et al., "On deep generative models with applications to recognition," in Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on. IEEE, 2011, pp. 2857–2864.
  211. 211.V. Mnih, J. M. Susskind, G. E. Hinton et al., "Modeling natural images using gated mrfs," IEEE transactions on pattern analysis and machine intelligence, vol. 35, no. 9, pp. 2206–2222, 2013.
  212. 212.V. Mnih, G. E. Hinton et al., "Generating more realistic images using gated mrf's," in Advances in Neural Information Processing Systems, 2010, pp. 2002–2010.
  213. 213.Y. Cheng, B. Jiang, and K. Jia, "A deep structure for facial expression recognition under partial occlusion," in Intelligent Information Hiding and Multimedia Signal Processing (IIH-MSP), 2014 Tenth International Conference on. IEEE, 2014, pp. 211–214.
  214. 214.M. Xu, W. Cheng, Q. Zhao, L. Ma, and F. Xu, "Facial expression recognition based on transfer learning from deep convolutional networks," in Natural Computation (ICNC), 2015 11th International Conference on. IEEE, 2015, pp. 702–708.
  215. 215.Y. Liu, J. Zeng, S. Shan, and Z. Zheng, "Multi-channel pose-aware convolution neural networks for multi-view facial expression recognition," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 458–465.
  216. 216.S. He, S. Wang, W. Lan, H. Fu, and Q. Ji, "Facial expression recognition using deep boltzmann machine from thermal infrared images," in Affective Computing and Intelligent Interaction (ACII), 2013 Humaine Association Conference on. IEEE, 2013, pp. 239–244.
  217. 217.Z. Wu, T. Chen, Y. Chen, Z. Zhang, and G. Liu, "Nirexpnet: Threestream 3d convolutional neural network for near infrared facial expression recognition," Applied Sciences, vol. 7, no. 11, p. 1184, 2017.
  218. 218.E. P. Ijjina and C. K. Mohan, "Facial expression recognition using kinect depth sensor and convolutional neural networks," in Machine Learning and Applications (ICMLA), 2014 13th International Conference on. IEEE, 2014, pp. 392–396.
  219. 219.M. Z. Uddin, M. M. Hassan, A. Almogren, M. Zuair, G. Fortino, and J. Torresen, "A facial expression recognition system using robust face features from depth videos and deep learning," Computers & Electrical Engineering, vol. 63, pp. 114–125, 2017.
  220. 220.M. Z. Uddin, W. Khaksar, and J. Torresen, "Facial expression recognition using salient features and convolutional neural network," IEEE Access, vol. 5, pp. 26 146–26 161, 2017.
  221. 221.W. Li, D. Huang, H. Li, and Y. Wang, "Automatic 4d facial expression recognition using dynamic geometrical image network," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 24–30.
  222. 222.F.-J. Chang, A. T. Tran, T. Hassner, I. Masi, R. Nevatia, and G. Medioni, "Expnet: Landmark-free, deep, 3d facial expressions," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 122–129.
  223. 223.O. K. Oyedotun, G. Demisse, A. E. R. Shabayek, D. Aouada, and B. Ottersten, "Facial expression recognition via joint deep learning of rgb-depth map latent representations," in 2017 IEEE International Conference on Computer Vision Workshop (ICCVW), 2017.
  224. 224.H. Li, J. Sun, Z. Xu, and L. Chen, "Multimodal 2d+ 3d facial expression recognition with deep fusion convolutional neural network," IEEE Transactions on Multimedia, vol. 19, no. 12, pp. 2816–2831, 2017.
  225. 225.A. Jan, H. Ding, H. Meng, L. Chen, and H. Li, "Accurate facial parts localization and deep learning for 3d facial expression recognition," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 466–472.
  226. 226.X. Wei, H. Li, J. Sun, and L. Chen, "Unsupervised domain adaptation with regularized optimal transport for multimodal 2d+ 3d facial expression recognition," in Automatic Face & Gesture Recognition (FG 2018), 2018 13th IEEE International Conference on. IEEE, 2018, pp. 31–37.
  227. 227.J. M. Susskind, G. E. Hinton, J. R. Movellan, and A. K. Anderson, "Generating facial expressions with deep belief nets," in Affective Computing. InTech, 2008.
  228. 228.M. Sabzevari, S. Toosizadeh, S. R. Quchani, and V. Abrishami, "A fast and accurate facial expression synthesis system for color face images using face graph and deep belief network," in Electronics and Information Engineering (ICEIE), 2010 International Conference On, vol. 2. IEEE, 2010, pp. V2–354.
  229. 229.R. Yeh, Z. Liu, D. B. Goldman, and A. Agarwala, "Semantic facial expression editing using autoencoded flow," arXiv preprint arXiv:1611.09961, 2016.
  230. 230.Y. Zhou and B. E. Shi, "Photorealistic facial expression synthesis by the conditional difference adversarial autoencoder," in Affective Computing and Intelligent Interaction (ACII), 2017 Seventh International Conference on. IEEE, 2017, pp. 370–376.
  231. 231.L. Song, Z. Lu, R. He, Z. Sun, and T. Tan, "Geometry guided adversarial facial expression synthesis," arXiv preprint arXiv:1712.03474, 2017.
  232. 232.H. Ding, K. Sricharan, and R. Chellappa, "Exprgan: Facial expression editing with controllable expression intensity," in AAAI, 2018, p. 67816788.
  233. 233.F. Qiao, N. Yao, Z. Jiao, Z. Li, H. Chen, and H. Wang, "Geometry-contrastive generative adversarial network for facial expression synthesis," arXiv preprint arXiv:1802.01822, 2018.
  234. 234.I. Masi, A. T. Tran, T. Hassner, J. T. Leksut, and G. Medioni, "Do we really need to collect millions of faces for effective face recognition?" in European Conference on Computer Vision. Springer, 2016, pp. 579–596.
  235. 235.N. Mousavi, H. Siqueira, P. Barros, B. Fernandes, and S. Wermter, "Understanding how deep neural networks learn face expressions," in Neural Networks (IJCNN), 2016 International Joint Conference on. IEEE, 2016, pp. 227–234.
  236. 236.R. Breuer and R. Kimmel, "A deep learning perspective on the origin of facial expressions," arXiv preprint arXiv:1705.01842, 2017.
  237. 237.M. D. Zeiler and R. Fergus, "Visualizing and understanding convolutional networks," in European conference on computer vision. Springer, 2014, pp. 818–833.
  238. 238.I. Lusi, J. C. J. Junior, J. Gorbova, X. Baró, S. Escalera, H. Demirel, J. Allik, C. Ozcinar, and G. Anbarjafari, "Joint challenge on dominant and complementary emotion recognition using micro emotion features and head-pose estimation: Databases," in Automatic Face & Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on. IEEE, 2017, pp. 809–813.
  239. 239.J. Wan, S. Escalera, X. Baro, H. J. Escalante, I. Guyon, M. Madadi, J. Allik, J. Gorbova, and G. Anbarjafari, "Results and analysis of chalearn lap multi-modal isolated and continuous gesture recognition, and real versus fake expressed emotions challenges," in ChaLearn LaP, Action, Gesture, and Emotion Recognition Workshop and Competitions: Large Scale Multimodal Gesture Recognition and Real versus Fake expressed emotions, ICCV, vol. 4, no. 6, 2017.
  240. 240.Y.-G. Kim and X.-P. Huynh, "Discrimination between genuine versus fake emotion using long-short term memory with parametric bias and facial landmarks," in Computer Vision Workshop (ICCVW), 2017 IEEE International Conference on. IEEE, 2017, pp. 3065–3072.
  241. 241.L. Li, T. Baltrusaitis, B. Sun, and L.-P. Morency, "Combining sequential geometry and texture features for distinguishing genuine and deceptive emotions," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 3147–3153.
  242. 242.J. Guo, S. Zhou, J. Wu, J. Wan, X. Zhu, Z. Lei, and S. Z. Li, "Multimodality network with visual and geometrical information for micro emotion recognition," in Automatic Face & Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on. IEEE, 2017, pp. 814–819.
  243. 243.I. Song, H.-J. Kim, and P. B. Jeon, "Deep learning for real-time robust facial expression recognition on a smartphone," in Consumer Electronics (ICCE), 2014 IEEE International Conference on. IEEE, 2014, pp. 564–567.
  244. 244.S. Bazrafkan, T. Nedelcu, P. Filipczuk, and P. Corcoran, "Deep learning for facial expression recognition: A step closer to a smartphone that knows your moods," in Consumer Electronics (ICCE), 2017 IEEE International Conference on. IEEE, 2017, pp. 217–220.
  245. 245.S. Hickson, N. Dufour, A. Sud, V. Kwatra, and I. Essa, "Eyemotion: Classifying facial expressions in vr using eye-tracking cameras," arXiv preprint arXiv:1707.07204, 2017.
  246. 246.S. A. Ossia, A. S. Shamsabadi, A. Taheri, H. R. Rabiee, N. Lane, and H. Haddadi, "A hybrid deep learning architecture for privacy-preserving mobile analytics," arXiv preprint arXiv:1703.02952, 2017.
  247. 247.K. Kulkarni, C. A. Corneanu, I. Ofodile, S. Escalera, X. Baro, S. Hyniewska, J. Allik, and G. Anbarjafari, "Automatic recognition of facial displays of unfelt emotions," arXiv preprint arXiv:1707.04061, 2017.
  248. 248.X. Zhou, K. Jin, Y. Shang, and G. Guo, "Visually interpretable representation learning for depression recognition from facial images," IEEE Transactions on Affective Computing, pp. 1–1, 2018.
  249. 249.E. Barsoum, C. Zhang, C. C. Ferrer, and Z. Zhang, "Training deep networks for facial expression recognition with crowd-sourced label distribution," in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 279–283.
  250. 250.J. A. Russell, "A circumplex model of affect." Journal of personality and social psychology, vol. 39, no. 6, p. 1161, 1980.
  251. 251.S. Li and W. Deng, "Deep emotion transfer network for cross-database facial expression recognition," in Pattern Recognition (ICPR), 2018 26th International Conference. IEEE, 2018, pp. 3092–3099.
  252. 252.M. Valstar, J. Gratch, B. Schuller, F. Ringeval, D. Lalanne, M. Torres Torres, S. Scherer, G. Stratou, R. Cowie, and M. Pantic, "Avec 2016: Depression, mood, and emotion recognition workshop and challenge," in Proceedings of the 6th International Workshop on Audio/Visual Emotion Challenge. ACM, 2016, pp. 3–10.
  253. 253.F. Ringeval, B. Schuller, M. Valstar, J. Gratch, R. Cowie, S. Scherer, S. Mozgai, N. Cummins, M. Schmitt, and M. Pantic, "Avec 2017: Real-life depression, and affect recognition workshop and challenge," in Proceedings of the 7th Annual Workshop on Audio/Visual Emotion Challenge. ACM, 2017, pp. 3–9.

Citation

MLA
Li, S., and W. Deng. “Deep Facial Expression Recognition: A Survey”. IEEE Transactions on Affective Computing, vol. 13, no. 3, 2022, pp. 1195–215, https://doi.org/10.1109/TAFFC.2020.2981446.
APA
Li, S., & Deng, W. (2022). Deep Facial Expression Recognition: A Survey. IEEE Transactions on Affective Computing, 13(3), 1195–1215. https://doi.org/10.1109/TAFFC.2020.2981446
Chicago
Li, S., and W. Deng. 2022. “Deep Facial Expression Recognition: A Survey”. IEEE Transactions on Affective Computing 13 (3): 1195–1215. https://doi.org/10.1109/TAFFC.2020.2981446.
Harvard
Li, S. and Deng, W. (2022) “Deep Facial Expression Recognition: A Survey”, IEEE Transactions on Affective Computing, 13(3), pp. 1195–1215. Available at: https://doi.org/10.1109/TAFFC.2020.2981446.
Vancouver
1. Li S, Deng W (2022) Deep Facial Expression Recognition: A Survey. IEEE Transactions on Affective Computing 13:1195–1215

BibTeX

@article{Li_2022, title={Deep Facial Expression Recognition: A Survey}, volume={13}, ISSN={2371-9850}, url={http://dx.doi.org/10.1109/TAFFC.2020.2981446}, DOI={10.1109/taffc.2020.2981446}, number={3}, journal={IEEE Transactions on Affective Computing}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Li, Shan and Deng, Weihong}, year={2022}, month=July, pages={1195–1215} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF