Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild

Shan LiWeihong DengJunping Du

article2017CVPR1,395 citations

Presents RAF-DB, a large-scale real-world facial expression database labeled through reliable crowdsourcing, alongside a deep locality-preserving CNN that significantly improves in-the-wild emotion recognition across basic and compound expressions.

Listen

Facial expression recognition is vital for modern human-computer interaction, affective computing, and social media analysis. However, most legacy systems rely on lab-controlled datasets where subjects display posed, stereotypical emotions under uniform conditions. In real-world environments, spontaneous expressions exhibit substantial ambiguity, wide variations in lighting and pose, and compound emotions that mix multiple affective states. The article aims to build a reliable, large-scale facial expression database reflecting real-world conditions and to develop a deep learning model capable of handling complex, multimodal emotion distributions.

To achieve this, the article introduces the Real-world Affective Faces Database (RAF-DB), comprising 29,672 facial images gathered from the internet. The images were labeled via crowdsourcing with 315 annotators, ensuring each image received approximately 40 independent evaluations. An Expectation-Maximization algorithm filtered out annotator noise and estimated label reliability, dividing the data into single-label basic emotions and two-label compound emotions. To address the significant intra-class variation of real-world faces, the article introduces a Deep Locality-Preserving Convolutional Neural Network (DLP-CNN). This architecture combines traditional classification loss with a locality-preserving loss to pull nearby samples of the same emotion class together in feature space, preserving natural intensity transitions.

Key findings demonstrate that real-world facial behaviors diverge markedly from lab-controlled settings. First, facial action units in the wild display much greater diversity; cross-database tests between RAF-DB and the lab-controlled CK+ dataset showed that a model trained on real-world data achieved 62% average accuracy on lab images, whereas a model trained on lab data dropped to 39% accuracy on real-world images. Second, baseline handcrafted visual descriptors struggled in the wild, dropping from roughly 88–92% accuracy on lab datasets to 56–65% on RAF-DB basic emotions and only 28–36% on compound emotions. Third, the proposed DLP-CNN significantly outperformed traditional methods and standard deep networks, achieving a benchmark score of 74.20% on basic emotions and 44.55% on compound emotions. Finally, features extracted by DLP-CNN generalized exceptionally well without fine-tuning, reaching 95.78% accuracy on CK+ and 51.05% on the challenging SFEW 2.0 dataset.

These results imply that relying on lab-trained emotion recognition models creates substantial operational risks and performance degradation when deployed in real-world systems. Real-world affective displays require architectures that accommodate compound emotions and multimodal distributions rather than forcing rigid, single-label categories. The article recommends that practitioners adopt large-scale, crowdsourced datasets like RAF-DB for benchmarking and employ locality-preserving deep learning architectures to improve feature discrimination.

Decision-makers should note that recognizing compound emotions remains difficult, with current baseline performances falling below 45% due to the scarcity of training samples per compound class. While confidence in the basic emotion detection capabilities of DLP-CNN is high, further research and data collection are needed to expand sample sizes for rare compound expressions before deploying fully autonomous emotion-recognition systems in critical applications.

  • Paper: Automatic Analysis of Facial Expressions: The State of the Art, Maja Pantic et al. (2000). This foundational survey establishes the structural limitations of laboratory-controlled facial expression recognition pipelines that the source paper directly aims to overcome in unconstrained real-world settings.
  • Paper: Recognizing Action Units for Facial Expression Analysis, Ying-li Tian et al. (2001). This paper establishes the classification of facial action units and combinations under controlled settings, providing essential foundational concepts for the source's analysis of action unit diversity in the wild.
  • Paper: Joint Face Detection and Alignment Using Multitask Cascaded Convolutional Networks, Kaipeng Zhang et al. (2016). This work introduces multi-task cascaded convolutional networks for face detection and landmark alignment, establishing key preprocessing and alignment practices essential for real-world facial expression recognition.
  • Paper: Facial Landmark Detection by Deep Multi-task Learning, Zhanpeng Zhang et al. (2014). This paper demonstrates deep multi-task learning for facial landmark detection under unconstrained conditions, providing essential context for handling wild variations in pose and expression.
  • Paper: Deep Learning Face Attributes in the Wild, Ziwei Liu et al. (2015). This work develops deep convolutional networks for recognizing facial attributes in unconstrained web images, pioneering feature learning methodologies for faces in the wild.
  • Paper: OpenFace: An open source facial behavior analysis toolkit, Tadas Baltrusaitis et al. (2016). This study introduces an open-source framework for real-time facial landmark detection, head pose estimation, and Action Unit recognition, establishing core tooling for facial behavior analysis.
  • Paper: Deep Face Recognition, Omkar M. Parkhi et al. (2015). This paper presents effective strategies for collecting, filtering, and training deep convolutional networks on large-scale web-retrieved face datasets, providing key architectural and dataset curation precedents.
  • Paper: Learning Face Representation from Scratch, Dong Yi et al. (2014). This foundational work demonstrates training deep representation models from scratch on large-scale web-harvested facial imagery.
Cover for Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild

Abstract

Past research on facial expressions have used relatively limited datasets, which makes it unclear whether current methods can be employed in real world. In this paper, we present a novel database, RAF-DB, which contains about 30000 facial images from thousands of individuals. Each image has been individually labeled about 40 times, then EM algorithm was used to filter out unreliable labels. Crowdsourcing reveals that real-world faces often express compound emotions, or even mixture ones. For all we know, RAF-DB is the first database that contains compound expressions in the wild. Our cross-database study shows that the action units of basic emotions in RAF-DB are much more diverse than, or even deviate from, those of lab-controlled ones. To address this problem, we propose a new DLP-CNN (Deep Locality-Preserving CNN) method, which aims to enhance the discriminative power of deep features by preserving the locality closeness while maximizing the inter-class scatters. The benchmark experiments on the 7-class basic expressions and 11-class compound expressions, as well as the additional experiments on SFEW and CK+ databases, show that the proposed DLP-CNN outperforms the state-of-the-art handcrafted features and deep learning based methods for the expression recognition in the wild.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Expression image datasets
  • 2.2. The framework for expression recognition
  • 2.3. Deep learning for expression recognition
  • 3. Real-world Expression Database: RAF-DB
  • 3.1. Creating RAF-DB
  • 3.2. CK+ and RAF Cross-Database Study
  • 4. Deep Locality-Preserving Feature Learning
  • 5. Baseline System
  • 6. Deep Learning System
  • 7. Conclusions and Future Work
  • 8. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Deep Locality-Preserving Loss for Facial Expression Recognition

    model/method

    To address significant intra-class variation and multi-modal class distributions in unconstrained facial expression recognition, Deep Locality-Preserving CNN (DLP-CNN) augments standard softmax loss with a Locality Preserving (LP) loss. Given deep feature activations xi∈Rdx_i \in \mathbb{R}^d extracted from the penultimate fully connected layer for sample ii in a mini-batch of size nn, the locality preserving loss is defined as:

    Llp=12∑i=1n∥xi−1k∑x∈Nk(xi)x∥22L_{lp} = \frac{1}{2} \sum_{i=1}^n \left\| x_i - \frac{1}{k} \sum_{x \in \mathcal{N}_k(x_i)} x \right\|_2^2

    where Nk(xi)\mathcal{N}_k(x_i) denotes the set of the kk nearest neighbors of sample xix_i that share the same class label as xix_i.

    The gradient of LlpL_{lp} with respect to the feature vector xix_i is:

    ∂Llp∂xi=xi−1k∑x∈Nk(xi)x\frac{\partial L_{lp}}{\partial x_i} = x_i - \frac{1}{k} \sum_{x \in \mathcal{N}_k(x_i)} x

    The total training objective jointly minimizes the cross-entropy softmax loss LsL_s (which enforces inter-class separability) and the locality preserving loss LlpL_{lp} (which enforces intra-class local compactness):

    L=Ls+λLlpL = L_s + \lambda L_{lp}

    where λ≥0\lambda \ge 0 is a weighting hyperparameter balancing global class separation and local intra-class neighborhood preservation.

    When k=nc−1k = n_c - 1 (where ncn_c is the number of training samples in class cc), the LP loss reduces to the standard center loss. Unlike center loss, which forces all samples of a class toward a single centroid, LP loss preserves local neighborhood topology and accommodates multi-modal distributions arising from varying expression intensities (e.g., smile versus laugh) and compound emotions.

  2. Knowl 2 — Optimization Algorithm of Deep Locality-Preserving CNN

    algorithm

    DLP-CNN is trained end-to-end via mini-batch stochastic gradient descent. For each mini-batch, the algorithm identifies the class-conditional kk-nearest neighbors for every sample in the feature space, computes the local neighborhood center, and jointly backpropagates the gradients from both the softmax loss and the locality preserving loss.

    Input: Mini-batch training data {xi}i=1n\{x_i\}_{i=1}^n with labels {yi}i=1n\{y_i\}_{i=1}^n, learning rate μ\mu, loss weighting hyperparameter λ\lambda, neighborhood size kk, initial network parameters WW, initial softmax loss parameters θ\theta.
    Output: Optimized network parameters WW and classifier parameters θ\theta.
    Initialize iteration counter t:=0t := 0
    repeat
        t:=t+1t := t + 1
        for each sample i∈{1,…,n}i \in \{1, \dots, n\} do
            Identify the kk nearest neighbors of xitx_i^t in the mini-batch sharing label yiy_i, represented by similarity matrix entry Sijt=1S_{ij}^t = 1 if xjtx_j^t is among the kk nearest neighbors of xitx_i^t (or vice versa) and yj=yiy_j = y_i, and Sijt=0S_{ij}^t = 0 otherwise.
            Compute the local neighborhood centroid for xitx_i^t:
            Cit:=1k∑j=1nxjtSijtC_i^t := \frac{1}{k} \sum_{j=1}^n x_j^t S_{ij}^t
        end for
        Update softmax loss parameters:
        θt+1:=θt−μt∂Lst∂θt\theta^{t+1} := \theta^t - \mu^t \frac{\partial L_s^t}{\partial \theta^t}
        Compute backpropagation error with respect to feature xitx_i^t:
        ∂Lt∂xit:=∂Lst∂xit+λ(xit−Cit)\frac{\partial L^t}{\partial x_i^t} := \frac{\partial L_s^t}{\partial x_i^t} + \lambda \left( x_i^t - C_i^t \right)
        Update network convolutional and fully connected layer parameters:
        Wt+1:=Wt−μt∑i=1n∂Lt∂xit∂xit∂WtW^{t+1} := W^t - \mu^t \sum_{i=1}^n \frac{\partial L^t}{\partial x_i^t} \frac{\partial x_i^t}{\partial W^t}
    until convergence
  3. Knowl 3 — Base Convolutional Neural Network Architecture for DLP-CNN

    model/method

    The base deep convolutional neural network architecture (baseDCNN) underlying DLP-CNN processes aligned facial images through an 18-layer structure:

    • Layer 1 (Conv): Kernel size 3×33 \times 3, 64 output channels, stride 1, padding 1
    • Layer 2 (ReLU): Rectified linear unit activation
    • Layer 3 (MPool): Max pooling, kernel size 2×22 \times 2, stride 2, padding 0
    • Layer 4 (Conv): Kernel size 3×33 \times 3, 96 output channels, stride 1, padding 1
    • Layer 5 (ReLU): Rectified linear unit activation
    • Layer 6 (MPool): Max pooling, kernel size 2×22 \times 2, stride 2, padding 0
    • Layer 7 (Conv): Kernel size 3×33 \times 3, 128 output channels, stride 1, padding 1
    • Layer 8 (ReLU): Rectified linear unit activation
    • Layer 9 (Conv): Kernel size 3×33 \times 3, 128 output channels, stride 1, padding 1
    • Layer 10 (ReLU): Rectified linear unit activation
    • Layer 11 (MPool): Max pooling, kernel size 2×22 \times 2, stride 2, padding 0
    • Layer 12 (Conv): Kernel size 3×33 \times 3, 256 output channels, stride 1, padding 1
    • Layer 13 (ReLU): Rectified linear unit activation
    • Layer 14 (Conv): Kernel size 3×33 \times 3, 256 output channels, stride 1, padding 1
    • Layer 15 (ReLU): Rectified linear unit activation
    • Layer 16 (FC): Fully connected layer producing a 2000-dimensional deep feature representation (x∈R2000x \in \mathbb{R}^{2000})
    • Layer 17 (ReLU): Rectified linear unit activation
    • Layer 18 (FC): Fully connected classification layer producing 7 class logits before the softmax layer
  4. Knowl 4 — Real-world Affective Faces Database (RAF-DB) Structure and Emotion Subsets

    definition

    The Real-world Affective Faces Database (RAF-DB) consists of 29,672 facial images downloaded from Flickr using keyword queries spanning six basic emotions and neutral emotion. Each image is annotated by approximately 40 independent crowd workers from a pool of 315 trained annotators. After filtering via an Expectation-Maximization reliability estimation framework, each image jj has a 7-dimensional ground-truth emotion distribution Gj={g1,g2,…,g7}G_j = \{g_1, g_2, \dots, g_7\}, where gk=∑i=1Rαi1(tji=k)g_k = \sum_{i=1}^R \alpha_i \mathbf{1}(t_j^i = k) for basic emotion categories k∈{1:surprise,2:fear,3:disgust,4:happiness,5:sadness,6:anger,7:neutral}k \in \{1: \text{surprise}, 2: \text{fear}, 3: \text{disgust}, 4: \text{happiness}, 5: \text{sadness}, 6: \text{anger}, 7: \text{neutral}\}, weighted by annotator reliability αi\alpha_i.

    RAF-DB is structured into two main subsets based on the mean value gmean=17∑k=17gkg_{\text{mean}} = \frac{1}{7} \sum_{k=1}^7 g_k:

    • Single-label Subset (Basic Emotions): 15,339 images with a single valid label satisfying gk>gmeang_k > g_{\text{mean}}. Class distribution: Happiness (5,957; 38.84%), Sadness (2,460; 16.04%), Surprise (1,619; 10.55%), Disgust (877; 5.72%), Anger (867; 5.65%), Fear (355; 2.31%), and Neutral.
    • Two-tab Subset (Compound Emotions): 3,954 images with bimodal emotion distributions (two non-neutral categories exceeding gmeang_{\text{mean}}) across 11 classes (after excluding the rare "Fearfully Disgusted" category with N=8N=8): Angrily Disgusted (841; 21.23%), Sadly Disgusted (738; 18.63%), Happily Surprised (697; 17.59%), Fearfully Surprised (560; 14.13%), Happily Disgusted (266; 6.71%), Angrily Surprised (176; 4.44%), Sadly Angry (163; 4.11%), Fearfully Angry (150; 3.79%), Disgustedly Surprised (148; 3.74%), Sadly Fearful (129; 3.26%), and Sadly Surprised (86; 2.17%).

    Metadata includes bounding boxes, 5 manually annotated facial landmark locations, 37 automatic Face++ landmarks, head pose angles (pitch, yaw, roll), and demographic attributes (52% female, 43% male, 5% unsure; 77% Caucasian, 15% Asian, 8% African-American; ages 0–70 years).

  5. Knowl 5 — Expectation-Maximization Label Reliability Estimation for Crowdsourced Emotion Annotations

    algorithm

    To filter subjective and noisy annotations across RR independent labelers for nn images, an Expectation-Maximization (EM) framework estimates the unobserved true emotion label yj∈{1,…,7}y_j \in \{1, \dots, 7\} for image xjx_j, annotator reliability αi>0\alpha_i > 0, and image difficulty 1/βj1/\beta_j (with βj>0\beta_j > 0). The probability of annotator ii assigning label tjit_j^i given true label yjy_j is modeled as a sigmoid function:

    p(tji=yj∣αi,βj)=11+exp⁡(−αiβj)p(t_j^i = y_j \mid \alpha_i, \beta_j) = \frac{1}{1 + \exp(-\alpha_i \beta_j)}

    The EM algorithm iterates over the following steps:

    Input: Dataset D={(xj,tj1,tj2,…,tjR)}j=1nD = \{(x_j, t_j^1, t_j^2, \dots, t_j^R)\}_{j=1}^n with labels tji∈{1,…,7}t_j^i \in \{1, \dots, 7\}.
    Output: Estimated reliability αi∗\alpha_i^* for each annotator i∈{1,…,R}i \in \{1, \dots, R\}.
    Initialize for each j∈{1,…,n}j \in \{1, \dots, n\}:
        True label yjy_j via majority voting
        βj:=−∑i=1Rp(tji)ln⁡p(tji)\beta_j := - \sum_{i=1}^R p(t_j^i) \ln p(t_j^i) (empirical label entropy)
    Initialize for each i∈{1,…,R}i \in \{1, \dots, R\}:
        αi:=1.0\alpha_i := 1.0
    repeat
        E-step:
            for each image j∈{1,…,n}j \in \{1, \dots, n\} and label y∈{1,…,7}y \in \{1, \dots, 7\} do
                Qj(y):=∏i=1Rp(y∣tji,αi,βj)=p(tj,y∣α,β)∑y′p(tj,y′∣α,β)Q_j(y) := \prod_{i=1}^R p(y \mid t_j^i, \alpha_i, \beta_j) = \frac{p(t_j, y \mid \alpha, \beta)}{\sum_{y'} p(t_j, y' \mid \alpha, \beta)}
            end for
        M-step:
            for each annotator i∈{1,…,R}i \in \{1, \dots, R\} do
                αi:=arg⁡max⁡αi∑j=1n∑yjQj(yj)ln⁡p(tj,yj∣αi,βj)Qj(yj)\alpha_i := \arg\max_{\alpha_i} \sum_{j=1}^n \sum_{y_j} Q_j(y_j) \ln \frac{p(t_j, y_j \mid \alpha_i, \beta_j)}{Q_j(y_j)}
            end for
            Update βj\beta_j via gradient ascent on the log-likelihood function
    until convergence

    Retaining annotators whose reliability exceeds quality filtering leaves 285 trusted annotators (out of 315) with an overall Cronbach's Alpha score of 0.966 across annotations.

  6. Knowl 6 — Deep Locality-Preserving CNN Expression Recognition Performance on RAF-DB

    data/table

    Expression classification performance of deep neural networks trained from scratch on RAF-DB basic emotions and evaluated on 7-class basic and 11-class compound expression sets. 2000-dimensional deep features extracted from the penultimate fully connected layer are classified using multiclass Support Vector Machines (mSVM) or Linear Discriminant Analysis (LDA). The performance metric is the mean diagonal value of the confusion matrix.

    Classifier Model Basic Emotions (%) Compound (%)
    Anger Disgust Fear Happiness Sadness Surprise Neutral Average Average
    mSVM VGG 68.52 27.50 35.13 85.32 64.85 66.32 59.88 58.22 31.63
    AlexNet 58.64 21.87 39.19 86.16 60.88 62.31 60.15 55.60 28.22
    baseDCNN 70.99 52.50 50.00 92.91 77.82 79.64 83.09 72.42 40.17
    Center Loss 68.52 53.13 54.05 93.08 78.45 79.63 83.24 72.87 39.97
    DLP-CNN 71.60 52.15 62.16 92.83 80.13 81.16 80.29 74.20 44.55
    LDA VGG 66.05 25.00 37.84 73.08 51.46 53.49 47.21 50.59 16.27
    AlexNet 43.83 27.50 37.84 75.78 39.33 61.70 48.53 47.79 15.56
    baseDCNN 66.05 47.50 51.35 89.45 74.27 76.90 77.50 69.00 28.23
    Center Loss 64.81 49.38 54.05 92.41 74.90 76.29 77.21 69.86 27.33
    DLP-CNN 77.51 55.41 52.50 90.21 73.64 74.07 73.53 70.98 32.29

    DLP-CNN achieves 74.20% mean diagonal accuracy on basic emotions and 44.55% on compound emotions with mSVM. Center loss improves over baseDCNN slightly on basic emotions (72.87% vs. 72.42%) but declines on compound emotions (39.97% vs. 40.17%), whereas DLP-CNN yields a +4.38% improvement on compound emotions over baseDCNN due to preserving multi-modal intra-class local manifolds.

  7. Knowl 7 — Cross-Database Action Unit Divergence between Real-World and Lab-Controlled Expressions

    empirical result

    Facial Action Coding System (FACS) analysis of sub-sampled RAF-DB expressions reveals that real-world Action Unit (AU) occurrences deviate substantially from lab-controlled datasets such as CK+:

    (%) AU1 AU2 AU4 AU5 AU6 AU7 AU9 AU10 AU12 AU15 AU17 AU20 AU25 AU26 AU27
    Surprise 97 97 – – – – – – – – – – 84 98 53*
    Fear 78 42 74 79 50 – – – – – – 30* – 61* 43*
    Disgust – – 51 – – 34* 89* – – – – – 82 26 55*
    Happiness – – – – 98 – – – 85 – – – 97 23 –
    Sadness 88 – 84 – – – – – – 21* 54 – – 49* –
    Anger – – 96 72* – 94 – 36 – – 87 – 79* 72* –

    (Probabilities <10%<10\% are omitted. Asterisks () indicate AU occurrence probability differing by ≥40%\ge 40\% from CK+.)*

    In cross-database transfer experiments using HOG features and an RBF-kernel Support Vector Machine:

    • Training on real-world RAF-DB and testing on CK+ achieves an average diagonal confusion matrix accuracy of 62.0%.
    • Training on lab-controlled CK+ and testing on RAF-DB achieves only 39.0% average diagonal confusion matrix accuracy.

    This discrepancy demonstrates that real-world expressions exhibit significantly higher AU diversity and multi-modality than posed lab-controlled expressions.

  8. Knowl 8 — Classification Performance of Handcrafted Descriptors on RAF-DB, CK+, and JAFFE

    data/table

    Evaluation of handcrafted appearance features (LBP: 5,900-D from 59-bin uniform LBP8,2u2LBP_{8,2}^{u2} on 10×1010 \times 10 cells; HOG: 4,000-D from 10×1010 \times 10 blocks of four 5×55 \times 5 cells with 10 bins; Gabor: 4,000-D from 40 filters at 5 scales and 8 orientations) using multiclass SVM (mSVM with RBF kernel) and LDA with kk-nearest neighbors (LDA+kNN). Performance on RAF-DB is measured using 5-fold cross-validation with the mean diagonal value of the confusion matrix as the evaluation metric.

    Classifier Feature Basic Emotions (%) Compound Emotions (%)
    CK+ JAFFE RAF-DB RAF-DB
    mSVM LBP 88.92 78.81 55.98 28.84
    HOG 90.50 84.76 58.45 33.65
    Gabor 91.98 88.95 65.12 35.76
    LDA+kNN LBP 85.84 77.74 50.97 22.89
    HOG 91.77 80.12 51.36 24.01
    Gabor 92.33 83.45 56.93 23.81

    Handcrafted feature recognition rates drop sharply from over 88–92% on lab-controlled datasets (CK+) to 55.98–65.12% on unconstrained basic expressions in RAF-DB, and further down to 28.84–35.76% on compound expressions. A naive random classifier achieves 16.07% on basic emotions and 5.79% on compound emotions.

  9. Knowl 9 — Cross-Dataset Transfer Performance of Pre-trained DLP-CNN Features on CK+ and SFEW 2.0

    empirical result

    To evaluate feature generalization, fixed-length 2000-dimensional deep activations from DLP-CNN (trained solely on RAF-DB without fine-tuning) were evaluated on the lab-controlled CK+ dataset and the spontaneous movie dataset SFEW 2.0 (under the EmotiW 2015 evaluation protocol):

    Dataset AUDN FP+SAE Mollahosseini et al. SFEW Best (EmotiW'15) DLP-CNN (No Finetuning)
    CK+ 93.70% 91.11% 93.2% – 95.78%
    SFEW 2.0 30.14% – 47.7% 52.5% 51.05%

    Without fine-tuning on the target datasets, the RAF-DB pre-trained DLP-CNN features achieve 95.78% recognition accuracy on CK+ (outperforming specialized architectures such as AUDN at 93.70% and FP+SAE at 91.11%) and 51.05% on SFEW 2.0 (approaching the EmotiW 2015 winning ensemble of 52.5%, which was trained with extra SFEW data).

Coverage note — All primary contributions—including the RAF-DB dataset creation and subsets, crowdsourcing EM reliability algorithm, FACS AU cross-database analysis, DLP-CNN locality-preserving loss formulation, optimization algorithm, base architecture, baseline experiments, deep learning benchmarks, and cross-dataset transfer experiments—have been fully captured.

References

  1. 1.C. F. Benitez-Quiroz, R. Srinivasan, and A. M. Martinez. Emotionet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild. In Proceedings of IEEE International Conference on Computer Vision & Pattern Recognition (CVPR16), Las Vegas, NV, USA, 2016.
  2. 2.V. Bettadapura. Face expression recognition and analysis: the state of the art. arXiv preprint arXiv:1203.6722, 2012.
  3. 3.J. C. Borod. The neuropsychology of emotion. Oxford University Press New York, 2000.
  4. 4.C.-C. Chang and C.-J. Lin. LIBSVM: A library for support vector machines. ACM Transactions on Intelligent Systems and Technology, 2:27:1–27:27, 2011. Software available at http://www.csie.ntu.edu.tw/ ~cjlin/libsvm.
  5. 5.N. Dalal and B. Triggs. Histograms of oriented gradients for human detection. In Computer Vision and Pattern Recognition, 2005. CVPR 2005. IEEE Computer Society Conference on, volume 1, pages 886–893. IEEE, 2005.
  6. 6.A. Dhall, R. Goecke, J. Joshi, M. Wagner, and T. Gedeon. Emotion recognition in the wild challenge 2013. In Proceedings of the 15th ACM on International conference on multimodal interaction, pages 509–516. ACM, 2013.
  7. 7.A. Dhall, R. Goecke, S. Lucey, and T. Gedeon. Static facial expression analysis in tough conditions: Data, evaluation protocol and benchmark. In Computer Vision Workshops (ICCV Workshops), 2011 IEEE International Conference on, pages 2106–2112. IEEE, 2011.
  8. 8.A. Dhall, O. Ramana Murthy, R. Goecke, J. Joshi, and T. Gedeon. Video and image based emotion recognition challenges in the wild: Emotiw 2015. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction, pages 423–426. ACM, 2015.
  9. 9.J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. In ICML, pages 647–655, 2014.
  10. 10.S. Du, Y. Tao, and A. M. Martinez. Compound facial expressions of emotion. Proceedings of the National Academy of Sciences, 111(15):E1454–E1462, 2014.
  11. 11.P. Ekman. Facial expression and emotion. American psychologist, 48(4):384, 1993.
  12. 12.P. Ekman and W. V. Friesen. Facial action coding system. 1977.
  13. 13.P. Ekman, W. V. Friesen, M. O’Sullivan, A. Chan, I. Diacoyanni-Tarlatzis, K. Heider, R. Krause, W. A. LeCompte, T. Pitcairn, P. E. Ricci-Bitti, et al. Universals and cultural differences in the judgments of facial expressions of emotion. Journal of personality and social psychology, 53(4):712, 1987.
  14. 14.B. Fasel and J. Luettin. Automatic facial expression analysis: a survey. Pattern recognition, 36(1):259–275, 2003.
  15. 15.C. Ferri, J. Hernandez-Orallo, and R. Modroiu. An experimental comparison of performance measures for classification. Pattern Recognition Letters, 30(1):27–38, 2009.
  16. 16.I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cukierski, Y. Tang, D. Thaler, D.-H. Lee, et al. Challenges in representation learning: A report on three machine learning contests. In Neural information processing, pages 117–124. Springer, 2013.
  17. 17.X. He and P. Niyogi. Locality preserving projections. In NIPS, volume 16, 2003.
  18. 18.M. Inc. Face++ research toolkit. www.faceplusplus.com, Dec. 2013.
  19. 19.S. E. Kahou, C. Pal, X. Bouthillier, P. Froumenty, C¸ . Gulc¸ehre, R. Memisevic, P. Vincent, A. Courville, Y. Bengio, R. C. Ferrari, et al. Combining modality specific deep neural networks for emotion recognition in video. In Proceedings of the 15th ACM on International conference on multimodal interaction, pages 543–550. ACM, 2013.
  20. 20.B.-K. Kim, J. Roh, S.-Y. Dong, and S.-Y. Lee. Hierarchical committee of deep convolutional neural networks for robust facial expression recognition. Journal on Multimodal User Interfaces, pages 1–17, 2016.
  21. 21.A. Krizhevsky, I. Sutskever, and G. E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems, pages 1097–1105, 2012.
  22. 22.G. Levi and T. Hassner. Emotion recognition in the wild via convolutional neural networks and mapped binary patterns. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction, pages 503–510. ACM, 2015.
  23. 23.C. Liu and H. Wechsler. Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition. Image processing, IEEE Transactions on, 11(4):467–476, 2002.
  24. 24.M. Liu, S. Li, S. Shan, and X. Chen. Au-aware deep networks for facial expression recognition. In Automatic Face and Gesture Recognition (FG), 2013 10th IEEE International Conference and Workshops on, pages 1–6. IEEE, 2013.
  25. 25.M. Liu, S. Li, S. Shan, and X. Chen. Au-inspired deep networks for facial expression feature learning. Neurocomputing, 159:126–136, 2015.
  26. 26.P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews. The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2010 IEEE Computer Society Conference on, pages 94–101. IEEE, 2010.
  27. 27.Y. Lv, Z. Feng, and C. Xu. Facial expression recognition via deep learning. In Smart Computing (SMARTCOMP), 2014 International Conference on, pages 303–308. IEEE, 2014.
  28. 28.M. J. Lyons, S. Akamatsu, M. Kamachi, J. Gyoba, and J. Budynek. The japanese female facial expression (jaffe) database. 1998.
  29. 29.M. J. Lyons, J. Budynek, and S. Akamatsu. Automatic classification of single facial images. IEEE Transactions on Pattern Analysis & Machine Intelligence, (12):1357–1362, 1999.
  30. 30.D. McDuff, R. El Kaliouby, T. Senechal, M. Amr, J. F. Cohn, and R. Picard. Affectiva-mit facial expression dataset (amfed): Naturalistic and spontaneous facial expressions collected in-the-wild. In Computer Vision and Pattern Recognition Workshops (CVPRW), 2013 IEEE Conference on, pages 881–888. IEEE, 2013.
  31. 31.A. Mollahosseini, D. Chan, and M. H. Mahoor. Going deeper in facial expression recognition using deep neural networks. In 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1–10. IEEE, 2016.
  32. 32.H.-W. Ng, V. D. Nguyen, V. Vonikakis, and S. Winkler. Deep learning for emotion recognition on small datasets using transfer learning. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction, pages 443–449. ACM, 2015.
  33. 33.T. Ojala, M. Pietikainen, and T. M ¨ aenp ¨ a¨a. Multiresolution gray-scale and rotation invariant texture classification with local binary patterns. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 24(7):971–987, 2002.
  34. 34.M. Pantic and L. J. M. Rothkrantz. Automatic analysis of facial expressions: the state of the art. IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(12):1424–1445, Dec 2000.
  35. 35.M. Pardas and A. Bonafonte. Facial animation parameters extraction and expression recognition using hidden markov models. Signal Processing: Image Communication, 17(9):675–688, 2002.
  36. 36.X. Peng, Z. Xia, L. Li, and X. Feng. Towards facial expression recognition in the wild: A new database and deep recognition system. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 93–99, 2016.
  37. 37.J. A. Russell. Is there universal recognition of emotion from facial expressions? a review of the cross-cultural studies. Psychological bulletin, 115(1):102, 1994.
  38. 38.C. Shan, S. Gong, and P. W. McOwan. Facial expression recognition based on local binary patterns: A comprehensive study. Image and Vision Computing, 27(6):803–816, 2009.
  39. 39.A. Sharif Razavian, H. Azizpour, J. Sullivan, and S. Carlsson. Cnn features off-the-shelf: an astounding baseline for recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pages 806–813, 2014.
  40. 40.K. Simonyan and A. Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014.
  41. 41.Y. Tang. Deep learning using linear support vector machines. arXiv preprint arXiv:1306.0239, 2013.
  42. 42.Y. I. Tian, T. Kanade, and J. F. Cohn. Recognizing action units for facial expression analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence, 23(2):97–115, Feb 2001.
  43. 43.Y. Wen, K. Zhang, Z. Li, and Y. Qiao. A discriminative feature learning approach for deep face recognition. In European Conference on Computer Vision, pages 499–515. Springer, 2016.
  44. 44.J. Whitehill and C. W. Omlin. Haar features for facs au recognition. In Automatic Face and Gesture Recognition, 2006. FGR 2006. 7th International Conference on, page 5–pp. IEEE, IEEE, 2006.
  45. 45.J. Whitehill, T.-f. Wu, J. Bergsma, J. R. Movellan, and P. L. Ruvolo. Whose vote should count more: Optimal integration of labels from labelers of unknown expertise. In Advances in neural information processing systems, pages 2035–2043, 2009.
  46. 46.Z. Zeng, M. Pantic, G. I. Roisman, and T. S. Huang. A survey of affect recognition methods: Audio, visual, and spontaneous expressions. Pattern Analysis and Machine Intelligence, IEEE Transactions on, 31(1):39–58, 2009.
  47. 47.X. Zhang, L. Yin, J. F. Cohn, S. Canavan, M. Reale, A. Horowitz, P. Liu, and J. M. Girard. Bp4d-spontaneous: a high-resolution spontaneous 3d dynamic facial expression database. Image and Vision Computing, 32(10):692–706, 2014.

Citation

MLA
Li, S., et al. “Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild”. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2584–93, https://doi.org/10.1109/CVPR.2017.277.
APA
Li, S., Deng, W., & Du, J. (2017). Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2584–2593. https://doi.org/10.1109/CVPR.2017.277
Chicago
Li, S., W. Deng, and J. Du. 2017. “Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild”. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2584–93. https://doi.org/10.1109/CVPR.2017.277.
Harvard
Li, S., Deng, W. and Du, J. (2017) “Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 2584–2593. Available at: https://doi.org/10.1109/CVPR.2017.277.
Vancouver
1. Li S, Deng W, Du J (2017) Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild. In: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 2584–2593

BibTeX

@inproceedings{Li_2017, title={Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild}, url={http://dx.doi.org/10.1109/CVPR.2017.277}, DOI={10.1109/cvpr.2017.277}, booktitle={2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Li, Shan and Deng, Weihong and Du, JunPing}, year={2017}, month=July, pages={2584–2593} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF