AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild

Ali MollahosseiniBehzad HasaniMohammad H. Mahoor

article2017IEEE Transactions on Affective Computing2,268 citationsT-AC Most Influential Paper Award

Presents AffectNet, a large-scale in-the-wild facial expression database annotated for both categorical emotions and continuous valence-arousal dimensions, establishing deep learning baselines to advance automated affective computing.

Listen

Automated facial expression recognition is critical for developing responsive human-machine interfaces, yet existing systems struggle to perform reliably in real-world, uncontrolled environments. Progress has been hindered because available training datasets are relatively small, often captured in artificial laboratory settings, and largely restricted to discrete emotion categories (such as happiness or anger). Crucially, existing datasets lack sufficient coverage of continuous dimensional modelsnamely valence (how positive or negative an emotion is) and arousal (the level of excitation or calmness)—which are essential for capturing subtle variations and emotional intensity.

The article introduces AffectNet, a massive in-the-wild facial affect database, and demonstrates its utility for training automated systems to recognize both discrete emotion categories and continuous dimensional states.

To build AffectNet, researchers collected over one million facial images from the web using 1,250 emotion-related search queries across six languages. Twelve trained human annotators manually labeled 450,000 images for discrete emotion categories, continuous valence and arousal levels, and facial occlusions, while a subset of 36,000 images was labeled by two independent annotators to assess human agreement. The authors then developed baseline deep convolutional neural networks to classify discrete emotions and predict continuous valence and arousal values, comparing these models against conventional machine learning methods and an established commercial system.

The investigation produced three main findings. First, human agreement on facial affect in uncontrolled settings is moderately low: annotators agreed on discrete emotion categories in only 60.7% of cases and demonstrated higher consistency when rating valence than arousal. Second, among various deep learning strategies used to tackle data imbalance, a weighted-loss approach achieved the highest skew-normalized performance, delivering an overall multi-class accuracy of 0.63 and an F1-score of 0.62, compared to 0.37 accuracy and 0.31 F1-score for conventional support vector machines. Third, the deep learning baseline outperformed both support vector regression and an off-the-shelf commercial tool (Microsoft Cognitive Services Emotion API), which scored an overall accuracy of 0.48 and struggled significantly on under-represented classes like contempt, disgust, and fear.

These findings indicate that large-scale, in-the-wild datasets coupled with deep learning architectures substantially reduce the risk of system failure in unconstrained environments. Because the commercial off-the-shelf system and baseline models struggled on minority and nuanced emotions, organizations relying on facial recognition for human-computer interaction or monitoring must account for performance disparities across different emotional states. The research shows that continuous dimensional modeling effectively captures complex states that discrete categories miss, though automated detection remains inherently more challenging for arousal than valence.

Organizations developing emotion-aware technologies should adopt large-scale in-the-wild benchmarks like AffectNet for model training and benchmarking. Teams should utilize weighted-loss formulations or advanced balancing strategies during training to prevent severe performance drops on minority emotional classes. Furthermore, researchers and practitioners should investigate co-training methods that leverage both categorical and dimensional labels simultaneously within the same corpus to boost overall prediction robustness.

A primary limitation of this work is the intrinsic skew of web-sourced data, which heavily favors positive and neutral expressions over rarer emotions like disgust and contempt. Additionally, the moderate rate of human annotator agreement highlights inherent subjectivity in emotional perception, meaning automated models trained on these labels will carry baseline uncertainty, particularly when predicting emotional arousal.

arXiv: 1708.03985
Cover for AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild

Abstract

Automated affective computing in the wild setting is a challenging problem in computer vision. Existing annotated databases of facial expressions in the wild are small and mostly cover discrete emotions (aka the categorical model). There are very limited annotated facial databases for affective computing in the continuous dimensional model (e.g., valence and arousal). To meet this need, we collected, annotated, and prepared for public distribution a new database of facial emotions in the wild (called AffectNet). AffectNet contains more than 1,000,000 facial images from the Internet by querying three major search engines using 1250 emotion related keywords in six different languages. About half of the retrieved images were manually annotated for the presence of seven discrete facial expressions and the intensity of valence and arousal. AffectNet is by far the largest database of facial expression, valence, and arousal in the wild enabling research in automated facial expression recognition in two different emotion models. Two baseline deep neural networks are used to classify images in the categorical model and predict the intensity of valence and arousal. Various evaluation metrics show that our deep neural network baselines can perform better than conventional machine learning methods and off-the-shelf facial expression recognition systems.

Table of Contents

  • I Introduction
  • II Related Work
  • II-A Existing databases
  • II-B Evaluation Metrics
  • II-C Existing Algorithms
  • III AffectNet
  • III-A Facial Images from the Web
  • III-B Annotation
  • III-B1 Categorical Model Annotation
  • III-B2 Dimensional (Valence & Arousal) Annotation
  • III-C Annotation Agreement
  • IV Baseline
  • IV-A Test, Validation, and Training Sets
  • IV-B Categorical Model Baseline
  • IV-C Dimensional Model (Valence and Arousal) Baseline
  • V Conclusion
  • References
  • A

Knowls

  1. Knowl 1 — AffectNet Dataset Construction and Specifications

    definition

    AffectNet is a large-scale database of facial expressions in the wild containing over 1,000,0001{,}000{,}000 facial images collected by querying three search engines (Google, Bing, Yahoo) using 1,250 emotion-related keywords across six languages: English, Spanish, Portuguese, German, Arabic, and Farsi. Queries combined emotional terms with gender, age, or ethnicity qualifiers, filtered against non-human objects and commercial stock-photo watermarks.

    From approximately 1,800,0001{,}800{,}000 retrieved URLs, face detection (via OpenCV) and 66 facial landmark localizations (using regression of local binary features trained on the 300W dataset) were applied. Over 1,000,0001{,}000{,}000 images containing at least one face with detected landmarks were retained. A subset of 450,000450{,}000 images was manually annotated by 12 trained human experts across both discrete categorical expressions and continuous dimensional affect (valence and arousal).

    Key dataset properties include:

    • Image Resolution: Mean face resolution of 425×425425 \times 425 pixels with a standard deviation of 349×349349 \times 349 pixels.
    • Demographics & Attributes: 49%49\% male, mean estimated age of 33.0133.01 years (standard deviation 16.9616.96 years). Forehead, mouth, and eye occlusions are present in 4.5%4.5\%, 1.08%1.08\%, and 0.49%0.49\% of images, respectively. 9.63%9.63\% wear glasses; 51.07%51.07\% and 41.40%41.40\% feature eye and lip cosmetics.
    • Head Pose: Mean estimated pitch, yaw, and roll are 0.00.0^\circ, 0.7-0.7^\circ, and 1.19-1.19^\circ.
  2. Knowl 2 — AffectNet Affective Annotation Taxonomy and Protocol

    definition

    AffectNet annotates facial affect in two complementary psychological frameworks:

    1. Categorical Model (11 Categories):

      • Basic/Compound Expressions (8 classes): Neutral, Happy, Sad, Surprise, Fear, Disgust, Anger, Contempt.
      • None: Facial expressions conveying non-primary emotional states (e.g., sleepy, bored, tired, flirtatious/seducing, confused, ashamed, focused) that annotators could not assign to the primary eight categories, but for which valence and arousal can still be measured.
      • Uncertain: Images where annotators could not confidently identify an expression.
      • Non-Face: Images containing no face, watermark occlusions over the face, bounding box detection failures, artwork/drawings/paintings, or faces distorted beyond normal appearance.
    2. Dimensional Model (Continuous Valence and Arousal):

      • Affect is mapped onto Russell's 2D circumplex coordinate system where the horizontal axis represents valence (positivity vs. negativity, [1.0,1.0][-1.0, 1.0]) and the vertical axis represents arousal (calm/soothing vs. exciting/agitating, [1.0,1.0][-1.0, 1.0]).
      • Annotations are constrained within the circumplex disk: Cartesian coordinates (v,a)(v, a) satisfy polar bounds 0r10 \le r \le 1 with r=v2+a2r = \sqrt{v^2 + a^2} and angle 0θ<3600 \le \theta < 360^\circ.
      • Bounding verification regions for valence-arousal pairs were mapped to each categorical expression in the annotation interface (e.g., Happy required valence in (0.0,1.0](0.0, 1.0] and arousal in [0.2,0.5][-0.2, 0.5]), raising an alert if an annotator selected a point outside the expected region to prevent annotation errors.
  3. Knowl 3 — Inter-Annotator Agreement in Wild Facial Expression and Affect Modeling

    empirical result

    To evaluate human reliability on in-the-wild facial affect, 36,00036{,}000 AffectNet images were annotated independently and blindly by two expert annotators.

    Categorical Agreement:

    • Overall category agreement across annotators was 60.7%60.7\%.
    • Agreement was highest for Happy (79.6%79.6\%) and Non-Face (83.9%83.9\%), and lowest for None (9.6%9.6\%) and Uncertain (20.6%20.6\%).
    • Among specific emotion categories, agreements were: Neutral (50.8%50.8\%), Sad (69.7%69.7\%), Surprise (66.5%66.5\%), Fear (61.1%61.1\%), Disgust (67.6%67.6\%), Anger (62.3%62.3\%), and Contempt (66.9%66.9\%).

    Dimensional Agreement: Agreement was evaluated across two groups: when annotators agreed on the categorical label versus across all 36,00036{,}000 images.

    Metric Same Category All Images
    Valence Arousal Valence Arousal
    Root Mean Square Error (RMSE) 0.190 0.261 0.340 0.362
    Pearson Correlation (CORR) 0.951 0.766 0.823 0.567
    Sign Agreement Metric (SAGR) 0.906 0.709 0.815 0.667
    Concordance Correlation (CCC) 0.951 0.746 0.821 0.551

    Human agreement is markedly higher on valence than on arousal across all metrics, indicating that perceived emotional valence is substantially less subjective than perceived arousal in static facial imagery.

  4. Knowl 4 — Continuous Affect Evaluation Metrics Formulation

    equation

    Continuous prediction of valence and arousal across nn samples, where θ^i\hat{\theta}_i is the predicted value and θi\theta_i is the ground-truth value for sample ii with θ^i,θi[1,1]\hat{\theta}_i, \theta_i \in [-1, 1], is evaluated using four standard metrics:

    1. Root Mean Square Error (RMSE): RMSE=1ni=1n(θ^iθi)2\text{RMSE} = \sqrt{\frac{1}{n} \sum_{i=1}^n (\hat{\theta}_i - \theta_i)^2}

    2. Pearson's Correlation Coefficient (CC / CORR): CC=COV{θ^,θ}σθ^σθ=E[(θ^μθ^)(θμθ)]σθ^σθ\text{CC} = \frac{\text{COV}\{\hat{\theta}, \theta\}}{\sigma_{\hat{\theta}} \sigma_\theta} = \frac{\mathbb{E}[(\hat{\theta} - \mu_{\hat{\theta}})(\theta - \mu_\theta)]}{\sigma_{\hat{\theta}} \sigma_\theta} where μθ^,μθ\mu_{\hat{\theta}}, \mu_\theta and σθ^,σθ\sigma_{\hat{\theta}}, \sigma_\theta denote the means and standard deviations of the predictions and ground truths.

    3. Concordance Correlation Coefficient (CCC): ρc=2ρσθ^σθσθ^2+σθ2+(μθ^μθ)2\rho_c = \frac{2 \rho \sigma_{\hat{\theta}} \sigma_\theta}{\sigma_{\hat{\theta}}^2 + \sigma_\theta^2 + (\mu_{\hat{\theta}} - \mu_\theta)^2} where ρ\rho is the Pearson correlation coefficient. Unlike CC, CCC penalizes predictions that correlate well with ground truth but have shifted mean or variance.

    4. Sign Agreement Metric (SAGR): SAGR=1ni=1nδ(sign(θ^i),sign(θi))\text{SAGR} = \frac{1}{n} \sum_{i=1}^n \delta(\text{sign}(\hat{\theta}_i), \text{sign}(\theta_i)) where δ(a,b)=1\delta(a, b) = 1 if a=ba = b and 00 otherwise. SAGR captures whether the predicted valence or arousal shares the correct emotional polarity.

  5. Knowl 5 — Weighted Cross-Entropy Loss for Imbalanced Facial Expression Classification

    model/method

    To counter extreme class imbalance in wild emotion datasets, a cost-sensitive weighted cross-entropy loss (Infogain loss) is formulated. For a training example with feature representation XX, ground-truth class label l{1,,K}l \in \{1, \dots, K\}, and predicted softmax probabilities p^i=exp(xi)j=1Kexp(xj)\hat{p}_i = \frac{\exp(x_i)}{\sum_{j=1}^K \exp(x_j)}, the loss is defined as:

    E=i=1KHl,ilog(p^i)E = - \sum_{i=1}^K H_{l, i} \log(\hat{p}_i)

    where HH is a K×KK \times K diagonal penalty matrix whose non-zero entries are set inversely proportional to class frequency relative to the minimum class size:

    Hij={fifmin,if i=j0,otherwiseH_{ij} = \begin{cases} \frac{f_i}{f_{\min}}, & \text{if } i = j \\ 0, & \text{otherwise} \end{cases}

    where fif_i is the number of training samples belonging to class ii, and fmin=minkfkf_{\min} = \min_k f_k is the sample count of the most under-represented class (Disgust in AffectNet). When HH is the identity matrix, this reduces to standard cross-entropy.

    The derivative of the loss with respect to the network logit xkx_k is:

    Exk=(i=1KHl,i)p^kHl,k\frac{\partial E}{\partial x_k} = \left( \sum_{i=1}^K H_{l, i} \right) \hat{p}_k - H_{l, k}

    This formulation strongly penalizes classification errors on rare classes (e.g., Contempt, Disgust, Fear) while preventing frequent classes (Happy, Neutral) from dominating the gradient updates.

  6. Knowl 6 — AffectNet Benchmark Partitions and Skew-Normalized Evaluation Protocol

    experimental setup

    The AffectNet dataset is partitioned into three standardized splits for training, validation, and testing:

    • Test Set: Formed from the 36,00036{,}000 double-annotated images. Categorical ground truth is determined by selecting the queried keyword when one annotator matched the query (29.5%29.5\% of disagreements) and picking randomly between annotators otherwise. Continuous dimensional labels are assigned by randomly picking one of the two annotations. The test set reflects natural web skew (e.g., >11,000>11{,}000 happy vs. 1,000\sim 1{,}000 contempt images).
    • Validation Set: Constructed by sampling 500 images per category at random, yielding a perfectly balanced set across the 8 primary categorical classes for hyperparameter tuning.
    • Training Set: Comprises all remaining annotated images (heavily imbalanced across categories).

    Skew-Normalized Evaluation Protocol: Because extreme class imbalance heavily distorts standard metrics like Accuracy, F1-Score, Cohen's Kappa, Krippendorff's Alpha, and AUC-PR, AffectNet defines a skew-normalization procedure for test evaluation. Test classes are randomly under-sampled to the size of the smallest class, metrics are computed on the balanced subset, and this process is repeated over 200 random trials. The reported skew-normalized score is the mean across the 200 balanced trials.

  7. Knowl 7 — Categorical Emotion Recognition Baselines and Comparison on AffectNet

    empirical result

    AlexNet architectures trained with four class-imbalance strategies (Imbalanced, Down-Sampling to max 15,00015{,}000/class, Up-Sampling to match Happy class, and Weighted-Loss) were evaluated on the 8-class AffectNet test set against a linear Support Vector Machine (SVM on 256×256256 \times 256 crops with HOG cell size 8 and PCA retaining 95%95\% variance, yielding 6,697 features) and the commercial Microsoft Cognitive Services Emotion API.

    All metrics are macro-averaged across 8 categories (Accuracy computed multi-class; others one-vs-rest). Both Original (Orig) and Skew-Normalized (Norm) scores over 200 balanced trials are reported:

    Metric Imbalanced Down-Sample Up-Sample Weighted-Loss Linear SVM MS Cognitive
    Orig Norm Orig Norm Orig Norm Orig Norm Orig Norm Orig Norm
    Accuracy 0.72 0.54 0.68 0.58 0.68 0.57 0.64 0.63 0.60 0.37 0.68 0.48
    F1-Score 0.57 0.52 0.56 0.57 0.56 0.55 0.55 0.62 0.37 0.31 0.51 0.45
    Cohen κ\kappa 0.53 0.46 0.51 0.51 0.52 0.49 0.50 0.57 0.32 0.25 0.46 0.40
    Kripp. α\alpha 0.52 0.45 0.51 0.51 0.51 0.48 0.50 0.57 0.31 0.22 0.46 0.37
    AUC 0.85 0.80 0.82 0.85 0.82 0.84 0.86 0.86 0.77 0.70 0.83 0.77
    AUC-PR 0.56 0.55 0.54 0.57 0.55 0.56 0.58 0.64 0.39 0.37 0.52 0.50

    The Weighted-Loss AlexNet baseline attains the highest skew-normalized F1-score (0.620.62), skew-normalized Accuracy (0.630.63), Kappa (0.570.57), Alpha (0.570.57), and AUC-PR (0.640.64), significantly outperforming SVM and MS Cognitive API, especially on minority classes (Disgust, Fear, Contempt).

  8. Knowl 8 — Continuous Valence and Arousal Prediction Baselines on AffectNet

    empirical result

    Continuous dimensional affect baseline prediction was evaluated on the AffectNet test set using two independent AlexNet regression models (one for valence, one for arousal) where the final fully-connected classification layer was replaced with a single linear regression neuron trained with Euclidean (L2L_2) loss:

    E=12Nn=1Ny^nyn22E = \frac{1}{2N} \sum_{n=1}^N \|\hat{y}_n - y_n\|_2^2

    with initial learning rate 0.0010.001, batch size 256, and momentum 0.90.9 trained over 16 epochs. These were benchmarked against Support Vector Regression (SVR with linear kernel via Liblinear on HOG features reduced to 6,697 dimensions via PCA retaining 95%95\% variance):

    Metric CNN (AlexNet) SVR (HOG + PCA)
    Valence Arousal Valence Arousal
    RMSE 0.394 0.402 0.494 0.400
    CORR 0.602 0.539 0.429 0.360
    SAGR 0.728 0.670 0.619 0.748
    CCC 0.541 0.450 0.340 0.199

    AlexNet substantially outperforms SVR in CORR and CCC for both dimensions. Spatial error maps across the valence-arousal circumplex show that prediction error is lowest at the center of the circumplex, whereas low-valence mid-arousal and low-arousal mid-valence regions (corresponding to contempt, boredom, and sleepiness) exhibit the highest regression errors.

  9. Knowl 9 — Natural Skew and Query Hit Rates of In-The-Wild Facial Affect

    data/table

    The category breakdown across the 450,000450{,}000 manually annotated AffectNet images reveals the severe intrinsic distribution skew of facial affect on the Internet:

    Category Image Count Category Image Count
    Neutral 80,276 Disgust 5,264
    Happy 146,198 Contempt 5,135
    Sad 29,487 None 35,322
    Surprise 16,288 Uncertain 13,163
    Fear 8,191 Non-Face 88,895
    Anger 28,130 Total 450,000

    Query Hit Rates:

    • Queries intended for Happy had by far the highest precision hit rate: 48.9%48.9\% of returned images were verified as happy.
    • For all other emotion queries, the hit rate for the queried emotion was below 20%20\%: Sad (15.7%15.7\%), Surprise (16.0%16.0\%), Anger (12.2%12.2\%), Fear (4.0%4.0\%), Disgust (2.7%2.7\%), and Contempt (2.4%2.4\%).
    • Across all query types, approximately 15%15\% of returned images were annotated as Neutral and 16%20%16\%\text{--}20\% as Non-Face, reflecting an overwhelming natural bias toward publishing positive and neutral facial expressions on web platforms.

Coverage note — None was omitted; all primary contributions, dataset construction procedures, annotation taxonomies, inter-rater analyses, baseline loss formulations, and comprehensive experimental results for both categorical and dimensional affect are covered.

References

  1. 1.J. Tao and T. Tan, “Affective computing: A review,” in International Conference on Affective Computing and Intelligent Interaction. Springer, 2005, pp. 981–995.
  2. 2.P. Ekman and W. V. Friesen, “Constants across cultures in the face and emotion.” Journal of personality and social psychology, vol. 17, no. 2, p. 124, 1971.
  3. 3.J. A. Russell, “A circumplex model of affect,” Journal of Personality and Social Psychology, vol. 39, no. 6, pp. 1161–1178, 1980.
  4. 4.P. Ekman and W. V. Friesen, “Facial action coding system,” 1977.
  5. 5.W. V. Friesen and P. Ekman, “Emfacs-7: Emotional facial action coding system,” Unpublished manuscript, University of California at San Francisco, vol. 2, p. 36, 1983.
  6. 6.S. Du, Y. Tao, and A. M. Martinez, “Compound facial expressions of emotion,” Proceedings of the National Academy of Sciences, vol. 111, no. 15, pp. E1454–E1462, 2014.
  7. 7.M. Lyons, S. Akamatsu, M. Kamachi, and J. Gyoba, “Coding facial expressions with gabor wavelets,” in Automatic Face and Gesture Recognition, 1998. Proceedings. Third IEEE International Conference on. IEEE, 1998, pp. 200–205.
  8. 8.Y.-I. Tian, T. Kanade, and J. F. Cohn, “Recognizing action units for facial expression analysis,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 23, no. 2, pp. 97–115, 2001.
  9. 9.P. Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews, “The extended cohn-kanade dataset (ck+): A complete dataset for action unit and emotion-specified expression,” in Computer Vision and Pattern Recognition Workshops (CVPRW), 2010 IEEE Computer Society Conference on. IEEE, 2010, pp. 94–101.
  10. 10.M. Pantic, M. Valstar, R. Rademaker, and L. Maat, “Web-based database for facial expression analysis,” in Multimedia and Expo, 2005. ICME 2005. IEEE International Conference on. IEEE, 2005, pp. 5–pp.
  11. 11.R. Gross, I. Matthews, J. Cohn, T. Kanade, and S. Baker, “Multi-pie,” Image and Vision Computing, vol. 28, no. 5, pp. 807–813, 2010.
  12. 12.J. F. Cohn and K. L. Schmidt, “The timing of facial motion in posed and spontaneous smiles,” International Journal of Wavelets, Multiresolution and Information Processing, vol. 2, no. 02, pp. 121–132, 2004.
  13. 13.M. F. Valstar, M. Pantic, Z. Ambadar, and J. F. Cohn, “Spontaneous vs. posed facial behavior: automatic analysis of brow actions,” in Proceedings of the 8th international conference on Multimodal interfaces. ACM, 2006, pp. 162–170.
  14. 14.S. M. Mavadati, M. H. Mahoor, K. Bartlett, P. Trinh, and J. F. Cohn, “Disfa: A spontaneous facial action intensity database,” Affective Computing, IEEE Transactions on, vol. 4, no. 2, pp. 151–160, 2013.
  15. 15.D. McDuff, R. Kaliouby, T. Senechal, M. Amr, J. Cohn, and R. Picard, “Affectiva-mit facial expression dataset (am-fed): Naturalistic and spontaneous facial expressions collected,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2013, pp. 881–888.
  16. 16.I. Sneddon, M. McRorie, G. McKeown, and J. Hanratty, “The belfast induced natural emotion database,” IEEE Transactions on Affective Computing, vol. 3, no. 1, pp. 32–41, 2012.
  17. 17.J. F. Grafsgaard, J. B. Wiggins, K. E. Boyer, E. N. Wiebe, and J. C. Lester, “Automatically recognizing facial expression: Predicting engagement and frustration.” in EDM, 2013, pp. 43–50.
  18. 18.A. Dhall, R. Goecke, J. Joshi, M. Wagner, and T. Gedeon, “Emotion recognition in the wild challenge 2013,” in Proceedings of the 15th ACM on International conference on multimodal interaction. ACM, 2013, pp. 509–516.
  19. 19.I. J. Goodfellow, D. Erhan, P. L. Carrier, A. Courville, M. Mirza, B. Hamner, W. Cukierski, Y. Tang, D. Thaler, D.-H. Lee et al., “Challenges in representation learning: A report on three machine learning contests,” Neural Networks, vol. 64, pp. 59–63, 2015.
  20. 20.A. Mollahosseini, B. Hasani, M. J. Salvador, H. Abdollahi, D. Chan, and M. H. Mahoor, “Facial expression recognition from world wild web,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, June 2016.
  21. 21.S. Zafeiriou, A. Papaioannou, I. Kotsia, M. A. Nicolaou, and G. Zhao, “Facial affect “in-the-wild”: A survey and a new database,” in International Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Affect ”in-the-wild” Workshop, June 2016.
  22. 22.C. F. Benitez-Quiroz, R. Srinivasan, and A. M. Martinez, “Emotionet: An accurate, real-time algorithm for the automatic annotation of a million facial expressions in the wild,” in Proceedings of IEEE International Conference on Computer Vision & Pattern Recognition (CVPR16), Las Vegas, NV, USA, 2016.
  23. 23.M. A. Nicolaou, H. Gunes, and M. Pantic, “Audio-visual classification and fusion of spontaneous affective data in likelihood space,” in Pattern Recognition (ICPR), 2010 20th International Conference on. IEEE, 2010, pp. 3695–3699.
  24. 24.——, “Continuous prediction of spontaneous affect from multiple cues and modalities in valence-arousal space,” IEEE Transactions on Affective Computing, vol. 2, no. 2, pp. 92–105, 2011.
  25. 25.F. Ringeval, A. Sonderegger, J. Sauer, and D. Lalanne, “Introducing the recola multimodal corpus of remote collaborative and affective interactions,” in Automatic Face and Gesture Recognition (FG), 2013 10th IEEE International Conference and Workshops on. IEEE, 2013, pp. 1–8.
  26. 26.S. Koelstra, C. Muhl, M. Soleymani, J.-S. Lee, A. Yazdani, T. Ebrahimi, T. Pun, A. Nijholt, and I. Patras, “Deap: A database for emotion analysis; using physiological signals,” IEEE Transactions on Affective Computing, vol. 3, no. 1, pp. 18–31, 2012.
  27. 27.A. Dhall, R. Goecke, S. Lucey, and T. Gedeon, “Static facial expression analysis in tough conditions: Data, evaluation protocol and benchmark,” in Computer Vision Workshops (ICCV Workshops), 2011 IEEE International Conference on. IEEE, 2011, pp. 2106–2112.
  28. 28.A. Mollahosseini and M. H. Mahoor, “Bidirectional warping of active appearance model,” in Computer Vision and Pattern Recognition Workshops (CVPRW), 2013 IEEE Conference on. IEEE, 2013, pp. 875–880.
  29. 29.S. Ren, X. Cao, Y. Wei, and J. Sun, “Face alignment at 3000 fps via regressing local binary features,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2014, pp. 1685–1692.
  30. 30.L. Yu, “face-alignment-in-3000fps,” https://github.com/yulequan/face-alignment-in-3000fps, 2016.
  31. 31.G. A. Miller, “Wordnet: a lexical database for english,” Communications of the ACM, vol. 38, no. 11, pp. 39–41, 1995.
  32. 32.P. Viola and M. J. Jones, “Robust real-time face detection,” International journal of computer vision, vol. 57, no. 2, pp. 137–154, 2004.
  33. 33.D. You, O. C. Hamsici, and A. M. Martinez, “Kernel optimization in discriminant analysis,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 33, no. 3, pp. 631–638, 2011.
  34. 34.P. Lucey, J. F. Cohn, K. M. Prkachin, P. E. Solomon, and I. Matthews, “Painful data: The unbc-mcmaster shoulder pain expression archive database,” in Automatic Face & Gesture Recognition and Workshops (FG 2011), 2011 IEEE International Conference on. IEEE, 2011, pp. 57–64.
  35. 35.M. M. Bradley and P. J. Lang, “Measuring emotion: the self-assessment manikin and the semantic differential,” Journal of behavior therapy and experimental psychiatry, vol. 25, no. 1, pp. 49–59, 1994.
  36. 36.B. Schuller, M. Valstar, F. Eyben, G. McKeown, R. Cowie, and M. Pantic, “Avec 2011–the first international audio/visual emotion challenge,” in International Conference on Affective Computing and Intelligent Interaction. Springer, 2011, pp. 415–424.
  37. 37.B. Schuller, M. Valster, F. Eyben, R. Cowie, and M. Pantic, “Avec 2012: the continuous audio/visual emotion challenge,” in Proceedings of the 14th ACM international conference on Multimodal interaction. ACM, 2012, pp. 449–456.
  38. 38.M. Valstar, B. Schuller, K. Smith, F. Eyben, B. Jiang, S. Bilakhia, S. Schnieder, R. Cowie, and M. Pantic, “Avec 2013: the continuous audio/visual emotion and depression recognition challenge,” in Proceedings of the 3rd ACM international workshop on Audio/visual emotion challenge. ACM, 2013, pp. 3–10.
  39. 39.M. Valstar, B. Schuller, K. Smith, T. Almaev, F. Eyben, J. Krajewski, R. Cowie, and M. Pantic, “Avec 2014: 3d dimensional affect and depression recognition challenge,” in Proceedings of the 4th International Workshop on Audio/Visual Emotion Challenge. ACM, 2014, pp. 3–10.
  40. 40.F. Ringeval, B. Schuller, M. Valstar, R. Cowie, and M. Pantic, “Avec 2015: The 5th international audio/visual emotion challenge and workshop,” in Proceedings of the 23rd ACM international conference on Multimedia. ACM, 2015, pp. 1335–1336.
  41. 41.M. Valstar, J. Gratch, B. Schuller, F. Ringeval, D. Lalanne, M. T. Torres, S. Scherer, G. Stratou, R. Cowie, and M. Pantic, “Avec 2016-depression, mood, and emotion recognition workshop and challenge,” arXiv preprint arXiv:1605.01600, 2016.
  42. 42.G. McKeown, M. Valstar, R. Cowie, M. Pantic, and M. Schroder, “The semaine database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent,” IEEE Transactions on Affective Computing, vol. 3, no. 1, pp. 5–17, 2012.
  43. 43.A. Mollahosseini, D. Chan, and M. H. Mahoor, “Going deeper in facial expression recognition using deep neural networks,” IEEE Winter Conference on Applications of Computer Vision (WACV), 2016.
  44. 44.C. Shan, S. Gong, and P. W. McOwan, “Facial expression recognition based on local binary patterns: A comprehensive study,” Image and Vision Computing, vol. 27, no. 6, pp. 803–816, 2009.
  45. 45.X. Zhang, M. H. Mahoor, and S. M. Mavadati, “Facial expression recognition using {l} {p}-norm mkl multiclass-svm,” Machine Vision and Applications, pp. 1–17, 2015.
  46. 46.L. He, D. Jiang, L. Yang, E. Pei, P. Wu, and H. Sahli, “Multimodal affective dimension prediction using deep bidirectional long short-term memory recurrent neural networks,” in Proceedings of the 5th International Workshop on Audio/Visual Emotion Challenge. ACM, 2015, pp. 73–80.
  47. 47.Y. Fan, X. Lu, D. Li, and Y. Liu, “Video-based emotion recognition using cnn-rnn and c3d hybrid networks,” in Proceedings of the 18th ACM International Conference on Multimodal Interaction. ACM, 2016, pp. 445–450.
  48. 48.Y. Tang, “Deep learning using linear support vector machines,” arXiv preprint arXiv:1306.0239, 2013.
  49. 49.M. Sokolova, N. Japkowicz, and S. Szpakowicz, “Beyond accuracy, f-score and roc: a family of discriminant measures for performance evaluation,” in Australasian Joint Conference on Artificial Intelligence. Springer, 2006, pp. 1015–1021.
  50. 50.J. Cohen, “A coefficient of agreement for nominal scales,” Educational and Psychological Measurement, vol. 20, no. 1, p. 37, 1960.
  51. 51.K. Krippendorff, “Estimating the reliability, systematic error and random error of interval data,” Educational and Psychological Measurement, vol. 30, no. 1, pp. 61–70, 1970.
  52. 52.P. E. Shrout and J. L. Fleiss, “Intraclass correlations: uses in assessing rater reliability.” Psychological bulletin, vol. 86, no. 2, p. 420, 1979.
  53. 53.L. A. Jeni, J. F. Cohn, and F. De La Torre, “Facing imbalanced data–recommendations for the use of performance metrics,” in Affective Computing and Intelligent Interaction (ACII), 2013 Humaine Association Conference on. IEEE, 2013, pp. 245–251.
  54. 54.S. Bermejo and J. Cabestany, “Oriented principal component analysis for large margin classifiers,” Neural Networks, vol. 14, no. 10, pp. 1447–1461, 2001.
  55. 55.M. Mohammadi, E. Fatemizadeh, and M. H. Mahoor, “Pca-based dictionary building for accurate facial expression recognition via sparse representation,” Journal of Visual Communication and Image Representation, vol. 25, no. 5, pp. 1082–1092, 2014.
  56. 56.C. Liu and H. Wechsler, “Gabor feature based classification using the enhanced fisher linear discriminant model for face recognition,” IEEE Trans. Image Process., vol. 11, no. 4, pp. 467–476, 2002.
  57. 57.A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems, 2012, pp. 1097–1105.
  58. 58.A. Toshev and C. Szegedy, “Deeppose: Human pose estimation via deep neural networks,” in Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on. IEEE, 2014, pp. 1653–1660.
  59. 59.Y. Taigman, M. Yang, M. Ranzato, and L. Wolf, “Deepface: Closing the gap to human-level performance in face verification,” in Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on. IEEE, 2014, pp. 1701–1708.
  60. 60.C. Sagonas, E. Antonakos, G. Tzimiropoulos, S. Zafeiriou, and M. Pantic, “300 faces in-the-wild challenge: Database and results,” Image and Vision Computing, vol. 47, pp. 3–18, 2016.
  61. 61.G. Paltoglou and M. Thelwall, “Seeing stars of valence and arousal in blog posts,” IEEE Transactions on Affective Computing, vol. 4, no. 1, pp. 116–123, 2013.
  62. 62.C. Cortes and V. Vapnik, “Support-vector networks,” Machine learning, vol. 20, no. 3, pp. 273–297, 1995.
  63. 63.A. Smola and V. Vapnik, “Support vector regression machines,” Advances in neural information processing systems, vol. 9, pp. 155–161, 1997.
  64. 64.G. Caridakis, K. Karpouzis, and S. Kollias, “User and context adaptive neural networks for emotion recognition,” Neurocomputing, vol. 71, no. 13, pp. 2553–2562, 2008.
  65. 65.H. He and E. A. Garcia, “Learning from imbalanced data,” IEEE Trans. Knowl. Data Eng., vol. 21, no. 9, pp. 1263–1284, 2009.
  66. 66.Y. Jia, E. Shelhamer, J. Donahue, S. Karayev, J. Long, R. Girshick, S. Guadarrama, and T. Darrell, “Caffe: Convolutional architecture for fast feature embedding,” in Proceedings of the 22nd ACM international conference on Multimedia. ACM, 2014, pp. 675–678.
  67. 67.N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), vol. 1. IEEE, 2005, pp. 886–893.
  68. 68.R.-E. Fan, K.-W. Chang, C.-J. Hsieh, X.-R. Wang, and C.-J. Lin, “Liblinear: A library for large linear classification,” Journal of machine learning research, vol. 9, no. Aug, pp. 1871–1874, 2008.
  69. 69.“Microsoft cognitive services - emotion api,” https://www.microsoft.com/cognitive-services/en-us/emotion-api, (Accessed on 12/01/2016).

Citation

MLA
Mollahosseini, A., et al. “AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild”. IEEE Transactions on Affective Computing, vol. 10, no. 1, 2019, pp. 18–31, https://doi.org/10.1109/TAFFC.2017.2740923.
APA
Mollahosseini, A., Hasani, B., & Mahoor, M. H. (2019). AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild. IEEE Transactions on Affective Computing, 10(1), 18–31. https://doi.org/10.1109/TAFFC.2017.2740923
Chicago
Mollahosseini, A., B. Hasani, and M. H. Mahoor. 2019. “AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild”. IEEE Transactions on Affective Computing 10 (1): 18–31. https://doi.org/10.1109/TAFFC.2017.2740923.
Harvard
Mollahosseini, A., Hasani, B. and Mahoor, M.H. (2019) “AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild”, IEEE Transactions on Affective Computing, 10(1), pp. 18–31. Available at: https://doi.org/10.1109/TAFFC.2017.2740923.
Vancouver
1. Mollahosseini A, Hasani B, Mahoor MH (2019) AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild. IEEE Transactions on Affective Computing 10:18–31

BibTeX

@article{Mollahosseini_2019, title={AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild}, volume={10}, ISSN={2371-9850}, url={http://dx.doi.org/10.1109/TAFFC.2017.2740923}, DOI={10.1109/taffc.2017.2740923}, number={1}, journal={IEEE Transactions on Affective Computing}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Mollahosseini, Ali and Hasani, Behzad and Mahoor, Mohammad H.}, year={2019}, month=Jan, pages={18–31} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF