Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition

Hanyang WangBo LiShuang WuSiyuan ShenFeng LiuShouhong DingAimin Zhou

article2023CVPR96 citations

Proposes a multi-instance learning framework for dynamic facial expression recognition that treats non-target video frames as weakly supervised data and balances short- and long-term temporal dependencies to achieve state-of-the-art accuracy using a standard 3D CNN backbone.

Listen

Dynamic facial expression recognition in video is critical for applications in human-computer interaction, driver monitoring, and healthcare diagnostics. However, recognizing expressions in unconstrained real-world video remains difficult because video clips are typically labeled as a whole without specifying the exact frame intervals where emotions occur. Previous methods treated non-target frames merely as noise and failed to account for the difference between strong short-term facial movements and weaker long-term temporal connections.

The article demonstrates that video emotion recognition is fundamentally a weakly supervised learning problem and proposes a unified framework named Multi-3D Dynamic Facial Expression Learning to address inexact labels and temporal imbalances.

To evaluate this framework, the authors used a multi-instance learning approach where each video clip is treated as a collection of short 3D visual segments. The model extracts local motion features from multi-frame segments using a standard 3D convolutional network, models long-term temporal context using bidirectional recurrent networks and self-attention, and stabilizes feature representations using dynamic normalization. The system was trained and evaluated on two benchmark datasets containing over 16,000 and 38,000 video clips respectively, focusing on standard recognition recall metrics.

The experimental findings show that the proposed framework achieved top performance on both benchmarks, outperforming previous leading methods by 1.06 to 1.70 percentage points in overall recognition accuracy. Furthermore, segmenting videos into 3D multi-frame instances yielded substantially better accuracy than treating individual frames independently or processing entire videos as single sequences. The dynamic instance aggregation and normalization modules also provided clear performance gains over standard pooling baselines while maintaining high computational efficiency.

These results indicate that treating dynamic facial expression recognition as a weakly supervised multi-instance task offers superior accuracy and efficiency compared to standard video architectures. This design lowers computational requirements and reduces the reliance on costly, precise frame-by-frame annotations. The authors recommend adopting multi-instance 3D learning structures for video emotion analytics and exploring transfer learning, self-supervised pre-training, and micro-expression techniques to address remaining challenges with severe class imbalance and low-intensity facial expressions.

  • Paper: Micron-BERT: BERT-Based Facial Micro-Expression Recognition, Xuan-Bac Nguyen et al. (2023). Extends fine-grained temporal facial video analysis by applying self-supervised bidirectional transformer architectures to recognize subtle facial micro-expressions without manual landmark supervision.
Cover for Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition

Abstract

Dynamic Facial Expression Recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video format. Previous research has considered non-target frames as noisy frames, but we propose that it should be treated as a weakly supervised problem. We also identify the imbalance of short- and long-term temporal relationships in DFER. Therefore, we introduce the Multi-3D Dynamic Facial Expression Learning (M3DFEL) framework, which utilizes Multi-Instance Learning (MIL) to handle inexact labels. M3DFEL generates 3D-instances to model the strong short-term temporal relationship and utilizes 3DCNNs for feature extraction. The Dynamic Long-term Instance Aggregation Module (DLIAM) is then utilized to learn the long-term temporal relationships and dynamically aggregate the instances. Our experiments on DFEW and FERV39K datasets show that M3DFEL outperforms existing state-of-the-art approaches with a vanilla R3D18 backbone. The source code is available at https://github.com/faceeyes/M3DFEL.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Dynamic Facial Expression Recognition
  • 2.2. Multi-Instance Learning
  • 3. Method
  • 3.1. Overview
  • 3.2. Proposed Method
  • 4. Experiments
  • 4.1. Datasets
  • 4.2. Implementation Details
  • 4.3. Comparison with the State-of-the-art Methods
  • 4.4. Ablation Study
  • 4.5. Visualization
  • 5. Discussion
  • 6. Conclusion
  • 7. Acknowledgments
  • References

Knowls

  1. Knowl 1 — Multi-3D Dynamic Facial Expression Learning (M3DFEL) Framework

    model/method

    Dynamic Facial Expression Recognition (DFER) is formulated as a weakly supervised Multi-Instance Learning (MIL) problem to address videos containing inexact emotion labels and non-target facial expressions (such as neutral frames or speech movements). The M3DFEL framework decomposes the learning pipeline into four stages:

    1. 3D-Instance Generation: A video clip of TT frames is partitioned along the temporal dimension into NN short 3D clips (instances) rather than individual 2D frames, enabling the extraction of local temporal facial motions.
    2. Instance Feature Extraction: A 3D convolutional neural network (vanilla R3D-18) extracts feature representations for each 3D instance independently, producing an instance feature sequence F∈RN×CF \in \mathbb{R}^{N \times C}, where NN is the number of instances in the bag and CC is the feature channel dimension.
    3. Dynamic Long-term Instance Aggregation Module (DLIAM): A Bidirectional Long Short-Term Memory (BiLSTM) network and Multi-Head Self-Attention (MHSA) model weak long-term inter-instance dependencies across the video. Dynamic Multi-Instance Normalization (DMIN) stabilizes instance weights by enforcing consistency across instances and the overall bag.
    4. Classification: Instance features are weighted by the normalized attention scores, fused via a 1D convolution layer into a bag-level representation Z∈RN×CZ \in \mathbb{R}^{N \times C}, and mapped to class logits via a fully connected layer supervised with Cross-Entropy Loss and label smoothing.
  2. Knowl 2 — 3D-Instance Generation for Multi-Instance Learning in DFER

    model/method

    In Dynamic Facial Expression Recognition (DFER), framing video-level emotion recognition as standard 2D frame-level Multi-Instance Learning (MIL) fails because isolated frames (e.g., during speech or transitional facial movements) appear ambiguous or neutral without temporal context.

    To capture strong short-term temporal relationships, 3D-Instance Generation splits a video VV containing TT frames into a bag of NN 3D instances:

    I=[I1,I2,…,IN]I = [I_1, I_2, \dots, I_N]

    where each instance In∈RCin×Tsub×H×WI_n \in \mathbb{R}^{C_{in} \times T_{sub} \times H \times W} is a spatio-temporal sub-clip of Tsub=T/NT_{sub} = T / N consecutive frames (CinC_{in} channels, spatial height HH, width WW). Extracting features from these 3D instances using a 3D CNN enables the model to capture subtle facial muscle motions across consecutive frames and distinguish consistent emotional expressions from conversational movement.

  3. Knowl 3 — Dynamic Long-term Instance Aggregation Module (DLIAM)

    model/method

    The Dynamic Long-term Instance Aggregation Module (DLIAM) aggregates instance-level features into a bag-level representation while modeling weak long-term temporal dependencies across instances. Given an extracted bag of instance features F∈RN×CF \in \mathbb{R}^{N \times C} (where NN is the number of 3D instances and CC is the channel dimension):

    1. Temporal Context Encoding: A Bidirectional LSTM (BiLSTM) processes the sequence of instance representations FF to capture sequential transitions across time, generating sequence representations X∈RN×CX \in \mathbb{R}^{N \times C}.
    2. Inter-Instance Attention: A Multi-Head Self-Attention (MHSA) layer computes relational attention weights A∈RN×CA \in \mathbb{R}^{N \times C} across the instances.
    3. Dynamic Normalization: Dynamic Multi-Instance Normalization (DMIN) normalizes attention tensor AA across both bag and instance dimensions to yield A^∈RN×C\hat{A} \in \mathbb{R}^{N \times C}.
    4. Gated Aggregation: The normalized attention weights are passed through a Sigmoid activation and multiplied element-wise with the instance features XX. A 1D convolutional layer then reduces the weighted sequence into the bag-level feature representation Z∈RN×CZ \in \mathbb{R}^{N \times C}:

    Z=Conv1D(X⊙Sigmoid(A^))Z = \text{Conv1D}(X \odot \text{Sigmoid}(\hat{A}))

    where ⊙\odot denotes element-wise multiplication.

  4. Knowl 4 — Dynamic Multi-Instance Normalization (DMIN)

    equation

    To prevent unstable instance predictions across short time intervals and enforce temporal consistency, Dynamic Multi-Instance Normalization (DMIN) dynamically blends bag-level and instance-level statistics. For an attention matrix A∈RN×CA \in \mathbb{R}^{N \times C} (where NN is the number of instances and CC is the channel dimension), the normalized attention value A^nc\hat{A}_{nc} for the cc-th channel of the nn-th instance is:

    A^nc=Anc−∑k∈Kwkμk∑k∈Kwk′σk2+ϵ⋅γ+β\hat{A}_{nc} = \frac{A_{nc} - \sum_{k \in \mathcal{K}} w_k \mu_k}{\sqrt{\sum_{k \in \mathcal{K}} w'_k \sigma_k^2} + \epsilon} \cdot \gamma + \beta

    where K={bn,in}\mathcal{K} = \{bn, in\} contains the bag-level normalizer (bnbn) and instance-level normalizer (inin), ϵ>0\epsilon > 0 is a small constant for numerical stability, and γ,β\gamma, \beta are learnable affine parameters.

    The bag-level statistics compute the mean and variance across both instance (NN) and channel (CC) dimensions:

    μbn=1NC∑n=1N∑c=1CAnc,σbn2=1NC∑n=1N∑c=1C(Anc−μbn)2\mu_{bn} = \frac{1}{NC} \sum_{n=1}^{N} \sum_{c=1}^{C} A_{nc}, \quad \sigma_{bn}^2 = \frac{1}{NC} \sum_{n=1}^{N} \sum_{c=1}^{C} (A_{nc} - \mu_{bn})^2

    where μbn,σbn∈R1\mu_{bn}, \sigma_{bn} \in \mathbb{R}^1.

    The instance-level statistics compute the mean and variance across the instance dimension (NN) independently for each channel:

    μin=1N∑n=1NAnc,σin2=1N∑n=1N(Anc−μin)2\mu_{in} = \frac{1}{N} \sum_{n=1}^{N} A_{nc}, \quad \sigma_{in}^2 = \frac{1}{N} \sum_{n=1}^{N} (A_{nc} - \mu_{in})^2

    where μin,σin∈RC\mu_{in}, \sigma_{in} \in \mathbb{R}^C.

    The importance weights wkw_k and wk′w'_k are dynamically adjusted using learnable parameters λk,λk′\lambda_k, \lambda'_k via softmax:

    wk=eλk∑j∈Keλj,∑k∈Kwk=1,∑k∈Kwk′=1w_k = \frac{e^{\lambda_k}}{\sum_{j \in \mathcal{K}} e^{\lambda_j}}, \quad \sum_{k \in \mathcal{K}} w_k = 1, \quad \sum_{k \in \mathcal{K}} w'_k = 1

  5. Knowl 5 — Experimental Setup and Implementation Details for M3DFEL

    experimental setup

    The M3DFEL framework is evaluated on two in-the-wild dynamic facial expression recognition datasets:

    • DFEW: Comprises over 16,000 video clips from >1,500 movies annotated with seven basic emotion classes (Happy, Sad, Neutral, Angry, Surprise, Disgust, Fear) and evaluated using 5-fold cross-validation.
    • FERV39K: Comprises 38,935 video clips across 4 scenarios and 22 fine-grained scenes with the same seven emotion categories, evaluated on official train/test splits.

    Implementation Details:

    • Backbone: Vanilla R3D-18 initialized with pre-trained Torchvision weights.
    • Input Sampling: 16 frames sampled per video clip, divided into N=4N=4 instances of 4 frames each.
    • Data Augmentation: Random cropping, horizontal flipping, and color jitter with an intensity factor of 0.4.
    • Optimization: AdamW optimizer with cosine learning rate decay for 300 epochs (including 20 warmup epochs); initial learning rate is 5×10−45 \times 10^{-4}, minimum learning rate is 5×10−65 \times 10^{-6}, and weight decay is 0.050.05.
    • Batch Size and Regularization: Batch size of 256; label smoothing parameter of 0.10.1.
    • Metrics: Weighted Average Recall (WAR, overall accuracy) and Unweighted Average Recall (UAR, balanced mean per-class accuracy).
  6. Knowl 6 — Comparison of M3DFEL with State-of-the-Art Methods on DFEW Benchmark

    data/table

    Performance comparison on the 5-fold cross-validation split of the DFEW dataset across individual emotion categories (Happy, Sad, Neutral, Angry, Surprise, Disgust, Fear), Weighted Average Recall (WAR), Unweighted Average Recall (UAR), and computational complexity in GFLOPs:

    Method Hap. Sad Neu. Ang. Sur. Dis. Fea. WAR(%) UAR(%) FLOPs(G)
    C3D 75.17 39.49 55.11 62.49 45.00 1.38 20.51 53.54 42.74 38.57
    P3D 74.85 43.40 54.18 60.42 50.99 0.69 23.28 54.47 43.97 -
    I3D 78.61 44.19 56.69 55.87 45.88 2.07 20.51 54.27 43.40 6.99
    R(2+1)D18 79.67 39.07 57.66 50.39 48.26 3.45 21.06 53.22 42.79 42.36
    3D ResNet18 73.13 48.26 50.51 64.75 50.10 0.00 26.39 54.98 44.73 8.32
    ResNet18+LSTM 78.00 40.65 53.77 56.83 45.00 4.14 21.62 53.08 42.86 7.78
    EC-STFL 79.18 49.05 57.85 60.98 46.15 2.76 21.51 56.51 45.35 8.32
    FormerDFER 84.05 62.57 67.52 70.03 56.43 3.45 31.78 65.70 53.69 9.11
    STT 87.36 67.90 64.97 71.24 53.10 3.49 34.04 66.45 54.58 -
    DPCNet - - - - - - - 66.32 55.02 9.52
    NR-DFERNet 88.47 64.84 70.03 75.09 61.60 0.00 19.43 68.19 54.21 6.33
    M3DFEL (Ours) 89.59 68.38 67.88 74.24 59.69 0.00 31.63 69.25 56.10 1.65

    Using a standard vanilla R3D-18 backbone, M3DFEL achieves state-of-the-art accuracy, outperforming NR-DFERNet by +1.06%+1.06\% in WAR and +1.89%+1.89\% in UAR while requiring only 1.65 GFLOPs1.65\text{ GFLOPs} (compared to 6.33 GFLOPs6.33\text{ GFLOPs} for NR-DFERNet and 9.11 GFLOPs9.11\text{ GFLOPs} for FormerDFER).

  7. Knowl 7 — Comparison of M3DFEL with State-of-the-Art Methods on FERV39K Benchmark

    data/table

    Performance comparison on the large-scale FERV39K benchmark dataset evaluated by Weighted Average Recall (WAR) and Unweighted Average Recall (UAR):

    Method WAR(%) UAR(%)
    C3D 31.69 22.68
    P3D 33.39 23.20
    I3D 38.78 30.17
    R(2+1)D18 41.28 31.55
    3D ResNet18 37.57 26.67
    R18+LSTM 42.95 30.92
    2R18+LSTM 43.20 31.28
    NR-DFERNet 45.97 33.99
    M3DFEL (Ours) 47.67 35.94

    M3DFEL achieves the highest recognition performance on FERV39K, outperforming NR-DFERNet by +1.70%+1.70\% WAR and +1.95%+1.95\% UAR. In addition, M3DFEL surpasses baseline architectures with identical backbones (3D ResNet18 and R18+LSTM) by +10.10%+10.10\% and +4.72%+4.72\% WAR, respectively.

  8. Knowl 8 — Effect of Bag Size and 3D Temporal Partitioning in M3DFEL

    data/table

    Ablation experiments on the DFEW dataset evaluating the impact of bag size NN when sampling 16 frames per video:

    Bag Size (NN) WAR(%) UAR(%)
    1 68.04 55.36
    2 68.55 55.92
    4 69.25 56.10
    8 68.24 55.32
    16 66.36 53.56
    • Bag Size 1: The entire 16-frame video is passed directly to the 3D CNN without MIL instance partitioning, bypassing the aggregation module and achieving 68.04%68.04\% WAR.
    • Bag Size 16: Degrading the 3DMIL model to 2D frame-level MIL (ResNet18 on individual frames) yields the lowest performance (66.36%66.36\% WAR, 53.56%53.56\% UAR) due to the lack of short-term temporal motion capture.
    • Bag Size 4: Partitioning 16 frames into 4 instances of 4 frames each provides the best trade-off between strong short-term temporal feature extraction and dynamic long-term aggregation (69.25%69.25\% WAR, 56.10%56.10\% UAR).
  9. Knowl 9 — Ablation Study of Dynamic Long-term Instance Aggregation Module (DLIAM)

    data/table

    Ablation analysis on the DFEW dataset examining the performance of different components within the Dynamic Long-term Instance Aggregation Module (DLIAM):

    Setting Module Configuration WAR(%) UAR(%)
    a Baseline (Average Pooling) 68.23 55.44
    b w/o LSTM 68.63 55.62
    c w/o DMIN 68.91 56.03
    d w/o MHSA 69.13 56.21
    e M3DFEL (Full DLIAM) 69.25 56.10
    • Replacing standard Average Pooling (baseline) with the full DLIAM module yields an increase of +1.02%+1.02\% in WAR and +0.66%+0.66\% in UAR.
    • Removing the BiLSTM (setting b) causes the largest performance degradation among DLIAM submodules (68.63%68.63\% WAR), demonstrating that sequence-aware temporal modeling across instances is superior to time-independent attention.
    • Adding Dynamic Multi-Instance Normalization (DMIN) improves accuracy by +0.34%+0.34\% WAR over setting c (69.25%69.25\% vs. 68.91%68.91\%), showing that dynamic feature normalization stabilizes instance aggregation.
  10. Knowl 10 — Limitations and Error Analysis in In-The-Wild Dynamic Facial Expression Recognition

    limitation

    Evaluation and confusion matrix analyses of M3DFEL on DFEW highlight critical unresolved challenges in in-the-wild DFER:

    1. Severe Class Imbalance: In DFEW, Disgust represents only 1.22%1.22\% and Fear represents 8.14%8.14\% of the data. M3DFEL achieves 0.00%0.00\% accuracy on Disgust and 31.63%31.63\% on Fear across 5 folds, as the loss objective heavily prioritizes majority classes.
    2. Neutral Prediction Bias on Subtle Expressions: Low-intensity emotional expressions (micro-expressions) exhibit subtle muscle movements that resemble Neutral faces. The model frequently defaults to predicting Neutral for ambiguous clips as a lower-risk prediction under classification uncertainty.
    3. Classification Stage Bottleneck: Failure case analysis indicates that MIL instance aggregation successfully identifies non-neutral segments within predominantly neutral video bags, but the backend classifier frequently misclassifies the specific emotion category (e.g., predicting Surprise instead of Fear).

Coverage note — Visual t-SNE feature cluster plots and qualitative single-instance prediction strips were omitted in favor of the quantitative benchmarks, full ablation tables, and the synthesized error analysis knowl.

References

  1. 1.Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017.
  2. 2.Zhanli Chen, Rashid Ansari, and Diana J. Wilkie. Learning pain from action unit combinations: A weakly supervised approach via multiple instance learning. IEEE Transactions on Affective Computing, 13(1):135–146, 2022.
  3. 3.Zhaoyu Chen, Bo Li, Shuang Wu, Jianghe Xu, Shouhong Ding, and Wenqiang Zhang. Shape matters: Deformable patch attack. In Shai Avidan, Gabriel J. Brostow, Moustapha Cisse, Giovanni Maria Farinella, and Tal Hassner, editors, Computer Vision - ECCV 2022. Springer, 2022.
  4. 4.Zhaoyu Chen, Bo Li, Jianghe Xu, Shuang Wu, Shouhong Ding, and Wenqiang Zhang. Towards practical certifiable patch defense with vision transformer. In Proceedings of the IEEE/CVF Conference on CVPR, pages 15148–15158, June 2022.
  5. 5.Yin Fan, Xiangju Lu, Dian Li, and Yuanliu Liu. Video-based emotion recognition using cnn-rnn and c3d hybrid networks. In Proceedings of the 18th ACM international conference on multimodal interaction, pages 445–450, 2016.
  6. 6.Zhixin Fang, Libai Cai, and Gang Wang. Metahuman creator the starting point of the metaverse. In 2021 International Symposium on Computer Technology and Information Science (ISCTIS), pages 154–157. IEEE, 2021.
  7. 7.Xiaoxu Feng, Xiwen Yao, Gong Cheng, and Junwei Han. Weakly supervised rotation-invariant aerial object detection network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14146–14155, 2022.
  8. 8.Michael Gadermayr and Maximilian Tschuchnig. Multiple instance learning for digital pathology: A review on the state-of-the-art, limitations & future potential. arXiv preprint arXiv:2206.04425, 2022.
  9. 9.Shuyong Gao, Wei Zhang, Yan Wang, Qianyu Guo, Chenglong Zhang, Yangji He, and Wenqiang Zhang. Weakly-supervised salient object detection using point supervison. arXiv preprint arXiv:2203.11652, 2022.
  10. 10.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Can spatiotemporal 3d cnns retrace the history of 2d cnns and imagenet? In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 6546–6555, 2018.
  11. 11.Xingxun Jiang, Yuan Zong, Wenming Zheng, Chuangao Tang, Wanchuang Xia, Cheng Lu, and Jiateng Liu. Dfew: A large-scale database for recognizing dynamic facial expressions in the wild. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2881–2889, 2020.
  12. 12.Ziv Lautman and Shahar Lev-Ari. The use of smart devices for mental health diagnosis and care, 2022.
  13. 13.Jiyoung Lee, Seungryong Kim, Sunok Kim, Jungin Park, and Kwanghoon Sohn. Context-aware emotion recognition networks. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10143–10152, 2019.
  14. 14.Min Kyu Lee, Dong Yoon Choi, Dae Ha Kim, and Byung Cheol Song. Visual scene-aware hybrid neural network architecture for video-based facial expression recognition. In 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019), pages 1–8, 2019.
  15. 15.Bo Li, Zhengxing Sun, and Yuqi Guo. Supervae: Superpixelwise variational autoencoder for salient object detection. In The Thirty-Third AAAI Conference, 2019.
  16. 16.Bo Li, Zhengxing Sun, Qian Li, Yunjie Wu, and Anqi Hu. Group-wise deep object co-segmentation with co-attention recurrent neural network. In 2019 IEEE/CVF International Conference on ICCV, pages 8518–8527. IEEE, 2019.
  17. 17.Bo Li, Zhengxing Sun, Lv Tang, and Anqi Hu. Two-b-real net: Two-branch network for real-time salient object detection. In IEEE International Conference on ICASSP. IEEE, 2019.
  18. 18.Bo Li, Zhengxing Sun, Lv Tang, Yunhan Sun, and Jinlong Shi. Detecting robust co-saliency with recurrent co-attention neural network. In Sarit Kraus, editor, IJCAI, 2019.
  19. 19.Bo Li, Zhengxing Sun, Quan Wang, and Qian Li. Co-saliency detection based on hierarchical consistency. In Laurent Amsaleg, Benoit Huet, Martha A. Larson, Guillaume Gravier, Hayley Hung, Chong-Wah Ngo, and Wei Tsang Ooi, editors, Proceedings of the 27th ACM International Conference on MM, pages 1392–1400. ACM, 2019.
  20. 20.Bo Li, Lv Tang, Senyun Kuang, Mofei Song, and Shouhong Ding. Toward stable co-saliency detection and object co-segmentation. IEEE Trans. Image Process., 31:6532–6547, 2022.
  21. 21.Hanting Li, Hongjing Niu, Zhaoqing Zhu, and Feng Zhao. Intensity-aware loss for dynamic facial expression recognition in the wild. arXiv preprint arXiv:2208.10335, 2022.
  22. 22.Hanting Li, Mingzhe Sui, Zhaoqing Zhu, et al. Nr-dfernet: Noise-robust network for dynamic facial expression recognition. arXiv preprint arXiv:2206.04975, 2022.
  23. 23.Hangyu Li, Nannan Wang, Xi Yang, Xiaoyu Wang, and Xinbo Gao. Towards semi-supervised deep facial expression recognition with an adaptive confidence margin. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4166–4175, 2022.
  24. 24.Zuojin Li, Liukui Chen, Ling Nie, and Simon X Yang. A novel learning model of driver fatigue features representation for steering wheel angle. IEEE Transactions on Vehicular Technology, 71(1):269–281, 2021.
  25. 25.Feng Liu, Si-Yuan Shen, Zi-Wang Fu, Han-Yang Wang, Ai-Min Zhou, and Jia-Yin Qi. Lgcct: A light gated and crossed complementation transformer for multimodal speech emotion recognition. Entropy, 24(7):1010, 2022.
  26. 26.Feng Liu, Hanyang Wang, Jiahao Zhang, Ziwang Fu, Aimin Zhou, Jiayin Qi, and Zhibin Li. Evogan: An evolutionary computation assisted gan. Neurocomputing, 469:81–90, 2022.
  27. 27.Feng Liu, Han-Yang Wang, Si-Yuan Shen, Xun Jia, Jing-Yi Hu, Jia-Hao Zhang, Xi-Yi Wang, Ying Lei, Ai-Min Zhou, Jia-Yin Qi, et al. Opo-fcm: A computational affection based occ-pad-ocean federation cognitive modeling approach. IEEE Transactions on Computational Social Systems, 2022.
  28. 28.Cheng Lu, Wenming Zheng, Chaolong Li, Chuangao Tang, Suyuan Liu, Simeng Yan, and Yuan Zong. Multiple spatio-temporal feature learning for video-based emotion recognition in the wild. In Proceedings of the 20th ACM International Conference on Multimodal Interaction, ICMI ’18, page 646–652, New York, NY, USA, 2018. Association for Computing Machinery.
  29. 29.Ping Luo, Jiamin Ren, Zhanglin Peng, Ruimao Zhang, and Jingyu Li. Differentiable learning-to-normalize via switchable normalization. In International Conference on Learning Representations, 2018.
  30. 30.Zhekun Luo, Devin Guillory, Baifeng Shi, Wei Ke, Fang Wan, Trevor Darrell, and Huijuan Xu. Weakly-supervised action localization with expectation-maximization multi-instance learning. In European conference on computer vision, pages 729–745. Springer, 2020.
  31. 31.Fuyan Ma, Bin Sun, and Shutao Li. Spatio-temporal transformer for dynamic facial expression recognition in the wild. arXiv preprint arXiv:2205.04749, 2022.
  32. 32.Zhaofan Qiu, Ting Yao, and Tao Mei. Learning spatio-temporal representation with pseudo-3d residual networks. In proceedings of the IEEE International Conference on Computer Vision, pages 5533–5541, 2017.
  33. 33.Luca Romeo, Andrea Cavallo, Lucia Pepa, Nadia Bianchi-Berthouze, and Massimiliano Pontil. Multiple instance learning for emotion recognition using physiological signals. IEEE Transactions on Affective Computing, 13(1):389–407, 2022.
  34. 34.Siyuan Shen, Feng Liu, and Aimin Zhou. Mingling or misalignment? temporal shift for speech emotion recognition with pre-trained representations. arXiv preprint arXiv:2302.13277, 2023.
  35. 35.Lv Tang and Bo Li. CLASS: cross-level attention and supervision for salient objects detection. In Hiroshi Ishikawa, Cheng-Lin Liu, Tomas Pajdla, and Jianbo Shi, editors, ACCV 2020, 2020.
  36. 36.Lv Tang and Bo Li. Cosformer: Detecting co-salient object with transformers. arXiv preprint arXiv:2104.14729, 2021.
  37. 37.Lv Tang, Bo Li, Senyun Kuang, Mofei Song, and Shouhong Ding. Re-thinking the relations in co-saliency detection. IEEE Transactions on Circuits and Systems for Video Technology, 2022.
  38. 38.Lv Tang, Bo Li, Yijie Zhong, Shouhong Ding, and Mofei Song. Disentangled high quality salient object detection. In 2021 IEEE/CVF ICCV, pages 3560–3570. IEEE, 2021.
  39. 39.Peng Tang, Xinggang Wang, Xiang Bai, and Wenyu Liu. Multiple instance detection network with online instance classifier refinement. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2843–2851, 2017.
  40. 40.Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. Learning spatiotemporal features with 3d convolutional networks. In Proceedings of the IEEE international conference on computer vision, pages 4489–4497, 2015.
  41. 41.Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. A closer look at spatiotemporal convolutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 6450–6459, 2018.
  42. 42.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  43. 43.Kai Wang, Xiaojiang Peng, Jianfei Yang, Shijian Lu, and Yu Qiao. Suppressing uncertainties for large-scale facial expression recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6897–6906, 2020.
  44. 44.Weijie Wang, Nicu Sebe, and Bruno Lepri. Rethinking the learning paradigm for facial expression recognition. arXiv preprint arXiv:2209.15402, 2022.
  45. 45.Yan Wang, Wei Song, Wei Tao, Antonio Liotta, Dawei Yang, Xinlei Li, Shuyong Gao, Yixuan Sun, Weifeng Ge, Wei Zhang, et al. A systematic review on affective computing: Emotion models, databases, and recent advances. Information Fusion, 2022.
  46. 46.Yan Wang, Yixuan Sun, Yiwen Huang, Zhongying Liu, Shuyong Gao, Wei Zhang, Weifeng Ge, and Wenqiang Zhang. Ferv39k: A large-scale multi-scene dataset for facial expression recognition in videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20922–20931, 2022.
  47. 47.Yan Wang, Yixuan Sun, Wei Song, Shuyong Gao, Yiwen Huang, Zhaoyu Chen, Weifeng Ge, and Wenqiang Zhang. Dpcnet: Dual path multi-excitation collaborative network for facial expression representation learning in videos. In Proceedings of the 30th ACM International Conference on Multimedia, pages 101–110, 2022.
  48. 48.Chongliang Wu, Shangfei Wang, and Qiang Ji. Multi-instance hidden markov model for facial expression recognition. In 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), volume 1, pages 1–6, 2015.
  49. 49.Hongrun Zhang, Yanda Meng, Yitian Zhao, Yihong Qiao, Xiaoyun Yang, Sarah E Coupland, and Yalin Zheng. Dtfd-mil: Double-tier feature distillation multiple instance learning for histopathology whole slide image classification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18802–18812, 2022.
  50. 50.Jie Zhang, Chen Chen, Bo Li, Lingjuan Lyu, Shuang Wu, Shouhong Ding, Chunhua Shen, and Chao Wu. DENSE: Data-free one-shot federated learning. In Advances in NeurIPS, 2022.
  51. 51.Jie Zhang, Zhiqi Li, Bo Li, Jianghe Xu, Shuang Wu, Shouhong Ding, and Chao Wu. Federated learning with label distribution skew via logits calibration. In Proceedings of the ICML. PMLR, 2022.
  52. 52.Tong Zhang, Wenming Zheng, Zhen Cui, Yuan Zong, and Yang Li. Spatial–temporal recurrent neural network for emotion recognition. IEEE transactions on cybernetics, 49(3):839–847, 2018.
  53. 53.Yu Zhao, Fan Yang, Yuqi Fang, Hailing Liu, Niyun Zhou, Jun Zhang, Jiarui Sun, Sen Yang, Bjoern Menze, Xinjuan Fan, et al. Predicting lymph node metastasis using histopathological images based on multiple instance learning with deep graph convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4837–4846, 2020.
  54. 54.Zengqun Zhao and Qingshan Liu. Former-dfer: Dynamic facial expression recognition transformer. In Proceedings of the 29th ACM International Conference on Multimedia, pages 1553–1561, 2021.
  55. 55.Jiawen Zheng, Bo Li, ShengChuan Zhang, Shuang Wu, Liujuan Cao, and Shouhong Ding. Attack can benefit: An adversarial approach to recognizing facial expressions under noisy annotations. In AAAI Conference, 2023.
  56. 56.Yijie Zhong, Bo Li, Lv Tang, Senyun Kuang, Shuang Wu, and Shouhong Ding. Detecting camouflaged object in frequency domain. In Proceedings of the IEEE/CVF Conference on CVPR, pages 4504–4513, June 2022.
  57. 57.Yijie Zhong, Bo Li, Lv Tang, Hao Tang, and Shouhong Ding. Highly efficient natural image matting. CoRR, abs/2110.12748, 2021.

Citation

MLA
Wang, H., et al. “Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 17958–68, https://doi.org/10.1109/CVPR52729.2023.01722.
APA
Wang, H., Li, B., Wu, S., Shen, S., Liu, F., Ding, S., & Zhou, A. (2023). Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17958–17968. https://doi.org/10.1109/CVPR52729.2023.01722
Chicago
Wang, H., B. Li, S. Wu, et al. 2023. “Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 17958–68. https://doi.org/10.1109/CVPR52729.2023.01722.
Harvard
Wang, H. et al. (2023) “Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 17958–17968. Available at: https://doi.org/10.1109/CVPR52729.2023.01722.
Vancouver
1. Wang H, Li B, Wu S, Shen S, Liu F, Ding S, Zhou A (2023) Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 17958–17968

BibTeX

@inproceedings{Wang_2023, title={Rethinking the Learning Paradigm for Dynamic Facial Expression Recognition}, url={http://dx.doi.org/10.1109/CVPR52729.2023.01722}, DOI={10.1109/cvpr52729.2023.01722}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wang, Hanyang and Li, Bo and Wu, Shuang and Shen, Siyuan and Liu, Feng and Ding, Shouhong and Zhou, Aimin}, year={2023}, month=June, pages={17958–17968} }
Metadata:Crossref

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE