Micron-BERT: BERT-Based Facial Micro-Expression Recognition
Xuan-Bac NguyenChi Nhan DuongXin LiSusan GauchHan-Seok SeoKhoa Luu
Proposes a self-supervised BERT-based framework that captures subtle facial micro-expressions between video frames without manual annotations by pairing diagonal micro-attention with an unsupervised patch-of-interest detector.
Facial micro-expressions are brief, involuntary muscle movements lasting between a fraction of a second and half a second. Because they reveal genuine emotional states that individuals often attempt to suppress, recognizing them is critical in high-stakes fields such as criminal analysis and lie detection. However, automated micro-expression recognition is exceptionally challenging because these movements are subtle and short-lived, while existing computer vision models struggle to isolate tiny facial changes from background noise and head movements without requiring costly manual data labeling.
The article aims to introduce and evaluate Micron-BERT, a deep learning framework based on bidirectional transformers that automatically detects, localizes, and classifies facial micro-expressions without requiring pre-labeled facial landmarks or bounding boxes.
To achieve this, the authors designed a self-supervised approach that trains on unlabeled video frames. The method uses a patch-swapping strategy combined with two specialized components: Diagonal Micro-Attention, which pinpoints minute pixel differences across video frames, and a Patch of Interest module, which automatically forces the model to focus on salient facial regions while ignoring background interference. The system was pre-trained on an unlabeled dataset of eight million video frames and evaluated across four widely recognized benchmark datasets (CASME3, CASME II, SAMM, and SMIC) as well as a composite dataset spanning subjects of diverse ages, genders, and ethnicities.
The evaluation produced several notable findings. First, Micron-BERT significantly outperformed all prior state-of-the-art methods across all tested benchmarks. On the challenging CASME3 dataset, it achieved double-digit accuracy and F1 score improvements over previous baselines, demonstrating roughly 14 to 22 percentage-point gains across various emotion-classification settings. Second, on established benchmarks like CASME II, SAMM, and SMIC, the system consistently set new performance records, reaching an unweighted F1 score of 90.34% on CASME II and 85.50% on SMIC. Third, ablation testing demonstrated that combining both the micro-attention and facial region-focusing modules yielded an approximate 10 percentage-point performance jump over foundational transformer baselines, confirming that isolating facial regions from noisy backgrounds is essential for accurate recognition.
These findings demonstrate that self-supervised deep learning can master micro-movement detection directly from raw video, eliminating the heavy cost and time associated with manual data labeling and landmark pre-processing. By effectively suppressing background distractions and focusing on true facial movements, this approach provides a robust and scalable foundation for automated emotion and deception analysis in real-world environments.
Organizations developing automated behavioral and facial analysis tools should consider adopting self-supervised transformer architectures that incorporate localized attention mechanisms rather than relying on standard whole-image models. Future development should focus on enhancing the model's robustness against extreme lighting shifts, as unmoving facial areas exposed to sudden illumination changes can occasionally be misinterpreted as movement. Overall, the extensive experimental validation across diverse public datasets provides high confidence in the framework's effectiveness under standard video capture conditions.
- Paper: BEiT: BERT Pre-Training of Image Transformers, Hangbo Bao et al. (2022). BEiT establishes the masked image modeling framework that adapts BERT pre-training to vision transformers, forming the conceptual foundation for Micron-BERT's architecture.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). BERT introduces the bidirectional Transformer pre-training methodology that Micron-BERT adapts and customizes for unsupervised facial micro-expression recognition.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Masked Autoencoders provide foundational principles of self-supervised visual representation learning through patch masking and reconstruction that motivate Micron-BERT's design.
- Paper: Deep Facial Expression Recognition: A Survey, Shan Li et al. (2018). This survey provides essential background on deep learning architectures, spatio-temporal modeling, and benchmark challenges in facial expression recognition.
- Paper: Recognizing Action Units for Facial Expression Analysis, Ying-li Tian et al. (2001). This paper establishes the foundational principles of recognizing subtle, fine-grained localized facial muscle actions essential for understanding micro-expressions.
- Paper: ConvNeXt V2: Co-designing and Scaling ConvNets with Masked Autoencoders, Sanghyun Woo et al. (2023). ConvNeXt V2 advances self-supervised masked autoencoder architectures by incorporating global response normalization to prevent feature collapse, offering complementary architectural insights for visual representation learning.
- Paper: Efficient Multi-Scale Attention Module with Cross-Spatial Learning, Daliang Ouyang et al. (2023). This work explores efficient multi-scale attention and cross-spatial learning modules that complement and extend the spatial attention mechanisms developed in Micron-BERT.
