Micron-BERT: BERT-Based Facial Micro-Expression Recognition

Xuan-Bac NguyenChi Nhan DuongXin LiSusan GauchHan-Seok SeoKhoa Luu

article2023CVPR126 citations

Proposes a self-supervised BERT-based framework that captures subtle facial micro-expressions between video frames without manual annotations by pairing diagonal micro-attention with an unsupervised patch-of-interest detector.

Listen

Facial micro-expressions are brief, involuntary muscle movements lasting between a fraction of a second and half a second. Because they reveal genuine emotional states that individuals often attempt to suppress, recognizing them is critical in high-stakes fields such as criminal analysis and lie detection. However, automated micro-expression recognition is exceptionally challenging because these movements are subtle and short-lived, while existing computer vision models struggle to isolate tiny facial changes from background noise and head movements without requiring costly manual data labeling.

The article aims to introduce and evaluate Micron-BERT, a deep learning framework based on bidirectional transformers that automatically detects, localizes, and classifies facial micro-expressions without requiring pre-labeled facial landmarks or bounding boxes.

To achieve this, the authors designed a self-supervised approach that trains on unlabeled video frames. The method uses a patch-swapping strategy combined with two specialized components: Diagonal Micro-Attention, which pinpoints minute pixel differences across video frames, and a Patch of Interest module, which automatically forces the model to focus on salient facial regions while ignoring background interference. The system was pre-trained on an unlabeled dataset of eight million video frames and evaluated across four widely recognized benchmark datasets (CASME3, CASME II, SAMM, and SMIC) as well as a composite dataset spanning subjects of diverse ages, genders, and ethnicities.

The evaluation produced several notable findings. First, Micron-BERT significantly outperformed all prior state-of-the-art methods across all tested benchmarks. On the challenging CASME3 dataset, it achieved double-digit accuracy and F1 score improvements over previous baselines, demonstrating roughly 14 to 22 percentage-point gains across various emotion-classification settings. Second, on established benchmarks like CASME II, SAMM, and SMIC, the system consistently set new performance records, reaching an unweighted F1 score of 90.34% on CASME II and 85.50% on SMIC. Third, ablation testing demonstrated that combining both the micro-attention and facial region-focusing modules yielded an approximate 10 percentage-point performance jump over foundational transformer baselines, confirming that isolating facial regions from noisy backgrounds is essential for accurate recognition.

These findings demonstrate that self-supervised deep learning can master micro-movement detection directly from raw video, eliminating the heavy cost and time associated with manual data labeling and landmark pre-processing. By effectively suppressing background distractions and focusing on true facial movements, this approach provides a robust and scalable foundation for automated emotion and deception analysis in real-world environments.

Organizations developing automated behavioral and facial analysis tools should consider adopting self-supervised transformer architectures that incorporate localized attention mechanisms rather than relying on standard whole-image models. Future development should focus on enhancing the model's robustness against extreme lighting shifts, as unmoving facial areas exposed to sudden illumination changes can occasionally be misinterpreted as movement. Overall, the extensive experimental validation across diverse public datasets provides high confidence in the framework's effectiveness under standard video capture conditions.

Cover for Micron-BERT: BERT-Based Facial Micro-Expression Recognition

Abstract

Micro-expression recognition is one of the most challenging topics in affective computing. It aims to recognize tiny facial movements difficult for humans to perceive in a brief period, i.e., 0.25 to 0.5 seconds. Recent advances in pre-training deep Bidirectional Transformers (BERT) have significantly improved self-supervised learning tasks in computer vision. However, the standard BERT in vision problems is designed to learn only from full images or videos, and the architecture cannot accurately detect details of facial micro-expressions. This paper presents Micron-BERT (µ-BERT), a novel approach to facial micro-expression recognition. The proposed method can automatically capture these movements in an unsupervised manner based on two key ideas. First, we employ Diagonal Micro-Attention (DMA) to detect tiny differences between two frames. Second, we introduce a new Patch of Interest (PoI) module to localize and highlight micro-expression interest regions and simultaneously reduce noisy backgrounds and distractions. By incorporating these components into an end-to-end deep network, the proposed µ-BERT significantly outperforms all previous work in various micro-expression tasks. µ-BERT can be trained on a large-scale unlabeled dataset, i.e., up to 8 million images, and achieves high accuracy on new unseen facial micro-expression datasets. Empirical experiments show µ-BERT consistently outperforms state-of-the-art performance on four micro-expression benchmarks, including SAMM, CASME II, SMIC, and CASME3, by significant margins. Code will be available at https://github.com/uark-cviu/Micron-BERT

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. BERT Revisited
  • 3.1. BERT in Vision Problems
  • 3.2. Limitations of BERT in Vision Problems
  • 4. The Proposed µ-BERT Approach
  • 4.1. Non-overlapping Patches Representation
  • 4.2. µ-Encoder
  • 4.3. µ-Decoder
  • 4.4. Blockwise Swapping
  • 4.5. Diagonal Micro Attention (DMA)
  • 4.6. Patch of Interest (POI)
  • 4.7. Loss Functions
  • 5. Experimental Results
  • 5.1. Datasets and Protocols
  • 5.2. Micro-Expression Self-Training
  • 5.3. Micro-Expression Recognition
  • 5.4. Results
  • 5.5. How µ-BERT perceives micro-movements
  • 5.6. Ablation studies
  • 6. Conclusions and Discussions
  • References

Knowls

  1. Knowl 1 — Micron-BERT Architecture for Facial Micro-Expression Recognition

    model/method

    Micron-BERT (μ\mu-BERT) is an end-to-end self-supervised bidirectional transformer framework designed to capture subtle facial micro-movements across consecutive video frames without requiring tokenizers or explicit facial landmark annotations. Given two video frames It,It+δ∈RH×W×CI_t, I_{t+\delta} \in \mathbb{R}^{H \times W \times C} separated by a temporal index δ\delta, each image is partitioned into Np=HW/ps2N_p = HW / ps^2 non-overlapping patches of size ps×psps \times ps. Each patch ptip_t^i is linearly projected into a latent vector zti=α(pti)+e(i)∈R1×d\mathbf{z}_t^i = \alpha(p_t^i) + \mathbf{e}(i) \in \mathbb{R}^{1 \times d}, where α\alpha is a projection embedding network and e(i)\mathbf{e}(i) is a fixed positional embedding, forming the latent sequence Zt=[zt0,zt1,…,ztNp−1]∈RNp×d\mathbf{Z}_t = [\mathbf{z}_t^0, \mathbf{z}_t^1, \dots, \mathbf{z}_t^{N_p-1}] \in \mathbb{R}^{N_p \times d}.

    The μ\mu-BERT architecture consists of five core components:

    1. μ\mu-Encoder (E\mathcal{E}): A stack of LeL_e transformer blocks composed of alternating Multi-Head Attention (MHA) and Multi-Layer Perceptron (MLP) layers with Layer Normalization (LN), producing latent representations Pt=E(Zt)∈RNp×d\mathbf{P}_t = \mathcal{E}(\mathbf{Z}_t) \in \mathbb{R}^{N_p \times d}.
    2. Blockwise Swapping: A data augmentation step that randomly exchanges rectangular blocks of corresponding patches between ItI_t and It+δI_{t+\delta} to yield a perturbed composite image It/sI_{t/s} containing patches Pt/s\mathcal{P}_{t/s}.
    3. Diagonal Micro-Attention (DMA): An attention mechanism that measures patch-wise correlations between the perturbed frame It/sI_{t/s} and the future frame It+δI_{t+\delta}, extracting the diagonal attention weights to identify swapped and micro-movement regions.
    4. Patch of Interest (POI): A self-supervised saliency module that learns to focus on facial regions and filter out background noise using a contextual token and a contextual agreement loss.
    5. μ\mu-Decoder (D\mathcal{D}): A symmetric transformer decoder with LdL_d blocks followed by a linear projection and reshaping operation that reconstructs the original frame ItI_t from the processed latent representations Pt\mathbf{P}_t as yt′∈RH×W×C\mathbf{y}'_t \in \mathbb{R}^{H \times W \times C}.
  2. Knowl 2 — Diagonal Micro-Attention Mechanism

    model/method

    Diagonal Micro-Attention (DMA) is an attention mechanism that detects micro-disparities between two frames ItI_t and It+δI_{t+\delta} after Blockwise Swapping creates a swapped patch sequence Pt/s\mathbf{P}_{t/s}. The sequence Pt/s\mathbf{P}_{t/s} contains both unchanged patches from ItI_t and swapped patches originating from It+δI_{t+\delta}.

    To identify and weigh the patches containing micro-movements without requiring supervised ground-truth swap indices, an attention matrix A^∈RNp×Np\hat{A} \in \mathbb{R}^{N_p \times N_p} is computed between the query projections of Pt+δ\mathbf{P}_{t+\delta} and the key projections of Pt/s\mathbf{P}_{t/s}: A^=softmax(Q(Pt+δ)⊗K(Pt/s)T),∑j=0Np−1A^(i,j)=1\hat{A} = \text{softmax}\left(Q(\mathbf{P}_{t+\delta}) \otimes K(\mathbf{P}_{t/s})^T\right), \quad \sum_{j=0}^{N_p-1} \hat{A}(i,j) = 1 where Q(⋅)Q(\cdot) and K(⋅)K(\cdot) are linear query and key projection functions, ⊗\otimes denotes matrix multiplication, and NpN_p is the number of image patches.

    Because corresponding patches from the same spatial position in It+δI_{t+\delta} and It/sI_{t/s} yield high correlation scores when a patch has been swapped into It/sI_{t/s} from It+δI_{t+\delta}, the diagonal elements satisfy A^(i,i)>A^(j,j)\hat{A}(i,i) > \hat{A}(j,j) for all swapped patches pt/si∈Pt+δp_{t/s}^i \in \mathcal{P}_{t+\delta} relative to unswapped patches pt/sj∈Ptp_{t/s}^j \in \mathcal{P}_t. The diagonal vector diag(A^)∈RNp×1\text{diag}(\hat{A}) \in \mathbb{R}^{N_p \times 1} is used as an importance weight vector to modulate the value projections V(Pt/s)V(\mathbf{P}_{t/s}): Pdma=diag(A^)×V(Pt/s)\mathbf{P}_{dma} = \text{diag}(\hat{A}) \times V(\mathbf{P}_{t/s}) where ×\times denotes element-wise multiplication across patch tokens.

  3. Knowl 3 — Patch of Interest Module and Saliency Weighting

    model/method

    The Patch of Interest (POI) module isolates salient facial regions from background noise without using bounding boxes, facial landmark detectors, or segmentation masks. It prepends a learnable Contextual Token zCT\mathbf{z}^{CT} to the patch sequence before passing through the transformer encoder E\mathcal{E}. As the sequence traverses the transformer layers, zCT\mathbf{z}^{CT} aggregates contextual information from all spatial patches zti\mathbf{z}_t^i.

    To force zCT\mathbf{z}^{CT} to capture subject-centric facial features rather than background context, an agreement loss enforces consistency between the contextual representation of frame It+δI_{t+\delta} and a randomly cropped version Crop(It+δ)\text{Crop}(I_{t+\delta}): Lagg=MSE(pt+δCT,pt+δ/cropCT)\mathcal{L}_{agg} = \text{MSE}\left(\mathbf{p}^{CT}_{t+\delta}, \mathbf{p}^{CT}_{t+\delta/crop}\right) where pt+δCT\mathbf{p}^{CT}_{t+\delta} and pt+δ/cropCT\mathbf{p}^{CT}_{t+\delta/crop} denote the output contextual token features for the full frame and cropped frame, respectively, and MSE\text{MSE} is the Mean Squared Error.

    The facial saliency score vector St+δ\mathbf{S}_{t+\delta} is extracted directly from the attention map AA of the final self-attention layer in the encoder E\mathcal{E} corresponding to the contextual token: St+δ=A[0,:]=[st+δ0,st+δ1,…,st+δNp−1],∑i=0Np−1st+δi=1\mathbf{S}_{t+\delta} = A[0, :] = \left[s^0_{t+\delta}, s^1_{t+\delta}, \dots, s^{N_p-1}_{t+\delta}\right], \quad \sum_{i=0}^{N_p-1} s^i_{t+\delta} = 1 where a higher score st+δis^i_{t+\delta} indicates that patch ii contains richer facial contextual information.

    To jointly prioritize micro-movements and suppress background noise, the Diagonal Micro-Attention features Pdma\mathbf{P}_{dma} are weighted by the combined matrix WW: W=diag(A^)×St+δW = \text{diag}(\hat{A}) \times \mathbf{S}_{t+\delta} Pdma=W×V(Pt/s)\mathbf{P}_{dma} = W \times V(\mathbf{P}_{t/s}) where diag(A^)\text{diag}(\hat{A}) is the diagonal correlation vector between It+δI_{t+\delta} and swapped frame It/sI_{t/s}, and V(⋅)V(\cdot) is the value projection.

  4. Knowl 4 — Blockwise Swapping Algorithm

    algorithm

    Blockwise Swapping is an unsupervised data augmentation procedure that creates a synthetic composite image patch set Pt/s\mathcal{P}_{t/s} by iteratively selecting 2D rectangular blocks of patches from frame It+δI_{t+\delta} and replacing the corresponding patches in frame ItI_t. The swapping continues until the total count of swapped patches reaches a target swapping ratio rs×Npr_s \times N_p.

    Input: Patch sequence Pt\mathcal{P}_t, Patch sequence Pt+δ\mathcal{P}_{t+\delta} (where Np=h×wN_p = h \times w), swapping ratio rsr_s (default: 0.5), minimum block size min_bsmin\_bs (default: 16), minimum aspect ratio min_armin\_ar (default: 0.3)
    Output: Swapped patch sequence Pt/s\mathcal{P}_{t/s}
    Pt/s←Pt\mathcal{P}_{t/s} \leftarrow \mathcal{P}_t
    Np←∣Pt∣N_p \leftarrow |\mathcal{P}_t|
    c←0c \leftarrow 0
    while c≤rs×Npc \leq r_s \times N_p do
        bs←rnd(min_bs,rs×Np−c)bs \leftarrow \text{rnd}(min\_bs, r_s \times N_p - c)
        ar←rnd(min_ar,1/min_ar)ar \leftarrow \text{rnd}(min\_ar, 1 / min\_ar)
        m←⌊bs⋅ar⌋m \leftarrow \lfloor \sqrt{bs \cdot ar} \rfloor
        n←⌊bs/ar⌋n \leftarrow \lfloor \sqrt{bs / ar} \rfloor
        p←rnd(0,h−m)p \leftarrow \text{rnd}(0, h - m)
        q←rnd(0,w−n)q \leftarrow \text{rnd}(0, w - n)
        for i∈[p,p+m)i \in [p, p + m) do
            for j∈[q,q+n)j \in [q, q + n) do
                k←i×w+jk \leftarrow i \times w + j
                Pt/s(k)←Pt+δ(k)\mathcal{P}_{t/s}(k) \leftarrow \mathcal{P}_{t+\delta}(k)
            end for
        end for
        c←c+m×nc \leftarrow c + m \times n
    end while
    return Pt/s\mathcal{P}_{t/s}
  5. Knowl 5 — Micron-BERT Self-Supervised Training Objective

    equation

    The self-supervised training objective of μ\mu-BERT balances image reconstruction of the original frame from the swapped frame and contextual agreement for facial region localization: L=γ×Lr+β×Lagg\mathcal{L} = \gamma \times \mathcal{L}_r + \beta \times \mathcal{L}_{agg} where γ\gamma and β\beta are weighting hyperparameters.

    The reconstruction loss Lr\mathcal{L}_r evaluates the pixel-level Mean Squared Error between the decoder's reconstructed output image yt′∈RH×W×C\mathbf{y}'_t \in \mathbb{R}^{H \times W \times C} and the target ground-truth original frame It∈RH×W×CI_t \in \mathbb{R}^{H \times W \times C}: Lr=MSE(yt′,It)\mathcal{L}_r = \text{MSE}(\mathbf{y}'_t, I_t)

    The contextual agreement loss Lagg\mathcal{L}_{agg} measures the Mean Squared Error between the contextual token features pt+δCT\mathbf{p}^{CT}_{t+\delta} of frame It+δI_{t+\delta} and pt+δ/cropCT\mathbf{p}^{CT}_{t+\delta/crop} of its cropped variant Crop(It+δ)\text{Crop}(I_{t+\delta}): Lagg=MSE(pt+δCT,pt+δ/cropCT)\mathcal{L}_{agg} = \text{MSE}\left(\mathbf{p}^{CT}_{t+\delta}, \mathbf{p}^{CT}_{t+\delta/crop}\right)

  6. Knowl 6 — Downstream Micro-Expression Recognition Setup and Metrics

    model/method

    For downstream Micro-Expression Recognition (MER), the pre-trained μ\mu-BERT encoder E\mathcal{E} and Diagonal Micro-Attention (DMA) module are used as the backbone. The input to the network is a pair of frames: the onset frame (representing ItI_t) and the apex frame (representing It+δI_{t+\delta}).

    The feature representation Pdma\mathbf{P}_{dma}, which encodes the micro-movements and texture changes occurring between the onset and apex frames, is fed into a classification head to predict the micro-expression class.

    Evaluation is conducted using Leave-One-Out Cross-Validation (LOOCV) under the MEGC 2019 challenge standard protocols using two metrics:

    1. Unweighted F1-score (UF1): UF1=1C∑i=0C−12×TPi2×TPi+FPi+FNi\text{UF1} = \frac{1}{C} \sum_{i=0}^{C-1} \frac{2 \times \text{TP}_i}{2 \times \text{TP}_i + \text{FP}_i + \text{FN}_i}
    2. Unweighted Average Recall (UAR): UAR=1C∑i=0C−1TPiNi\text{UAR} = \frac{1}{C} \sum_{i=0}^{C-1} \frac{\text{TP}_i}{N_i} where CC is the total number of micro-expression classes, NiN_i is the total number of samples in class ii, and TPi,FPi,FNi\text{TP}_i, \text{FP}_i, \text{FN}_i are true positives, false positives, and false negatives for class ii.
  7. Knowl 7 — Ablation Study of Self-Supervised Baselines and Micron-BERT Modules on CASME3

    data/table

    Evaluating pre-training methods and progressive component configurations of μ\mu-BERT on the 7-class CASME3 benchmark demonstrates the cumulative contributions of Blockwise Swapping, Diagonal Micro-Attention (DMA), and Patch of Interest (POI). Generic self-supervised vision baselines (MoCo V3, BEiT, MAE) pre-trained on CASME3 improve over ImageNet-pretrained ViT-S, but remain limited by single-image context learning. Introducing Blockwise Swapping alone (MB1) outperforms MAE by ~1.4% UF1 and ~2.1% UAR. Adding DMA (MB2) guides the model to micro-movements, adding ~2.1% UF1 and ~3.2% UAR. Incorporating POI (MB3) to filter background noise provides an additional ~5.3% UF1 and ~6.4% UAR boost.

    Method Pre-train DMA POI UF1 (%) UAR (%)
    ViT-S ImageNet 20.34 18.76
    MoCo V3 - R50 CASME3 19.12 17.36
    MoCo V3 - R101 CASME3 20.14 18.52
    MoCo V3 CASME3 22.13 19.34
    BEiT CASME3 23.54 19.89
    MAE CASME3 23.86 20.87
    -BERT (MB1) CASME3 25.27 22.96
    -BERT (MB2) CASME3 ✓ 27.35 26.18
    -BERT (MB3) CASME3 ✓ ✓ 32.64 32.54
  8. Knowl 8 — Benchmark Recognition Performance of Micron-BERT Across Datasets

    data/table

    Micron-BERT (μ\mu-BERT) consistently outperforms prior state-of-the-art micro-expression recognition methods across multiple standard benchmarks under Leave-One-Out Cross-Validation (LOOCV):

    • CASME3 (Table 1): On 3-class recognition, μ\mu-BERT achieves 56.04% UF1 and 61.25% UAR (vs. RCN-A at 39.28% / 38.93%). On 4-class, it achieves 47.18% UF1 and 49.13% UAR (vs. Baseline (+Depth) at 30.01% / 29.82%). On 7-class, it achieves 32.64% UF1 and 32.54% UAR (vs. Baseline (+Depth) at 17.73% / 18.29%).
    • CASME II (Table 2): On 5-class evaluation, μ\mu-BERT attains 85.53% UF1 and 83.48% UAR (outperforming TSCNN at 80.70% UF1 and SMA-STN at 82.59% UAR). On 3-class evaluation, μ\mu-BERT attains 90.34% UF1 and 89.14% UAR (outperforming MAE at 88.03% and OFF-ApexNet at 88.28% UAR).
    • SAMM (Table 3, 5-class): μ\mu-BERT attains 83.86% UF1 and 84.75% UAR (outperforming MAE at 80.40% UF1 and MiMaNet at 76.70% UAR).
    • SMIC (Table 4, 3-class): μ\mu-BERT achieves 85.50% UF1 and 83.84% UAR (outperforming MAE at 81.86% UF1 / 80.82% UAR).
    • Composite Dataset (MEGC2019, Table 5, 3-class): μ\mu-BERT achieves 89.03% UF1 and 88.42% UAR (outperforming MAE at 88.50% UF1 and MiMaNet at 88.30% UF1 / 87.60% UAR).
    Dataset # Classes Best Prior Method Prior Best (UF1 / UAR %) -BERT (UF1 / UAR %)
    CASME3 3 RCN-A 39.28 / 38.93 56.04 / 61.25
    CASME3 4 Baseline (+Depth) 30.01 / 29.82 47.18 / 49.13
    CASME3 7 Baseline (+Depth) 17.73 / 18.29 32.64 / 32.54
    CASME II 5 TSCNN / SMA-STN 80.70 / 82.59 85.53 / 83.48
    CASME II 3 MAE / OFF-ApexNet 88.03 / 88.28 90.34 / 89.14
    SAMM 5 MAE 80.40 / 88.98 83.86 / 84.75
    SMIC 3 MAE 81.86 / 80.82 85.50 / 83.84
    Composite (MEGC2019) 3 MAE / MiMaNet 88.50 / 87.60 89.03 / 88.42
  9. Knowl 9 — Implementation and Self-Supervised Pre-Training Setup

    experimental setup

    The self-supervised pre-training of μ\mu-BERT is performed using an unlabelled pool of 8 million video frames extracted from the raw footage of CASME3 (excluding test-set frames), without utilizing emotion labels, apex frames, onset frames, or offset annotations.

    Key architectural and training specifications include:

    • Image and Patch Dimensions: Input resolution H=W=224H = W = 224 with C=3C = 3 color channels. Patch size ps=8ps = 8, yielding Np=784N_p = 784 non-overlapping patches.
    • Transformer Architecture: Latent feature dimension d=512d = 512 for all projected vectors. Encoder depth Le=4L_e = 4 transformer blocks; decoder depth Ld=4L_d = 4 transformer blocks.
    • Swapping and Temporal Hyperparameters: Temporal frame gap δ\delta is sampled uniformly at random between a lower bound of 5 and an upper bound of 11 (δ∈[5,11]\delta \in [5, 11]). Swapping ratio rs=0.5r_s = 0.5 (50% of patches swapped between frames).
    • Optimization: Trained using PyTorch with the AdamW optimizer for 100 epochs. Learning rate starts at 10−410^{-4} and decays to 0 using a CosineLinear schedule. The batch size is 64 per GPU across 32 ×\times NVIDIA A100 GPUs (40GB each), completing pre-training in approximately 3 days.
  10. Knowl 10 — Sensitivity of Micro-Difference Representations to Non-Motion Facial Illumination Variations

    limitation

    While the Patch of Interest (POI) module effectively suppresses non-facial background regions and their associated lighting noise, it is limited when lighting and illumination shifts occur directly on static facial regions (such as the forehead) in the absence of actual muscle movement. In such scenarios, illumination-induced pixel disparities between the two frames can be mistakenly detected by the network as facial micro-movement features, potentially introducing noise into the downstream recognition pipeline.

Coverage note — None was omitted; all key contributions including framework architecture, Blockwise Swapping algorithm, Diagonal Micro-Attention, Patch of Interest module, loss functions, pre-training details, experimental results across five benchmarks, ablation studies, and stated limitations are fully covered.

References

  1. 1.Hangbo Bao, Li Dong, and Furu Wei. BEiT: BERT pre-training of image transformers. 2021. 2, 8
  2. 2.Xianye Ben, Yi Ren, Junping Zhang, Su-Jing Wang, Kidiyo Kpalma, Weixiao Meng, and Yong-Jin Liu. Video-based facial micro-expression analysis: A survey of datasets, features and algorithms. IEEE transactions on pattern analysis and machine intelligence, 2021. 1
  3. 3.Bin Chen, Kun-Hong Liu, Yong Xu, Qing-Qiang Wu, and Jun-Feng Yao. Block division convolutional network with implicit deep features augmentation for micro-expression recognition. IEEE Transactions on Multimedia, 2022. 8
  4. 4.Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9640–9649, 2021. 8
  5. 5.Ciprian Adrian Corneanu, Marc Oliu Simon, Jeffrey F. Cohn, and Sergio Escalera Guerrero. Survey on rgb, 3d, thermal, and multimodal approaches for facial expression recognition: History, trends, and affect-related applications. IEEE Transactions on Pattern Analysis and Machine Intelligence, 38(8):1548–1568, 2016. 1
  6. 6.Adrian K. Davison, Cliff Lansley, Nicholas Costen, Kevin Tan, and Moi Hoon Yap. Samm: A spontaneous micro-facial movement dataset. IEEE Transactions on Affective Computing, 9(1):116–129, 2018. 1, 2, 5
  7. 7.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 2, 8
  8. 8.Y.S. Gan, S.T. Liong, W.C. Yau, Y.C. Huang, and L.K. Tan. Off-apexnet on micro-expression recognition system. Signal Processing: Image Communication, 74:129–139, 2019. 6, 7
  9. 9.Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022. 2, 7, 8
  10. 10.Ankith Jain Rakesh Kumar and Bir Bhanu. Micro-expression classification based on landmark relations with graph attention convolutional network. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 1511–1520, June 2021. 7
  11. 11.Ling Lei, Tong Chen, Shigang Li, and Jianfeng Li. Micro-expression recognition based on facial graph representation learning and facial action unit fusion. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1571–1580, 2021. 1, 2
  12. 12.Ling Lei, Tong Chen, Shigang Li, and Jianfeng Li. Micro-expression recognition based on facial graph representation learning and facial action unit fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 1571–1580, June 2021. 7, 8
  13. 13.Ling Lei, Jianfeng Li, Tong Chen, and Shigang Li. A novel Graph-TCN with a graph structured representation for micro-expression recognition. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2237–2245, 2020. 7
  14. 14.Jingting Li, Zizhao Dong, Shaoyuan Lu, Su-Jing Wang, Wen-Jing Yan, Yinhuan Ma, Ye Liu, Changbing Huang, and Xiaolan Fu. Cas(me)<sup>3</sup>: A third generation facial spontaneous micro-expression database with depth information and high ecological validity. IEEE Transactions on Pattern Analysis and Machine Intelligence, pages 1–1, 2022. 2, 5, 6, 7, 8
  15. 15.Xiaobai Li, Tomas Pfister, Xiaohua Huang, Guoying Zhao, and Matti Pietik¨ainen. A spontaneous micro-expression database: Inducement, collection and baseline. In 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), pages 1–6, 2013. 1, 2, 5
  16. 16.Yante Li, Xiaohua Huang, and Guoying Zhao. Micro-expression action unit detection with spatial and channel attention. Neurocomputing, 436:221–231, 2021. 2
  17. 17.S.T. Liong, Y.S. Gan, J. See, H. Khor, and Y. Huang. Shallow triple stream three-dimensional cnn (ststnet) for micro-expression recognition. In 2019 14th IEEE International Conference on Automatic Face and Gesture Recognition (FG), pages 1–5. IEEE, 2019. 7
  18. 18.Sze-Teng Liong, Yee Siang Gan, John See, Huai-Qian Khor, and Yen-Chang Huang. Shallow triple stream three-dimensional cnn (ststnet) for micro-expression recognition. In 2019 14th IEEE international conference on automatic face & gesture recognition (FG 2019), pages 1–5. IEEE, 2019. 7
  19. 19.Jiateng Liu, Wenming Zheng, and Yuan Zong. SMA-STN: Segmented movement-attending spatiotemporal network for micro-expression recognition. arXiv preprint arXiv:2010.09342, 2020. 6, 7
  20. 20.Yuchi Liu, Heming Du, Liang Zheng, and Tom Gedeon. A neural micro-expression recognizer. In 2019 14th IEEE International Conference on Automatic Face and Gesture Recognition (FG), pages 1–4. IEEE, 2019. 8
  21. 21.Yanju Liu, Yange Li, Xinhai Yi, Zuojin Hu, Huiyu Zhang, and Yanzhong Liu. Lightweight vit model for micro-expression recognition enhanced by transfer learning. Frontiers in Neurorobotics, 16, 2022. 2
  22. 22.Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, et al. Swin transformer v2: Scaling up capacity and resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12009–12019, 2022. 2
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10012–10022, 2021. 2
  24. 24.Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:1608.03983, 2016. 6
  25. 25.Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 6
  26. 26.Xuan-Bac Nguyen, Duc Toan Bui, Chi Nhan Duong, Tien D Bui, and Khoa Luu. Clusformer: A transformer based clustering approach to unsupervised large-scale face and visual landmark recognition. 2
  27. 27.Xuan-Bac Nguyen, Guee Sang Lee, Soo Hyung Kim, and Hyung Jeong Yang. Self-supervised learning based on spatial awareness for medical image analysis. IEEE Access, 8:162973–162981, 2020. 2
  28. 28.Xuan Nie, Madhumita A Takalkar, Mengyang Duan, Haimin Zhang, and Min Xu. GEME: Dual-stream multi-task gender-based micro-expression recognition. Neurocomputing, 427:13–28, 2021. 7
  29. 29.Tae-Hyun Oh, Ronnachai Jaroensri, Changil Kim, Mohamed Elgharib, Fredo Durand, William T. Freeman, and Wojciech Matusik. Learning-based video motion magnification, 2018. 1, 2, 7
  30. 30.Kha Gia Quach, Ngan Le, Chi Nhan Duong, Ibsa Jalata, Kaushik Roy, and Khoa Luu. Non-volume preserving-based fusion to group-level emotion recognition on crowd videos. Pattern Recognition, 128:108646, 2022. 2
  31. 31.Ankith Jain Rakesh Kumar and Bir Bhanu. Micro-expression classification based on landmark relations with graph attention convolutional network. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pages 1511–1520, 2021. 2
  32. 32.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, pages 8821–8831. PMLR, 2021. 2
  33. 33.Tal Ridnik, Emanuel Ben-Baruch, Asaf Noy, and Lihi Zelnik-Manor. Imagenet-21k pretraining for the masses. arXiv preprint arXiv:2104.10972, 2021. 2
  34. 34.John See, Moi Hoon Yap, Jingting Li, Xiaopeng Hong, and Su-Jing Wang. Megc 2019 – the second facial micro-expressions grand challenge. In 2019 14th IEEE International Conference on Automatic Face Gesture Recognition (FG 2019), pages 1–5, 2019. 6
  35. 35.B. Song, K. Li, Y. Zong, J. Zhu, W. Zheng, J. Shi, and L. Zhao. Recognizing spontaneous micro-expression using a three-stream convolutional neural network. IEEE Access, 7:184537–184551, 2019. 6, 7
  36. 36.Bo Sun, Siming Cao, Dongliang Li, Jun He, and Lejun Yu. Dynamic micro-expression recognition using knowledge distillation. IEEE Transactions on Affective Computing, 2020. 7
  37. 37.Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In European conference on computer vision, pages 402–419. Springer, 2020. 7
  38. 38.Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Herve Jegou. Training data-efficient image transformers amp; distillation through attention. In International Conference on Machine Learning, volume 139, pages 10347–10357, July 2021. 2
  39. 39.Thuong-Khanh Tran, Quang-Nhat Vo, Xiaopeng Hong, Xiaobai Li, and Guoying Zhao. Micro-expression spotting: A new benchmark. Neurocomputing, 443:356–368, 2021. 2
  40. 40.Thanh-Dat Truong, Quoc-Huy Bui, Chi Nhan Duong, Han-Seok Seo, Son Lam Phung, Xin Li, and Khoa Luu. Direcformer: A directed attention in transformer approach to robust action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 20030–20040, June 2022. 2
  41. 41.Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham, Bhiksha Raj, Ngan Le, and Khoa Luu. The right to talk: An audio-visual transformer approach. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1105–1114, October 2021. 2
  42. 42.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 3
  43. 43.Su-Jing Wang, Ying He, Jingting Li, and Xiaolan Fu. Mesnet: A convolutional neural network for spotting multi-scale micro-expression intervals in long videos. IEEE Transactions on Image Processing, 30:3956–3969, 2021. 2
  44. 44.Yan Wang, Yikun Huang, Can Liu, Xiaoying Gu, Dandan Yang, Shuopeng Wang, and Bo Zhang. Micro expression recognition via dual-stream spatiotemporal attention network. Journal of Healthcare Engineering, 2021, 2021. 7
  45. 45.Yandan Wang, John See, Yee-Hui Oh, Raphael C.-W. Phan, Yogachandran Rahulamathavan, Huo-Chong Ling, Su-Wei Tan, and Xujie Li. Effective recognition of facial micro-expressions with video motion magnification. Multimedia Tools and Applications, 76(20):21665–21690, 2016. 2
  46. 46.Mengting Wei, Wenming Zheng, Yuan Zong, Xingxun Jiang, Cheng Lu, and Jiateng Liu. A novel micro-expression recognition approach using attention-based magnification-adaptive networks. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2420–2424. IEEE, 2022. 7
  47. 47.Bin Xia and Shangfei Wang. Micro-expression recognition enhanced by macro-expression from spatial-temporal domain. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, pages 1186–1193, 2021. 6, 7, 8
  48. 48.Bin Xia, Weikang Wang, Shangfei Wang, and Enhong Chen. Learning from macro-expression: a micro-expression recognition framework. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2936–2944, 2020. 7, 8
  49. 49.Zhaoqiang Xia, Wei Peng, Huai-Qian Khor, Xiaoyi Feng, and Guoying Zhao. Revealing the invisible with model and data shrinking for composite-database micro-expression recognition. IEEE Transactions on Image Processing, 29:8590–8605, 2020. 6, 7
  50. 50.Wen-Jing Yan, Xiaobai Li, Su-Jing Wang, Guoying Zhao, Yong-Jin Liu, Yu-Hsin Chen, and Xiaolan Fu. Casme ii: An improved spontaneous micro-expression database and the baseline evaluation. PLOS ONE, 9(1):1–8, 01 2014. 2, 5
  51. 51.Wen-Jing Yan, Qi Wu, Yong-Jin Liu, Su-Jing Wang, and Xiaolan Fu. Casme database: A dataset of spontaneous micro-expressions collected from neutralized faces. In 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG), pages 1–7, 2013. 1
  52. 52.Jianhui Yu, Chaoyi Zhang, Yang Song, and Weidong Cai. ICE-GAN: Identity-aware and capsule-enhanced gan for micro-expression recognition and synthesis. arXiv preprint arXiv:2005.04370, 2020. 8
  53. 53.Ling Zhou, Qirong Mao, Xiaohua Huang, Feifei Zhang, and Zhihong Zhang. Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition. arXiv preprint arXiv:2101.04838, 2021. 8
  54. 54.Ling Zhou, Qirong Mao, Xiaohua Huang, Feifei Zhang, and Zhihong Zhang. Feature refinement: An expression-specific feature learning and fusion method for micro-expression recognition. Pattern Recognition, 122:108275, 2022. 7
  55. 55.L. Zhou, Q. Mao, and L. Xue. Dual-inception network for cross-database micro-expression recognition. In 2019 14th IEEE International Conference on Automatic Face and Gesture Recognition (FG), pages 1–5. IEEE, 2019. 8

Citation

MLA
Nguyen, X.-B., et al. “Micron-BERT: BERT-based Facial Micro-Expression Recognition”. arXiv, 2023, http://arxiv.org/abs/2304.03195v1.
APA
Nguyen, X.-B., Duong, C. N., Li, X., Gauch, S., Seo, H.-S., & Luu, K. (2023). Micron-BERT: BERT-based Facial Micro-Expression Recognition. arXiv. http://arxiv.org/abs/2304.03195v1
Chicago
Nguyen, X.-B., C. N. Duong, X. Li, S. Gauch, H.-S. Seo, and K. Luu. 2023. “Micron-BERT: BERT-based Facial Micro-Expression Recognition”. arXiv. http://arxiv.org/abs/2304.03195v1.
Harvard
Nguyen, X.-B. et al. (2023) “Micron-BERT: BERT-based Facial Micro-Expression Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.03195v1.
Vancouver
1. Nguyen X-B, Duong CN, Li X, Gauch S, Seo H-S, Luu K (2023) Micron-BERT: BERT-based Facial Micro-Expression Recognition. arXiv

BibTeX

@article{nguyen2023micron,
  title = {Micron-BERT: BERT-based Facial Micro-Expression Recognition},
  author = {Nguyen, Xuan-Bac and Duong, Chi Nhan and Li, Xin and Gauch, Susan and Seo, Han-Seok and Luu, Khoa},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.03195v1},
  eprint = {2304.03195}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE