Built independently by an author, for readers. Read the story and support ChapterPal

keyword

facial graph representation learning

Facial graph representation learning is a computational approach in computer vision and machine learning that models a human face as a network of interconnected points to learn compact, meaningful feature representations of facial structure, geometry, and dynamics. In this framework, facial components—such as key landmarks, specific regions of interest, or action units—are formulated as nodes, while the spatial, semantic, or temporal relationships between them are encoded as edges. By applying graph-based neural architectures to process these networks, the approach effectively captures both localized muscle movements and global topological dependencies across the entire face. This structural formulation offers enhanced robustness against common real-world visual variations, such as changes in head pose, partial occlusions, and uneven illumination, making it widely useful for facial expression analysis, micro-expression recognition, face parsing, and affective computing.

1 item

Micron-BERT: BERT-Based Facial Micro-Expression Recognition

Micron-BERT: BERT-Based Facial Micro-Expression Recognition

Xuan-Bac Nguyen, Chi Nhan Duong, Xin Li, Susan Gauch, Han-Seok Seo, Khoa Luu

OrganizationsConcordia UniversityUniversity of ArkansasWest Virginia University

Why you should read this

Proposes a self-supervised BERT-based framework that captures subtle facial micro-expressions between video frames without manual annotations by pairing diagonal micro-attention with an unsupervised patch-of-interest detector.

Micro-expression recognition is one of the most challenging topics in affective computing. It aims to recognize tiny facial movements difficult for humans to perceive in a brief period, i.e., 0.25 to 0.5 seconds. Recent advances in pre-training deep Bidirectional Transformers (BERT) have significantly improved self-supervised learning tasks in computer vision. However, the standard BERT in vision problems is designed to learn only from full images or videos, and the architecture cannot accurately detect details of facial micro-expressions. This paper presents Micron-BERT (µ-BERT), a novel approach to facial micro-expression recognition. The proposed method can automatically capture these movements in an unsupervised manner based on two key ideas. First, we employ Diagonal Micro-Attention (DMA) to detect tiny differences between two frames. Second, we introduce a new Patch of Interest (PoI) module to localize and highlight micro-expression interest regions and simultaneously reduce noisy backgrounds and distractions. By incorporating these components into an end-to-end deep network, the proposed µ-BERT significantly outperforms all previous work in various micro-expression tasks. µ-BERT can be trained on a large-scale unlabeled dataset, i.e., up to 8 million images, and achieves high accuracy on new unseen facial micro-expression datasets. Empirical experiments show µ-BERT consistently outperforms state-of-the-art performance on four micro-expression benchmarks, including SAMM, CASME II, SMIC, and CASME3, by significant margins. Code will be available at https://github.com/uark-cviu/Micron-BERT

Added

2026-09-26