keyword
Deep Bidirectional Transformers
Deep bidirectional transformers are neural network architectures based on multi-layer self-attention mechanisms that simultaneously analyze information from both preceding and following contexts across all hidden layers. Unlike traditional sequence models that process data strictly in a single direction or combine separate forward and backward passes at the output, deep bidirectional architectures allow each token or feature to jointly condition on its entire surrounding context throughout the full depth of the model. Typically pre-trained on large unlabeled datasets through self-supervised tasks, such as predicting masked elements within an input sequence, these models generate rich contextual representations that can be fine-tuned with minimal modifications for a wide variety of downstream applications in natural language processing, computer vision, and multimodal learning.
2 items

Micron-BERT: BERT-Based Facial Micro-Expression Recognition
Xuan-Bac Nguyen, Chi Nhan Duong, Xin Li, Susan Gauch, Han-Seok Seo, Khoa Luu
Why you should read this
Proposes a self-supervised BERT-based framework that captures subtle facial micro-expressions between video frames without manual annotations by pairing diagonal micro-attention with an unsupervised patch-of-interest detector.
Micro-expression recognition is one of the most challenging topics in affective computing. It aims to recognize tiny facial movements difficult for humans to perceive in a brief period, i.e., 0.25 to 0.5 seconds. Recent advances in pre-training deep Bidirectional Transformers (BERT) have significantly improved self-supervised learning tasks in computer vision. However, the standard BERT in vision problems is designed to learn only from full images or videos, and the architecture cannot accurately detect details of facial micro-expressions. This paper presents Micron-BERT (µ-BERT), a novel approach to facial micro-expression recognition. The proposed method can automatically capture these movements in an unsupervised manner based on two key ideas. First, we employ Diagonal Micro-Attention (DMA) to detect tiny differences between two frames. Second, we introduce a new Patch of Interest (PoI) module to localize and highlight micro-expression interest regions and simultaneously reduce noisy backgrounds and distractions. By incorporating these components into an end-to-end deep network, the proposed µ-BERT significantly outperforms all previous work in various micro-expression tasks. µ-BERT can be trained on a large-scale unlabeled dataset, i.e., up to 8 million images, and achieves high accuracy on new unseen facial micro-expression datasets. Empirical experiments show µ-BERT consistently outperforms state-of-the-art performance on four micro-expression benchmarks, including SAMM, CASME II, SMIC, and CASME3, by significant margins. Code will be available at https://github.com/uark-cviu/Micron-BERT
Added
2026-09-26

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova
Why you should read this
Introduces BERT, a conceptually simple yet empirically powerful model that revolutionized NLP by enabling state-of-the-art performance across diverse tasks through bidirectional representation pre-training, significantly reducing the need for task-specific architecture modifications.
We introduce a new language representation model called BERT, which stands for Bidirectional Encoder Representations from Transformers. Unlike recent language representation models, BERT is designed to pre-train deep bidirectional representations from unlabeled text by jointly conditioning on both left and right context in all layers. As a result, the pre-trained BERT model can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of tasks, such as question answering and language inference, without substantial task-specific architecture modifications. BERT is conceptually simple and empirically powerful. It obtains new state-of-the-art results on eleven natural language processing tasks, including pushing the GLUE score to 80.5% (7.7% point absolute improvement), MultiNLI accuracy to 86.7% (4.6% absolute improvement), SQuAD v1.1 question answering Test F1 to 93.2 (1.5 point absolute improvement) and SQuAD v2.0 Test F1 to 83.1 (5.1 point absolute improvement).
Added
2026-02-14

