Built independently by an author, for readers. Read the story and support ChapterPal

topic

action unit

An action unit is a standardized, fundamental visual component used in computer vision and affective computing to describe and quantify discrete facial muscle movements. Derived from the Facial Action Coding System, each action unit corresponds to the observable activation of a specific muscle or muscle group, such as raising an eyebrow, tightening the eyelids, or pulling the corners of the lips. In automated visual analysis, computer vision algorithms detect the presence and intensity of individual action units across image and video sequences to objectively reconstruct complex facial expressions. By decomposing facial behavior into these granular, anatomically grounded elements, systems can reliably interpret emotions, pain, and cognitive states without relying solely on subjective holistic categories.

4 items

Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild

Reliable Crowdsourcing and Deep Locality-Preserving Learning for Expression Recognition in the Wild

Shan Li, Weihong Deng, Junping Du

OrganizationsBeijing University of Posts and Telecommunications

Why you should read this

Presents RAF-DB, a large-scale real-world facial expression database labeled through reliable crowdsourcing, alongside a deep locality-preserving CNN that significantly improves in-the-wild emotion recognition across basic and compound expressions.

Past research on facial expressions have used relatively limited datasets, which makes it unclear whether current methods can be employed in real world. In this paper, we present a novel database, RAF-DB, which contains about 30000 facial images from thousands of individuals. Each image has been individually labeled about 40 times, then EM algorithm was used to filter out unreliable labels. Crowdsourcing reveals that real-world faces often express compound emotions, or even mixture ones. For all we know, RAF-DB is the first database that contains compound expressions in the wild. Our cross-database study shows that the action units of basic emotions in RAF-DB are much more diverse than, or even deviate from, those of lab-controlled ones. To address this problem, we propose a new DLP-CNN (Deep Locality-Preserving CNN) method, which aims to enhance the discriminative power of deep features by preserving the locality closeness while maximizing the inter-class scatters. The benchmark experiments on the 7-class basic expressions and 11-class compound expressions, as well as the additional experiments on SFEW and CK+ databases, show that the proposed DLP-CNN outperforms the state-of-the-art handcrafted features and deep learning based methods for the expression recognition in the wild.

Added

2026-09-25

Recognizing Action Units for Facial Expression Analysis

Recognizing Action Units for Facial Expression Analysis

Ying-li Tian, T. Kanade, Jeffrey F. Cohn

Why you should read this

Develops an automated facial expression analysis system that accurately recognizes subtle and combined Facial Action Coding System action units by combining multi-state tracking of permanent features and transient wrinkles with neural network classification.

Most automatic expression analysis systems attempt to recognize a small set of prototypic expressions (e.g. happiness and anger). Such prototypic expressions, however, occur infrequently. Human emotions and intentions are communicated more often by changes in one or two discrete facial features. We develop an automatic system to analyze subtle changes in facial expressions based on both permanent facial features (brows, eyes, mouth) and transient facial features (deepening of facial furrows) in a nearly frontal image sequence. Unlike most existing systems, our system attempts to recognize fine-grained changes in facial expression based on Facial Action Coding System (FACS) action units (AUs), instead of six basic expressions (e.g. happiness and anger). Multi-state face and facial component models are proposed for tracking and modeling different facial features, including lips, eyes, brows, cheeks, and their related wrinkles and facial furrows. Then we convert the results of tracking to detailed parametric descriptions of the facial features. With these features as the inputs, 11 lower face action units (AUs) and 7 upper face AUs are recognized by a neural network algorithm. A recognition rate of 96.7% for lower face AUs and 95% for upper face AUs is obtained respectively. The recognition results indicate that our system can identify action units regardless of whether they occurred singly or in combinations.

Added

2026-09-18

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook

Xinlei Yu, Zhangquan Chen, Yongbo He, Tianyu Fu, Cheng Yang, Chengming Xu, Yue Ma, Xiaobin Hu, Zhe Cao, Jie Xu, Guibin Zhang, Jiale Tao, Jiayi Zhang, Siyuan Ma, Kaituo Feng, Haojie Huang, Youxing Li, Ronghao Chen, Huacan Wang, Chenglin Wu, Zikun Su, Xiaogang Xu, Kelu Yao, Kun Wang, Chen Gao, Yue Liao, Ruqi Huang, Tao Jin, Zhucun Xue, Cheng Tan, Jiangning Zhang, Wenqi Ren, Yanwei Fu, Yong Liu, Yu Wang, Xiangyu Yue, Yu-Gang Jiang, Shuicheng Yan

OrganizationsBeijing University of Posts and TelecommunicationsDeepWisdomFudan UniversityNanjing UniversityNanyang Technological UniversityNational University of SingaporeQuantaAlphaRenmin University of ChinaShanghai Artificial Intelligence LaboratoryShanghai Jiao Tong UniversitySun Yat-sen UniversityTencentThe Chinese University of Hong KongThe Hong Kong University of Science and TechnologyTsinghua UniversityUniversity of Chinese Academy of SciencesZhejiang LabZhejiang University

Why you should read this

Presents a unified and comprehensive landscape of latent space in language-based models, offering a foundational understanding of its evolution, mechanisms, and diverse abilities, crucial for anyone seeking to grasp the underpinnings of next-generation AI.

Latent space is rapidly emerging as a native substrate for language-based models. While modern systems are still commonly understood through explicit token-level generation, an increasing body of work shows that many critical internal processes are more naturally carried out in continuous latent space than in human-readable verbal traces. This shift is driven by the structural limitations of explicit-space computation, including linguistic redundancy, discretization bottlenecks, sequential inefficiency, and semantic loss. This survey aims to provide a unified and up-to-date landscape of latent space in language-based models. We organize the survey into five sequential perspectives: Foundation, Evolution, Mechanism, Ability, and Outlook. We begin by delineating the scope of latent space, distinguishing it from explicit or verbal space and from the latent spaces commonly studied in generative visual models. We then trace the field's evolution from early exploratory efforts to the current large-scale expansion. To organize the technical landscape, we examine existing work through the complementary lenses of mechanism and ability. From the perspective of Mechanism, we identify four major lines of development: Architecture, Representation, Computation, and Optimization. From the perspective of Ability, we show how latent space supports a broad capability spectrum spanning Reasoning, Planning, Modeling, Perception, Memory, Collaboration, and Embodiment. Beyond consolidation, we discuss the key open challenges, and outline promising directions for future research. We hope this survey serves not only as a reference for existing work, but also as a foundation for understanding latent space as a general computational and systems paradigm for next-generation intelligence.

Added

2026-04-20

License

Published with permission