keyword
face recognition
Face recognition is a computer vision and biometric technology that automatically identifies or verifies the identity of an individual from digital images or video frames of their face. The process typically involves detecting a human face within an image, extracting distinctive geometric or deep learned feature representations, and comparing these facial representations against stored profiles or known identity databases using similarity metrics. Widely used in security systems, user authentication, automated surveillance, and photo organization, face recognition systems operate across both closed-set scenarios, where all subjects are predefined, and open-set scenarios with previously unseen individuals. The core technical challenge in developing robust systems lies in reliably distinguishing facial identities despite variations in lighting conditions, camera viewpoints, head pose, facial expressions, aging, and occlusions.
26 items

Sibling-Attack: Rethinking Transferable Adversarial Attacks against Face Recognition
Zexin Li, Bangjie Yin, Taiping Yao, Junfeng Guo, Shouhong Ding, Simin Chen, Cong Liu
Why you should read this
Proposes a multi-task adversarial framework that leverages gradient information from face attribute recognition to substantially improve black-box attack transferability against commercial face recognition systems.
A hard challenge in developing practical face recognition (FR) attacks is due to the black-box nature of the target FR model, i.e., inaccessible gradient and parameter information to attackers. While recent research took an important step towards attacking black-box FR models through leveraging transferability, their performance is still limited, especially against online commercial FR systems that can be pessimistic (e.g., a less than 50% ASR–attack success rate on average). Motivated by this, we present Sibling-Attack, a new FR attack technique for the first time explores a novel multi-task perspective (i.e., leveraging extra information from multi-correlated tasks to boost attacking transferability). Intuitively, Sibling-Attack selects a set of tasks correlated with FR and picks the Attribute Recognition (AR) task as the task used in Sibling-Attack based on theoretical and quantitative analysis. Sibling-Attack then develops an optimization framework that fuses adversarial gradient information through (1) constraining the cross-task features to be under the same space, (2) a joint-task meta optimization framework that enhances the gradient compatibility among tasks, and (3) a cross-task gradient stabilization method which mitigates the oscillation effect during attacking. Extensive experiments demonstrate that Sibling-Attack outperforms state-of-the-art FR attack techniques by a non-trivial margin, boosting ASR by 12.61% and 55.77% on average on state-of-the-art pre-trained FR models and two well-known, widely used commercial FR systems.
Added
2026-09-26

A Study of Face Obfuscation in ImageNet
Kaiyu Yang, Jacqueline H. Yau, Li Fei-Fei, Jia Deng, Olga Russakovsky
Why you should read this
Demonstrates that obfuscating incidental human faces in ImageNet protects individual privacy with virtually no drop in classification accuracy or downstream feature transferability.
Face obfuscation (blurring, mosaicing, etc.) has been shown to be effective for privacy protection; nevertheless, object recognition research typically assumes access to complete, unobfuscated images. In this paper, we explore the effects of face obfuscation on the popular ImageNet challenge visual recognition benchmark. Most categories in the ImageNet challenge are not people categories; however, many incidental people appear in the images, and their privacy is a concern. We first annotate faces in the dataset. Then we demonstrate that face obfuscation has minimal impact on the accuracy of recognition models. Concretely, we benchmark multiple deep neural networks on obfuscated images and observe that the overall recognition accuracy drops only slightly (≤ 1.0%). Further, we experiment with transfer learning to 4 downstream tasks (object recognition, scene recognition, face attribute classification, and object detection) and show that features learned on obfuscated images are equally transferable. Our work demonstrates the feasibility of privacy-aware visual recognition, improves the highly-used ImageNet challenge benchmark, and suggests an important path for future visual datasets. Data and code are available at https://github.com/princetonvisualai/imagenet-face-obfuscation.
Added
2026-09-26

Face Recognition: The Problem of Compensating for Changes in Illumination Direction
Yael Adini, Y. Moses, S. Ullman
Why you should read this
Demonstrates through systematic empirical testing that standard illumination-invariant image representations such as edge maps, intensity derivatives, and 2D Gabor filters fail to overcome lighting direction changes in face recognition, establishing the need for richer three-dimensional or model-based recognition strategies.
A face recognition system must recognize a face from a novel image despite the variations between images of the same face. A common approach to overcoming image variations because of changes in the illumination conditions is to use image representations that are relatively insensitive to these variations. Examples of such representations are edge maps, image intensity derivatives, and images convolved with 2D Gabor-like filters. Here we present an empirical study that evaluates the sensitivity of these representations to changes in illumination, as well as viewpoint and facial expression. Our findings indicated that none of the representations considered is sufficient by itself to overcome image variations because of a change in the direction of illumination. Similar results were obtained for changes due to viewpoint and expression. Image representations that emphasized the horizontal features were found to be less sensitive to changes in the direction of illumination. However, systems based only on such representations failed to recognize up to 20 percent of the faces in our database. Humans performed considerably better under the same conditions. We discuss possible reasons for this superiority and alternative methods for overcoming illumination effects in recognition.
Added
2026-09-25

Region Covariance: A Fast Descriptor for Detection and Classification
Oncel Tuzel, Fatih Porikli, Peter Meer
Why you should read this
Proposes a fast, low-dimensional region covariance descriptor computed via integral images and matched using a Riemannian distance metric, enabling highly accurate and efficient object detection and texture classification that naturally resists variations in rotation and illumination.
We describe a new region descriptor and apply it to two problems, object detection and texture classification. The covariance of d-features, e.g., the three-dimensional color vector, the norm of first and second derivatives of intensity with respect to x and y, etc., characterizes a region of interest. We describe a fast method for computation of covariances based on integral images. The idea presented here is more general than the image sums or histograms, which were already published before, and with a series of integral images the covariances are obtained by a few arithmetic operations. Covariance matrices do not lie on Euclidean space, therefore we use a distance metric involving generalized eigenvalues which also follows from the Lie group structure of positive definite matrices. Feature matching is a simple nearest neighbor search under the distance metric and performed extremely rapidly using the integral images. The performance of the covariance features is superior to other methods, as it is shown, and large rotations and illumination changes are also absorbed by the covariance matrix.
Added
2026-09-25

A Trainable System for Object Detection
CONSTANTINE PAPAGEORGIOU, TOMASO POGGIO
Why you should read this
Presents a general framework for object detection in cluttered scenes by combining an overcomplete dictionary of multiscale Haar wavelets with support vector machines to achieve real-time detection across multiple visual domains.
This paper presents a general, trainable system for object detection in unconstrained, cluttered scenes. The system derives much of its power from a representation that describes an object class in terms of an overcomplete dictionary of local, oriented, multiscale intensity differences between adjacent regions, efficiently computable as a Haar wavelet transform. This example-based learning approach implicitly derives a model of an object class by training a support vector machine classifier using a large set of positive and negative examples. We present results on face, people, and car detection tasks using the same architecture. In addition, we quantify how the representation affects detection performance by considering several alternate representations including pixels and principal components. We also describe a real-time application of our person detection system as part of a driver assistance system.
Added
2026-09-25

Toward Open Set Recognition
W. Scheirer, A. Rocha, Archana Sapkota, T. Boult
Why you should read this
Formalizes the open set recognition problem and introduces the 1-vs-Set Machine to limit classification risk in unconstrained spaces where unseen classes emerge at test time.
To date, almost all experimental evaluations of machine learning-based recognition algorithms in computer vision have taken the form of "closed set" recognition, whereby all testing classes are known at training time. A more realistic scenario for vision applications is "open set" recognition, where incomplete knowledge of the world is present at training time, and unknown classes can be submitted to an algorithm during testing. This article explores the nature of open set recognition, and formalizes its definition as a constrained minimization problem. The open set recognition problem is not well addressed by existing algorithms because it requires strong generalization. As a step towards a solution, we introduce a novel "1-vs-Set Machine," which sculpts a decision space from the marginal distances of a 1-class or binary SVM with a linear kernel. This methodology applies to several different applications in computer vision where open set recognition is a challenging problem, including object recognition and face verification. We consider both in this work, with large scale experiments performed over data from the Caltech 256, ImageNet, and Labeled Faces in the Wild sets. The experiments highlight the effectiveness of machines adapted for open set evaluation compared to existing 1-class and binary SVMs for the same tasks.
Added
2026-09-24

PCANet: A Simple Deep Learning Baseline for Image Classification?
Tsung-Han Chan, Kui Jia, Shenghua Gao, Jiwen Lu, Zinan Zeng, Yi Ma
Why you should read this
Presents PCANet, an ultra-simple deep learning architecture built from cascaded principal component analysis, binary hashing, and block histograms that matches or exceeds complex deep neural networks across multiple standard image classification benchmarks.
In this work, we propose a very simple deep learning network for image classification which comprises only the very basic data processing components: cascaded principal component analysis (PCA), binary hashing, and block-wise histograms. In the proposed architecture, PCA is employed to learn multistage filter banks. It is followed by simple binary hashing and block histograms for indexing and pooling. This architecture is thus named as a PCA network (PCANet) and can be designed and learned extremely easily and efficiently. For comparison and better understanding, we also introduce and study two simple variations to the PCANet, namely the RandNet and LDANet. They share the same topology of PCANet but their cascaded filters are either selected randomly or learned from LDA. We have tested these basic networks extensively on many benchmark visual datasets for different tasks, such as LFW for face verification, MultiPIE, Extended Yale B, AR, FERET datasets for face recognition, as well as MNIST for hand-written digits recognition. Surprisingly, for all tasks, such a seemingly naive PCANet model is on par with the state of the art features, either prefixed, highly hand-crafted or carefully learned (by DNNs). Even more surprisingly, it sets new records for many classification tasks in Extended Yale B, AR, FERET datasets, and MNIST variations. Additional experiments on other public datasets also demonstrate the potential of the PCANet serving as a simple but highly competitive baseline for texture classification and object recognition.
Added
2026-09-24

Attribute and simile classifiers for face verification
Neeraj Kumar, A. Berg, P. Belhumeur, S. Nayar
Why you should read this
Proposes novel attribute and simile classifiers that capture high-level visual traits and reference similarities, dramatically cutting face verification error rates on unconstrained benchmarks without requiring image pair alignment.
We present two novel methods for face verification. Our first method – “attribute” classifiers – uses binary classifiers trained to recognize the presence or absence of describable aspects of visual appearance (e.g., gender, race, and age). Our second method – “simile” classifiers – removes the manual labeling required for attribute classification and instead learns the similarity of faces, or regions of faces, to specific reference people. Neither method requires costly, often brittle, alignment between image pairs; yet, both methods produce compact visual descriptions, and work on real-world images. Furthermore, both the attribute and simile classifiers improve on the current state-of-the-art for the LFW data set, reducing the error rates compared to the current best by 23.92% and 26.34%, respectively, and 31.68% when combined. For further testing across pose, illumination, and expression, we introduce a new data set – termed PubFig – of real-world images of public figures (celebrities and politicians) acquired from the internet. This data set is both larger (60,000 images) and deeper (300 images per individual) than existing data sets of its kind. Finally, we present an evaluation of human performance.
Added
2026-09-24

Active Appearance Models Revisited
Iain Matthews, Simon Baker
Why you should read this
Proposes a fast, analytical Active Appearance Model fitting algorithm based on inverse compositional image alignment that projects out appearance variations to substantially improve convergence speed and alignment accuracy over traditional heuristic methods.
Active Appearance Models (AAMs) and the closely related concepts of Morphable Models and Active Blobs are generative models of a certain visual phenomenon. Although linear in both shape and appearance, overall, AAMs are nonlinear parametric models in terms of the pixel intensities. Fitting an AAM to an image consists of minimizing the error between the input image and the closest model instance; i.e. solving a nonlinear optimization problem. We propose an efficient fitting algorithm for AAMs based on the inverse compositional image alignment algorithm. We show how the appearance variation can be “projected out” using this algorithm and how the algorithm can be extended to include a “shape normalizing” warp, typically a 2D similarity transformation. We evaluate our algorithm to determine which of its novel aspects improve AAM fitting performance.
Added
2026-09-18

Using Discriminant Eigenfeatures for Image Retrieval
D. Swets, J. Weng
Why you should read this
Proposes a Discriminant Karhunen-Loève projection framework that combines principal component analysis with linear discriminant analysis to eliminate non-informative variations like lighting and substantially improve content-based image retrieval across diverse object classes.
This paper describes the automatic selection of features from an image training set using the theories of multidimensional discriminant analysis and the associated optimal linear projection. We demonstrate the effectiveness of these Most Discriminating Features for view-based class retrieval from a large database of widely varying real-world objects presented as “well-framed” views, and compare it with that of the principal component analysis.
Added
2026-09-18

Robust Object Recognition with Cortex-Like Mechanisms
Thomas Serre, Lior Wolf, S. Bileschi, M. Riesenhuber, T. Poggio
Why you should read this
Presents a biologically motivated hierarchical vision architecture that alternates between template matching and max-pooling to achieve high-accuracy object and scene recognition from minimal training data.
We introduce a new general framework for the recognition of complex visual scenes, which is motivated by biology: We describe a hierarchical system that closely follows the organization of visual cortex and builds an increasingly complex and invariant feature representation by alternating between a template matching and a maximum pooling operation. We demonstrate the strength of the approach on a range of recognition tasks: From invariant single object recognition in clutter to multiclass categorization problems and complex scene understanding tasks that rely on the recognition of both shape-based as well as texture-based objects. Given the biological constraints that the system had to satisfy, the approach performs surprisingly well: It has the capability of learning from only a few training examples and competes with state-of-the-art systems. We also discuss the existence of a universal, redundant dictionary of features that could handle the recognition of most object categories. In addition to its relevance for computer vision, the success of this approach suggests a plausibility proof for a class of feedforward models of object recognition in cortex.
Added
2026-09-18

The CMU Pose, Illumination, and Expression Database
Terence Sim, Simon Baker, Maan Bsat
Why you should read this
Presents the CMU Pose, Illumination, and Expression database, providing an extensive benchmark of over 40,000 systematically varied facial images across 68 subjects to advance face recognition and 3D modeling research.
In the Fall of 2000 we collected a database of over 40,000 facial images of 68 people. Using the CMU 3D Room we imaged each person across 13 different poses, under 43 different illumination conditions, and with 4 different expressions. We call this the CMU Pose, Illumination, and Expression (PIE) database. We describe the imaging hardware, the collection procedure, the organization of the images, several possible uses, and how to obtain the database.
Added
2026-09-18

Face Recognition Based on Fitting a 3D Morphable Model
V. Blanz, T. Vetter
Why you should read this
Develops a 3D morphable model framework that reconstructs 3D facial shape and texture from single 2D images, enabling accurate face recognition across wide variations in pose and lighting.
This paper presents a method for face recognition across variations in pose, ranging from frontal to profile views, and across a wide range of illuminations, including cast shadows and specular reflections. To account for these variations, the algorithm simulates the process of image formation in 3D space, using computer graphics, and it estimates 3D shape and texture of faces from single images. The estimate is achieved by fitting a statistical, morphable model of 3D faces to images. The model is learned from a set of textured 3D scans of heads. We describe the construction of the morphable model, an algorithm to fit the model to images, and a framework for face identification. In this framework, faces are represented by model parameters for 3D shape and texture. We present results obtained with 4,488 images from the publicly available CMU-PIE database and 1,940 images from the FERET database.
Added
2026-09-16

Acquiring linear subspaces for face recognition under variable lighting
Kuang-chih Lee, J. Ho, D. Kriegman
Why you should read this
Shows that low-dimensional linear subspaces for face recognition under variable lighting can be constructed directly from five to nine real images taken under specific point-source directions, eliminating the need for 3D reconstruction or large training datasets.
Previous work has demonstrated that the image variation of many objects (human faces in particular) under variable lighting can be effectively modeled by low-dimensional linear spaces, even when there are multiple light sources and shadowing. Basis images spanning this space are usually obtained in one of three ways: A large set of images of the object under different lighting conditions is acquired, and principal component analysis (PCA) is used to estimate a subspace. Alternatively, synthetic images are rendered from a 3D model (perhaps reconstructed from images) under point sources and, again, PCA is used to estimate a subspace. Finally, images rendered from a 3D model under diffuse lighting based on spherical harmonics are directly used as basis images. In this paper, we show how to arrange physical lighting so that the acquired images of each object can be directly used as the basis vectors of a low-dimensional linear space and that this subspace is close to those acquired by the other methods. More specifically, there exist configurations of k point light source directions, with k typically ranging from 5 to 9, such that, by taking k images of an object under these single sources, the resulting subspace is an effective representation for recognition under a wide range of lighting conditions. Since the subspace is generated directly from real images, potentially complex and/or brittle intermediate steps such as 3D reconstruction can be completely avoided; nor is it necessary to acquire large numbers of training images or to physically construct complex diffuse (harmonic) light fields. We validate the use of subspaces constructed in this fashion within the context of face recognition.
Added
2026-09-14

Overview of the face recognition grand challenge
P. Phillips, P. Flynn, W. T. Scruggs, K. Bowyer, Jin Chang, Kevin Hoffman, Joe Marques, Jaesik Min, W. Worek
Why you should read this
Establishes the Face Recognition Grand Challenge benchmark with a 50,000-image dataset of 3D scans and high-resolution stills across six experimental protocols to drive an order-of-magnitude reduction in face recognition error rates.
Over the last couple of years, face recognition researchers have been developing new techniques. These developments are being fueled by advances in computer vision techniques, computer design, sensor design, and interest in fielding face recognition systems. Such advances hold the promise of reducing the error rate in face recognition systems by an order of magnitude over Face Recognition Vendor Test (FRVT) 2002 results. The Face Recognition Grand Challenge (FRGC) is designed to achieve this performance goal by presenting to researchers a six-experiment challenge problem along with data corpus of 50,000 images. The data consists of 3D scans and high resolution still imagery taken under controlled and uncontrolled conditions. This paper describes the challenge problem, data corpus, and presents baseline performance and preliminary results on natural statistics of facial imagery.
Added
2026-09-14

Face Recognition: Features Versus Templates
R. Brunelli, T. Poggio
Why you should read this
Demonstrates through systematic experimental comparison that gradient-based template matching significantly outperforms geometric feature vectors in automated frontal face recognition, achieving perfect classification on a 47-person benchmark.
Over the last twenty years several different techniques have been proposed for computer recognition of human faces. The purpose of this paper is to compare two simple but general strategies on a common database (frontal images of faces of 47 people, 26 males and 21 females, four images per person). We have developed and implemented two new algorithms, the first one based on the computation of a set of geometrical features, such as nose width and length, mouth position and chin shape, and the second one based on almost-grey-level template matching. The results obtained on the testing sets, about 90% correct recognition using geometrical features and perfect recognition using template matching, favour our implementation of the template matching approach.
Source
https://cse.msu.edu/~rossarun/BiometricsTextBook/Papers/Face/Brunelli_FeaturesVsTemplates_PAMI93.pdfAdded
2026-09-14

VGGFace2: A Dataset for Recognising Faces across Pose and Age
Qiong Cao, Li Shen, Weidi Xie, Omkar M. Parkhi, Andrew Zisserman
Why you should read this
Introduces VGGFace2, a large-scale dataset of 3.31 million images across 9,131 identities designed to train deep neural networks across wide variations in pose and age, achieving state-of-the-art accuracy on the IARPA Janus benchmarks.
In this paper, we introduce a new large-scale face dataset named VGGFace2. The dataset contains 3.31 million images of 9131 subjects, with an average of 362.6 images for each subject. Images are downloaded from Google Image Search and have large variations in pose, age, illumination, ethnicity and profession (e.g. actors, athletes, politicians). The dataset was collected with three goals in mind: (i) to have both a large number of identities and also a large number of images for each identity; (ii) to cover a large range of pose, age and ethnicity; and (iii) to minimize the label noise. We describe how the dataset was collected, in particular the automated and manual filtering stages to ensure a high accuracy for the images of each identity. To assess face recognition performance using the new dataset, we train ResNet-50 (with and without Squeeze-and-Excitation blocks) Convolutional Neural Networks on VGGFace2, on MS- Celeb-1M, and on their union, and show that training on VGGFace2 leads to improved recognition performance over pose and age. Finally, using the models trained on these datasets, we demonstrate state-of-the-art performance on all the IARPA Janus face recognition benchmarks, e.g. IJB-A, IJB-B and IJB-C, exceeding the previous state-of-the-art by a large margin. Datasets and models are publicly available.
Added
2026-09-12


SphereFace: Deep Hypersphere Embedding for Face Recognition
Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, Le Song
Why you should read this
Introduces the angular softmax loss to enforce adjustable angular margins on hypersphere embeddings, establishing a geometric formulation for learning highly discriminative features in open-set face recognition.
This paper addresses deep face recognition (FR) problem under open-set protocol, where ideal face features are expected to have smaller maximal intra-class distance than minimal inter-class distance under a suitably chosen metric space. However, few existing algorithms can effectively achieve this criterion. To this end, we propose the angular softmax (A-Softmax) loss that enables convolutional neural networks (CNNs) to learn angularly discriminative features. Geometrically, A-Softmax loss can be viewed as imposing discriminative constraints on a hypersphere manifold, which intrinsically matches the prior that faces also lie on a manifold. Moreover, the size of angular margin can be quantitatively adjusted by a parameter . We further derive specific to approximate the ideal feature criterion. Extensive analysis and experiments on Labeled Face in the Wild (LFW), Youtube Faces (YTF) and MegaFace Challenge show the superiority of A-Softmax loss in FR tasks. The code has also been made publicly available.
Added
2026-09-12

Face recognition using Laplacianfaces
Xiaofei He, Shuicheng Yan, Yuxiao Hu, P. Niyogi, HongJiang Zhang
Why you should read this
Proposes Laplacianfaces, a linear subspace learning method based on Locality Preserving Projections that preserves local neighborhood structure to outperform standard Eigenface and Fisherface techniques under varying poses, lighting, and facial expressions.
We propose an appearance-based face recognition method called the Laplacianface approach. By using Locality Preserving Projections (LPP), the face images are mapped into a face subspace for analysis. Different from Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA) which effectively see only the Euclidean structure of face space, LPP finds an embedding that preserves local information, and obtains a face subspace that best detects the essential face manifold structure. The Laplacianfaces are the optimal linear approximations to the eigenfunctions of the Laplace Beltrami operator on the face manifold. In this way, the unwanted variations resulting from changes in lighting, facial expression, and pose may be eliminated or reduced. Theoretical analysis shows that PCA, LDA, and LPP can be obtained from different graph models. We compare the proposed Laplacianface approach with Eigenface and Fisherface methods on three different face data sets. Experimental results suggest that the proposed Laplacianface approach provides a better representation and achieves lower error rates in face recognition.
Added
2026-09-11

Detecting Faces in Images: A Survey
Ming-Hsuan Yang, D. Kriegman, N. Ahuja
Why you should read this
Categorizes over 150 face detection approaches into a structured four-part taxonomy while critically evaluating benchmark datasets, evaluation criteria, and performance trade-offs across varied imaging conditions.
Images containing faces are essential to intelligent vision-based human computer interaction, and research efforts in face processing include face recognition, face tracking, pose estimation, and expression recognition. However, many reported methods assume that the faces in an image or an image sequence have been identified and localized. To build fully automated systems that analyze the information contained in face images, robust and efficient face detection algorithms are required. Given a single image, the goal of face detection is to identify all image regions which contain a face regardless of its three-dimensional position, orientation, and the lighting conditions. Such a problem is challenging because faces are nonrigid and have a high degree of variability in size, shape, color, and texture. Numerous techniques have been developed to detect faces in a single image, and the purpose of this paper is to categorize and evaluate these algorithms. We also discuss relevant issues such as data collection, evaluation metrics, and benchmarking. After analyzing these algorithms and identifying their limitations, we conclude with several promising directions for future research.
Added
2026-09-11
