PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition

Chien-Yi WangYu-Ding LuShang-Ta YangShang-Hong Lai

article2022CVPR146 citations

Proposes PatchNet, a face anti-spoofing framework that reformulates presentation attack detection as fine-grained patch recognition across capture devices and materials, using asymmetric margin and self-supervised losses to achieve superior generalization on unseen spoof types without auxiliary pixel-wise supervision.

Listen

Facial recognition systems are widely deployed for biometric security, but they remain vulnerable to physical presentation attacks, such as printed photos or digital video replays. Existing face anti-spoofing techniques often suffer from poor generalization when confronted with unseen spoof mediums or varying camera sensors. Current models typically rely on whole-face binary classification or complex auxiliary supervision—like synthetic depth or reflection maps—which are computationally expensive and prone to overfitting to dataset collection biases.

The article introduces PatchNet, a streamlined framework designed to evaluate and demonstrate whether reformulating face anti-spoofing into a fine-grained local patch recognition task improves detection accuracy and model generalization. Rather than evaluating an entire resized face, PatchNet trains a model to classify specific combinations of camera capture devices and spoofing materials directly from undistorted local image patches.

To establish this framework, the authors evaluated image inputs using fixed-size patches cropped directly from raw, uncompressed frames across five public benchmark datasets (OULU-NPU, SiW, CASIA-FASD, Replay-Attack, and MSU-MFSD). The model employs a ResNet-18 feature encoder trained with two specialized loss functions: an asymmetric angular margin loss that enforces tight clustering for live samples while accommodating broad spoof diversity, and a self-supervised similarity loss to ensure feature consistency across different patch views of the same capture. Performance was assessed through standard intra-dataset, cross-dataset, and domain generalization protocols.

The experimental findings show significant performance advantages. First, moving from a standard binary classification baseline to fine-grained patch cropping reduced the average classification error rate on the OULU-NPU benchmark from 6.25% to 1.88%, ultimately reaching a 0.0% error rate with full loss regularization. Second, PatchNet matched or outperformed state-of-the-art methods across standard intra-dataset protocols, achieving a 0.0% error rate across multiple sub-protocols on OULU-NPU and SiW. Third, in domain generalization benchmarks where models were trained on three datasets and tested on an unseen fourth, PatchNet achieved top-tier competitive results, including an area under the curve of up to 98.46%. Finally, the resulting embedding space enabled practical few-shot reference adaptation, boosting performance on specific low-quality sensors from 88.49% to 90.7% with only ten live sample references.

These results indicate that spoofing cues are inherently local and material-specific rather than global facial features. By eliminating the need for pseudo-ground-truth auxiliary maps or complex adversarial domain adaptation, PatchNet substantially lowers implementation complexity and training overhead. This approach reduces security risks in biometric deployments by offering robust cross-environment reliability without demanding high-capacity neural network architectures.

Organizations deploying facial biometric systems should consider patch-level material classification strategies and test security performance on individual camera sensors rather than aggregate dataset averages. Where feasible, teams can implement few-shot live reference enrollment to quickly adapt deployed security systems to new camera hardware. For future development, additional work is recommended to test the framework on broader variations of unconstrained capture environments and explore cross-domain material perception datasets.

The authors note limitations regarding extreme sensor noise and heavy image compression, which degraded feature discrimination on certain low-quality devices. Furthermore, patch cropping requires sufficient spatial resolution, as very small patches (e.g., 64 pixels) degrade accuracy. Nevertheless, the extensive multi-benchmark validation supports high confidence in PatchNet’s efficacy as an efficient and generalizable face anti-spoofing framework.

arXiv: 2203.14325
Cover for PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition

Abstract

Face anti-spoofing (FAS) plays a critical role in securing face recognition systems from different presentation attacks. Previous works leverage auxiliary pixel-level supervision and domain generalization approaches to address unseen spoof types. However, the local characteristics of image captures, i.e., capturing devices and presenting materials, are ignored in existing works and we argue that such information is required for networks to discriminate between live and spoof images. In this work, we propose PatchNet which reformulates face anti-spoofing as a fine-grained patch-type recognition problem. To be specific, our framework recognizes the combination of capturing devices and presenting materials based on the patches cropped from non-distorted face images. This reformulation can largely improve the data variation and enforce the network to learn discriminative feature from local capture patterns. In addition, to further improve the generalization ability of the spoof feature, we propose the novel Asymmetric Margin-based Classification Loss and Self-supervised Similarity Loss to regularize the patch embedding space. Our experimental results verify our assumption and show that the model is capable of recognizing unseen spoof types robustly by only looking at local regions. Moreover, the fine-grained and patch-level reformulation of FAS outperforms the existing approaches on intra-dataset, cross-dataset, and domain generalization benchmarks. Furthermore, our PatchNet framework can enable practical applications like Few-Shot Reference-based FAS and facilitate future exploration of spoof-related intrinsic cues.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Proposed Method
  • 3.1. Overview
  • 3.2. Patch Features Extraction
  • 3.3. Fine-Grained Patch Recognition
  • 3.3.1. Preliminaries
  • 3.3.2. Asymmetric AM-Softmax Loss
  • 3.4. Self-Supervised Similarity Loss
  • 3.5. Training and Testing
  • 3.5.1. Total Loss
  • 3.5.2. Testing Strategy
  • 4. Experiments
  • 4.1. Datasets and Protocols
  • 4.2. Implementation Details
  • 4.3. Intra-Dataset Testing
  • 4.3.1. Results on Oulu-NPU
  • 4.3.2. Results on SiW
  • 4.4. Ablation Study
  • 4.5. Cross-Dataset Testing
  • 4.5.1. Experiments between C and I
  • 4.5.2. Domain Generalization Experiments
  • 4.6. Visualizations
  • 4.7. Applications
  • 5. Conclusions and Future Work
  • References

Knowls

  1. Knowl 1 — Fine-Grained Patch-Type Formulation for Face Anti-Spoofing

    model/method

    Conventional face anti-spoofing (FAS) models frame presentation attack detection as binary classification (live versus spoof) and resize cropped face images to fixed input dimensions. Resizing face crops distorts fine-grained sensor characteristics, screen moiré patterns, and material micro-textures. PatchNet reformulates FAS as a multi-class fine-grained patch-type recognition task operating directly on un-resized local face patches.

    Classes are defined as the Cartesian product of capture devices (sensors and resolutions) and presentation mediums (live human skin, print paper types, screen display types). Given raw face bounding boxes detected from video frames, un-distorted square patches of fixed pixel size (such as 160×160160 \times 160) are cropped directly without spatial scaling. An encoder network EθE_\theta (such as ResNet-18) extracts representations for each patch, followed by an L2L_2-normalization layer that projects features onto the unit hypersphere:

    f=Eθ(x)∥Eθ(x)∥2f = \frac{E_\theta(x)}{\|E_\theta(x)\|_2}

    where xx is an input patch and f∈Rdf \in \mathbb{R}^d is the normalized patch embedding satisfying ∥f∥2=1\|f\|_2 = 1.

  2. Knowl 2 — Asymmetric Angular Margin Softmax Loss

    equation

    In face anti-spoofing patch classification, bona fide (live) facial skin exhibits compact intra-class distributions, whereas spoof presentation attacks encompass diverse materials and displays that span a wider feature distribution. To account for this asymmetry, the Asymmetric Additive Margin Softmax (AAMS) loss applies different angular margin penalties to live and spoof classes on a normalized hypersphere.

    Let NN denote the total number of fine-grained classes, partitioned into a set of kk live classes L={L1,…,Lk}L = \{L_1, \dots, L_k\} and N−kN-k spoof classes S={S1,…,SN−k}S = \{S_1, \dots, S_{N-k}\}. For an L2L_2-normalized feature vector fi∈Rdf_i \in \mathbb{R}^d (∥fi∥2=1\|f_i\|_2 = 1) with ground-truth class label yi∈{1,…,N}y_i \in \{1, \dots, N\}, and normalized linear classification weight vectors Wj∈RdW_j \in \mathbb{R}^d (∥Wj∥2=1\|W_j\|_2 = 1) for j∈{1,…,N}j \in \{1, \dots, N\}, the sample loss is defined as:

    LAAMS(fi)={−log⁡exp⁡(s⋅(WyiTfi−ml))exp⁡(s⋅(WyiTfi−ml))+∑j=1,j≠yiNexp⁡(s⋅WjTfi),yi∈L−log⁡exp⁡(s⋅(WyiTfi−ms))exp⁡(s⋅(WyiTfi−ms))+∑j=1,j≠yiNexp⁡(s⋅WjTfi),yi∈S\mathcal{L}_{AAMS}(f_i) = \begin{cases} -\log \frac{\exp\left(s \cdot (W_{y_i}^T f_i - m_l)\right)}{\exp\left(s \cdot (W_{y_i}^T f_i - m_l)\right) + \sum_{j=1, j \neq y_i}^{N} \exp\left(s \cdot W_j^T f_i\right)}, & y_i \in L \\ -\log \frac{\exp\left(s \cdot (W_{y_i}^T f_i - m_s)\right)}{\exp\left(s \cdot (W_{y_i}^T f_i - m_s)\right) + \sum_{j=1, j \neq y_i}^{N} \exp\left(s \cdot W_j^T f_i\right)}, & y_i \in S \end{cases}

    where s>0s > 0 is a fixed scaling hyperparameter (set to s=30.0s = 30.0), ml≥0m_l \ge 0 is the angular margin penalty for live classes (set to ml=0.4m_l = 0.4), and ms≥0m_s \ge 0 is the angular margin penalty for spoof classes (set to ms=0.1m_s = 0.1). Imposing ml>msm_l > m_s forces the encoder to form tighter hyperspherical clusters for live captures while maintaining broad class boundaries for presentation attacks.

    For a mini-batch of nn face images, each generating two stochastic patch views xit1x_i^{t_1} and xit2x_i^{t_2} with normalized embeddings fit1f_i^{t_1} and fit2f_i^{t_2}, the batch asymmetric classification loss is:

    LAsym=1n∑i=1n(LAAMS(fit1)+LAAMS(fit2))\mathcal{L}_{Asym} = \frac{1}{n}\sum_{i=1}^n \left( \mathcal{L}_{AAMS}(f_i^{t_1}) + \mathcal{L}_{AAMS}(f_i^{t_2}) \right)

  3. Knowl 3 — Self-Supervised Spatial Similarity Loss for Patch Representations

    equation

    Because presentation medium characteristics (such as paper texture, display grid artifacts, or natural skin reflectance) are distributed across the entire face region, local patch representations extracted from different positions of the same face image should exhibit spatial location and rotation invariance.

    For an un-distorted face crop xix_i, two augmented patch views xit1=t1(xi)x_i^{t_1} = t_1(x_i) and xit2=t2(xi)x_i^{t_2} = t_2(x_i) are sampled via stochastic transformations t1,t2∼Tt_1, t_2 \sim \mathcal{T} comprising random horizontal flipping, random rotation, and fixed-size cropping. Passing these patches through an encoder and L2L_2-normalization yields feature vectors fit1,fit2∈Rdf_i^{t_1}, f_i^{t_2} \in \mathbb{R}^d where ∥fit1∥2=∥fit2∥2=1\|f_i^{t_1}\|_2 = \|f_i^{t_2}\|_2 = 1. The self-supervised similarity loss enforces feature consistency between paired views:

    LSim(fit1,fit2)=1n∑i=1n∥fit1−fit2∥2\mathcal{L}_{Sim}(f_i^{t_1}, f_i^{t_2}) = \frac{1}{n} \sum_{i=1}^n \|f_i^{t_1} - f_i^{t_2}\|_2

    where nn is the batch size and ∥⋅∥2\|\cdot\|_2 denotes the Euclidean distance. The complete multi-task training objective combines the classification and self-supervised consistency terms:

    L=α1LAsym+α2LSim\mathcal{L} = \alpha_1 \mathcal{L}_{Asym} + \alpha_2 \mathcal{L}_{Sim}

    where α1=1.0\alpha_1 = 1.0 and α2=1.0\alpha_2 = 1.0 in all implementations.

  4. Knowl 4 — Patch-Level Multi-Crop Inference and Score Aggregation

    algorithm

    During inference on a full face image, PatchNet extracts a regular grid of un-distorted overlapping patches, computes fine-grained class posterior probabilities on each patch, and aggregates the probabilities corresponding to all bona fide (live) classes across all patches to produce a final liveness score.

    Input: Face crop image XX of size W×HW \times H, patch size K=160K = 160, number of anchor points per axis G=3G = 3, feature encoder EθE_\theta, normalized class weight matrix W∈RN×dW \in \mathbb{R}^{N \times d}, live class index set L⊂{1,…,N}L \subset \{1, \dots, N\}, scale factor s=30.0s = 30.0
    Output: Overall liveness probability score LiveProb∈[0,1]\text{LiveProb} \in [0, 1]
    Set P=G×GP = G \times G
    Compute grid coordinate intervals:
        x-anchors=linspace(K/2,W−K/2,G)x\text{-anchors} = \text{linspace}(K / 2, W - K / 2, G)
        y-anchors=linspace(K/2,H−K/2,G)y\text{-anchors} = \text{linspace}(K / 2, H - K / 2, G)
    Initialize score accumulator Stotal=0S_{\text{total}} = 0
    for each (cx,cy)(c_x, c_y) in x-anchors×y-anchorsx\text{-anchors} \times y\text{-anchors} do
        Crop patch xpx_p of size K×KK \times K centered at (cx,cy)(c_x, c_y) from XX without resizing
        Compute normalized feature fp=Eθ(xp)∥Eθ(xp)∥2f_p = \frac{E_\theta(x_p)}{\|E_\theta(x_p)\|_2}
        Compute scaled class logits zj=s⋅(WjTfp)z_j = s \cdot (W_j^T f_p) for all j∈{1,…,N}j \in \{1, \dots, N\}
        Compute patch-level live probability:
            plive=∑y∈Lexp⁡(zy)∑j=1Nexp⁡(zj)p_{\text{live}} = \sum_{y \in L} \frac{\exp(z_y)}{\sum_{j=1}^N \exp(z_j)}
        Stotal=Stotal+pliveS_{\text{total}} = S_{\text{total}} + p_{\text{live}}
    end for
    return LiveProb=StotalP\text{LiveProb} = \frac{S_{\text{total}}}{P}
  5. Knowl 5 — Intra-Dataset FAS Performance on OULU-NPU and SiW

    data/table

    Intra-dataset evaluations test models against specific variations, such as unseen environments, attack mediums, and capture devices. OULU-NPU evaluates 4 protocols (Prot. 1: illumination/location shifts; Prot. 2: unseen attack mediums; Prot. 3: unseen camera sensors; Prot. 4: all combined shifts). SiW evaluates across 3 protocols (Prot. 1: variations in pose and expression; Prot. 2: varying spoof mediums; Prot. 3: unseen spoof mediums with sub-protocols 3-1 and 3-2). Performance is reported using Attack Presentation Classification Error Rate (APCER, %), Bona Fide Presentation Classification Error Rate (BPCER, %), and Average Classification Error Rate (ACER, %).

    Dataset Protocol Method APCER (%) BPCER (%) ACER (%)
    OULU-NPU Prot. 1 CDCN 0.4 1.7 1.0
    NAS-FAS 0.4 0.0 0.2
    PatchNet 0.0 0.0 0.0
    OULU-NPU Prot. 2 CDCN 1.5 1.4 1.5
    NAS-FAS 1.5 0.8 1.2
    PatchNet 1.1 1.2 1.2
    OULU-NPU Prot. 3 CDCN 2.4±1.32.4 \pm 1.3 2.2±2.02.2 \pm 2.0 2.3±1.42.3 \pm 1.4
    NAS-FAS 2.1±1.32.1 \pm 1.3 1.4±1.11.4 \pm 1.1 1.7±0.61.7 \pm 0.6
    PatchNet 1.8±1.471.8 \pm 1.47 0.56±1.240.56 \pm 1.24 1.18±1.261.18 \pm 1.26
    OULU-NPU Prot. 4 CDCN 4.6±4.64.6 \pm 4.6 9.2±8.09.2 \pm 8.0 6.9±2.96.9 \pm 2.9
    NAS-FAS 4.2±5.34.2 \pm 5.3 1.7±2.61.7 \pm 2.6 2.9±2.82.9 \pm 2.8
    PatchNet 2.5±3.812.5 \pm 3.81 3.33±3.733.33 \pm 3.73 2.9±3.02.9 \pm 3.0
    SiW Prot. 1 NAS-FAS 0.07 0.17 0.12
    PatchNet 0.00 0.00 0.00
    SiW Prot. 2 NAS-FAS 0.00±0.000.00 \pm 0.00 0.09±0.100.09 \pm 0.10 0.04±0.050.04 \pm 0.05
    PatchNet 0.00±0.000.00 \pm 0.00 0.00±0.000.00 \pm 0.00 0.00±0.000.00 \pm 0.00
    SiW Prot. 3 NAS-FAS 1.58±0.231.58 \pm 0.23 1.46±0.081.46 \pm 0.08 1.52±0.131.52 \pm 0.13
    PatchNet 3.06±1.13.06 \pm 1.1 1.83±0.831.83 \pm 0.83 2.45±0.452.45 \pm 0.45

    PatchNet attains 0.0% ACER on OULU-NPU Protocol 1 and SiW Protocols 1 and 2, matching or outperforming methods that rely on pixel-wise auxiliary supervision (depth/reflection maps) and neural architecture search.

  6. Knowl 6 — Cross-Dataset and Domain Generalization FAS Performance

    data/table

    To evaluate generalization across distinct imaging domains without using domain-adversarial or meta-learning training objectives, PatchNet is evaluated on cross-dataset transfers between CASIA-MFSD (C) and Replay-Attack (I), as well as on the standard four-dataset domain generalization leave-one-out benchmark across OULU-NPU (O), CASIA-MFSD (C), Replay-Attack (I), and MSU-MFSD (M). In the leave-one-out benchmark, fine-grained classes from the three training datasets are combined into 18, 17, 22, and 21 classes, respectively. Performance is reported in Half Total Error Rate (HTER, %) and Area Under the ROC Curve (AUC, %).

    OCI →\to M OMI →\to C OCM →\to I ICM →\to O
    Method HTER(%) AUC(%) HTER(%) AUC(%) HTER(%) AUC(%) HTER(%) AUC(%)
    MADDG 17.69 88.06 24.50 84.51 22.19 84.99 27.89 80.02
    RFM 13.89 93.98 20.27 88.16 17.30 90.48 16.45 91.16
    SSDG-R 7.38 97.17 10.44 95.94 11.71 96.59 15.61 91.54
    DRDG 15.56 91.79 12.43 95.81 19.05 88.79 15.63 91.75
    PatchNet 7.10 98.46 11.33 94.58 13.40 95.67 11.82 95.07

    In the two-dataset cross-evaluation between CASIA-MFSD (C) and Replay-Attack (I), PatchNet achieves 9.9% HTER on C →\to I (compared to 6.0% for DC-CDN and 15.5% for CDCN) and 26.2% HTER on I →\to C (compared to 30.1% for DC-CDN and 32.6% for CDCN).

  7. Knowl 7 — Ablation Analysis of Class Granularity, Patch Sizing, and Loss Formulations

    data/table

    Ablation experiments conducted on OULU-NPU Protocol 1 (measuring ACER, %) and multi-dataset domain generalization benchmarks evaluate the contribution of fine-grained class formulation, un-distorted patch cropping, asymmetric angular margins (LAsym\mathcal{L}_{Asym}), and self-supervised consistency (LSim\mathcal{L}_{Sim}).

    Class Definition Input Processing LAsym\mathcal{L}_{Asym} LSim\mathcal{L}_{Sim} OULU-NPU Prot. 1 ACER (%)
    Binary Resize (256×256256 \times 256) – – 6.25
    Fine-Grained Resize (256×256256 \times 256) – – 3.54
    Binary PatchCrop (160×160160 \times 160) – – 5.63
    Fine-Grained PatchCrop (160×160160 \times 160) – – 1.88
    Fine-Grained PatchCrop (160×160160 \times 160) ✓\checkmark (ml=ms=0m_l=m_s=0) – 1.46
    Fine-Grained PatchCrop (160×160160 \times 160) ✓\checkmark (ml=0.4,ms=0.1m_l=0.4, m_s=0.1) – 0.63
    Fine-Grained PatchCrop (160×160160 \times 160) ✓\checkmark (ml=0.4,ms=0.1m_l=0.4, m_s=0.1) ✓\checkmark 0.00

    Varying angular margins (ml,ms)(m_l, m_s) on OULU-NPU Protocol 1 demonstrates that symmetric margins (0.0,0.0)→1.46%(0.0, 0.0) \to 1.46\%, (0.2,0.2)→0.83%(0.2, 0.2) \to 0.83\%, and (0.4,0.4)→0.63%(0.4, 0.4) \to 0.63\% underperform the asymmetric setting (ml=0.4,ms=0.1)(m_l=0.4, m_s=0.1), which reaches 0.00%0.00\% ACER (with LSim\mathcal{L}_{Sim}). Applying an excessively large margin to spoof classes (ms>0.1m_s > 0.1) impairs performance due to the heterogeneous nature of presentation attacks. Furthermore, replacing fine-grained classes with SSDG-style 4-class coarse labeling (one live class, and one spoof class per training dataset) degrades O&C&I →\to M HTER from 7.10% to 10.24% and I&C&M →\to O HTER from 11.82% to 16.26%.

  8. Knowl 8 — Device-Specific Evaluation and Few-Shot Reference-Based Face Anti-Spoofing

    empirical result

    Because camera sensor quality influences the learned patch embedding, evaluating presentation attack detection performance on individual sensor devices reveals performance discrepancies masked by aggregate dataset metrics. In cross-dataset testing on MSU-MFSD (devices M1 and M2) and CASIA-MFSD (devices C1, C2, C3), high-quality sensors (C3, M1, M2) achieve superior baseline detection compared to low-quality/high-compression sensors (C2).

    Providing KK-shot live reference face images (K∈{5,10}K \in \{5, 10\}) from the target device enables few-shot reference-based FAS. Test patch embeddings are matched against target live reference embeddings using cosine distance, where higher similarity indicates bona fide live presentation.

    OCI →\to M OMI →\to C
    Method / Protocol M1 AUC (%) M2 AUC (%) C1 AUC (%) C2 AUC (%) C3 AUC (%)
    PatchNet (0-shot) 99.54 98.63 94.29 88.49 98.13
    PatchNet (5-shot reference) 99.70 99.60 94.30 89.80 98.60
    PatchNet (10-shot reference) 99.80 99.60 95.30 90.70 99.20

    Averaged across 10 independent experimental runs, incorporating 10 live reference samples improves AUC on the challenging, low-quality C2 sensor from 88.49% to 90.70%, and boosts overall target device discriminability.

  9. Knowl 9 — Patch-Type Feature Retrieval for Training Class Refinement

    model/method

    The normalized patch embedding space learned by PatchNet enables patch-type retrieval between target domain test samples and source domain training class prototypes (classification weight vectors WjW_j).

    When a model trained on 21 fine-grained classes from MSU-MFSD (M), CASIA-MFSD (C), and Replay-Attack (I) is queried using a live bona fide patch from OULU-NPU (O), cosine similarity ranking reveals that high-quality live classes (e.g., M2_L, C3_L) rank highest, but low-quality live classes from degraded sensors (e.g., C2_L, C1_L, I_L) rank below certain high-quality spoof types (e.g., M1_S2). This indicates that severe sensor degradation introduces artifacts resembling presentation attacks.

    To mitigate this cross-domain confusion when deploying on a target dataset like OULU-NPU, the learned embedding space allows two refinement strategies:

    1. Ambiguous class removal: Excluding ambiguous low-quality live classes (C2_L, C1_L, I_L) from the classification layer during inference increases target domain AUC from 95.07% to 95.27%.
    2. Class re-labeling: Re-defining low-quality live classes (C2_L, C1_L, I_L) as spoof categories during source model training increases target domain AUC to 95.87%.

Coverage note — None was omitted; all key contributions including framework formulation, loss equations, multi-crop inference algorithm, benchmark results, ablation studies, and few-shot/retrieval applications are covered.

References

  1. 1.Z. Boulkenafet, J. Komulainen, L. Li, X. Feng, and A. Hadid. Oulu-npu: A mobile face presentation attack database with real-world variations. In IEEE International Conference on Automatic Face Gesture Recognition (FG), 2017.
  2. 2.R. Cai, H. Li, S. Wang, C. Chen, and A. C. Kot. Drl-fas: A novel framework based on deep reinforcement learning for face anti-spoofing. IEEE Transactions on Information Forensics and Security, 2021.
  3. 3.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020.
  4. 4.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. arXiv preprint arXiv:2011.10566, 2020.
  5. 5.Ivana Chingovska, Andre Anjos, and Sebastien Marcel. On the effectiveness of local binary patterns in face antispoofing. In BIOSIG-proceedings of the international conference of biometrics special interest group (BIOSIG). IEEE, 2012.
  6. 6.Debayan Deb and Anil K Jain. Look locally infer globally: A generalizable face anti-spoofing approach. IEEE Transactions on Information Forensics and Security, 2020.
  7. 7.Jiankang Deng, Jia Guo, Evangelos Ververas, Irene Kotsia, and Stefanos Zafeiriou. Retinaface: Single-shot multi-level face localisation in the wild. In CVPR, 2020.
  8. 8.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In CVPR, 2019.
  9. 9.Haocheng Feng, Zhibin Hong, Haixiao Yue, Yang Chen, Keyao Wang, Junyu Han, Jingtuo Liu, and Errui Ding. Learning generalized spoof cues for face anti-spoofing. arXiv preprint arXiv:2005.03922, 2020.
  10. 10.Jean-Bastien Grill, Florian Strub, Florent Altche, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent: A new approach to self-supervised learning. arXiv preprint arXiv:2006.07733, 2020.
  11. 11.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, 2016.
  12. 12.Yunpei Jia, Jie Zhang, Shiguang Shan, and Xilin Chen. Single-side domain generalization for face anti-spoofing. In CVPR, 2020.
  13. 13.Taewook Kim, YongHyun Kim, Inhan Kim, and Daijin Kim. Basn: Enriching feature representation using bipartite auxiliary supervisions for face anti-spoofing. In CVPR Workshops, 2019.
  14. 14.Shubao Liu, Ke-Yue Zhang, Taiping Yao, Mingwei Bi, Shouhong Ding, Jilin Li, Feiyue Huang, and Lizhuang Ma. Adaptive normalized representation learning for generalizable face anti-spoofing. In Proceedings of the 29th ACM International Conference on Multimedia, 2021.
  15. 15.Shubao Liu, Ke-Yue Zhang, Taiping Yao, Kekai Sheng, Shouhong Ding, Ying Tai, Jilin Li, Yuan Xie, and Lizhuang Ma. Dual reweighting domain generalization for face presentation attack detection. arXiv preprint arXiv:2106.16128, 2021.
  16. 16.Weiyang Liu, Yandong Wen, Zhiding Yu, Ming Li, Bhiksha Raj, and Le Song. Sphereface: Deep hypersphere embedding for face recognition. In CVPR, 2017.
  17. 17.Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang. Large-margin softmax loss for convolutional neural networks. In ICML, 2016.
  18. 18.Yaojie Liu, Amin Jourabloo, and Xiaoming Liu. Learning deep models for face anti-spoofing: Binary or auxiliary supervision. In CVPR, 2018.
  19. 19.Yaojie Liu, Joel Stehouwer, and Xiaoming Liu. On disentangling spoof trace for generic face anti-spoofing. In ECCV, 2020.
  20. 20.Rui Shao, Xiangyuan Lan, Jiawei Li, and Pong C Yuen. Multi-adversarial discriminative deep domain generalization for face presentation attack detection. In CVPR, 2019.
  21. 21.Rui Shao, Xiangyuan Lan, and Pong C Yuen. Regularized fine-grained meta face anti-spoofing. In AAAI, 2020.
  22. 22.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 2008.
  23. 23.Feng Wang, Jian Cheng, Weiyang Liu, and Haijun Liu. Additive margin softmax for face verification. IEEE Signal Processing Letters, 2018.
  24. 24.Guoqing Wang, Hu Han, Shiguang Shan, and Xilin Chen. Cross-domain face presentation attack detection via multidomain disentangled representation learning. In CVPR, 2020.
  25. 25.Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu. Cosface: Large margin cosine loss for deep face recognition. In CVPR, 2018.
  26. 26.Yu-Chun Wang, Chien-Yi Wang, and Shang-Hong Lai. Disentangled representation with dual-stage feature learning for face anti-spoofing. In WACV, 2022.
  27. 27.D. Wen, H. Han, and A. K. Jain. Face spoof detection with image distortion analysis. IEEE Transactions on Information Forensics and Security, 2015.
  28. 28.Xiao Yang, Wenhan Luo, Linchao Bao, Yuan Gao, Dihong Gong, Shibao Zheng, Zhifeng Li, and Wei Liu. Face antispoofing: Model matters, so does data. In CVPR, 2019.
  29. 29.Zitong Yu, Xiaobai Li, Xuesong Niu, Jingang Shi, and Guoying Zhao. Face anti-spoofing with human material perception. In ECCV, 2020.
  30. 30.Zitong Yu, Yunxiao Qin, Hengshuang Zhao, Xiaobai Li, and Guoying Zhao. Dual-cross central difference network for face anti-spoofing. arXiv preprint arXiv:2105.01290, 2021.
  31. 31.Z Yu, J Wan, Y Qin, X Li, SZ Li, and G Zhao. Nas-fas: Static-dynamic central difference network search for face anti-spoofing. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2020.
  32. 32.Zitong Yu, Chenxu Zhao, Zezheng Wang, Yunxiao Qin, Zhuo Su, Xiaobai Li, Feng Zhou, and Guoying Zhao. Searching central difference convolutional networks for face anti-spoofing. In CVPR, 2020.
  33. 33.Ke-Yue Zhang, Taiping Yao, Jian Zhang, Ying Tai, Shouhong Ding, Jilin Li, Feiyue Huang, Haichuan Song, and Lizhuang Ma. Face anti-spoofing via disentangled representation learning. In ECCV, 2020.
  34. 34.Z. Zhang, J. Yan, S. Liu, Z. Lei, D. Yi, and S. Z. Li. A face antispoofing database with diverse attacks. In IAPR International Conference on Biometrics (ICB), 2012.

Citation

MLA
Wang, C.-Y., et al. “PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition”. arXiv, 2022, http://arxiv.org/abs/2203.14325v1.
APA
Wang, C.-Y., Lu, Y.-D., Yang, S.-T., & Lai, S.-H. (2022). PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition. arXiv. http://arxiv.org/abs/2203.14325v1
Chicago
Wang, C.-Y., Y.-D. Lu, S.-T. Yang, and S.-H. Lai. 2022. “PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition”. arXiv. http://arxiv.org/abs/2203.14325v1.
Harvard
Wang, C.-Y. et al. (2022) “PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.14325v1.
Vancouver
1. Wang C-Y, Lu Y-D, Yang S-T, Lai S-H (2022) PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition. arXiv

BibTeX

@article{wang2022patchnet,
  title = {PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition},
  author = {Wang, Chien-Yi and Lu, Yu-Ding and Yang, Shang-Ta and Lai, Shang-Hong},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.14325v1},
  eprint = {2203.14325}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE