Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment

Yiyou SunYaojie LiuXiaoming LiuYixuan LiWen-Sheng Chu

article2023CVPR97 citations

Proposes a face anti-spoofing framework that preserves domain-specific signals while aligning live-to-spoof transition directions through invariant risk minimization, outperforming conventional domain-invariant feature learning approaches on cross-domain benchmarks.

Listen

Face recognition systems are critical infrastructure for mobile authentication, payment security, and identity verification. However, these systems remain vulnerable to presentation attacks using printed photos or digital replays. While conventional face anti-spoofing models perform well under controlled conditions, they frequently fail when deployed across new environments, varied camera sensors, and fluctuating image resolutions. Most current solutions attempt to remove these domain-specific differences to create a single, domain-invariant feature space. The article demonstrates that this standard strategy is flawed: attempting to eliminate domain signals often forces models to rely on misleading correlations—such as confusing image blur with spoofing patterns—which degrades detection reliability on unseen test systems.

The article's main objective is to establish a new face anti-spoofing framework that maintains domain-specific variations while learning a domain-invariant decision boundary across all deployment scenarios. To achieve this, the article evaluates a strategy termed separability and alignment, implemented on a standard ResNet-18 architecture and tested across four widely recognized public face anti-spoofing benchmark datasets.

The approach combines two complementary mechanisms. First, it uses supervised contrastive learning to enforce separability, ensuring that samples from distinct environments and attack classes occupy well-defined, distinct clusters rather than being blended together. Second, it aligns the live-to-spoof transition trajectory across all domains using an optimization algorithm called projected gradient invariant risk minimization. Instead of attempting to optimize a single, rigid boundary that frequently leads to convergence failures, the algorithm iteratively updates separate boundaries for each training domain and projects them toward a unified global classifier.

The findings show that this approach substantially outperforms existing state-of-the-art methods. In standard benchmark evaluations, the proposed method lowered the error rate across cross-domain testing protocols, achieving an error reduction of over 25% relative to the best baseline on one challenging test split. Furthermore, when evaluated under a realistic convergence protocol using the final ten training epochs rather than an artificially selected best-performing snapshot, the framework sustained low error rates and high detection stability. Existing baseline methods experienced severe performance drops under stable convergence testing, confirming that prior benchmark conventions overestimated real-world robustness.

These results demonstrate that maintaining environment-specific characteristics in feature representations, rather than attempting to eliminate them, enables more dependable and transferable biometric security. For operational systems, this reduces vulnerability to presentation attacks and avoids the cost of continuously retraining models for new camera hardware. Organizations deploying facial verification should transition away from adversarial domain-invariance approaches toward separability and alignment frameworks. Future work should expand validation to 3D mask attacks and broader edge devices, as current evaluations focus primarily on print and digital replay datasets.

Cover for Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment

Abstract

This work studies the generalization issue of face anti-spoofing (FAS) models on domain gaps, such as image resolution, blurriness and sensor variations. Most prior works regard domain-specific signals as a negative impact, and apply metric learning or adversarial losses to remove them from feature representation. Though learning a domain-invariant feature space is viable for the training data, we show that the feature shift still exists in an unseen test domain, which backfires on the generalizability of the classifier. In this work, instead of constructing a domain-invariant feature space, we encourage domain separability while aligning the live-to-spoof transition (i.e., the trajectory from live to spoof) to be the same for all domains. We formulate this FAS strategy of separability and alignment (SA-FAS) as a problem of invariant risk minimization (IRM), and learn domain-variant feature representation but domain-invariant classifier. We demonstrate the effectiveness of SA-FAS on challenging cross-domain FAS datasets and establish state-of-the-art performance. Code is available at https://github.com/sunyiyou/SAFAS.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Method
  • 3.1. Problem Setup
  • 3.2. Separability
  • 3.3. Alignment
  • 3.4. Training and inference
  • 4. Experiments
  • 4.1. Experimental setups
  • 4.2. Cross-domain performance
  • 5. Ablation and Discussion
  • 5.1. Effectiveness of loss components
  • 5.2. Separability and alignment analysis
  • 5.3. UMAP visualization
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — SA-FAS Framework for Cross-Domain Face Anti-Spoofing

    model/method

    The Separability and Alignment for Face Anti-Spoofing (SA-FAS) framework addresses cross-domain generalization in face presentation attack detection (PAD). While conventional methods aim to learn domain-invariant representations by eliminating domain discrepancy, SA-FAS retains domain-specific signals and constructs a domain-variant feature space coupled with a domain-invariant decision boundary.

    The framework consists of two core principles:

    1. Separability: Features belonging to different source domains E={e(1),e(2),…,e(E)}\mathcal{E} = \{e^{(1)}, e^{(2)}, \dots, e^{(E)}\} and binary classes Y={0 (live),1 (spoof)}\mathcal{Y} = \{0\text{ (live)}, 1\text{ (spoof)}\} are encouraged to occupy distinct, compact clusters in the feature space via Supervised Contrastive Learning (LsepL_{\text{sep}}).
    2. Alignment: The live-to-spoof transitions across domains are aligned along the same directional trajectory by formulating the classifier learning as Invariant Risk Minimization (IRM) optimized via Projected Gradient optimization (PG-IRM).

    The complete model comprises a feature encoder ϕ:X→Rm\phi: \mathcal{X} \to \mathbb{R}^m producing l2l_2-normalized feature representations z=ϕ(x)z = \phi(x), and domain-specific linear classifiers (hyperplanes) βe:Rm→R\beta_e: \mathbb{R}^m \to \mathbb{R} for each training domain e∈Ee \in \mathcal{E}. The combined training objective is:

    min⁡ϕ,βe(1),…,βe(E)Lalign+λLsep\min_{\phi, \beta_{e^{(1)}}, \dots, \beta_{e^{(E)}}} L_{\text{align}} + \lambda L_{\text{sep}}

    s.t. ∀e∈E,  ∃βe∈Ωe(ϕ),  βe∈Υα(βe)\text{s.t. } \forall e \in \mathcal{E}, \; \exists \beta_e \in \Omega_e(\phi), \; \beta_e \in \Upsilon_\alpha(\beta_e)

    where λ\lambda is a balancing hyperparameter, Ωe(ϕ)\Omega_e(\phi) represents the set of domain-optimal classifiers, and Υα(βe)\Upsilon_\alpha(\beta_e) is the α\alpha-adjacency set enforcing consistency across domains. At inference time on an unseen target domain sample x∈Xx \in \mathcal{X}, the prediction score is given by ensembling across the domain classifiers:

    f(x)=Ee∈E[βeTϕ(x)]f(x) = \mathbb{E}_{e \in \mathcal{E}} [\beta_e^T \phi(x)]

  2. Knowl 2 — Equivalence of Invariant Risk Minimization and the Projected Gradient IRM Objective

    theoretical result

    In cross-domain generalization, standard Invariant Risk Minimization (IRM) seeks an encoder ϕ:X→Rm\phi: \mathcal{X} \to \mathbb{R}^m and a single classifier β∗:Rm→R\beta^*: \mathbb{R}^m \to \mathbb{R} that minimizes the empirical risk across training domains E\mathcal{E} while remaining simultaneously optimal for every domain e∈Ee \in \mathcal{E}:

    min⁡ϕ,β∗1∣E∣∑e∈ERe(ϕ,β∗)s.t.β∗∈arg⁡min⁡βRe(ϕ,β),  ∀e∈E\min_{\phi, \beta^*} \frac{1}{|\mathcal{E}|} \sum_{e \in \mathcal{E}} R^e(\phi, \beta^*) \quad \text{s.t.} \quad \beta^* \in \arg\min_\beta R^e(\phi, \beta), \; \forall e \in \mathcal{E}

    where Re(ϕ,β)=E(x,y,e)∼D[ℓ(f(x;ϕ,β),y)]R^e(\phi, \beta) = \mathbb{E}_{(x, y, e) \sim \mathcal{D}} [\ell(f(x; \phi, \beta), y)].

    Theorem (PG-IRM Objective Equivalence): For all α∈(0,1)\alpha \in (0, 1), the bi-level IRM objective is equivalent to optimizing individual domain-wise hyperplanes βe\beta_e via the projected gradient objective:

    min⁡ϕ,βe(1),…,βe(E)1∣E∣∑e∈ERe(ϕ,βe)\min_{\phi, \beta_{e^{(1)}}, \dots, \beta_{e^{(E)}}} \frac{1}{|\mathcal{E}|} \sum_{e \in \mathcal{E}} R^e(\phi, \beta_e)

    s.t.∀e∈E,  ∃βe∈Ωe(ϕ),  βe∈Υα(βe)\text{s.t.} \quad \forall e \in \mathcal{E}, \; \exists \beta_e \in \Omega_e(\phi), \; \beta_e \in \Upsilon_\alpha(\beta_e)

    where the parametric constrained set for environment ee is defined as Ωe(ϕ)=arg⁡min⁡βRe(ϕ,β)\Omega_e(\phi) = \arg\min_\beta R^e(\phi, \beta), and the α\alpha-adjacency set Υα(βe)\Upsilon_\alpha(\beta_e) is defined as:

    Υα(βe)={β  |  max⁡e′∈E∖{e}min⁡βe′∈Ωe′(ϕ)∥β−βe′∥2≤αmax⁡e′∈E∖{e}min⁡βe′∈Ωe′(ϕ)∥βe−βe′∥2}\Upsilon_\alpha(\beta_e) = \left\{ \beta \;\middle|\; \max_{e' \in \mathcal{E} \setminus \{e\}} \min_{\beta_{e'} \in \Omega_{e'}(\phi)} \|\beta - \beta_{e'}\|_2 \le \alpha \max_{e' \in \mathcal{E} \setminus \{e\}} \min_{\beta_{e'} \in \Omega_{e'}(\phi)} \|\beta_e - \beta_{e'}\|_2 \right\}

    This reformulation replaces the potentially empty single-hyperplane constraint set with a non-empty, projectable domain-wise constraint set.

  3. Knowl 3 — Supervised Contrastive Loss for Domain-Class Separability

    equation

    To achieve feature separability across both domain identities and live/spoof classes, Supervised Contrastive Learning (SupCon) is formulated over mini-batches augmented to size 2b2b, denoted as {(x~i,y~i,e~i)}i=12b\{(\tilde{x}_i, \tilde{y}_i, \tilde{e}_i)\}_{i=1}^{2b} with corresponding l2l_2-normalized feature embeddings zi=ϕ(x~i)∈Rmz_i = \phi(\tilde{x}_i) \in \mathbb{R}^m.

    The per-batch separability loss LsepL_{\text{sep}} is defined as:

    Lsep=∑i=12b−1∣S(i)∣∑j∈S(i)log⁡exp⁡(zi⋅zj/τ)∑t=1,t≠i2bexp⁡(zi⋅zt/τ)L_{\text{sep}} = \sum_{i=1}^{2b} \frac{-1}{|S(i)|} \sum_{j \in S(i)} \log \frac{\exp(z_i \cdot z_j / \tau)}{\sum_{t=1, t \neq i}^{2b} \exp(z_i \cdot z_t / \tau)}

    where:

    • τ>0\tau > 0 is a scalar temperature parameter;
    • ii is the anchor index;
    • S(i)={j∈{1,…,2b}:j≠i,y~j=y~i,e~j=e~i}S(i) = \{j \in \{1, \dots, 2b\} : j \neq i, \tilde{y}_j = \tilde{y}_i, \tilde{e}_j = \tilde{e}_i\} is the set of positive indices containing samples that have both the identical class label (live vs. spoof) and the identical domain origin as anchor ii;
    • ∣S(i)∣|S(i)| is the cardinality of S(i)S(i).

    All samples in the batch with either a different domain or a different class label act as negative pairs, forcing the feature space to retain domain-specific characteristics in separated clusters.

  4. Knowl 4 — Projected Gradient Training Algorithm for SA-FAS

    algorithm

    The training pipeline optimizes the shared feature encoder ϕ\phi alongside domain-specific hyperplanes {βe}e∈E\{\beta_e\}_{e \in \mathcal{E}} by alternating between stochastic gradient descent and Euclidean projection via linear interpolation onto the α\alpha-adjacency set.

    Input: Training dataset D={(xi,yi,ei)}i=1ND = \{(x_i, y_i, e_i)\}_{i=1}^N, network encoder ϕ\phi, classifiers {βe}e∈E\{\beta_e\}_{e \in \mathcal{E}}, learning rate γ\gamma, alignment parameter α\alpha, alignment starting epoch TaT_a, total epochs TT.
    Output: Optimized encoder ϕ\phi and aligned classifiers {βe}e∈E\{\beta_e\}_{e \in \mathcal{E}}.
    for t=0,1,…,Tt = 0, 1, \dots, T do
        Sample and augment mini-batch to obtain pairs {(x~i,y~i,e~i)}i=12b\{(\tilde{x}_i, \tilde{y}_i, \tilde{e}_i)\}_{i=1}^{2b}
        Calculate total loss Lall=Lalign+λLsepL_{\text{all}} = L_{\text{align}} + \lambda L_{\text{sep}}
        for e∈Ee \in \mathcal{E} do
            β~et+1=βet−γ∇βetLall\tilde{\beta}_e^{t+1} = \beta_e^t - \gamma \nabla_{\beta_e^t} L_{\text{all}}
            Select farthest domain classifier eˉ=arg⁡max⁡e′∈E∖{e}∥β~et+1−βe′t∥2\bar{e} = \arg\max_{e' \in \mathcal{E} \setminus \{e\}} \|\tilde{\beta}_e^{t+1} - \beta_{e'}^t\|_2
            α′=1−1t>Ta(1−α)\alpha' = 1 - \mathbf{1}_{t > T_a} (1 - \alpha)
            βet+1=α′β~et+1+(1−α′)βeˉt\beta_e^{t+1} = \alpha' \tilde{\beta}_e^{t+1} + (1 - \alpha') \beta_{\bar{e}}^t
        end for
        ϕt+1=ϕt−γ∇ϕtLall\phi^{t+1} = \phi^t - \gamma \nabla_{\phi^t} L_{\text{all}}
    end for

    Key execution parameters in the framework are set to α=0.995\alpha = 0.995, λ=0.1\lambda = 0.1, Ta=20T_a = 20, with base learning rate γ=5×10−3\gamma = 5\times 10^{-3} and weight decay 5×10−45\times 10^{-4}.

  5. Knowl 5 — Separability and Alignment Quantitative Metrics

    definition

    To evaluate the geometric properties of the learned feature space and its decision boundaries on an unseen test domain without relying solely on classification error, two quantitative metrics are defined:

    1. Separability Score (SsepS_{\text{sep}}): Measures the angular separation between the live and spoof feature representations in the test domain:

    Ssep=1−cos⁡(Espoof[z],  Elive[z])S_{\text{sep}} = 1 - \cos\left(\mathbb{E}_{\text{spoof}}[z], \; \mathbb{E}_{\text{live}}[z]\right)

    where z=ϕ(x)∈Rmz = \phi(x) \in \mathbb{R}^m is the l2l_2-normalized feature embedding, and Espoof[z]\mathbb{E}_{\text{spoof}}[z] and Elive[z]\mathbb{E}_{\text{live}}[z] denote the mean feature vectors of spoof and live test samples, respectively. A higher SsepS_{\text{sep}} indicates greater class separation.

    1. Alignment Score (SalignS_{\text{align}}): Measures the cosine similarity between the learned domain hyperplane normal vectors βe\beta_e and the oracle transition vector from live to spoof in the test domain:

    Salign=Ee∈E[cos⁡(βe,  Espoof[z]−Elive[z])]S_{\text{align}} = \mathbb{E}_{e \in \mathcal{E}} \left[ \cos\left(\beta_e, \; \mathbb{E}_{\text{spoof}}[z] - \mathbb{E}_{\text{live}}[z]\right) \right]

    where Espoof[z]−Elive[z]\mathbb{E}_{\text{spoof}}[z] - \mathbb{E}_{\text{live}}[z] serves as the oracle live-to-spoof transition direction in the embedding space. A higher SalignS_{\text{align}} indicates that the classifiers are aligned with the true attack transition trajectory.

  6. Knowl 6 — Cross-Domain Face Anti-Spoofing Benchmark Protocols and Evaluation Setup

    experimental setup

    Cross-domain performance is evaluated across four standard public datasets: CASIA-FASD (C), Idiap Replay-Attack (I), MSU-MFSD (M), and Oulu-NPU (O). A leave-one-out cross-validation protocol evaluates generalization to unseen domains: OCI→M\text{OCI}\to\text{M}, OMI→C\text{OMI}\to\text{C}, OCM→I\text{OCM}\to\text{I}, and ICM→O\text{ICM}\to\text{O} (where letters denote training sets and the arrow indicates the held-out test set).

    Implementation Details:

    • Face images are cropped via MTCNN and resized to 256×256256 \times 256.
    • Backbone architecture is ResNet-18 producing l2l_2-normalized penultimate embeddings.
    • Optimization uses SGD with an initial learning rate of 5×10−35\times 10^{-3} (decayed by a factor of 2 at epochs 40 and 80 for 100 total epochs; ICM→O\text{ICM}\to\text{O} runs for 300 epochs, decaying at 120 and 240).
    • Weight decay is 5×10−45\times 10^{-4} and per-domain batch size is 96.
    • SA-FAS hyperparameters: α=0.995\alpha = 0.995, λ=0.1\lambda = 0.1, and alignment start epoch Ta=20T_a = 20.

    Evaluation Metrics:

    • Half Total Error Rate (HTER=FAR+FRR2\text{HTER} = \frac{\text{FAR} + \text{FRR}}{2}, lower is better);
    • Area Under the Receiver Operating Characteristic Curve (AUC\text{AUC}, higher is better);
    • True Positive Rate at a False Positive Rate of 5% (TPR95\text{TPR95}, higher is better).
  7. Knowl 7 — Cross-Domain FAS Performance under Converged Evaluation Protocol

    data/table

    Reporting cross-domain PAD metrics using the best-performing test snapshot introduces selection bias and high variance. The converged evaluation protocol reports the mean and standard deviation over the last 10 training epochs upon convergence (where the binary loss drops below 10−310^{-3} for 10 consecutive epochs or maximum epochs are reached).

    Method (%) OCI→\toM OMI→\toC OCM→\toI ICM→\toO
    HTER↓\downarrow / AUC↑\uparrow / TPR95↑\uparrow HTER↓\downarrow / AUC↑\uparrow / TPR95↑\uparrow HTER↓\downarrow / AUC↑\uparrow / TPR95↑\uparrow HTER↓\downarrow / AUC↑\uparrow / TPR95↑\uparrow
    SSDG-R 14.651.21^{1.21} / 91.931.35^{1.35} / 53.682.56^{2.56} 28.760.89^{0.89} / 80.911.10^{1.10} / 41.472.68^{2.68} 22.841.14^{1.14} / 78.671.31^{1.31} / 50.805.95^{5.95} 15.831.29^{1.29} / 92.130.96^{0.96} / 66.544.00^{4.00}
    SSAN-R 21.793.68^{3.68} / 84.063.78^{3.78} / 51.914.28^{4.28} 26.442.91^{2.91} / 78.842.83^{2.83} / 45.364.29^{4.29} 35.398.04^{8.04} / 70.139.03^{9.03} / 64.002.70^{2.70} 25.723.74^{3.74} / 79.374.69^{4.69} / 36.755.19^{5.19}
    PatchNet 25.921.13^{1.13} / 83.430.87^{0.87} / 38.758.31^{8.31} 36.261.98^{1.98} / 71.381.89^{1.89} / 19.223.85^{3.85} 29.752.76^{2.76} / 80.531.35^{1.35} / 54.252.18^{2.18} 23.491.80^{1.80} / 84.621.92^{1.92} / 39.396.83^{6.83}
    SA-FAS (Ours) 14.361.10^{1.10} / 92.060.53^{0.53} / 55.714.82^{4.82} 19.400.66^{0.66} / 88.690.67^{0.67} / 50.533.60^{3.60} 11.481.10^{1.10} / 95.740.55^{0.55} / 77.053.26^{3.26} 11.290.32^{0.32} / 95.230.24^{0.24} / 73.381.64^{1.64}

    Under this evaluation setting:

    1. Error rates for all methods are noticeably higher than single-snapshot best numbers, demonstrating that conventional metrics overestimate generalization.
    2. Adversarial domain-generalization methods (e.g., SSAN-R) exhibit large standard deviations across late epochs.
    3. SA-FAS demonstrates consistent performance across all four cross-domain benchmarks, achieving the lowest HTER and highest AUC and TPR95 with lower variance.
  8. Knowl 8 — Best-Snapshot Cross-Domain FAS Performance Comparison

    data/table

    Cross-domain face anti-spoofing performance evaluated under the conventional best-test-snapshot protocol across CASIA (C), Idiap Replay (I), MSU-MFSD (M), and Oulu-NPU (O):

    Method (%) OCI→\toM OMI→\toC OCM→\toI ICM→\toO
    HTER↓\downarrow AUC↑\uparrow HTER↓\downarrow AUC↑\uparrow HTER↓\downarrow AUC↑\uparrow HTER↓\downarrow AUC↑\uparrow
    MMD-AAE 27.08 83.19 44.59 58.29 31.58 75.18 40.98 63.08
    MADDG 17.69 88.06 24.50 84.51 22.19 84.99 27.98 80.02
    SSDG-M 16.67 90.47 23.11 85.45 18.21 94.61 25.17 81.83
    DR-MD-Net 17.02 90.10 19.68 87.43 20.87 86.72 25.02 81.47
    RFMeta 13.89 93.98 20.27 88.16 17.30 90.48 16.45 91.16
    NAS-FAS 19.53 88.63 16.54 90.18 14.51 93.84 13.80 93.43
    D2AM 12.70 95.66 20.98 85.58 15.43 91.22 15.27 90.87
    SDA 15.40 91.80 24.50 84.40 15.60 90.10 23.10 84.30
    DRDG 12.43 95.81 19.05 88.79 15.56 91.79 15.63 91.75
    ANRL 10.83 96.75 17.83 89.26 16.03 91.04 15.67 91.90
    SSAN-M 10.42 94.76 16.47 90.81 14.00 94.58 19.51 88.17
    SSDG-R 7.38 97.17 10.44 95.94 11.71 96.59 15.61 91.54
    SSAN-R 6.67 98.75 10.00 96.67 8.88 96.79 13.72 93.63
    PatchNet 7.10 98.46 11.33 94.58 13.40 95.67 11.82 95.07
    SA-FAS (Ours) 5.95 96.55 8.78 95.37 6.58 97.54 10.00 96.23

    SA-FAS achieves the lowest HTER across all four leave-one-out protocols, outperforming the previous best baseline SSAN-R on OCM→I\text{OCM}\to\text{I} by 2.30% in HTER (a relative error reduction exceeding 25%).

  9. Knowl 9 — Ablation Study of Loss Formulations for Separability and Invariant Risk Minimization

    data/table

    Ablation performance averaged over all four cross-domain FAS benchmarks (OCI→\toM, OMI→\toC, OCM→\toI, ICM→\toO) compares various representation learning losses and IRM formulation strategies:

    Method HTER (%) ↓\downarrow AUC (%) ↑\uparrow TPR95 (%) ↑\uparrow
    SimCLR 22.531.31^{1.31} 84.421.04^{1.04} 51.143.44^{3.44}
    SimSiam 18.890.97^{0.97} 89.930.80^{0.80} 56.622.88^{2.88}
    Triplet 18.752.31^{2.31} 88.112.30^{2.30} 50.538.76^{8.76}
    SupCon (SSDG) 17.911.05^{1.05} 90.100.68^{0.68} 61.982.87^{2.87}
    SupCon 17.031.73^{1.73} 90.681.29^{1.29} 56.725.06^{5.06}
    ERM 17.221.26^{1.26} 90.211.38^{1.38} 58.623.77^{3.77}
    DANN 17.931.02^{1.02} 90.660.56^{0.56} 58.663.14^{3.14}
    IRM-v1 17.410.77^{0.77} 91.160.52^{0.52} 60.982.10^{2.10}
    VREx 25.021.92^{1.92} 80.652.20^{2.20} 45.123.78^{3.78}
    IB-IRM 17.570.74^{0.74} 91.710.51^{0.51} 62.162.35^{2.35}
    PG-IRM (Ours) 15.580.96^{0.96} 92.030.62^{0.62} 63.312.59^{2.59}
    SA-FAS (Ours) 14.250.79^{0.79} 92.930.49^{0.49} 64.163.33^{3.33}

    Key takeaways:

    1. Separability losses: Supervised Contrastive Learning (SupCon) outperforms self-supervised (SimCLR, SimSiam) and metric learning (Triplet) alternatives by forming distinct clusters per domain-class pair.
    2. Alignment/Invariance losses: Lagrangian-penalty IRM approximations (IRM-v1, VREx, IB-IRM) yield marginal improvements or instability in deep networks, whereas PG-IRM directly enforces consistency among hyperplanes via projected gradient descent, reducing average HTER to 15.58%.
    3. Combined Framework: Combining SupCon and PG-IRM in SA-FAS yields the best overall performance (14.25% HTER, 92.93% AUC, 64.16% TPR95).

Coverage note — None was omitted; all key contributions including framework formulation, theoretical equivalence, training algorithm, diagnostic metrics, convergence evaluation, benchmark comparisons, and loss ablations are fully covered.

References

  1. 1.Kartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet, Yoshua Bengio, Ioannis Mitliagkas, and Irina Rish. Invariance principle meets information bottleneck for out-of-distribution generalization. Advances in Neural Information Processing Systems, 34:3438–3450, 2021. 3, 7
  2. 2.Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz. Invariant risk minimization. arXiv preprint arXiv:1907.02893, 2019. 3, 4, 7
  3. 3.Yousef Atoum, Yaojie Liu, Amin Jourabloo, and Xiaoming Liu. Face anti-spoofing using patch and depth-based CNNs. In 2017 IEEE International Joint Conference on Biometrics (IJCB), pages 319–328. IEEE, 2017. 1, 2
  4. 4.Shai Ben-David, John Blitzer, Koby Crammer, and Fernando Pereira. Analysis of representations for domain adaptation. Advances in neural information processing systems, 19, 2006. 2
  5. 5.Gilles Blanchard, Aniket Anand Deshmukh, Ürun Dogan, Gyemin Lee, and Clayton Scott. Domain generalization by marginal transfer learning. The Journal of Machine Learning Research, 22(1):46–100, 2021. 3, 8
  6. 6.Zinelabidine Boulkenafet, Jukka Komulainen, and Abdenour Hadid. Face anti-spoofing based on color texture analysis. In 2015 IEEE international conference on image processing (ICIP), pages 2636–2640. IEEE, 2015. 1, 2
  7. 7.Zinelabinde Boulkenafet, Jukka Komulainen, Lei Li, Xiaoyi Feng, and Abdenour Hadid. Oulu-npu: A mobile face presentation attack database with real-world variations. In 2017 12th IEEE international conference on automatic face & gesture recognition (FG 2017), pages 612–618. IEEE, 2017. 6
  8. 8.Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised learning of visual features by contrasting cluster assignments. Proceedings of Advances in Neural Information Processing Systems, 2020. 3
  9. 9.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In Proceedings of the international conference on machine learning, pages 1597–1607. PMLR, 2020. 3, 7
  10. 10.Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15750–15758, 2021. 3, 7
  11. 11.Zhihong Chen, Taiping Yao, Kekai Sheng, Shouhong Ding, Ying Tai, Jilin Li, Feiyue Huang, and Xinyu Jin. Generalizable representation learning for mixture domain face anti-spoofing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1132–1139, 2021. 2, 6
  12. 12.Girija Chetty. Biometric liveness checking using multimodal fuzzy fusion. In International Conference on Fuzzy Systems, pages 1–8. IEEE, 2010. 2
  13. 13.Ivana Chingovska, André Anjos, and Sébastien Marcel. On the effectiveness of local binary patterns in face anti-spoofing. In 2012 BIOSIG-proceedings of the international conference of biometrics special interest group (BIOSIG), pages 1–7. IEEE, 2012. 6
  14. 14.Yo Joong Choe, Jiyeon Ham, and Kyubyong Park. An empirical study of invariant risk minimization. arXiv preprint arXiv:2004.05007, 2020. 3
  15. 15.Debayan Deb, Xiaoming Liu, and Anil Jain. Unified detection of digital and physical face attacks. In Proceedings of the 17th International Conference on Automatic Face and Gesture Recognition (FG), 2023. 2
  16. 16.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4690–4699, 2019. 1
  17. 17.Tiago de Freitas Pereira, André Anjos, José Mario De Martino, and Sébastien Marcel. LBP-TOP based countermeasure against face spoofing attacks. In Asian Conference on Computer Vision, pages 121–132. Springer, 2012. 1, 2
  18. 18.Chuang Gan, Tianbao Yang, and Boqing Gong. Learning attributes equals multi-source domain generalization. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 87–97, 2016. 3
  19. 19.Yaroslav Ganin and Victor Lempitsky. Unsupervised domain adaptation by backpropagation. In International conference on machine learning, pages 1180–1189. PMLR, 2015. 3
  20. 20.Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural networks. The journal of machine learning research, 17(1):2096–2030, 2016. 3
  21. 21.Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. Domain-adversarial training of neural networks. The journal of machine learning research, 17(1):2096–2030, 2016. 7, 8
  22. 22.Anjith George and Sébastien Marcel. Deep pixel-wise binary supervision for face presentation attack detection. In 2019 International Conference on Biometrics (ICB), pages 1–8. IEEE, 2019. 3
  23. 23.Anjith George and Sébastien Marcel. On the effectiveness of vision transformers for zero-shot face anti-spoofing. In 2021 IEEE International Joint Conference on Biometrics (IJCB), pages 1–8. IEEE, 2021. 2
  24. 24.Muhammad Ghifary, David Balduzzi, W Bastiaan Kleijn, and Mengjie Zhang. Scatter component analysis: A unified framework for domain adaptation and domain generalization. IEEE transactions on pattern analysis and machine intelligence, 39(7):1414–1430, 2016. 3
  25. 25.Rui Gong, Wen Li, Yuhua Chen, and Luc Van Gool. Dlow: Domain flow for adaptation and generalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2477–2486, 2019. 3
  26. 26.Thomas Grubinger, Adriana Birlutiu, Holger Schöner, Thomas Natschläger, and Tom Heskes. Domain generalization based on transfer component analysis. In International Work-Conference on Artificial Neural Networks, pages 325–334. Springer, 2015. 3
  27. 27.Xiao Guo, Yaojie Liu, Anil Jain, and Xiaoming Liu. Multi-domain learning for updating face anti-spoofing models. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XIII, pages 230–249. Springer, 2022. 2
  28. 28.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 9729–9738, 2020. 3
  29. 29.Hsin-Ping Huang, Deqing Sun, Yaojie Liu, Wen-Sheng Chu, Taihong Xiao, Jinwei Yuan, Hartwig Adam, and Ming-Hsuan Yang. Adaptive transformers for robust few-shot cross-domain face anti-spoofing. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XIII, pages 37–54. Springer, 2022. 2
  30. 30.Yunpei Jia, Jie Zhang, Shiguang Shan, and Xilin Chen. Single-side domain generalization for face anti-spoofing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8484–8493, 2020. 2, 3, 6, 7, 8, 15, 16
  31. 31.Amin Jourabloo, Yaojie Liu, and Xiaoming Liu. Face de-spoofing: Anti-spoofing via noise modeling. In Proceedings of the European conference on computer vision (ECCV), pages 290–306, 2018. 2
  32. 32.Pritish Kamath, Akilesh Tangella, Danica Sutherland, and Nathan Srebro. Does invariant risk minimization capture invariance? In International Conference on Artificial Intelligence and Statistics, pages 4069–4077. PMLR, 2021. 3, 4, 7
  33. 33.Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. Supervised contrastive learning. In Advances in Neural Information Processing Systems, volume 33, pages 18661–18673, 2020. 2, 3, 7
  34. 34.Taewook Kim and Yonghyun Kim. Suppressing spoof-irrelevant factors for domain-agnostic face anti-spoofing. IEEE Access, 9:86966–86974, 2021. 2
  35. 35.Taewook Kim, YongHyun Kim, Inhan Kim, and Daijin Kim. Basn: Enriching feature representation using bipartite auxiliary supervisions for face anti-spoofing. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 1, 2
  36. 36.Klaus Kollreider, Hartwig Fronthaler, Maycel Isaac Faraj, and Josef Bigun. Real-time face detection and motion analysis with application in “liveness” assessment. IEEE Transactions on Information Forensics and Security, 2(3):548–558, 2007. 2
  37. 37.David Krueger, Ethan Caballero, Joern-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Remi Le Priol, and Aaron Courville. Out-of-distribution generalization via risk extrapolation (rex). In International Conference on Machine Learning, pages 5815–5826. PMLR, 2021. 3, 7
  38. 38.David Krueger, Ethan Caballero, Joern-Henrik Jacobsen, Amy Zhang, Jonathan Binas, Dinghuai Zhang, Remi Le Priol, and Aaron Courville. Out-of-distribution generalization via risk extrapolation (rex). In International Conference on Machine Learning, pages 5815–5826. PMLR, 2021. 7
  39. 39.Haoliang Li, Wen Li, Hong Cao, Shiqi Wang, Feiyue Huang, and Alex C Kot. Unsupervised domain adaptation for face anti-spoofing. IEEE Transactions on Information Forensics and Security, 13(7):1794–1809, 2018. 2
  40. 40.Haoliang Li, Sinno Jialin Pan, Shiqi Wang, and Alex C Kot. Domain generalization with adversarial feature learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5400–5409, 2018. 3, 6
  41. 41.Lei Li, Xiaoyi Feng, Zinelabidine Boulkenafet, Zhaoqiang Xia, Mingming Li, and Abdenour Hadid. An original face anti-spoofing approach using partial convolutional neural network. In 2016 Sixth International Conference on Image Processing Theory, Tools and Applications (IPTA), pages 1–6. IEEE, 2016. 1
  42. 42.Ya Li, Xinmei Tian, Mingming Gong, Yajing Liu, Tongliang Liu, Kun Zhang, and Dacheng Tao. Deep domain generalization via conditional invariant adversarial networks. In Proceedings of the European Conference on Computer Vision (ECCV), pages 624–639, 2018. 3
  43. 43.Shubao Liu, Ke-Yue Zhang, Taiping Yao, Mingwei Bi, Shouhong Ding, Jilin Li, Feiyue Huang, and Lizhuang Ma. Adaptive normalized representation learning for generalizable face anti-spoofing. In Proceedings of the 29th ACM International Conference on Multimedia, pages 1469–1477, 2021. 2, 6
  44. 44.Shubao Liu, Ke-Yue Zhang, Taiping Yao, Kekai Sheng, Shouhong Ding, Ying Tai, Jilin Li, Yuan Xie, and Lizhuang Ma. Dual reweighting domain generalization for face presentation attack detection. arXiv preprint arXiv:2106.16128, 2021. 2, 6
  45. 45.Yaojie Liu, Amin Jourabloo, and Xiaoming Liu. Learning deep models for face anti-spoofing: Binary or auxiliary supervision. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 389–398, 2018. 1, 2, 3
  46. 46.Yaojie Liu and Xiaoming Liu. Spoof trace disentanglement for generic face anti-spoofing. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3):3813–3830, 2023. 2
  47. 47.Yaojie Liu, Joel Stehouwer, Amin Jourabloo, and Xiaoming Liu. Deep tree learning for zero-shot face anti-spoofing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4680–4689, 2019. 1, 2
  48. 48.Yaojie Liu, Joel Stehouwer, and Xiaoming Liu. On disentangling spoof trace for generic face anti-spoofing. In European Conference on Computer Vision, pages 406–422. Springer, 2020. 2
  49. 49.Leland McInnes, John Healy, and James Melville. UMAP: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018. 8
  50. 50.Jovana Mitrovic, Brian McWilliams, Jacob Walker, Lars Buesing, and Charles Blundell. Representation learning via invariant causal mechanisms. arXiv preprint arXiv:2010.07922, 2020. 3
  51. 51.Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf. Domain generalization via invariant feature representation. In International Conference on Machine Learning, pages 10–18. PMLR, 2013. 3
  52. 52.Jorge Nocedal and Stephen J Wright. Numerical optimization. Springer, 1999. 4
  53. 53.Gang Pan, Lin Sun, Zhaohui Wu, and Shihong Lao. Eyeblink-based anti-spoofing in face recognition from a generic webcamera. In 2007 IEEE 11th international conference on computer vision, pages 1–8. IEEE, 2007. 2
  54. 54.Vardan Papyan, XY Han, and David L Donoho. Prevalence of neural collapse during the terminal phase of deep learning training. Proceedings of the National Academy of Sciences, 117(40):24652–24663, 2020. 2
  55. 55.Keyurkumar Patel, Hu Han, and Anil K Jain. Secure face unlock: Spoof detection on smartphones. IEEE transactions on information forensics and security, 11(10):2268–2283, 2016. 2
  56. 56.Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktashmotlagh, and Sridha Sridharan. Correlation-aware adversarial domain adaptation and generalization. Pattern Recognition, 100:107124, 2020. 3
  57. 57.Elan Rosenfeld, Pradeep Kumar Ravikumar, and Andrej Risteski. The risks of invariant risk minimization. In International Conference on Learning Representations, 2020. 3, 4, 7
  58. 58.Suman Saha, Wenhao Xu, Menelaos Kanakis, Stamatios Georgoulis, Yuhua Chen, Danda Pani Paudel, and Luc Van Gool. Domain agnostic feature learning for image and video based face anti-spoofing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 802–803, 2020. 2
  59. 59.Florian Schroff, Dmitry Kalenichenko, and James Philbin. Facenet: A unified embedding for face recognition and clustering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 815–823, 2015. 7
  60. 60.Rui Shao, Xiangyuan Lan, Jiawei Li, and Pong C Yuen. Multi-adversarial discriminative deep domain generalization for face presentation attack detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10023–10031, 2019. 2, 3, 6
  61. 61.Rui Shao, Xiangyuan Lan, and Pong C Yuen. Regularized fine-grained meta face anti-spoofing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 11974–11981, 2020. 2, 6
  62. 62.Anoopkumar Sonar, Vincent Pacelli, and Anirudha Majumdar. Invariant policy optimization: Towards stronger generalization in reinforcement learning. In Learning for Dynamics and Control, pages 21–33. PMLR, 2021. 3
  63. 63.Joel Stehouwer, Amin Jourabloo, Yaojie Liu, and Xiaoming Liu. Noise modeling, synthesis and classification for generic object anti-spoofing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 7294–7303, 2020. 2
  64. 64.Aaron Van den Oord, Yazhe Li, and Oriol Vinyals. Representation learning with contrastive predictive coding. arXiv e-prints, pages arXiv–1807, 2018. 3
  65. 65.Vladimir Vapnik. Principles of risk minimization for learning theory. Advances in neural information processing systems, 4, 1991. 3
  66. 66.Chien-Yi Wang, Yu-Ding Lu, Shang-Ta Yang, and Shang-Hong Lai. PatchNet: A simple face anti-spoofing framework via fine-grained patch recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20281–20290, 2022. 1, 2, 6, 7, 15
  67. 67.Guoqing Wang, Hu Han, Shiguang Shan, and Xilin Chen. Improving cross-database face presentation attack detection via adversarial domain adaptation. In 2019 International Conference on Biometrics (ICB), pages 1–8. IEEE, 2019. 2
  68. 68.Guoqing Wang, Hu Han, Shiguang Shan, and Xilin Chen. Cross-domain face presentation attack detection via multi-domain disentangled representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6678–6687, 2020. 6
  69. 69.Guoqing Wang, Hu Han, Shiguang Shan, and Xilin Chen. Unsupervised adversarial domain adaptation for cross-domain face presentation attack detection. IEEE Transactions on Information Forensics and Security, 16:56–69, 2020. 2
  70. 70.Jindong Wang, Cuiling Lan, Chang Liu, Yidong Ouyang, Tao Qin, Wang Lu, Yiqiang Chen, Wenjun Zeng, and Philip Yu. Generalizing to unseen domains: A survey on domain generalization. IEEE Transactions on Knowledge and Data Engineering, 2022. 2
  71. 71.Jingjing Wang, Jingyi Zhang, Ying Bian, Youyi Cai, Chunmao Wang, and Shiliang Pu. Self-domain adaptation for face anti-spoofing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 2746–2754, 2021. 2, 6
  72. 72.Zhuo Wang, Zezheng Wang, Zitong Yu, Weihong Deng, Jiahong Li, Tingting Gao, and Zhongyuan Wang. Domain generalization via shuffled style assembly for face anti-spoofing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4123–4133, 2022. 2, 3, 6, 7, 15
  73. 73.Di Wen, Hu Han, and Anil K Jain. Face spoof detection with image distortion analysis. IEEE Transactions on Information Forensics and Security, 10(4):746–761, 2015. 6
  74. 74.Jianwei Yang, Zhen Lei, and Stan Z Li. Learn convolutional neural network for face anti-spoofing. arXiv preprint arXiv:1408.5601, 2014. 1, 2
  75. 75.Jianwei Yang, Zhen Lei, Shengcai Liao, and Stan Z Li. Face liveness detection with component dependent descriptor. In 2013 International Conference on Biometrics (ICB), pages 1–6. IEEE, 2013. 2
  76. 76.Zitong Yu, Xiaobai Li, Xuesong Niu, Jingang Shi, and Guoying Zhao. Face anti-spoofing with human material perception. In European conference on computer vision, pages 557–575. Springer, 2020. 1, 2
  77. 77.Zitong Yu, Jun Wan, Yunxiao Qin, Xiaobai Li, Stan Z Li, and Guoying Zhao. NAS-FAS: Static-dynamic central difference network search for face anti-spoofing. IEEE transactions on pattern analysis and machine intelligence, 43(9):3005–3023, 2020. 6
  78. 78.Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE signal processing letters, 23(10):1499–1503, 2016. 6
  79. 79.Zhiwei Zhang, Junjie Yan, Sifei Liu, Zhen Lei, Dong Yi, and Stan Z Li. A face antispoofing database with diverse attacks. In 2012 5th IAPR international conference on Biometrics (ICB), pages 26–31. IEEE, 2012. 6
  80. 80.Qianyu Zhou, Ke-Yue Zhang, Taiping Yao, Ran Yi, Kekai Sheng, Shouhong Ding, and Lizhuang Ma. Generative domain adaptation for face anti-spoofing. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part V, pages 335–356. Springer, 2022. 2

Citation

MLA
Sun, Y., et al. “Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment”. arXiv, 2023, http://arxiv.org/abs/2303.13662v1.
APA
Sun, Y., Liu, Y., Liu, X., Li, Y., & Chu, W.-S. (2023). Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment. arXiv. http://arxiv.org/abs/2303.13662v1
Chicago
Sun, Y., Y. Liu, X. Liu, Y. Li, and W.-S. Chu. 2023. “Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment”. arXiv. http://arxiv.org/abs/2303.13662v1.
Harvard
Sun, Y. et al. (2023) “Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2303.13662v1.
Vancouver
1. Sun Y, Liu Y, Liu X, Li Y, Chu W-S (2023) Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment. arXiv

BibTeX

@article{sun2023rethinking,
  title = {Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment},
  author = {Sun, Yiyou and Liu, Yaojie and Liu, Xiaoming and Li, Yixuan and Chu, Wen-Sheng},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2303.13662v1},
  eprint = {2303.13662}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE