Rethinking Domain Generalization for Face Anti-spoofing: Separability and Alignment
Yiyou SunYaojie LiuXiaoming LiuYixuan LiWen-Sheng Chu
Proposes a face anti-spoofing framework that preserves domain-specific signals while aligning live-to-spoof transition directions through invariant risk minimization, outperforming conventional domain-invariant feature learning approaches on cross-domain benchmarks.
Face recognition systems are critical infrastructure for mobile authentication, payment security, and identity verification. However, these systems remain vulnerable to presentation attacks using printed photos or digital replays. While conventional face anti-spoofing models perform well under controlled conditions, they frequently fail when deployed across new environments, varied camera sensors, and fluctuating image resolutions. Most current solutions attempt to remove these domain-specific differences to create a single, domain-invariant feature space. The article demonstrates that this standard strategy is flawed: attempting to eliminate domain signals often forces models to rely on misleading correlations—such as confusing image blur with spoofing patterns—which degrades detection reliability on unseen test systems.
The article's main objective is to establish a new face anti-spoofing framework that maintains domain-specific variations while learning a domain-invariant decision boundary across all deployment scenarios. To achieve this, the article evaluates a strategy termed separability and alignment, implemented on a standard ResNet-18 architecture and tested across four widely recognized public face anti-spoofing benchmark datasets.
The approach combines two complementary mechanisms. First, it uses supervised contrastive learning to enforce separability, ensuring that samples from distinct environments and attack classes occupy well-defined, distinct clusters rather than being blended together. Second, it aligns the live-to-spoof transition trajectory across all domains using an optimization algorithm called projected gradient invariant risk minimization. Instead of attempting to optimize a single, rigid boundary that frequently leads to convergence failures, the algorithm iteratively updates separate boundaries for each training domain and projects them toward a unified global classifier.
The findings show that this approach substantially outperforms existing state-of-the-art methods. In standard benchmark evaluations, the proposed method lowered the error rate across cross-domain testing protocols, achieving an error reduction of over 25% relative to the best baseline on one challenging test split. Furthermore, when evaluated under a realistic convergence protocol using the final ten training epochs rather than an artificially selected best-performing snapshot, the framework sustained low error rates and high detection stability. Existing baseline methods experienced severe performance drops under stable convergence testing, confirming that prior benchmark conventions overestimated real-world robustness.
These results demonstrate that maintaining environment-specific characteristics in feature representations, rather than attempting to eliminate them, enables more dependable and transferable biometric security. For operational systems, this reduces vulnerability to presentation attacks and avoids the cost of continuously retraining models for new camera hardware. Organizations deploying facial verification should transition away from adversarial domain-invariance approaches toward separability and alignment frameworks. Future work should expand validation to 3D mask attacks and broader edge devices, as current evaluations focus primarily on print and digital replay datasets.
- Paper: Domain Generalization: A Survey, Kaiyang Zhou et al. (2021). Provides a comprehensive survey and theoretical foundation of domain generalization paradigms and statistical alignment approaches that the source analyzes and rethinks.
- Paper: PatchNet: A Simple Face Anti-Spoofing Framework via Fine-Grained Patch Recognition, Chien-Yi Wang et al. (2022). Presents a foundational baseline framework for cross-dataset face anti-spoofing using ResNet-18 and margin-based representation learning across the standard presentation attack benchmarks.
- Paper: Simultaneous Deep Transfer Across Domains and Tasks, Eric Tzeng et al. (2015). Establishes classic domain alignment and domain-invariant representation learning principles that the source paper directly critiques and seeks to replace.
- Paper: Decompose, Adjust, Compose: Effective Normalization by Playing with Frequency for Domain Generalization, Sangrok Lee et al. (2023). Builds upon domain generalization concepts by tackling feature content distortion in normalization layers to preserve intrinsic domain-invariant signals.
- Paper: Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning, Chuangchuang Tan et al. (2024). Extends generalization in facial forgery detection by leveraging frequency-space learning across unseen generative models.
- Paper: Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection, Chuangchuang Tan et al. (2024). Advances generalizable facial forensic detection across novel generators by isolating universal spatial artifact patterns from generative up-sampling operations.
