Implicit Identity Driven Deepfake Face Swapping Detection

Baojin HuangZhongyuan WangJifan YangJiaxin AiQin ZouQian WangDengpan Ye

article2023CVPR197 citations

Proposes an implicit identity-driven framework that detects deepfake face swaps by identifying discrepancies between a face's explicit visual identity and its underlying target identity.

Listen

The rapid advancement of deep learning has made face-swapping manipulation technologies capable of creating highly realistic fake videos of public figures and private individuals. These realistic manipulations pose serious security, political, and commercial risks, increasing the demand for reliable digital forensics. Existing detection tools typically treat forgery identification as basic image classification or focus on narrow visual artifacts like local noise and blending boundaries. Consequently, these systems often fail to generalize when encountering unfamiliar manipulation methods, compressed video files, or real-world operational environments.

The article demonstrates an implicit identity driven detection framework designed to improve face-swapping identification across unseen manipulations. The primary objective is to evaluate whether contrasting a face's visible outward appearance with its residual target identity can reliably separate authentic media from synthetic manipulations.

The approach introduces two concepts: explicit identity, which is what the face visibly looks like, and implicit identity, which represents the underlying identity traits left behind by the target face. In an authentic face, explicit and implicit identities match, whereas swapped faces exhibit a divergence between the two. The framework extracts visible features using a standard face recognition system and trains a neural network backbone to map implicit features using two specialized guidance objectives. The first objective pulls authentic faces closer to their visible identity while pushing synthetic faces away. The second objective explores target identity traits by grouping known fake faces around their underlying target labels and ensuring identity consistency across video frames. The final prediction relies on the similarity distance between these explicit and implicit representations.

Key experimental results demonstrate strong generalization advantages. In cross-dataset evaluations on the highly challenging Deepfake Detection Challenge benchmark, the model achieved an area under the curve of 81.23%, outperforming the previous leading baseline by 4.52 percentage points. When trained on standard data and tested against unseen, high-fidelity manipulation techniques such as FaceShifter, the model delivered a 5.04 percentage point improvement over previous methods. Furthermore, under low-quality compressed video evaluations on the Celeb-DF benchmark, it improved detection performance by 4.69 percentage points over leading alternatives. Ablation experiments verified that combining both identity contrast and implicit exploration losses improved base classification accuracy on benchmark datasets by roughly 9 to 10 percentage points.

These findings indicate that identity-inconsistency detection captures fundamental, manipulation-agnostic evidence of tampering rather than superficial surface artifacts. Deploying identity-driven architectures lowers operational risk by significantly reducing false acceptance rates when organizations face emerging manipulation tools. For security and compliance leaders, this shift offers a more resilient, future-proof defense against deceptive media campaigns without requiring constant retraining on every newly published generation algorithm.

Organizations tasked with media verification and identity security should consider testing identity-discrepancy detection methods alongside traditional artifact-based detectors in multi-layered screening pipelines. Before operational deployment, teams should conduct internal pilot evaluations on their own operational video streams to identify optimal similarity thresholds. Caution is warranted because the model exhibits slight performance trade-offs on in-domain training data compared to specialized narrow detectors, and expression-only manipulations (such as facial reenactment) do not involve identity swapping and therefore require modified constraint settings.

Cover for Implicit Identity Driven Deepfake Face Swapping Detection

Abstract

In this paper, we consider the face swapping detection from the perspective of face identity. Face swapping aims to replace the target face with the source face and generate the fake face that the human cannot distinguish between real and fake. We argue that the fake face contains the explicit identity and implicit identity, which respectively corresponds to the identity of the source face and target face during face swapping. Note that the explicit identities of faces can be extracted by regular face recognizers. Particularly, the implicit identity of real face is consistent with the its explicit identity. Thus the difference between explicit and implicit identity of face facilitates face swapping detection. Following this idea, we propose a novel implicit identity driven framework for face swapping detection. Specifically, we design an explicit identity contrast (EIC) loss and an implicit identity exploration (IIE) loss, which supervises a CNN backbone to embed face images into the implicit identity space. Under the guidance of EIC, real samples are pulled closer to their explicit identities, while fake samples are pushed away from their explicit identities. Moreover, IIE is derived from the margin-based classification loss function, which encourages the fake faces with known target identities to enjoy intra-class compactness and inter-class diversity. Extensive experiments and visualizations on several datasets demonstrate the generalization of our method against the state-of-the-art counterparts.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Face Swapping
  • 2.2. Face Forgery Detection
  • 3. Proposed Method
  • 3.1. Explicit Identity Contrast
  • 3.2. Implicit Identity Exploration
  • 3.3. Overall Loss Function
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Ablation Study
  • 4.3. Quantitative Results
  • 4.4. Visualization
  • 5. Conclusion
  • 6. Acknowledgement
  • References

Knowls

  1. Knowl 1 — Explicit and Implicit Identity Formulation for Face Swapping Detection

    definition

    In the context of face swapping detection, a facial image is characterized by two distinct identity representations:

    1. Explicit Identity: The visual identity perceived from the face image, corresponding to the source face identity in a manipulated (fake) face or the authentic subject identity in a pristine (real) face. Explicit identity embeddings Fex(x)F_{ex}(x) are extracted using an off-the-shelf, fixed face recognition model (such as CosFace trained on WebFace).

    2. Implicit Identity: The underlying target face identity retained in a manipulated face due to imperfect identity decoupling by generative face swapping algorithms. For an authentic face, the implicit identity is identical to its explicit identity. For a swapped face, the implicit identity corresponds to the original target face onto which the source face was composited.

    Face swapping detection is thereby formulated as evaluating identity consistency: when the explicit identity embedding Fex(x)F_{ex}(x) and the learned implicit identity embedding Fim(x)F_{im}(x) of an input face xx are close in feature space, the face is determined to be real; when they diverge, the face is determined to be fake.

  2. Knowl 2 — Explicit Identity Contrast Loss

    equation

    The Explicit Identity Contrast (EIC) loss supervises an implicit identity embedding network FimF_{im} using explicit identity features Fex(xi)F_{ex}(x_i) extracted by a fixed generic face recognition network as a reference:

    Leic=1NF∑i∈Fδ(Fim(xi),Fex(xi))−1NR∑i∈Rδ(Fim(xi),Fex(xi))\mathcal{L}_{\mathrm{eic}} = \frac{1}{N_{F}} \sum_{i \in F} \delta\left(F_{im}(x_i), F_{ex}(x_i)\right) - \frac{1}{N_{R}} \sum_{i \in R} \delta\left(F_{im}(x_i), F_{ex}(x_i)\right)

    where:

    • RR and FF denote the subsets of real and fake face samples in a training batch, with sample counts NR=∣R∣N_R = |R| and NF=∣F∣N_F = |F|, respectively.
    • xi∈Rh×w×3x_i \in \mathbb{R}^{h \times w \times 3} is an aligned face image.
    • δ(u,v)=u∥u∥⋅v∥v∥\delta(u, v) = \frac{u}{\|u\|} \cdot \frac{v}{\|v\|} calculates the cosine similarity between feature vectors uu and vv.

    By minimizing this loss, real samples are pulled closer to their explicit identity features in the implicit feature space, while fake samples are pushed away from their explicit identity features to isolate explicit-identity-irrelevant target identity cues.

  3. Knowl 3 — Implicit Identity Exploration Loss with Progressive Margin for Known Target Identities

    model/method

    To train the implicit identity embedding network FimF_{im} on face images with known target identities, fake faces are labeled with the target face identity yiy_i (implicit identity), and real faces are labeled with their own identity yiy_i. For the set K\mathcal{K} of samples with known implicit identities, the margin-based classification loss Liie+\mathcal{L}_{iie}^{+} is formulated as:

    Liie+=−E(xi,yi)∼K[log⁡es(cos⁡(θyi)−m)es(cos⁡(θyi)−m)+∑j≠yiescos⁡θj]\mathcal{L}_{iie}^{+} = -\mathbb{E}_{(x_i, y_i) \sim \mathcal{K}} \left[ \log \frac{e^{s(\cos(\theta_{y_i}) - m)}}{e^{s(\cos(\theta_{y_i}) - m)} + \sum_{j \neq y_i} e^{s \cos \theta_j}} \right]

    where:

    • θj\theta_j is the angle between the normalized implicit feature Fim(xi)∥Fim(xi)∥\frac{F_{im}(x_i)}{\|F_{im}(x_i)\|} and the normalized weight proxy of the jj-th identity class on the hypersphere.
    • ss is the feature rescaling parameter.
    • mm is the margin hyperparameter.

    The margin is differentiated across real and fake samples:

    • For real samples, the margin is fixed to mreal=0.4m_{\mathrm{real}} = 0.4.
    • For fake samples, a dynamic progressive margin mfakem_{\mathrm{fake}} is computed based on real sample convergence:

    mfake=α⋅1Nr∑i∈Rminicos⁡(θyi)m_{\mathrm{fake}} = \alpha \cdot \frac{1}{N_r} \sum_{i \in R_{\mathrm{mini}}} \cos(\theta_{y_i})

    where RminiR_{\mathrm{mini}} denotes the set of real samples in the mini-batch, Nr=∣Rmini∣N_r = |R_{\mathrm{mini}}|, and α\alpha is a scaling factor set to 0.50.5.

    In early training, mfakem_{\mathrm{fake}} is near zero, allowing the network to first fit the implicit identities of real target faces before progressively pulling fake faces toward their target identities.

  4. Knowl 4 — Implicit Identity Exploration Loss for Unknown Target Identities

    model/method

    For fake video datasets where the target face identity is unannotated, inter-frame identity consistency within the same manipulated video is enforced using a dynamic feature lookup table.

    Let U\mathcal{U} be the set of fake samples with unknown target identities, where each sample xi∈Ux_i \in \mathcal{U} shares an unknown implicit identity index yi∗y_i^* with other frames from the same video. A lookup table V∈RD×QV \in \mathbb{R}^{D \times Q} stores normalized DD-dimensional implicit identity features for all QQ unknown video identities.

    The probability of classifying xix_i into identity yi∗y_i^* is computed using cosine similarity scaled by temperature τ=0.1\tau = 0.1:

    Liie−=−E(xi,yi∗)∼U[log⁡e(vyi∗TFim(xi)/τ)∑j=1Qe(vjTFim(xi)/τ)]\mathcal{L}_{iie}^{-} = -\mathbb{E}_{(x_i, y_i^*) \sim \mathcal{U}} \left[ \log \frac{e^{(v_{y_i^*}^T F_{im}(x_i) / \tau)}}{\sum_{j=1}^{Q} e^{(v_j^T F_{im}(x_i) / \tau)}} \right]

    During backward propagation, the corresponding table entry is updated with momentum:

    vyi∗←βvyi∗+(1−β)Fim(xi)v_{y_i^*} \leftarrow \beta v_{y_i^*} + (1 - \beta) F_{im}(x_i)

    where β∈[0,1]\beta \in [0, 1].

    The complete Implicit Identity Exploration (IIE) loss is given by:

    Liie=Liie++Liie−\mathcal{L}_{iie} = \mathcal{L}_{iie}^{+} + \mathcal{L}_{iie}^{-}

  5. Knowl 5 — Overall Loss Function for Implicit Identity Driven Deepfake Detection

    equation

    The overall training objective L\mathcal{L} of the Implicit Identity Driven (IID) detection framework combines classification supervision with explicit contrast and implicit exploration losses:

    L=Lbce+λ1Leic+λ2Liie\mathcal{L} = \mathcal{L}_{\mathrm{bce}} + \lambda_{1} \mathcal{L}_{\mathrm{eic}} + \lambda_{2} \mathcal{L}_{\mathrm{iie}}

    where:

    • Lbce\mathcal{L}_{\mathrm{bce}} is the standard binary cross-entropy loss computed on the output of a fully connected classification layer taking the difference between implicit and explicit identity representations.
    • Leic\mathcal{L}_{\mathrm{eic}} is the Explicit Identity Contrast loss.
    • Liie=Liie++Liie−\mathcal{L}_{\mathrm{iie}} = \mathcal{L}_{iie}^+ + \mathcal{L}_{iie}^- is the composite Implicit Identity Exploration loss.
    • λ1=0.05\lambda_{1} = 0.05 and λ2=0.1\lambda_{2} = 0.1 are balancing hyperparameters.

    For expression-swapping manipulation samples (e.g., Face2Face and NeuralTextures), where face identity is not swapped, only the Liie\mathcal{L}_{\mathrm{iie}} constraint is applied while the Leic\mathcal{L}_{\mathrm{eic}} constraint is omitted.

  6. Knowl 6 — Component Ablation of the Implicit Identity Driven Detection Framework

    data/table

    An ablation study on Celeb-DF and DFDC evaluates the individual and joint contributions of the Explicit Identity Contrast (Leic\mathcal{L}_{\mathrm{eic}}) and Implicit Identity Exploration (Liie\mathcal{L}_{\mathrm{iie}}) losses in terms of Accuracy (ACC, %) and Area Under the ROC Curve (AUC, %):

    Model Leic\mathcal{L}_{\mathrm{eic}} Liie\mathcal{L}_{\mathrm{iie}} Celeb-DF DFDC
    ACC (%) AUC (%) ACC (%) AUC (%)
    A (Baseline) 70.34 74.09 69.85 72.65
    B ✓ 77.76 82.24 76.39 78.80
    C ✓ 76.40 81.46 74.95 77.22
    D (Full IID) ✓ ✓ 79.16 83.80 79.37 81.23

    Key takeaways:

    • Incorporating Leic\mathcal{L}_{\mathrm{eic}} alone (Model B) improves baseline performance by 7.42% ACC / 8.15% AUC on Celeb-DF and 6.54% ACC / 6.15% AUC on DFDC, demonstrating the effectiveness of referencing explicit identity features.
    • Using Liie\mathcal{L}_{\mathrm{iie}} alone without Leic\mathcal{L}_{\mathrm{eic}} (Model C) results in inferior performance compared to Model B because Liie\mathcal{L}_{\mathrm{iie}} requires the explicit identity anchor to ground the real samples' implicit representations.
    • The combined framework (Model D) achieves the highest accuracy and AUC across both benchmarks.
  7. Knowl 7 — Cross-Dataset Generalization of Implicit Identity Driven Detection

    data/table

    When trained on high-quality FaceForensics++ (FF++ C23), the Implicit Identity Driven (IID) model is evaluated on unseen cross-dataset benchmarks (Celeb-DF, DFD, DFDC) using AUC (%) and Equal Error Rate (EER, %):

    Method FF++ (C23) Celeb-DF DFD DFDC
    AUC (%) EER (%) AUC (%) EER (%) AUC (%) EER (%) AUC (%) EER (%)
    Xception 99.09 3.77 65.27 38.77 87.86 21.04 69.90 35.41
    EN-b4 99.22 3.36 68.52 35.61 87.37 21.99 70.12 34.54
    Face X-ray 87.40 - 74.20 - 85.60 - 70.00 -
    MLDG 98.99 3.46 74.56 30.81 88.14 21.34 71.86 34.44
    F3-Net 98.10 3.58 71.21 34.03 86.10 26.17 72.88 33.38
    MAT(EN-b4) 99.27 3.35 76.65 32.83 87.58 21.73 67.34 38.31
    GFF 98.36 3.85 75.31 32.48 85.51 25.64 71.58 34.77
    LTW 99.17 3.32 77.14 29.34 88.56 20.57 74.58 33.81
    Local-relation 99.46 3.01 78.26 29.67 89.24 20.32 76.53 32.41
    DCL 99.30 3.26 82.30 26.53 91.66 16.63 76.71 31.97
    UIA-ViT 99.33 - 82.41 - 94.68 - 75.80 -
    IID (Ours) 99.32 2.99 83.80 24.85 93.92 14.01 81.23 26.80

    Under low-quality training on FF++ (C40), IID achieves 96.79% AUC on intra-testing and 82.04% AUC on Celeb-DF, outperforming SPSL (76.88%) and ITA (77.35%). On the challenging DFDC benchmark, IID reaches 81.23% AUC, surpassing DCL by 4.52%.

  8. Knowl 8 — Cross-Manipulation Generalization Performance

    data/table

    Cross-manipulation generalization is evaluated by training on one manipulation subset of FF++ (C23) or FaceShifter (FST) and evaluating on the other two unseen manipulation datasets across DeepFakes (DF), FaceSwap (FS), and FaceShifter (FST). Metric is AUC (%):

    Train Dataset Method DF (%) FS (%) FST (%) Mean AUC (%)
    DF EN-b4 99.97 46.24 51.26 65.82
    MAT 99.92 40.61 45.39 61.97
    GFF 99.87 47.21 51.93 66.34
    DCL 99.98 61.01 68.45 76.48
    IID (Ours) 99.51 63.83 73.49 78.94
    FS EN-b4 69.25 99.89 60.76 76.63
    MAT 64.13 99.67 57.37 73.72
    GFF 70.21 99.85 61.29 77.12
    DCL 74.80 99.90 64.86 79.85
    IID (Ours) 75.39 99.73 66.18 80.43
    FST EN-b4 61.11 56.19 99.52 72.27
    MAT 58.15 55.03 99.16 70.78
    GFF 61.48 56.17 99.41 72.35
    DCL 63.98 58.43 99.49 73.97
    IID (Ours) 65.42 59.50 99.50 74.81

    When trained on DF and tested on FST, IID achieves 73.49% AUC (a 5.04% gain over DCL at 68.45%). Across all three settings, IID yields the highest mean AUC on unseen manipulation types (78.94%, 80.43%, and 74.81%, respectively).

  9. Knowl 9 — Multi-Source Manipulation Generalization Performance

    data/table

    In multi-source manipulation evaluation, models are trained on three manipulation methods in FF++ and evaluated on the remaining unseen manipulation method under high-quality (C23) and low-quality (C40) compression using EfficientNet-b0:

    Method GID-DF (C23) GID-DF (C40) GID-F2F (C23) GID-F2F (C40)
    ACC (%) AUC (%) ACC (%) AUC (%) ACC (%) AUC (%) ACC (%) AUC (%)
    EfficientNet 82.40 91.11 67.60 75.30 63.32 80.10 61.41 67.40
    Focalloss 81.33 90.31 67.47 74.95 60.80 79.80 61.00 67.21
    ForensicTransfer 72.01 - 68.20 - 64.50 - 55.00 -
    Multi-task 70.30 - 66.76 - 58.74 - 56.50 -
    MLDG 84.21 91.82 67.15 73.12 63.46 77.10 58.12 61.70
    LTW 85.60 92.70 69.15 75.60 65.60 80.20 65.70 72.40
    DCL 87.70 94.90 75.90 83.82 68.40 82.93 67.85 75.07
    IID (Ours) 88.21 95.03 76.90 84.55 69.36 84.37 67.99 74.80

    where GID-DF indicates training on FaceSwap, Face2Face, and NeuralTextures and testing on DeepFakes; GID-F2F indicates training on DeepFakes, FaceSwap, and NeuralTextures and testing on Face2Face. IID achieves state-of-the-art performance across all benchmarks, reaching 88.21% ACC / 95.03% AUC on GID-DF (C23) and 76.90% ACC / 84.55% AUC on GID-DF (C40).

  10. Knowl 10 — Explicit-Implicit Identity Similarity Distribution Properties

    empirical result

    The explicit-implicit identity similarity (EIIS) is defined as the cosine similarity δ(Fim(x),Fex(x))\delta(F_{im}(x), F_{ex}(x)) between the extracted implicit embedding Fim(x)F_{im}(x) and the explicit embedding Fex(x)F_{ex}(x) of an input face xx.

    Empirical distributions on DeepFakes (FF++ C23) and FaceShifter show clear separation between authentic and manipulated faces:

    • For authentic faces, EIIS values cluster tightly in the range [0.5,1.0][0.5, 1.0].
    • For fake faces, EIIS values cluster in the range [−0.3,0.5][-0.3, 0.5].

    In identity verification tests on source-target-fake triplets:

    • Positive pairs (fake face paired with the true target face, sharing the same implicit identity) exhibit positive cosine similarities with a narrow distribution variance.
    • Negative pairs (fake face paired with the source face, having differing implicit identities) center around zero or negative cosine similarities with significantly wider distribution variance.

    This confirms that the implicit identity network reliably isolates and preserves target identity representations from manipulated face images.

Coverage note — None. All primary methodological formulations, loss designs, ablation studies, and empirical benchmark evaluations are fully covered.

References

  1. 1.Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. Mesonet: a compact facial video forgery detection network. In IEEE WIFS, pages 1–7, 2018. 1, 3, 7
  2. 2.Jianmin Bao, Dong Chen, Fang Wen, Houqiang Li, and Gang Hua. Towards open-set identity preserving face synthesis. In IEEE CVPR, pages 6713–6722, 2018. 2
  3. 3.Junyi Cao, Chao Ma, Taiping Yao, Shen Chen, Shouhong Ding, and Xiaokang Yang. End-to-end reconstruction-classification learning for face forgery detection. In IEEE CVPR, pages 4113–4122, 2022. 3, 4
  4. 4.Shenhao Cao, Qin Zou, Xiuqing Mao, Dengpan Ye, and Zhongyuan Wang. Metric learning for anti-compression facial forgery detection. In ACM MM, pages 1929–1937, 2021. 4
  5. 5.Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. Everybody dance now. In IEEE ICCV, pages 5933–5942, 2019. 1
  6. 6.Renwang Chen, Xuanhong Chen, Bingbing Ni, and Yanhao Ge. Simswap: An efficient framework for high fidelity face swapping. In ACM MM, pages 2003–2011, 2020. 2
  7. 7.Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. Local relation learning for face forgery detection. In AAAI, volume 35, pages 1081–1088, 2021. 1, 7
  8. 8.Franc¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In IEEE CVPR, pages 1251–1258, 2017. 3
  9. 9.Davide Cozzolino, Justus Thies, Andreas Rossler, Christian Riess, Matthias Nießner, and Luisa Verdoliva. Forensictransfer: Weakly-supervised domain adaptation for forgery detection. arXiv preprint arXiv:1812.02510, 2018. 8
  10. 10.Hao Dang, Feng Liu, Joel Stehouwer, Xiaoming Liu, and Anil K Jain. On the detection of digital face manipulation. In IEEE CVPR, pages 5781–5790, 2020. 1, 3
  11. 11.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition. In IEEE CVPR, pages 4685–4694, 2018. 2
  12. 12.Jiankang Deng, Jia Guo, Zhou Yuxiang, Jinke Yu, Irene Kotsia, and Stefanos Zafeiriou. Retinaface: Single-stage dense face localisation in the wild. arXiv: Computer Vision and Pattern Recognition, 2019. 5
  13. 13.Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397, 2020. 5
  14. 14.Joel Frank, Thorsten Eisenhofer, Lea Schonherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz. Leveraging frequency analysis for deep fake image recognition. In ICML, pages 3247–3258. PMLR, 2020. 3
  15. 15.Gege Gao, Huaibo Huang, Chaoyou Fu, Zhaoyang Li, and Ran He. Information bottleneck disentanglement for identity swapping. In IEEE CVPR, pages 3404–3413, 2021. 2
  16. 16.Yue Gao, Fangyun Wei, Jianmin Bao, Shuyang Gu, Dong Chen, Fang Wen, and Zhouhui Lian. High-fidelity and arbitrary face editing. In IEEE CVPR, pages 16115–16124, 2021. 1
  17. 17.Deepfakes github. Deepfakes. http://github.com/deepfakes/faceswap, 2017. 2, 5
  18. 18.FaceSwap github. Faceswap. https://github.com/MarekKowalski/FaceSwap, 2017. 5
  19. 19.Qiqi Gu, Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, and Ran Yi. Exploiting fine-grained face forgery clues via progressive enhancement learning. In AAAI, volume 36, pages 735–743, 2022. 1
  20. 20.Zhihao Gu, Yang Chen, Taiping Yao, Shouhong Ding, Jilin Li, Feiyue Huang, and Lizhuang Ma. Spatiotemporal inconsistency learning for deepfake video detection. In ACM MM, pages 3473–3481, 2021. 3, 7
  21. 21.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In IEEE CVPR, pages 770–778, 2016. 5
  22. 22.Yuge Huang, Yuhan Wang, Ying Tai, Xiaoming Liu, Pengcheng Shen, Shaoxin Li, Jilin Li, and Feiyue Huang. Curricularface: adaptive curriculum learning loss for deep face recognition. In IEEE CVPR, pages 5901–5910, 2020. 2
  23. 23.Iryna Korshunova, Wenzhe Shi, Joni Dambre, and Lucas Theis. Fast face-swap using convolutional neural networks. In IEEE ICCV, pages 3677–3685, 2017. 2
  24. 24.Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy Hospedales. Learning to generalize: Meta-learning for domain generalization. In AAAI, volume 32, 2018. 7, 8
  25. 25.Jiaming Li, Hongtao Xie, Jiahong Li, Zhongyuan Wang, and Yongdong Zhang. Frequency-aware discriminative feature learning supervised by single-center loss for face forgery detection. In IEEE CVPR, pages 6458–6467, 2021. 3
  26. 26.Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. Faceshifter: Towards high fidelity and occlusion aware face swapping. arXiv preprint arXiv:1912.13457, 2019. 2, 5
  27. 27.Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. Face x-ray for more general face forgery detection. In IEEE CVPR, pages 5001–5010, 2020. 1, 3, 6, 7
  28. 28.Yuezun Li and Siwei Lyu. Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656, 2018. 7
  29. 29.Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. Celeb-df: A large-scale challenging dataset for deepfake forensics. In IEEE CVPR, pages 3207–3216, 2020. 5
  30. 30.Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. Focal loss for dense object detection. In IEEE CVPR, pages 2980–2988, 2017. 8
  31. 31.Honggu Liu, Xiaodan Li, Wenbo Zhou, Yuefeng Chen, Yuan He, Hui Xue, Weiming Zhang, and Nenghai Yu. Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. In IEEE CVPR, pages 772–781, 2021. 7
  32. 32.Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. Generalizing face forgery detection with high-frequency features. In IEEE CVPR, pages 16317–16326, 2021. 7
  33. 33.Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, and Wael AbdAlmageed. Two-branch recurrent network for isolating deepfakes in videos. In ECCV, pages 667–684. Springer, 2020. 3, 7
  34. 34.R Natsume, T Yatagawa, and S Morishima. Rsgan: Face swapping and editing using face and hair representation in latent spaces. arxiv 2018. arXiv preprint arXiv:1804.03447. 2
  35. 35.Ryota Natsume, Tatsuya Yatagawa, and Shigeo Morishima. Fsnet: An identity-aware generative model for image-based face swapping. In ACCV, pages 117–132. Springer, 2018. 2
  36. 36.Huy H Nguyen, Fuming Fang, Junichi Yamagishi, and Isao Echizen. Multi-task learning for detecting and segmenting manipulated facial images and videos. In IEEE International Conference on Biometrics Theory, Applications and Systems, pages 1–8, 2019. 3, 7, 8
  37. 37.Huy H Nguyen, Junichi Yamagishi, and Isao Echizen. Capsule-forensics: Using capsule networks to detect forged images and videos. In IEEE ICASSP, pages 2307–2311, 2019. 1, 3
  38. 38.Google Research Nick Dufour and Jigsaw Andrew Gully. Deep fake detection dataset. https : / / ai . googleblog . com / 2019 / 09 / contributing - data-to-deepfake-detection.html, 2019. 5
  39. 39.Yuval Nirkin, Yosi Keller, and Tal Hassner. Fsgan: Subject agnostic face swapping and reenactment. In IEEE ICCV, pages 7184–7193, 2019. 2
  40. 40.Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. NeurIPS Workshop, 2017. 5
  41. 41.Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In ECCV, pages 86–103. Springer, 2020. 1, 3
  42. 42.Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. In IEEE ICCV, pages 1–11, 2019. 1, 3, 4, 5, 6, 7
  43. 43.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. 3
  44. 44.Ke Sun, Hong Liu, Taiping Yao, Xiaoshuai Sun, Shen Chen, Shouhong Ding, and Rongrong Ji. An information theoretic approach for attention-driven face forgery detection. In ECCV, pages 111–127. Springer, 2022. 7
  45. 45.Ke Sun, Hong Liu, Qixiang Ye, Yue Gao, Jianzhuang Liu, Ling Shao, and Rongrong Ji. Domain general face forgery detection by learning to weight. In AAAI, volume 35, pages 2638–2646, 2021. 6, 7, 8
  46. 46.Ke Sun, Taiping Yao, Shen Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. Dual contrastive learning for general face forgery detection. In AAAI, volume 36, pages 2316–2324, 2022. 3, 4, 6, 7, 8
  47. 47.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In ICML, pages 6105–6114. PMLR, 2019. 7, 8
  48. 48.Justus Thies, Michael Zollhofer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. ACM TOG, 38(4):1–12, 2019. 1, 5
  49. 49.Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. Face2face: Real-time face capture and reenactment of rgb videos. In IEEE CVPR, pages 2387–2395, 2016. 5
  50. 50.Chengrui Wang and Weihong Deng. Representative forgery mining for fake face detection. In IEEE CVPR, pages 14923–14932, 2021. 1, 3
  51. 51.Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu. Cosface: Large margin cosine loss for deep face recognition. In IEEE CVPR, pages 5265–5274, 2018. 1, 2, 4, 5, 7
  52. 52.Jun Wei, Shuhui Wang, and Qingming Huang. F3 net: fusion, feedback and focus for salient object detection. In AAAI, volume 34, pages 12321–12328, 2020. 6, 7
  53. 53.Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. Multi-attentional deepfake detection. In IEEE CVPR, pages 2185–2194, 2021. 1, 3, 7
  54. 54.Peng Zhou, Xintong Han, Vlad I Morariu, and Larry S Davis. Two-stream neural networks for tampered face detection. In IEEE CVPRW, pages 1831–1839, 2017. 3
  55. 55.Wanyi Zhuang, Qi Chu, Zhentao Tan, Qiankun Liu, Haojie Yuan, Changtao Miao, Zixiang Luo, and Nenghai Yu. Uia-vit: Unsupervised inconsistency-aware method based on vision transformer for face forgery detection. In ECCV. Springer, 2022. 3, 6, 7
  56. 56.Bojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma, and Yu-Gang Jiang. Wilddeepfake: A challenging real-world dataset for deepfake detection. In ACM MM, pages 2382–2390, 2020. 3

Citation

MLA
Huang, B., et al. “Implicit Identity Driven Deepfake Face Swapping Detection”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 4490–99, https://doi.org/10.1109/CVPR52729.2023.00436.
APA
Huang, B., Wang, Z., Yang, J., Ai, J., Zou, Q., Wang, Q., & Ye, D. (2023). Implicit Identity Driven Deepfake Face Swapping Detection. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4490–4499. https://doi.org/10.1109/CVPR52729.2023.00436
Chicago
Huang, B., Z. Wang, J. Yang, et al. 2023. “Implicit Identity Driven Deepfake Face Swapping Detection”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 4490–99. https://doi.org/10.1109/CVPR52729.2023.00436.
Harvard
Huang, B. et al. (2023) “Implicit Identity Driven Deepfake Face Swapping Detection”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 4490–4499. Available at: https://doi.org/10.1109/CVPR52729.2023.00436.
Vancouver
1. Huang B, Wang Z, Yang J, Ai J, Zou Q, Wang Q, Ye D (2023) Implicit Identity Driven Deepfake Face Swapping Detection. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 4490–4499

BibTeX

@inproceedings{Huang_2023, title={Implicit Identity Driven Deepfake Face Swapping Detection}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00436}, DOI={10.1109/cvpr52729.2023.00436}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Huang, Baojin and Wang, Zhongyuan and Yang, Jifan and Ai, Jiaxin and Zou, Qin and Wang, Qian and Ye, Dengpan}, year={2023}, month=June, pages={4490–4499} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE