Protecting Celebrities from DeepFake with Identity Consistency Transformer

Xiaoyi DongJianmin BaoDongdong ChenTing ZhangWeiming ZhangNenghai YuDong ChenFang WenBaining Guo

article2022CVPR170 citations

Proposes an identity-consistency vision transformer that detects deepfakes by exposing semantic identity mismatches between inner and outer face regions, achieving high detection generalization under real-world video degradations without relying on fragile low-level texture artifacts.

Listen

Rapid advancements in deepfake technology make it easy to generate highly convincing manipulated facial media. These realistic forgeries increasingly target public figures, corporate leaders, and politicians, posing severe risks to media trust, organizational security, and public information integrity. Existing detection methods largely rely on identifying subtle, low-level visual artifacts or texture anomalies. However, as synthesis quality improves and videos undergo real-world compression, resizing, or transmission noise, these low-level traces disappear, causing conventional detectors to fail.

To overcome these vulnerabilities, the article introduces the Identity Consistency Transformer, a high-level semantic detection framework. The objective of the article is to demonstrate that detecting identity discrepancies between the inner face (such as eyes, nose, and mouth) and the outer face (such as face shape and contour) provides a far more robust and generalizable defense against face-swapping manipulations than analyzing pixel artifacts.

The researchers implemented a unified vision transformer model that simultaneously learns distinct identity representations for the inner and outer regions of a suspect face. To train the system without needing fake images from specific deepfake generation tools, the authors used a large-scale public dataset of real faces and created artificial swaps via mask blending and color correction. A dedicated consistency loss function pulls inner and outer identity representations together when they belong to the same person and pushes them apart when they differ. Furthermore, the framework incorporates an authentic reference set of known public identities, creating a reference-assisted variant that cross-checks suspect faces against verified images.

Evaluations across multiple unseen standard benchmarks (including FaceForensics++, DeepFake Detection, DeeperForensics, and Celeb-DF) revealed three critical findings. First, the core framework achieved an average area under the curve score of 87.01% on unseen datasets, consistently outperforming conventional baseline models, which often scored around 50% to 79%. Second, when enhanced with verified reference identities, the system's average performance jumped to 96.34% across benchmark datasets and achieved 100% accuracy on real-world, carefully crafted internet deepfake videos. Third, while traditional artifact-based detectors degraded rapidly under common image disruptions like compression, blur, pixelation, and contrast shifts, the identity-based transformer maintained stable and superior detection accuracy.

These findings indicate that shifting from fragile low-level texture analysis to high-level semantic identity verification creates a resilient, future-proof defense against modern manipulation techniques. Organizations protecting high-profile leaders or managing digital media platforms can achieve high detection reliability without constantly retraining models on new generative tools. The authors recommend adopting identity consistency frameworks to safeguard public figures and incorporating automated authentic reference databases where media archives exist. However, decision-makers should note that this method specifically targets face-swapping manipulations; it is not designed to detect face reenactments where identity remains unchanged, and its performance can degrade under extreme Gaussian noise or when authentic reference databases are too sparse.

  • Paper: Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics, Yuezun Li et al. (2019). This paper establishes the Celeb-DF benchmark and high-quality deepfake evaluation protocol utilized directly in the target study to test cross-dataset forgery detection.
  • Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). This foundational work introduces the FaceForensics++ benchmark and standard manipulation baselines against which the source paper measures its semantic identity consistency detector.
  • Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). This study analyzes how conventional binary classifiers rely on low-level synthetic artifacts across generative models, setting up the exact generalization failure modes the source paper solves via high-level identity modeling.
  • Paper: Deep Face Recognition: A Survey, Mei Wang et al. (2018). This survey provides essential background on deep metric learning, loss functions, and facial feature representations that underpin the source paper's inner-outer identity consistency formulation.
  • Paper: MesoNet: a Compact Facial Video Forgery Detection Network, Darius Afchar et al. (2018). This paper introduces compact artifact-based neural networks for facial video tampering detection, representing the traditional low-level detection paradigm that the source study seeks to surpass.
  • Paper: Face2Face: Real-Time Face Capture and Reenactment of RGB Videos, Justus Thies et al. (2016). This seminal paper introduces RGB-based facial capture and reenactment techniques, providing critical domain context on facial manipulation paradigms evaluated in deepfake forensics.
Cover for Protecting Celebrities from DeepFake with Identity Consistency Transformer

Abstract

In this work we propose Identity Consistency Transformer, a novel face forgery detection method that focuses on high-level semantics, specifically identity information, and detecting a suspect face by finding identity inconsistency in inner and outer face regions. The Identity Consistency Transformer incorporates a consistency loss for identity consistency determination. We show that Identity Consistency Transformer exhibits superior generalization ability not only across different datasets but also across various types of image degradation forms found in real-world applications including deepfake videos. The Identity Consistency Transformer can be easily enhanced with additional identity information when such information is available, and for this reason it is especially well-suited for detecting face forgeries involving celebrities.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Identity Consistency Transformer
  • 3.1. Identity Extraction Model
  • 3.2. Identity Consistency Detection
  • 3.3. Benefits and Limitations and Impacts
  • 4. Experiment
  • 4.1. Comparison with State-of-the-art Methods
  • 4.2. Analysis of ICT
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Identity Consistency Transformer Architecture

    model/method

    The Identity Consistency Transformer (ICT) is a deepfake detection model designed to extract and compare semantic identity representations from both the inner face and outer face regions simultaneously using a single Vision Transformer backbone.

    Given an input RGB face image I∈RH×W×CI \in \mathbb{R}^{H \times W \times C}, the image is partitioned into non-overlapping patches of spatial size P×PP \times P, producing N=HW/P2N = HW / P^2 patch tokens. Each flattened patch is linearly projected into a token embedding of dimension dd. Two learnable identity embedding tokens—an inner token and an outer token—are prepended to the sequence of patch embeddings, along with learnable 1D position embeddings.

    The full token sequence is processed through a Transformer encoder comprising 12 stacked blocks. Each block consists of a multi-head self-attention module (12 attention heads) and a multi-layer perceptron (MLP) module, preceded by layer normalization. The global self-attention mechanism allows the inner and outer tokens to selectively route and attend to different facial regions across the entire image.

    The final outputs of the inner and outer tokens from the last Transformer block yield two identity feature vectors, fin∈Rdf^{\text{in}} \in \mathbb{R}^d and fout∈Rdf^{\text{out}} \in \mathbb{R}^d. Both vectors are passed through a shared linear projection classification head followed by batch normalization to output identity classification logits. The default model uses input resolution 112×112112 \times 112, patch resolution 14×1414 \times 14 (N=64N = 64 patches), embedding dimension d=384d = 384, and contains 21.45 million parameters.

  2. Knowl 2 — Identity Classification and Consistency Loss Formulations

    equation

    The Identity Consistency Transformer is trained using a multi-task loss combining an ArcFace cosine-margin classification loss Lident\mathcal{L}_{\text{ident}} and an identity consistency loss Lconsist\mathcal{L}_{\text{consist}}.

    For a training batch containing face-swapped images IijI_{ij} (where the inner face originates from subject IiI_i with identity label yiy_i and the outer face originates from subject IjI_j with identity label yjy_j), the identity classification loss applied to both inner and outer feature vectors f∈Rdf \in \mathbb{R}^d is:

    Lident=−log⁡es⋅cos⁡(θk,y+m)es⋅cos⁡(θk,y+m)+∑c≠yes⋅cos⁡(θk,c)\mathcal{L}_{\text{ident}} = -\log \frac{e^{s \cdot \cos(\theta_{k, y} + m)}}{e^{s \cdot \cos(\theta_{k, y} + m)} + \sum_{c \neq y} e^{s \cdot \cos(\theta_{k, c})}}

    where θk,c\theta_{k, c} is the angle between the shared projection head weight vector Wc∈RdW_c \in \mathbb{R}^d for class cc and feature fkf_k, ss is a hypersphere radius scale factor (set to s=64s = 64), and mm is an additive angular margin parameter (set to m=0m = 0 at the first epoch and linearly increased to m=0.3m = 0.3 after 10 epochs).

    To explicitly constrain corresponding identity vectors from paired swapped samples (Iij,Iji)(I_{ij}, I_{ji}) to match in embedding space, the identity consistency loss is defined as:

    Lconsist=I(pijin=pjiout=yi)∥fijin−fjiout∥22+I(pijout=pjiin=yj)∥fijout−fjiin∥22\mathcal{L}_{\text{consist}} = \mathbb{I}\left(p^{\text{in}}_{ij} = p^{\text{out}}_{ji} = y_i\right) \left\|f^{\text{in}}_{ij} - f^{\text{out}}_{ji}\right\|_2^2 + \mathbb{I}\left(p^{\text{out}}_{ij} = p^{\text{in}}_{ji} = y_j\right) \left\|f^{\text{out}}_{ij} - f^{\text{in}}_{ji}\right\|_2^2

    where fijin,fijoutf^{\text{in}}_{ij}, f^{\text{out}}_{ij} denote the extracted inner and outer identity vectors for image IijI_{ij}, pijin,pijoutp^{\text{in}}_{ij}, p^{\text{out}}_{ij} are the predicted class labels, and I(⋅)\mathbb{I}(\cdot) is the indicator function. The consistency loss activates exclusively when the classifier predictions for both tokens are correct, preventing erroneous gradient propagation.

    The combined objective function is:

    L=Lident+ηLconsist\mathcal{L} = \mathcal{L}_{\text{ident}} + \eta \mathcal{L}_{\text{consist}}

    where η\eta is a weighting hyperparameter initialized to η=4.0\eta = 4.0 at the first epoch and increased by 0.50.5 after each subsequent training epoch.

  3. Knowl 3 — Reference-Assisted Identity Consistency Detection (ICT-Ref)

    model/method

    When authentic reference images of protected individuals (such as public figures or celebrities) are accessible, the Identity Consistency Transformer can be extended into a reference-assisted verification mode denoted ICT-Ref.

    An authentic reference set F={(fnin,fnout)=T(In)}n=1N\mathcal{F} = \{(f_n^{\text{in}}, f_n^{\text{out}}) = T(I_n)\}_{n=1}^N is constructed by applying the pre-trained identity extraction model TT to NN genuine face images. Given a suspect image II with extracted features (fin,fout)=T(I)(f^{\text{in}}, f^{\text{out}}) = T(I):

    1. Inner-guided nearest neighbor query: kin=arg⁡min⁡n∈{1,…,N}d(fin,fnin),αin=d(fin,fkinin)k^{\text{in}} = \arg\min_{n \in \{1, \dots, N\}} d(f^{\text{in}}, f_n^{\text{in}}), \quad \alpha_{\text{in}} = d(f^{\text{in}}, f_{k^{\text{in}}}^{\text{in}}) Dout=d(fout,fkinout)D_{\text{out}} = d(f^{\text{out}}, f_{k^{\text{in}}}^{\text{out}})

    2. Outer-guided nearest neighbor query: kout=arg⁡min⁡n∈{1,…,N}d(fout,fnout),αout=d(fout,fkoutout)k^{\text{out}} = \arg\min_{n \in \{1, \dots, N\}} d(f^{\text{out}}, f_n^{\text{out}}), \quad \alpha_{\text{out}} = d(f^{\text{out}}, f_{k^{\text{out}}}^{\text{out}}) Din=d(fin,fkoutin)D_{\text{in}} = d(f^{\text{in}}, f_{k^{\text{out}}}^{\text{in}})

    3. Combined consistency metric: Dref=λDin_out+ω(αin)Dout+ω(αout)DinD_{\text{ref}} = \lambda D_{\text{in\_out}} + \omega(\alpha_{\text{in}}) D_{\text{out}} + \omega(\alpha_{\text{out}}) D_{\text{in}}

    where d(⋅,⋅)d(\cdot, \cdot) is the Euclidean distance, Din_out=d(fin,fout)D_{\text{in\_out}} = d(f^{\text{in}}, f^{\text{out}}), and the adaptive weighting function is given by: ω(x)=11+exp⁡(x−θτ)\omega(x) = \frac{1}{1 + \exp\left(\frac{x - \theta}{\tau}\right)} with identity threshold θ=0.5\theta = 0.5, scale factor τ=0.5\tau = 0.5, and dynamic weight λ=2.5−ω(αout)−ω(αin)\lambda = 2.5 - \omega(\alpha_{\text{out}}) - \omega(\alpha_{\text{in}}). Higher DrefD_{\text{ref}} scores indicate higher confidence of face forgery.

  4. Knowl 4 — Reference-Free Identity Consistency Metric

    equation

    In the standalone (reference-free) mode of the Identity Consistency Transformer, face forgery detection is evaluated directly on a single suspect image II without requiring external reference images.

    Given the identity extraction Transformer T(I)=(fin,fout)T(I) = (f^{\text{in}}, f^{\text{out}}), where fin∈Rdf^{\text{in}} \in \mathbb{R}^d is the inner face identity vector and fout∈Rdf^{\text{out}} \in \mathbb{R}^d is the outer face identity vector, the reference-free forgery score Din_outD_{\text{in\_out}} is computed as the Euclidean distance between the two vectors:

    Din_out=d(fin,fout)=∥fin−fout∥2D_{\text{in\_out}} = d(f^{\text{in}}, f^{\text{out}}) = \left\|f^{\text{in}} - f^{\text{out}}\right\|_2

    Because the shared classification head and consistency loss enforce identical representation vectors for inner and outer regions originating from the same individual, a low Din_outD_{\text{in\_out}} value indicates high identity consistency (authentic face), whereas a high Din_outD_{\text{in\_out}} value indicates identity discrepancy across regions (manipulated/swapped face).

  5. Knowl 5 — Synthetic Training Data Generation via Face Landmark Swapping

    algorithm

    The Identity Consistency Transformer does not require fake images generated by existing deepfake manipulation algorithms during training. Instead, synthetic training pairs are constructed exclusively from a dataset of real face images with identity annotations (MS-Celeb-1M).

    Input: Real face dataset D={(Im,ym)}m=1M\mathcal{D} = \{(I_m, y_m)\}_{m=1}^M with identity labels ymy_m
    Output: Synthetic training pair (Iij,yi,yj)(I_{ij}, y_i, y_j) and (Iji,yj,yi)(I_{ji}, y_j, y_i)
    Sample image IiI_i with identity label yiy_i
    Find candidate image IjI_j with identity label yj≠yiy_j \neq y_i having matching head pose via facial landmark alignment
    Extract facial landmark coordinates for IiI_i and IjI_j
    Compute the convex hull of facial landmarks on IjI_j
    Apply random geometric deformation to the convex hull to generate inner face mask MM
    Adjust the color distribution of the inner region of IiI_i by matching the local channel-wise RGB means to the inner region of IjI_j
    Blend color-corrected inner region of IiI_i into IjI_j using mask MM to produce swapped face IijI_{ij}
    Repeat swapping procedure in reverse using inner region of IjI_j into IiI_i to produce counterpart IjiI_{ji}
    return (Iij,yi,yj),(Iji,yj,yi)(I_{ij}, y_i, y_j), (I_{ji}, y_j, y_i)
  6. Knowl 6 — Cross-Dataset Generalization on Benchmark Datasets

    data/table

    The table below compares frame-level Area Under the Receiver Operating Characteristic curve (AUC in %) of Identity Consistency Transformer (ICT) and reference-assisted ICT (ICT-Ref) against competing methods evaluated on unseen test datasets: DeepFake Detection (DFD), FaceForensics++ (FF++), DeeperForensics (Deeper), Celeb-DeepFake v1 (CD1), and Celeb-DeepFake v2 (CD2). The 'Avg' column denotes the average open-set AUC.

    Method DFD FF++ Deeper CD1 CD2 Avg
    Multi-task 65.21 72.23 65.32 72.28 61.06 65.96
    MesoInc4 59.06 63.41 51.41 42.26 53.60 53.95
    Capsule 69.70 96.50 68.44 69.98 63.65 67.94
    Xcep-c0 89.05 99.26 57.76 48.08 50.37 61.32
    Xcep-c23 95.60 98.54 69.85 74.97 77.82 79.56
    FWA 80.59 74.82 45.46 72.88 64.87 67.72
    DSP-FWA 90.99 81.90 60.00 78.51 81.41 78.56
    CNNDetect 60.12 71.08 57.16 56.12 57.17 60.33
    Patch-Foren 49.91 73.75 55.35 59.66 57.16 55.52
    FFD 76.61 92.32 46.64 74.15 77.80 73.50
    Face X-ray 94.14 98.44 72.35 74.76 75.39 79.16
    Two-Branch - - - - 73.41 -
    PCL+I2G - - - - 81.80 -
    Nirkin et al. - 99.70 - - 66.00 -
    ICT (Ours) 84.13 90.22 93.57 81.43 85.71 87.01
    ICT-Ref (Ours) 93.17 98.56 99.25 96.41 94.43 96.34

    While artifact-based baselines degrade on newer, high-fidelity datasets exhibiting fewer blending boundaries (e.g., Deeper and Celeb-DF v2), ICT achieves 87.01% open-set average AUC, outperforming all baseline models. ICT-Ref achieves state-of-the-art performance with 96.34% average AUC across all benchmark datasets.

  7. Knowl 7 — Robustness Across Image and Video Degradation Modalities

    empirical result

    When subjected to eight realistic image and video degradation types across five severity levels—saturation shifts, contrast variations, block-wise distortion, Gaussian noise, Gaussian blur, pixelation, JPEG compression (quality factors 90, 70, 50, 30, 20), and video compression codecs—the Identity Consistency Transformer (ICT) and reference-assisted ICT (ICT-Ref) demonstrate significantly higher resilience than low-level artifact detectors (Xception, DSP-FWA, Face X-ray, and FFD).

    Low-level artifact detectors suffer severe performance drops (AUC falling from ~80–90% down to near 50–60%) as degradation severity increases, because image operations corrupt high-frequency blending boundaries and pixel-level generation artifacts. In contrast, ICT and ICT-Ref maintain steady frame-level AUC curves near or above 85–95% across increasing severities for seven of the eight degradation types. Only extreme severities of additive Gaussian noise cause noticeable degradation in ICT performance, demonstrating that semantic identity representations are substantially more robust to visual corruption than low-level texture cues.

  8. Knowl 8 — DeepFake Detection Performance on Real-World Celebrity Videos

    data/table

    The table below presents the detection performance (frame-level AUC in %) on highly photorealistic, real-world deepfake celebrity videos collected from the YouTube channel 'Ctrl Shift Face'.

    Method Video1 Video2 Video3 Avg
    Multi-task 62.50 26.45 78.50 55.81
    DSP-FWA 23.08 33.14 51.81 36.01
    Xception-c23 44.56 83.85 99.31 75.90
    FFD 56.04 71.16 79.07 68.75
    Face X-ray 66.49 94.03 87.10 82.54
    ICT (Ours) 82.27 97.78 98.05 94.36
    ICT-Ref (Ours) 100.0 100.0 100.0 100.0

    Existing detectors struggle on crafted real-world videos (e.g., DSP-FWA averages 36.01% AUC and Multi-task averages 55.81% AUC). In contrast, reference-free ICT achieves 94.36% average AUC, while reference-assisted ICT-Ref achieves 100.0% AUC across all evaluated videos.

  9. Knowl 9 — Comparison of ViT and Convolutional Backbones for Dual Identity Learning

    data/table

    The table below evaluates architectural choices for extracting dual identity vectors (inner and outer) from a single face image, comparing Vision Transformer against ResNet-50 configurations evaluated by frame-level AUC (%) on DeeperForensics (Deeper) and Celeb-DeepFake v2 (CD2).

    Model #Params Deeper CD2
    Res50 43.79M NaN NaN
    Res50-split-image 43.79M 79.42 73.71
    Res50-split-model 2 ×\times 43.79M 91.48 84.66
    ICT 21.45M 93.57 85.71

    A standard single ResNet-50 with shared convolutional feature maps and dual classification heads fails to converge (loss evaluates to NaN) because the shared spatial convolutions pull the inner and outer representations together too strongly when applied to conflicting identities. Splitting input images prior to feeding them into Res50 (Res50-split-image) yields inferior accuracy (73.71% on CD2). Training two distinct ResNet-50 networks for inner and outer faces (Res50-split-model) achieves 84.66% on CD2 at the expense of 4×4\times the parameter count of ICT (87.58M vs. 21.45M). The single-backbone Vision Transformer in ICT leverages global multi-head self-attention to independently route inner and outer tokens across the image, converging stably and achieving the highest AUC.

  10. Knowl 10 — Ablation on Training Components and Metric Distances

    data/table

    The table below presents ablation experiments evaluating the effect of the identity consistency loss, synthetic mask deformation, and color correction on frame-level AUC (%) across four unseen test datasets.

    Variant DFD FF++ Deeper CD2
    w/o consistency loss 60.82 62.24 52.90 51.48
    w/o mask deformation 83.59 83.47 79.06 83.44
    w/o color correction 80.43 90.11 92.15 82.86
    Full ICT 84.13 90.22 93.57 85.71

    Removing the consistency loss leads to a performance collapse of 23% to 40% across benchmarks (e.g., dropping to 51.48% on CD2), verifying that consistency loss is critical for forcing identical representations for same-identity regions and separation for swapped regions.

    Additionally, decomposing the individual distance terms in the reference-assisted framework on Celeb-DeepFake v2 shows:

    • Standalone ICT (Din_outD_{\text{in\_out}}): 85.71% AUC
    • Querying reference with outer token only (DinD_{\text{in}}): 92.52% AUC
    • Querying reference with inner token only (DoutD_{\text{out}}): 87.45% AUC
    • Combined reference metric (DrefD_{\text{ref}}): 94.43% AUC

    Using DinD_{\text{in}} outperforms DoutD_{\text{out}} by approximately 5.0% because the outer face is typically unaltered in face swapping, allowing the nearest neighbor search with foutf^{\text{out}} to retrieve the true reference identity of the subject more accurately.

  11. Knowl 11 — Scope Limitation to Identity-Inconsistent Manipulations

    limitation

    The Identity Consistency Transformer (ICT) relies fundamentally on detecting identity discrepancy between the inner facial features and the outer facial context. Consequently, the method is designed specifically for face-swapping manipulations. It is not designed to detect face reenactment, expression manipulation, or facial attribute editing where the inner face identity remains identical to the original subject.

Coverage note — None was omitted; all key contributed models, loss formulations, training algorithms, experimental benchmarks, degradation evaluations, architectural ablations, and stated limitations are fully represented.

References

  1. 1.https://github.com/iperov/DeepFaceLab. 1
  2. 2.https://github.com/dfaker/df. 1
  3. 3.https://www.malavida.com/en/soft/fakeapp/. 1
  4. 4.https://github.com/deepfakes/faceswap. 1, 2
  5. 5.https://github.com/shaoanlu/faceswap-GAN. 1
  6. 6.https://ai.googleblog.com/2019/09/contributing-data-to-deepfake-detection.html. 5
  7. 7.Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. Mesonet: a compact facial video forgery detection network, 2018. 1, 2, 6
  8. 8.Sandipan Banerjee, John S Bernhard, Walter J Scheirer, Kevin W Bowyer, and Patrick J Flynn. Srefi: Synthesis of realistic example face images. In 2017 IEEE International Joint Conference on Biometrics (IJCB), pages 37–45. IEEE, 2017. 2
  9. 9.Jianmin Bao, Dong Chen, Fang Wen, Houqiang Li, and Gang Hua. Towards open-set identity preserving face synthesis, 2018. 1, 2
  10. 10.Jawadul H Bappy, Cody Simons, Lakshmanan Nataraj, BS Manjunath, and Amit K Roy-Chowdhury. Hybrid lstm and encoder–decoder architecture for detection of image forgeries. IEEE Transactions on Image Processing, 28(7):3286–3300, 2019. 1, 2
  11. 11.Dmitri Bitouk, Neeraj Kumar, Samreen Dhillon, Peter Belhumeur, and Shree K Nayar. Face swapping: automatically replacing faces in photographs. In ACM Transactions on Graphics (TOG), volume 27, page 39. ACM, 2008. 2
  12. 12.Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020. 3
  13. 13.Chen Cao, Yanlin Weng, Shun Zhou, Yiying Tong, and Kun Zhou. Facewarehouse: A 3d facial expression database for visual computing. IEEE Transactions on Visualization and Computer Graphics, 20(3):413–425, 2013. 2
  14. 14.Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. What makes fake images detectable? understanding properties that generalize, 2020. 6
  15. 15.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Hee-woo Jun, David Luan, and Ilya Sutskever. Generative pre-training from pixels. In International Conference on Machine Learning, pages 1691–1703. PMLR, 2020. 3
  16. 16.Yinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu, Xiaoyi Dong, Lu Yuan, and Zicheng Liu. Mobile-former: Bridging mobilenet and transformer. arXiv preprint arXiv:2108.05895, 2021. 3
  17. 17.Yi-Ting Cheng, Virginia Tzeng, Yu Liang, Chuan-Chang Wang, Bing-Yu Chen, Yung-Yu Chuang, and Ming Ouhyoung. 3d-model-based face replacement in video. In SIGGRAPH’09: Posters, page 29. ACM, 2009. 2
  18. 18.Joni Adamson Clarke. Bringing the past to lif, 1996. 2
  19. 19.Davide Cozzolino, Andreas Rössler, Justus Thies, Matthias Nießner, and Luisa Verdoliva. Id-reveal: Identity-aware deepfake video detection. arXiv preprint arXiv:2012.02512, 2020. 3
  20. 20.Kevin Dale, Kalyan Sunkavalli, Micah K Johnson, Daniel Vlasic, Wojciech Matusik, and Hanspeter Pfister. Video face replacement. In ACM Transactions on Graphics (TOG), volume 30, page 130. ACM, 2011. 2
  21. 21.Hao Dang, Feng Liu, Joel Stehouwer, Xiaoming Liu, and Anil K Jain. On the detection of digital face manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5781–5790, 2020. 6
  22. 22.Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou. Arcface: Additive angular margin loss for deep face recognition, 2019. 2, 4
  23. 23.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 3, 4
  24. 24.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. Cswin transformer: A general vision transformer backbone with cross-shaped windows. arXiv preprint arXiv:2107.00652, 2021. 3
  25. 25.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. 3, 4
  26. 26.Zhenglin Geng, Chen Cao, and Sergey Tulyakov. 3d guided fine-grained face manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9821–9830, 2019. 2
  27. 27.Yandong Guo, Lei Zhang, Yuxiao Hu, Xiaodong He, and Jianfeng Gao. Ms-celeb-1m: A dataset and benchmark for large-scale face recognition. In European conference on computer vision, pages 87–102. Springer, 2016. 4, 5
  28. 28.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. 8
  29. 29.Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2889–2898, 2020. 5, 6
  30. 30.Iryna Korshunova, Wenzhe Shi, Joni Dambre, and Lucas Theis. Fast face-swap using convolutional neural networks. In Proceedings of the IEEE International Conference on Computer Vision, pages 3677–3685, 2017. 2
  31. 31.Trung-Nghia Le, Huy H. Nguyen, Junichi Yamagishi, and Isao Echizen. Openforensics: Large-scale challenging dataset for multi-face forgery detection and segmentation in-the-wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10117–10127, October 2021. 2
  32. 32.Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. Faceshifter: Towards high fidelity and occlusion aware face swapping. arXiv preprint arXiv:1912.13457, 2019. 1, 2
  33. 33.Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. Face x-ray for more general face forgery detection, 2020. 1, 2, 4, 5, 6
  34. 34.Yuezun Li, Ming-Ching Chang, and Siwei Lyu. In ictu oculi: Exposing ai generated fake face videos by detecting eye blinking, 2018. 2
  35. 35.Yuezun Li and Siwei Lyu. Exposing deepfake videos by detecting face warping artifacts, 2019. 1, 2, 6
  36. 36.Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. Celeb-df: A large-scale challenging dataset for deepfake forensics, 2020. 5, 7
  37. 37.Yuan Lin, Shengjin Wang, Qian Lin, and Feng Tang. Face swapping under large pose variations: A 3d model based approach. In 2012 IEEE International Conference on Multimedia and Expo, pages 333–338. IEEE, 2012. 2
  38. 38.Yaqi Liu, Qingxiao Guan, Xianfeng Zhao, and Yun Cao. Image forgery localization based on multi-scale convolutional neural networks. In Proceedings of the 6th ACM Workshop on Information Hiding and Multimedia Security, pages 85–90, 2018. 1, 2
  39. 39.Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, and Wael AbdAlmageed. Two-branch recurrent network for isolating deepfakes in videos. In European Conference on Computer Vision, pages 667–684. Springer, 2020. 6
  40. 40.Iacopo Masi, Anh Tuan Tran, Tal Hassner, Jatuporn Toy Leksut, and Gerard Medioni. Do we really need to collect millions of faces for effective face recognition? In European conference on computer vision, pages 579–596. Springer, 2016. 2
  41. 41.Iacopo Masi, Anh Tuan Tran, Tal Hassner, Gozde Sahin, and Gerard Medioni. Face-specific data augmentation for unconstrained face recognition. International Journal of Computer Vision, 127(6):642–667, 2019. 2
  42. 42.Falko Matern, Christian Riess, and Marc Stamminger. Exploiting visual artifacts to expose deepfakes and face manipulations. In 2019 IEEE Winter Applications of Computer Vision Workshops (WACVW), pages 83–92. IEEE, 2019. 1, 2
  43. 43.Ryota Natsume, Tatsuya Yatagawa, and Shigeo Morishima. Rsgan: face swapping and editing using face and hair representation in latent spaces. arXiv preprint arXiv:1804.03447, 2018. 2
  44. 44.Huy H. Nguyen, Fuming Fang, Junichi Yamagishi, and Isao Echizen. Multi-task learning for detecting and segmenting manipulated facial images and videos, 2019. 1, 2, 6
  45. 45.Huy H. Nguyen, Junichi Yamagishi, and Isao Echizen. Use of a capsule network to detect fake images and videos, 2019. 1, 2, 6
  46. 46.Yuval Nirkin, Yosi Keller, and Tal Hassner. Fsgan: Subject agnostic face swapping and reenactment, 2019. 1, 2
  47. 47.Yuval Nirkin, Iacopo Masi, Anh Tran Tuan, Tal Hassner, and Gerard Medioni. On face segmentation, face swapping, and face perception. In 2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018), pages 98–105. IEEE, 2018. 1, 2
  48. 48.Yuval Nirkin, Lior Wolf, Yosi Keller, and Tal Hassner. Deepfake detection based on the discrepancy between the face and its context. arXiv preprint arXiv:2008.12262, 2020. 3, 6
  49. 49.Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues, 2020. 1, 2
  50. 50.Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018. 3
  51. 51.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. 3
  52. 52.Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images, 2019. 1, 2, 5, 6
  53. 53.Branko Samarzija and Slobodan Ribaric. An approach to the de-identification of faces in different poses. In 2014 37th International Convention on Information and Communication Technology, Electronics and Microelectronics (MIPRO), pages 1246–1251. IEEE, 2014. 2
  54. 54.Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics (ToG), 36(4):1–13, 2017. 2
  55. 55.Justus Thies, Michael Zollhöfer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. ACM Transactions on Graphics (TOG), 38(4):1–12, 2019. 5
  56. 56.Justus Thies, Michael Zollhöfer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2387–2395, 2016. 2, 5
  57. 57.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017. 3, 4
  58. 58.Ziyu Wan, Jingbo Zhang, Dongdong Chen, and Jing Liao. High-fidelity pluralistic image completion with transformers. In ICCV, 2021. 3
  59. 59.Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu. Cosface: Large margin cosine loss for deep face recognition, 2018. 2
  60. 60.Hong-Xia Wang, Chunhong Pan, Haifeng Gong, and Huai-Yu Wu. Facial image composition based on active appearance model. In 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 893–896. IEEE, 2008. 2
  61. 61.Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. Bevt: Bert pretraining of video transformers. arXiv preprint arXiv:2112.01529, 2021. 3
  62. 62.Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. Cnn-generated images are surprisingly easy to spot... for now. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, volume 7, 2020. 6
  63. 63.Xin Yang, Yuezun Li, and Siwei Lyu. Exposing deep fakes using inconsistent head poses, 2018. 2
  64. 64.Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. Multi-attentional deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. 2
  65. 65.Tianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding, Yuanjun Xiong, and Wei Xia. Learning to recognize patch-wise consistency for deepfake detection. arXiv preprint arXiv:2012.09311, 2020. 6
  66. 66.Yinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng, and Fang Wen. Exploring temporal coherence for more general video face forgery detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15044–15054, October 2021. 2
  67. 67.Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. Learning rich features for image manipulation detection, 2018. 1, 2
  68. 68.Peng Zhou, Xintong Han, Vlad I. Morariu, and Larry S. Davis. Two-stream neural networks for tampered face detection, 2018. 1, 2

Citation

MLA
Dong, X., et al. “Protecting Celebrities from DeepFake with Identity Consistency Transformer”. arXiv, 2022, http://arxiv.org/abs/2203.01318v3.
APA
Dong, X., Bao, J., Chen, D., Zhang, T., Zhang, W., Yu, N., Chen, D., Wen, F., & Guo, B. (2022). Protecting Celebrities from DeepFake with Identity Consistency Transformer. arXiv. http://arxiv.org/abs/2203.01318v3
Chicago
Dong, X., J. Bao, D. Chen, et al. 2022. “Protecting Celebrities from DeepFake with Identity Consistency Transformer”. arXiv. http://arxiv.org/abs/2203.01318v3.
Harvard
Dong, X. et al. (2022) “Protecting Celebrities from DeepFake with Identity Consistency Transformer”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.01318v3.
Vancouver
1. Dong X, Bao J, Chen D, Zhang T, Zhang W, Yu N, Chen D, Wen F, Guo B (2022) Protecting Celebrities from DeepFake with Identity Consistency Transformer. arXiv

BibTeX

@article{dong2022protecting,
  title = {Protecting Celebrities from DeepFake with Identity Consistency Transformer},
  author = {Dong, Xiaoyi and Bao, Jianmin and Chen, Dongdong and Zhang, Ting and Zhang, Weiming and Yu, Nenghai and Chen, Dong and Wen, Fang and Guo, Baining},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.01318v3},
  eprint = {2203.01318}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE