Delving into Sequential Patches for Deepfake Detection

Jiazhi GuanHang ZhouZhibin HongErrui DingJingdong WangChengbin QuanYoujian Zhao

article2022NeurIPS84 citations

Proposes a local and temporal-aware transformer framework that captures fine-grained temporal inconsistencies across independent spatial patch sequences to improve deepfake detection generalization across unseen generation methods and video degradations.

Listen

Rapid advancements in artificial facial manipulation have made deepfake videos increasingly realistic and visually untraceable, posing serious threats to information security, organizational reputation, and individual privacy. Existing detection systems struggle with two operational flaws: they fail to generalize to new, unseen generation methods because they overfit to specific visual flaws, and their accuracy degrades significantly when videos undergo common post-processing operations like compression or blurring. Reliable detection requires identifying fundamental manipulation traces that persist across different synthesis techniques and transmission degradations.

The article aims to design and evaluate a novel deepfake detection system—termed the Local- and Temporal-aware Transformer-based Deepfake Detection (LTTD) framework—that reliably spots video manipulations by capturing subtle, low-level temporal inconsistencies across local spatial patches.

To evaluate this framework, the authors conducted extensive cross-dataset and robustness experiments. The model divides incoming video clips into localized spatial patch sequences, models their frame-to-frame consistency using shallow three-dimensional convolutional enhancements alongside self-attention mechanisms, and aggregates this information to identify global contrasts between authentic and modified regions. Training was performed on the standard FaceForensics++ benchmark dataset, and generalization was evaluated across four challenging, unseen datasets totaling thousands of real and manipulated videos. Additional stress tests evaluated performance across seven real-world image perturbations (including compression, noise, blur, and color distortion) across five severity levels.

The experimental findings show that the proposed approach achieves state-of-the-art generalization and robustness. First, across four unseen target datasets, the framework achieved an average area under the curve (AUC) metric of 91.9%, outperforming competing detectors by substantial margins—such as achieving 80.4% AUC on the highly unconstrained Deepfake Detection Challenge dataset where many existing methods scored below 75%. Second, across all seven common video distortions, the model maintained an average AUC of 95.0%, suffering an overall performance drop of only 4.3% compared to pristine video, compared to drops of 7.3% and 22.3% in leading baselines. Third, ablation analyses confirmed that isolating localized patches effectively avoids overfitting to method-specific artifacts, grouping all unseen manipulation types into a unified feature distribution. Finally, consecutive frame sampling proved essential, as sparser temporal sampling degraded performance.

These results demonstrate that focusing on subtle, short-span temporal inconsistencies in localized regions provides a practical, highly generalizable foundation for digital content verification. Organizations can deploy such architectures with lower risk of failure when processing video degraded by internet bandwidth constraints, compression algorithms, or social media distribution pipelines. Unlike traditional models that require continuous retraining for every new facial manipulation algorithm, this framework offers more durable protection against evolving threats.

Decision-makers should consider piloting localized temporal-inconsistency detectors within automated content moderation, digital forensics, and media verification workflows. However, direct operational deployment should be approached in phases alongside human oversight. Further development should focus on establishing decision confidence calibration, evaluating computational efficiency on edge hardware, and stress-testing the model against future adversaries who may explicitly design generation algorithms to counter low-level temporal detection.

arXiv: 2207.02803
Cover for Delving into Sequential Patches for Deepfake Detection

Abstract

Recent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have identified the importance of local low-level cues and temporal information in pursuit to generalize well across deepfake methods, however, they still suffer from robustness problem against post-processings. In this work, we propose the Local- & Temporal-aware Transformer-based Deepfake Detection (LTTD) framework, which adopts a local-to-global learning protocol with a particular focus on the valuable temporal information within local sequences. Specifically, we propose a Local Sequence Transformer (LST), which models the temporal consistency on sequences of restricted spatial regions, where low-level information is hierarchically enhanced with shallow layers of learned 3D filters. Based on the local temporal embeddings, we then achieve the final classification in a global contrastive way. Extensive experiments on popular datasets validate that our approach effectively spots local forgery cues and achieves state-of-the-art performance.

Table of Contents

  • 1 Introduction
  • 2 Related work
  • 2.1 Deepfake detection
  • 2.2 Vision transformer
  • 3 Approach
  • 3.1 Problem statement
  • 3.2 Local Sequence Transformer
  • 3.3 Cross-Patch Inconsistency loss and Cross-Patch Aggregation
  • 4 Experiments
  • 4.1 Setups
  • 4.2 Generalizability evaluation
  • 4.3 Robustness to perturbations
  • 4.4 Ablations
  • 5 Conclusion and discussion
  • References
  • Checklist

Knowls

  1. Knowl 1 — Local- & Temporal-Aware Transformer-Based Deepfake Detection Framework

    model/method

    The Local- & Temporal-aware Transformer-based Deepfake Detection (LTTD) framework is designed to detect manipulated facial videos by isolating low-level temporal inconsistencies across local spatial regions. Given an input video clip v∈RT×C×H×Wv \in \mathbb{R}^{T \times C \times H \times W} of TT successive frames with spatial dimensions H×WH \times W and CC channels, each frame is divided into non-overlapping spatial patches of size P×PP \times P, producing N=HW/P2N = HW/P^2 patch locations. Patches from identical spatial positions across all TT frames are grouped into local patch sequences si={xt,i∈RC⋅P2∣t=1,2,…,T}s^i = \{x^{t,i} \in \mathbb{R}^{C \cdot P^2} \mid t = 1, 2, \dots, T\} for i∈{1,2,…,N}i \in \{1, 2, \dots, N\}.

    The framework operates in three cascaded stages:

    1. Independent Local Sequence Modeling: Each patch sequence sis^i is processed independently through a weight-shared Local Sequence Transformer (LST) module that combines shallow 3D convolutions and self-attention layers to extract location-specific temporal embeddings {zt,i}t=1T\{z^{t,i}\}_{t=1}^T and a location temporal token ztempiz^i_{\text{temp}}.
    2. Cross-Patch Inconsistency (CPI) Supervision: An intra-frame contrastive loss LCPI\mathcal{L}_{\text{CPI}} is computed across spatial regions to penalize discrepancies between pairwise cosine similarities of temporal embeddings and a ground-truth modification similarity matrix.
    3. Cross-Patch Aggregation (CPA): The location temporal tokens {ztempi}i=1N\{z^i_{\text{temp}}\}_{i=1}^N are prepended with a learnable class token xclassx_{\text{class}} and aggregated across all spatial positions via additional Transformer blocks, followed by a fully connected classification head that outputs the real-versus-fake prediction.
  2. Knowl 2 — Local Sequence Transformer (LST) Architecture

    model/method

    The Local Sequence Transformer (LST) processes patch sequences si={xt,i∈RC⋅P2∣t=1,…,T}s^i = \{x^{t,i} \in \mathbb{R}^{C \cdot P^2} \mid t=1,\dots,T\} at spatial location i∈{1,…,N}i \in \{1,\dots,N\} through a Local Sequence Embedding stage followed by three Low-level Enhanced Transformer (LET) stages.

    Local Sequence Embedding

    The input patches are embedded via a linear projection Es∈RD×(C⋅P2)E_s \in \mathbb{R}^{D \times (C \cdot P^2)} and concurrently filtered through a 3D convolution with max-pooling (kernel size k=2k=2, output channels Ct=64C_t=64):

    zs0i=[Esx1,i;Esx2,i;… ;EsxT,i]z_{s0}^i = [E_s x^{1,i}; E_s x^{2,i}; \dots; E_s x^{T,i}]

    [y01,i;y02,i;… ;y0T,i]=Maxpool(Conv3d([x1,i;x2,i;… ;xT,i]);k)[y_0^{1,i}; y_0^{2,i}; \dots; y_0^{T,i}] = \text{Maxpool}(\text{Conv3d}([x^{1,i}; x^{2,i}; \dots; x^{T,i}]); k)

    zt0i=[E0y01,i;E0y02,i;… ;E0y0T,i],E0∈RD×(Ct⋅(P/k)2)z_{t0}^i = [E_0 y_0^{1,i}; E_0 y_0^{2,i}; \dots; E_0 y_0^{T,i}], \quad E_0 \in \mathbb{R}^{D \times (C_t \cdot (P/k)^2)}

    z0i=[xtempi;zs0i(1)⋅σ(zt0i(1));zs0i(2)⋅σ(zt0i(2));… ;zs0i(T)⋅σ(zt0i(T))]+Eposz_0^i = [x_{\text{temp}}^i; z_{s0}^i(1) \cdot \sigma(z_{t0}^i(1)); z_{s0}^i(2) \cdot \sigma(z_{t0}^i(2)); \dots; z_{s0}^i(T) \cdot \sigma(z_{t0}^i(T))] + E_{\text{pos}}

    where xtempi∈RDx_{\text{temp}}^i \in \mathbb{R}^D is a learnable temporal token, σ(⋅)\sigma(\cdot) is the sigmoid function, and Epos∈R(T+1)×DE_{\text{pos}} \in \mathbb{R}^{(T+1) \times D} is a learnable temporal position embedding.

    Low-level Enhanced Transformer (LET) Stages

    For stages l∈{1,2,3}l \in \{1, 2, 3\}, given embedding zl−1i=[ztempi;z1,i;… ;zT,i]z_{l-1}^i = [z_{\text{temp}}^i; z^{1,i}; \dots; z^{T,i}] and feature maps yl−1iy_{l-1}^i:

    zsli=[ztempi′;zs1,i;… ;zsT,i]=Trans3(zl−1i)z_{sl}^i = [z_{\text{temp}}^{i\prime}; z_s^{1,i}; \dots; z_s^{T,i}] = \text{Trans}^3(z_{l-1}^i)

    [yl1,i;yl2,i;… ;ylT,i]=Maxpool(Conv3d([yl−11,i;yl−12,i;… ;yl−1T,i]);k)[y_l^{1,i}; y_l^{2,i}; \dots; y_l^{T,i}] = \text{Maxpool}(\text{Conv3d}([y_{l-1}^{1,i}; y_{l-1}^{2,i}; \dots; y_{l-1}^{T,i}]); k)

    ztli=[Elyl1,i;Elyl2,i;… ;ElylT,i],El∈RD×(Ct⋅(P/kl+1)2)z_{tl}^i = [E_l y_l^{1,i}; E_l y_l^{2,i}; \dots; E_l y_l^{T,i}], \quad E_l \in \mathbb{R}^{D \times (C_t \cdot (P / k^{l+1})^2)}

    zli=[ztempi′;zsli(1)⋅σ(ztli(1));… ;zsli(T)⋅σ(ztli(T))]z_l^i = [z_{\text{temp}}^{i\prime}; z_{sl}^i(1) \cdot \sigma(z_{tl}^i(1)); \dots; z_{sl}^i(T) \cdot \sigma(z_{tl}^i(T))]

    where Trans3(⋅)\text{Trans}^3(\cdot) denotes three cascaded standard Transformer encoder blocks operating along the temporal dimension. Independent processing reduces the computational complexity from O(T2⋅N2)O(T^2 \cdot N^2) (full spatiotemporal self-attention) to O(T⋅N2)O(T \cdot N^2).

  3. Knowl 3 — Cross-Patch Inconsistency (CPI) Loss

    equation

    The Cross-Patch Inconsistency (CPI) loss supervises pairwise similarity between spatial patch temporal representations based on frame manipulation masks. For spatial regions p,q∈{1,2,…,N}p, q \in \{1, 2, \dots, N\}, temporal representations are temporally averaged:

    fp=1T∑t=1Tzt,p,fq=1T∑t=1Tzt,qf^p = \frac{1}{T} \sum_{t=1}^T z^{t,p}, \quad f^q = \frac{1}{T} \sum_{t=1}^T z^{t,q}

    Dense cosine similarities between regions are computed as:

    simp,q=⟨fp,fq⟩∥fp∥⋅∥fq∥\text{sim}^{p,q} = \frac{\langle f^p, f^q \rangle}{\|f^p\| \cdot \|f^q\|}

    Given the ground-truth manipulation mask sequence mo∈RT×H×Wm_o \in \mathbb{R}^{T \times H \times W} (obtained by subtracting manipulated frames from original frames and normalizing to [0,1][0, 1]), the temporal mean and spatially pooled mask values are obtained via:

    mα=1T∑t=1Tmot∈RH×W,mβ=Interpolate(mα)∈RN×N,m=Flatten(mβ)∈RNm_\alpha = \frac{1}{T} \sum_{t=1}^T m_o^t \in \mathbb{R}^{H \times W}, \quad m_\beta = \text{Interpolate}(m_\alpha) \in \mathbb{R}^{\sqrt{N} \times \sqrt{N}}, \quad m = \text{Flatten}(m_\beta) \in \mathbb{R}^N

    The ground-truth similarity between patches pp and qq is defined as:

    simgtp,q=2⋅(1−∣mp−mq∣)−1∈[−1,1]\text{sim}_{\text{gt}}^{p,q} = 2 \cdot (1 - |m^p - m^q|) - 1 \in [-1, 1]

    For authentic (real) video clips without manipulation, simgt=1N×N\text{sim}_{\text{gt}} = \mathbf{1}^{N \times N}. The CPI loss is formulated as:

    LCPI=∑p=1N∑q=1N(max⁡(∣simp,q−simgtp,q∣−μ,0))2\mathcal{L}_{\text{CPI}} = \sum_{p=1}^N \sum_{q=1}^N \left( \max\left( |\text{sim}^{p,q} - \text{sim}_{\text{gt}}^{p,q}| - \mu, 0 \right) \right)^2

    where μ=0.1\mu = 0.1 is a tolerance margin allowing intra-class feature diversity.

  4. Knowl 4 — Cross-Patch Aggregation (CPA) and End-to-End Objective

    model/method

    After the Local Sequence Transformer (LST) yields spatial temporal tokens {ztempi∈RD∣i=1,…,N}\{z_{\text{temp}}^i \in \mathbb{R}^D \mid i=1, \dots, N\}, the Cross-Patch Aggregation (CPA) module aggregates global consistency across all NN spatial locations. A learnable class token xclass∈RDx_{\text{class}} \in \mathbb{R}^D is prepended:

    ZCPA(0)=[xclass;ztemp1;ztemp2;… ;ztempN]Z_{\text{CPA}}^{(0)} = [x_{\text{class}}; z_{\text{temp}}^1; z_{\text{temp}}^2; \dots; z_{\text{temp}}^N]

    This sequence is passed through three cascaded Transformer encoder blocks across the spatial dimension N+1N+1. The output embedding corresponding to the class token is fed into a fully connected (FC) layer with sigmoid activation to output the binary probability y^∈[0,1]\hat{y} \in [0, 1] of the clip being fake.

    The entire network is trained end-to-end using the joint objective:

    L=LBCE+λ⋅LCPI\mathcal{L} = \mathcal{L}_{\text{BCE}} + \lambda \cdot \mathcal{L}_{\text{CPI}}

    where LBCE\mathcal{L}_{\text{BCE}} is the binary cross-entropy loss between y^\hat{y} and the ground-truth label y∈{0,1}y \in \{0, 1\}, and λ=10−3\lambda = 10^{-3} balances the Cross-Patch Inconsistency loss.

  5. Knowl 5 — Cross-Dataset Generalizability Evaluation

    data/table

    Models were trained on FaceForensics++ (FF++ HQ) and tested without fine-tuning on four unseen datasets: CelebDF-V2 (CelebDF), DeepFake Detection Challenge (DFDC private test set, >3000 videos), FaceShift (FaceSh), and DeeperForensics (DeepFo). Performance is reported as video-level Area Under the Receiver Operating Characteristic Curve (AUC %):

    Method CelebDF DFDC FaceSh DeepFo Average
    CNN-GRU 69.8 68.9 80.8 74.1 73.4
    Multi-task 75.5 68.1 66.0 77.7 71.9
    PatchForensics 69.6 65.6 57.8 81.8 68.7
    FWA 69.5 67.3 65.5 50.2 63.1
    Face X-ray 79.5 65.5 92.8 86.8 81.2
    PCL+I2G 90.0 67.5 - 99.4 85.6
    SBI+EB4 89.9 74.9 97.4 77.7 85.0
    LipForensics 82.4 73.5 97.1 97.6 87.7
    FTCN-TT 86.9 74.0 98.8 98.8 89.6
    LTTD (ours) 89.3 80.4 99.5 98.5 91.9

    LTTD achieves a new state-of-the-art cross-dataset average AUC of 91.9%, outperforming semantic temporal models (FTCN-TT: 89.6%, LipForensics: 87.7%) and spatial boundary methods (Face X-ray: 81.2%), particularly on the challenging DFDC dataset where LTTD reaches 80.4% AUC (+6.4% over FTCN-TT).

  6. Knowl 6 — Robustness Evaluation Across Common Video Perturbations

    data/table

    Models were trained on clean FF++ and evaluated on seven perturbation types across five severity levels: Color Saturation (CS), Color Contrast (CC), Block-Wise Noise (BW), Gaussian Noise (GNC), Gaussian Blur (GB), Pixelation (PX), and Video Compression (VC). Video-level AUC (%) averaged across all 5 severity levels and average performance drop relative to clean evaluation are reported:

    Method Clean CS CC BW GNC GB PX VC Avg Drop
    Face X-ray 99.8 97.6 88.5 99.1 49.8 63.8 88.6 55.2 77.5 -22.3
    LipForensics 99.9 99.9 99.6 87.4 73.8 96.1 95.6 95.6 92.6 -7.3
    LTTD (ours) 99.4 98.9 96.4 96.1 82.6 97.5 98.6 95.0 95.0 -4.3

    LTTD exhibits highest overall robustness with an average perturbed AUC of 95.0% and the smallest degradation (-4.3%), retaining high detection capability under strong video compression (95.0% AUC) and pixelation (98.6% AUC), where spatial boundary methods like Face X-ray fail (55.2% and 88.6% AUC respectively).

  7. Knowl 7 — Ablation Study of LTTD Architectural Components

    data/table

    Ablation experiments evaluated on models trained on FF++ (HQ) demonstrate the impact of the Local Sequence Transformer (LST), Cross-Patch Inconsistency (CPI) loss, and Cross-Patch Aggregation (CPA) module on in-dataset performance (FF++) and unseen cross-dataset generalization (FaceSh, DFDC, DeepFo):

    FF++ FaceSh DFDC DeepFo Cross-Avg
    Method ACC AUC ACC AUC ACC AUC ACC AUC ACC AUC
    Xception 96.08 99.38 72.47 78.60 60.47 67.36 69.21 83.28 67.38 76.41
    ViT 95.00 97.92 62.86 65.56 64.81 72.89 71.85 83.24 66.51 73.90
    ViViT 94.71 97.92 63.21 77.40 67.52 74.16 60.12 82.86 63.62 78.14
    LTTD w/o LST 95.57 98.57 90.71 97.44 60.40 70.06 87.68 97.94 79.60 88.48
    LTTD w/o CPI 97.29 99.23 95.00 98.86 67.95 77.03 90.62 97.68 84.52 91.19
    LTTD w/o CPA 96.14 99.00 91.79 99.35 65.29 70.64 87.39 96.62 81.49 88.87
    LTTD 97.72 99.52 96.55 99.51 71.34 80.39 92.53 98.50 86.81 92.80

    Removing LST (replaced by standard patch embedding and global Transformer blocks) causes a 4.32% drop in Cross-Avg AUC; omitting CPI loss causes a 1.61% drop; omitting CPA (replaced by spatial average pooling and FC) causes a 3.93% drop, confirming that each module contributes significantly to cross-domain generalization.

  8. Knowl 8 — Ablation of Temporal Clip Length and Frame Sampling Interval

    data/table

    The effect of input clip length (CL, number of frames) and frame sampling space (FSS, sampling stride) on detection performance (AUC %) when trained on FF++:

    CL FSS FF++ DFDC DeepFo Average
    8 1 99.26 77.25 96.82 91.11
    16 1 99.52 80.39 98.50 92.80
    32 1 99.47 80.92 98.41 92.93
    64 1 99.38 79.23 98.49 92.37
    16 2 99.15 76.17 97.32 90.88
    16 4 98.51 73.02 91.33 87.62

    Varying clip length from 16 to 64 consecutive frames (FSS=1) produces nearly identical average AUC (92.80% vs 92.93%), indicating that long video spans are unnecessary to capture local temporal inconsistency. In contrast, sparse sampling (increasing FSS to 2 or 4) degrades cross-dataset average AUC significantly from 92.80% to 87.62%, proving that contiguous frame-to-frame transitions are essential for low-level temporal forgery pattern learning.

  9. Knowl 9 — Jitter-Free Video Preprocessing Protocol

    experimental setup

    Standard per-frame face bounding box cropping introduces artificial temporal jitter due to independent frame-level bounding box predictions, which corrupts genuine low-level temporal patterns.

    To prevent this artifact:

    1. A random temporal clip window of length T=16T=16 is selected on-the-fly from the video.
    2. A single static bounding box encompassing the entire facial region across all TT frames is computed using MTCNN face detection on the clip.
    3. All TT frames in the clip are cropped using this identical bounding box and resized to spatial dimensions H×W=224×224H \times W = 224 \times 224.
    4. Each frame is divided into N=196N = 196 non-overlapping patches of size P×P=16×16P \times P = 16 \times 16.
    5. At evaluation time, the first 128 frames of each video are divided into 8 clips of 16 frames, and clip predictions are averaged to produce the final video-level score.
  10. Knowl 10 — Limitations of LTTD Deepfake Detection

    limitation

    The LTTD approach exhibits the following limitations:

    1. High Gaussian Noise Sensitivity: Under heavy Gaussian noise perturbations, low-level spatial-temporal cues are degraded, causing detection performance to drop to 82.6% AUC.
    2. Vulnerability to Future Generative Pipelines: Deepfake generation methods that explicitly incorporate adversarial training on low-level temporal features or enforce strict frame-to-frame temporal regularity could circumvent detection.
    3. Real-World Calibration and Deployment: The confidence scores of the detector are not explicitly calibrated for open-world deployment where domain shifts, mixed compression formats, and unseen camera hardware occur simultaneously.

Coverage note — Qualitative Class Activation Map (CAM) forgery localization figures and t-SNE feature manifold visualizations were omitted as standalone knowls because their qualitative insights are fully covered by the architectural and empirical ablation knowls.

References

  1. 1.Deepfakes github. https://github.com/deepfakes/faceswap, 2022. Accessed: 2022-02-28.
  2. 2.Deepfakes github. https://github.com/MarekKowalski/FaceSwap, 2022. Accessed: 2022-02-28.
  3. 3.Dfdc challenge. https://www.kaggle.com/c/deepfake-detection-challenge, 2022. Accessed: 2022-04-04.
  4. 4.Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. Mesonet: a compact facial video forgery detection network. In 2018 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–7. IEEE, 2018.
  5. 5.Shruti Agarwal and Hany Farid. Detecting deep-fake videos from aural and oral dynamics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 981–989, 2021.
  6. 6.Shruti Agarwal, Hany Farid, Tarek El-Gaaly, and Ser-Nam Lim. Detecting deep-fake videos from appearance and behavior. In 2020 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–6. IEEE, 2020.
  7. 7.Irene Amerini, Leonardo Galteri, Roberto Caldelli, and Alberto Del Bimbo. Deepfake video detection through optical flow based cnn. In Proceedings of the IEEE/CVF international conference on computer vision workshops, pages 0–0, 2019.
  8. 8.Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6836–6846, 2021.
  9. 9.Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. What makes fake images detectable? understanding properties that generalize. In European Conference on Computer Vision, pages 103–120. Springer, 2020.
  10. 10.Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. Local relation learning for face forgery detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1081–1088, 2021.
  11. 11.François Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1251–1258, 2017.
  12. 12.Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, and Michael Ying Yang. Spatial-temporal transformer for dynamic scene graph generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 16372–16382, 2021.
  13. 13.Hao Dang, Feng Liu, Joel Stehouwer, Xiaoming Liu, and Anil K Jain. On the detection of digital face manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5781–5790, 2020.
  14. 14.Sowmen Das, Selim Seferbekov, Arup Datta, Md Islam, Md Amin, et al. Towards solving the deepfake problem: An analysis on improving deepfake detection using dynamic face augmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3776–3785, 2021.
  15. 15.Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397, 2020.
  16. 16.Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Ting Zhang, Weiming Zhang, Nenghai Yu, Dong Chen, Fang Wen, and Baining Guo. Protecting celebrities with identity consistency transformer. arXiv preprint arXiv:2203.01318, 2022.
  17. 17.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
  18. 18.Brendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi, and Graham W Taylor. Sstvos: Sparse spatiotemporal transformers for video object segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5912–5921, 2021.
  19. 19.Qiqi Gu, Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, and Ran Yi. Exploiting fine-grained face forgery clues via progressive enhancement learning. arXiv preprint arXiv:2112.13977, 2021.
  20. 20.Jiazhi Guan, Hang Zhou, Mingming Gong, Youjian Zhao, Errui Ding, and Jingdong Wang. Detecting deepfake by creating spatio-temporal regularity disruption. arXiv preprint arXiv:2207.10402, 2022.
  21. 21.Alexandros Haliassos, Konstantinos Vougioukas, Stavros Petridis, and Maja Pantic. Lips don’t lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5039–5049, 2021.
  22. 22.Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. Learning spatio-temporal features with 3d residual networks for action recognition. In Proceedings of the IEEE International Conference on Computer Vision Workshops, pages 3154–3160, 2017.
  23. 23.Xinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu, Chen Change Loy, Xun Cao, and Feng Xu. Audio-driven emotional video portraits. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  24. 24.Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2889–2898, 2020.
  25. 25.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 4401–4410, 2019.
  26. 26.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8110–8119, 2020.
  27. 27.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  28. 28.Mohammad Rami Koujan, Michail Christos Doukas, Anastasios Roussos, and Stefanos Zafeiriou. Head2head: Video-based neural head synthesis. arXiv preprint arXiv:2005.10954, 2020.
  29. 29.Jiaming Li, Hongtao Xie, Jiahong Li, Zhongyuan Wang, and Yongdong Zhang. Frequency-aware discriminative feature learning supervised by single-center loss for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6458–6467, 2021.
  30. 30.Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. Advancing high fidelity identity swapping for forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5074–5083, 2020.
  31. 31.Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. Face x-ray for more general face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5001–5010, 2020.
  32. 32.Xiangyu Li, Yonghong Hou, Pichao Wang, Zhimin Gao, Mingliang Xu, and Wanqing Li. Trear: Transformer-based rgb-d egocentric action recognition. IEEE Transactions on Cognitive and Developmental Systems, 2021.
  33. 33.Yuezun Li, Ming-Ching Chang, and Siwei Lyu. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In 2018 IEEE International Workshop on Information Forensics and Security (WIFS), pages 1–7. IEEE, 2018.
  34. 34.Yuezun Li and Siwei Lyu. Exposing deepfake videos by detecting face warping artifacts. arXiv preprint arXiv:1811.00656, 2018.
  35. 35.Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. Celeb-df: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3207–3216, 2020.
  36. 36.Honggu Liu, Xiaodan Li, Wenbo Zhou, Yuefeng Chen, Yuan He, Hui Xue, Weiming Zhang, and Nenghai Yu. Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 772–781, 2021.
  37. 37.Xian Liu, Yinghao Xu, Qianyi Wu, Hang Zhou, Wayne Wu, and Bolei Zhou. Semantic-aware implicit neural audio-driven video portrait generation. ECCV, 2022.
  38. 38.Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu. Video swin transformer. arXiv preprint arXiv:2106.13230, 2021.
  39. 39.Zhengzhe Liu, Xiaojuan Qi, and Philip HS Torr. Global texture enhancement for fake face detection in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8060–8069, 2020.
  40. 40.Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. Generalizing face forgery detection with high-frequency features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16317–16326, 2021.
  41. 41.Pingchuan Ma, Brais Martinez, Stavros Petridis, and Maja Pantic. Towards practical lipreading with distilled and efficient models. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7608–7612. IEEE, 2021.
  42. 42.Huy H Nguyen, Fuming Fang, Junichi Yamagishi, and Isao Echizen. Multi-task learning for detecting and segmenting manipulated facial images and videos. In 2019 IEEE 10th International Conference on Biometrics Theory, Applications and Systems (BTAS), pages 1–8. IEEE, 2019.
  43. 43.Chiara Plizzari, Marco Cannici, and Matteo Matteucci. Spatial temporal transformer network for skeleton-based action recognition. In International Conference on Pattern Recognition, pages 694–701. Springer, 2021.
  44. 44.Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In European Conference on Computer Vision, pages 86–103. Springer, 2020.
  45. 45.Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 1–11, 2019.
  46. 46.Ekraam Sabir, Jiaxin Cheng, Ayush Jaiswal, Wael AbdAlmageed, Iacopo Masi, and Prem Natarajan. Recurrent convolutional strategies for face manipulation detection in videos. Interfaces (GUI), 3(1):80–87, 2019.
  47. 47.Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision, pages 618–626, 2017.
  48. 48.Kaede Shiohara and Toshihiko Yamasaki. Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18720–18729, 2022.
  49. 49.Ke Sun, Taiping Yao, Shen Chen, Shouhong Ding, Rongrong Ji, et al. Dual contrastive learning for general face forgery detection. arXiv preprint arXiv:2112.13522, 2021.
  50. 50.Zekun Sun, Yujie Han, Zeyu Hua, Na Ruan, and Weijia Jia. Improving the efficiency and robustness of deepfakes detection through precise geometric features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3609–3618, 2021.
  51. 51.Shahroz Tariq, Sangyup Lee, and Simon S Woo. A convolutional lstm based residual network for deepfake video detection. arXiv preprint arXiv:2009.07480, 2020.
  52. 52.Justus Thies, Michael Zollhöfer, and Matthias Nießner. Deferred neural rendering: Image synthesis using neural textures. ACM Transactions on Graphics (TOG), 38(4):1–12, 2019.
  53. 53.Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2387–2395, 2016.
  54. 54.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008.
  55. 55.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in neural information processing systems, pages 5998–6008, 2017.
  56. 56.Chengrui Wang and Weihong Deng. Representative forgery mining for fake face detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14923–14932, 2021.
  57. 57.Junke Wang, Zuxuan Wu, Jingjing Chen, and Yu-Gang Jiang. M2tr: Multi-modal multi-scale transformers for deepfake detection. arXiv preprint arXiv:2104.09770, 2021.
  58. 58.Xin Yang, Yuezun Li, and Siwei Lyu. Exposing deep fakes using inconsistent head poses. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8261–8265. IEEE, 2019.
  59. 59.Junbo Yin, Jianbing Shen, Chenye Guan, Dingfu Zhou, and Ruigang Yang. Lidar-based online 3d video object detection with graph-based message passing and spatiotemporal transformer attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11495–11504, 2020.
  60. 60.Daichi Zhang, Chenyu Li, Fanzhao Lin, Dan Zeng, and Shiming Ge. Detecting deepfake videos with temporal dropout 3dcnn. IJCAI, 2021.
  61. 61.Kaipeng Zhang, Zhanpeng Zhang, Zhifeng Li, and Yu Qiao. Joint face detection and alignment using multitask cascaded convolutional networks. IEEE Signal Processing Letters, 23(10):1499–1503, 2016.
  62. 62.Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe. Vidtr: Video transformer without convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 13577–13587, 2021.
  63. 63.Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. Multi-attentional deepfake detection. arXiv preprint arXiv:2103.02406, 2021.
  64. 64.Tianchen Zhao, Xiang Xu, Mingze Xu, Hui Ding, Yuanjun Xiong, and Wei Xia. Learning self-consistency for deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15023–15033, 2021.
  65. 65.Yinglin Zheng, Jianmin Bao, Dong Chen, Ming Zeng, and Fang Wen. Exploring temporal coherence for more general video face forgery detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15044–15054, 2021.
  66. 66.Hang Zhou, Yasheng Sun, Wayne Wu, Chen Change Loy, Xiaogang Wang, and Ziwei Liu. Pose-controllable talking face generation by implicitly modularized audio-visual representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  67. 67.Peng Zhou, Xintong Han, Vlad I Morariu, and Larry S Davis. Two-stream neural networks for tampered face detection. In 2017 IEEE conference on computer vision and pattern recognition workshops (CVPRW), pages 1831–1839. IEEE, 2017.
  68. 68.Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159, 2020.
  69. 69.Bojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma, and Yu-Gang Jiang. Wilddeepfake: A challenging real-world dataset for deepfake detection. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2382–2390, 2020.

Citation

MLA
Guan, J., et al. “Delving into Sequential Patches for Deepfake Detection”. Advances in Neural Information Processing Systems, vol. 35, 2022, pp. 4517–30, https://proceedings.neurips.cc/paper_files/paper/2022/file/1d051fb631f104cb2a621451f37676b9-Paper-Conference.pdf.
APA
Guan, J., Zhou, H., Hong, Z., Ding, E., Wang, J., Quan, C., & Zhao, Y. (2022). Delving into Sequential Patches for Deepfake Detection. Advances in Neural Information Processing Systems, 35, 4517–4530. https://proceedings.neurips.cc/paper_files/paper/2022/file/1d051fb631f104cb2a621451f37676b9-Paper-Conference.pdf
Chicago
Guan, J., H. Zhou, Z. Hong, et al. 2022. “Delving into Sequential Patches for Deepfake Detection”. Advances in Neural Information Processing Systems 35: 4517–30. https://proceedings.neurips.cc/paper_files/paper/2022/file/1d051fb631f104cb2a621451f37676b9-Paper-Conference.pdf.
Harvard
Guan, J. et al. (2022) “Delving into Sequential Patches for Deepfake Detection”, Advances in Neural Information Processing Systems. Curran Associates, Inc., pp. 4517–4530. Available at: https://proceedings.neurips.cc/paper_files/paper/2022/file/1d051fb631f104cb2a621451f37676b9-Paper-Conference.pdf.
Vancouver
1. Guan J, Zhou H, Hong Z, Ding E, Wang J, Quan C, Zhao Y (2022) Delving into Sequential Patches for Deepfake Detection. In: Advances in Neural Information Processing Systems. Curran Associates, Inc., pp 4517–4530

BibTeX

@inproceedings{guan2022delving,
  title = {Delving into Sequential Patches for Deepfake Detection},
  author = {Guan, Jiazhi and Zhou, Hang and Hong, Zhibin and Ding, Errui and Wang, Jingdong and Quan, Chengbin and Zhao, Youjian},
  year = {2022},
  booktitle = {Advances in Neural Information Processing Systems},
  publisher = {Curran Associates, Inc.},
  volume = {35},
  pages = {4517-4530},
  url = {https://proceedings.neurips.cc/paper_files/paper/2022/file/1d051fb631f104cb2a621451f37676b9-Paper-Conference.pdf}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors