Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection

Yuan WangKun YuChen ChenXiyuan HuSilong Peng

article2023CVPR183 citations

Proposes a Spatial-Frequency Dynamic Graph framework for deepfake detection that extracts content-adaptive frequency artifacts and reasons over high-order cross-domain relationships using dynamic graph learning to generalize across diverse face manipulation benchmarks.

Listen

Rapid advancements in digital face synthesis tools allow individuals to manipulate facial identities and expressions with unprecedented realism. This technology poses severe risks to security, public trust, and compliance by enabling deceptive media, political disinformation, and non-consensual imagery. While early automated detection tools focused on obvious visual flaws or hand-crafted frequency patterns, modern counterfeit methods easily evade these static defenses. Existing systems struggle primarily because they analyze visual textures and frequency data in isolation without understanding the complex, content-driven relationships between them.

The article demonstrates an advanced deepfake detection framework named Spatial-Frequency Dynamic Graph that captures content-aware frequency clues and reasons about the complex interactions between visual image content and underlying frequency distributions. To accomplish this, the authors introduce a three-stage architecture that adaptively extracts fine and coarse frequency clues using visual masks, generates multi-scale spatial and frequency attention maps to retain rich context, and applies a dynamic graph convolutional network with multi-layer perceptron mixers to evaluate the relationships between both domains.

The system was evaluated across six major industry benchmark datasets—including FaceForensics++, WildDeepfake, Celeb-DF, and Deepfake Detection Challenge—under varying image qualities and cross-dataset testing. The evaluation produced four key findings. First, the proposed framework consistently outperformed leading deepfake detection baselines, achieving 92.28% accuracy and an area under the curve score of 95.98% on low-quality manipulated video tests. Second, in cross-dataset evaluations where the model was trained on one source and tested against unseen manipulation techniques, it attained strong generalization with area under the curve scores of 75.83% on Celeb-DF, 88.00% on DFD, and 73.64% on DFDC, markedly exceeding rival approaches. Third, perturbation stress testing demonstrated superior noise resilience, dropping by only 10.10% accuracy under severe salt-and-pepper noise where competing systems suffered declines between 31% and 49%. Fourth, ablation and parameter studies confirmed that combining content-adaptive extraction with multi-scale attention and a ten-neighbor dynamic graph construction produces the most robust and accurate classification.

These findings indicate that integrating adaptive frequency extraction with relational graph learning significantly reduces the risk of models overfitting to specific, known counterfeit generation tools. By capturing subtle discrepancies across wide semantic areas—such as backgrounds and hair—alongside localized facial anomalies, this approach substantially lowers false-positive and false-negative detection rates in compressed and degraded real-world video pipelines. Consequently, this architecture provides a viable, high-accuracy foundation for operational content moderation and automated media verification systems.

Organizations evaluating this technology should consider moving toward dual-domain, graph-based verification pipelines rather than relying solely on visual or hand-crafted frequency filters. Stakeholders should conduct pilot evaluations on proprietary, real-world media streams to assess computational trade-offs, particularly regarding graph neighbor parameter tuning. While the findings provide strong confidence in controlled and cross-dataset benchmarks, operational deployment must account for potential latency constraints from dynamic graph computations and the need for continued testing against emerging, post-generation adversarial attacks.

Cover for Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection

Abstract

With the springing up of face synthesis techniques, it is prominent in need to develop powerful face forgery detection methods due to security concerns. Some existing methods attempt to employ auxiliary frequency-aware information combined with CNN backbones to discover the forged clues. Due to the inadequate information interaction with image content, the extracted frequency features are thus spatially irrelevant, struggling to generalize well on increasingly realistic counterfeit types. To address this issue, we propose a Spatial-Frequency Dynamic Graph method to exploit the relation-aware features in spatial and frequency domains via dynamic graph learning. To this end, we introduce three well-designed components: 1) Content-guided Adaptive Frequency Extraction module to mine the content-adaptive forged frequency clues. 2) Multiple Domains Attention Map Learning module to enrich the spatial-frequency contextual features with multiscale attention maps. 3) Dynamic Graph Spatial-Frequency Feature Fusion Network to explore the high-order relation of spatial and frequency features. Extensive experiments on several benchmark show that our proposed method sustainedly exceeds the state-of-the-arts by a considerable margin.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 3. Method
  • 3.1. Content-aware Frequency Extraction
  • 3.2. Multiple Domains Attention Map Learning
  • 3.3. Dynamic Graph Learning
  • 3.4. Loss Function
  • 4. Experiment
  • 4.1. Experimental Setup
  • 4.2. Experimental Results
  • 4.3. Ablation Study
  • 4.4. Visualization
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Spatial-Frequency Dynamic Graph Architecture for Deepfake Detection

    model/method

    The Spatial-Frequency Dynamic Graph (SFDG) framework is an end-to-end face forgery detection model designed to exploit high-order relations between spatial representations and content-adaptive frequency features. The framework consists of three principal modules:

    1. Content-guided Adaptive Frequency Extraction (CAFÉ): Decomposes input RGB images into coarse-grained frequency bands and fine-grained sliding-window frequency representations, merging them adaptively using a spatial content mask predicted by a U-Net submodule.
    2. Multiple Domains Attention Map Learning (MDAML): Leverages an EfficientNet-b4 backbone partitioned into low-, mid-, and high-level feature streams. It employs a Multi-Scale Attention Ensemble (MSAE) combined with Attention Map Refinement Blocks (AMRB) to generate multiscale spatial and frequency attention maps, which are subsequently aggregated via Bilinear Attention Pooling (BAP) with low-level texture features and high-level semantics.
    3. Dynamic Graph Spatial-Frequency Feature Fusion Network (DG-SF3^3Net): Constructs dynamic kk-nearest neighbor (kNNk\text{NN}) graphs across concatenated spatial and frequency feature channels, executing dynamic graph convolutions and graph-weighted MLP-Mixer interaction layers to discover cross-domain semantic correspondences.

    The framework is trained end-to-end using a joint objective combining standard Cross-Entropy loss with Region Independent Loss (RIL).

  2. Knowl 2 — Content-guided Adaptive Frequency Extraction

    model/method

    The Content-guided Adaptive Frequency Extraction (CAFÉ) module extracts content-aware frequency representations from an input RGB image xs∈R3×H×Wx^s \in \mathbb{R}^{3 \times H \times W}.

    First, a spatial content mask Ms∈R1×H×WM^s \in \mathbb{R}^{1 \times H \times W} is estimated using a U-Net architecture to preserve content-relevant spatial regions.

    Second, coarse-grained frequency decomposition is performed by applying NfN_f manually designed binary frequency-band filters {fi}i=1Nf\{f_i\}_{i=1}^{N_f} across the 2D Discrete Cosine Transform (DCT) domain to separate low-, middle-, and high-frequency components: xcoarsef,i=H−1(H(xs)⊙fi),i∈{1,…,Nf}x_{\text{coarse}}^{f, i} = \mathcal{H}^{-1}\left( \mathcal{H}(x^s) \odot f_i \right), \quad i \in \{1, \dots, N_f\} where H\mathcal{H} and H−1\mathcal{H}^{-1} denote the 2D DCT and inverse 2D DCT operators, respectively, and ⊙\odot represents the Hadamard product. Stacking across filters produces the coarse frequency representation xcoarsef∈R3Nf×H×Wx_{\text{coarse}}^f \in \mathbb{R}^{3N_f \times H \times W}.

    Third, fine-grained localized frequency extraction slices xsx^s into l×ll \times l non-overlapping patches pm,n∈R3×l×lp_{m, n} \in \mathbb{R}^{3 \times l \times l}, applies DCT to each patch to obtain localized frequency coefficients dm,nf∈R3×l×ld_{m, n}^f \in \mathbb{R}^{3 \times l \times l}, repeats the channels to match 3Nf3N_f, concatenates all patches into df∈R3Nf×H×Wd^f \in \mathbb{R}^{3N_f \times H \times W}, and processes them through Conv2d-BatchNorm-ReLU blocks to produce xfinef∈R3Nf×H×Wx_{\text{fine}}^f \in \mathbb{R}^{3N_f \times H \times W}.

    Finally, the adaptive frequency head xfh∈R3Nf×H×Wx^{fh} \in \mathbb{R}^{3N_f \times H \times W} is formed via content-mask modulation: xfh=(1−Ms)⊙xfinef+Ms⊙xcoarsefx^{fh} = (1 - M^s) \odot x_{\text{fine}}^f + M^s \odot x_{\text{coarse}}^f Strided convolutions project xfhx^{fh} into a final frequency feature tensor xf∈RCf×Hf×Wfx^f \in \mathbb{R}^{C_f \times H_f \times W_f} with CfC_f channels.

  3. Knowl 3 — Multiple Domains Attention Map Learning and Bilinear Attention Pooling

    model/method

    The Multiple Domains Attention Map Learning (MDAML) module enriches spatial and frequency features with multiscale semantic contextual attention. The network uses Multi-Scale Attention Ensemble (MSAE) blocks and Attention Map Refinement Blocks (AMRB) in both spatial and frequency streams.

    MSAE downsamples feature maps via convolutional layers and global average pooling into hierarchical scales to expand receptive fields. AMRB captures global context per scale using global average pooling and a Sigmoid activation, generating attention weights that rescale the multiscale maps before spatial upsampling and summation into refined attention maps Fs∈RTs×H×WF^s \in \mathbb{R}^{T_s \times H \times W} (spatial) and Ff∈RTf×H×WF^f \in \mathbb{R}^{T_f \times H \times W} (frequency), where TsT_s and TfT_f denote the number of attention channels in each domain.

    To capture low-level manipulated artifacts, a Textural Feature Enhancement block extracts texture features Ftex∈RCtex×Htex×WtexF_{\text{tex}} \in \mathbb{R}^{C_{\text{tex}} \times H_{\text{tex}} \times W_{\text{tex}}}. For each spatial attention map FksF_k^s (k∈{1,…,Ts}k \in \{1, \dots, T_s\}), partial texture features Fktex=Ftex⊙FksF_{k}^{\text{tex}} = F_{\text{tex}} \odot F_k^s are computed. Bilinear Attention Pooling (BAP) produces the kk-th spatial attention vector pks∈R1×Ctexp_k^s \in \mathbb{R}^{1 \times C_{\text{tex}}}: pks=∑m=1Htex∑n=1WtexFk,m,ntex∥∑m=1Htex∑n=1WtexFk,m,ntex∥2p_k^s = \frac{\sum_{m=1}^{H_{\text{tex}}} \sum_{n=1}^{W_{\text{tex}}} F_{k, m, n}^{\text{tex}}}{\left\| \sum_{m=1}^{H_{\text{tex}}} \sum_{n=1}^{W_{\text{tex}}} F_{k, m, n}^{\text{tex}} \right\|_2} Stacking {pks}k=1Ts\{p_k^s\}_{k=1}^{T_s} forms the spatial relation feature matrix Ps∈RTs×NsP^s \in \mathbb{R}^{T_s \times N_s}. Analogously, BAP combines frequency attention maps FfF^f with high-level spatial backbone features to generate the frequency relation feature matrix Pf∈RTf×NfP^f \in \mathbb{R}^{T_f \times N_f}.

  4. Knowl 4 — Dynamic Graph Spatial-Frequency Feature Fusion Network

    model/method

    The Dynamic Graph Spatial-Frequency Feature Fusion Network (DG-SF3^3Net) reasons about high-order cross-domain interactions between spatial and frequency features.

    The spatial and frequency feature matrices Ps∈RTs×NsP^s \in \mathbb{R}^{T_s \times N_s} and Pf∈RTf×NfP^f \in \mathbb{R}^{T_f \times N_f} are concatenated along the channel dimension to form the initial node matrix V(0)∈RT(0)×N(0)V^{(0)} \in \mathbb{R}^{T^{(0)} \times N^{(0)}}, where T(0)=Ts+TfT^{(0)} = T_s + T_f and N(0)=max⁡(Ns,Nf)N^{(0)} = \max(N_s, N_f). Each column vector vi(0)v_i^{(0)} acts as a graph node.

    At layer tt, a dynamic kk-nearest neighbor (kNNk\text{NN}) graph G(t)=(V(t),E(t))G^{(t)} = (V^{(t)}, E^{(t)}) is constructed based on pairwise Euclidean distances: N(t)(i)={vjm(t)∣vjm(t)∈kNN(vi(t)),m=1,…,k}\mathcal{N}^{(t)}(i) = \left\{ v_{j_m}^{(t)} \mid v_{j_m}^{(t)} \in k\text{NN}\left(v_i^{(t)}\right), m = 1, \dots, k \right\} An adjacency matrix A~(t)\tilde{A}^{(t)} with self-loops is computed, and node representations are updated via dynamic graph convolution: V(t+1)=ReLU(D~(t)−12A~(t)D~(t)−12V(t)W(t))V^{(t+1)} = \text{ReLU}\left( \tilde{D}^{(t)-\frac{1}{2}} \tilde{A}^{(t)} \tilde{D}^{(t)-\frac{1}{2}} V^{(t)} W^{(t)} \right) where D~(t)∈RT(t)×T(t)\tilde{D}^{(t)} \in \mathbb{R}^{T^{(t)} \times T^{(t)}} is the degree matrix of A~(t)\tilde{A}^{(t)} and W(t)∈RN(t)×N(t+1)W^{(t)} \in \mathbb{R}^{N^{(t)} \times N^{(t+1)}} is a learnable projection matrix.

    Following graph convolutions, Graph Information Interaction (GI2^2) blocks apply interleaved channel-mixing MLPs and token-mixing MLPs (based on MLP-Mixer). The token-mixing MLP activations are explicitly weighted by the adjacency matrix from the first GCN layer to preserve localized neighborhood relationships across nodes.

  5. Knowl 5 — Region Independent Loss for Spatial-Frequency Attention Features

    equation

    The SFDG model is optimized with a joint objective combining standard binary cross-entropy loss LCE\mathcal{L}_{CE} and Region Independent Loss LRIL\mathcal{L}_{RIL}: L=λ1LCE+λ2LRIL\mathcal{L} = \lambda_1 \mathcal{L}_{CE} + \lambda_2 \mathcal{L}_{RIL} where λ1=1\lambda_1 = 1 and λ2=1\lambda_2 = 1.

    Given the post-graph semantic feature matrix V~∈RM×N\tilde{V} \in \mathbb{R}^{M \times N} with MM attention maps and NN feature dimensions, LRIL\mathcal{L}_{RIL} enforces intra-class compactness and inter-class separability across a batch of size BB: LRIL=∑i=1B∑j=1MReLU(∥V~ji−cjt∥22−min(yi))+∑i≠jMReLU(mout−∥cit−cjt∥22)\mathcal{L}_{RIL} = \sum_{i=1}^B \sum_{j=1}^M \text{ReLU}\left( \left\|\tilde{V}_j^i - c_j^t\right\|_2^2 - m_{\text{in}}(y_i) \right) + \sum_{i \neq j}^M \text{ReLU}\left( m_{\text{out}} - \left\|c_i^t - c_j^t\right\|_2^2 \right) where c∈RM×Nc \in \mathbb{R}^{M \times N} are attention feature centers updated iteratively during training, yi∈{0,1}y_i \in \{0, 1\} is the binary ground-truth forgery label of sample ii, min(yi)∈[0.05,0.1]m_{\text{in}}(y_i) \in [0.05, 0.1] is the intra-class margin constraint, and mout=0.2m_{\text{out}} = 0.2 is the inter-class margin constraint.

  6. Knowl 6 — Face Forgery Detection Performance Across Standard Benchmarks

    data/table

    The SFDG method was evaluated against state-of-the-art face forgery detection methods across four benchmark configurations: FaceForensics++ low quality (FF++ LQ, c40 compression), FaceForensics++ high quality (FF++ HQ, c23 compression), WildDeepfake, and Celeb-DF. Evaluation metrics are binary classification Accuracy (Acc, %) and Area Under the ROC Curve (AUC, %).

    Method FF++ (LQ) FF++ (HQ) WildDeepfake Celeb-DF
    Acc AUC Acc AUC Acc AUC Acc AUC
    Xception 86.86 89.30 95.73 96.30 79.99 88.86 97.90 99.73
    EfficientNet-b4 86.67 88.20 96.63 99.18 82.33 90.12 98.19 99.83
    Add-Net 87.50 91.01 96.78 97.74 76.25 86.17 96.93 99.55
    SCL 89.00 92.40 96.69 99.30 – – – –
    MADD 88.69 90.40 97.60 99.29 82.62 90.71 97.92 99.94
    F3Net 90.43 93.30 97.52 98.10 80.66 87.53 95.95 98.93
    PEL 90.52 94.28 97.63 99.32 84.14 91.62 – –
    RECCE 91.03 95.02 97.06 99.32 83.25 92.02 98.59 99.94
    Local Relation 91.47 95.21 97.59 99.46 – – – –
    M2TR 92.35 94.22 98.23 99.48 – – – –
    SFDG (Xception) 91.08 94.49 97.61 99.45 83.36 92.15 98.95 99.94
    SFDG (Ours) 92.28 95.98 98.19 99.53 84.41 92.57 99.22 99.96

    SFDG (with EfficientNet-b4 backbone) achieves the highest AUC on all four benchmarks (95.98% on FF++ LQ, 99.53% on FF++ HQ, 92.57% on WildDeepfake, and 99.96% on Celeb-DF) and outperforms prior spatial-frequency methods (such as F3Net and PEL) particularly under heavy compression (FF++ LQ).

  7. Knowl 7 — Cross-Dataset Generalization of Deepfake Detectors Trained on FF++

    data/table

    Models trained exclusively on the FaceForensics++ (FF++ LQ) training set were evaluated on five unseen test datasets without fine-tuning: Celeb-DF, DeepFake Detection (DFD), WildDeepfake, Deepfake Detection Challenge preview (DFDC), and DeeperForensics-1.0 (DF-v1.0). Metrics reported are AUC (%) and Equal Error Rate (EER, where lower ↓\downarrow is better).

    Method Celeb-DF DFD WildDeepfake DFDC DF-v1.0
    AUC EER↓\downarrow AUC EER↓\downarrow AUC EER↓\downarrow AUC EER↓\downarrow AUC EER↓\downarrow
    Xception 60.05 0.432 65.43 0.393 60.59 0.619 55.65 0.461 80.27 0.265
    EfficientNet-b4 64.29 0.419 83.17 0.235 64.27 0.376 60.12 0.428 85.31 0.228
    Add-Net 57.83 0.444 57.16 0.453 54.21 0.462 51.60 0.548 – –
    MADD 68.64 0.371 74.18 0.327 65.65 0.397 63.02 0.410 89.34 0.173
    F3Net 67.95 0.368 69.50 0.354 60.49 0.434 57.87 0.442 82.27 0.246
    PEL 69.18 0.357 75.86 0.308 67.39 0.383 63.31 0.404 – –
    SFDG (Ours) 75.83 0.303 88.00 0.197 69.27 0.377 73.64 0.337 92.10 0.151

    SFDG achieves the highest AUC and lowest EER across all five out-of-domain test sets, demonstrating that reasoning over dynamic spatial-frequency graph relations prevents overfitting to manipulation-specific artifacts present in the training set.

  8. Knowl 8 — Robustness of Deepfake Detection Against Common Visual Perturbations

    data/table

    The robustness of detectors was measured by applying three common image corruptions—Gaussian Noise, Salt & Pepper Noise, and Gaussian Blur—to test images from FF++ (LQ) and WildDeepfake. The metric reported is the accuracy degradation (ΔAcc\Delta\text{Acc}, in percentage points relative to unperturbed accuracy; smaller negative magnitude indicates higher robustness).

    Method +GaussianNoise +SaltPepperNoise +GaussianBlur
    Δ\DeltaAcc(FF) Δ\DeltaAcc(Wild) Δ\DeltaAcc(FF) Δ\DeltaAcc(Wild) Δ\DeltaAcc(FF) Δ\DeltaAcc(Wild)
    Xception -2.65% -0.98% -32.44% -27.80% -6.22% -12.71%
    Add-Net -41.51% -11.66% -11.28% -18.21% -11.28% -12.91%
    F3Net -9.86% -1.17% -31.08% -43.57% -11.08% -12.43%
    MADD -1.79% -0.99% -49.30% -29.47% -12.23% -14.86%
    PEL -0.10% -0.86% -9.39% -4.25% -7.41% -10.88%
    SFDG (Ours) -0.10% -0.75% -10.10% -3.74% -3.76% -5.12%

    SFDG exhibits minimal performance degradation under all perturbation types, exhibiting the smallest accuracy drop under Gaussian noise on WildDeepfake (−0.75%-0.75\%) and under Gaussian blur on both datasets (−3.76%-3.76\% on FF++, −5.12%-5.12\% on WildDeepfake).

  9. Knowl 9 — Ablation Study of SFDG Core Architectural Components

    data/table

    An ablation study on FaceForensics++ (FF++ LQ) evaluates the relative contribution of each proposed component: the Content-guided Adaptive Frequency Extraction (CAFÉ) module, the Multiple Domains Attention Map Learning (MDAML) module, and the Dynamic Graph Spatial-Frequency Feature Fusion Network (DG-SF3^3Net). Baseline model (a) corresponds to the MADD architecture.

    ID CAFÉ MDAML DG-SF3^3Net Acc (LQ, %) AUC (LQ, %)
    (a) 88.69 90.40
    (b) ✓ 89.52 93.37
    (c) ✓ ✓ 91.48 94.79
    (d) ✓ ✓ 92.05 95.45
    Ours ✓ ✓ ✓ 92.28 95.98
    • Adding CAFÉ to baseline (b vs. a) improves accuracy by +0.83%+0.83\% and AUC by +2.97%+2.97\%.
    • Adding MDAML to CAFÉ (c vs. b) yields +1.96%+1.96\% accuracy and +1.42%+1.42\% AUC gains by supplying multiscale contextual attention.
    • Incorporating DG-SF3^3Net enables high-order relation discovery, with the full model achieving peak performance at 92.28%92.28\% accuracy and 95.98%95.98\% AUC.
  10. Knowl 10 — Sensitivity of DG-SF$^3$Net to Graph Neighborhood Size $k$

    data/table

    The number of nearest neighbors kk used in the dynamic kNNk\text{NN} graph construction inside DG-SF3^3Net governs the receptive scope and aggregation density of graph convolutions. The impact of varying k∈{5,10,15,20}k \in \{5, 10, 15, 20\} was evaluated on FF++ (LQ) and WildDeepfake.

    #Neighbors (kk) FF++ (LQ) WildDeepfake
    Acc (%) AUC (%) Acc (%) AUC (%)
    5 92.10 95.60 82.03 90.57
    10 92.28 95.98 84.41 92.57
    15 92.28 95.49 83.57 91.15
    20 92.02 91.59 81.46 90.38

    Setting k=10k = 10 yields optimal performance on both benchmarks. When k<10k < 10, the sparse connectivity limits cross-domain information exchange; when k>10k > 10, excessive connectivity degrades the regional separability of spatial-frequency attention maps and leads to overfitting.

Coverage note — None was omitted. All principal contributions—the SFDG framework, CAFÉ module, MDAML module, DG-SF3Net dynamic graph reasoning, training objective, intra-dataset benchmark comparisons, cross-dataset generalization experiments, corruption robustness tests, component ablation, and hyperparameter sensitivity analyses—are fully captured.

References

  1. 1.Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen. Mesonet: a compact facial video forgery detection network. In 2018 IEEE International Workshop on Information Forensics and Security, pages 1–7, 2018.
  2. 2.Junyi Cao, Chao Ma, Taiping Yao, Shen Chen, Shouhong Ding, and Xiaokang Yang. End-to-end reconstruction-classification learning for face forgery detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4113–4122, 2022.
  3. 3.Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. Local relation learning for face forgery detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 1081–1088, 2021.
  4. 4.Franc¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1251–1258, 2017.
  5. 5.Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, and Cristian Canton Ferrer. The deepfake detection challenge (dfdc) preview dataset. arXiv preprint arXiv:1910.08854, 2019.
  6. 6.Nick Dufour and Andrew Gully. Contributing data to deepfake detection research. Google AI Blog, 1(3), 2019.
  7. 7.Nguyen et al. Capsule-forensics: Using capsule networks to detect forged images and videos. In 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2307–2311, 2019.
  8. 8.Jessica J. Fridrich and Jan Kodovsky. Rich models for steganalysis of digital images. IEEE Trans. Inf. Forensics Secur., 7(3):868–882, 2012.
  9. 9.Qiqi Gu, Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, and Ran Yi. Exploiting fine-grained face forgery clues via progressive enhancement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 735–743, 2022.
  10. 10.Kai Han, Yunhe Wang, Jianyuan Guo, Yehui Tang, and Enhua Wu. Vision gnn: An image is worth graph of nodes. arXiv preprint arXiv:2206.00272, 2022.
  11. 11.Yinan He, Bei Gan, Siyu Chen, Yichun Zhou, Guojun Yin, Luchuan Song, Lu Sheng, Jing Shao, and Ziwei Liu. Forgerynet: A versatile benchmark for comprehensive forgery analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4360–4369, 2021.
  12. 12.Fa-Ting Hong, Longhao Zhang, Li Shen, and Dan Xu. Depth-aware generative adversarial network for talking head video generation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3397–3406, 2022.
  13. 13.Liming Jiang, Ren Li, Wayne Wu, Chen Qian, and Chen Change Loy. Deeperforensics-1.0: A large-scale dataset for real-world face forgery detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2889–2898, 2020.
  14. 14.Hyunsu Kim, Yunjey Choi, Junho Kim, Sungjoo Yoo, and Youngjung Uh. Exploiting spatial dimensions of latent in GAN for real-time image editing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 852–861, 2021.
  15. 15.Minha Kim, Shahroz Tariq, and Simon S Woo. Fretal: Generalizing deepfake detection using knowledge distillation and representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1001–1012, 2021.
  16. 16.Thomas N. Kipf and Max Welling. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net, 2017.
  17. 17.M Kowalski. Faceswap. https://github.com/marekkowalski/faceswap. Accessed: 2020-08-01, 2018.
  18. 18.Jiaming Li, Hongtao Xie, Jiahong Li, Zhongyuan Wang, and Yongdong Zhang. Frequency-aware discriminative feature learning supervised by single-center loss for face forgery detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6458–6467, 2021.
  19. 19.Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu. Visual semantic reasoning for image-text matching. In Proceedings of the IEEE International Conference on Computer Vision, pages 4654–4662, 2019.
  20. 20.Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo. Face x-ray for more general face forgery detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 5001–5010, 2020.
  21. 21.Yuezun Li, Ming-Ching Chang, and Siwei Lyu. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In 2018 IEEE International Workshop on Information Forensics and Security, pages 1–7, 2018.
  22. 22.Yuezun Li, Xin Yang, Pu Sun, Honggang Qi, and Siwei Lyu. Celeb-df: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3207–3216, 2020.
  23. 23.Lin and Maji. Bilinear cnn models for fine-grained visual recognition. In Proceedings of the IEEE International Conference on Computer Vision, pages 1449–1457, 2015.
  24. 24.Honggu Liu, Xiaodan Li, Wenbo Zhou, Yuefeng Chen, Yuan He, Hui Xue, Weiming Zhang, and Nenghai Yu. Spatial-phase shallow learning: rethinking face forgery detection in frequency domain. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 772–781, 2021.
  25. 25.Yuchen Luo, Yong Zhang, Junchi Yan, and Wei Liu. Generalizing face forgery detection with high-frequency features. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 16317–16326, 2021.
  26. 26.Huy H Nguyen, Fuming Fang, Junichi Yamagishi, and Isao Echizen. Multi-task learning for detecting and segmenting manipulated facial images and videos. In 2019 IEEE 10th International Conference on Biometrics Theory, Applications and Systems, pages 1–8, 2019.
  27. 27.Ivan Perov, Daiheng Gao, Nikolay Chervoniy, Kunlin Liu, Sugasa Marangonda, Chris Ume, Mr Dpfks, Carl Shift Facenheim, et al. Deepfacelab: A simple, flexible and extensible face swapping framework. 2020.
  28. 28.Yuyang Qian, Guojun Yin, Lu Sheng, Zixuan Chen, and Jing Shao. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In Proceedings of the IEEE Conference on European Conference on Computer Vision, pages 86–103. Springer, 2020.
  29. 29.Ronneberger. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-assisted Intervention, pages 234–241. Springer, 2015.
  30. 30.Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE International Conference on Computer Vision, pages 1–11, 2019.
  31. 31.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3):211–252, 2015.
  32. 32.Christos Sagonas, Epameinondas Antonakos, Georgios Tzimiropoulos, Stefanos Zafeiriou, and Maja Pantic. 300 faces in-the-wild challenge: Database and results. Image and Vision Computing, 47:3–18, 2016.
  33. 33.R. R. Selvaraju, A. Das, R. Vedantam, M. Cogswell, D. Parikh, and D. Batra. Grad-cam: Why did you say that? visual explanations from deep networks via gradient-based localization. arXiv e-prints, 2016.
  34. 34.Ke Sun, Hong Liu, Qixiang Ye, Yue Gao, Jianzhuang Liu, Ling Shao, and Rongrong Ji. Domain general face forgery detection by learning to weight. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 2638–2646, 2021.
  35. 35.Ke Sun, Taiping Yao, Shen Chen, Shouhong Ding, Jilin Li, and Rongrong Ji. Dual contrastive learning for general face forgery detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 2316–2324, 2022.
  36. 36.Supasorn Suwajanakorn and Seitz. Synthesizing obama: learning lip sync from audio. ACM Transactions on Graphics (TOG), 36(4):1–13, 2017.
  37. 37.Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning, pages 6105–6114. PMLR, 2019.
  38. 38.Thies. Deferred neural rendering: Image synthesis using neural textures. ACM Transactions on Graphics (TOG), 38(4):1–12, 2019.
  39. 39.Justus Thies, Michael Zollhofer, Marc Stamminger, Christian Theobalt, and Matthias Nießner. Face2face: Real-time face capture and reenactment of rgb videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2387–2395, 2016.
  40. 40.Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, et al. Mlp-mixer: An all-mlp architecture for vision. Advances in Neural Information Processing Systems, 34:24261–24272, 2021.
  41. 41.M Tora. Deepfakes, 2018. https://github.com/deepfakes/faceswap/tree/v2.0.0. Accessed: 2021-03-29.
  42. 42.Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(11), 2008.
  43. 43.Junke Wang, Zuxuan Wu, Wenhao Ouyang, Xintong Han, Jingjing Chen, Yu-Gang Jiang, and Ser-Nam Li. M2tr: Multi-modal multi-scale transformers for deepfake detection. In Proceedings of the 2022 International Conference on Multimedia Retrieval, pages 615–623, 2022.
  44. 44.Wayne Wu, Yunxuan Zhang, Cheng Li, Chen Qian, and Chen Change Loy. Reenactgan: Learning to reenact faces via boundary transfer. In Proceedings of the European Conference on Computer Vision, pages 603–619, 2018.
  45. 45.Xi Wu, Zhen Xie, YuTao Gao, and Yu Xiao. Sstnet: Detecting manipulated faces through spatial, steganalysis and temporal features. In 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2952–2956. IEEE, 2020.
  46. 46.Yanbo Xu, Yueqin Yin, Liming Jiang, Qianyi Wu, Chengyao Zheng, Chen Change Loy, Bo Dai, and Wayne Wu. Transeditor: Transformer-based dual-space gan for highly controllable facial editing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 7683–7692, 2022.
  47. 47.Xin Yang, Yuezun Li, and Siwei Lyu. Exposing deep fakes using inconsistent head poses. In 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 8261–8265, 2019.
  48. 48.Changqian Yu, Jingbo Wang, Chao Peng, Changxin Gao, Gang Yu, and Nong Sang. Bisenet: Bilateral segmentation network for real-time semantic segmentation. In Proceedings of the European Conference on Computer Vision, pages 325–341, 2018.
  49. 49.Yulun Zhang, Chen Fang, Yilin Wang, Zhaowen Wang, Zhe Lin, Yun Fu, and Jimei Yang. Multimodal style transfer via graph cuts. In Proceedings of the IEEE International Conference on Computer Vision, pages 5943–5951, 2019.
  50. 50.Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. Multi-attentional deepfake detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2185–2194, 2021.
  51. 51.Yifan Zhao, Ke Yan, Feiyue Huang, and Jia Li. Graph-based high-order relation discovery for fine-grained recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 15079–15088, 2021.
  52. 52.Bojia Zi, Minghao Chang, Jingjing Chen, Xingjun Ma, and Yu-Gang Jiang. Wilddeepfake: A challenging real-world dataset for deepfake detection. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2382–2390, 2020.

Citation

MLA
Wang, Y., et al. “Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 7278–87, https://doi.org/10.1109/CVPR52729.2023.00703.
APA
Wang, Y., Yu, K., Chen, C., Hu, X., & Peng, S. (2023). Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7278–7287. https://doi.org/10.1109/CVPR52729.2023.00703
Chicago
Wang, Y., K. Yu, C. Chen, X. Hu, and S. Peng. 2023. “Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection”. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7278–87. https://doi.org/10.1109/CVPR52729.2023.00703.
Harvard
Wang, Y. et al. (2023) “Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection”, 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 7278–7287. Available at: https://doi.org/10.1109/CVPR52729.2023.00703.
Vancouver
1. Wang Y, Yu K, Chen C, Hu X, Peng S (2023) Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection. In: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 7278–7287

BibTeX

@inproceedings{Wang_2023, title={Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection}, url={http://dx.doi.org/10.1109/CVPR52729.2023.00703}, DOI={10.1109/cvpr52729.2023.00703}, booktitle={2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Wang, Yuan and Yu, Kun and Chen, Chen and Hu, Xiyuan and Peng, Silong}, year={2023}, month=June, pages={7278–7287} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE