Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning

Chuangchuang TanYao ZhaoShikui WeiGuanghua GuPing LiuYunchao Wei

article2024AAAI357 citations

Presents FreqNet, a compact deepfake detector that learns source-agnostic representations by applying convolutions to both phase and amplitude spectra in the frequency domain, achieving a 9.8% improvement across 17 unseen generation models with only 1.9 million parameters.

Listen

Rapid advances in generative artificial intelligence have made synthetic images, commonly known as deepfakes, increasingly realistic and difficult for humans to detect. While early detection tools identified frequency artifacts produced by image synthesis, newer generative architectures create unique patterns that cause existing detectors to fail on unfamiliar sources. Most state-of-the-art systems overfit to the specific models they were trained on, creating severe security and trust vulnerabilities across media, corporate, and public sectors.

The article demonstrates that shifting detector training directly into the frequency domain significantly improves the generalizability of deepfake detection. The authors introduce FreqNet, a lightweight framework designed to identify forgeries across unseen generative models even when trained on very limited data.

The researchers evaluated their approach through extensive cross-model experiments. The detector was trained on constrained image categories from a single generator and then tested against a comprehensive benchmark of 17 distinct generative models, comprising over 36,000 real-world scene test images and 80,000 face images. FreqNet achieves domain-agnostic detection by combining two high-level mechanisms: extracting high-frequency components from images and internal feature representations across spatial and channel dimensions, and applying specialized frequency convolutional layers to phase and amplitude spectra within the network.

The experimental findings show substantial improvements in both detection accuracy and computational efficiency. Across the 17 tested generative models, FreqNet achieved an overall mean accuracy of 92.8%, outperforming the prior state-of-the-art benchmark by 9.8 percentage points. On self-synthesized test sets from nine distinct models, FreqNet achieved a 94.0% mean accuracy, surpassing the leading baseline by 16.4 percentage points. When applied to high-resolution face images from unseen generators, it maintained high accuracy rates between 98.7% and 99.5%. Crucially, FreqNet achieved these results using only 1.9 million parameters, compared to 304 million parameters required by the top competing baseline.

These results demonstrate that deepfake detectors do not need massive parameter scales or exhaustive multi-generator training to achieve strong cross-model generalization. Operating directly within the frequency spectrum allows detectors to isolate structural forgery indicators rather than model-specific visual artifacts. This approach substantially reduces the compute costs and deployment footprints required for deepfake defense while improving resilience against newly emerging synthesis methods.

Organizations seeking to implement synthetic media moderation should consider lightweight frequency-domain architectures as cost-effective, high-performing options for real-time verification pipelines. Decision-makers should validate FreqNet in pilot deployments on target data streams and explore combining frequency-based plugins with existing enterprise monitoring tools. Further evaluations on heavily compressed or post-processed digital media will help establish operational boundaries before broad deployment.

Cover for Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning

Abstract

This research addresses the challenge of developing a universal deepfake detector that can effectively identify unseen deepfake images despite limited training data. Existing frequency-based paradigms have relied on frequency-level artifacts introduced during the up-sampling in GAN pipelines to detect forgeries. However, the rapid advancements in synthesis technology have led to specific artifacts for each generation model. Consequently, these detectors have exhibited a lack of proficiency in learning the frequency domain and tend to overfit to the artifacts present in the training data, leading to suboptimal performance on unseen sources. To address this issue, we introduce a novel frequency-aware approach called FreqNet, centered around frequency domain learning, specifically designed to enhance the generalizability of deepfake detectors. Our method forces the detector to continuously focus on high-frequency information, exploiting high-frequency representation of features across spatial and channel dimensions. Additionally, we incorporate a straightforward frequency domain learning module to learn source-agnostic features. It involves convolutional layers applied to both the phase spectrum and amplitude spectrum between the Fast Fourier Transform (FFT) and Inverse Fast Fourier Transform (iFFT). Extensive experimentation involving 17 GANs demonstrates the effectiveness of our proposed method, showcasing state-of-the-art performance (+9.8%) while requiring fewer parameters. The code is available at https://github.com/chuangchuangtan/FreqNet-DeepfakeDetection.

Table of Contents

  • Introduction
  • Related Work
  • Image-based Deepfake Detection
  • Frequency-based Deepfake Detection
  • Methodology
  • Problem Definition
  • Overall Architecture
  • Experiments
  • Datasets
  • Implementation Details
  • Deepfake Performance on Real-world Scene
  • Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — FreqNet Architecture and Domain-Invariant Deepfake Detection Framework

    model/method

    FreqNet is a lightweight deepfake detection network designed to generalize to unseen generative adversarial networks (GANs) despite being trained on constrained data. Unlike methods that classify frequency-level image artifacts directly, FreqNet enforces frequency-space feature learning throughout its convolutional backbone (comprising 1.9 million parameters without pretraining).

    Given an RGB input image x∈RW×H×3x \in \mathbb{R}^{W \times H \times 3} and binary label y∈{0,1}y \in \{0, 1\} (00 for real, 11 for fake), the detection framework optimizes a binary cross-entropy loss ℓ(⋅)\ell(\cdot):

    Dfreq=arg⁡min⁡θℓ(Dfreq(xh;θ),y)D_{\text{freq}} = \arg\min_\theta \ell(D_{\text{freq}}(x_h; \theta), y)

    where θ\theta denotes network parameters, and xhx_h is the spatial high-frequency representation of the input image xx.

    The overall architecture sequentially processes features via:

    1. Input Image →\rightarrow High-Frequency Representation of Image (HFRI) module →\rightarrow Convolutional layer;
    2. High-Frequency Representation of Feature across Channel (HFRFC) →\rightarrow Frequency Conv Layer (FCL);
    3. High-Frequency Representation of Feature across Spatial dimensions (HFRFS) →\rightarrow HFRFC →\rightarrow FCL →\rightarrow Residual Convolutional block;
    4. HFRFS →\rightarrow FCL →\rightarrow HFRFS →\rightarrow Two Convolutional layers →\rightarrow FCL →\rightarrow Residual Convolutional block;
    5. Global Average Pooling →\rightarrow Fully Connected (FC) classification layer.
  2. Knowl 2 — High-Frequency Representation of Image (HFRI)

    model/method

    The High-Frequency Representation of Image (HFRI) module isolates high-frequency spatial cues from the input image before feeding it into the classifier backbone, removing low-frequency visual content that can bias detectors toward source-specific image categories.

    For an input image x∈RW×H×3x \in \mathbb{R}^{W \times H \times 3}, 2D Fast Fourier Transform (FFT) F\mathcal{F} converts the image into the frequency domain. The zero-frequency component is centered, and an ideal 2D high-pass filter Bh(⋅)B_h(\cdot) zeros out low-frequency center coefficients:

    fh=Bh(F(x))f_h = B_h(\mathcal{F}(x))

    Bh(fi,j)={fi,j,otherwise0,if ∣i∣<W/4 and ∣j∣<H/4B_h(f_{i,j}) = \begin{cases} f_{i,j}, & \text{otherwise} \\ 0, & \text{if } |i| < W/4 \text{ and } |j| < H/4 \end{cases}

    where (i,j)(i, j) index frequency coordinates with the center of the image spectrum as the origin. The filtered spectrum fh∈RW×H×3f_h \in \mathbb{R}^{W \times H \times 3} is then mapped back to the spatial image space via 2D Inverse Fast Fourier Transform (iFFT) IF\mathcal{IF}:

    xh=IF(fh)x_h = \mathcal{IF}(f_h)

    The resulting filtered image xhx_h serves as the input to the downstream detection layers.

  3. Knowl 3 — High-Frequency Representation of Feature (HFRF) across Spatial and Channel Dimensions

    model/method

    The High-Frequency Representation of Feature (HFRF) module compels intermediate feature representations to retain high-frequency cues, preventing the neural network from overfitting to low-frequency artifacts tied to specific training generators.

    For an intermediate feature tensor Mk∈RH×W×CM^k \in \mathbb{R}^{H \times W \times C} produced by the kk-th convolutional stage, HFRF applies spectral filtering along both spatial dimensions (W,H)(W, H) and the channel dimension CC:

    Mhk(dim)={IFW,H(Bh(FW,H(Mk))),dim=(W,H)IFC(Bh(FC(Mk))),dim=CM^k_h(\text{dim}) = \begin{cases} \mathcal{IF}_{W,H}(B_h(\mathcal{F}_{W,H}(M^k))), & \text{dim} = (W, H) \\ \mathcal{IF}_C(B_h(\mathcal{F}_C(M^k))), & \text{dim} = C \end{cases}

    where:

    • FW,H\mathcal{F}_{W,H} and IFW,H\mathcal{IF}_{W,H} denote the 2D FFT and iFFT along the spatial dimensions (W,H)(W, H).
    • FC\mathcal{F}_C and IFC\mathcal{IF}_C denote the 1D FFT and iFFT along the channel dimension CC.
    • Bh(⋅)B_h(\cdot) represents a centered high-pass filter setting low-frequency components around the zero-frequency center to zero.

    These two component extractors (HFRFS across spatial dimensions and HFRFC across channel dimensions) are placed at different stages of the network to enforce sensitivity to high-frequency variations across both spatial structures and feature channel activations.

  4. Knowl 4 — Frequency Convolutional Layer (FCL) on Phase and Amplitude Spectra

    model/method

    The Frequency Convolutional Layer (FCL) enables direct feature learning in the Fourier domain by applying dedicated convolutional operations to the amplitude and phase spectra independently, learning source-agnostic forgery representations.

    Given intermediate feature maps Mk∈RH×W×CM^k \in \mathbb{R}^{H \times W \times C} from the kk-th layer, 2D FFT FW,H\mathcal{F}_{W,H} converts the features into the complex frequency domain, decomposed into amplitude spectrum famf_{\text{am}} and phase spectrum fphf_{\text{ph}}:

    f=fam+fph=FW,H(Mk)f = f_{\text{am}} + f_{\text{ph}} = \mathcal{F}_{W,H}(M^k)

    Separate convolutional operations LconvL_{\text{conv}} are applied to famf_{\text{am}} and fphf_{\text{ph}}:

    f~am=Lconv(fam)\tilde{f}_{\text{am}} = L_{\text{conv}}(f_{\text{am}})

    f~ph=Lconv(fph)\tilde{f}_{\text{ph}} = L_{\text{conv}}(f_{\text{ph}})

    The transformed spectra are combined and mapped back to the spatial feature domain via 2D iFFT IFW,H\mathcal{IF}_{W,H}:

    M~k=IFW,H(f~am+f~ph)\tilde{M}^k = \mathcal{IF}_{W,H}(\tilde{f}_{\text{am}} + \tilde{f}_{\text{ph}})

    This frequency-domain processing allows the network to learn invariant spectral manipulation cues across diverse generative pipelines.

  5. Knowl 5 — Experimental Setup and Training Protocol for Generalizable Deepfake Detection

    experimental setup

    FreqNet is evaluated under constrained training settings to measure zero-shot generalization across 17 distinct GAN architectures.

    Training Set:

    • Sourced from ForenSynths: ProGAN-generated synthetic images and LSUN real images (18,000 real and 18,000 fake images per category).
    • Three standard training subsets:
      • 1-class: horse
      • 2-class: chair, horse
      • 4-class: car, cat, chair, horse

    Evaluation Sets:

    • ForenSynths Test Set: 8 generative models (ProGAN, StyleGAN, StyleGAN2, BigGAN, CycleGAN, StarGAN, GauGAN, Deepfake) paired with real images from 6 datasets (LSUN, ImageNet, CelebA, CelebA-HQ, COCO, FaceForensics++).
    • Wild/Self-Synthesis Test Set: 9 additional GAN models (AttGAN, BEGAN, CramerGAN, InfoMaxGAN, MMDGAN, RelGAN, S3GAN, SNGAN, STGAN), containing 36,000 images balanced between real and fake.
    • Face Test Set: 20,000 real images from CelebA-HQ and 60,000 fake images from ProGAN, StyleGAN, and StyleGAN2.

    Training Hyperparameters:

    • Optimizer: Adam with initial learning rate 2×10−22 \times 10^{-2}.
    • Learning rate schedule: reduced by 20% every 10 epochs.
    • Batch size: 32; Total epochs: 100.
    • Metrics: Accuracy (Acc., %) and Average Precision (A.P., %).
  6. Knowl 6 — Cross-Model Deepfake Detection Performance on ForenSynths Benchmark

    data/table

    The table below details cross-generator classification accuracy (Acc.) and average precision (A.P.) across 8 unseen generative models on the ForenSynths benchmark, comparing FreqNet with prior image-, frequency-, feature-, and gradient-based detectors under 1-class, 2-class, and 4-class ProGAN training setups.

    Methods Settings ProGAN StyleGAN StyleGAN2 BigGAN CycleGAN StarGAN GauGAN Deepfake Mean
    Input #n Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P.
    Wang (2020) Img 1 50.4 63.8 50.4 79.3 68.2 94.7 50.2 61.3 50.0 52.9 50.0 48.2 50.3 67.6 50.1 51.5 52.5 64.9
    Frank (2020) Freq 1 78.9 77.9 69.4 64.8 67.4 64.0 62.3 58.6 67.4 65.4 60.5 59.5 67.5 69.1 52.4 47.3 65.7 63.3
    Durall (2020) Freq 1 85.1 79.5 59.2 55.2 70.4 63.8 57.0 53.9 66.7 61.4 99.8 99.6 58.7 54.8 53.0 51.9 68.7 65.0
    F3Net (2020) Freq 1 96.9 99.9 86.3 99.8 80.5 99.8 66.6 72.2 76.7 84.0 99.1 100.0 59.1 60.6 61.2 82.3 78.3 87.3
    BiHPF (2022a) Freq 1 82.5 81.4 68.0 62.8 68.8 63.6 67.0 62.5 75.5 74.2 90.1 90.1 73.6 92.1 51.6 49.9 72.1 72.1
    FrePGAN (2022c) Img 1 95.5 99.4 80.6 90.6 77.4 93.0 63.5 60.5 59.4 59.9 99.6 100.0 53.0 49.1 70.4 81.5 74.9 79.3
    LGrad (2023) Grad 1 99.4 99.9 96.0 99.6 93.8 99.4 79.5 88.9 84.7 94.4 99.5 100.0 70.9 81.8 66.7 77.9 86.3 92.7
    Ojha (2023) Fea 1 99.1 100.0 77.2 95.9 69.8 95.8 94.5 99.0 97.1 99.9 98.0 100.0 95.7 100.0 82.4 91.7 89.2 97.8
    FreqNet Freq 1 98.0 99.9 92.0 98.7 89.5 97.9 85.5 93.1 96.1 99.1 94.2 98.4 91.8 99.6 69.8 94.4 89.6 97.6
    Wang (2020) Img 2 64.6 92.7 52.8 82.8 75.7 96.6 51.6 70.5 58.6 81.5 51.2 74.3 53.6 86.6 50.6 51.5 57.3 79.6
    Frank (2020) Freq 2 85.7 81.3 73.1 68.5 75.0 70.9 76.9 70.8 86.5 80.8 85.0 77.0 67.3 65.3 50.1 55.3 75.0 71.2
    Durall (2020) Freq 2 79.0 73.9 63.6 58.8 67.3 62.1 69.5 62.9 65.4 60.8 99.4 99.4 67.0 63.0 50.5 50.2 70.2 66.4
    F3Net (2020) Freq 2 97.9 100.0 84.5 99.5 82.2 99.8 65.5 73.4 81.2 89.7 100.0 100.0 57.0 59.2 59.9 83.0 78.5 88.1
    BiHPF (2022a) Freq 2 87.4 87.4 71.6 74.1 77.0 81.1 82.6 80.6 86.0 86.6 93.8 80.8 75.3 88.2 53.7 54.0 78.4 79.1
    FrePGAN (2022c) Img 2 99.0 99.9 80.8 92.0 72.2 94.0 66.0 61.8 69.1 70.3 98.5 100.0 53.1 51.0 62.2 80.6 75.1 81.2
    LGrad (2023) Grad 2 99.8 100.0 94.8 99.7 92.4 99.6 82.5 92.4 85.9 94.7 99.7 99.9 73.7 83.2 60.6 67.8 86.2 92.2
    Ojha (2023) Fea 2 99.7 100.0 78.8 97.4 75.4 96.7 91.2 99.0 91.9 99.8 96.3 99.9 91.9 100.0 80.0 89.4 88.1 97.8
    FreqNet Freq 2 99.6 100.0 90.4 98.9 85.8 98.1 89.0 96.0 96.7 99.8 97.5 100.0 88.0 98.8 80.7 92.0 91.0 97.9
    Wang (2020) Img 4 91.4 99.4 63.8 91.4 76.4 97.5 52.9 73.3 72.7 88.6 63.8 90.8 63.9 92.2 51.7 62.3 67.1 86.9
    High-Freq Freq 4 98.9 100.0 74.4 98.3 68.8 97.3 75.2 92.1 71.0 87.9 92.7 100.0 75.5 86.5 57.0 74.9 76.7 92.1
    Frank (2020) Freq 4 90.3 85.2 74.5 72.0 73.1 71.4 88.7 86.0 75.5 71.2 99.5 99.5 69.2 77.4 60.7 49.1 78.9 76.5
    Durall (2020) Freq 4 81.1 74.4 54.4 52.6 66.8 62.0 60.1 56.3 69.0 64.0 98.1 98.1 61.9 57.4 50.2 50.0 67.7 64.4
    F3Net (2020) Freq 4 99.4 100.0 92.6 99.7 88.0 99.8 65.3 69.9 76.4 84.3 100.0 100.0 58.1 56.7 63.5 78.8 80.4 86.2
    BiHPF (2022a) Freq 4 90.7 86.2 76.9 75.1 76.2 74.7 84.9 81.7 81.9 78.9 94.4 94.4 69.5 78.1 54.4 54.6 78.6 77.9
    FrePGAN (2022c) Img 4 99.0 99.9 80.7 89.6 84.1 98.6 69.2 71.1 71.1 74.4 99.9 100.0 60.3 71.7 70.9 91.9 79.4 87.2
    LGrad (2023) Grad 4 99.9 100.0 94.8 99.9 96.0 99.9 82.9 90.7 85.3 94.0 99.6 100.0 72.4 79.3 58.0 67.9 86.1 91.5
    Ojha (2023) Fea 4 99.7 100.0 89.0 98.7 83.9 98.4 90.5 99.1 87.9 99.8 91.4 100.0 89.9 100.0 80.2 90.2 89.1 98.3
    FreqNet Freq 4 99.6 100.0 90.2 99.7 88.0 99.5 90.5 96.0 95.8 99.6 85.7 99.8 93.4 98.6 88.9 94.4 91.5 98.5

    Under the 4-class setting, FreqNet achieves a top mean accuracy of 91.5% and mean A.P. of 98.5%, outperforming LGrad (86.1% Acc.) and Ojha (89.1% Acc.).

  7. Knowl 7 — Cross-Model Deepfake Detection Performance on 9-GAN Wild Synthesis Dataset

    data/table

    The table below presents the detection performance (Accuracy and Average Precision in %) on an independent evaluation set consisting of 9 unseen GAN models, evaluated using the 4-class trained models.

    Method AttGAN BEGAN CramerGAN InfoMaxGAN MMDGAN RelGAN S3GAN SNGAN STGAN Mean
    Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P.
    Wang (2020) 51.1 83.7 50.2 44.9 81.5 97.5 71.1 94.7 72.9 94.4 53.3 82.1 55.2 66.1 62.7 90.4 63.0 92.7 62.3 82.9
    F3Net (2020) 85.2 94.8 87.1 97.5 89.5 99.8 67.1 83.1 73.7 99.6 98.8 100.0 65.4 70.0 51.6 93.6 60.3 99.9 75.4 93.1
    LGrad (2023) 68.6 93.8 69.9 89.2 50.3 54.0 71.1 82.0 57.5 67.3 89.1 99.1 78.5 86.0 78.0 87.4 54.8 68.0 68.6 80.8
    Ojha (2023) 78.5 98.3 72.0 98.9 77.6 99.8 77.6 98.9 77.6 99.7 78.2 98.7 85.2 98.1 77.6 98.7 74.2 97.8 77.6 98.8
    FreqNet 89.8 98.8 98.8 100.0 95.2 98.2 94.5 97.3 95.2 98.2 100.0 100.0 88.3 94.3 85.4 90.5 98.8 100.0 94.0 97.5

    On these 9 wild-scene GAN models, FreqNet achieves a mean accuracy of 94.0%, which is a 16.4% absolute improvement over Ojha (2023) (77.6% Acc.) and an 18.6% improvement over F3Net (75.4% Acc.).

  8. Knowl 8 — Model Complexity and 17-GAN Mean Accuracy Comparison

    data/table

    The table compares parameter count and overall mean classification accuracy (mAcc) across all 17 evaluated generative adversarial networks (8 from ForenSynths and 9 from the self-synthesis dataset).

    Methods Parameters ↓\downarrow mAcc. ↑\uparrow of 17 models
    F3Net (2020) 48.9 M 77.8%
    LGrad (2023) 46.6 M 76.8%
    Ojha (2023) 304.0 M 83.0%
    FreqNet 1.9 M 92.8% (+9.8%)

    FreqNet requires only 1.9 million parameters (a ∼\sim160×\times parameter reduction relative to the 304.0M parameter model of Ojha et al., 2023) while achieving 92.8% mean accuracy across 17 GANs, outperforming Ojha et al. by 9.8% absolute accuracy.

  9. Knowl 9 — Ablation Study of FreqNet Architectural Components

    data/table

    The contribution of each frequency-domain component in FreqNet is evaluated on the ForenSynths dataset (mean accuracy in % across test generators).

    HFRI HFRFS HFRFC FCL Mean Acc. (%)
    ✓ ✓ ✓ 84.3
    ✓ ✓ ✓ 85.3
    ✓ ✓ ✓ 87.8
    ✓ ✓ 82.0
    ✓ ✓ ✓ 83.8
    ✓ ✓ ✓ ✓ 91.5

    Where:

    • HFRI: High-Frequency Representation of Image input module.
    • HFRFS: High-Frequency Representation of Feature across Spatial dimensions (W,H)(W, H).
    • HFRFC: High-Frequency Representation of Feature across Channel dimension CC.
    • FCL: Frequency Convolutional Layer on amplitude and phase spectra.

    Removing any single module degrades detection accuracy: omitting HFRI reduces mean accuracy from 91.5% to 84.3% (-7.2%), omitting FCL reduces it to 83.8% (-7.7%), and removing both spatial and channel HFRF modules drops accuracy to 82.0% (-9.5%).

  10. Knowl 10 — Class Activation Mapping Behavior and Cross-Category Generalization to Face Images

    empirical result

    Class Activation Map (CAM) analysis shows distinct activation behaviors between authentic and synthesized images:

    • For real images, the CAMs highlight broad, holistic image regions.
    • For GAN-generated fake images, the CAMs concentrate on localized artifact regions.

    Furthermore, even when trained exclusively on non-facial object categories (ProGAN classes: car, cat, chair, horse), FreqNet transfers effectively to facial deepfake benchmarks, achieving classification accuracies of:

    • 98.7% on ProGAN faces,
    • 99.0% on StyleGAN faces,
    • 99.5% on StyleGAN2 faces, evaluated against 20,000 real CelebA-HQ faces and 60,000 synthetic faces.

Coverage note — No substantial contributed material was omitted. All architectural modules, equations, benchmark comparisons, parameter analyses, ablation results, and CAM insights are fully covered.

References

  1. 1.Bellemare, M. G.; et al. 2017. The cramer distance as a solution to biased wasserstein gradients. arXiv preprint arXiv:1705.10743.
  2. 2.Berthelot, D.; et al. 2017. Began: Boundary equilibrium generative adversarial networks. arXiv preprint arXiv:1703.10717.
  3. 3.Brock, A.; et al. 2018. Large Scale GAN Training for High Fidelity Natural Image Synthesis. In International Conference on Learning Representations.
  4. 4.Cao, J.; et al. 2022. End-to-end reconstruction-classification learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4113–4122.
  5. 5.Chai, L.; et al. 2020. What makes fake images detectable? understanding properties that generalize. In European conference on computer vision, 103–120. Springer.
  6. 6.Chen, L.; et al. 2022. Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 18710–18719.
  7. 7.Choi, Y.; et al. 2018. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In Proceedings of the IEEE conference on computer vision and pattern recognition, 8789–8797.
  8. 8.Chollet, F. 2017. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1251–1258.
  9. 9.Cooley, J. W.; et al. 1969. The fast Fourier transform and its applications. IEEE Transactions on Education, 12(1): 27–34.
  10. 10.Durall, R.; et al. 2020. Watch your up-convolution: Cnn based generative deep neural networks are failing to reproduce spectral distributions. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 7890–7899.
  11. 11.Frank, J.; et al. 2020. Leveraging frequency analysis for deep fake image recognition. In International conference on machine learning, 3247–3258. PMLR.
  12. 12.Goodfellow, I. J.; et al. 2014. Generative Adversarial Nets. In NIPS.
  13. 13.Haliassos, A.; et al. 2021. Lips don’t lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 5039–5049.
  14. 14.He, Y.; et al. 2021. Beyond the Spectrum: Detecting Deepfakes via Re-Synthesis. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, 2534–2541. International Joint Conferences on Artificial Intelligence Organization.
  15. 15.He, Z.; et al. 2019. AttGAN: Facial Attribute Editing by Only Changing What You Want. IEEE Transactions on Image Processing, 28(11): 5464–5478.
  16. 16.Jeong, Y.; et al. 2022a. BiHPF: Bilateral High-Pass Filters for Robust Deepfake Detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 48–57.
  17. 17.Jeong, Y.; et al. 2022b. FingerprintNet: Synthesized Fingerprints for Generated Image Detection. In European Conference on Computer Vision, 76–94. Springer.
  18. 18.Jeong, Y.; et al. 2022c. FrePGAN: robust deepfake detection using frequency-level perturbations. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 1060–1068.
  19. 19.Karras, T.; et al. 2018. Progressive Growing of GANs for Improved Quality, Stability, and Variation. In International Conference on Learning Representations.
  20. 20.Karras, T.; et al. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4401–4410.
  21. 21.Karras, T.; et al. 2020. Analyzing and improving the image quality of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8110–8119.
  22. 22.Kingma, D. P.; et al. 2015. Adam: A Method for Stochastic Optimization. In ICLR (Poster).
  23. 23.Lee, K. S.; et al. 2021. Infomax-gan: Improved adversarial image generation via information maximization and contrastive learning. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, 3942–3952.
  24. 24.Li, C.; Huang, Z.; Paudel, D. P.; Wang, Y.; Shahbazi, M.; Hong, X.; and Van Gool, L. 2023. A continual deepfake detection benchmark: Dataset, methods, and essentials. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 1339–1349.
  25. 25.Li, C.-L.; et al. 2017. Mmd gan: Towards deeper understanding of moment matching network. Advances in neural information processing systems, 30.
  26. 26.Li, J.; et al. 2021. Frequency-aware discriminative feature learning supervised by single-center loss for face forgery detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 6458–6467.
  27. 27.Li, Y.; et al. 2018. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In 2018 IEEE International workshop on information forensics and security (WIFS), 1–7. IEEE.
  28. 28.Lin, T.-Y.; et al. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, 740–755. Springer.
  29. 29.Liu, M.; et al. 2019. Stgan: A unified selective transfer network for arbitrary image attribute editing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 3673–3682.
  30. 30.Liu, Z.; et al. 2015. Deep learning face attributes in the wild. In Proceedings of the IEEE international conference on computer vision, 3730–3738.
  31. 31.Lučić, M.; et al. 2019. High-fidelity image generation with fewer labels. In International conference on machine learning, 4183–4192. PMLR.
  32. 32.Luo, Y.; et al. 2021. Generalizing face forgery detection with high-frequency features. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 16317–16326.
  33. 33.Masi, I.; et al. 2020. Two-branch recurrent network for isolating deepfakes in videos. In European conference on computer vision, 667–684. Springer.
  34. 34.Miyato, T.; et al. 2018. Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957.
  35. 35.Nie, W.; et al. 2019. Relgan: Relational generative adversarial networks for text generation. In International conference on learning representations.
  36. 36.Ojha, U.; et al. 2023. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 24480–24489.
  37. 37.Park, T.; et al. 2019. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2337–2346.
  38. 38.Paszke, A.; et al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems, 32.
  39. 39.Qian, Y.; et al. 2020. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In European conference on computer vision, 86–103. Springer.
  40. 40.Rossler, A.; et al. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF international conference on computer vision, 1–11.
  41. 41.Russakovsky, O.; et al. 2015. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3): 211–252.
  42. 42.Shiohara, K.; et al. 2022. Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18720–18729.
  43. 43.Tan, C.; et al. 2023. Learning on Gradients: Generalized Artifacts Representation for GAN-Generated Images Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 12105–12114.
  44. 44.Wang, C.; et al. 2021. Representative forgery mining for fake face detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 14923–14932.
  45. 45.Wang, S.-Y.; et al. 2020. CNN-generated images are surprisingly easy to spot... for now. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8695–8704.
  46. 46.Wang, Y.; et al. 2023. Dynamic Graph Learning With Content-Guided Spatial-Frequency Relation Reasoning for Deepfake Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 7278–7287.
  47. 47.Woo, S.; et al. 2022. ADD: Frequency Attention and Multi-View Based Knowledge Distillation to Detect Low-Quality Compressed Deepfake Images. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, 122–130.
  48. 48.Yu, F.; et al. 2015. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365.
  49. 49.Yu, Y.; et al. 2020. Mining generalized features for detecting ai-manipulated fake faces. arXiv preprint arXiv:2010.14129.
  50. 50.Zhou, B.; et al. 2016. Learning deep features for discriminative localization. In Proceedings of the IEEE conference on computer vision and pattern recognition, 2921–2929.
  51. 51.Zhu, J.-Y.; et al. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE international conference on computer vision, 2223–2232.

Citation

MLA
Tan, C., et al. “Frequency-Aware Deepfake Detection: Improving Generalizability Through Frequency Space Learning”. arXiv, 2024, http://arxiv.org/abs/2403.07240v1.
APA
Tan, C., Zhao, Y., Wei, S., Gu, G., Liu, P., & Wei, Y. (2024). Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Learning. arXiv. http://arxiv.org/abs/2403.07240v1
Chicago
Tan, C., Y. Zhao, S. Wei, G. Gu, P. Liu, and Y. Wei. 2024. “Frequency-Aware Deepfake Detection: Improving Generalizability Through Frequency Space Learning”. arXiv. http://arxiv.org/abs/2403.07240v1.
Harvard
Tan, C. et al. (2024) “Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.07240v1.
Vancouver
1. Tan C, Zhao Y, Wei S, Gu G, Liu P, Wei Y (2024) Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Learning. arXiv

BibTeX

@article{tan2024frequency,
  title = {Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Learning},
  author = {Tan, Chuangchuang and Zhao, Yao and Wei, Shikui and Gu, Guanghua and Liu, Ping and Wei, Yunchao},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.07240v1},
  eprint = {2403.07240}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF