DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection
Zhiyuan YanYong ZhangXinhang YuanSiwei LyuBaoyuan Wu
Presents DeepfakeBench, the first open-source standardized benchmark for deepfake detection, unifying data pipelines, training frameworks, and evaluation protocols across 15 state-of-the-art detectors and 9 datasets to enable truly fair and reproducible performance comparisons.
Rapid advances in facial manipulation technology have raised severe risks of disinformation, privacy violations, and declining trust in digital media. While researchers have developed numerous deepfake detection tools, the field has suffered from a lack of standardized data pipelines, disparate experimental protocols, and unreleased source code. These inconsistencies prevent reliable performance comparisons and can lead to misleading conclusions about model effectiveness. The article introduces DeepfakeBench, the first unified and extensible benchmarking platform designed to establish standardized, transparent evaluations of deepfake detection methods.
The article evaluates 15 state-of-the-art detection algorithms across nine prominent deepfake datasets using consistent frame sampling, image alignment, training configurations, and standardized evaluation metrics. The tested tools span three primary categories: simple classifiers, spatial detectors that focus on localized visual artifacts, and frequency-based detectors that examine spectral distributions. Models were trained under uniform conditions and assessed for within-domain accuracy, cross-dataset transferability, and resilience across diverse manipulation techniques.
The analysis yielded several critical findings regarding real-world model effectiveness. First, most models perform well on familiar data—achieving within-domain accuracy scores above 94%—but experience sharp performance drops of 15% to 35% when evaluated against unfamiliar datasets or novel manipulation methods. Second, simple standard classifiers matched or nearly matched the performance of highly complex specialized detectors once training conditions, such as pre-training and data augmentations, were held constant. Third, underlying architectural choices proved crucial: network backbones incorporating depthwise separable convolutions consistently outperformed alternatives like standard residual networks of comparable parameter size. Fourth, standard training enhancements such as data augmentations presented trade-offs, sometimes improving general robustness against video compression while degrading performance on fine-grained frequency artifacts.
These findings indicate that many perceived performance advantages in previous research stemmed from inconsistent experimental setups rather than genuine algorithmic superiority. For organizations deploying forensic defenses, relying on specialized models trained on limited scenarios poses a significant risk of detection failure in the wild. Real-world systems require broad generalization capabilities rather than optimization for specific manipulation artifacts. Consequently, decision-makers should avoid over-investing in complex architectures when simpler, well-tuned models on robust backbones achieve comparable results.
To advance effective deepfake defense, practitioners and researchers should adopt standardized benchmarks for model validation, prioritize network backbones featuring separable convolutions, and incorporate phase-spectrum features into blending detectors to improve cross-dataset generalization. Current evaluations remain subject to boundaries, as testing was conducted at the frame level rather than across full video sequences, and focused primarily on face-swapping techniques. Future initiatives must expand into temporal video analysis and address emerging generative image formats, such as diffusion-generated content, to maintain robust protection against evolving digital threats.
- Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). Introduces the seminal FaceForensics++ benchmark and standard forensic evaluation protocols that serve as foundational datasets and baselines directly evaluated within DeepfakeBench.
- Paper: Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics, Yuezun Li et al. (2019). Presents Celeb-DF, one of the key high-quality face-swapping datasets systematically integrated into DeepfakeBench's unified cross-dataset benchmark.
- Paper: MesoNet: a Compact Facial Video Forgery Detection Network, Darius Afchar et al. (2018). Introduces MesoNet, a classic compact baseline architecture whose frame-level forensic detection capabilities and efficiency are benchmarked in DeepfakeBench.
- Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). Demonstrates how standard CNN classifiers trained with data augmentations can generalize across synthetic images, motivating DeepfakeBench's findings on standard classifiers versus complex detectors.
- Paper: End-to-End Reconstruction-Classification Learning for Face Forgery Detection, Junyi Cao et al. (2022). Introduces RECCE, an end-to-end reconstruction-classification detector designed for unseen face forgery generalizability that is representative of the specialized methods analyzed in DeepfakeBench.
- Paper: Delving into Sequential Patches for Deepfake Detection, Jiazhi Guan et al. (2022). Presents a transformer-based sequential patch detection model, representing the category of spatial and patch-level detectors evaluated under uniform conditions in DeepfakeBench.
- Paper: Exploiting Fine-Grained Face Forgery Clues via Progressive Enhancement Learning, Qiqi Gu et al. (2022). Introduces progressive spatial-frequency enhancement learning for face forgery detection, establishing techniques directly relevant to DeepfakeBench's frequency-based model evaluations.
- Paper: Protecting Celebrities from DeepFake with Identity Consistency Transformer, Xiaoyi Dong et al. (2022). Proposes the Identity Consistency Transformer for high-level semantic face forgery detection, exemplifying the advanced semantic approaches examined in standardized benchmark settings.
- Paper: Dynamic Graph Learning with Content-guided Spatial-Frequency Relation Reasoning for Deepfake Detection, Yuan Wang et al. (2023). Extends spatial and frequency-based forensic analysis by using dynamic graph learning to jointly reason over spatial-frequency interactions across standardized benchmarks.
- Paper: Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning, Chuangchuang Tan et al. (2024). Applies domain-agnostic frequency space learning (FreqNet) to advance the generalizability bottlenecks across unseen generative models identified in DeepfakeBench.
- Paper: Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection, Chuangchuang Tan et al. (2024). Develops a generalizable deepfake detector focused on up-sampling spatial artifacts to address the cross-generator degradation highlighted by unified benchmark evaluations.
- Paper: AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection, Trevine Oorloff et al. (2024). Expands detection beyond the frame-level visual evaluations of DeepfakeBench by incorporating multi-modal audio-visual synchronization and feature fusion.
- Paper: Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images, Zeyu Lu et al. (2023). Contrasts model detection efficacy with human perceptual limits across modern generative datasets, answering DeepfakeBench's call to explore newer generative architectures.
- Paper: Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks, Mehrdad Saberi et al. (2024). Investigates the adversarial vulnerabilities and fundamental robustness limits of deepfake detectors and watermarks under realistic evasion attacks.
- Paper: Hierarchical Fine-Grained Image Forgery Detection and Localization, Xiao Guo et al. (2023). Builds upon unified detection evaluations by introducing a hierarchical multi-branch architecture for simultaneous detection, pixel-level localization, and generator attribution.
