DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection

Zhiyuan YanYong ZhangXinhang YuanSiwei LyuBaoyuan Wu

article2023NeurIPS196 citations

Presents DeepfakeBench, the first open-source standardized benchmark for deepfake detection, unifying data pipelines, training frameworks, and evaluation protocols across 15 state-of-the-art detectors and 9 datasets to enable truly fair and reproducible performance comparisons.

Listen

Rapid advances in facial manipulation technology have raised severe risks of disinformation, privacy violations, and declining trust in digital media. While researchers have developed numerous deepfake detection tools, the field has suffered from a lack of standardized data pipelines, disparate experimental protocols, and unreleased source code. These inconsistencies prevent reliable performance comparisons and can lead to misleading conclusions about model effectiveness. The article introduces DeepfakeBench, the first unified and extensible benchmarking platform designed to establish standardized, transparent evaluations of deepfake detection methods.

The article evaluates 15 state-of-the-art detection algorithms across nine prominent deepfake datasets using consistent frame sampling, image alignment, training configurations, and standardized evaluation metrics. The tested tools span three primary categories: simple classifiers, spatial detectors that focus on localized visual artifacts, and frequency-based detectors that examine spectral distributions. Models were trained under uniform conditions and assessed for within-domain accuracy, cross-dataset transferability, and resilience across diverse manipulation techniques.

The analysis yielded several critical findings regarding real-world model effectiveness. First, most models perform well on familiar data—achieving within-domain accuracy scores above 94%—but experience sharp performance drops of 15% to 35% when evaluated against unfamiliar datasets or novel manipulation methods. Second, simple standard classifiers matched or nearly matched the performance of highly complex specialized detectors once training conditions, such as pre-training and data augmentations, were held constant. Third, underlying architectural choices proved crucial: network backbones incorporating depthwise separable convolutions consistently outperformed alternatives like standard residual networks of comparable parameter size. Fourth, standard training enhancements such as data augmentations presented trade-offs, sometimes improving general robustness against video compression while degrading performance on fine-grained frequency artifacts.

These findings indicate that many perceived performance advantages in previous research stemmed from inconsistent experimental setups rather than genuine algorithmic superiority. For organizations deploying forensic defenses, relying on specialized models trained on limited scenarios poses a significant risk of detection failure in the wild. Real-world systems require broad generalization capabilities rather than optimization for specific manipulation artifacts. Consequently, decision-makers should avoid over-investing in complex architectures when simpler, well-tuned models on robust backbones achieve comparable results.

To advance effective deepfake defense, practitioners and researchers should adopt standardized benchmarks for model validation, prioritize network backbones featuring separable convolutions, and incorporate phase-spectrum features into blending detectors to improve cross-dataset generalization. Current evaluations remain subject to boundaries, as testing was conducted at the frame level rather than across full video sequences, and focused primarily on face-swapping techniques. Future initiatives must expand into temporal video analysis and address emerging generative image formats, such as diffusion-generated content, to maintain robust protection against evolving digital threats.

Cover for DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection

Abstract

A critical yet frequently overlooked challenge in the field of deepfake detection is the lack of a standardized, unified, comprehensive benchmark. This issue leads to unfair performance comparisons and potentially misleading results. Specifically, there is a lack of uniformity in data processing pipelines, resulting in inconsistent data inputs for detection models. Additionally, there are noticeable differences in experimental settings, and evaluation strategies and metrics lack standardization. To fill this gap, we present the first comprehensive benchmark for deepfake detection, called DeepfakeBench, which offers three key contributions: 1) a unified data management system to ensure consistent input across all detectors, 2) an integrated framework for state-of-the-art methods implementation, and 3) standardized evaluation metrics and protocols to promote transparency and reproducibility. Featuring an extensible, modular-based codebase, DeepfakeBench contains 15 state-of-the-art detection methods, 9 deepfake datasets, a series of deepfake detection evaluation protocols and analysis tools, as well as comprehensive evaluations. Moreover, we provide new insights based on extensive analysis of these evaluations from various perspectives (e.g., data augmentations, backbones). We hope that our efforts could facilitate future research and foster innovation in this increasingly critical domain. All codes, evaluations, and analyses of our benchmark are publicly available at https://github.com/SCLBD/DeepfakeBench.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Our Benchmark
  • 3.1 Datasets and Detectors
  • 3.2 Codebase
  • 4 Evaluations and Analysis
  • 4.1 Experimental Setup
  • 4.2 Evaluations
  • 4.3 Analysis
  • 5 Conclusions, Future Plans, and Societal Impacts
  • 6 Contents in Appendix
  • 7 Acknowledgement
  • References
  • A Appendix
  • A.1 Details of Data Processing
  • A.2 Details of Algorithms Implementation and Visualizations
  • A.3 Training Details and Full Experimental Results
  • A.4 Other Analysis Results

Citation

MLA
Yan, Z., et al. “DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection”. arXiv, 2023, http://arxiv.org/abs/2307.01426v2.
APA
Yan, Z., Zhang, Y., Yuan, X., Lyu, S., & Wu, B. (2023). DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. arXiv. http://arxiv.org/abs/2307.01426v2
Chicago
Yan, Z., Y. Zhang, X. Yuan, S. Lyu, and B. Wu. 2023. “DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection”. arXiv. http://arxiv.org/abs/2307.01426v2.
Harvard
Yan, Z. et al. (2023) “DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2307.01426v2.
Vancouver
1. Yan Z, Zhang Y, Yuan X, Lyu S, Wu B (2023) DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. arXiv

BibTeX

@article{yan2023deepfakebench,
  title = {DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection},
  author = {Yan, Zhiyuan and Zhang, Yong and Yuan, Xinhang and Lyu, Siwei and Wu, Baoyuan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2307.01426v2},
  eprint = {2307.01426}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors