Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection
Chuangchuang TanHuan LiuYao ZhaoShikui WeiGuanghua GuPing LiuYunchao Wei
Reveals that up-sampling operations in generative networks induce structural local pixel dependencies, introducing a Neighboring Pixel Relationships method that improves generalizable deepfake detection across 28 generative architectures by 11.6%.
The rapid advancement of artificial intelligence image generation tools poses growing security, political, and economic risks. Existing detection systems struggle to identify synthetic images created by newer or unfamiliar generation tools because detection models often overfit to specific training sources. Current detection techniques frequently analyze entire images in the frequency domain to identify artificial signatures, but these patterns vary widely across different generators and fail to provide reliable, universal detection.
The article demonstrates that examining local pixel relationships introduced during the image generation process provides a universal signature for detecting synthetic images. Specifically, the article evaluates whether focusing on spatial artifacts left by standard up-sampling operations—a process used by virtually all generative systems to scale up low-resolution features—can accurately identify deepfakes across diverse, unseen generative models.
To evaluate this approach, the authors developed a representation method called Neighboring Pixel Relationships, which captures local relative pixel differences within small image patches. The authors trained a lightweight classifier on synthetic images from a single generative model (ProGAN) across four object categories (cars, cats, chairs, and horses) and evaluated its detection performance across five benchmark datasets comprising 28 distinct generative models, including both generative adversarial networks and modern diffusion architectures.
The evaluation yielded several key findings. First, the proposed approach achieved a mean accuracy of 92.2% across 38 sub-test sets encompassing all 28 generative models, outperforming the best existing baseline methods by approximately 11.6% to 12.4%. Second, the detector maintained an average accuracy of 93.2% across 9 previously unseen generative adversarial network models. Third, despite training only on an older generative adversarial network, the system achieved a 95.3% mean accuracy on diffusion models in standard benchmarks and 80.1% accuracy on high-step diffusion models such as Midjourney and DALL-E, demonstrating strong cross-architecture transferability. Finally, hyperparameter tests confirmed that local two-by-two pixel patch analysis aligned best with standard up-sampling configurations across generative pipelines.
These findings indicate that generative image systems leave consistent, exploitable structural traces at the local pixel level regardless of overall image realism or generation framework. For organizations managing digital media integrity, fraud risks, and regulatory compliance, this local representation method offers a cost-effective, high-performing detection mechanism that does not require continuous retraining on every emerging generative tool.
Organizations and developers seeking to enhance synthetic media detection should consider adopting local spatial relationship representations as a core detection feature. When implementing this framework, utilizing two-by-two spatial grids provides the most consistent baseline performance across diverse image sources. Before operational deployment, teams should conduct internal evaluations on image streams subjected to heavy social media compression or downstream editing, as real-world degradation may impact fine-grained pixel traces.
- Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). This paper establishes the foundational benchmark demonstrating that CNN-based generative models leave universal, cross-model artifacts that facilitate generalizable synthetic image detection.
- Paper: Alias-Free Generative Adversarial Networks, Tero Karras et al. (2021). This work analyzes how standard upsampling and convolutional operations introduce aliasing and structural artifacts in CNN generators, providing key architectural context for understanding generative upsampling flaws.
- Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). This study establishes the standard benchmarking methodology and evaluation protocols for detecting manipulated and synthetic facial imagery across diverse forgery pipelines.
- Paper: Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics, Yuezun Li et al. (2019). This paper introduces a high-quality deepfake forensic benchmark that highlights the challenge of detecting subtle, realistic synthesis artifacts beyond standard low-level visual flaws.
- Paper: Analyzing and Improving the Image Quality of StyleGAN, Tero Karras et al. (2020). This paper examines the architectural origins of characteristic localized artifacts in deep generative networks, directly informing the structural analysis of CNN generator components.
No sufficiently relevant recommendations were found.
