Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection

Chuangchuang TanHuan LiuYao ZhaoShikui WeiGuanghua GuPing LiuYunchao Wei

article2024CVPR431 citations

Reveals that up-sampling operations in generative networks induce structural local pixel dependencies, introducing a Neighboring Pixel Relationships method that improves generalizable deepfake detection across 28 generative architectures by 11.6%.

Listen

The rapid advancement of artificial intelligence image generation tools poses growing security, political, and economic risks. Existing detection systems struggle to identify synthetic images created by newer or unfamiliar generation tools because detection models often overfit to specific training sources. Current detection techniques frequently analyze entire images in the frequency domain to identify artificial signatures, but these patterns vary widely across different generators and fail to provide reliable, universal detection.

The article demonstrates that examining local pixel relationships introduced during the image generation process provides a universal signature for detecting synthetic images. Specifically, the article evaluates whether focusing on spatial artifacts left by standard up-sampling operations—a process used by virtually all generative systems to scale up low-resolution features—can accurately identify deepfakes across diverse, unseen generative models.

To evaluate this approach, the authors developed a representation method called Neighboring Pixel Relationships, which captures local relative pixel differences within small image patches. The authors trained a lightweight classifier on synthetic images from a single generative model (ProGAN) across four object categories (cars, cats, chairs, and horses) and evaluated its detection performance across five benchmark datasets comprising 28 distinct generative models, including both generative adversarial networks and modern diffusion architectures.

The evaluation yielded several key findings. First, the proposed approach achieved a mean accuracy of 92.2% across 38 sub-test sets encompassing all 28 generative models, outperforming the best existing baseline methods by approximately 11.6% to 12.4%. Second, the detector maintained an average accuracy of 93.2% across 9 previously unseen generative adversarial network models. Third, despite training only on an older generative adversarial network, the system achieved a 95.3% mean accuracy on diffusion models in standard benchmarks and 80.1% accuracy on high-step diffusion models such as Midjourney and DALL-E, demonstrating strong cross-architecture transferability. Finally, hyperparameter tests confirmed that local two-by-two pixel patch analysis aligned best with standard up-sampling configurations across generative pipelines.

These findings indicate that generative image systems leave consistent, exploitable structural traces at the local pixel level regardless of overall image realism or generation framework. For organizations managing digital media integrity, fraud risks, and regulatory compliance, this local representation method offers a cost-effective, high-performing detection mechanism that does not require continuous retraining on every emerging generative tool.

Organizations and developers seeking to enhance synthetic media detection should consider adopting local spatial relationship representations as a core detection feature. When implementing this framework, utilizing two-by-two spatial grids provides the most consistent baseline performance across diverse image sources. Before operational deployment, teams should conduct internal evaluations on image streams subjected to heavy social media compression or downstream editing, as real-world degradation may impact fine-grained pixel traces.

  • Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). This paper establishes the foundational benchmark demonstrating that CNN-based generative models leave universal, cross-model artifacts that facilitate generalizable synthetic image detection.
  • Paper: Alias-Free Generative Adversarial Networks, Tero Karras et al. (2021). This work analyzes how standard upsampling and convolutional operations introduce aliasing and structural artifacts in CNN generators, providing key architectural context for understanding generative upsampling flaws.
  • Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). This study establishes the standard benchmarking methodology and evaluation protocols for detecting manipulated and synthetic facial imagery across diverse forgery pipelines.
  • Paper: Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics, Yuezun Li et al. (2019). This paper introduces a high-quality deepfake forensic benchmark that highlights the challenge of detecting subtle, realistic synthesis artifacts beyond standard low-level visual flaws.
  • Paper: Analyzing and Improving the Image Quality of StyleGAN, Tero Karras et al. (2020). This paper examines the architectural origins of characteristic localized artifacts in deep generative networks, directly informing the structural analysis of CNN generator components.

No sufficiently relevant recommendations were found.

Cover for Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection

Abstract

Recently, the proliferation of highly realistic synthetic images, facilitated through a variety of GANs and Diffusions, has significantly heightened the susceptibility to misuse. While the primary focus of deepfake detection has traditionally centered on the design of detection algorithms, an investigative inquiry into the generator architectures has remained conspicuously absent in recent years. This paper contributes to this lacuna by rethinking the architectures of CNN-based generators, thereby establishing a generalized representation of synthetic artifacts. Our findings illuminate that the up-sampling operator can, beyond frequency-based artifacts, produce generalized forgery artifacts. In particular, the local interdependence among image pixels caused by upsampling operators is significantly demonstrated in synthetic images generated by GAN or diffusion. Building upon this observation, we introduce the concept of Neighboring Pixel Relationships(NPR) as a means to capture and characterize the generalized structural artifacts stemming from up-sampling operations. A comprehensive analysis is conducted on an open-world dataset, comprising samples generated by \tft{28 distinct generative models}. This analysis culminates in the establishment of a novel state-of-the-art performance, showcasing a remarkable \tft{11.6%} improvement over existing methods. The code is available at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Image-based Fake Detection
  • 2.2 Frequency-based Fake Detection
  • 3 Methodology
  • 3.1 Problem setup
  • 3.2 Up-sampling operations in generator pipeline
  • 3.3 Neighboring pixel relationships
  • 4 Experiments
  • 4.1 Settings
  • 4.2 Generalization capability evaluation
  • 4.2.1 GAN-Sources Evaluation
  • 4.2.2 Diffusion-Sources Evaluation
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — Neighboring Pixel Relationships (NPR) Representation

    model/method

    Neighboring Pixel Relationships (NPR) is an artifact representation designed to capture local, spatial up-sampling signatures left by CNN-based generator pipelines. Given an input RGB image I∈R(l⋅W)×(l⋅H)×3I \in \mathbb{R}^{(l \cdot W) \times (l \cdot H) \times 3} produced after an up-sampling layer with scale factor ll and subsequent convolutional layers, the image is partitioned into W×HW \times H non-overlapping spatial grids, where each grid vIcv_I^c consists of an l×ll \times l patch of pixels:

    vIc={w1,w2,…,wn},n=l×lv_I^c = \{w_1, w_2, \dots, w_n\}, \quad n = l \times l

    For each grid vIcv_I^c, the relative neighboring pixel relationship v^Ic\hat{v}_I^c is computed by subtracting a chosen reference pixel wjw_j (1≤j≤n1 \le j \le n) from every pixel in the grid:

    v^Ic={w1−wj, w2−wj, …, wn−wj}\hat{v}_I^c = \{w_1 - w_j, \, w_2 - w_j, \, \dots, \, w_n - w_j\}

    The NPR representation of the entire image is the full spatial assembly of all transformed grid vectors v^Ic\hat{v}_I^c. In the standard configuration, the scaling factor is set to l=2l = 2 (2×22 \times 2 pixel patches, n=4n = 4) and the reference anchor is chosen as the first element j=1j = 1 (the top-left pixel of the patch).

  2. Knowl 2 — Mechanism of Spatial Up-Sampling Artifacts in Generative Networks

    theoretical result

    In CNN-based generative models (both GANs and U-Net-based diffusion models), up-sampling operations transform low-resolution latent feature representations x∈RW×H×Cx \in \mathbb{R}^{W \times H \times C} into higher-resolution representations x^=up(x)∈R(l⋅W)×(l⋅H)×C\hat{x} = \text{up}(x) \in \mathbb{R}^{(l \cdot W) \times (l \cdot H) \times C} prior to final image synthesis through convolutional layers I=conv(x^)I = \text{conv}(\hat{x}).

    When up-sampling (such as nearest-neighbor interpolation) is applied, adjacent elements within each l×ll \times l patch of x^\hat{x} share identical or strongly correlated values. Because the subsequent convolutional mapping conv(⋅)\text{conv}(\cdot) is fixed during inference and exhibits translation invariance, this local feature correlation is deterministically propagated into local pixel interdependencies in the synthesized image II.

    Unlike global frequency-domain signatures that vary widely across generative architectures and cause detectors to overfit to specific training sources, local spatial interdependencies extracted within l×ll \times l pixel neighborhoods constitute a source-invariant artifact that generalizes across diverse up-sampling methods (e.g., nearest-neighbor and bilinear interpolation) and across different generative model families (GANs and Diffusion models).

  3. Knowl 3 — Cross-Source Deepfake Detection Performance Across 28 Generative Models

    data/table

    When trained exclusively on 4 object classes (car, cat, chair, horse) synthesized by ProGAN and evaluated on 38 cross-source sub-testsets spanning 28 distinct generative models (comprising 17 GANs, 1 Deepfake method, and 10 Diffusion models), the NPR-based detector achieves a mean classification accuracy of 92.2%, outperforming all prior deepfake detection methods by at least 11.6 percentage points.

    Method Mean Acc. of 38 sub-testsets (%)
    CNNDetection 57.3
    Frank 56.8
    Durall 56.6
    Patchfor 80.6
    F3Net 78.1
    SelfBland 61.2
    GANDetection 59.5
    LGrad 80.5
    Ojha 79.8
    NPR (Ours) 92.2

    The evaluation includes GAN models (ProGAN, StyleGAN, StyleGAN2, BigGAN, CycleGAN, StarGAN, GauGAN, Deepfake, AttGAN, BEGAN, CramerGAN, InfoMaxGAN, MMDGAN, RelGAN, S3GAN, SNGAN, STGAN) and Diffusion models (DDPM, IDDPM, ADM, LDM, PNDM, VQ-Diffusion, Stable Diffusion v1, Stable Diffusion v2, Glide, Midjourney, and DALLE).

  4. Knowl 4 — Cross-Diffusion Detection Generalization from GAN-Trained Detectors

    data/table

    Detectors trained strictly on ProGAN GAN-generated images generalize effectively to images generated by modern diffusion models when using Neighboring Pixel Relationships (NPR) as the input representation. Detection performance is evaluated using classification Accuracy (Acc.) and Average Precision (A.P.) on three diffusion benchmark suites.

    On the DiffusionForensics benchmark:

    Method ADM DDPM IDDPM LDM Mean
    Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P. Acc. A.P.
    CNNDetection 53.9 71.8 62.7 76.6 50.2 82.7 50.4 78.7 51.0 79.8
    Patchfor 77.5 93.9 62.3 97.1 50.0 91.6 99.5 100.0 78.1 97.8
    F3Net 80.9 96.9 84.7 99.4 74.7 98.9 100.0 100.0 85.8 99.0
    LGrad 86.4 97.5 99.9 100.0 66.1 92.8 99.7 100.0 88.2 98.5
    Ojha 78.4 92.1 72.9 78.8 75.0 92.8 82.2 97.1 74.4 91.7
    NPR (Ours) 88.6 98.9 99.8 100.0 91.8 99.8 100.0 100.0 95.3 99.8

    On the Ojha diffusion benchmark (including DALLE, Glide variations, ADM, and LDM variations), NPR achieves a mean accuracy of 95.2% and mean A.P. of 97.4%, compared to 90.9% / 97.2% for LGrad and 86.9% / 94.5% for Ojha.

    On the Self-Synthesis diffusion benchmark with 1000 diffusion steps (DDPM, IDDPM, ADM) and web-sourced text-to-image models (Midjourney, DALLE), NPR achieves a mean accuracy of 80.1% and mean A.P. of 84.8%, outperforming LGrad (64.2% / 74.3%) and Ojha (66.4% / 73.9%).

  5. Knowl 5 — Cross-GAN Generalization Performance on ForenSynths and Self-Synthesized Datasets

    data/table

    NPR demonstrates robust generalization across diverse GAN architectures when trained only on 4 classes of ProGAN. Performance is measured by Accuracy (Acc. %) and Average Precision (A.P. %).

    On the ForenSynths test set (8 generation models):

    Method ProGAN StyleGAN StyleGAN2 BigGAN CycleGAN StarGAN GauGAN Deepfake Mean
    Acc. / A.P. Acc. / A.P. Acc. / A.P. Acc. / A.P. Acc. / A.P. Acc. / A.P. Acc. / A.P. Acc. / A.P. Acc. / A.P.
    CNNDetection 91.4 / 99.4 63.8 / 91.4 76.4 / 97.5 52.9 / 73.3 72.7 / 88.6 63.8 / 90.8 63.9 / 92.2 51.7 / 62.3 67.1 / 86.9
    Patchfor 97.8 / 100.0 82.6 / 93.1 83.6 / 98.5 64.7 / 69.5 74.5 / 87.2 100.0 / 100.0 57.2 / 55.4 85.0 / 93.2 80.7 / 87.1
    F3Net 99.4 / 100.0 92.6 / 99.7 88.0 / 99.8 65.3 / 69.9 76.4 / 84.3 100.0 / 100.0 58.1 / 56.7 63.5 / 78.8 80.4 / 86.2
    LGrad 99.9 / 100.0 94.8 / 99.9 96.0 / 99.9 82.9 / 90.7 85.3 / 94.0 99.6 / 100.0 72.4 / 79.3 58.0 / 67.9 86.1 / 91.5
    Ojha 99.7 / 100.0 89.0 / 98.7 83.9 / 98.4 90.5 / 99.1 87.9 / 99.8 91.4 / 100.0 89.9 / 100.0 80.2 / 90.2 89.1 / 98.3
    NPR (Ours) 99.8 / 100.0 96.3 / 99.8 97.3 / 100.0 87.5 / 94.5 95.0 / 99.5 99.7 / 100.0 86.6 / 88.8 77.4 / 86.2 92.5 / 96.1

    On the Self-Synthesis 9 GANs dataset (AttGAN, BEGAN, CramerGAN, InfoMaxGAN, MMDGAN, RelGAN, S3GAN, SNGAN, STGAN), NPR achieves a mean accuracy of 93.2% and mean A.P. of 96.6%, outperforming Patchfor (90.1% / 95.4%), Ojha (77.6% / 98.8%), and LGrad (68.6% / 80.8%).

  6. Knowl 6 — NPR Deepfake Detection Architecture and Training Protocol

    model/method

    The deepfake detection framework using Neighboring Pixel Relationships consists of two primary stages:

    1. Representation Extraction: For any input RGB image XiX_i, the extractor f(⋅)f(\cdot) computes the spatial NPR representation Pi=f(Xi)P_i = f(X_i) by local grid partitioning (l=2l=2) and reference pixel subtraction (j=1j=1).
    2. Classification Network: A lightweight convolutional neural network built with standard convolution layers and ResNet residual blocks, containing 1.44 million trainable parameters, takes PiP_i as input and outputs a binary real/fake classification score.

    Training Protocol:

    • Training Data: The detector is trained strictly on the 4-class subset (car, cat, chair, horse) of the ForenSynths ProGAN dataset. Each class contains 18,000 synthetic ProGAN images and 18,000 real images from the LSUN dataset (totaling 72,000 real and 72,000 fake training samples).
    • Optimizer and Hyperparameters: Adam optimizer with an initial learning rate of 2×10−42 \times 10^{-4} and a batch size of 32.
    • Loss Function: Binary cross-entropy loss.
  7. Knowl 7 — Impact of Patch Size and Reference Operators on NPR Detection

    empirical result

    The effectiveness of NPR depends on the local patch scale ll and the reference baseline subtracted from each grid vIcv_I^c.

    Evaluating on the ForenSynths benchmark reveals the following performance across hyperparameter variations:

    1. Patch scale l×ll \times l: Grids of size 2×22 \times 2 (l=2l = 2) consistently outperform 3×33 \times 3 grids (l=3l = 3). For instance, with top-left anchor w1w_1, 2×22 \times 2 yields 92.5% mean accuracy / 96.1% mean A.P., whereas 3×33 \times 3 yields 89.4% mean accuracy / 94.1% mean A.P. This aligns with generator designs that predominantly utilize scale-2 up-sampling layers.
    2. Anchor pixel selection (wjw_j): For 2×22 \times 2 grids, setting wj=w1w_j = w_1 achieves the highest mean accuracy (92.5%), followed by w4w_4 (91.4%), w3w_3 (91.0%), and w2w_2 (87.5%).
    3. Aggregation alternatives: Replacing the single anchor pixel wjw_j with patch-level statistics—such as the grid average avg(vIc)\text{avg}(v_I^c) or grid maximum max(vIc)\text{max}(v_I^c)—yields comparable detection performance (88.3% mean Acc / 91.2% mean A.P. for avg\text{avg}, and 90.9% mean Acc / 95.6% mean A.P. for max\text{max} on 2×22 \times 2 patches). This confirms that the relative variance and local dependencies across 2×22 \times 2 pixel blocks capture the discriminative artifact.
  8. Knowl 8 — Class Activation Mapping (CAM) Analysis of NPR Detectors

    empirical result

    Qualitative analysis using Class Activation Mapping (CAM) on the NPR-based detector demonstrates distinct activation patterns between authentic and AI-synthesized images:

    1. Authentic Images: For real images (e.g., from ImageNet), CAM heatmaps show broadly distributed activations spread across wide contextual regions of the image.
    2. Synthetic Images: For fake images generated by text-to-image diffusion models (e.g., Midjourney and DALLE), the CAM heatmaps concentrate intensely on localized fine-detail regions—such as hair strands, eye boundaries, facial hair, and sharp texture transitions.

    Even though the detector is trained exclusively on four specific object categories (cars, cats, chairs, and horses), the CAM activations remain sensitive to fine-grained spatial synthesis traces on unseen categories and human faces, illustrating that the detector learns source-invariant up-sampling artifacts rather than semantic class-specific features.

Coverage note — None was omitted; all key theoretical analyses, method formulations, ablation studies, and comprehensive multi-dataset benchmark results are included.

References

  1. 1.Marc G Bellemare et al. The cramer distance as a solution to biased wasserstein gradients. arXiv preprint arXiv:1705.10743, 2017. 5
  2. 2.David Berthelot et al. Began: Boundary equilibrium generative adversarial networks. arXiv preprint arXiv:1703.10717, 2017. 5
  3. 3.Andrew Brock et al. Large scale gan training for high fidelity natural image synthesis. In ICLR, 2018. 5
  4. 4.Junyi Cao et al. End-to-end reconstruction-classification learning for face forgery detection. In CVPR, pages 4113–4122, 2022. 3
  5. 5.Lucy Chai et al. What makes fake images detectable? understanding properties that generalize. In ECCV, pages 103–120. Springer, 2020. 3, 5, 6, 7
  6. 6.Liang Chen et al. Ost: Improving generalization of deepfake detection via one-shot test-time training. In Advances in Neural Information Processing Systems, 2022. 2
  7. 7.Liang Chen et al. Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection. In CVPR, pages 18710–18719, 2022. 3
  8. 8.Yunjey Choi et al. Stargan: Unified generative adversarial networks for multi-domain image-to-image translation. In CVPR, pages 8789–8797, 2018. 5
  9. 9.Franc¸ois Chollet. Xception: Deep learning with depthwise separable convolutions. In CVPR, pages 1251–1258, 2017. 1, 2
  10. 10.Prafulla Dhariwal et al. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021. 2, 5
  11. 11.Ricard Durall et al. Watch your up-convolution: Cnn based generative deep neural networks are failing to reproduce spectral distributions. In CVPR, pages 7890–7899, 2020. 1, 2, 3, 4, 5, 6, 7
  12. 12.Joel Frank et al. Leveraging frequency analysis for deep fake image recognition. In ICML, pages 3247–3258. PMLR, 2020. 1, 2, 3, 4, 5, 6, 7
  13. 13.Ian J Goodfellow et al. Generative adversarial nets. In NIPS, 2014. 1
  14. 14.Shuyang Gu et al. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10696–10706, 2022. 5
  15. 15.Alexandros Haliassos et al. Lips don’t lie: A generalisable and robust approach to face forgery detection. In CVPR, pages 5039–5049, 2021. 3
  16. 16.Kaiming He et al. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 5
  17. 17.Yang He et al. Beyond the spectrum: Detecting deepfakes via re-synthesis. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21, pages 2534–2541. International Joint Conferences on Artificial Intelligence Organization, 2021. 3
  18. 18.Zhenliang He et al. Attgan: Facial attribute editing by only changing what you want. IEEE Transactions on Image Processing, 28(11):5464–5478, 2019. 5
  19. 19.Jonathan Ho et al. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33: 6840–6851, 2020. 1, 2, 5
  20. 20.Yonghyun Jeong et al. Bihpf: Bilateral high-pass filters for robust deepfake detection. In WACV, pages 48–57, 2022. 2, 3, 5, 6
  21. 21.Yonghyun Jeong et al. Fingerprintnet: Synthesized fingerprints for generated image detection. In ECCV, pages 76–94. Springer, 2022. 2, 3
  22. 22.Yonghyun Jeong et al. Frepgan: robust deepfake detection using frequency-level perturbations. In AAAI, pages 1060–1068, 2022. 2, 3, 4, 5, 6
  23. 23.Yan Ju et al. Fusing global and local features for generalized ai-synthesized image detection. In 2022 IEEE International Conference on Image Processing (ICIP), pages 3465–3469. IEEE, 2022. 3
  24. 24.Tero Karras et al. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2018. 1, 5
  25. 25.Tero Karras et al. A style-based generator architecture for generative adversarial networks. In CVPR, pages 4401–4410, 2019. 1, 5
  26. 26.Tero Karras et al. Analyzing and improving the image quality of stylegan. In CVPR, pages 8110–8119, 2020. 5
  27. 27.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR (Poster), 2015. 5
  28. 28.Kwot Sin Lee et al. Infomax-gan: Improved adversarial image generation via information maximization and contrastive learning. In WACV, pages 3942–3952, 2021. 5
  29. 29.Chun-Liang Li et al. Mmd gan: Towards deeper understanding of moment matching network. Advances in neural information processing systems, 30, 2017. 5
  30. 30.Yuezun Li et al. In ictu oculi: Exposing ai created fake videos by detecting eye blinking. In 2018 IEEE International workshop on information forensics and security (WIFS), pages 1–7. IEEE, 2018. 3
  31. 31.Tsung-Yi Lin et al. Microsoft coco: Common objects in context. In ECCV, pages 740–755. Springer, 2014. 5
  32. 32.Luping Liu et al. Pseudo numerical methods for diffusion models on manifolds. In ICLR, 2022. 5
  33. 33.Ming Liu et al. Stgan: A unified selective transfer network for arbitrary image attribute editing. In CVPR, pages 3673–3682, 2019. 5
  34. 34.Ziwei Liu et al. Deep learning face attributes in the wild. In ICCV, pages 3730–3738, 2015. 5
  35. 35.Mario Luˇci´c et al. High-fidelity image generation with fewer labels. In ICML, pages 4183–4192. PMLR, 2019. 5
  36. 36.Yuchen Luo et al. Generalizing face forgery detection with high-frequency features. In CVPR, pages 16317–16326, 2021. 3
  37. 37.Sara Mandelli et al. Detecting gan-generated images by orthogonal training of multiple cnns. In 2022 IEEE International Conference on Image Processing (ICIP), pages 3091–3095. IEEE, 2022. 5, 6, 7
  38. 38.Francesco Marra et al. Do gans leave artificial fingerprints? In 2019 IEEE conference on multimedia information processing and retrieval (MIPR), pages 506–511. IEEE, 2019. 3
  39. 39.Iacopo Masi et al. Two-branch recurrent network for isolating deepfakes in videos. In ECCV, pages 667–684. Springer, 2020. 3
  40. 40.Takeru Miyato et al. Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957, 2018. 5
  41. 41.Alex Nichol et al. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. 5
  42. 42.Alexander Quinn Nichol et al. Improved denoising diffusion probabilistic models. In ICML, pages 8162–8171. PMLR, 2021. 5
  43. 43.Weili Nie et al. Relgan: Relational generative adversarial networks for text generation. In ICLR, 2019. 5
  44. 44.Utkarsh Ojha et al. Towards universal fake image detectors that generalize across generative models. In CVPR, pages 24480–24489, 2023. 2, 3, 5, 6, 7, 8
  45. 45.Taesung Park et al. Semantic image synthesis with spatially-adaptive normalization. In CVPR, pages 2337–2346, 2019. 5
  46. 46.Yuyang Qian et al. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In ECCV, pages 86–103. Springer, 2020. 2, 3, 5, 6, 7
  47. 47.Aditya Ramesh et al. Zero-shot text-to-image generation. In ICML, pages 8821–8831. PMLR, 2021. 5
  48. 48.Robin Rombach et al. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1, 2, 5
  49. 49.Olaf Ronneberger et al. U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pages 234–241. Springer, 2015. 4
  50. 50.Andreas Rossler et al. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the ICCV, pages 1–11, 2019. 2, 5
  51. 51.Olga Russakovsky et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3):211–252, 2015. 5
  52. 52.Christoph Schuhmann et al. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. In NeurIPS Workshop Datacentric AI, number FZJ-2022-00923. Julich Super-computing Center, 2021. 5
  53. 53.Kaede Shiohara et al. Detecting deepfakes with self-blended images. In CVPR, pages 18720–18729, 2022. 3, 5, 6, 7
  54. 54.Chuangchuang Tan et al. Learning on gradients: Generalized artifacts representation for gan-generated images detection. In CVPR (CVPR), pages 12105–12114, 2023. 2, 3, 5, 6, 7
  55. 55.Chengrui Wang et al. Representative forgery mining for fake face detection. In CVPR, pages 14923–14932, 2021. 3
  56. 56.Sheng-Yu Wang et al. Cnn-generated images are surprisingly easy to spot... for now. In CVPR, pages 8695–8704, 2020. 3, 5, 6, 7, 8
  57. 57.Zhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang, and Houqiang Li. Altfreezing for more general video face forgery detection. In CVPR, pages 4129–4138, 2023. 3
  58. 58.Zhendong Wang et al. Dire for diffusion-generated image detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22445–22455, 2023. 5, 6, 7, 8
  59. 59.Simon Woo et al. Add: Frequency attention and multi-view based knowledge distillation to detect low-quality compressed deepfake images. In AAAI, pages 122–130, 2022. 3
  60. 60.Fisher Yu et al. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015. 5
  61. 61.Ning Yu et al. Attributing fake images to gans: Learning and analyzing gan fingerprints. In Proceedings of the ICCV, pages 7556–7566, 2019. 3
  62. 62.Xu Zhang et al. Detecting and simulating artifacts in gan fake images. In 2019 IEEE international workshop on information forensics and security (WIFS), pages 1–6. IEEE, 2019. 2
  63. 63.Bolei Zhou et al. Learning deep features for discriminative localization. In CVPR, pages 2921–2929, 2016. 8
  64. 64.Jiapeng Zhu et al. In-domain gan inversion for real image editing. In ECCV, pages 592–608. Springer, 2020. 1
  65. 65.Jun-Yan Zhu et al. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV, pages 2223–2232, 2017. 5

Citation

MLA
Tan, C., et al. “Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection”. arXiv, 2023, http://arxiv.org/abs/2312.10461v2.
APA
Tan, C., Liu, H., Zhao, Y., Wei, S., Gu, G., Liu, P., & Wei, Y. (2023). Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection. arXiv. http://arxiv.org/abs/2312.10461v2
Chicago
Tan, C., H. Liu, Y. Zhao, et al. 2023. “Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection”. arXiv. http://arxiv.org/abs/2312.10461v2.
Harvard
Tan, C. et al. (2023) “Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2312.10461v2.
Vancouver
1. Tan C, Liu H, Zhao Y, Wei S, Gu G, Liu P, Wei Y (2023) Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection. arXiv

BibTeX

@article{tan2023rethinking,
  title = {Rethinking the Up-Sampling Operations in CNN-based Generative Network for Generalizable Deepfake Detection},
  author = {Tan, Chuangchuang and Liu, Huan and Zhao, Yao and Wei, Shikui and Gu, Guanghua and Liu, Ping and Wei, Yunchao},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2312.10461v2},
  eprint = {2312.10461}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE