DRCT: Diffusion Reconstruction Contrastive Training towards Universal Detection of Diffusion Generated Images

Baoying ChenJishen ZengJianquan YangRui Yang

article2024ICML187 citations

Proposes a contrastive training framework using diffusion-reconstructed hard samples and a two-million image benchmark across 16 generators to significantly improve the generalization of AI-generated image detectors against unseen diffusion models.

Listen

Rapid advancements in diffusion-based artificial intelligence have enabled the creation of photorealistic synthetic images, raising critical concerns regarding digital misinformation, copyright infringement, and election interference. While existing synthetic image detectors perform effectively on image formats seen during training, their accuracy collapses when evaluated against new, unseen generation architectures. The article addresses this generalization bottleneck by developing a universal training framework, Diffusion Reconstruction Contrastive Training (DRCT), designed to reliably identify synthetic images across diverse generative tools.

To develop and evaluate the method, the authors constructed DRCT-2M, a benchmark dataset comprising 2 million synthetic images spanning 16 diffusion architectures alongside an evaluation set of 136,000 real-world samples gathered from online platforms. The proposed framework operates on the premise of hard sample classification: by forcing a detector to distinguish authentic images from nearly identical reconstructed counterparts that carry subtle generative traces, the system learns more universal representations. The framework pairs diffusion-based reconstruction of real and generated images with a combined objective of contrastive and classification losses, and it was benchmarked across standard datasets against leading baseline detectors.

Key findings show substantial performance gains across multiple evaluation environments. On cross-model evaluations within the DRCT-2M benchmark, equipping standard detector backbones with DRCT elevated average detection accuracy from approximately 79–83% to between 91% and 97%. On independent GenImage cross-dataset tests, the framework improved baseline accuracy by 7 to 15 percentage points over standard approaches. In real-world wild tests, where conventional detectors experienced severe accuracy degradation down to 11–64%, the enhanced models sustained 82% to 97% detection accuracy. Furthermore, ablation experiments confirmed that both reconstructed sample training and contrastive loss integration provided distinct performance lifts of roughly 6.5% each, while maintaining up to 99% accuracy against image compression and resizing.

These results demonstrate that detectors trained on subtle generative artifacts rather than high-level semantic features achieve greater resilience and lower operational risk when deployed in production security environments. Organizations implementing media forensics, digital watermarking, or trust-and-safety filters can adopt this framework to bolster defenses without discarding existing detector architectures.

Moving forward, stakeholders deploying synthetic image detection tools should incorporate diffusion reconstruction and contrastive training pipelines into active defense monitoring, paired with access controls to counter potential adversarial exploitation. Future research should prioritize expanding the framework to detect localized image manipulations, refining detection for non-diffusion generators such as Generative Adversarial Networks (GANs), and enhancing feature interpretability.

No sufficiently relevant recommendations were found.

Cover for DRCT: Diffusion Reconstruction Contrastive Training towards Universal Detection of Diffusion Generated Images

Abstract

Diffusion models have made significant strides in visual content generation but also raised increasing demands on generated image detection. Existing detection methods have achieved considerable progress, but they usually suffer a significant decline in accuracy when detecting images generated by an unseen diffusion model. In this paper, we seek to address the generalizability of generated image detectors from the perspective of hard sample classification. The basic idea is that if a classifier can distinguish generated images that closely resemble real ones, then it can also effectively detect less similar samples, potentially even those produced by a different diffusion model. Based on this idea, we propose Diffusion Reconstruction Contrastive Learning (DRCT), a universal framework to enhance the generalizability of the existing detectors. DRCT generates hard samples by high-quality diffusion reconstruction and adopts contrastive training to guide the learning of diffusion artifacts. In addition, we have built a million-scale dataset, DRCT-2M, including 16 types diffusion models for the evaluation of generalizability of detection methods. Extensive experimental results show that detectors enhanced with DRCT achieve over a 10% accuracy improvement in cross-set tests. The code, models, and dataset will soon be available at https://github.com/beibuwandeluori/DRCT.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. Image Generation with Diffusion Models
  • 2.2. Generated Image Detection
  • 3. Diffusion Reconstruction Contrastive Training
  • 3.1. The DRCT Framework
  • 3.2. Diffusion Reconstruction
  • 3.3. Contrastive Training
  • 3.4. DRCT-2M Dataset
  • 4. Experiments
  • 4.1. Experimental Setup
  • 4.2. Comparisons of Detection Accuracies
  • 4.3. Comparisons of Generalizability
  • 4.4. Robustness against Post-Processing
  • 4.5. Ablation Studies
  • 5. Discussions
  • Acknowledgements
  • Impact Statement
  • References
  • A. More Analysis of the Proposed DRCT Framework
  • A.1. More Evaluation Metrics on DRCT-2M
  • A.2. More Comparison Results on GenImage
  • A.3. In-Depth Analysis of Reconstructed Image Detectability
  • A.4. The Effect of Reconstruction Step
  • A.5. The Sensitivity of the Detector to the λ Parameter in Loss Function
  • A.6. The Sensitivity of the Detector to the m Parameter in Loss Function
  • A.7. The Influence of Image Category on Detection Performance
  • B. Additional Details of DRCT-2M and DRCT-2M-wild Datasets
  • B.1. More Details of DRCT-2M
  • B.2. More Details of DRCT-2M-wild

Knowls

  1. Knowl 1 — DRCT trains detectors on reconstructed hard examples

    model/method

    Diffusion Reconstruction Contrastive Training (DRCT) enhances a binary real-versus-generated image detector by training it on four image types: original real images, real images reconstructed by a diffusion model, original generated images, and generated images reconstructed by a diffusion model. The detector labels only the original real images as Real; the other three types are treated as Fake. The intended hard examples are reconstructed real images: they retain a real image’s visual content but acquire subtle traces of diffusion generation, so they can resemble real images while carrying generated-image artifacts. The method uses both a classification objective and a contrastive objective to train the detector. The authors’ feature visualizations show reconstructed real images close to real images before DRCT training, supporting their use as difficult training examples.

  2. Knowl 2 — Latent diffusion reconstruction creates the training examples

    model/method

    DRCT reconstructs real and generated images using a pretrained Stable Diffusion model in latent space. For an input image, a variational autoencoder (VAE) encoder first produces a latent representation. Noise is added to that latent using a forward diffusion process; a denoising network then applies deterministic DDIM steps to obtain a reconstructed latent, which the VAE decoder converts back to an image. This differs from reconstructing directly in pixel space: the method operates on the encoded latent representation. One or more selected diffusion models can be used to produce reconstructions. In experiments, the reconstruction model is typically SDv1 or SDv2.

  3. Knowl 3 — DRCT combines margin-based contrastive loss with classification loss

    equation

    DRCT uses a margin-based contrastive loss on pairs of detector feature vectors and combines it with binary cross-entropy for image classification. Let NpN_p be the number of feature pairs; let Yj=1Y_j=1 when pair jj has the same Real/Fake class label and Yj=0Y_j=0 otherwise; let DjD_j be the Euclidean distance between the pair’s feature vectors; and let m>0m>0 be the negative-pair margin. The pairwise loss is

    Lcontrastive=1Np∑j=1Np[YjDj2+(1−Yj)max⁡(0,m−Dj)2].L_{\mathrm{contrastive}}=\frac{1}{N_p}\sum_{j=1}^{N_p}\left[Y_jD_j^2+(1-Y_j)\max(0,m-D_j)^2\right].

    For a batch of NsN_s images, let yi∈{0,1}y_i\in\{0,1\} be the binary target for image ii and pip_i the detector’s predicted probability of its positive class. The classification loss and total objective are

    Lcross-entropy=−∑i=1Ns[yilog⁡(pi)+(1−yi)log⁡(1−pi)],Ltotal=λLcontrastive+(1−λ)Lcross-entropy.L_{\mathrm{cross\text{-}entropy}}=-\sum_{i=1}^{N_s}\left[y_i\log(p_i)+(1-y_i)\log(1-p_i)\right],\qquad L_{\mathrm{total}}=\lambda L_{\mathrm{contrastive}}+(1-\lambda)L_{\mathrm{cross\text{-}entropy}}.

    The paper’s default experimental settings are margin m=1.0m=1.0 and loss balance λ=0.3\lambda=0.3, with λ∈[0,1)\lambda\in[0,1).

  4. Knowl 4 — DRCT-2M and DRCT-2M-Wild provide two million synthetic images and real-world samples

    data/table

    DRCT-2M contains approximately two million generated images, with 123,287 images for each of 16 diffusion-model types. Ten types are text-to-image models: LDM, SDv1.4, SDv1.5, SDv2, SDXL, SDXL-Refiner, SD-Turbo, SDXL-Turbo, LCM-SDv1.5, and LCM-SDXL. The other six are image-to-image models: three ControlNet variants (SDv1-Ctrl, SDv2-Ctrl, and SDXL-Ctrl) and three diffusion-reconstruction variants (SDv1-DR, SDv2-DR, and SDXL-DR). Text-to-image prompts come from MSCOCO captions; ControlNet inputs use captions and Canny edge maps from real images; reconstruction-model inputs use real images, empty prompts, and zero masks. The real-image source is MSCOCO.

    DRCT-2M-Wild adds 136,817 internet-collected generated images from eight sources: DreamShaper XL10 (8,696), Niji Special Edition (3,615), Realistic Vision v5.1 (36,999), Deep Negative v1.x (11,600), Detail Tweaker Lora (24,293), MajicMix Realistic (41,496), rMada Merge (769), and Midjourney (9,349). The first seven sources were collected from CIVITAI and Midjourney images from Discord. The two dataset components support evaluation across controlled diffusion-model variants and images encountered in real-world settings.

  5. Knowl 5 — Reconstructed real images are difficult examples with diffusion-related traces

    empirical result

    The paper’s feature-space and frequency analyses support the premise that diffusion-reconstructed real images are useful hard examples. In a t-SNE visualization of detector features, real-reconstructed images cluster near original real images, whereas SDv1.4-generated images lie farther from the real-image cluster; the reconstructed-real samples are therefore harder to distinguish from real images than ordinary generated samples. After DRCT fine-tuning, the visualization shows improved separation of the relevant groups. In a separate analysis averaging Fourier amplitude spectra over 5,000 images per category, the spectra of SDv1.4 images, SDv1.4-reconstructed images, and SDv2 images share a pattern distinct from real images, while the spectrum of real-reconstructed images is more similar to real images. Together these observations indicate that reconstructed real images preserve real-like appearance while containing generation-related characteristics.

  6. Knowl 6 — DRCT improves accuracy across diffusion variants in DRCT-2M

    empirical result

    On DRCT-2M, detectors were evaluated across 16 generated-image subsets after training on SDv1.4-generated images and MSCOCO real images; DRCT variants additionally used reconstructions, with the original fake-image source matched to the reconstruction model (SDv1.4 for SDv1 reconstruction and SDv2 for SDv2 reconstruction). Average accuracy (ACC) across the 16 subsets was 79.11% for the Conv-B baseline, compared with 90.79% for DRCT/Conv-B using SDv1 reconstruction and 96.55% using SDv2 reconstruction. For UnivFD, the corresponding results were 83.46%, 90.49%, and 91.35%. The authors report that conventional detectors often lose accuracy on substantially altered or unseen variants such as SDXL and on diffusion-reconstructed images, whereas DRCT improves average performance across the evaluated variants.

  7. Knowl 7 — DRCT improves transfer to GenImage subsets

    empirical result

    On GenImage, DRCT-enhanced detectors improved performance under both same-benchmark and cross-dataset training protocols. When trained on GenImage’s SDv1.4 subset and evaluated across its eight subsets, UnivFD achieved 79.45% average ACC and DRCT/UnivFD achieved 89.49%; Conv-B rose from 74.98% to 82.08%. In a separate cross-dataset evaluation, detectors trained on DRCT-2M/SDv1.4 were tested on GenImage: DRCT/UnivFD achieved 87.67% average ACC, compared with 67.73% for UnivFD and 72.83% for the strongest non-DRCT comparator, F3Net. In this cross-dataset setting, DRCT/Conv-B scored 83.53%, versus 68.98% for Conv-B. These results cover diffusion generators and, in GenImage, the non-diffusion BigGAN subset as well.

  8. Knowl 8 — DRCT transfers to internet-collected generated images

    empirical result

    On DRCT-2M-Wild, detectors trained on DRCT-2M/SDv1.4 were evaluated on images collected from eight online generation sources. The baseline UnivFD achieved 63.55% average ACC, while DRCT/UnivFD using SDv1 reconstruction achieved 87.23% and using SDv2 reconstruction achieved 96.90%. Conv-B scored 44.14% without DRCT, compared with 82.06% for DRCT/Conv-B using SDv1 reconstruction and 93.16% using SDv2 reconstruction. The results show that DRCT-enhanced detectors transferred substantially better to these collected images than their corresponding baseline detectors.

  9. Knowl 9 — Ablations support the use of reconstruction and contrastive supervision

    empirical result

    In a Conv-B ablation trained on DRCT-2M/SDv1.4 and evaluated on GenImage, the full DRCT configuration achieved 83.53% average ACC, compared with 68.98% for the initial baseline configuration. The reported ablation analysis found that adding reconstructed real images improved average accuracy by 6.55 percentage points; adding back original SDv1.4 fake images gave a further 2.5-point gain, and adding reconstructed fake images gave a further 0.96-point gain. Adding contrastive loss to the classification-only configuration increased average ACC from 76.99% to 83.53%, a 6.54-point improvement. The paper also reports that, for DRCT/Conv-B on GenImage, average accuracy first increased and then decreased as either loss weight λ\lambda or margin mm increased; the best tested settings were λ=0.3\lambda=0.3 and m=1.0m=1.0.

  10. Knowl 10 — DRCT improves robustness to resizing and JPEG compression

    empirical result

    The robustness evaluation applied resizing at scales 0.50.5, 0.750.75, 1.01.0, 1.251.25, and 1.51.5, and JPEG compression at quality factors 60, 70, 80, 90, and 100, to real and generated test images. Compared with the corresponding baseline detectors, DRCT-enhanced Conv-B detectors showed higher robustness, with reported accuracy reaching up to 99% under resizing and JPEG compression. The authors found DRCT/Conv-B more robust than DRCT/UnivFD in these tests and attributed the difference mainly to Conv-B tuning all network weights, while UnivFD tunes only its final fully connected layer.

  11. Knowl 11 — Reported limitations concern non-diffusion and localized generation

    limitation

    The authors report that DRCT improves detection of non-diffusion-generated images such as GAN outputs, but the improvement is less marked than for diffusion-generated images, which they attribute to differences in generation processes and artifacts. Their evaluation covers globally generated images, not localization or detection of locally generated regions; they identify small manipulated regions as a particularly challenging case and leave this setting for future work.

Coverage note — Omitted the supplementary category-by-category detection analysis and detailed image-quality score comparisons because they are secondary to the DRCT method, benchmark construction, and generalization results.

References

  1. 1.Bird, J. J. and Lotfi, A. Cifake: Image classification and explainable identification of ai-generated synthetic images. ArXiv, abs/2303.14126, 2023. URL https://api.semanticscholar.org/CorpusID:257757303.
  2. 2.Brock, A., Donahue, J., and Simonyan, K. Large scale gan training for high fidelity natural image synthesis. ArXiv, abs/1809.11096, 2018. URL https://api.semanticscholar.org/CorpusID:52889459.
  3. 3.Canny, J. A computational approach to edge detection. IEEE Transactions on pattern analysis and machine intelligence, (6):679–698, 1986.
  4. 4.Corvi, R., Cozzolino, D., Zingarini, G., Poggi, G., Nagano, K., and Verdoliva, L. On the detection of synthetic images generated by diffusion models. ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 1–5, 2022.
  5. 5.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. ArXiv, abs/2105.05233, 2021. URL https://api.semanticscholar.org/CorpusID:234357997.
  6. 6.Epstein, D. C., Jain, I., Wang, O., and Zhang, R. Online detection of ai-generated images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 382–392, 2023.
  7. 7.Frank, J. C., Eisenhofer, T., Schönherr, L., Fischer, A., Kolossa, D., and Holz, T. Leveraging frequency analysis for deep fake image recognition. ArXiv, abs/2003.08685, 2020. URL https://api.semanticscholar.org/CorpusID:213175447.
  8. 8.Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A. C., and Bengio, Y. Generative adversarial networks. Communications of the ACM, 63:139 – 144, 2014.
  9. 9.Guarnera, L., Giudice, O., and Battiato, S. Level up the deepfake detection: a method to effectively discriminate images generated by gan architectures and diffusion models. ArXiv, abs/2303.00608, 2023. URL https://api.semanticscholar.org/CorpusID:257255351.
  10. 10.Guo, X., Liu, X., Ren, Z., Grosz, S., Masi, I., and Liu, X. Hierarchical fine-grained image forgery detection and localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3155–3165, 2023.
  11. 11.Hadsell, R., Chopra, S., and LeCun, Y. Dimensionality reduction by learning an invariant mapping. 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), 2:1735–1742, 2006. URL https://api.semanticscholar.org/CorpusID:8281592.
  12. 12.Higgins, I., Matthey, L., Pal, A., Burgess, C. P., Glorot, X., Botvinick, M. M., Mohamed, S., and Lerchner, A. beta-vae: Learning basic visual concepts with a constrained variational framework. In International Conference on Learning Representations, 2016. URL https://api.semanticscholar.org/CorpusID:46798026.
  13. 13.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. ArXiv, abs/2006.11239, 2020. URL https://api.semanticscholar.org/CorpusID:219955663.
  14. 14.Hou, X., Shen, L., Sun, K., and Qiu, G. Deep feature consistent variational autoencoder. 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), pp. 1133–1141, 2016. URL https://api.semanticscholar.org/CorpusID:5257869.
  15. 15.Ju, Y., Jia, S., Ke, L., Xue, H., Nagano, K., and Lyu, S. Fusing global and local features for generalized ai-synthesized image detection. 2022 IEEE International Conference on Image Processing (ICIP), pp. 3465–3469, 2022. URL https://api.semanticscholar.org/CorpusID:247762136.
  16. 16.Karras, T., Aila, T., Laine, S., and Lehtinen, J. Progressive growing of gans for improved quality, stability, and variation. ArXiv, abs/1710.10196, 2017. URL https://api.semanticscholar.org/CorpusID:3568073.
  17. 17.Karras, T., Laine, S., and Aila, T. A style-based generator architecture for generative adversarial networks. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 4396–4405, 2018. URL https://api.semanticscholar.org/CorpusID:54482423.
  18. 18.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. CoRR, abs/1312.6114, 2013. URL https://api.semanticscholar.org/CorpusID:216078090.
  19. 19.Langley, P. Crafting papers on machine learning. In Langley, P. (ed.), Proceedings of the 17th International Conference on Machine Learning (ICML 2000), pp. 1207–1216, Stanford, CA, 2000. Morgan Kaufmann.
  20. 20.Li, J., Li, D., Xiong, C., and Hoi, S. C. H. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, 2022. URL https://api.semanticscholar.org/CorpusID:246411402.
  21. 21.Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 2014. URL https://api.semanticscholar.org/CorpusID:14113767.
  22. 22.Liu, B., Yang, F., Bi, X., Xiao, B., Li, W., and Gao, X. Detecting generated images by real images. In European Conference on Computer Vision, 2022a. URL https://api.semanticscholar.org/CorpusID:253121047.
  23. 23.Liu, L., Ren, Y., Lin, Z., and Zhao, Z. Pseudo numerical methods for diffusion models on manifolds. ArXiv, abs/2202.09778, 2022b. URL https://api.semanticscholar.org/CorpusID:247011732.
  24. 24.Liu, Z., Qi, X., Jia, J., and Torr, P. H. S. Global texture enhancement for fake face detection in the wild. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8057–8066, 2020. URL https://api.semanticscholar.org/CorpusID:211010785.
  25. 25.Liu, Z., Mao, H., Wu, C., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11966–11976, 2022c. URL https://api.semanticscholar.org/CorpusID:245837420.
  26. 26.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. ArXiv, abs/2206.00927, 2022. URL https://api.semanticscholar.org/CorpusID:249282317.
  27. 27.Luo, S., Tan, Y., Patil, S., Gu, D., von Platen, P., Passos, A., Huang, L., Li, J., and Zhao, H. Lcm-lora: A universal stable-diffusion acceleration module. ArXiv, abs/2311.05556, 2023. URL https://api.semanticscholar.org/CorpusID:265067414.
  28. 28.Ma, R., Duan, J., Kong, F., Shi, X., and Xu, K. Exposing the fake: Effective diffusion-generated images detection. In The Second Workshop on New Frontiers in Adversarial Machine Learning, 2023. URL https://openreview.net/forum?id=7R62e4Wgim.
  29. 29.Nichol, A. and Dhariwal, P. Improved denoising diffusion probabilistic models. ArXiv, abs/2102.09672, 2021. URL https://api.semanticscholar.org/CorpusID:231979499.
  30. 30.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, 2021. URL https://api.semanticscholar.org/CorpusID:245335086.
  31. 31.Ojha, U., Li, Y., and Lee, Y. J. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 24480–24489, 2023.
  32. 32.Park, T., Liu, M.-Y., Wang, T.-C., and Zhu, J.-Y. Semantic image synthesis with spatially-adaptive normalization. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2332–2341, 2019. URL https://api.semanticscholar.org/CorpusID:81981856.
  33. 33.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Muller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. ArXiv, abs/2307.01952, 2023. URL https://api.semanticscholar.org/CorpusID:259341735.
  34. 34.Qian, Y., Yin, G., Sheng, L., Chen, Z., and Shao, J. Thinking in frequency: Face forgery detection by mining frequency-aware clues. ArXiv, abs/2007.09355, 2020. URL https://api.semanticscholar.org/CorpusID:220647499.
  35. 35.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, 2021. URL https://api.semanticscholar.org/CorpusID:231591445.
  36. 36.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. ArXiv, abs/2204.06125, 2022. URL https://api.semanticscholar.org/CorpusID:248097655.
  37. 37.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10674–10685, 2021. URL https://api.semanticscholar.org/CorpusID:245335280.
  38. 38.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M. Photorealistic text-to-image diffusion models with deep language understanding. ArXiv, abs/2205.11487, 2022. URL https://api.semanticscholar.org/CorpusID:248986576.
  39. 39.Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R. Adversarial diffusion distillation. ArXiv, abs/2311.17042, 2023. URL https://api.semanticscholar.org/CorpusID:265466173.
  40. 40.Sha, Z., Li, Z., Yu, N., and Zhang, Y. De-fake: Detection and attribution of fake images generated by text-to-image generation models. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 3418–3432, 2023.
  41. 41.Sohn, K., Lee, H., and Yan, X. Learning structured output representation using deep conditional generative models. In Neural Information Processing Systems, 2015. URL https://api.semanticscholar.org/CorpusID:13936837.
  42. 42.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. ArXiv, abs/2010.02502, 2020. URL https://api.semanticscholar.org/CorpusID:222140788.
  43. 43.Tan, C., Zhao, Y., Wei, S., Gu, G., and Wei, Y. Learning on gradients: Generalized artifacts representation for gan-generated images detection. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12105–12114, 2023. URL https://api.semanticscholar.org/CorpusID:259226993.
  44. 44.van den Oord, A., Vinyals, O., and Kavukcuoglu, K. Neural discrete representation learning. ArXiv, abs/1711.00937, 2017. URL https://api.semanticscholar.org/CorpusID:20282961.
  45. 45.von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S., and Wolf, T. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022.
  46. 46.Wang, S.-Y., Wang, O., Zhang, R., Owens, A., and Efros, A. A. Cnn-generated images are surprisingly easy to spot. . . for now. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 8692–8701, 2019.
  47. 47.Wang, Z., Bao, J., Zhou, W., Wang, W., Hu, H., Chen, H., and Li, H. Dire for diffusion-generated image detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 22445–22455, October 2023.
  48. 48.Wu, H., Zhou, J., and Zhang, S. Generalizable synthetic image detection via language-guided contrastive learning. ArXiv, abs/2305.13800, 2023a. URL https://api.semanticscholar.org/CorpusID:258841383.
  49. 49.Wu, X., Hao, Y., Sun, K., Chen, Y., Zhu, F., Zhao, R., and Li, H. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. ArXiv, abs/2306.09341, 2023b. URL https://api.semanticscholar.org/CorpusID:259171771.
  50. 50.Xi, Z., Huang, W., Wei, K., Luo, W., and Zheng, P. Ai-generated image detection using a cross-attention enhanced dual-stream network. In 2023 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), pp. 1463–1470. IEEE, 2023.
  51. 51.Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., and Dong, Y. Imagereward: Learning and evaluating human preferences for text-to-image generation. ArXiv, abs/2304.05977, 2023. URL https://api.semanticscholar.org/CorpusID:258079316.
  52. 52.Zhang, L., Rao, A., and Agrawala, M. Adding conditional control to text-to-image diffusion models. ArXiv, abs/2302.05543, 2023. URL https://api.semanticscholar.org/CorpusID:256827727.
  53. 53.Zhao, S., Song, J., and Ermon, S. Infovae: Information maximizing variational autoencoders. ArXiv, abs/1706.02262, 2017. URL https://api.semanticscholar.org/CorpusID:34051459.
  54. 54.Zhong, N., Xu, Y., Qian, Z., and Zhang, X. Rich and poor texture contrast: A simple yet effective approach for ai-generated image detection. ArXiv, abs/2311.12397, 2023. URL https://api.semanticscholar.org/CorpusID:265309137.
  55. 55.Zhu, J.-Y., Park, T., Isola, P., and Efros, A. A. Unpaired image-to-image translation using cycle-consistent adversarial networks. 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2242–2251, 2017. URL https://api.semanticscholar.org/CorpusID:206770979.
  56. 56.Zhu, M., Chen, H., Yan, Q., Huang, X., Lin, G., Li, W., Tu, Z., Hu, H., Hu, J., and Wang, Y. Genimage: A million-scale benchmark for detecting ai-generated image. ArXiv, abs/2306.08571, 2023. URL https://api.semanticscholar.org/CorpusID:259164965.

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/