Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images

Zeyu LuDi HuangLei BaiJingjing QuChengyue WuXihui LiuWanli Ouyang

article2023NeurIPS106 citations

Presents a two-million-image dataset alongside comprehensive human and algorithmic benchmarks, revealing that people fail to identify state-of-the-art AI-generated images nearly 39% of the time while specialized detection models reduce this error rate to 13%.

Listen

Photographs have historically served as a reliable medium for recording factual events and personal experiences. However, rapid advancements in artificial intelligence (AI) text-to-image synthesis now enable the generation of hyper-realistic visual content, raising significant risks regarding misinformation, fraudulent media, and the erosion of public trust. The article addresses the urgent problem of determining whether humans and automated detection algorithms can reliably distinguish state-of-the-art AI-generated images from authentic photographs.

The main objective of the article is to benchmark human discernment against cutting-edge AI detection models when identifying modern, high-quality synthetic images. To support this evaluation, the researchers developed Fake2M, a comprehensive public dataset containing over 2 million synthetic images generated by prominent diffusion, generative adversarial, and autoregressive models alongside matched real photographs. Using this resource, the article establishes HPBench to test human perception across a diverse group of 50 participants and MPBench to assess multiple automated vision backbones across 11 validation sets.

The findings reveal that modern AI-generated images readily deceive human vision, resulting in an overall human misclassification rate of 38.7% (61.3% accuracy). Notably, participants correctly classified synthetic images as fake only 55.8% of the time, and their confidence in authentic photography was also degraded, correctly sorting real online images only 66.9% of the time. Human participants performed best when evaluating portraits and multi-person scenes (up to 67.5% accuracy) due to perceptual sensitivity to human anatomical details, but dropped to 50.8% accuracy on non-human objects. In contrast, automated algorithms significantly outperformed humans, with the top-performing model achieving an 87% accuracy (a 13% error rate) on the same evaluation benchmark. However, model performance varied substantially across validation sets, and no single architecture consistently dominated across all data configurations.

These results demonstrate that human visual inspection alone is no longer an adequate defense against AI-generated misinformation. While automated detection systems offer superior accuracy, their sensitivity to specific generative architectures and sampling parameters creates operational risks and vulnerabilities to novel synthesis techniques. Organizations relying on visual evidence for policy, security, or compliance must recognize that unassisted staff cannot reliably authenticate images, and existing algorithmic safeguards remain imperfect.

To mitigate these risks, decision-makers should deploy multi-model automated screening tools trained on highly diverse generative datasets, rather than relying on single-architecture detectors. Organizations should also provide targeted verification training to staff while establishing formal digital provenance tracking. Further research and technical pilots are needed to improve detector generalization across unseen generative techniques and to develop balanced architectures that effectively spot both authentic and synthetic visual media.

The study's conclusions should be interpreted in light of certain limitations, including the controlled 50-participant human sample and the rapid evolution of generative models, which may soon eliminate current visual artifacts such as texture over-smoothing. Nevertheless, there is high confidence in the core finding that modern generative tools have surpassed human ability to discern fake media without automated assistance.

Cover for Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images

Abstract

Photos serve as a way for humans to record what they experience in their daily lives, and they are often regarded as trustworthy sources of information. However, there is a growing concern that the advancement of artificial intelligence (AI) technology may produce fake photos, which can create confusion and diminish trust in photographs. This study aims to comprehensively evaluate agents for distinguishing state-of-the-art AI-generated visual content. Our study benchmarks both human capability and cutting-edge fake image detection AI algorithms, using a newly collected large-scale fake image dataset Fake2M. In our human perception evaluation, titled HPBench, we discovered that humans struggle significantly to distinguish real photos from AI-generated ones, with a misclassification rate of 38.7%. Along with this, we conduct the model capability of AI-Generated images detection evaluation MPBench and the top-performing model from MPBench achieves a 13% failure rate under the same setting used in the human evaluation. We hope that our study can raise awareness of the potential risks of AI-generated images and facilitate further research to prevent the spread of false information. More information can refer to https://github.com/Inf-imagine/Sentry.

Table of Contents

  • 1 Introduction
  • 2 Dataset Collection and Generation
  • 2.1 Collect Data for Human Evaluation
  • 2.2 Collect Data for Model Evaluation
  • 2.3 Comparison with Other Datasets
  • 3 HPBench: Human Perception of AI-Generated Images Evaluation
  • 3.1 Evaluation Setup
  • 3.2 Results and Analysis
  • 3.2.1 Overall Ability to Distinguish Real and AI-generated Images
  • 3.2.2 Distinguishing Abilities of Participants with Various Personal Backgrounds
  • 3.2.3 Distinguishability of different photo categories
  • 3.2.4 Results of the Judgment Criteria and Analysis of the AIGC Defects
  • 4 MPBench: Model Perception of AI-Generated Images Evaluation
  • 4.1 Experiments Setup
  • 4.2 Results and Analysis
  • 4.2.1 Comparative Analysis of Accuracy Across Various Models
  • 4.2.2 Comparative Analysis of Accuracy Across Various Training Datasets
  • 4.2.3 Comparative Analysis of Accuracy Across Various Validation Datasets
  • 4.2.4 Evaluate the best model under the same setting used in HPBench.
  • 5 Related Work
  • 6 Conclusion
  • References

Knowls

  1. Knowl 1 — Fake2M Dataset Construction and Structure

    experimental setup

    Fake2M is a large-scale dataset created for training and evaluating AI-generated image detection systems. It contains over 2 million synthetic images generated by modern generative models alongside corresponding real image datasets.

    Training Sets

    The training component contains approximately 2 million fake images and matching real images:

    1. SD-V1.5Real-dpms-25: 1 million fake images synthesized from the first 1M captions of Conceptual Captions (CC3M) using Stable Diffusion v1.5 Realistic Vision V2.0 with DPM-Solver across 25 steps.
    2. IF-V1.0-dpms++-25: 1 million fake images generated from the first 1M captions of CC3M using DeepFloyd IF v1.0 with DPM-Solver++ across 25 steps.
    3. StyleGAN3: 87,300 fake images matching GAN domains, consisting of 35K from StyleGAN3-t-ffhq, 35K from StyleGAN3-r-ffhq, 650 from StyleGAN3-t-metfaces, 650 from StyleGAN3-r-metfaces, 8K from StyleGAN3-t-afhqv2, and 8K from StyleGAN3-r-afhqv2.
    4. Real Training Sets: 1M real images from CC3M-Train and 87,300 real images combined from FFHQ, AFHQv2, and MetFaces (StyleGAN3-Train).
  2. Knowl 2 — MPBench Benchmark Setup and Dataset Configurations

    experimental setup

    MPBench is a model perception benchmark consisting of 11 fake validation datasets (~257K images) and 3 real validation datasets (139K images) spanning diffusion models, generative adversarial networks (GANs), autoregressive models, and proprietary text-to-image engines.

    Validation Datasets

    • Real Validation Datasets: ImageNet-Test (100K images from ImageNet-1k test set), CelebA-HQ-Train (24K images), CC3M-Val (15K images).
    • Diffusion Fake Datasets: Generated using the 15K captions of CC3M validation split:
      • SD-V2.1-dpms-25 (15K images, Stable Diffusion v2.1, DPM-Solver 25 steps)
      • SD-V1.5-dpms-25 (15K images, Stable Diffusion v1.5, DPM-Solver 25 steps)
      • SD-V1.5Real-dpms-25 (15K images, Realistic Vision v2.0, DPM-Solver 25 steps)
      • IF-V1.0-dpms++-10 (15K images, IF v1.0, DPM-Solver++ 10 steps)
      • IF-V1.0-dpms++-25 (15K images, IF v1.0, DPM-Solver++ 25 steps)
      • IF-V1.0-dpms++-50 (15K images, IF v1.0, DPM-Solver++ 50 steps)
      • IF-V1.0-ddim-50 (15K images, IF v1.0, DDIM 50 steps)
      • IF-V1.0-ddpm-50 (15K images, IF v1.0, DDPM 50 steps)
    • Autoregressive Fake Dataset: CogView2 (22K images generated from CC3M validation captions).
    • Proprietary Generative Model Dataset: Midjourney (5.5K images crawled from the community).
    • GAN Fake Dataset: StyleGAN3 (60K images generated using StyleGAN3 variants on FFHQ, MetFaces, and AFHQv2).

    Model Training Settings

    Evaluated models are trained under four dataset configurations (each totaling balanced real and fake images):

    • Setting A: 1M fake images from SD-V1.5Real-dpms-25 and 1M real images from CC3M-Train.
    • Setting B: 1M fake images from IF-V1.0-dpms++-25 and 1M real images from CC3M-Train.
    • Setting C: 87K fake images from StyleGAN3 and 87K real images from StyleGAN3-Train.
    • Setting D (Multi-generator): 460K fake images from SD-V1.5Real-dpms-25, 460K fake images from IF-V1.0-dpms++-25, 87K fake images from StyleGAN3, 1M real images from CC3M-Train, and 87K real images from StyleGAN3-Train.
  3. Knowl 3 — HPBench Evaluation Protocol and Dataset Setup

    experimental setup

    HPBench evaluates human perception in identifying state-of-the-art AI-generated images versus real photographs.

    Data Construction

    • Synthetic Images: 151 photorealistic images generated using Midjourney-V5 across eight categories: Multiperson (10), Landscape (27), Man (17), Woman (30), Record (15), Plant (13), Animal (29), and Object (10). Generation prompts used authenticating suffixes (such as "normal color, aesthetic, shocking, very detailed, photographic, 8K, HDR"). Annotators filtered out low-quality generations containing obvious visual artifacts.
    • Real Images: 244 photographs sourced from 500px and Google Images using identical descriptive prompts as the AI images, categorized across the same eight domains.

    Participant Setup

    A cohort of 50 diverse participants completed an untimed questionnaire containing 100 randomized images (each labeled as real or AI-generated). Demographics included:

    • Gender: 31 Male, 19 Female.
    • Background: 27 with generative AI experience (w/ AIGC), 23 without generative AI experience (w/o AIGC).
    • Age: 42 participants aged 20–29; 8 participants aged 30–45.
  4. Knowl 4 — Human Performance in Discerning AI-Generated Images (HPBench)

    data/table

    In HPBench, positive samples represent AI-generated images and negative samples represent real photographs. Performance across 50 participants evaluated on 100-question image sets is summarized below:

    Metric All Multiperson Landscape Man Woman Record Plant Animal Object
    Accuracy ↑\uparrow 0.6134 0.6750 0.5650 0.6433 0.6637 0.6233 0.5983 0.6133 0.5083
    Precision ↑\uparrow 0.6278 0.7075 0.5657 0.6666 0.6765 0.6340 0.6213 0.6156 0.5112
    Recall ↑\uparrow 0.5577 0.5966 0.5714 0.5733 0.6275 0.5833 0.5033 0.6033 0.3800
    FOR ↓\downarrow 0.3981 0.3487 0.4358 0.3742 0.3473 0.3858 0.4173 0.3888 0.4933

    where FOR=False NegativesFalse Negatives+True Negatives\text{FOR} = \frac{\text{False Negatives}}{\text{False Negatives} + \text{True Negatives}} is the False Omission Rate.

    Key Perceptual Findings

    1. Average Accuracy: The overall human accuracy is 61.34%61.34\% (misclassification rate of 38.66%38.66\%). Individual accuracy ranged from 40%40\% (lowest) to 73%73\% (highest).
    2. Real vs. Fake Bias: Humans correctly identified real photos with 66.9%66.9\% probability, but misclassified AI-generated images as real with a probability of 44.2%44.2\%.
    3. Category Variance: Categories featuring humans (Multiperson at 67.50%67.50\%, Woman at 66.37%66.37\%, Man at 64.33%64.33\%) showed higher distinguishability than inanimate categories (Object at 50.83%50.83\%, Landscape at 56.50%56.50\%), revealing a 16.67%16.67\% gap between easiest and hardest categories.
    4. Experience Effect: Participants with prior AIGC experience scored 2.7%2.7\% higher overall and 3.7%3.7\% higher on detecting AI-generated images compared to participants without AIGC experience.
  5. Knowl 5 — Human Visual Judgment Criteria for Detecting AI Artifacts

    empirical result

    Analysis of subjective criteria reported by human participants when successfully identifying AI-generated images in HPBench shows the distribution of perceived defects:

    Judgment Criterion Selection Count Percentage
    Detail Problem 332 28%
    Smooth Problem 205 17%
    Intuition (Unspecified) 169 14%
    Blur 142 12%
    Color Artifacts 122 10%
    Shadow Light 95 8%
    Daub 59 5%
    Rationality 57 5%

    Fine detail inconsistencies (28%28\%) and unnatural skin or surface smoothness (17%17\%) represent the primary visual cues utilized by human observers. 14%14\% of correct decisions were attributed purely to intuition without identifiable visual flaws.

  6. Knowl 6 — Quantitative Evaluation of Fake Detection Models in MPBench

    data/table

    MPBench evaluates five model architectures across four training dataset configurations (Settings A, B, C, D) over 14 validation benchmarks (3 real datasets, 11 fake datasets). Detection classification accuracy (%) is detailed below:

    Training Setting Model Architecture Real Acc. Fake Acc. Total Avg. Acc.
    Setting A ConvNeXt-S (Blur+JPEG 0.1) 98.3 62.5 70.2
    (SD-V1.5Real 1M) ConvNeXt-S (Blur+JPEG 0.5) 98.5 60.4 68.6
    ResNet-50 (Blur+JPEG 0.1) 94.6 44.3 55.1
    ResNet-50 (Blur+JPEG 0.5) 94.8 42.1 53.4
    CLIP-ViT-L (Linear Classifier) 70.6 85.1 82.0
    Setting B ConvNeXt-S (Blur+JPEG 0.1) 95.8 41.3 52.9
    (IF-V1.0 1M) ConvNeXt-S (Blur+JPEG 0.5) 95.8 44.1 55.2
    ResNet-50 (Blur+JPEG 0.1) 93.4 51.2 60.3
    ResNet-50 (Blur+JPEG 0.5) 93.6 54.4 62.8
    CLIP-ViT-L (Linear Classifier) 85.9 52.9 60.0
    Setting C ConvNeXt-S (Blur+JPEG 0.1) 67.4 44.7 49.6
    (StyleGAN3 87K) ConvNeXt-S (Blur+JPEG 0.5) 77.1 35.4 44.3
    ResNet-50 (Blur+JPEG 0.1) 28.3 68.6 60.0
    ResNet-50 (Blur+JPEG 0.5) 54.5 45.6 47.5
    CLIP-ViT-L (Linear Classifier) 64.5 86.4 81.7
    Setting D ConvNeXt-S (Blur+JPEG 0.1) 98.4 77.1 81.6
    (Multi-generator) ConvNeXt-S (Blur+JPEG 0.5) 99.1 78.7 83.0
    ResNet-50 (Blur+JPEG 0.1) 87.5 78.2 80.2
    ResNet-50 (Blur+JPEG 0.5) 90.4 76.4 79.4
    CLIP-ViT-L (Linear Classifier) 75.4 83.3 81.6

    ConvNeXt-S with Blur+JPEG augmentation (0.5 probability) trained on Setting D achieved the highest overall accuracy across the benchmark (83.0%83.0\%).

  7. Knowl 7 — Model Perception Trade-offs and Generator Generalization

    empirical result

    Analysis of the MPBench benchmark results reveals two key patterns in automated AI image detection:

    1. Real vs. Fake Accuracy Asymmetry: Vision backbone fine-tuning (ConvNeXt-S) achieves superior real image detection (up to 99.1%99.1\% accuracy under Setting D), but exhibits lower cross-model fake generalization. In contrast, linear probing on frozen CLIP features (CLIP-ViT-L) achieves higher cross-generator fake image classification (85.1%85.1\% in Setting A, 86.4%86.4\% in Setting C, 83.3%83.3\% in Setting D) but suffers lower accuracy on real images (70.6%70.6\% to 75.4%75.4\%).
    2. Training Diversity Benefits: Models trained on single generator distributions (Settings A, B, C) perform poorly when evaluated on out-of-domain generators (e.g., ConvNeXt-S trained on IF-V1.0 achieves only 2.2%2.2\% accuracy on SD-V2.1 fake images). Multi-generator training (Setting D) yields balanced performance across diverse generative architectures (Diffusion, GAN, Autoregressive, Midjourney).
  8. Knowl 8 — Performance Discrepancy Between Model and Human Perception

    empirical result

    When evaluating the highest-performing automated detection model from MPBench—ConvNeXt-S with Blur and JPEG augmentation (0.5 probability) trained under Dataset Setting D—on the identical test split structure as HPBench (50 real images and 50 AI-generated images):

    • Top Model Accuracy: The ConvNeXt-S model achieved an accuracy of 87.0%87.0\% (13.0%13.0\% failure/misclassification rate).
    • Human Observer Accuracy: The human cohort on HPBench achieved an average accuracy of 61.3%61.3\% (38.7%38.7\% failure/misclassification rate).

    While contemporary AI detectors outperform human perception by 25.725.7 percentage points under matched evaluation conditions, the 13.0%13.0\% failure rate demonstrates that detecting cutting-edge generative content remains challenging for existing machine classifiers.

Coverage note — None was omitted; all primary benchmark designs (Fake2M, HPBench, MPBench), experimental results, tables, and comparative analyses are fully covered.

References

  1. 1.500px. https://500px.com/. Accessed: 2023-04-17. 3
  2. 2.bbc news: Fake trump arrest photos: How to spot an ai-generated image. https://www.bbc.com/news/world-us-canada-65069316. Accessed: 2023-04-18. 2
  3. 3.bbc news: Sony world photography award 2023: Winner refuses award after revealing ai creation. https://www.bbc.com/news/entertainment-arts-65296763. Accessed: 2023-04-17. 2
  4. 4.Deepfloyd. if. https://github.com/deep-floyd/IF. Accessed: 2023-06-7. 3, 9
  5. 5.Google images. https://images.google.com/. Accessed: 2023-04-17. 3
  6. 6.Midjourney. https://www.midjourney.com/. Accessed: 2023-04-17. 3, 9
  7. 7.photutorial: Number of photos (2023): Statistics, facts, & predictions. https://photutorial.com/. Accessed: 2023-04-18. 2
  8. 8.Stable diffusion v1.5. realistic vision v2.0. https://civitai.com/models/4201/realistic-vision-v20. Accessed: 2023-06-7. 3
  9. 9.Vishal Asnani, Xi Yin, Tal Hassner, Sijia Liu, and Xiaoming Liu. Proactive image manipulation detection. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 9
  10. 10.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022. 2
  11. 11.Jordan J Bird and Ahmad Lotfi. Cifake: Image classification and explainable identification of ai-generated synthetic images. arXiv preprint arXiv:2303.14126, 2023. 5
  12. 12.Andrew Brock, Jeff Donahue, and Karen Simonyan. Large scale GAN training for high fidelity natural image synthesis. International Conference on Learning Representations, 2019. 2, 9
  13. 13.Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. What makes fake images detectable? understanding properties that generalize. European Conference on Computer Vision, 2020. 9
  14. 14.Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha. Stargan v2: Diverse image synthesis for multiple domains. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. 4
  15. 15.Davide Cozzolino, Giovanni Poggi, and Luisa Verdoliva. Splicebuster: A new blind image splicing detector. IEEE International Workshop on Information Forensics and Security, 2015. 9
  16. 16.Hao Dang, Feng Liu, Joel Stehouwer, Xiaoming Liu, and Anil K. Jain. On the detection of digital face manipulation. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. 5
  17. 17.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. 2009. 4, 5
  18. 18.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems. 2
  19. 19.Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems, 2021. 3, 9
  20. 20.Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. Structure and content-guided video synthesis with diffusion models. arXiv preprint arXiv:2302.03011, 2023. 2
  21. 21.Joel Frank, Thorsten Eisenhofer, Lea Schonherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz. Leveraging frequency analysis for deep fake image recognition. International Conference on Machine Learning, 2020. 9
  22. 22.Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scene-based text-to-image generation with human priors. European Conference on Computer Vision, 2022. 9
  23. 23.Apurva Gandhi and Shomik Jain. Adversarial perturbations fool deepfake detectors. International Joint Conference on Neural Networks, 2020. 5
  24. 24.Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial networks. arXiv preprint arXiv:1406.2661, 2014. 2, 9
  25. 25.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2016. 7
  26. 26.Yinan He, Bei Gan, Siyu Chen, Yichun Zhou, Guojun Yin, Luchuan Song, Lu Sheng, Jing Shao, and Ziwei Liu. Forgerynet: A versatile benchmark for comprehensive forgery analysis. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. 5
  27. 27.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey A. Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022. 2
  28. 28.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 2020. 2, 4
  29. 29.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. International Conference on Learning Representations, 2018. 2
  30. 30.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Advances in Neural Information Processing Systems, 2022. 2
  31. 31.Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. Advances in Neural Information Processing Systems, 2021. 2, 3, 4
  32. 32.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019. 4, 5, 9
  33. 33.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. 2
  34. 34.Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. Maskgan: Towards diverse and interactive facial image manipulation. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. 4
  35. 35.Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 7
  36. 36.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. Advances in Neural Information Processing Systems, 2022. 4
  37. 37.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. arXiv preprint arXiv:2211.01095, 2022. 3
  38. 38.Zeyu Lu, Chengyue Wu, Xinyuan Chen, Yaohui Wang, Lei Bai, Yu Qiao, and Xihui Liu. Hierarchical diffusion autoencoders and disentangled image manipulation. arXiv preprint arXiv:2304.11829, 2023. 2
  39. 39.Francesco Marra, Diego Gragnaniello, Davide Cozzolino, and Luisa Verdoliva. Detection of gan-generated fake images over social networks. IEEE Multimedia Information Processing and Retrieval, 2018. 9
  40. 40.Lakshmanan Nataraj, Tajuddin Manhar Mohammed, B. S. Manjunath, Shivkumar Chandrasekaran, Arjuna Flenner, Jawadul H. Bappy, and Amit K. Roy-Chowdhury. Detecting GAN generated fake images using co-occurrence matrices. Media Watermarking, Security, and Forensics, 2019. 9
  41. 41.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. International Conference on Machine Learning, 2021. 2
  42. 42.Utkarsh Ojha, Yuheng Li, and Yong Jae Lee. Towards universal fake image detectors that generalize across generative models. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 7, 9
  43. 43.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022. 2
  44. 44.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. International Conference on Machine Learning, 2021. 7
  45. 45.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with CLIP latents. Advances in Neural Information Processing Systems, 2022. 2, 9
  46. 46.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 3, 9
  47. 47.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv preprint arXiv:2208.12242, 2022.
  48. 48.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 2022. 2, 9
  49. 49.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. LAION-5B: an open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 2022. 9
  50. 50.Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. Association for Computational Linguistics, 2018. 4, 5
  51. 51.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. International Conference on Learning Representations, 2020. 2, 4
  52. 52.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 2021.
  53. 53.Yang Song and Stefano Ermon. Improved techniques for training score-based generative models. Advances in Neural Information Processing Systems, 2020.
  54. 54.Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. International Conference on Learning Representations, 2021. 2
  55. 55.Luisa Verdoliva, Davide Cozzolino, and Koki Nagano. 2022 ieee image and video processing cup synthetic image detection. 2022. 5
  56. 56.Run Wang, Felix Juefei-Xu, Lei Ma, Xiaofei Xie, Yihao Huang, Jian Wang, and Yang Liu. Fakespotter: A simple yet robust baseline for spotting ai-synthesized fake faces. International Joint Conference on Artificial Intelligence, 2020. 5
  57. 57.Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A. Efros. Cnn-generated images are surprisingly easy to spot... for now. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020. 5, 7, 9
  58. 58.Yabin Wang, Zhiwu Huang, and Xiaopeng Hong. Benchmarking deepart detection. arXiv preprint arXiv:2302.14475, 2023. 5
  59. 59.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content-rich text-to-image generation. Advances in Neural Information Processing Systems, 2022. 9
  60. 60.Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. Cross-modal contrastive learning for text-to-image generation. IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021. 2, 9
  61. 61.Lvmin Zhang and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. arXiv preprint arXiv:2302.05543, 2023. 2
  62. 62.Xu Zhang, Svebor Karaman, and Shih-Fu Chang. Detecting and simulating artifacts in GAN fake images. IEEE International Workshop on Information Forensics and Security, 2019. 9

Citation

MLA
Lu, Z., et al. “Seeing Is Not Always Believing: Benchmarking Human and Model Perception of AI-Generated Images”. arXiv, 2023, http://arxiv.org/abs/2304.13023v3.
APA
Lu, Z., Huang, D., Bai, L., Qu, J., Wu, C., Liu, X., & Ouyang, W. (2023). Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images. arXiv. http://arxiv.org/abs/2304.13023v3
Chicago
Lu, Z., D. Huang, L. Bai, et al. 2023. “Seeing Is Not Always Believing: Benchmarking Human and Model Perception of AI-Generated Images”. arXiv. http://arxiv.org/abs/2304.13023v3.
Harvard
Lu, Z. et al. (2023) “Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2304.13023v3.
Vancouver
1. Lu Z, Huang D, Bai L, Qu J, Wu C, Liu X, Ouyang W (2023) Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images. arXiv

BibTeX

@article{lu2023seeing,
  title = {Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images},
  author = {Lu, Zeyu and Huang, Di and Bai, Lei and Qu, Jingjing and Wu, Chengyue and Liu, Xihui and Ouyang, Wanli},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2304.13023v3},
  eprint = {2304.13023}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors