Seeing is not always believing: Benchmarking Human and Model Perception of AI-Generated Images
Zeyu LuDi HuangLei BaiJingjing QuChengyue WuXihui LiuWanli Ouyang
Presents a two-million-image dataset alongside comprehensive human and algorithmic benchmarks, revealing that people fail to identify state-of-the-art AI-generated images nearly 39% of the time while specialized detection models reduce this error rate to 13%.
Photographs have historically served as a reliable medium for recording factual events and personal experiences. However, rapid advancements in artificial intelligence (AI) text-to-image synthesis now enable the generation of hyper-realistic visual content, raising significant risks regarding misinformation, fraudulent media, and the erosion of public trust. The article addresses the urgent problem of determining whether humans and automated detection algorithms can reliably distinguish state-of-the-art AI-generated images from authentic photographs.
The main objective of the article is to benchmark human discernment against cutting-edge AI detection models when identifying modern, high-quality synthetic images. To support this evaluation, the researchers developed Fake2M, a comprehensive public dataset containing over 2 million synthetic images generated by prominent diffusion, generative adversarial, and autoregressive models alongside matched real photographs. Using this resource, the article establishes HPBench to test human perception across a diverse group of 50 participants and MPBench to assess multiple automated vision backbones across 11 validation sets.
The findings reveal that modern AI-generated images readily deceive human vision, resulting in an overall human misclassification rate of 38.7% (61.3% accuracy). Notably, participants correctly classified synthetic images as fake only 55.8% of the time, and their confidence in authentic photography was also degraded, correctly sorting real online images only 66.9% of the time. Human participants performed best when evaluating portraits and multi-person scenes (up to 67.5% accuracy) due to perceptual sensitivity to human anatomical details, but dropped to 50.8% accuracy on non-human objects. In contrast, automated algorithms significantly outperformed humans, with the top-performing model achieving an 87% accuracy (a 13% error rate) on the same evaluation benchmark. However, model performance varied substantially across validation sets, and no single architecture consistently dominated across all data configurations.
These results demonstrate that human visual inspection alone is no longer an adequate defense against AI-generated misinformation. While automated detection systems offer superior accuracy, their sensitivity to specific generative architectures and sampling parameters creates operational risks and vulnerabilities to novel synthesis techniques. Organizations relying on visual evidence for policy, security, or compliance must recognize that unassisted staff cannot reliably authenticate images, and existing algorithmic safeguards remain imperfect.
To mitigate these risks, decision-makers should deploy multi-model automated screening tools trained on highly diverse generative datasets, rather than relying on single-architecture detectors. Organizations should also provide targeted verification training to staff while establishing formal digital provenance tracking. Further research and technical pilots are needed to improve detector generalization across unseen generative techniques and to develop balanced architectures that effectively spot both authentic and synthetic visual media.
The study's conclusions should be interpreted in light of certain limitations, including the controlled 50-participant human sample and the rapid evolution of generative models, which may soon eliminate current visual artifacts such as texture over-smoothing. Nevertheless, there is high confidence in the core finding that modern generative tools have surpassed human ability to discern fake media without automated assistance.
- Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). This foundational work establishes the early baseline for cross-generator synthetic image detection, providing the technical context against which HPBench and MPBench evaluate modern diffusion-era models.
- Paper: FaceForensics++: Learning to Detect Manipulated Facial Images, Andreas Rössler et al. (2019). It introduced early benchmarks and human baseline testing for manipulated visual media, establishing standard evaluation methodologies adapted and expanded by modern AI-generation detection studies.
- Paper: Harmonizing the object recognition strategies of deep neural networks with humans, Thomas Fel et al. (2022). It details how human visual feature recognition fundamentally diverges from deep neural network strategies, explaining why human and algorithmic error profiles differ on synthetic imagery.
- Paper: Toward Verifiable and Reproducible Human Evaluation for Text-to-Image Generation, Mayu Otani et al. (2023). It systematically explores the limitations of automated visual metrics compared to standardized human perceptual evaluation in text-to-image synthesis.
- Paper: Rethinking the Up-Sampling Operations in CNN-Based Generative Network for Generalizable Deepfake Detection, Chuangchuang Tan et al. (2024). It tackles the generalization failure across unseen generative architectures highlighted by the source by proposing upsampling artifact detection in the spatial domain.
- Paper: Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning, Chuangchuang Tan et al. (2024). It directly extends efforts to overcome detector sensitivity to specific generation models by learning frequency-space representations across multi-generator benchmarks.
- Paper: Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks, Mehrdad Saberi et al. (2024). It evaluates the adversarial limits and practical evasion vulnerabilities of automated synthetic image detectors that organizations rely on to counter unassisted human fallibility.
- Paper: WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models, Changhoon Kim et al. (2024). It explores model weight fingerprinting as a complementary digital provenance defense to address the shortcomings of standalone visual perception and post-hoc detection.
- Paper: VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation, Max Ku et al. (2024). It leverages multimodal LLMs to develop explainable perceptual and semantic quality metrics that bridge the gap between automated detection and human assessment.
