DRCT: Diffusion Reconstruction Contrastive Training towards Universal Detection of Diffusion Generated Images
Baoying ChenJishen ZengJianquan YangRui Yang
Proposes a contrastive training framework using diffusion-reconstructed hard samples and a two-million image benchmark across 16 generators to significantly improve the generalization of AI-generated image detectors against unseen diffusion models.
Rapid advancements in diffusion-based artificial intelligence have enabled the creation of photorealistic synthetic images, raising critical concerns regarding digital misinformation, copyright infringement, and election interference. While existing synthetic image detectors perform effectively on image formats seen during training, their accuracy collapses when evaluated against new, unseen generation architectures. The article addresses this generalization bottleneck by developing a universal training framework, Diffusion Reconstruction Contrastive Training (DRCT), designed to reliably identify synthetic images across diverse generative tools.
To develop and evaluate the method, the authors constructed DRCT-2M, a benchmark dataset comprising 2 million synthetic images spanning 16 diffusion architectures alongside an evaluation set of 136,000 real-world samples gathered from online platforms. The proposed framework operates on the premise of hard sample classification: by forcing a detector to distinguish authentic images from nearly identical reconstructed counterparts that carry subtle generative traces, the system learns more universal representations. The framework pairs diffusion-based reconstruction of real and generated images with a combined objective of contrastive and classification losses, and it was benchmarked across standard datasets against leading baseline detectors.
Key findings show substantial performance gains across multiple evaluation environments. On cross-model evaluations within the DRCT-2M benchmark, equipping standard detector backbones with DRCT elevated average detection accuracy from approximately 79–83% to between 91% and 97%. On independent GenImage cross-dataset tests, the framework improved baseline accuracy by 7 to 15 percentage points over standard approaches. In real-world wild tests, where conventional detectors experienced severe accuracy degradation down to 11–64%, the enhanced models sustained 82% to 97% detection accuracy. Furthermore, ablation experiments confirmed that both reconstructed sample training and contrastive loss integration provided distinct performance lifts of roughly 6.5% each, while maintaining up to 99% accuracy against image compression and resizing.
These results demonstrate that detectors trained on subtle generative artifacts rather than high-level semantic features achieve greater resilience and lower operational risk when deployed in production security environments. Organizations implementing media forensics, digital watermarking, or trust-and-safety filters can adopt this framework to bolster defenses without discarding existing detector architectures.
Moving forward, stakeholders deploying synthetic image detection tools should incorporate diffusion reconstruction and contrastive training pipelines into active defense monitoring, paired with access controls to counter potential adversarial exploitation. Future research should prioritize expanding the framework to detect localized image manipulations, refining detection for non-diffusion generators such as Generative Adversarial Networks (GANs), and enhancing feature interpretability.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This foundational account of diffusion denoising provides the generative-model background needed to understand DRCT’s diffusion-reconstruction hard-sample generation.
- Paper: CNN-Generated Images Are Surprisingly Easy to Spot… for Now, Sheng-Yu Wang et al. (2019). Its single-generator detector and cross-generator evaluation establish the generalization problem that DRCT tackles for diffusion-generated images.
No sufficiently relevant recommendations were found.
