DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models
Zijie J. WangEvan MontoyaDavid MunechikaHaoyang YangBenjamin HooverDuen Horng Chau
Introduces a public dataset of 14 million Stable Diffusion images and 1.8 million user-written prompts alongside generation hyperparameters to enable research on prompt engineering, generative model behavior, safety filtering, and deepfake detection.
Recent advances in artificial intelligence have enabled users to generate high-quality images from natural language descriptions. However, achieving desired visual results remains difficult because users often do not understand how generative models interpret prompts or how model settings influence generation quality. This trial-and-error approach limits usability and creates challenges in assessing real-world model behavior, safety risks, and misuse.
The article introduces and evaluates DIFFUSIONDB, the first large-scale, open-source dataset designed to systematically study text-to-image prompt engineering, model failures, and user interactions. The authors analyze the structural properties of real-world prompts, identify configurations that cause generation failures, and explore evidence of harmful usage.
To construct the dataset, the researchers collected 14 million images, 1.8 million unique prompts, and associated generation parameters generated by users of the Stable Diffusion platform on Discord in August 2022. The team organized this 6.5-terabyte repository into modular sub-folders, paired each image with complete technical metadata, and integrated automated safety classifiers to detect not-suitable-for-work (NSFW) content. They subsequently performed syntactic parsing, semantic embedding analysis, and statistical regression to evaluate prompt characteristics and failure modes.
The analysis produced several critical findings. First, prompt usage is heavily concentrated in short inputs (6 to 12 tokens) and dominated by English (98.3%), though a notable spike at the 75-token limit indicates that many users exceed model constraints without realizing their inputs are truncated. Second, statistical tests demonstrated that lower parameter values—specifically small step counts, small image dimensions, and negative prompt-weighting scales—are significantly correlated (p < 0.0001) with generation failures where images diverge completely from the input prompt. Non-English and extremely short prompts also trigger severe alignment errors. Third, prompt semantic representations diverge noticeably from the resulting image representations, highlighting model limitations in reproducing specific visual concepts like photorealistic human faces. Finally, the analysis revealed clear evidence of misuse, including tens of thousands of prompts targeting political figures and generating nonconsensual explicit content or visual disinformation.
These findings indicate that current text-to-image interfaces fail to adequately guide user input, leading to wasted compute cycles and suboptimal outputs. Furthermore, the persistence of toxic, nonconsensual, and misleading imagery despite platform moderation underscores significant compliance, reputational, and safety risks for organizations deploying generative models.
The article recommends developing intelligent user interfaces equipped with prompt autocompletion, parameter guardrails, and quality feedback to prevent common generation errors. Organizations should also leverage large prompt-image corpora to build automated deepfake detectors, train better-aligned models, and improve safety filtering mechanisms. Next steps should focus on establishing human-rated benchmarks for visual quality and expanding multilingual training data.
Confidence in these findings is strong regarding the observed user behavior and technical failure modes within Stable Diffusion. However, readers should consider key limitations: the dataset reflects early adopters on a single platform and may not fully represent novice behaviors or generalize across competing generative architectures.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). Introduces Latent Diffusion Models and Stable Diffusion, the core text-to-image architecture whose prompt behavior, hyperparameters, and generated outputs form the foundation of DiffusionDB.
- Paper: LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs, Christoph Schuhmann et al. (2021). Provides the foundational open web-scale image-text dataset used to train models like Stable Diffusion, contextualizing how web data informs the visual priors queried by user prompts.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). Demonstrates how cross-attention mechanisms in diffusion models map specific text tokens to spatial image regions, providing key background on how prompts directly govern generated image features.
- Paper: GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models, Alexander Quinn Nichol et al. (2022). Pioneers text-guided diffusion models with classifier-free guidance, establishing the standard conditioning techniques analyzed across prompt hyperparameter settings in DiffusionDB.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the fundamental mathematical and algorithmic framework of Denoising Diffusion Probabilistic Models underlying modern text-to-image synthesis.
- Paper: Optimizing Prompts for Text-to-Image Generation, Yaru Hao et al. (2023). Uses real-world human-engineered prompt galleries to train language models that automatically optimize user prompts for Stable Diffusion.
- Paper: VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models, Wenhao Wang et al. (2024). Extends the large-scale prompt-gallery paradigm from static text-to-image diffusion to dynamic text-to-video models by curating and analyzing millions of user prompts.
- Paper: Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models, Patrick Schramowski et al. (2023). Investigates safety vulnerabilities and inappropriate degeneration in Stable Diffusion using real-world user prompts, building directly on the safety concerns highlighted in DiffusionDB.
- Paper: What the DAAM: Interpreting Stable Diffusion Using Cross Attention, Raphael Tang et al. (2023). Develops cross-attention attribution maps to interpret word-level influences and semantic failure modes in Stable Diffusion across varied prompt syntaxes.
- Paper: Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models, Gowthami Somepalli et al. (2023). Investigates content memorization and replication risks in Stable Diffusion, directly addressing the copyright and generative fidelity issues raised in large-scale prompt studies.
- Paper: SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis, Dustin Podell et al. (2024). Advances the underlying Stable Diffusion architecture to higher resolutions and improved prompt adherence, directly addressing the fidelity and hyperparameter limitations analyzed in DiffusionDB.
- Paper: Initno: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization, Xiefan Guo et al. (2024). Addresses semantic alignment and subject omission errors in text-to-image diffusion by optimizing initial noise based on cross-attention responses to text prompts.
