DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

Zijie J. WangEvan MontoyaDavid MunechikaHaoyang YangBenjamin HooverDuen Horng Chau

article2023ACL524 citationsBest Paper Award, Honorable Mention

Introduces a public dataset of 14 million Stable Diffusion images and 1.8 million user-written prompts alongside generation hyperparameters to enable research on prompt engineering, generative model behavior, safety filtering, and deepfake detection.

Listen

Recent advances in artificial intelligence have enabled users to generate high-quality images from natural language descriptions. However, achieving desired visual results remains difficult because users often do not understand how generative models interpret prompts or how model settings influence generation quality. This trial-and-error approach limits usability and creates challenges in assessing real-world model behavior, safety risks, and misuse.

The article introduces and evaluates DIFFUSIONDB, the first large-scale, open-source dataset designed to systematically study text-to-image prompt engineering, model failures, and user interactions. The authors analyze the structural properties of real-world prompts, identify configurations that cause generation failures, and explore evidence of harmful usage.

To construct the dataset, the researchers collected 14 million images, 1.8 million unique prompts, and associated generation parameters generated by users of the Stable Diffusion platform on Discord in August 2022. The team organized this 6.5-terabyte repository into modular sub-folders, paired each image with complete technical metadata, and integrated automated safety classifiers to detect not-suitable-for-work (NSFW) content. They subsequently performed syntactic parsing, semantic embedding analysis, and statistical regression to evaluate prompt characteristics and failure modes.

The analysis produced several critical findings. First, prompt usage is heavily concentrated in short inputs (6 to 12 tokens) and dominated by English (98.3%), though a notable spike at the 75-token limit indicates that many users exceed model constraints without realizing their inputs are truncated. Second, statistical tests demonstrated that lower parameter values—specifically small step counts, small image dimensions, and negative prompt-weighting scales—are significantly correlated (p < 0.0001) with generation failures where images diverge completely from the input prompt. Non-English and extremely short prompts also trigger severe alignment errors. Third, prompt semantic representations diverge noticeably from the resulting image representations, highlighting model limitations in reproducing specific visual concepts like photorealistic human faces. Finally, the analysis revealed clear evidence of misuse, including tens of thousands of prompts targeting political figures and generating nonconsensual explicit content or visual disinformation.

These findings indicate that current text-to-image interfaces fail to adequately guide user input, leading to wasted compute cycles and suboptimal outputs. Furthermore, the persistence of toxic, nonconsensual, and misleading imagery despite platform moderation underscores significant compliance, reputational, and safety risks for organizations deploying generative models.

The article recommends developing intelligent user interfaces equipped with prompt autocompletion, parameter guardrails, and quality feedback to prevent common generation errors. Organizations should also leverage large prompt-image corpora to build automated deepfake detectors, train better-aligned models, and improve safety filtering mechanisms. Next steps should focus on establishing human-rated benchmarks for visual quality and expanding multilingual training data.

Confidence in these findings is strong regarding the observed user behavior and technical failure modes within Stable Diffusion. However, readers should consider key limitations: the dataset reflects early adopters on a single platform and may not fully represent novice behaviors or generalize across competing generative architectures.

arXiv: 2210.14896poloclub/diffusiondb
Cover for DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

Abstract

With recent advancements in diffusion models, users can generate high-quality images by writing text prompts in natural language. However, generating images with desired details requires proper prompts, and it is often unclear how a model reacts to different prompts or what the best prompts are. To help researchers tackle these critical challenges, we introduce DIFFUSIONDB, the first large-scale text-to-image prompt dataset totaling 6.5TB, containing 14 million images generated by Stable Diffusion, 1.8 million unique prompts, and hyperparameters specified by real users. We analyze the syntactic and semantic characteristics of prompts. We pinpoint specific hyperparameter values and prompt styles that can lead to model errors and present evidence of potentially harmful model usage, such as the generation of misinformation. The unprecedented scale and diversity of this human-actuated dataset provide exciting research opportunities in understanding the interplay between prompts and generative models, detecting deepfakes, and designing human-AI interaction tools to help users more easily use these models. DIFFUSIONDB is publicly available at: https://poloclub.github.io/diffusiondb.

Table of Contents

  • 1 Introduction
  • 2 Constructing DIFFUSIONDB
  • 2.1 Collecting User Generated Images
  • 2.2 Extracting Image Metadata
  • 2.3 Identifying NSFW Content
  • 2.4 Organizing DIFFUSIONDB
  • 2.5 Distributing DIFFUSIONDB
  • 3 Data Analysis
  • 3.1 Prompt Length
  • 3.2 Prompt Language
  • 3.3 Characterizing Prompts
  • 3.3.1 Prompt Syntactic Features
  • 3.3.2 Prompt Semantic Features
  • 3.4 Characterizing Images
  • 3.5 Stable Diffusion Error Analysis
  • 3.6 Potentially Harmful Uses
  • 4 Enabling New Research Directions
  • 5 Related Work
  • 6 Conclusion
  • 7 Limitations
  • 8 Ethics Statement
  • Acknowledgements
  • References
  • A Data Sheet for DIFFUSIONDB
  • ACL 2023 Responsible NLP Checklist

Knowls

  1. Knowl 1 — DiffusionDB Dataset Specification and Schema

    definition

    DiffusionDB is a large-scale dataset comprising 14 million images generated by Stable Diffusion (Version 1) based on 1,819,808 unique natural language prompts and hyperparameter configurations submitted by users on the public Stable Diffusion Discord server. The complete dataset totals approximately 6.5 TB and is released under a CC0 1.0 Universal Public Domain Dedication.

    The dataset is organized into 14,000 sub-folders, each containing 1,000 lossless WebP images and an accompanying JSON file mapping image identifiers to their metadata. A unified metadata table is provided in Apache Parquet format containing 13 fields per record:

    1. image_name: Unique filename generated via UUID Version 4.
    2. image_path: Path to the image file.
    3. prompt: The user-entered text prompt.
    4. seed: Random generation seed.
    5. cfg_scale: Classifier-Free Guidance scale.
    6. step: Number of denoising diffusion steps.
    7. sampler: Sampling algorithm chosen by the user.
    8. width: Image width in pixels.
    9. height: Image height in pixels.
    10. user_name_hash: SHA-256 hash of the Discord creator's username.
    11. timestamp: Generation timestamp.
    12. image_nsfw_score: Probability score of NSFW visual content.
    13. prompt_nsfw_score: Probability score of NSFW text content.

    A smaller subset, DiffusionDB-2M, samples one random image per multi-image generation grid (collage) to maximize prompt uniqueness and reduce storage footprint.

  2. Knowl 2 — Hyperparameter-Induced Generation Failures in Stable Diffusion

    empirical result

    To identify prompt-image misalignment, CLIP embeddings from a frozen ViT-L/14 model are computed for all prompts and images in DiffusionDB. The cosine distance dd between prompt text embedding vpv_p and generated image embedding viv_i,

    d=1−vp⋅vi∥vp∥∥vi∥d = 1 - \frac{v_p \cdot v_i}{\|v_p\| \|v_i\|}

    follows an empirical normal distribution N(μ=0.7123,σ2=0.04132)\mathcal{N}(\mu = 0.7123, \sigma^2 = 0.0413^2). Prompt-image pairs with d>μ+4σd > \mu + 4\sigma (d>0.8775d > 0.8775) that were not blurred by Stable Diffusion are designated as semantic failure cases, identifying 13,411 instances.

    A logistic regression analysis reveals that four primary generation hyperparameters—Classifier-Free Guidance (CFG) scale, diffusion denoising steps, image width, and image height—are each statistically significantly and negatively correlated with the probability of generating a failed image (p<0.0001p < 0.0001 for all four variables). The distribution of chosen sampling algorithms for failure cases also differs significantly from the overall population distribution (χ2=40873.11\chi^2 = 40873.11, p<0.0001p < 0.0001).

    Specific failure modes include:

    • Negative CFG scale values (e.g., CFG=−1\text{CFG} = -1), which invert conditional guidance and cause the model to generate imagery unrelated to the prompt (e.g., a bowl of soup instead of 'superman').
    • Excessively low step counts (e.g., step=2\text{step} = 2), leading to underdeveloped, blurry images.
    • Highly distorted aspect ratios or undersized dimensions (e.g., 64×51264 \times 512 pixels).
  3. Knowl 3 — Automated NSFW and Blurred Content Detection Pipeline

    model/method

    DiffusionDB annotates each prompt-image pair with safety scores to filter out Not-Safe-For-Work (NSFW) content that bypassed Discord moderation and Stable Diffusion's built-in filter:

    1. Prompt Toxicity Scoring: A pre-trained multilingual transformer predicts probabilities across six toxic categories: toxic, obscene, threat, insult, identity attack, and sexually explicit. The text NSFW score is computed as: Scoretext=max⁡(P(toxic),P(sexually explicit))\text{Score}_{\text{text}} = \max(P(\text{toxic}), P(\text{sexually explicit}))

    2. Image NSFW Scoring: A pre-trained EfficientNet classifier predicts probabilities over five classes: P(drawing)P(\text{drawing}), P(hentai)P(\text{hentai}), P(neutral)P(\text{neutral}), P(sexual)P(\text{sexual}), and P(porn)P(\text{porn}). The image NSFW score is computed as: Scoreimage=P(hentai)+P(sexual)+P(porn)\text{Score}_{\text{image}} = P(\text{hentai}) + P(\text{sexual}) + P(\text{porn})

    3. Built-in Blur Detection: Images pre-blurred by Stable Diffusion's filter are detected by applying a Laplacian convolution kernel with a variance threshold of 10. Pre-blurred images are assigned a fixed Scoreimage=2.0\text{Score}_{\text{image}} = 2.0 (achieving 100% precision and 100% recall on 50,000 randomly sampled test images).

    Evaluation on a manually annotated balanced sample of 5,000 images and 2,000 prompts yields:

    • Prompt NSFW detector: Precision = 0.36040.3604, unadjusted recall = 0.95650.9565, sampling-bias-adjusted recall = 0.66610.6661.
    • Image NSFW detector: Precision = 0.31500.3150, unadjusted recall = 0.97220.9722, sampling-bias-adjusted recall = 0.30370.3037.
  4. Knowl 4 — Prompt-Induced Semantic Failures Under Default Generation Settings

    empirical result

    When holding all generation hyperparameters close to default Stable Diffusion values, 1,100 unique severe prompt-image failure cases (d>μ+4σd > \mu + 4\sigma) remain in DiffusionDB.

    Statistical analysis indicates that these failures are heavily driven by prompt characteristics:

    • The token length of failure-inducing prompts is significantly shorter than the dataset average (one-tailed t=−23.7203t = -23.7203, p<0.0001p < 0.0001).
    • The proportion of English prompts among failure cases is significantly lower than the general dataset population (χ2=1024.56\chi^2 = 1024.56, p<0.0001p < 0.0001).
    • Prompts consisting primarily of non-English text, very few tokens, or emojis produce disproportionately high rates of semantic divergence from the generated image.
  5. Knowl 5 — Syntactic and Linguistic Characteristics of User Diffusion Prompts

    empirical result

    Analysis of the 1,819,808 unique prompts in DiffusionDB yields the following syntactic and linguistic distributions:

    • Length Distribution: Prompt lengths peak between 6 and 12 tokens. A distinct secondary spike occurs at 75 tokens, corresponding to the hard truncation limit of Stable Diffusion's CLIP tokenizer (which discards tokens beyond 75, excluding start and end tokens).
    • Language Distribution: English accounts for 98.3% of unique prompts. Among non-English prompts, 34 languages contain at least 100 prompts, led by German (5.2k prompts), French (4.6k), Italian (3.2k), and Spanish (3.0k).
    • Syntactic Composition: Prompts are predominantly composed of comma-separated fragments. Parsing via Named Entity Recognition (NER) and dependency parsing (noun phrases and root extraction) reveals standardized modifier vocabularies:
      • Quality boosters: 'highly detailed', 'intricate', 'sharp focus', '8k'.
      • Artist signatures: 'greg rutkowski', 'artstation'.
      • Medium and lighting styles: 'digital painting', 'oil painting', 'portrait painting', 'volumetric lighting', 'atmospheric lighting', 'studio lighting'.
  6. Knowl 6 — Semantic Embedding Shift Between Prompt and Image Representations

    empirical result

    Visualizing 768-dimensional CLIP ViT-L/14 text embeddings of 1.8 million prompts via UMAP (n_neighbors=60,min_dist=0.1n\_\text{neighbors} = 60, \text{min}\_\text{dist} = 0.1) combined with Kernel Density Estimation (multivariate Gaussian kernel, Silverman bandwidth) identifies two primary clusters: art-related prompts and photography-related prompts. Photography prompts further separate into non-human objects and celebrity subjects.

    When projecting 2 million corresponding generated image embeddings into the identical UMAP space, the distribution shifts relative to prompt space. Prompts containing cinematic or film keywords ('movie' cluster) map directly into art and portrait clusters in the image embedding space. This indicates a semantic gap between text prompt intent and generated visual output, hypothesized to stem from the model's difficulty in synthesizing photorealistic faces without shifting toward artistic/portrait styles.

  7. Knowl 7 — Empirical Identification of Malicious and Harmful Generative Prompts

    empirical result

    Named entity recognition and keyword analysis across DiffusionDB prompts identify real-world usage patterns involving harmful and deceptive generations:

    • Political Misrepresentation: Prompts frequently target prominent political figures, including over 65,000 image generations referencing 'Donald Trump' and over 48,000 referencing 'Joe Biden', often generating derogatory or incriminating depictions (e.g., depicted in handcuffs or as grotesque creatures).
    • Nonconsensual Sexual Content: Female celebrities appear frequently in prompt logs (ranking immediately after artists and politicians), with a substantial subset requesting sexually explicit or nonconsensual pornography.
    • Disinformation and Propaganda: Keyword queries reveal prompt generations designed to produce deceptive visual artifacts, such as conspiratorial medical claims ('scientists putting microchips into a vaccine') and falsified war imagery regarding the Russo-Ukrainian War.
  8. Knowl 8 — Limitations of DiffusionDB

    limitation

    The DiffusionDB dataset and its empirical analyses have four principal limitations:

    1. Residual Unsafe Content: Despite automated filtering and Discord moderation, some NSFW, violent, or toxic prompt-image pairs remain in the dataset.
    2. Data Source and Demographic Skew: Prompts originate from early Discord beta users who were predominantly AI art enthusiasts with prior generative modeling experience; prompt distributions and styles may not represent novice users or domain-specific disciplines (e.g., clinical or medical imaging).
    3. Image Quality Metric Constraints: Joint CLIP text-image cosine distance measures semantic alignment between prompt and image, but does not provide an objective measure of perceptual quality, visual realism, aesthetic fidelity, or generation artifacts.
    4. Cross-Model Generalizability: Prompt structures (such as heavy reliance on comma-separated modifier tags) reflect heuristic conventions developed specifically for Stable Diffusion and may not transfer directly to other text-to-image architectures like DALL-E 2 or Midjourney.

Coverage note — None was omitted; all key contributions including dataset construction, NSFW detection pipelines, syntactic/semantic analyses, error analyses, harmful use analyses, and limitations are fully covered.

References

  1. 1.Apache. 2013. Apache Parquet: Open Source, Column-oriented Data File Format Designed for Efficient Data Storage and Retrieval.
  2. 2.Stephen Bach, Victor Sanh, Zheng Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-david, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Fries, Maged Alshaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-jian Jiang, and Alexander Rush. 2022. PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations.
  3. 3.Ali Borji. 2022. Generated Faces in the Wild: Quantitative Comparison of Stable Diffusion, Midjourney and DALL-E 2. arXiv 2210.00586.
  4. 4.Gwern Branwen. 2020. GPT-3 Creative Fiction.
  5. 5.Pierre Chambon, Christian Bluethgen, Curtis P. Langlotz, and Akshay Chaudhari. 2022. Adapting Pre-trained Vision-Language Foundational Models to Medical Imaging Domains. arXiv 2210.04133.
  6. 6.Alex Clark. 2015. Pillow: Python Imaging Library (Fork).
  7. 7.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daume III, and Kate Crawford. 2020. Datasheets for Datasets. arXiv:1803.09010 [cs].
  8. 8.Google. 2010. Comparative Study of WebP, JPEG and JPEG 2000.
  9. 9.Thomas Griffiths, Michael Jordan, Joshua Tenenbaum, and David Blei. 2003. Hierarchical topic models and the nested chinese restaurant process. In Advances in Neural Information Processing Systems, volume 16.
  10. 10.Laura Hanu and Unitary team. 2020. Detoxify: Toxic Comment Classification with Pytorch Lightning and Transformers.
  11. 11.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. 2022. Imagen Video: High Definition Video Generation with Diffusion Models. arXiv 2210.02303.
  12. 12.Oleksii Holub. 2017. DiscordChatExporter: Exports Discord Chat Logs to a File.
  13. 13.David Holz. 2022. Midjourney: Exploring New Mediums of Thought and Expanding the Imaginative Powers of the Human Species.
  14. 14.Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020. spaCy: Industrial-strength natural language processing in python.
  15. 15.Harold Hotelling. 1936. Relations Between Two Sets of Variates. Biometrika, 28.
  16. 16.Eero Hyvonen and Eetu Makela. 2006. Semantic Autocompletion. In The Semantic Web – ASWC 2006, volume 4185.
  17. 17.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers.
  18. 18.P. Leach, M. Mealling, and R. Salz. 2005. A Universally Unique IDentifier (UUID) URN Namespace. Technical report, RFC Editor.
  19. 19.Li-Jia Li, Chong Wang, Yongwhan Lim, David M. Blei, and Li Fei-Fei. 2010. Building and using a semantivisual image hierarchy. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition.
  20. 20.Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2022. Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing. ACM Computing Surveys.
  21. 21.Vivian Liu and Lydia B Chilton. 2022. Design Guidelines for Prompt Engineering Text-to-Image Generative Models. In CHI Conference on Human Factors in Computing Systems.
  22. 22.Maria Teresa Llano, Mark d’Inverno, Matthew Yee-King, Jon McCormack, Alon Ilsar, Alison Pease, and Simon Colton. 2022. Explainable Computational Creativity. arXiv 2205.05682.
  23. 23.Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).
  24. 24.Scott M. Lundberg and Su-In Lee. 2017. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17.
  25. 25.Leland McInnes, John Healy, and James Melville. 2020. UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction. arXiv:1802.03426 [cs, stat].
  26. 26.Yisroel Mirsky and Wenke Lee. 2022. The Creation and Detection of Deepfakes: A Survey. ACM Computing Surveys, 54.
  27. 27.Jonas Oppenlaender. 2022. A Taxonomy of Prompt Modifiers for Text-To-Image Generation. arXiv 2204.13988.
  28. 28.Nikita Pavlichenko and Dmitry Ustalov. 2022. Best Prompts for Text-to-Image Models and How to Find Them. arXiv 2209.11711.
  29. 29.Patrick Von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf. 2022. Diffusers: State-of-the-art diffusion models.
  30. 30.Han Qiao, Vivian Liu, and Lydia Chilton. 2022. Initial Images: Using Image Prompts to Improve Subject Representation in Multimodal AI Generated Art. In Creativity and Cognition.
  31. 31.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research.
  32. 32.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv 2204.06125.
  33. 33.Laria Reynolds and Kyle McDonell. 2021. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems.
  34. 34.Leonard Richardson. 2007. Beautiful Soup Documentation.
  35. 35.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  36. 36.Kevin Roose. 2022. An A.I.-Generated Picture Won an Art Prize. Artists Aren’t Happy.
  37. 37.Murray Rosenblatt. 1956. Remarks on Some Nonparametric Estimates of a Density Function. The Annals of Mathematical Statistics, 27.
  38. 38.Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning To Retrieve Prompts for In-Context Learning. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies.
  39. 39.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. arXiv 2205.11487.
  40. 40.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. 2022. LAION-5B: An open large-scale dataset for training next generation image-text models. arXiv 2210.08402.
  41. 41.Sharif Shameem. 2022. Lexica: Building a Creative Tool for the Future.
  42. 42.Bernard W Silverman. 2018. Density Estimation for Statistics and Data Analysis.
  43. 43.StabilityAI. 2022a. Stable Diffusion Discord Server Rules.
  44. 44.StabilityAI. 2022b. Stable Diffusion Dream Studio beta Terms of Service.
  45. 45.Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of Machine Learning Research, 9.
  46. 46.Weixin Wang, Hui Wang, Guozhong Dai, and Hongan Wang. 2006. Visualization of large hierarchical data by circle packing. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems.
  47. 47.Kyle Wiggers. 2022. Deepfakes for all: Uncensored AI art model prompts ethics questions.
  48. 48.Simon Willison, Adam Stacoviak, and Jerod Stacoviak. 2022. Stable Diffusion Breaks the Internet.

Citation

MLA
Wang, Z. J., et al. “DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 893–911, https://doi.org/10.18653/v1/2023.acl-long.51.
APA
Wang, Z. J., Montoya, E., Munechika, D., Yang, H., Hoover, B., & Chau, D. H. (2023). DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 893–911. https://doi.org/10.18653/v1/2023.acl-long.51
Chicago
Wang, Z. J., E. Montoya, D. Munechika, H. Yang, B. Hoover, and D. H. Chau. 2023. “DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 893–911. https://doi.org/10.18653/v1/2023.acl-long.51.
Harvard
Wang, Z.J. et al. (2023) “DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 893–911. Available at: https://doi.org/10.18653/v1/2023.acl-long.51.
Vancouver
1. Wang ZJ, Montoya E, Munechika D, Yang H, Hoover B, Chau DH (2023) DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 893–911

BibTeX

@inproceedings{wang-etal-2023-diffusiondb,
    title = "{D}iffusion{DB}: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models",
    author = "Wang, Zijie J.  and
      Montoya, Evan  and
      Munechika, David  and
      Yang, Haoyang  and
      Hoover, Benjamin  and
      Chau, Duen Horng",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.51/",
    doi = "10.18653/v1/2023.acl-long.51",
    pages = "893--911"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/