Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models

Patrick SchramowskiManuel BrackBjörn DeiserothKristian Kersting

article2023CVPR636 citations

Proposes Safe Latent Diffusion, an inference-time guidance method that suppresses inappropriate and sexually explicit content in text-to-image diffusion models without requiring model retraining, external classifiers, or sacrificing image quality.

Listen

Modern artificial intelligence models that generate images from text descriptions are rapidly being deployed across commercial and public platforms. However, because these systems are trained on billions of uncurated image-text pairs scraped directly from the internet, they internalize and reproduce harmful societal biases and degenerated behaviors. These models frequently output inappropriate content—such as explicit nudity, violence, hate, self-harm, and severe ethnic stereotypes—even when user prompts contain no overtly toxic language. Standard post-generation filters often fail or are easily bypassed, creating substantial safety, regulatory, and reputational risks for organizations deploying generative systems.

The article demonstrates the extent of inappropriate degeneration in text-to-image systems and evaluates a new mitigation technique called Safe Latent Diffusion (SLD). The objective is to establish an effective method to remove and suppress harmful visual concepts directly during the image generation process without requiring expensive model retraining or degrading final image quality.

To conduct this evaluation, the researchers introduced the Inappropriate Image Prompts (I2P) benchmark, consisting of 4,703 real-world prompts covering seven distinct harmful categories (hate, harassment, violence, self-harm, sexual content, shocking imagery, and illegal acts). The study used Stable Diffusion to test image generation outcomes and evaluated outputs using automated classifiers, including Q16 and NudeNet. The team evaluated the proposed safety guidance approach against baseline image generation and alternative techniques across multiple parameter configurations, measuring both safety efficacy and visual quality via user studies and standardized image metrics.

The analysis yielded several critical findings. First, unmitigated Stable Diffusion generated inappropriate images at an overall rate of 39%, ranging up to 52% for shocking imagery, despite only 1.5% of the prompts being classified as toxic text. Second, training biases translated into pronounced ethnic disparities; for example, prompting the baseline model with the term "japanese body" produced explicit nudity over 75% of the time, compared to a 35% global average. Third, applying Safe Latent Diffusion reduced the overall probability of generating inappropriate content by over 75%, dropping the occurrence to just 9% under the strongest configuration. Finally, user preference studies confirmed that removing inappropriate elements via SLD did not harm image fidelity or text alignment, with over 59% to 63% of evaluators rating SLD outputs as equal to or better than baseline images.

These findings indicate that generative models can successfully self-correct by leveraging the conceptual knowledge acquired during pre-training. Rather than relying solely on post-hoc blocking classifiers or attempting to scrub all negative concepts from training corpora—which could prevent models from understanding what concepts to avoid—SLD steers the internal generation path away from defined unsafe concepts at inference time. This approach offers organizations a computationally efficient way to manage compliance and brand safety risks while maintaining generative capabilities.

Organizations deploying generative image systems should implement internal guidance mechanisms like SLD rather than relying solely on keyword filtering or output blockers. System operators should choose hyperparameter strengths tailored to their risk tolerance and user demographic, using moderate settings for adult creative tools and maximum suppression for sensitive environments, such as applications accessible to children. Furthermore, developers should transparently state which concepts are actively suppressed and combine inference-level safety steering with responsible data curation.

The study's primary limitations stem from the subjective nature of inappropriate imagery across different cultures, as well as the classifier's conservative false-positive rate when evaluating suppressed images. Additionally, while SLD significantly reduces ethnic and reporting biases, it does not completely eliminate them when minimal visual alteration settings are prioritized. Nevertheless, confidence in the core findings remains high, demonstrating that inference-time safety guidance is a viable, scalable, and effective operational safeguard for generative image models.

arXiv: 2211.05105
Cover for Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models

Abstract

Text-conditioned image generation models have recently achieved astonishing results in image quality and text alignment and are consequently employed in a fast-growing number of applications. Since they are highly data-driven, relying on billion-sized datasets randomly scraped from the internet, they also suffer, as we demonstrate, from degenerated and biased human behavior. In turn, they may even reinforce such biases. To help combat these undesired side effects, we present safe latent diffusion (SLD). Specifically, to measure the inappropriate degeneration due to unfiltered and imbalanced training sets, we establish a novel image generation test bed—inappropriate image prompts (I2P)—containing dedicated, real-world image-to-text prompts covering concepts such as nudity and violence. As our exhaustive empirical evaluation demonstrates, the introduced SLD removes and suppresses inappropriate image parts during the diffusion process, with no additional training required and no adverse effect on overall image quality or text alignment.1

Table of Contents

  • 1. Introduction
  • 2. Risks and Promises of Unfiltered Data
  • 3. Safe Latent Diffusion (SLD)
  • 4. Configuring Safe Latent Diffusion
  • 5. Inappropriate Image Prompts (I2P)
  • 6. Experimental Evaluation
  • 7. Discussion & Limitations
  • 8. Conclusion
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Safe Latent Diffusion Formulation

    model/method

    Safe Latent Diffusion (SLD) modifies the inference process of text-to-image diffusion models to steer generation away from undesirable or unsafe concepts without modifying model weights or using external classifiers. Standard classifier-free guidance adjusts the unconditional noise prediction ϵθ(zt)\epsilon_\theta(z_t) using a text prompt conditioning cpc_p:

    ϵ~θ(zt,cp)=ϵθ(zt)+sg(ϵθ(zt,cp)−ϵθ(zt))\tilde{\epsilon}_\theta(z_t, c_p) = \epsilon_\theta(z_t) + s_g(\epsilon_\theta(z_t, c_p) - \epsilon_\theta(z_t))

    where ztz_t is the latent representation at diffusion step tt, θ\theta denotes model parameters, and sg>0s_g > 0 is the prompt guidance scale.

    SLD incorporates an additional conditioning cSc_S representing an inappropriate safety concept (defined via text prompt SS) to push generation away from cSc_S:

    ϵˉθ(zt,cp,cS)=ϵθ(zt)+sg(ϵθ(zt,cp)−ϵθ(zt)−γ(zt,cp,cS))\bar{\epsilon}_\theta(z_t, c_p, c_S) = \epsilon_\theta(z_t) + s_g (\epsilon_\theta(z_t, c_p) - \epsilon_\theta(z_t) - \gamma(z_t, c_p, c_S))

    where the safety guidance adjustment γ(zt,cp,cS)\gamma(z_t, c_p, c_S) is defined as:

    γ(zt,cp,cS)=μ(cp,cS;sS,λ)(ϵθ(zt,cS)−ϵθ(zt))\gamma(z_t, c_p, c_S) = \mu(c_p, c_S; s_S, \lambda)(\epsilon_\theta(z_t, c_S) - \epsilon_\theta(z_t))

    Here, sSs_S is the safety guidance scale, and μ\mu applies element-wise thresholding and scaling:

    μ(cp,cS;sS,λ)={max⁡(1,∣Δ∣),where ϵθ(zt,cp)⊖ϵθ(zt,cS)<λ0,otherwise\mu(c_p, c_S; s_S, \lambda) = \begin{cases} \max(1, |\Delta|), & \text{where } \epsilon_\theta(z_t, c_p) \ominus \epsilon_\theta(z_t, c_S) < \lambda \\ 0, & \text{otherwise} \end{cases}

    Δ=sS(ϵθ(zt,cp)−ϵθ(zt,cS))\Delta = s_S(\epsilon_\theta(z_t, c_p) - \epsilon_\theta(z_t, c_S))

    The threshold λ\lambda acts as a boundary parameter in the latent space; only dimensions where the difference between prompt-conditioned and concept-conditioned estimates is below λ\lambda are actively shifted.

  2. Knowl 2 — Warm-up and Momentum Mechanisms in Safe Latent Diffusion

    model/method

    To preserve global composition while removing unsafe elements, Safe Latent Diffusion (SLD) introduces warm-up and momentum parameters into the safety guidance term γ\gamma:

    1. Warm-up period (δ\delta): Safety guidance is disabled during the first δ\delta diffusion steps, setting γ(zt,cp,cS)=0\gamma(z_t, c_p, c_S) = 0 when diffusion step t<δt < \delta. This allows the diffusion model to establish the global layout and structural composition before safety adjustments modify fine-grained details.

    2. Guidance momentum (νt\nu_t): To accelerate adjustments for dimensions consistently guided in the same direction, a momentum term νt\nu_t is added to the safety guidance γt\gamma_t:

    γt(zt,cp,cS)=μ(cp,cS;sS,λ)(ϵθ(zt,cS)−ϵθ(zt))+smνt\gamma_t(z_t, c_p, c_S) = \mu(c_p, c_S; s_S, \lambda)(\epsilon_\theta(z_t, c_S) - \epsilon_\theta(z_t)) + s_m \nu_t

    νt+1=βmνt+(1−βm)γt\nu_{t+1} = \beta_m \nu_t + (1 - \beta_m) \gamma_t

    where sm∈[0,1]s_m \in [0, 1] is the momentum scale, βm∈[0,1)\beta_m \in [0, 1) controls momentum volatility, and ν0=0\nu_0 = 0. Momentum accumulates even during the initial warm-up period (t<δt < \delta) to enable swift correction once the warm-up period concludes.

  3. Knowl 3 — Safe Latent Diffusion Hyperparameter Configurations

    model/method

    Safe Latent Diffusion defines a default text prompt for the unsafe concept SS:

    "hate, harassment, violence, suffering, humiliation, harm, suicide, sexual, nudity, bodily fluids, blood, obscene gestures, illegal activity, drug use, theft, vandalism, weapons, child abuse, brutality, cruelty"

    Four standard hyperparameter configurations provide varying levels of intervention strength:

    Configuration Warm-up (δ\delta) Scale (sSs_S) Threshold (λ\lambda) Momentum Scale (sms_m) Momentum βm\beta_m
    Hyp-Weak 15 200 0.0 0.0 -
    Hyp-Medium 10 1000 0.01 0.3 0.4
    Hyp-Strong 7 2000 0.025 0.5 0.7
    Hyp-Max 0 5000 1.0 0.5 0.7
    • Hyp-Weak handles superficial elements (e.g., light blood splatters) while leaving artistic representations and global layout largely untouched.
    • Hyp-Medium modifies prominent inappropriate content (e.g., altering exposed nudity into swimwear/bikinis).
    • Hyp-Strong aggressively suppresses unsafe content (e.g., generating full clothing/dresses) with minor overall image changes.
    • Hyp-Max applies immediate safety guidance across all diffusion steps to eliminate inappropriate content at the expense of greater deviation from the unguided image.
  4. Knowl 4 — Inappropriate Image Prompts (I2P) Benchmark

    definition

    The Inappropriate Image Prompts (I2P) benchmark is an evaluation dataset consisting of 4,703 unique, real-world text prompts designed to test the susceptibility of text-to-image diffusion models to inappropriate degeneration.

    Prompts were sourced by querying the Lexica database of user prompts for Stable Diffusion using 26 search terms spanning seven distinct categories of inappropriate content:

    1. Hate
    2. Harassment
    3. Violence
    4. Self-harm
    5. Sexual content
    6. Shocking images
    7. Illegal activity

    Each entry in the dataset includes the prompt text, associated generation hyperparameters (seed, guidance scale, image dimensions), category assignments, an estimate of inappropriate image generation probability, and a text toxicity score evaluated via the Perspective API.

  5. Knowl 5 — Quantitative Inappropriateness Mitigation on the I2P Benchmark

    data/table

    Safe Latent Diffusion (SLD) substantially reduces the probability of generating inappropriate images across all seven I2P categories when evaluated on Stable Diffusion (SD 1.4). Images were automatically evaluated by combining the Q16 classifier (broad inappropriate content) and NudeNet (detecting exposed genitalia).

    Category SD 1.4 Neg. Prompt Hyp-Weak Hyp-Medium Hyp-Strong Hyp-Max Exp. Max SD Exp. Max Hyp-Strong
    Hate 0.40 0.18 0.27 0.20 0.15 0.09 0.970.060.97_{0.06} 0.770.190.77_{0.19}
    Harassment 0.34 0.16 0.24 0.17 0.13 0.09 0.940.080.94_{0.08} 0.730.180.73_{0.18}
    Violence 0.43 0.24 0.36 0.23 0.17 0.14 0.890.040.89_{0.04} 0.790.130.79_{0.13}
    Self-harm 0.40 0.16 0.27 0.16 0.10 0.07 0.970.060.97_{0.06} 0.610.200.61_{0.20}
    Sexual 0.35 0.12 0.23 0.14 0.09 0.06 0.910.080.91_{0.08} 0.530.160.53_{0.16}
    Shocking 0.52 0.28 0.41 0.30 0.20 0.13 1.000.011.00_{0.01} 0.850.140.85_{0.14}
    Illegal activity 0.34 0.14 0.23 0.14 0.09 0.06 0.940.100.94_{0.10} 0.620.200.62_{0.20}
    Overall 0.39 0.18 0.29 0.19 0.13 0.09 0.960.070.96_{0.07} 0.720.190.72_{0.19}

    Overall generation of inappropriate content drops from 39% in the SD baseline to 9% under the Hyp-Max configuration (a relative reduction >75%). Standard negative prompting achieves a reduction to 18%, but SLD configurations (Hyp-Strong and Hyp-Max) achieve superior mitigation without restricting negative prompt usage for text conditioning.

  6. Knowl 6 — Image Quality and Text Alignment Impact of Safe Latent Diffusion

    data/table

    Safe Latent Diffusion achieves safety suppression while preserving image fidelity and text alignment relative to standard Stable Diffusion (SD). Evaluation includes Fréchet Inception Distance on MS-COCO (COCO FID-30k), CLIP similarity distance, and human user preference on the DrawBench benchmark:

    Configuration FID-30k ↓\downarrow User Fidelity (% ≥\ge SD) ↑\uparrow CLIP Distance ↓\downarrow User Alignment (% ≥\ge SD) ↑\uparrow
    SD Baseline 14.43 - 0.75 -
    Hyp-Weak 15.81 63.70 0.75 60.88
    Hyp-Medium 16.90 62.37 0.75 59.45
    Hyp-Strong 18.28 63.13 0.76 59.62
    Hyp-Max 18.76 63.60 0.76 60.58

    While automated FID scores increase moderately under stronger safety parameters (from 14.43 to 18.76), human evaluations on DrawBench show that over 62% of users judge SLD images as equal to or better than the SD baseline in image quality, and roughly 60% judge them equal to or better in text alignment.

  7. Knowl 7 — Ineffectiveness of Text-Level Toxicity Filtering for Image Generation Safety

    empirical result

    Filtering prompts based strictly on textual toxicity fails to prevent diffusion models from generating inappropriate images. Analysis of the 4,703 prompts in the I2P benchmark demonstrates:

    1. Only 1.5% of the prompts are classified as toxic according to the Perspective API.
    2. In contrast, unconstrained Stable Diffusion generates inappropriate imagery for 34% to 52% of these same prompts across different content categories.
    3. The correlation between prompt textual toxicity and the generation of inappropriate images is weak (Spearman rank correlation r=0.22r = 0.22).

    Consequently, text-based keyword blocking or toxicity filters on user prompts do not provide reliable guardrails against visual inappropriate degeneration.

  8. Knowl 8 — Mitigation of Dataset Reporting Biases and Demographic Disparities via SLD

    empirical result

    Unfiltered training data (LAION-5B) induces strong demographic reporting biases in Stable Diffusion. When evaluating the prompt template <country> body across 50 countries:

    • Baseline Stable Diffusion exhibits high variation in explicit nudity (measured by NudeNet detection of exposed genitalia), with a global average of 35.0% and extreme outliers such as Japan yielding >75% explicit nudity.
    • Applying Safe Latent Diffusion (Hyp-Strong) reduces overall explicit content by 75%, lowering Japan to 12.0% nudity and the global average to 9.25%.
    • However, a statistically significant correlation remains between per-country nudity rates with and without SLD (Spearman r=0.52,p=0.01r = 0.52, p = 0.01), indicating that inference-time guidance attenuates but does not completely remove deep representation biases present in the training data.
  9. Knowl 9 — Limitations and Dual-Use Risks of Safe Latent Diffusion

    limitation

    Safe Latent Diffusion has several notable limitations and dual-use considerations:

    1. Dependence on Model Knowledge: SLD relies entirely on the diffusion model's latent representation of the unsafe concept SS. If pre-training datasets are sanitized by completely removing all inappropriate concepts, the model loses the semantic representations required for SLD to target and steer away from them at inference time.
    2. Reversibility and Dual Use: The guidance vector γ(zt,cp,cS)\gamma(z_t, c_p, c_S) can be inverted by negating the safety guidance sign, creating a mechanism to deliberately maximize the generation of inappropriate content.
    3. Subjectivity and Potential Censorship: The definition of inappropriate content reflects specific normative frameworks; applying arbitrary concepts SS can be repurposed for broad and non-transparent content censorship.

Coverage note — None was omitted; all key contributions—including the SLD mathematical framework, hyperparameter configurations, the I2P benchmark, empirical results on inappropriateness, fidelity/alignment trade-offs, bias mitigation, and limitations—are covered.

References

  1. 1.Abubakar Abid, Maheen Farooqi, and James Zou. Persistent anti-muslim bias in large language models. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES), page 298–306. Association for Computing Machinery, 2021. 1, 2
  2. 2.Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. Text2live: Text-driven layered image and video editing. Preprint at https://arxiv.org/abs/2204.02491, 2022. 2
  3. 3.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of ACM Conference on Fairness, Accountability, and Transparency (FAccT), pages 610–623, 2021. 1, 2
  4. 4.Abeba Birhane and Vinay Uday Prabhu. Large image datasets: A pyrrhic win for computer vision? In Proceedings of IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1536–1546, 2021. 1, 2, 3
  5. 5.Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. Multimodal datasets: misogyny, pornography, and malignant stereotypes. CoRR, abs/2110.01963, 2021. 1, 2, 3
  6. 6.Tolga Bolukbasi, Kai-Wei Chang, James Y. Zou, Venkatesh Saligrama, and Adam Tauman Kalai. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings. In Proceedings of the Advances in Neural Information Processing Systems: Annual Conference on Neural Information Processing Systems (NeurIPS), pages 4349–4357. Curran Associates Inc., 2016. 1, 2
  7. 7.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners. In Proceedings of the Advances in Neural Information Processing Systems: Annual Conference on Neural Information Processing Systems (NeurIPS), 2020. 2
  8. 8.Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183–186, 2017. 1
  9. 9.Damien L. Crone, Stefan Bode, Carsten Murawski, and Simon M. Laham. The socio-moral image database (smid): A novel stimulus set for the study of social, moral and affective processes. PLOS ONE, 13(1), 2018. 13
  10. 10.Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole. On the genealogy of machine learning datasets: A critical history of imagenet. Big Data & Society, 8(2), 2021. 2
  11. 11.Shreyansh Gandhi, Samrat Kokkula, Abon Chaudhuri, Alessandro Magnani, Theban Stanley, Behzad Ahmadi, Venkatesh Kandaswamy, Omer Ovenc, and Shie Mannor. Scalable detection of offensive and non-compliant content / logo in product images. In Proceedings of IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 2020. 3
  12. 12.Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daume III, and Kate Crawford. Datasheets for datasets. Commun. ACM, 64(12):86–92, 2021. 6, 13
  13. 13.Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3356–3369. Association for Computational Linguistics, 2020. 1, 2, 12
  14. 14.Jonathan Gordon and Benjamin Van Durme. Reporting bias and knowledge acquisition. In Proceedings of the Workshop on Automated Knowledge Base Construction (AKBC), pages 25–30, 2013. 1
  15. 15.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. Preprint at https://arxiv.org/abs/2208.01626, 2022. 2
  16. 16.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Hugo Larochelle, Marc’Aurelio Ranzato, Raia Hadsell, Maria-Florina Balcan, and HsuanTien Lin, editors, Proceedings of the Advances in Neural Information Processing Systems: Annual Conference on Neural Information Processing Systems (NeurIPS), 2020. 4
  17. 17.Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. CoRR, abs/2207.12598, 2022. 4
  18. 18.Matthew Hutson. Robo-writers: the rise and risks of language-generating ai. Nature, 591:22–56, 2021. 1, 2
  19. 19.Abigail Z. Jacobs. Measurement and fairness. In Proceedings of ACM Conference on Fairness, Accountability, and Transparency (FAccT), pages 375–385. ACM, 2021. 2
  20. 20.Katrien Jacobs, Thomas Baudinette, and Alexandra Hambleton. Reflections on researching pornography across asia: voices from the region. Porn Studies, 7, 2020. 3
  21. 21.Sophie Jentzsch, Patrick Schramowski, Constantin A. Rothkopf, and Kristian Kersting. Semantics derived automatically from language corpora contain human-like moral choices. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES), pages 37–44, 2019. 2
  22. 22.Chen Karako and Putra Manggala. Using image fairness representations in diversity-based re-ranking for recommendations. In Tanja Mitrovic, Jie Zhang, Li Chen, and David Chin, editors, Adjunct Publication of the Conference on User Modeling, Adaptation and Personalization (UMAP), pages 23–28. ACM, 2018. 8
  23. 23.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. Preprint at https://arxiv.org/abs/2210.09276. 2
  24. 24.Agostina J. Larrazabal, Nicolas Nieto, Victoria Peterson, Diego H. Milone, and Enzo Ferrante. Gender imbalance in medical imaging datasets produces biased classifiers for computer-aided diagnosis. Proceedings of the National Academy of Sciences, 117(23):12592–12594, 2020. 2
  25. 25.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In Proceedings of the International Conference on Machine Learning (ICML). PMLR, 2022. 2, 3
  26. 26.Gaurav Parmar, Richard Zhang, and Jun-Yan Zhu. On aliased resizing and surprising subtleties in GAN evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 8
  27. 27.Fabio Petroni, Tim Rocktaschel, Sebastian Riedel, Patrick S. H. Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander H. Miller. Language models as knowledge bases? In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473. Association for Computational Linguistics, 2019. 1, 2
  28. 28.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML), 2021. 3, 6
  29. 29.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with CLIP latents. Preprint at https://arxiv.org/abs/2204.06125, 2022. 2, 6
  30. 30.Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426. Association for Computational Linguistics, 2020. 2
  31. 31.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. CoRR, abs/2205.11487, 2022. 2, 4, 8
  32. 32.Timo Schick, Sahana Udupa, and Hinrich Schutze. Self-diagnosis and self-debiasing: A proposal for reducing corpus-based bias in NLP. Transactions of the Association for Computational Linguistics (TACL), 9:1408–1424, 2021. 2
  33. 33.Patrick Schramowski, Christopher Tauchmann, and Kristian Kersting. Can machines help us answering question 16 in datasheets, and in turn reflecting on inappropriate content? In Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT), 2022. 3, 6, 13
  34. 34.Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A. Rothkopf, and Kristian Kersting. Large pre-trained language models contain human-like biases of what is right and wrong to do. Nature Machine Intelligence, 4(3), 2022. 2
  35. 35.Patrick Schramowski, Cigdem Turan, Sophie Jentzsch, Constantin A. Rothkopf, and Kristian Kersting. The moral choice machine. Frontiers Artif. Intell., 3:36, 2020. 2
  36. 36.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade W Gordon, Ross Wightman, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, Patrick Schramowski, Srivatsa R Kundurthy, Katherine Crowson, Ludwig Schmidt, Robert Kaczmarczyk, and Jenia Jitsev. Laion-5b: An open large-scale dataset for training next generation image-text models. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. 1, 2
  37. 37.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. LAION-400M: open dataset of clip-filtered 400 million image-text pairs. Preprint at https://arxiv.org/abs/2111.02114, 2021. 1, 2
  38. 38.Ryan Steed and Aylin Caliskan. Image representations learned with unsupervised pre-training contain human-like biases. In Proceedings of ACM Conference on Fairness, Accountability, and Transparency (FAccT), pages 701–713, 2021. 2
  39. 39.Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. Mitigating gender bias in natural language processing: Literature review. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), pages 1630–1640. Association for Computational Linguistics, 2019. 2
  40. 40.Dani Valevski, Matan Kalman, Yossi Matias, and Yaniv Leviathan. Unitune: Text-driven image editing by fine tuning an image generation model on a single image. Preprint at https://arxiv.org/abs/2210.09477. 2
  41. 41.Angelina Wang, Arvind Narayanan, and Olga Russakovsky. REVISE: A tool for measuring and mitigating bias in visual datasets. In Proceedings of European Conference on Computer Vision (ECCV), pages 733–751, 2020. 2
  42. 42.Robin Zheng. Why yellow fever isn’t flattering: A case against racial fetishes. Journal of the American Philosophical Association, 2, 2016. 3

Citation

MLA
Schramowski, P., et al. “Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models”. arXiv, 2022, http://arxiv.org/abs/2211.05105v4.
APA
Schramowski, P., Brack, M., Deiseroth, B., & Kersting, K. (2022). Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. arXiv. http://arxiv.org/abs/2211.05105v4
Chicago
Schramowski, P., M. Brack, B. Deiseroth, and K. Kersting. 2022. “Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models”. arXiv. http://arxiv.org/abs/2211.05105v4.
Harvard
Schramowski, P. et al. (2022) “Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.05105v4.
Vancouver
1. Schramowski P, Brack M, Deiseroth B, Kersting K (2022) Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models. arXiv

BibTeX

@article{schramowski2022safe,
  title = {Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models},
  author = {Schramowski, Patrick and Brack, Manuel and Deiseroth, Björn and Kersting, Kristian},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.05105v4},
  eprint = {2211.05105}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE