Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models
Patrick SchramowskiManuel BrackBjörn DeiserothKristian Kersting
Proposes Safe Latent Diffusion, an inference-time guidance method that suppresses inappropriate and sexually explicit content in text-to-image diffusion models without requiring model retraining, external classifiers, or sacrificing image quality.
Modern artificial intelligence models that generate images from text descriptions are rapidly being deployed across commercial and public platforms. However, because these systems are trained on billions of uncurated image-text pairs scraped directly from the internet, they internalize and reproduce harmful societal biases and degenerated behaviors. These models frequently output inappropriate content—such as explicit nudity, violence, hate, self-harm, and severe ethnic stereotypes—even when user prompts contain no overtly toxic language. Standard post-generation filters often fail or are easily bypassed, creating substantial safety, regulatory, and reputational risks for organizations deploying generative systems.
The article demonstrates the extent of inappropriate degeneration in text-to-image systems and evaluates a new mitigation technique called Safe Latent Diffusion (SLD). The objective is to establish an effective method to remove and suppress harmful visual concepts directly during the image generation process without requiring expensive model retraining or degrading final image quality.
To conduct this evaluation, the researchers introduced the Inappropriate Image Prompts (I2P) benchmark, consisting of 4,703 real-world prompts covering seven distinct harmful categories (hate, harassment, violence, self-harm, sexual content, shocking imagery, and illegal acts). The study used Stable Diffusion to test image generation outcomes and evaluated outputs using automated classifiers, including Q16 and NudeNet. The team evaluated the proposed safety guidance approach against baseline image generation and alternative techniques across multiple parameter configurations, measuring both safety efficacy and visual quality via user studies and standardized image metrics.
The analysis yielded several critical findings. First, unmitigated Stable Diffusion generated inappropriate images at an overall rate of 39%, ranging up to 52% for shocking imagery, despite only 1.5% of the prompts being classified as toxic text. Second, training biases translated into pronounced ethnic disparities; for example, prompting the baseline model with the term "japanese body" produced explicit nudity over 75% of the time, compared to a 35% global average. Third, applying Safe Latent Diffusion reduced the overall probability of generating inappropriate content by over 75%, dropping the occurrence to just 9% under the strongest configuration. Finally, user preference studies confirmed that removing inappropriate elements via SLD did not harm image fidelity or text alignment, with over 59% to 63% of evaluators rating SLD outputs as equal to or better than baseline images.
These findings indicate that generative models can successfully self-correct by leveraging the conceptual knowledge acquired during pre-training. Rather than relying solely on post-hoc blocking classifiers or attempting to scrub all negative concepts from training corpora—which could prevent models from understanding what concepts to avoid—SLD steers the internal generation path away from defined unsafe concepts at inference time. This approach offers organizations a computationally efficient way to manage compliance and brand safety risks while maintaining generative capabilities.
Organizations deploying generative image systems should implement internal guidance mechanisms like SLD rather than relying solely on keyword filtering or output blockers. System operators should choose hyperparameter strengths tailored to their risk tolerance and user demographic, using moderate settings for adult creative tools and maximum suppression for sensitive environments, such as applications accessible to children. Furthermore, developers should transparently state which concepts are actively suppressed and combine inference-level safety steering with responsible data curation.
The study's primary limitations stem from the subjective nature of inappropriate imagery across different cultures, as well as the classifier's conservative false-positive rate when evaluating suppressed images. Additionally, while SLD significantly reduces ethnic and reporting biases, it does not completely eliminate them when minimal visual alteration settings are prioritized. Nevertheless, confidence in the core findings remains high, demonstrating that inference-time safety guidance is a viable, scalable, and effective operational safeguard for generative image models.
- Paper: RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models, Samuel Gehman et al. (2020). Introduces the foundational benchmark and methodology for quantifying toxic degeneration and steering in generative models that Safe Latent Diffusion adapts to text-to-image systems.
- Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). Establishes classifier guidance and sampling mechanisms in diffusion models that underpin inference-time steering and safety guidance.
- Paper: GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models, Alexander Quinn Nichol et al. (2022). Develops text-guided diffusion models and classifier-free guidance, providing the core generative framework extended by Safe Latent Diffusion.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Presents the fundamental denoising diffusion probabilistic formulation on which latent diffusion architectures and their latent trajectories rely.
- Paper: Red Teaming Language Models with Language Models, Ethan Perez et al. (2022). Pioneers automated red-teaming paradigms for auditing generative model vulnerabilities that motivate the systematic prompt benchmarking in the source paper.
- Paper: Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient, Yongliang Wu et al. (2025). Extends the mitigation of inappropriate and sensitive concepts in diffusion models from inference-time guidance to permanent machine unlearning in model weights.
- Paper: OpenBias: Open-Set Bias Detection in Text-to-Image Generative Models, Moreno D'Incà et al. (2024). Generalizes the evaluation of systemic biases and inappropriate degeneration in text-to-image models to open-set, automated discovery without predefined prompt categories.
- Paper: RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural Prompts, Han Liu et al. (2023). Investigates adversarial prompt attacks designed to circumvent text moderation filters in text-to-image generation, exposing vulnerabilities that safety guidance mechanisms must defend against.
- Paper: On Provable Copyright Protection for Generative Models, Nikhil Vyas et al. (2023). Develops formal mathematical guarantees for inference-time protection against generating restricted training data concepts in diffusion models.
