SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders

Bartosz CywinskiKamil Deja

article2025ICML85 citations

Introduces SAeUron, a mechanistically interpretable method that trains sparse autoencoders across denoising timesteps to precisely identify and ablate unwanted visual concepts from text-to-image diffusion models without degrading overall generation quality or succumbing to adversarial attacks.

Listen

Modern text-to-image diffusion models can inadvertently create copyrighted, explicit, or harmful imagery, creating substantial safety, ethical, and legal risks. Retraining foundation models from scratch to remove unwanted content is computationally prohibitive, making machine unlearning an essential alternative. However, current unlearning techniques primarily rely on fine-tuning model weights or applying negative gradients. These methods function as opaque "black boxes" that frequently degrade overall image generation quality, struggle when removing multiple concepts sequentially, and often only superficially mask concepts, leaving models vulnerable to adversarial prompt attacks.

The article evaluates and demonstrates SAeUron, a transparent and interpretable unlearning framework that identifies and blocks specific visual concepts within text-to-image diffusion models using sparse autoencoders (SAEs), neural networks that decompose complex internal representations into distinct, human-interpretable components without modifying the base model's core weights.

To establish credibility, the evaluation was conducted across rigorous, standardized benchmarks using the Stable Diffusion architecture. The researchers trained lightweight SAEs on internal activations gathered across multiple denoising steps from key cross-attention blocks—specifically targeting an object-specialized block and a style-specialized block. Using an importance-scoring mechanism on validation prompt activations, SAeUron pinpoints concept-specific features and scales them down during inference. The system was benchmarked against leading unlearning methods on the UnlearnCanvas dataset (spanning 20 objects and 50 styles), the I2P inappropriate content benchmark, and the UnlearnDiffAtk adversarial evaluation framework.

The experimental findings demonstrate substantial improvements in both unlearning precision and operational stability. First, on the UnlearnCanvas benchmark, SAeUron achieved the highest overall average score of 94.03%, outperforming all existing methods in style unlearning (95.80% unlearning accuracy, 99.10% in-domain retain accuracy, and 99.40% cross-domain retain accuracy) while maintaining competitive performance in object unlearning (78.82% unlearning accuracy and over 95% retention). Second, SAeUron excelled at removing explicit content on the I2P benchmark, reducing total detected nudity instances to 18 (compared to 743 in the base model and 28 to 838 across competing baselines) while maintaining strong image quality. Third, under adversarial prompt attacks designed to force models into generating unlearned content, SAeUron exhibited high robustness, maintaining unlearning integrity with a post-attack success rate of only 1.40% on nudity prompts, whereas competing methods saw attack success rates surge up to 70–98%. Finally, SAeUron proved uniquely scalable for multi-concept erasure; in a stress test removing 49 out of 50 artistic styles simultaneously, the model retained 99.29% unlearning accuracy and over 95% preservation of the remaining style, avoiding the catastrophic degradation observed in fine-tuning approaches.

These findings indicate that generative concept removal can be achieved efficiently and transparently without retraining or destabilizing underlying foundation model parameters. By intervening only on interpretable latent features during inference, organizations can significantly reduce compute costs, storage demands (requiring only 0.2 GB of storage compared to 4 GB for fine-tuned checkpoints), and compliance risks. Furthermore, because target features can be inspected visually or annotated with language models prior to ablation, stakeholders gain mechanistic transparency over why and how content is blocked.

Based on these results, deploying inference-time sparse activation blocking is a highly effective strategy for closed-system image generation APIs and hosted platforms requiring rigorous safety filters and multi-concept unlearning. Organizations should prioritize mapping concept-specific feature dictionaries rather than maintaining separate fine-tuned model checkpoints for each safety constraint. Further work and pilot evaluations are recommended before full deployment, particularly to refine feature separation for closely related visual categories (such as visually overlapping animal classes) and to establish complementary guardrails for broad, abstract harms like hate or harassment that do not correspond to concrete visual geometries.

Decision-makers should note certain operational boundaries. SAeUron introduces a minor 1.92% latency overhead during image generation, and because it operates at inference time, it is primarily suitable for hosted or access-controlled platforms; open-source model releases could have the activation-blocking module bypassed by end users. Additionally, because the SAEs are trained on visual activations, the method's effectiveness is exceptionally high for well-defined visual elements (such as nudity, artistic styles, and objects) but diminishes for abstract concepts lacking consistent visual features.

No sufficiently relevant recommendations were found.

Cover for SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders

Abstract

Diffusion models, while powerful, can inadvertently generate harmful or undesirable content, raising significant ethical and safety concerns. Recent machine unlearning approaches offer potential solutions but often lack transparency, making it difficult to understand the changes they introduce to the base model. In this work, we introduce SAeUron, a novel method leveraging features learned by sparse autoencoders (SAEs) to remove unwanted concepts in text-to-image diffusion models. First, we demonstrate that SAEs, trained in an unsupervised manner on activations from multiple denoising timesteps of the diffusion model, capture sparse and interpretable features corresponding to specific concepts. Building on this, we propose a feature selection method that enables precise interventions on model activations to block targeted content while preserving overall performance. Our evaluation shows that SAeUron outperforms existing approaches on the UnlearnCanvas benchmark for concepts and style unlearning, and effectively eliminates nudity when evaluated with I2P. Moreover, we show that with a single SAE, we can remove multiple concepts simultaneously and that in contrast to other methods, SAeUron mitigates the possibility of generating unwanted content under adversarial attack. Code and checkpoints are available at GitHub.

Citation

MLA
Cywiński, B., and K. Deja. “SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders”. arXiv, 2025, http://arxiv.org/abs/2501.18052v3.
APA
Cywiński, B., & Deja, K. (2025). SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders. arXiv. http://arxiv.org/abs/2501.18052v3
Chicago
Cywiński, B., and K. Deja. 2025. “SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders”. arXiv. http://arxiv.org/abs/2501.18052v3.
Harvard
Cywiński, B. and Deja, K. (2025) “SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.18052v3.
Vancouver
1. Cywiński B, Deja K (2025) SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders. arXiv

BibTeX

@article{cywinski2025saeuron,
  title = {SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders},
  author = {Cywiński, Bartosz and Deja, Kamil},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.18052v3},
  eprint = {2501.18052}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/