EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers
Daiheng GaoShilin LuWenbo ZhouJiaming ChuJie ZhangMengxi JiaBang ZhangZhaoxin FanWeiming Zhang
Proposes a bi-level optimization framework combining LoRA tuning, attention map regularization, and self-contrastive learning to successfully remove unwanted concepts from modern flow-matching transformer models like Flux without degrading unrelated visual outputs.
Modern text-to-image generative models are trained on massive datasets from the open web, creating significant safety, copyright, and compliance risks when they inadvertently learn and generate inappropriate or sensitive content. While techniques to remove undesirable concepts have been widely developed for earlier image generation systems, recent architectural advances in modern models—such as the adoption of transformer blocks and sentence-level text encoders—have rendered these older removal methods ineffective. As organizations seek to deploy advanced image generation tools safely, they face a critical challenge: thoroughly removing unsafe content without degrading overall generation quality or breaking unrelated capabilities.
The article evaluates why existing concept removal methods fail on modern architectures and demonstrates a novel framework, EraseAnything, designed to eliminate targeted concepts while preserving the model's performance on unrelated concepts. The researchers formulate this challenge as a dual-level optimization process that simultaneously targets undesirable concept suppression and unrelated concept retention.
To establish credibility across diverse conditions, the researchers conducted extensive empirical evaluations using the modern Flux generative architecture across thousands of test prompts. The testing evaluated concrete objects, abstract artistic styles, social relationships, celebrity likenesses, and inappropriate material using standardized benchmarks such as the Inappropriate Image Prompt dataset, the MS-COCO captioning set, and the Ring-A-Bell adversarial prompt suite. In parallel, the team benchmarked EraseAnything against leading state-of-the-art unlearning methods and conducted a comprehensive user study involving twenty human evaluators assessing overall generation quality, diversity, and prompt adherence.
The article establishes several key findings. First, traditional unlearning methods transplanted directly to modern transformer architectures suffer from incomplete concept removal or severely harm unrelated generation quality, whereas the proposed method reduced explicit content from 605 instances in the baseline model to 199 while maintaining baseline-comparable image quality and text alignment. Second, the system achieved a 12.5% to 21.1% target concept accuracy rate across entities, styles, and relationships while keeping unrelated concept accuracy above 90%, outperforming existing ablative methods. Third, against adversarial prompt attacks designed to circumvent safety controls, the method maintained an attack success rate of only 2.5% on standard prompts and 11.9% on multi-step attacks, substantially lower than alternative approaches. Finally, human evaluation confirmed that the proposed framework achieved superior balance across cleanliness, preservation of unrelated details, and visual quality.
These findings indicate that generative image systems can be safely moderated and compliant with safety or copyright standards without costly retraining from scratch or risking performance degradation. The results show that older assumptions about feature localization no longer hold in modern transformer architectures, and successful safety filtering requires dynamic contrastive learning rather than static word-level suppression. This approach significantly lowers the operational risks and compliance barriers associated with deploying high-performance text-to-image systems.
Organizations deploying generative image models should adopt structured optimization frameworks to manage sensitive concepts, ensuring adapter weights are properly normalized when combining multiple concept filters. Further research and piloting are recommended to develop fine-grained intensity controls (such as adjustable sliders for concept strength) and to establish efficient integration methods when removing hundreds of distinct concepts simultaneously. Confidence in these results is high across evaluated tasks, though stakeholders should exercise caution when attempting massive-scale unlearning until combination limits are resolved in larger production environments.
- Paper: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2024). This paper introduces the rectified-flow transformer architecture that EraseAnything directly adapts, clarifying the model design and generation process its erasure method targets.
- Paper: Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient, Yongliang Wu et al. (2025). Its DoCo framework is a closely related prior approach to diffusion-model concept erasure, making its preservation and removal mechanisms useful context for EraseAnything’s alternative optimization.
- Paper: Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models, Patrick Schramowski et al. (2023). Safe Latent Diffusion establishes an important earlier approach to suppressing inappropriate content during image generation and provides context for EraseAnything’s safety evaluation.
- Paper: Linear Adversarial Concept Erasure, Shauli Ravfogel et al. (2022). Its formalization of concept erasure as suppressing target information while preserving useful representations provides foundational context for EraseAnything’s removal-retention objective.
- Paper: Cones: Concept Neurons in Diffusion Models for Customized Generation, Zhiheng Liu et al. (2023). Cones tests whether concepts reside in sparse diffusion-model neurons, a feature-localization assumption that helps frame EraseAnything’s challenge to static erasure strategies.
- Paper: MUSE: Machine Unlearning Six-Way Evaluation for Language Models, Weijia Shi et al. (2025). MUSE extends unlearning evaluation with a six-way test of forgetting, retained utility, leakage, scale, and repeated requests, offering a broader framework for assessing EraseAnything’s claims.
- Paper: RandAR: Decoder-only Autoregressive Visual Generation in Random Orders, Ziqi Pang et al. (2025). RandAR continues the study of transformer-based visual generation with a different sequence-ordering design, providing a useful follow-on test of how concept erasure interacts with evolving architectures.
