EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers

Daiheng GaoShilin LuWenbo ZhouJiaming ChuJie ZhangMengxi JiaBang ZhangZhaoxin FanWeiming Zhang

article2025ICML96 citations

Proposes a bi-level optimization framework combining LoRA tuning, attention map regularization, and self-contrastive learning to successfully remove unwanted concepts from modern flow-matching transformer models like Flux without degrading unrelated visual outputs.

Listen

Modern text-to-image generative models are trained on massive datasets from the open web, creating significant safety, copyright, and compliance risks when they inadvertently learn and generate inappropriate or sensitive content. While techniques to remove undesirable concepts have been widely developed for earlier image generation systems, recent architectural advances in modern models—such as the adoption of transformer blocks and sentence-level text encoders—have rendered these older removal methods ineffective. As organizations seek to deploy advanced image generation tools safely, they face a critical challenge: thoroughly removing unsafe content without degrading overall generation quality or breaking unrelated capabilities.

The article evaluates why existing concept removal methods fail on modern architectures and demonstrates a novel framework, EraseAnything, designed to eliminate targeted concepts while preserving the model's performance on unrelated concepts. The researchers formulate this challenge as a dual-level optimization process that simultaneously targets undesirable concept suppression and unrelated concept retention.

To establish credibility across diverse conditions, the researchers conducted extensive empirical evaluations using the modern Flux generative architecture across thousands of test prompts. The testing evaluated concrete objects, abstract artistic styles, social relationships, celebrity likenesses, and inappropriate material using standardized benchmarks such as the Inappropriate Image Prompt dataset, the MS-COCO captioning set, and the Ring-A-Bell adversarial prompt suite. In parallel, the team benchmarked EraseAnything against leading state-of-the-art unlearning methods and conducted a comprehensive user study involving twenty human evaluators assessing overall generation quality, diversity, and prompt adherence.

The article establishes several key findings. First, traditional unlearning methods transplanted directly to modern transformer architectures suffer from incomplete concept removal or severely harm unrelated generation quality, whereas the proposed method reduced explicit content from 605 instances in the baseline model to 199 while maintaining baseline-comparable image quality and text alignment. Second, the system achieved a 12.5% to 21.1% target concept accuracy rate across entities, styles, and relationships while keeping unrelated concept accuracy above 90%, outperforming existing ablative methods. Third, against adversarial prompt attacks designed to circumvent safety controls, the method maintained an attack success rate of only 2.5% on standard prompts and 11.9% on multi-step attacks, substantially lower than alternative approaches. Finally, human evaluation confirmed that the proposed framework achieved superior balance across cleanliness, preservation of unrelated details, and visual quality.

These findings indicate that generative image systems can be safely moderated and compliant with safety or copyright standards without costly retraining from scratch or risking performance degradation. The results show that older assumptions about feature localization no longer hold in modern transformer architectures, and successful safety filtering requires dynamic contrastive learning rather than static word-level suppression. This approach significantly lowers the operational risks and compliance barriers associated with deploying high-performance text-to-image systems.

Organizations deploying generative image models should adopt structured optimization frameworks to manage sensitive concepts, ensuring adapter weights are properly normalized when combining multiple concept filters. Further research and piloting are recommended to develop fine-grained intensity controls (such as adjustable sliders for concept strength) and to establish efficient integration methods when removing hundreds of distinct concepts simultaneously. Confidence in these results is high across evaluated tasks, though stakeholders should exercise caution when attempting massive-scale unlearning until combination limits are resolved in larger production environments.

Cover for EraseAnything: Enabling Concept Erasure in Rectified Flow Transformers

Abstract

Removing unwanted concepts from large-scale text-to-image (T2I) diffusion models while maintaining their overall generative quality remains an open challenge. This difficulty is especially pronounced in emerging paradigms, such as Stable Diffusion (SD) v3 and Flux, which incorporate flow matching and transformer-based architectures. These advancements limit the transferability of existing concept-erasure techniques that were originally designed for the previous T2I paradigm (e.g., SD v1.4). In this work, we introduce EraseAnything, the first method specifically developed to address concept erasure within the latest flow-based T2I framework. We formulate concept erasure as a bi-level optimization problem, employing LoRA-based parameter tuning and an attention map regularizer to selectively suppress undesirable activations. Furthermore, we propose a self-contrastive learning strategy to ensure that removing unwanted concepts does not inadvertently harm performance on unrelated ones. Experimental results demonstrate that EraseAnything successfully fills the research gap left by earlier methods in this new T2I paradigm, achieving SOTA performance across a wide range of concept erasure tasks.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 2.1. T2I Diffusion Models
  • 2.2. Concept Erasing
  • 2.3. Bi-level optimization (BO)
  • 3. Obstacles in migrating concept erasure methods to Flux
  • 4. Method
  • 4.1. Overview
  • 4.2. Bi-Level Finetuning Framework
  • 5. Experiments
  • 5.1. Implementation Details
  • 5.2. Results
  • 5.3. Ablation study
  • 6. Conclusion
  • Impact Statement
  • Ethical Aspects
  • Societal Consequences
  • Acknowledgment
  • References
  • A. Flux Architecture
  • B. Pattern of prompt & Black box attack
  • C. Prompt-related supplementary material
  • C.1. A heuristic cir sampling method
  • C.2. Complete list of Entity, Abstraction, Relationship
  • D. Derivative of Reverse Self-Contrastive Loss
  • E. User Study
  • F. Others
  • F.1. Celebrity
  • F.2. More Experimental Results
  • F.3. Limitations

Knowls

  1. Knowl 1 — EraseAnything frames concept removal and preservation as bi-level optimization

    model/method

    EraseAnything adapts concept erasure to Flux by tuning a LoRA adapter, denoted by Δθ\Delta\theta, with two linked objectives. The lower level suppresses a set of target concepts DunD_{\mathrm{un}} using an ESD-style erasure loss Lesd\mathcal{L}_{\mathrm{esd}} and a target-attention penalty Lattn\mathcal{L}_{\mathrm{attn}}. The upper level tunes the same adapter to preserve unrelated concepts DirD_{\mathrm{ir}} using an image-generation preservation loss Llora\mathcal{L}_{\mathrm{lora}} and a reverse self-contrastive loss Lrsc\mathcal{L}_{\mathrm{rsc}}:

    min⁡Δθ∗  Llora(Δθ∗;Dir)+Lrsc(Δθ∗;Dir),subject to Δθ∗∈arg⁡min⁡Δθ[Lesd(Δθ;Dun)+Lattn(Δθ;Dun)].\min_{\Delta\theta^*}\; \mathcal{L}_{\mathrm{lora}}(\Delta\theta^*;D_{\mathrm{ir}})+\mathcal{L}_{\mathrm{rsc}}(\Delta\theta^*;D_{\mathrm{ir}}), \qquad \text{subject to }\quad \Delta\theta^*\in\arg\min_{\Delta\theta}\left[\mathcal{L}_{\mathrm{esd}}(\Delta\theta;D_{\mathrm{un}})+\mathcal{L}_{\mathrm{attn}}(\Delta\theta;D_{\mathrm{un}})\right].

    Here DunD_{\mathrm{un}} is the collection of concepts to erase, and DirD_{\mathrm{ir}} contains concepts that should remain generatable. The lower-level update targets erasure; the upper-level update acts on the same LoRA adapter to limit collateral loss of unrelated generation.

  2. Knowl 2 — Target-token attention penalties strengthen the lower-level erasure objective

    equation

    In Flux’s dual-stream transformer blocks, EraseAnything uses an ESD-style objective to update LoRA parameters for the target concept. The method adds a penalty on attention assigned to the target word’s token positions in the prompt. Let WattnW_{\mathrm{attn}} be the attention tensor, with shape 24×1280×128024\times1280\times1280 in the paper’s 512-by-512 image-generation setting; let idxstart\mathrm{idx}_{\mathrm{start}} and idxend\mathrm{idx}_{\mathrm{end}} be the inclusive token-index range localized for the target word; and let FidxunF^{\mathrm{un}}_{\mathrm{idx}} be the corresponding attention slice. The added regularizer is

    Fidxun=Wattn[:,:,idx],Lattn=∑idx=idxstartidxendFidxun.F^{\mathrm{un}}_{\mathrm{idx}}=W_{\mathrm{attn}}[:,:,\mathrm{idx}],\qquad \mathcal{L}_{\mathrm{attn}}=\sum_{\mathrm{idx}=\mathrm{idx}_{\mathrm{start}}}^{\mathrm{idx}_{\mathrm{end}}}F^{\mathrm{un}}_{\mathrm{idx}}.

    Minimizing this term reduces the attention weight allocated to the target word’s token positions. The accompanying ESD-style loss uses the difference between the frozen original Flux model’s target-conditioned and empty-prompt velocity predictions, scaled by the negative-guidance factor η\eta, to direct the tuned model away from generating the target concept.

  3. Knowl 3 — Reverse self-contrastive learning supplies dynamically chosen preservation concepts

    model/method

    To preserve concepts beyond those explicitly represented in a prompt, EraseAnything contrasts attention features rather than requiring labeled preservation images for every unrelated concept. For a target concept, the method extracts a feature vector FunF^{\mathrm{un}} from its attention map, a synonym feature FsynF^{\mathrm{syn}}, and feature vectors FkF^{k} for KK sampled irrelevant concepts. It defines the reverse self-contrastive loss as

    Lrsc=log⁡(∑k=1Kexp⁡ ⁣(Fun⋅Fk/τ)exp⁡ ⁣(Fun⋅Fsyn/τ)),\mathcal{L}_{\mathrm{rsc}}= \log\left(\frac{\displaystyle\sum_{k=1}^{K}\exp\!\left(F^{\mathrm{un}}\cdot F^{k}/\tau\right)} {\exp\!\left(F^{\mathrm{un}}\cdot F^{\mathrm{syn}}/\tau\right)}\right),

    where ⋅\cdot is the dot product and τ\tau is a temperature; the reported setting is K=3K=3 and τ=0.07\tau=0.07. The authors describe this as a reverse contrastive objective intended to distinguish the erased concept from its synonym while steering its representation toward sampled irrelevant concepts. GPT-4o proposes candidate irrelevant words, which are grouped by their judged relatedness (no relation, far, or mid); one is randomly selected from each group by default. NLTK supplies a synonym of the target concept. The attention features are obtained from runs with a shared fixed starting latent, substituting the synonym and irrelevant words into the prompt.

  4. Knowl 4 — Flux’s joint text-image attention enables token-index localization

    model/method

    Flux lacks the explicit cross-attention layers used by many earlier text-to-image erasure methods. EraseAnything instead relies on the dual-stream blocks, where text and image features participate in joint attention. The authors report that columns of the resulting attention tensor associated with text-token positions correlate with localized image activations: identifying a target word’s token index therefore provides a prompt-specific attention feature that can be penalized or directly zeroed. In their analysis, directly zeroing the target-token columns could suppress a concept for an unmodified prompt, but was vulnerable to misspellings, added prefixes or suffixes, synonyms, and repeated target words. This motivated learning an adapter rather than relying on deterministic attention deletion.

  5. Knowl 5 — Training procedure and Flux implementation settings

    algorithm

    The training procedure alternates target-concept erasure updates with preservation updates on the same LoRA adapter. It operates on Flux.1 [dev] and tunes the text-related add_q_proj and add_k_proj parameters in the dual-stream blocks. The reported implementation uses AdamW, a 28-step flow-matching Euler sampler, 1,000 training steps, lower-level learning rate 0.0010.001, upper-level learning rate 0.00050.0005, and erasure guidance factor η=1\eta=1. The default irrelevant-concept count is three and the contrastive temperature is 0.070.07.

    Input: Frozen Flux.1 [dev], target concept set D_un
    Output: Fine-tuned LoRA adapter Δθ
    Initialize Δθ
    For each training step up to 1000:
        Sample target concept c_un from D_un
        Construct a meaningful prompt c containing c_un
        Shuffle the word order of c
        Locate the token-index range for c_un in the shuffled prompt
        Update Δθ with AdamW at learning rate 0.001 using:
            the ESD-style target-concept erasure objective
            plus the penalty on attention at the target-token indices
        Obtain a synonym c_syn with NLTK
        Ask GPT-4o for irrelevant concepts and select one from each of
            the no-relation, far, and mid relatedness groups
        Fix a shared initial latent and substitute c_syn and each selected
            irrelevant concept into c in separate runs
        Extract their attention features at higher flow timesteps
        Use generated prompt-conditioned images and the contrastive
            attention-feature objective for the preservation update
        Update the same Δθ with AdamW at learning rate 0.0005
    Return Δθ
  6. Knowl 6 — Nudity erasure reduces detector counts while largely retaining COCO generation quality

    data/table

    The nudity evaluation generated images for 4,703 prompts from the I2P benchmark and counted NudeNet detections at threshold 0.6. The authors also generated images for 10,000 MS-COCO validation captions and reported FID and CLIP scores. EraseAnything had the second-lowest total detected nudity count among the listed methods, behind UCE, while its MS-COCO scores were closer to the original Flux.1 [dev] scores than UCE’s were.

    Method Common Female Male Total ↓\downarrow FID ↓\downarrow CLIP ↑\uparrow
    CA (model-based) 253 65 26 344 22.66 29.05
    CA (noise-based) 290 72 28 390 23.07 28.73
    ESD 329 145 32 506 23.08 28.44
    UCE 122 39 12 173 30.71 24.56
    MACE 173 55 28 256 24.15 29.52
    EAP 287 86 13 386 22.30 29.86
    Meta-unlearning 355 140 26 521 22.69 29.91
    EraseAnything 129 48 22 199 21.75 30.24
    Flux.1 [dev] 406 161 38 605 21.32 30.87
  7. Knowl 7 — CLIP classification tests show erasure across entities, abstractions, and relationships

    data/table

    The authors tested 10 concepts in each of three categories: entities, abstractions, and relationships. CLIP classification accuracy was measured on the erased category (Acce\mathrm{Acc}_e, lower is better), unaffected categories (Accir\mathrm{Acc}_{ir}, higher is better), and synonyms of the erased category (Accg\mathrm{Acc}_g, lower is better). Across all three categories, EraseAnything improved on concept erasure, unrelated-concept retention, and synonym robustness compared with the corresponding CA baseline. Values are percentages.

    Method Acce↓\mathrm{Acc}_e\downarrow Accir↑\mathrm{Acc}_{ir}\uparrow Accg↓\mathrm{Acc}_g\downarrow
    CA (entity) 14.8 89.2 27.3
    CA (abstraction) 25.2 88.3 29.6
    CA (relationship) 22.7 88.6 23.1
    EraseAnything (entity) 12.5 91.7 18.6
    EraseAnything (abstraction) 21.1 90.5 24.7
    EraseAnything (relationship) 18.4 90.2 19.3
  8. Knowl 8 — Nudity erasure is more resistant than direct attention deletion to tested attacks

    data/table

    The attack evaluation used 285 revised nudity prompts from Ring-A-Bell-Nudity and NudeNet with detection threshold 0.6. The table reports the percentage of generated images detected as containing nudity, both without the tested attack and under MU-Attack at the initial velocity step or the first three velocity steps. EraseAnything had the lowest detection rate in each listed condition among the compared erased models, although the attacks increased its rate from 2.46% to 8.77% and 11.93%.

    Condition Flux.1 [dev] ESD CA EraseAnything
    Original 59.65% 7.36% 3.16% 2.46%
    MU-Attack (step 0) 64.56% 11.57% 15.44% 8.77%
    MU-Attack (steps 0, 1, 2) 65.96% 14.74% 16.49% 11.93%
  9. Knowl 9 — Ablations indicate that combining erasure and preservation losses gives the strongest reported scores

    data/table

    The loss ablation evaluated celebrity erasure using a subset of 100 CelebA identities that Flux could reconstruct, split into 50 target identities and 50 identities for retention. The reported metrics are classifier accuracies: Acce\mathrm{Acc}_e on the erased identities (lower is better) and Accir\mathrm{Acc}_{ir} on retained identities (higher is better). The full configuration achieved the lowest erased-identity accuracy and highest retained-identity accuracy among the listed loss combinations.

    Loss configuration Acce↓\mathrm{Acc}_e\downarrow Accir↑\mathrm{Acc}_{ir}\uparrow
    Lesd+Lattn\mathcal{L}_{esd}+\mathcal{L}_{attn} 15.3 82.1
    Lesd+Llora\mathcal{L}_{esd}+\mathcal{L}_{lora} 20.5 77.9
    Lesd+Lrsc\mathcal{L}_{esd}+\mathcal{L}_{rsc} 16.1 85.6
    Lattn+Lrsc\mathcal{L}_{attn}+\mathcal{L}_{rsc} 18.6 81.7
    Lattn+Llora+Lrsc\mathcal{L}_{attn}+\mathcal{L}_{lora}+\mathcal{L}_{rsc} 15.8 80.2
    Full configuration 14.9 88.5
  10. Knowl 10 — Normalized LoRA blending supports qualitative multi-concept erasure

    empirical result

    The authors tested combining adapters trained to erase different concepts. For NN adapters with weights WiW_i and LoRA parameter changes Δθi\Delta\theta_i, the merged adapter is Δθmul=∑i=1NWiΔθi\Delta\theta_{\mathrm{mul}}=\sum_{i=1}^{N}W_i\Delta\theta_i. In the tested normalized blending strategy, each weight is 1/N1/N and the weights sum to one. The paper reports qualitatively that this strategy retained generation of unrelated concepts close to the original Flux.1 [dev] model and allowed multiple concepts to be erased together. In contrast, unnormalized blending, with each weight equal to one, could overemphasize some concepts and impair image synthesis.

  11. Knowl 11 — The paper identifies scaling and erasure-strength control as limitations

    limitation

    EraseAnything’s reported limitations concern adapter composition and control of erasure intensity. When many concept-specific LoRAs are combined using a normalized sum, the contribution of each adapter decreases proportionally; the authors note that erasure impact can therefore weaken when erasing 10 or more concepts simultaneously. They identify efficient combination of very large numbers of adapters, such as more than 100, as an open problem. They also report that fine-tuning does not guarantee a chosen erasure strength, leaving fine-grained control—such as an interactive intensity slider—as future work.

Coverage note — The five-axis user study is omitted because the paper reports its outcomes graphically without recoverable aggregate scores; additional qualitative example grids are omitted because they do not add independently quantified results.

References

  1. 1.Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  2. 2.Bedapudi, P. Nudenet: Neural nets for nudity classification, detection and selective censoring, 2019.
  3. 3.Bird, S., Klein, E., and Loper, E. Natural language processing with Python: analyzing text with the natural language toolkit. ” O’Reilly Media, Inc.”, 2009.
  4. 4.Bui, A., Vuong, L., Doan, K., Le, T., Montague, P., Abraham, T., and Phung, D. Erasing undesirable concepts in diffusion models with adversarial preservation. arXiv preprint arXiv:2410.15618, 2024.
  5. 5.Colson, B., Marcotte, P., and Savard, G. An overview of bilevel optimization. Annals of operations research, 153: 235–256, 2007.
  6. 6.Esser, P., Kulal, S., Blattmann, A., Entezari, R., Muller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machine Learning, 2024.
  7. 7.Franceschi, L., Frasconi, P., Salzo, S., Grazzi, R., and Pontil, M. Bilevel programming for hyperparameter optimization and meta-learning. In International conference on machine learning, pp. 1568–1577. PMLR, 2018.
  8. 8.Gandikota, R., Materzynska, J., Fiotto-Kaufman, J., and Bau, D. Erasing concepts from diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2426–2436, 2023.
  9. 9.Gandikota, R., Orgad, H., Belinkov, Y., Materzynska, J., and Bau, D. Unified concept editing in diffusion models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 5111–5120, 2024.
  10. 10.Gao, H., Pang, T., Du, C., Hu, T., Deng, Z., and Lin, M. Meta-unlearning on diffusion models: Preventing relearning unlearned concepts. arXiv preprint arXiv:2410.12777, 2024.
  11. 11.Hao, Z., Ying, C., Su, H., Zhu, J., Song, J., and Cheng, Z. Bi-level physics-informed neural networks for pde constrained optimization using broyden’s hypergradients. arXiv preprint arXiv:2209.07075, 2022.
  12. 12.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 9729–9738, 2020.
  13. 13.Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022.
  14. 14.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  15. 15.Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021.
  16. 16.Huang, C.-P., Chang, K.-P., Tsai, C.-T., Lai, Y.-H., and Wang, Y.-C. F. Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasers. arXiv preprint arXiv:2311.17717, 2023.
  17. 17.Huang, Z., Wu, T., Jiang, Y., Chan, K. C., and Liu, Z. ReVersion: Diffusion-based relation inversion from images. In SIGGRAPH Asia 2024 Conference Papers, 2024.
  18. 18.Kingma, D. P. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  19. 19.Kumari, N., Zhang, B., Wang, S.-Y., Shechtman, E., Zhang, R., and Zhu, J.-Y. Ablating concepts in text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 22691–22702, 2023.
  20. 20.Li, L., Lu, S., Ren, Y., and Kong, A. W.-K. Set you straight: Auto-steering denoising trajectories to sidestep unwanted concepts. arXiv preprint arXiv:2504.12782, 2025.
  21. 21.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pp. 740–755. Springer, 2014.
  22. 22.Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022.
  23. 23.Liu, X., Gong, C., and Liu, Q. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022.
  24. 24.Liu, Y., An, J., Zhang, W., Li, M., Wu, D., Gu, J., Lin, Z., and Wang, W. Realera: Semantic-level concept erasure via neighbor-concept mining. arXiv preprint arXiv:2410.09140, 2024.
  25. 25.Liu, Z., Luo, P., Wang, X., and Tang, X. Large-scale celebfac es attributes (celeba) dataset. Retrieved August, 15 (2018):11, 2018.
  26. 26.Lorraine, J., Vicol, P., and Duvenaud, D. Optimizing millions of hyperparameters by implicit differentiation. In International conference on artificial intelligence and statistics, pp. 1540–1552. PMLR, 2020.
  27. 27.Loshchilov, I., Hutter, F., et al. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101, 5, 2017.
  28. 28.Lu, S., Liu, Y., and Kong, A. W.-K. Tf-icon: Diffusion-based training-free cross-domain image composition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 2294–2305, 2023.
  29. 29.Lu, S., Wang, Z., Li, L., Liu, Y., and Kong, A. W.-K. Mace: Mass concept erasure in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6430–6440, 2024a.
  30. 30.Lu, S., Zhou, Z., Lu, J., Zhu, Y., and Kong, A. W.-K. Robust watermarking using generative priors against image editing: From benchmarking to advances. arXiv preprint arXiv:2410.18775, 2024b.
  31. 31.Lyu, M., Yang, Y., Hong, H., Chen, H., Jin, X., He, Y., Xue, H., Han, J., and Ding, G. One-dimensional adapter to rule them all: Concepts diffusion models and erasing applications. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7559–7568, 2024.
  32. 32.Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft, 2022.
  33. 33.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  34. 34.Oord, A. v. d., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748, 2018.
  35. 35.Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Muller, J., Penna, J., and Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
  36. 36.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pp. 8748–8763. PMLR, 2021.
  37. 37.Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21 (140):1–67, 2020.
  38. 38.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. In International conference on machine learning, pp. 8821–8831. Pmlr, 2021.
  39. 39.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022.
  40. 40.Rando, J., Paleka, D., Lindner, D., Heim, L., and Tramer, F. Red-teaming the stable diffusion safety filter. arXiv preprint arXiv:2210.04610, 2022.
  41. 41.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  42. 42.Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18, pp. 234–241. Springer, 2015.
  43. 43.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems, 35:36479–36494, 2022.
  44. 44.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512, 2022.
  45. 45.Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22522–22531, 2023.
  46. 46.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022.
  47. 47.Shen, Q., Wang, Y., Yang, Z., Li, X., Wang, H., Zhang, Y., Scarlett, J., Zhu, Z., and Kawaguchi, K. Memory-efficient gradient unrolling for large-scale bi-level optimization. arXiv preprint arXiv:2406.14095, 2024.
  48. 48.Sinha, A., Malo, P., and Deb, K. A review on bilevel optimization: From classical to evolutionary approaches and applications. IEEE transactions on evolutionary computation, 22(2):276–295, 2017.
  49. 49.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020.
  50. 50.Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. Roformer: Enhanced transformer with rotary position embedding. Neurocomputing, 568:127063, 2024.
  51. 51.Vaswani, A. Attention is all you need. Advances in Neural Information Processing Systems, 2017.
  52. 52.von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., Nair, D., Paul, S., Berman, W., Xu, Y., Liu, S., and Wolf, T. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022.
  53. 53.Xie, J., Li, Y., Huang, Y., Liu, H., Zhang, W., Zheng, Y., and Shou, M. Z. Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 7452–7461, 2023.
  54. 54.Zhang, L., Liang, Y., and Xie, P. Blo-sam: Bi-level optimization based overfitting-preventing finetuning of sam. arXiv preprint arXiv:2402.16338, 2024a.
  55. 55.Zhang, Y., Chen, X., Jia, J., Zhang, Y., Fan, C., Liu, J., Hong, M., Ding, K., and Liu, S. Defensive unlearning with adversarial training for robust concept erasure in diffusion models. arXiv preprint arXiv:2405.15234, 2024b.
  56. 56.Zhang, Y., Jia, J., Chen, X., Chen, A., Zhang, Y., Liu, J., Ding, K., and Liu, S. To generate or not? safety-driven unlearned diffusion models are still easy to generate unsafe images... for now. In European Conference on Computer Vision, pp. 385–403. Springer, 2024c.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/