Cones: Concept Neurons in Diffusion Models for Customized Generation

Zhiheng LiuRuili FengKai ZhuYifei ZhangKecheng ZhengYu LiuDeli ZhaoJingren ZhouYang Cao

article2023ICML174 citations

Identifies subject-specific concept neurons within diffusion models to enable efficient multi-subject image generation by composing neuron clusters while cutting storage requirements by ninety percent.

Listen

Text-to-image artificial intelligence models often struggle to reliably depict specific, user-provided subjects in new scenes, particularly when combining several distinct subjects into a single coherent image. Existing customization techniques typically require fine-tuning large portions of the model or storing massive parameter files for every new subject. This approach creates severe storage bottlenecks, risks forgetting previously learned subjects, and leads to visual blending or missing details when multiple subjects are requested at once.

The article evaluates whether diffusion models localize individual subjects into specific network parameters, termed concept neurons, and demonstrates a method called Cones to identify and manipulate these neurons for customized image generation. The primary objective is to prove that pinpointing and deactivating these sparse, subject-specific parameters enables high-fidelity single- and multi-subject generation without requiring heavy retraining or massive storage overhead.

The authors analyze pre-trained diffusion models by measuring network gradient responses to target concepts, drawing inspiration from biological neuroscience techniques that map concept-specific brain activity. They isolate small clusters of concept neurons within the model's cross-attention layers using an adaptive sampling algorithm. The evaluation assesses visual similarity to target subjects and semantic alignment with text prompts across multiple image categories, comparing the approach against established customization baselines such as DreamBooth, Custom Diffusion, and Textual Inversion. A user study with 50 annotators further measures visual fidelity and prompt consistency.

The analysis reveals several critical findings. First, subjects are governed by highly sparse neuron clusters occupying only about 1.3% of the attention layer parameters for a single concept and roughly 7.0% for four concepts. Second, simply shutting off these identified neurons—operating at binary precision without further numerical optimization—reliably implants the target subject into newly generated scenes. Third, concept neurons exhibit strong disentanglement, sharing less than 2.5% of active neurons between distinct subjects; directly combining their neuron masks allows seamless multi-subject generation and is the first method demonstrated to composite up to four distinct subjects cleanly in a single image. Fourth, because the method only needs to store integer indices rather than dense floating-point weights, it reduces memory storage consumption by more than 90% compared to Custom Diffusion (requiring about 1.43 MB versus 72 MB) and over 99.9% compared to DreamBooth (3.3 GB).

These findings indicate that deep generative networks possess an interpretable, affine semantic structure within their parameter space. In practical applications, this translates to massive cost reductions for hosting personalized generative models and makes client-side customization viable on mobile and edge devices. Furthermore, the modularity of concept neurons avoids the catastrophic forgetting commonly encountered during sequential learning in baseline approaches.

Organizations deploying personalized image generation should consider transitioning from full-model fine-tuning to sparse neuron-masking architectures like Cones, particularly for mobile deployment and multi-subject generation pipelines. When deploying such systems, practitioners can leverage direct concatenation for fast, tuning-free composition, or apply targeted collaborative fine-tuning when higher visual precision is required.

The article notes several operational boundaries and limitations. While the technique scales reliably up to four simultaneous concepts, generating five or more subjects significantly increases failure rates. Additionally, pre-existing layout biases in underlying base models can occasionally cause positioning errors between juxtaposed subjects. Within these operating boundaries, however, the quantitative metrics and human evaluations provide high confidence in the method's ability to deliver efficient, lightweight, and high-fidelity customized generation.

No sufficiently relevant recommendations were found.

Cover for Cones: Concept Neurons in Diffusion Models for Customized Generation

Abstract

Human brains respond to semantic features of presented stimuli with different neurons. This raises the question of whether deep neural networks admit a similar behavior pattern. To investigate this phenomenon, this paper identifies a small cluster of neurons associated with a specific subject in a diffusion model. We call those neurons the concept neurons. They can be identified by statistics of network gradients to a stimulation connected with the given subject. The concept neurons demonstrate magnetic properties in interpreting and manipulating generation results. Shutting them can directly yield the related subject contextualized in different scenes. Concatenating multiple clusters of concept neurons can vividly generate all related concepts in a single image. Our method attains impressive performance for multi-subject customization, even four or more subjects. For large-scale applications, the concept neurons are environmentally friendly as we only need to store a sparse cluster of int index instead of dense float32 parameter values, reducing storage consumption by 90% compared with previous customized generation methods. Extensive qualitative and quantitative studies on diverse scenarios show the superiority of our

Table of Contents

  • 1. Introduction
  • 2. Preliminaries and Background
  • Diffusion Models.
  • Text-to-Image Diffusion Model.
  • Customized Generation.
  • 3. Method
  • 3.1. Concept Neurons for a Given Subject
  • 3.2. Interpretability of Concept Neurons
  • 3.3. Collaboratively Capturing Multiple Concepts
  • 3.4. Efficient Storage
  • 4. Experiments
  • 4.1. Implementation and Experiment Details
  • 4.2. Qualitative Evaluation
  • 4.3. Quantitative Evaluation and User Study
  • 5. Conclusion
  • Acknowledgments
  • References
  • Appendix
  • A. Proof
  • A.1. Proof to Theorem 3.1
  • A.2. Further Acceleration of Eq. (11)
  • B. Experiment Setups
  • B.1. Textual Inversion
  • B.2. Dreambooth
  • B.3. Custom Diffusion
  • B.4. Cones (Ours)
  • B.5. User Study
  • C. More Results
  • C.1. Sequential Training Comparison
  • C.2. Style Conversion
  • C.3. Editing Performance
  • C.4. Overfitting on the training prompt template
  • C.5. More results on multi subjects
  • C.6. Interpolation results between various subjects
  • C.7. Faliure modes

Knowls

  1. Knowl 1 — Gradient-sign criterion identifies subject concept neurons

    theoretical result

    For a scalar parameter θh\theta_h in a text-to-image diffusion model, let LconL_{\mathrm{con}} be the concept-implanting loss: a subject-reconstruction loss for reference images paired with a unique identifier, combined with a prior-preservation loss for other images from the same class. Consider scaling only θh\theta_h down slightly while leaving the other parameters fixed. If ρ>0\rho>0 is sufficiently small, scaling down decreases this loss exactly when

    θh∂Lcon∂θh>0.\theta_h\frac{\partial L_{\mathrm{con}}}{\partial \theta_h}>0.

    The local loss decrease is proportional to (θh∂Lcon∂θh)2\left(\theta_h\frac{\partial L_{\mathrm{con}}}{\partial \theta_h}\right)^2. The paper therefore defines a parameter satisfying this positive-sign condition as a concept neuron for the target subject. The criterion is used for parameters in the key- and value-projection attention layers.

  2. Knowl 2 — Gradient-based concept-neuron mask computation

    algorithm

    The mask-finding procedure evaluates the concept-implanting loss while progressively scaling model parameters down. Let θ∈Rn\boldsymbol{\theta}\in\mathbb{R}^n denote the vector of key- and value-attention parameters being searched, KK the maximum number of samples, ρ>0\rho>0 a small step size, and τ>0\tau>0 a threshold. The loss LconL_{\mathrm{con}} uses the target subject's reference images and identifier together with a prior-preservation term. Updates are elementwise and can be computed in parallel across parameters.

    Input: L_con, parameter vector θ, maximum sample count K, step size ρ, threshold τ
    Set θ^1 = θ
    For k = 1 to K - 1:
        Compute g^k = ∇_θ L_con(θ^k)
        Set θ^{k+1} = θ^k ⊙ (1 - ρ θ^k ⊙ g^k)
    Set Mp = sum over k = 1 to K of θ^k ⊙ ∇_θ L_con(θ^k)
    Set M = 1 - (Mp > τ)
    Output binary mask M, where 0 marks a concept neuron and 1 marks any other parameter

    This adaptive sequence takes progressively smaller scaling steps in regions where the gradient-based evidence is ambiguous. The paper also gives an accelerated implementation that optimizes a scaling vector ξ\boldsymbol{\xi} in Lcon(ξ⊙θ)L_{\mathrm{con}}(\boldsymbol{\xi}\odot\boldsymbol{\theta}), initialized at all ones, and estimates the accumulated statistic from the final scaling vector. The reported implementation uses an A100 GPU and batch size 2; it gives a base learning rate of 3×10−53\times10^{-5} and describes scaling the rate to 6×10−56\times10^{-5} with GPU count and batch size. For single-subject generation it reports a base learning rate of 2×10−52\times10^{-5} and 1,000 training steps. Numerical values for the mask threshold τ\tau and maximum sample count KK are not specified.

  3. Knowl 3 — Cones customizes generation by zeroing the identified parameters

    model/method

    Cones represents a target subject by the binary mask of its concept neurons in the pretrained model's key- and value-attention parameters. For a parameter vector θ\boldsymbol{\theta} and mask M\mathbf{M}, it uses M⊙θ\mathbf{M}\odot\boldsymbol{\theta}, where mask entries equal to 0 identify concept neurons and entries equal to 1 leave other parameters unchanged. This shuts off the identified parameters without further optimization; prompted with the subject's unique text identifier, the resulting model generates that subject in contexts specified by the prompt. Stable Diffusion V1.4 is the paper's default model.

  4. Knowl 4 — Unions of subject masks enable direct multi-subject generation

    empirical result

    Concept-neuron clusters for different subjects can be combined by taking the union of their identified parameter indices and shutting off that combined set. The modified diffusion model can then generate the corresponding subjects together when prompted with their identifiers, without fine-tuning the combined mask. The visual examples presented on page 5 include a cat and a wooden pot generated together; examples elsewhere in the paper extend this direct composition to additional subjects and backgrounds. This result is the empirical basis for treating concept-neuron clusters as composable subject representations.

  5. Knowl 5 — Joint-loss refinement improves concatenated multi-subject masks

    model/method

    To improve the quality of a directly combined set of subject concept neurons, Cones can refine the combination using a multi-concept-implanting loss formed by summing the concept-implanting losses for all included subjects. It reruns mask identification using this joint loss but restricts the search to the concatenated subject-specific concept-neuron positions, rather than searching all key- and value-attention parameters. The paper reports that this procedure slightly outperformed learning a multi-subject mask from scratch in its experiments, and attributes the benefit to reducing conflicts caused by inaccuracies in the separately identified masks. Using this refinement, the authors demonstrate generation of four distinct subjects in one image.

  6. Knowl 6 — Subject control is robust across parameter precision

    empirical result

    The paper tests changes to identified concept-neuron parameters at float32, float16, quaternary, and binary digital accuracy while freezing the other parameters. It reports close generation performance across these settings, indicating that subject control does not require high-precision changes to the identified parameters. In the binary case, the concept neurons are simply shut off, with no further tuning; this setting is the default used by Cones. The precision comparison is shown visually on page 5, and the paper does not report numerical scores for the four precision settings.

  7. Knowl 7 — Different subjects have sparse overlap in their concept neurons

    empirical result

    For the cat and wooden-pot subjects, the paper measures concept-neuron overlap in the Stable Diffusion V1.4 layer upblocks.2.attentions.1.transformerblocks.0.attn1.tov. The intersection of the two independently identified clusters is reported as only 2.42% of the total concept neurons, indicating that the clusters are largely distinct. The mask learned from a joint loss for both subjects shares 53.27% of its neurons with the direct concatenation of their separate clusters. These measurements support the observation that subject-specific masks can be combined while retaining largely separate parameter subsets.

  8. Knowl 8 — CLIP evaluation favors Cones as the number of subjects increases

    data/table

    The quantitative comparison on page 8 evaluates Textual Inversion, DreamBooth, Custom Diffusion, and Cones using CLIP cosine similarity. Image alignment compares generated images with the target subject; for multi-subject prompts it averages similarity over the included subjects. Text alignment compares generated-image embeddings with prompt embeddings after omitting the subject identifier. The evaluation uses 20 prompts per concept group and generates 50 images per prompt. Cones has the highest text-alignment score in each setting and the highest image-alignment score for two, three, and four subjects. For one subject, Textual Inversion has the highest image-alignment score (0.744), versus 0.725 for Cones.

    Subjects Method Text alignment Image alignment
    Single Textual Inversion 0.312 0.744
    Single DreamBooth 0.344 0.731
    Single Custom Diffusion 0.352 0.722
    Single Cones 0.361 0.725
    Two Textual Inversion 0.264 0.630
    Two DreamBooth 0.283 0.673
    Two Custom Diffusion 0.314 0.685
    Two Cones 0.337 0.698
    Three Textual Inversion 0.223 0.584
    Three DreamBooth 0.263 0.631
    Three Custom Diffusion 0.289 0.669
    Three Cones 0.301 0.685
    Four Textual Inversion 0.219 0.553
    Four DreamBooth 0.238 0.597
    Four Custom Diffusion 0.269 0.632
    Four Cones 0.285 0.653

    The scores show that Cones' advantage over the competing methods grows in the reported multi-subject settings, on both visual similarity and prompt alignment.

  9. Knowl 9 — Sparse masks reduce per-subject customization storage

    data/table

    Cones stores the indices of identified concept neurons rather than a dense set of fine-tuned floating-point model parameters. The storage and sparsity results reported on page 9 show increasing cost as more subjects are represented, while the selected parameters remain a minority of the attention-layer parameters. DreamBooth and Custom Diffusion are shown as comparison methods; sparsity is not reported for them. Under these reported settings, Cones requires less storage than either comparison method, including at four subjects.

    Method Storage Sparsity
    DreamBooth 3.3 GB –
    Custom Diffusion 72 MB –
    Cones (single subject) 1.43 MB ±\pm 0.34 MB 1.32% ±\pm 0.29%
    Cones (two subjects) 3.41 MB ±\pm 0.56 MB 2.43% ±\pm 0.44%
    Cones (three subjects) 4.96 MB ±\pm 0.70 MB 4.54% ±\pm 0.59%
    Cones (four subjects) 7.75 MB ±\pm 0.56 MB 7.01% ±\pm 0.26%

    Here sparsity is the percentage of concept neurons among the parameters in the attention layers. The paper reports storage savings of more than 90% compared with Custom Diffusion.

  10. Knowl 10 — Multi-subject generation becomes less reliable beyond four subjects

    limitation

    The paper reports failure cases when two subjects must be displayed juxtaposed and notes that the success rate generally falls as the number of included subjects increases. It observes substantially more failures when prompts involve five or more subjects. The demonstrated four-subject results therefore do not establish reliable generation for larger compositions; the juxtaposition issue is also described by the authors as a broader difficulty for Stable Diffusion.

Coverage note — The sequential-training comparison, in which Cones retained two learned subjects better than DreamBooth and Custom Diffusion, and the supplementary style-editing and interpolation examples are omitted because they are secondary analyses rather than core method or benchmark results.

References

  1. 1.Bausch, M., Niediek, J., Reber, T. P., Mackay, S., Bostrom, J., Elger, C. E., and Mormann, F. Concept neurons in the human medial temporal lobe flexibly represent abstract relations between concepts. Nature communications, 12(1):6164, 2021.
  2. 2.Chefer, H., Alaluf, Y., Vinker, Y., Wolf, L., and Cohen-Or, D. Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. arXiv preprint arXiv:2301.13826, 2023.
  3. 3.Dhariwal, P. and Nichol, A. Diffusion models beat GANs on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  4. 4.Feng, W., He, X., Fu, T.-J., Jampani, V., Akula, A., Narayana, P., Basu, S., Wang, X. E., and Wang, W. Y. Training-free structured diffusion guidance for compositional text-to-image synthesis. arXiv preprint arXiv:2212.05032, 2022.
  5. 5.Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-Or, D. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022.
  6. 6.Ho, J. and Salimans, T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  7. 7.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  8. 8.Huettel, S. A., Song, A. W., McCarthy, G., et al. Functional magnetic resonance imaging, volume 1. Sinauer Associates Sunderland, 2004.
  9. 9.Kumari, N., Zhang, B., Zhang, R., Shechtman, E., and Zhu, J.-Y. Multi-concept customization of text-to-image diffusion. arXiv preprint arXiv:2212.04488, 2022.
  10. 10.Kwong, K. K., Belliveau, J. W., Chesler, D. A., Goldberg, I. E., Weisskoff, R. M., Poncelet, B. P., Kennedy, D. N., Hoppel, B. E., Cohen, M. S., and Turner, R. Dynamic magnetic resonance imaging of human brain activity during primary sensory stimulation. Proceedings of the National Academy of Sciences, 89(12):5675–5679, 1992.
  11. 11.LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. nature, 521(7553):436–444, 2015.
  12. 12.Lee, J., Cho, K., and Kiela, D. Countering language drift via visual grounding. arXiv preprint arXiv:1909.04499, 2019.
  13. 13.Li, Y., Liu, H., Wu, Q., Mu, F., Yang, J., Gao, J., Li, C., and Lee, Y. J. Gligen: Open-set grounded text-to-image generation. arXiv preprint arXiv:2301.07093, 2023.
  14. 14.Liu, X., Park, D. H., Azadi, S., Zhang, G., Chopikyan, A., Hu, Y., Shi, H., Rohrbach, A., and Darrell, T. More control for free! image synthesis with semantic diffusion guidance. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 289–299, 2023.
  15. 15.Logothetis, N. K., Pauls, J., Augath, M., Trinath, T., and Oeltermann, A. Neurophysiological investigation of the basis of the fMRI signal. nature, 412(6843):150–157, 2001.
  16. 16.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. DPM-Solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022.
  17. 17.Lu, Y., Singhal, S., Strub, F., Courville, A., and Pietquin, O. Countering language drift with seeded iterated learning. In International Conference on Machine Learning, pp. 6437–6447. PMLR, 2020.
  18. 18.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021.
  19. 19.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pp. 8748–8763. PMLR, 2021.
  20. 20.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125, 2022.
  21. 21.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  22. 22.Rudin, W. et al. Principles of mathematical analysis, volume 3. McGraw-hill New York, 1976.
  23. 23.Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv preprint arXiv:2208.12242, 2022.
  24. 24.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  25. 25.Sharoh, D., Van Mourik, T., Bains, L. J., Segaert, K., Weber, K., Hagoort, P., and Norris, D. G. Laminar specific fMRI reveals directed interactions in distributed networks during language processing. Proceedings of the National Academy of Sciences, 116(42):21185–21190, 2019.
  26. 26.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020.
  27. 27.Thiebaut de Schotten, M. and Forkel, S. J. The emergent properties of the connected brain. Science, 378(6619):505–510, 2022.
  28. 28.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  29. 29.von Platen, P., Patil, S., Lozhkov, A., Cuenca, P., Lambert, N., Rasul, K., Davaadorj, M., and Wolf, T. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022.
  30. 30.Voynov, A., Aberman, K., and Cohen-Or, D. Sketch-guided text-to-image diffusion models. arXiv preprint arXiv:2211.13752, 2022.
  31. 31.Yu, J., Xu, Y., Koh, J. Y., Luong, T., Baid, G., Wang, Z., Vasudevan, V., Ku, A., Yang, Y., Ayan, B. K., et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2022.

Citation

MLA
Liu, Z., et al. “Cones: Concept Neurons in Diffusion Models for Customized Generation”. International Conference on Machine Learning, vol. 202, 2023, pp. 21548–66, https://proceedings.mlr.press/v202/liu23j.html.
APA
Liu, Z., Feng, R., Zhu, K., Zhang, Y., Zheng, K., Liu, Y., Zhao, D., Zhou, J., & Cao, Y. (2023). Cones: Concept Neurons in Diffusion Models for Customized Generation. International Conference on Machine Learning, 202, 21548–21566. https://proceedings.mlr.press/v202/liu23j.html
Chicago
Liu, Z., R. Feng, K. Zhu, et al. 2023. “Cones: Concept Neurons in Diffusion Models for Customized Generation”. International Conference on Machine Learning 202: 21548–66. https://proceedings.mlr.press/v202/liu23j.html.
Harvard
Liu, Z. et al. (2023) “Cones: Concept Neurons in Diffusion Models for Customized Generation”, International Conference on Machine Learning. PMLR, pp. 21548–21566. Available at: https://proceedings.mlr.press/v202/liu23j.html.
Vancouver
1. Liu Z, Feng R, Zhu K, Zhang Y, Zheng K, Liu Y, Zhao D, Zhou J, Cao Y (2023) Cones: Concept Neurons in Diffusion Models for Customized Generation. In: International Conference on Machine Learning. PMLR, pp 21548–21566

BibTeX

@InProceedings{pmlr-v202-liu23j,
  title = 	 {Cones: Concept Neurons in Diffusion Models for Customized Generation},
  author =       {Liu, Zhiheng and Feng, Ruili and Zhu, Kai and Zhang, Yifei and Zheng, Kecheng and Liu, Yu and Zhao, Deli and Zhou, Jingren and Cao, Yang},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {21548--21566},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/liu23j/liu23j.pdf},
  url = 	 {https://proceedings.mlr.press/v202/liu23j.html},
  abstract = 	 {Human brains respond to semantic features of presented stimuli with different neurons. This raises the question of whether deep neural networks admit a similar behavior pattern. To investigate this phenomenon, this paper identifies a small cluster of neurons associated with a specific subject in a diffusion model. We call those neurons the concept neurons. They can be identified by statistics of network gradients to a stimulation connected with the given subject. The concept neurons demonstrate magnetic properties in interpreting and manipulating generation results. Shutting them can directly yield the related subject contextualized in different scenes. Concatenating multiple clusters of concept neurons can vividly generate all related concepts in a single image. Our method attains impressive performance for multi-subject customization, even four or more subjects. For large-scale applications, the concept neurons are environmentally friendly as we only need to store a sparse cluster of int index instead of dense float32 parameter values, reducing storage consumption by 90% compared with previous customized generation methods. Extensive qualitative and quantitative studies on diverse scenarios show the superiority of our method in interpreting and manipulating diffusion models.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/