Multi-Concept Customization of Text-to-Image Diffusion

Nupur KumariBingliang ZhangRichard ZhangEli ShechtmanJun-Yan Zhu

article2022CVPR1,397 citations

Introduces Custom Diffusion, an efficient fine-tuning approach that updates key cross-attention weights in minutes to learn, merge, and compose multiple user-specified concepts in generative text-to-image models.

Listen

Text-to-image artificial intelligence models can generate diverse, high-quality visuals from natural language prompts, yet they struggle to represent specific, personal subjects—such as individual pets, proprietary items, or unique artistic styles. While retraining large models from scratch is cost-prohibitive and standard fine-tuning often causes catastrophic forgetting of existing knowledge, organizations require efficient, customizable image synthesis that can rapidly adopt new concepts without losing broad generative capabilities.

The article introduces and evaluates Custom Diffusion, a computationally efficient fine-tuning method that adapts text-to-image diffusion models to generate novel concepts and combine multiple new subjects within a single synthetic scene using as few as four sample images.

The researchers analyzed parameter shifts during fine-tuning across network layers and established that updating only the cross-attention key and value projection matrices—the layers mapping text input to visual features—is sufficient to acquire new concepts. They evaluated this approach on multiple datasets spanning pets, personal objects, and rare categories, including the newly introduced 101-concept dataset, CustomConcept101. To prevent the model from distorting related existing words, the framework incorporates a regularization dataset of approximately 200 real images retrieved using text similarity. The team also formulated a closed-form, constrained optimization algorithm allowing multiple independently fine-tuned concepts to be merged into a single model in roughly two seconds, alongside a standard joint training option. Performance was measured against leading baselines via visual similarity, text alignment, image quality metrics, and paired human evaluation studies.

The analysis produced several critical findings. First, Custom Diffusion achieves equal or superior visual fidelity and prompt alignment compared to existing methods while updating only about 3% of the total network parameters (75 MB of storage compared to 3 GB for full-model tuning). Second, it accelerates fine-tuning speeds substantially, completing adaptation in approximately 6 minutes on two A100 GPUs, which is 2 to 4 times faster than concurrent approaches. Third, in multi-concept compositional tasks, the method effectively avoids omitting subjects, earning a strong human preference of 81% to 87% over alternatives like Textual Inversion and 56% to 62% over DreamBooth. Fourth, the updated parameter weights can be compressed further via low-rank matrix approximation by up to five times without significant performance degradation.

These results demonstrate that targeted parameter updates drastically reduce the computational overhead and storage costs required to operationalize customized generative AI. For enterprise applications, this efficiency enables scalable personalization pipelines where thousands of individual concept models can be trained, stored, and combined dynamically without the infrastructure expense of full model replication. Furthermore, utilizing real retrieved images for regularization mitigates language drift, ensuring high output quality and preserving general pretraining capabilities.

Organizations implementing generative media should consider adopting selective attention-layer fine-tuning to minimize infrastructure expenditure and accelerate iteration cycles. For multi-concept workflows, teams should balance the trade-offs: joint training yields the highest visual consistency across complex compositions, while the closed-form merging method offers near-instantaneous integration for existing separate models. At the same time, teams must maintain awareness of limitations. Custom Diffusion inherits underlying base model weaknesses and faces degraded fidelity during challenging compositions involving similar semantic categories (such as multiple pets in the same frame) or when attempting to compose three or more distinct concepts simultaneously. Reliable synthetic media detection protocols should also be integrated to mitigate the governance and safety risks associated with realistic image personalization.

Cover for Multi-Concept Customization of Text-to-Image Diffusion

Abstract

While generative models produce high-quality images of concepts learned from a large-scale database, a user often wishes to synthesize instantiations of their own concepts (for example, their family, pets, or items). Can we teach a model to quickly acquire a new concept, given a few examples? Furthermore, can we compose multiple new concepts together? We propose Custom Diffusion, an efficient method for augmenting existing text-to-image models. We find that only optimizing a few parameters in the text-to-image conditioning mechanism is sufficiently powerful to represent new concepts while enabling fast tuning (~6 minutes). Additionally, we can jointly train for multiple concepts or combine multiple fine-tuned models into one via closed-form constrained optimization. Our fine-tuned model generates variations of multiple new concepts and seamlessly composes them with existing concepts in novel settings. Our method outperforms or performs on par with several baselines and concurrent works in both qualitative and quantitative evaluations while being memory and computationally efficient.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Single-Concept Fine-tuning
  • 3.2 Multiple-Concept Compositional Fine-tuning
  • 4 Experiments
  • 4.1 Single-Concept Fine-tuning Results
  • 4.2 Multiple-Concept Fine-tuning Results
  • 4.3 Human Preference Study
  • 4.4 Ablation and Applications
  • 5 Discussion and Limitations
  • References
  • A CustomConcept101
  • B Multi-Concept Optimization Based Method
  • C Experiments
  • D Evaluation
  • E Implementation and Experiment Details
  • F Societal Impact
  • G Change log

Citation

MLA
Kumari, N., et al. “Multi-Concept Customization of Text-to-Image Diffusion”. arXiv, 2022, http://arxiv.org/abs/2212.04488v2.
APA
Kumari, N., Zhang, B., Zhang, R., Shechtman, E., & Zhu, J.-Y. (2022). Multi-Concept Customization of Text-to-Image Diffusion. arXiv. http://arxiv.org/abs/2212.04488v2
Chicago
Kumari, N., B. Zhang, R. Zhang, E. Shechtman, and J.-Y. Zhu. 2022. “Multi-Concept Customization of Text-to-Image Diffusion”. arXiv. http://arxiv.org/abs/2212.04488v2.
Harvard
Kumari, N. et al. (2022) “Multi-Concept Customization of Text-to-Image Diffusion”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.04488v2.
Vancouver
1. Kumari N, Zhang B, Zhang R, Shechtman E, Zhu J-Y (2022) Multi-Concept Customization of Text-to-Image Diffusion. arXiv

BibTeX

@article{kumari2022multi,
  title = {Multi-Concept Customization of Text-to-Image Diffusion},
  author = {Kumari, Nupur and Zhang, Bingliang and Zhang, Richard and Shechtman, Eli and Zhu, Jun-Yan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.04488v2},
  eprint = {2212.04488}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/