FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition

Ganggui DingCanyu ZhaoWen WangZhen YangZide LiuHao ChenChunhua Shen

article2024CVPR69 citations

Proposes a tuning-free framework that composes multiple user-specified concepts into customized images using only a single reference image per subject, eliminating the need for test-time fine-tuning through multi-reference self-attention and weighted masking.

Listen

Generating customized images that seamlessly combine specific subjects or objects has become a key requirement across industries such as digital advertising, virtual try-on, and creative media. Existing artificial intelligence methods typically focus on single-subject customization and require time-consuming fine-tuning or retraining on extensive datasets. When applied to complex tasks involving multiple subjects, these models often suffer from identity distortion, conceptual confusion, and severe computational delays, limiting their practical deployment.

The article demonstrates FreeCustom, a tuning-free method for customized image generation that composes multiple distinct concepts using only a single reference image per concept. The primary objective is to evaluate how effectively this approach can preserve subject identities and align with text prompts across multiple base models without requiring any model retraining or parameter fine-tuning.

To achieve this, the authors designed a dual-path pipeline that extracts reference image features and integrates them during the standard image generation process. The core mechanism, termed multi-reference self-attention, injects features from reference images into deep network layers. A weighted masking strategy isolates the target subjects from their backgrounds to prevent unwanted visual clutter from bleeding into the output. The evaluation benchmarked the approach on a diverse dataset spanning animals, clothing, accessories, and human faces, comparing it against established tuning-based and tailored customization baselines across automated metrics and human user studies.

The findings show that FreeCustom matches or exceeds existing methods in image-text alignment and visual quality while eliminating preprocessing overhead. In multi-concept tasks, it achieved significantly higher user study ratings—scoring 4.40 out of 5 for text alignment, 4.65 for concept consistency, and 4.17 for image quality, compared to top baseline scores of 1.91, 2.53, and 2.48 respectively. Operationally, the method generated multi-concept images in 36 to 58 seconds with zero preprocessing time, whereas competing methods required up to several minutes of fine-tuning or days of prior model retraining. Additionally, providing input images that feature contextual interactions—such as a hat being worn rather than isolated on a blank background—substantially improved identity preservation.

These results demonstrate that complex, personalized visual content can be synthesized on demand without expensive retraining infrastructure or long turnaround times. The approach integrates directly into existing diffusion models in a plug-and-play manner, dramatically reducing operational compute costs and shortening product development cycles. Organizations looking to deploy scalable personalized media generation should adopt tuning-free attention injection architectures. However, decision-makers should note that the system currently lacks an explicit structural perception module to handle complex geometric relationships, and incorporating future enhancements like specialized image adapters is recommended for further accuracy.

Cover for FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition

Abstract

Benefiting from large-scale pre-trained text-to-image (T2I) generative models, impressive progress has been achieved in customized image generation, which aims to generate user-specified concepts. Existing approaches have extensively focused on single-concept customization and still encounter challenges when it comes to complex scenarios that involve combining multiple concepts. These approaches often require retraining/fine-tuning using a few images, leading to time-consuming training processes and impeding their swift implementation. Furthermore, the reliance on multiple images to represent a singular concept increases the difficulty of customization.

To this end, we propose FreeCustom, a novel tuning-free method to generate customized images of multi-concept composition based on reference concepts, using only one image per concept as input. Specifically, we introduce a new multi-reference self-attention (MRSA) mechanism and a weighted mask strategy that enables the generated image to access and focus more on the reference concepts. In addition, MRSA leverages our key finding that input concepts are better preserved when providing images with context interactions. Experiments show that our method's produced images are consistent with the given concepts and better aligned with the input text. Our method outperforms or performs on par with other training-based methods in terms of multi-concept composition and single-concept customization, but is simpler. Codes can be found here.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. Method
  • 4.1. Multi-Reference Self-Attention
  • 4.2. Weighted Mask
  • 4.3. Selective MRSA Replacement
  • 4.4. Preparing Images with Context Interaction
  • 5. Experiments
  • 5.1. Comparison with Existing Methods
  • 5.2. Ablation Studies
  • 5.3. More Applications
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Dual-Path FreeCustom Architecture for Multi-Concept Composition

    model/method

    FreeCustom is a tuning-free text-to-image customization framework that composes multiple user-specified concepts into a single generated image using only a single reference image per concept. The framework operates via a dual-path diffusion denoising pipeline comprising a concept reference path and a concept composition path.

    Let I={I1,I2,…,IN}\mathcal{I} = \{I_1, I_2, \dots, I_N\} be a set of NN reference images containing distinct concepts, with corresponding descriptive text prompts P={P1,P2,…,PN}\mathcal{P} = \{P_1, P_2, \dots, P_N\}, and let PP denote the target prompt for the composite image. Concept segmentation masks M={M1,M2,…,MN}\mathcal{M} = \{M_1, M_2, \dots, M_N\} are extracted from the reference images using a pre-trained open-set segmentation model (such as Grounded-Segment-Anything).

    1. Concept Reference Path: The reference images are converted to latent space representations z0′=Enc(I)z'_0 = \text{Enc}(\mathcal{I}) using a pre-trained VAE encoder Enc\text{Enc}. At each reverse denoising timestep tt, forward diffusion adds noise to obtain zt′z'_t. The noisy latents zt′z'_t and prompts P\mathcal{P} are fed into a standard frozen Stable Diffusion U-Net ϵθ\epsilon_\theta to extract intermediate query, key, and value features (Qn,Kn,Vn)(Q_n, K_n, V_n) across attention layers for each reference image n∈{1,…,N}n \in \{1, \dots, N\}. The final output noise prediction of ϵθ\epsilon_\theta is discarded.

    2. Concept Composition Path: A latent vector zT∼N(0,I)z_T \sim \mathcal{N}(0, \mathbf{I}) is initialized from Gaussian noise and iteratively denoised to z0z_0 conditioned on the target prompt PP using a modified U-Net ϵθ∗\epsilon^*_\theta. Within ϵθ∗\epsilon^*_\theta, standard self-attention modules are replaced by Multi-Reference Self-Attention (MRSA) mechanisms that inject the reference key-value features (Kn,Vn)(K_n, V_n) into the generation process. The final denoised latent z0z_0 is mapped back to pixel space via the VAE decoder Dec(z0)\text{Dec}(z_0) to produce the customized composite image.

  2. Knowl 2 — Multi-Reference Self-Attention with Weighted Concept Masking

    equation

    In FreeCustom, the standard self-attention mechanism in the diffusion U-Net is replaced with Multi-Reference Self-Attention (MRSA) combined with a weighted spatial mask strategy to inject reference concept features while suppressing irrelevant background information.

    MRSA⁡(Q,K′,V′,Mw)=Softmax⁡(Mw⊙(QK′T)d)V′\operatorname{MRSA}(\mathbf{Q}, \mathbf{K}', \mathbf{V}', \mathbf{M}_w) = \operatorname{Softmax}\left(\frac{\mathbf{M}_w \odot (\mathbf{Q} {\mathbf{K}'}^T)}{\sqrt{d}}\right) \mathbf{V}'

    where:

    • Q∈RHW×d\mathbf{Q} \in \mathbb{R}^{HW \times d} is the query feature matrix projected from the target latent ztz_t at a given attention layer with spatial resolution H×WH \times W and feature dimension dd.
    • K′=[K,K1,K2,…,KN]∈R(N+1)HW×d\mathbf{K}' = [\mathbf{K}, \mathbf{K}_1, \mathbf{K}_2, \dots, \mathbf{K}_N] \in \mathbb{R}^{(N+1)HW \times d} and V′=[V,V1,V2,…,VN]∈R(N+1)HW×d\mathbf{V}' = [\mathbf{V}, \mathbf{V}_1, \mathbf{V}_2, \dots, \mathbf{V}_N] \in \mathbb{R}^{(N+1)HW \times d} represent the concatenated key and value matrices formed by joining the composition path self-attention features (K,V)(\mathbf{K}, \mathbf{V}) with the NN reference features (Kn,Vn)n=1N(\mathbf{K}_n, \mathbf{V}_n)_{n=1}^N extracted from the reference path latents zt′z'_t.
    • Mw=[1,ω1M1,ω2M2,…,ωNMN]∈RHW×(N+1)HW\mathbf{M}_w = [\mathbf{1}, \omega_1 \mathbf{M}_1, \omega_2 \mathbf{M}_2, \dots, \omega_N \mathbf{M}_N] \in \mathbb{R}^{HW \times (N+1)HW} is the weighted spatial mask matrix. Here, 1\mathbf{1} is an all-ones matrix corresponding to the self-attention of the generated latent (weight 1), Mn\mathbf{M}_n is the binary segmentation mask of the nn-th reference concept downsampled to resolution H×WH \times W, and ωn∈R\omega_n \in \mathbb{R} is the concept scaling factor (typically chosen such that ωn∈[2,3]\omega_n \in [2, 3]).
    • ⊙\odot represents the element-wise Hadamard product.

    The mask Mw\mathbf{M}_w zeroes out attention to non-concept background regions in the reference images, and scaling factors ωn>1\omega_n > 1 boost attention towards specific concept features to preserve fine-grained concept details.

  3. Knowl 3 — Selective MRSA Replacement in Deep U-Net Blocks

    model/method

    Replacing the vanilla self-attention modules across all 7 basic blocks (16 layers) of the Stable Diffusion U-Net with Multi-Reference Self-Attention (MRSA) causes unnatural image synthesis, degraded conceptual coherence, and text-image misalignment. Because query features in the deeper layers of the U-Net govern spatial layout control and high-level semantic acquisition, FreeCustom restricts the replacement of self-attention with MRSA strictly to a selected subset of deeper U-Net blocks Ψ\Psi.

    Empirical evaluation demonstrates that setting Ψ=[5,6]\Psi = [5, 6] (the 5th and 6th basic blocks) produces optimal results. This selective replacement allows earlier layers to maintain standard structural and photorealistic generative priors while enabling deeper layers to query and bind concept-specific features from the reference images.

  4. Knowl 4 — Context Interaction Requirement for Reference Concept Fidelity

    empirical result

    The preservation of reference concept identity during multi-concept composition relies heavily on whether the concept in the reference image is depicted with contextual interaction (e.g., an item being worn or held, such as 'a cat wearing a hat') rather than presented as an isolated object against a plain background (e.g., an isolated hat).

    When input images contain only isolated concepts without context, the attention mechanism fails to inject distinct concept characteristics into the composite image, yielding unnatural and poorly aligned results. Providing reference images with context interactions—or manually creating reference contexts by pasting concept cutouts onto a proxy background or subject via a copy-paste strategy—enables the MRSA mechanism to capture global contextual interactions and generate high-fidelity, text-aligned customized compositions.

  5. Knowl 5 — Quantitative Evaluation of Multi-Concept and Single-Concept Image Customization

    data/table

    FreeCustom was evaluated against state-of-the-art single-concept customization methods (DreamBooth, NeTI, BLIP Diffusion) and multi-concept composition methods (Custom Diffusion, Perfusion) using Stable Diffusion V1.5 as the common base model. Image similarity was assessed via DINOv2 and CLIP-I; image-text alignment via CLIP-T and CLIP-T-L; and image quality via CLIP-IQA.

    Methods DINOv2 CLIP-I CLIP-T CLIP-T-L CLIP-IQA
    single-concept
    DreamBooth 0.8948 0.8906 27.3825 21.9413 0.7194
    NeTI 0.7839 0.8677 29.9023 25.1220 0.7239
    BLIP Diffusion 0.8734 0.8975 29.1278 24.0543 0.7655
    FreeCustom 0.8376 0.8755 32.0206 27.4440 0.7292
    multi-concept
    Custom Diffusion 0.6545 0.2393 29.0702 23.6657 0.8921
    Perfusion 0.6399 0.2277 22.1371 16.1719 0.8624
    FreeCustom 0.7625 0.2871 33.7826 27.8758 0.9002

    In multi-concept composition, FreeCustom achieves the highest performance across all 5 evaluation metrics, demonstrating substantial improvements in identity fidelity (DINOv2: 0.7625 vs 0.6545; CLIP-I: 0.2871 vs 0.2393), text alignment (CLIP-T: 33.7826 vs 29.0702), and overall visual quality (CLIP-IQA: 0.9002 vs 0.8921). In single-concept generation, FreeCustom delivers the highest text alignment scores (CLIP-T: 32.0206; CLIP-T-L: 27.4440) while maintaining competitive image similarity.

  6. Knowl 6 — Preprocessing and Inference Time Efficiency of Customization Methods

    data/table

    Computational efficiency was measured on an NVIDIA RTX 3090 GPU comparing FreeCustom against training-based and large-scale pre-trained customization baselines at 512×512512 \times 512 image resolution. Preprocessing time accounts for per-concept fine-tuning or dataset-scale model re-training.

    Methods Venue Preprocessing Inference Total
    single-concept
    DreamBooth CVPR'23 500s 3s 503s
    NeTI SIGGRAPH'23 420s 30s 450s
    BLIP Diffusion NeurIPS'23 6 days 3s 6 days
    FreeCustom this work 0 20s 20s
    multi-concept
    Custom Diffusion CVPR'23 287s 13s 300s
    Perfusion SIGGRAPH'23 821s 14s 835s
    FreeCustom (2 concepts) this work 0 36s 36s
    FreeCustom (3 concepts) this work 0 58s 58s

    Because FreeCustom is entirely tuning-free, its preprocessing time is 0 seconds. In single-concept customization, FreeCustom completes total generation in 20s, compared to 503s for DreamBooth and 6 days for BLIP Diffusion. In multi-concept composition, FreeCustom achieves total execution times of 36s (2 concepts) and 58s (3 concepts), eliminating the extensive 287s–821s fine-tuning steps required by Custom Diffusion and Perfusion.

  7. Knowl 7 — User Study on Text Alignment, Identity Consistency, and Image Quality

    data/table

    A user study evaluated generated customized images on a 5-point Likert scale (1 to 5, where 5 is best) across three criteria: text-to-image correspondence (Alignment), reference identity preservation (Consistency), and perceptual visual realism (Quality). The study comprised 23 questionnaires (20 questions evaluating 5 concepts) for single-concept customization and 42 questionnaires (30 questions evaluating 10 concept combinations of 2 to 4 concepts) for multi-concept composition.

    Methods Venue Alignment Consistency Quality
    single-concept
    DreamBooth CVPR'23 1.77 1.48 1.53
    NeTI SIGGRAPH'23 2.27 2.49 2.31
    BLIP Diffusion NeurIPS'23 3.50 3.04 2.85
    FreeCustom this work 3.62 3.11 2.89
    multi-concept
    Custom Diffusion CVPR'23 1.91 2.53 2.48
    Perfusion SIGGRAPH'23 1.50 1.70 1.73
    FreeCustom this work 4.40 4.65 4.17

    FreeCustom outperformed all competing methods in all three dimensions. In multi-concept composition, FreeCustom achieved 4.40 for text alignment, 4.65 for identity consistency, and 4.17 for image quality, markedly exceeding Custom Diffusion (1.91 / 2.53 / 2.48) and Perfusion (1.50 / 1.70 / 1.73).

  8. Knowl 8 — Model Portability, Appearance Transfer, and Extension to Conditional Frameworks

    empirical result

    FreeCustom operates without model fine-tuning and generalizes directly across multiple text-to-image generation tasks and architectures:

    1. Cross-Checkpoint Compatibility: FreeCustom applies directly without retraining to diverse diffusion base models and fine-tuned community checkpoints, including Stable Diffusion 1.4, Stable Diffusion 2.1, RealisticVision, CoffeeBreak, Anylora, and ReVAnimated.
    2. Plug-and-Play Integration with ControlNet and BLIP-Diffusion: Combining FreeCustom with ControlNet allows conditional layout controls (e.g., Canny edge or depth conditions) to be fulfilled while preserving the appearance and identity of reference concepts. Integrating FreeCustom into BLIP-Diffusion enhances its text prompt alignment and improves concept fidelity.
    3. Appearance Transfer: By querying features from an input concept and applying them to a distinct target category described in the prompt, FreeCustom transfers textures, patterns, and materials (e.g., goat fur, carpet patterns, or flower textures) onto newly generated objects (e.g., cats, tables, or sculptures).
  9. Knowl 9 — Lack of Explicit Geometric Structure Perception in FreeCustom

    limitation

    FreeCustom currently lacks an explicit module or structural mechanism to perceive and model the 3D or geometric structure of the input reference concepts. Consequently, when target compositions demand complex spatial reorientations or precise geometric transformations of the reference subjects, the framework relies solely on attention-based feature querying, which can limit geometric fidelity.

Coverage note — None was omitted; all key contributions including framework architecture, MRSA formulation, weighted masking, selective layer replacement, empirical/efficiency evaluations, user studies, cross-model applications, and limitations are fully covered.

References

  1. 1.Yuval Alaluf, Elad Richardson, Gal Metzer, and Daniel Cohen-Or. A neural space-time representation for text-to-image personalization. ACM Trans. Graph., 42(6), 2023. 1, 2, 6, 7
  2. 2.Omri Avrahami, Kfir Aberman, Ohad Fried, Daniel Cohen-Or, and Dani Lischinski. Break-a-scene: Extracting multiple concepts from a single image. In SIGGRAPH Asia 2023 Conference Papers, New York, NY, USA, 2023. Association for Computing Machinery. 3
  3. 3.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. corr, vol. abs/2211.01324 (2022), 2022. 2
  4. 4.Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22560–22570, 2023. 2, 3, 5
  5. 5.Kaixuan Chen, Jie Song, Shunyu Liu, Na Yu, Zunlei Feng, Gengshi Han, and Mingli Song. Distribution knowledge embedding for graph pooling. 2022. 2
  6. 6.Kaixuan Chen, Shunyu Liu, Tongtian Zhu, Ji Qiao, Yun Su, Yingjie Tian, et al. Improving expressivity of gnns with subgraph-specific factor embedded normalization. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2023. 2
  7. 7.Xi Chen, Lianghua Huang, Yu Liu, Yujun Shen, Deli Zhao, and Hengshuang Zhao. Anydoor: Zero-shot object-level image customization. arXiv preprint, 2023. 2
  8. 8.Timothee Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers, 2023. 6
  9. 9.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021. 1, 2, 3
  10. 10.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022. 1, 2, 6
  11. 11.Rinon Gal, Moab Arar, Yuval Atzmon, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Encoder-based domain tuning for fast personalization of text-to-image models. ACM Transactions on Graphics (TOG), 42(4):1–13, 2023. 2
  12. 12.Yuchao Gu, Xintao Wang, Jay Zhangjie Wu, Yujun Shi, Yunpeng Chen, Zihan Fan, Wuyou Xiao, Rui Zhao, Shuning Chang, Weijia Wu, et al. Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models. Advances in Neural Information Processing Systems, 36, 2024. 3
  13. 13.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022. 3, 6
  14. 14.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 1, 3
  15. 15.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022. 3
  16. 16.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6007–6017, 2023. 2
  17. 17.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4015–4026, 2023. 6
  18. 18.Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1931–1941, 2023. 1, 3, 6, 7
  19. 19.Dongxu Li, Junnan Li, and Steven Hoi. Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing. Advances in Neural Information Processing Systems, 36, 2024. 1, 2, 6, 7, 8
  20. 20.Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023. 2
  21. 21.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023. 6
  22. 22.Zhiheng Liu, Ruili Feng, Kai Zhu, Yifei Zhang, Kecheng Zheng, Yu Liu, Deli Zhao, Jingren Zhou, and Yang Cao. Cones: Concept neurons in diffusion models for customized generation. In International Conference on Machine Learning, 2023. 2, 3
  23. 23.Zhiheng Liu, Yifei Zhang, Yujun Shen, Kecheng Zheng, Kai Zhu, Ruili Feng, Yu Liu, Deli Zhao, Jingren Zhou, and Yang Cao. Cones 2: Customizable image synthesis with multiple subjects. arXiv preprint arXiv:2305.19327, 2023. 3
  24. 24.Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis. Towards end-to-end unified scene text detection and layout analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022. 2
  25. 25.Shangbang Long, Siyang Qin, Dmitry Panteleev, Alessandro Bissacco, Yasuhisa Fujii, and Michalis Raptis. Icdar 2023 competition on hierarchical text detection and recognition. arXiv preprint arXiv:2305.09750, 2023. 2
  26. 26.Alexander Quinn Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pages 8162–8171. PMLR, 2021. 1, 2, 3
  27. 27.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, pages 16784–16804. PMLR, 2022. 2
  28. 28.Maxime Oquab, Timothee Darcet, Theo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Russell Howes, Po-Yao Huang, Hu Xu, Vasu Sharma, Shang-Wen Li, Wojciech Galuba, Mike Rabbat, Mido Assran, Nicolas Ballas, Gabriel Synnaeve, Ishan Misra, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. Dinov2: Learning robust visual features without supervision, 2023. 6
  29. 29.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 6
  30. 30.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1 (2):3, 2022. 2
  31. 31.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 1, 2, 3
  32. 32.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500–22510, 2023. 1, 2, 6, 7
  33. 33.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Wei Wei, Tingbo Hou, Yael Pritch, Neal Wadhwa, Michael Rubinstein, and Kfir Aberman. Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models. arXiv preprint arXiv:2307.06949, 2023. 2, 3
  34. 34.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 1, 3
  35. 35.Yoad Tewel, Rinon Gal, Gal Chechik, and Yuval Atzmon. Key-locked rank one editing for text-to-image personalization. In ACM SIGGRAPH 2023 Conference Proceedings, 2023. 1, 2, 3, 6, 7
  36. 36.Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1921–1930, 2023. 3, 5, 6
  37. 37.Andrey Voynov, Qinghao Chu, Daniel Cohen-Or, and Kfir Aberman. p+: Extended textual conditioning in text-to-image generation. arXiv preprint arXiv:2303.09522, 2023. 2
  38. 38.Jianyi Wang, Kelvin CK Chan, and Chen Change Loy. Exploring clip for assessing the look and feel of images. In AAAI, 2023. 6
  39. 39.Wen Wang, kangyang Xie, Zide Liu, Hao Chen, Yue Cao, Xinlong Wang, and Chunhua Shen. Zero-shot video editing using off-the-shelf image diffusion models. arXiv preprint arXiv:2303.17599, 2023. 6
  40. 40.Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. ELITE: encoding visual concepts into textual embeddings for customized text-to-image generation. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 15897–15907. IEEE, 2023. 2
  41. 41.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7623–7633, 2023. 3
  42. 42.Zhen Yang, Ganggui Ding, Wen Wang, Hao Chen, Bohan Zhuang, and Chunhua Shen. Object-aware inversion and re-assembly for image editing. 2023. 6
  43. 43.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models, 2023. 8

Citation

MLA
Ding, G., et al. “FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 9089–98, https://doi.org/10.1109/CVPR52733.2024.00868.
APA
Ding, G., Zhao, C., Wang, W., Yang, Z., Liu, Z., Chen, H., & Shen, C. (2024). FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9089–9098. https://doi.org/10.1109/CVPR52733.2024.00868
Chicago
Ding, G., C. Zhao, W. Wang, et al. 2024. “FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition”. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 9089–98. https://doi.org/10.1109/CVPR52733.2024.00868.
Harvard
Ding, G. et al. (2024) “FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition”, 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp. 9089–9098. Available at: https://doi.org/10.1109/CVPR52733.2024.00868.
Vancouver
1. Ding G, Zhao C, Wang W, Yang Z, Liu Z, Chen H, Shen C (2024) FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition. In: 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, pp 9089–9098

BibTeX

@inproceedings{Ding_2024, title={FreeCustom: Tuning-Free Customized Image Generation for Multi-Concept Composition}, url={http://dx.doi.org/10.1109/CVPR52733.2024.00868}, DOI={10.1109/cvpr52733.2024.00868}, booktitle={2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, publisher={IEEE}, author={Ding, Ganggui and Zhao, Canyu and Wang, Wen and Yang, Zhen and Liu, Zide and Chen, Hao and Shen, Chunhua}, year={2024}, month=June, pages={9089–9098} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE