DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

Hong ChenYipeng ZhangSimin WuXin WangXuguang DuanYuwei ZhouWenwu Zhu

article2024ICLR81 citations

Proposes DisenBooth, a tuning framework for diffusion models that disentangles subject identity from background and pose features to prevent identity distortion and improve prompt fidelity in customized text-to-image generation.

Listen

Generating customized, photo-realistic images of specific subjects based on text descriptions is an increasingly important capability for creative industries, digital marketing, and virtual design. However, existing customization techniques struggle because they entangle the core identity of a subject with irrelevant details from reference photos, such as background clutter or specific poses. This entanglement leads to two common failure modes: the model either overfits to the original photo's background and ignores new text prompts, or it alters the subject's key visual identity.

The article introduces and evaluates DisenBooth, an identity-preserving framework designed to separate subject identity from extraneous visual details during the model tuning process. The main objective is to demonstrate that disentangling these visual features enables the generation of new, highly accurate scenes that strictly follow text instructions while keeping the subject's appearance intact.

The approach operates by dividing the image conditioning process into two distinct learning branches: a shared textual embedding that captures the core subject identity, and an image-specific visual embedding with a learnable feature-level mask that absorbs identity-irrelevant details like backgrounds and poses. To ensure effective separation without requiring manual image masks, the method applies two auxiliary training objectives: a weak denoising objective and a contrastive embedding objective. Credibility was established through extensive benchmarking on the standard 30-subject DreamBench dataset across 750 unique text prompts (generating 3,000 images total), parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) that trained only 2.9 million parameters instead of the full 865.9 million model parameters, and human evaluations involving 70 distinct users.

DisenBooth achieved the highest text-prompt fidelity among evaluated customization models with a CLIP-T score of 0.330, compared to DreamBooth's 0.319 and Textual Inversion's 0.318. While DreamBooth scored slightly higher in raw image similarity (a DINO score of 0.685 versus DisenBooth's 0.675), user evaluations revealed that DreamBooth's higher similarity was largely due to unwanted background overfitting. In human preference studies evaluating both identity preservation and overall generation quality, DisenBooth achieved top average rankings of 1.37 and 1.59 respectively, substantially outperforming DreamBooth (2.20 and 2.45), InstructPix2Pix (2.69 and 2.89), and Textual Inversion (3.73 and 3.07). DisenBooth also proved effective under constrained conditions, maintaining strong performance and identity disentanglement when fine-tuned on only a single reference image.

These findings indicate that separating core subject traits from background noise significantly improves generative image control without increasing computational complexity. By tuning only around 0.3% of the network parameters, the framework avoids the heavy operational and hardware costs of full-model retraining while preventing the prompt-ignoring failures common in previous tools. Additionally, because the visual details are cleanly separated, users can selectively blend reference poses or backgrounds into new generations with a controllable weighting parameter, providing greater creative flexibility.

Organizations developing personalized generative AI should adopt disentangled tuning architectures and parameter-efficient strategies like LoRA to reduce storage overhead and improve prompt adherence. Before commercial deployment, teams should conduct focused pilots to establish proper training weighting parameters (recommended between 0.001 and 0.01 for auxiliary loss terms) and evaluate system performance on non-centered or low-resolution subject images, potentially pairing the workflow with super-resolution pre-processing.

The findings are subject to certain boundaries: the framework inherits the baseline quality limitations of underlying base diffusion models, faces reduced detail accuracy when subjects occupy very few pixels or when customizing multiple subjects in a single scene, and does not separate background from pose within the irrelevant feature space. Confidence in the reported performance is high across standard single-subject customization benchmarks, but caution is warranted when applying the system to complex multi-object compositions or highly unconstrained input photography without prior image preprocessing.

Cover for DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation

Abstract

Subject-driven text-to-image generation aims to generate customized images of the given subject based on the text descriptions, which has drawn increasing attention. Existing methods mainly resort to finetuning a pretrained generative model, where the identity-relevant information (e.g., the boy) and the identity-irrelevant information (e.g., the background or the pose of the boy) are entangled in the latent embedding space. However, the highly entangled latent embedding may lead to the failure of subject-driven text-to-image generation as follows: (i) the identity-irrelevant information hidden in the entangled embedding may dominate the generation process, resulting in the generated images heavily dependent on the irrelevant information while ignoring the given text descriptions; (ii) the identity-relevant information carried in the entangled embedding can not be appropriately preserved, resulting in identity change of the subject in the generated images. To tackle the problems, we propose DisenBooth, an identity-preserving disentangled tuning framework for subject-driven text-to-image generation. Specifically, DisenBooth finetunes the pretrained diffusion model in the denoising process. Different from previous works that utilize an entangled embedding to denoise each image, DisenBooth instead utilizes disentangled embeddings to respectively preserve the subject identity and capture the identity-irrelevant information. We further design the novel weak denoising and contrastive embedding auxiliary tuning objectives to achieve the disentanglement. Extensive experiments show that our proposed DisenBooth framework outperforms baseline models for subject-driven text-to-image generation with the identity-preserved embedding. Additionally, by combining the identity-preserved embedding and identity-irrelevant embedding, DisenBooth demonstrates more generation flexibility and controllability

Citation

MLA
Chen, H., et al. “DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation”. arXiv, 2023, http://arxiv.org/abs/2305.03374v4.
APA
Chen, H., Zhang, Y., Wu, S., Wang, X., Duan, X., Zhou, Y., & Zhu, W. (2023). DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation. arXiv. http://arxiv.org/abs/2305.03374v4
Chicago
Chen, H., Y. Zhang, S. Wu, et al. 2023. “DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation”. arXiv. http://arxiv.org/abs/2305.03374v4.
Harvard
Chen, H. et al. (2023) “DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.03374v4.
Vancouver
1. Chen H, Zhang Y, Wu S, Wang X, Duan X, Zhou Y, Zhu W (2023) DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation. arXiv

BibTeX

@article{chen2023disenbooth,
  title = {DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation},
  author = {Chen, Hong and Zhang, Yipeng and Wu, Simin and Wang, Xin and Duan, Xuguang and Zhou, Yuwei and Zhu, Wenwu},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.03374v4},
  eprint = {2305.03374}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors