DisenBooth: Identity-Preserving Disentangled Tuning for Subject-Driven Text-to-Image Generation
Hong ChenYipeng ZhangSimin WuXin WangXuguang DuanYuwei ZhouWenwu Zhu
Proposes DisenBooth, a tuning framework for diffusion models that disentangles subject identity from background and pose features to prevent identity distortion and improve prompt fidelity in customized text-to-image generation.
Generating customized, photo-realistic images of specific subjects based on text descriptions is an increasingly important capability for creative industries, digital marketing, and virtual design. However, existing customization techniques struggle because they entangle the core identity of a subject with irrelevant details from reference photos, such as background clutter or specific poses. This entanglement leads to two common failure modes: the model either overfits to the original photo's background and ignores new text prompts, or it alters the subject's key visual identity.
The article introduces and evaluates DisenBooth, an identity-preserving framework designed to separate subject identity from extraneous visual details during the model tuning process. The main objective is to demonstrate that disentangling these visual features enables the generation of new, highly accurate scenes that strictly follow text instructions while keeping the subject's appearance intact.
The approach operates by dividing the image conditioning process into two distinct learning branches: a shared textual embedding that captures the core subject identity, and an image-specific visual embedding with a learnable feature-level mask that absorbs identity-irrelevant details like backgrounds and poses. To ensure effective separation without requiring manual image masks, the method applies two auxiliary training objectives: a weak denoising objective and a contrastive embedding objective. Credibility was established through extensive benchmarking on the standard 30-subject DreamBench dataset across 750 unique text prompts (generating 3,000 images total), parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) that trained only 2.9 million parameters instead of the full 865.9 million model parameters, and human evaluations involving 70 distinct users.
DisenBooth achieved the highest text-prompt fidelity among evaluated customization models with a CLIP-T score of 0.330, compared to DreamBooth's 0.319 and Textual Inversion's 0.318. While DreamBooth scored slightly higher in raw image similarity (a DINO score of 0.685 versus DisenBooth's 0.675), user evaluations revealed that DreamBooth's higher similarity was largely due to unwanted background overfitting. In human preference studies evaluating both identity preservation and overall generation quality, DisenBooth achieved top average rankings of 1.37 and 1.59 respectively, substantially outperforming DreamBooth (2.20 and 2.45), InstructPix2Pix (2.69 and 2.89), and Textual Inversion (3.73 and 3.07). DisenBooth also proved effective under constrained conditions, maintaining strong performance and identity disentanglement when fine-tuned on only a single reference image.
These findings indicate that separating core subject traits from background noise significantly improves generative image control without increasing computational complexity. By tuning only around 0.3% of the network parameters, the framework avoids the heavy operational and hardware costs of full-model retraining while preventing the prompt-ignoring failures common in previous tools. Additionally, because the visual details are cleanly separated, users can selectively blend reference poses or backgrounds into new generations with a controllable weighting parameter, providing greater creative flexibility.
Organizations developing personalized generative AI should adopt disentangled tuning architectures and parameter-efficient strategies like LoRA to reduce storage overhead and improve prompt adherence. Before commercial deployment, teams should conduct focused pilots to establish proper training weighting parameters (recommended between 0.001 and 0.01 for auxiliary loss terms) and evaluate system performance on non-centered or low-resolution subject images, potentially pairing the workflow with super-resolution pre-processing.
The findings are subject to certain boundaries: the framework inherits the baseline quality limitations of underlying base diffusion models, faces reduced detail accuracy when subjects occupy very few pixels or when customizing multiple subjects in a single scene, and does not separate background from pose within the irrelevant feature space. Confidence in the reported performance is high across standard single-subject customization benchmarks, but caution is warranted when applying the system to complex multi-object compositions or highly unconstrained input photography without prior image preprocessing.
- Paper: DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation, Nataniel Ruiz et al. (2023). DreamBooth establishes the baseline methodology for fine-tuning text-to-image diffusion models for subject-driven generation, which DisenBooth directly builds upon and enhances by addressing representation entanglement.
- Paper: An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion, Rinon Gal et al. (2022). Textual Inversion introduces the concept of personalizing text-to-image diffusion models via learned pseudo-word embeddings, providing foundational context for subject-driven customization techniques.
- Paper: Multi-Concept Customization of Text-to-Image Diffusion, Nupur Kumari et al. (2022). Custom Diffusion provides foundational methods for tuning cross-attention parameters to bind specific visual concepts in text-to-image models while preventing concept overfitting.
- Paper: High-Resolution Image Synthesis with Latent Diffusion Models, Robin Rombach et al. (2022). Latent Diffusion Models define the underlying text-to-image architecture and latent denoising dynamics that DisenBooth fine-tunes.
- Paper: Null-text Inversion for Editing Real Images using Guided Diffusion Models, Ron Mokady et al. (2022). Null-text Inversion demonstrates how intermediate diffusion embeddings can be optimized for identity and background fidelity during text-driven generation.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). Prompt-to-Prompt introduces cross-attention control mechanisms in diffusion models, offering foundational insights into separating textual attributes from visual identity.
- Paper: Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models, Chang Liu et al. (2024). StoryGen extends subject and identity consistency from single-image customized diffusion generation into open-ended, multi-frame visual storytelling sequences.
