Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Ling YangZhaochen YuChenlin MengMinkai XuStefano ErmonBin Cui

article2024ICML224 citations

Proposes a training-free framework that leverages multimodal large language models to decompose complex prompts into spatial sub-tasks for regional diffusion, outperforming DALL-E 3 and SDXL in multi-object compositional image generation and editing.

Listen

Modern text-to-image AI systems struggle to generate images from complex text prompts that involve multiple objects, detailed attributes, and intricate spatial relationships. When faced with dense prompts, existing models frequently mix up characteristics between subjects, miscount items, or place elements incorrectly. Current attempts to fix these issues typically rely on rigid spatial boxes or computationally expensive model retraining, both of which introduce performance trade-offs and added operational costs.

The article demonstrates and evaluates Recaption, Plan and Generate (RPG), a training-free framework designed to improve the compositionality and fidelity of text-to-image diffusion models without altering model weights. RPG leverages the chain-of-thought reasoning of multimodal large language models to act as a global task planner, breaking down complex visual generation and editing tasks into simpler regional subtasks.

The researchers evaluated RPG by pairing advanced multimodal language models (such as GPT-4) with leading image generation backbones (such as Stable Diffusion XL). The framework executes a three-stage workflow: first, it decomposes and expands user prompts into descriptive subprompts; second, it uses chain-of-thought reasoning to divide the image space into non-overlapping subregions; and third, it generates regional image representations in parallel before merging them through complementary regional diffusion. The team benchmarked this approach on standard industry evaluation datasets against leading models like SDXL, DALL-E 3, and specialized layout-guided systems, while also testing closed-loop image editing capabilities.

The evaluation revealed four key findings. First, RPG achieved the highest overall score on the compositional benchmark T2I-CompBench, outperforming all baseline models in attribute binding, spatial relationships, and complex compositions. In spatial relationship accuracy, RPG achieved a score of 0.4781, more than doubling SDXL's score of 0.2032. Second, RPG significantly outperformed commercial state-of-the-art models like DALL-E 3 and SDXL in numeric precision and attribute isolation, successfully preventing attribute bleeding across distinct objects. Third, the framework effectively unified text-guided image generation and iterative editing into a closed-loop system, achieving precise targeted modifications typically within three refinement rounds. Fourth, RPG demonstrated broad architectural flexibility, working successfully across various language models and diffusion backbones, as well as spatial conditioning tools like ControlNet.

These findings indicate that high-fidelity image composition does not necessarily require retraining massive generative models from scratch. Organizations can dramatically enhance image generation accuracy and reduce manual revisions by integrating intelligent language-based planners on top of existing generation backbones. Because the framework is training-free, it minimizes infrastructure and development costs while improving reliability in high-precision visual generation tasks.

Decision-makers should consider adopting modular, planning-guided architectures like RPG when deploying generative visual tools that demand high accuracy across complex prompts. Implementation teams should optimize the balance between global base prompts and localized subprompts depending on whether an image requires shared stylistic context or strict separation between distinct subjects.

While the results demonstrate high confidence and clear empirical advantages across multiple benchmarks, performance remains dependent on the underlying reasoning quality of the multimodal language model and the baseline rendering capacity of the chosen diffusion model. Users should note that excessive reliance on base prompt weighting can introduce semantic confusion, requiring careful calibration of blending parameters during production deployment.

Cover for Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Abstract

Diffusion models have exhibit exceptional performance in text-to-image generation and editing. However, existing methods often face challenges when handling complex text prompts that involve multiple objects with multiple attributes and relationships. In this paper, we propose a brand new training-free text-to-image generation/editing framework, namely Recaption, Plan and Generate (RPG), harnessing the powerful chain-of-thought reasoning ability of multimodal LLMs to enhance the compositionality of text-to-image diffusion models. Our approach employs the MLLM as a global planner to decompose the process of generating complex images into multiple simpler generation tasks within subregions. We propose complementary regional diffusion to enable region-wise compositional generation. Furthermore, we integrate text-guided image generation and editing within the proposed RPG in a closed-loop fashion, thereby enhancing generalization ability. Extensive experiments demonstrate our RPG outperforms state-of-the-art text-to-image models, including DALL-E 3 and SDXL, particularly in multi-category object composition and text-image semantic alignment. Notably, our RPG framework exhibits wide compatibility with various MLLM architectures and diffusion backbones. Our code is available at https://github.com/YangLing0818/RPG-DiffusionMaster.

Citation

MLA
Yang, L., et al. “Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs”. arXiv, 2024, http://arxiv.org/abs/2401.11708v3.
APA
Yang, L., Yu, Z., Meng, C., Xu, M., Ermon, S., & Cui, B. (2024). Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs. arXiv. http://arxiv.org/abs/2401.11708v3
Chicago
Yang, L., Z. Yu, C. Meng, M. Xu, S. Ermon, and B. Cui. 2024. “Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs”. arXiv. http://arxiv.org/abs/2401.11708v3.
Harvard
Yang, L. et al. (2024) “Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2401.11708v3.
Vancouver
1. Yang L, Yu Z, Meng C, Xu M, Ermon S, Cui B (2024) Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs. arXiv

BibTeX

@article{yang2024mastering,
  title = {Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs},
  author = {Yang, Ling and Yu, Zhaochen and Meng, Chenlin and Xu, Minkai and Ermon, Stefano and Cui, Bin},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2401.11708v3},
  eprint = {2401.11708}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/