Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs
Ling YangZhaochen YuChenlin MengMinkai XuStefano ErmonBin Cui
Proposes a training-free framework that leverages multimodal large language models to decompose complex prompts into spatial sub-tasks for regional diffusion, outperforming DALL-E 3 and SDXL in multi-object compositional image generation and editing.
Modern text-to-image AI systems struggle to generate images from complex text prompts that involve multiple objects, detailed attributes, and intricate spatial relationships. When faced with dense prompts, existing models frequently mix up characteristics between subjects, miscount items, or place elements incorrectly. Current attempts to fix these issues typically rely on rigid spatial boxes or computationally expensive model retraining, both of which introduce performance trade-offs and added operational costs.
The article demonstrates and evaluates Recaption, Plan and Generate (RPG), a training-free framework designed to improve the compositionality and fidelity of text-to-image diffusion models without altering model weights. RPG leverages the chain-of-thought reasoning of multimodal large language models to act as a global task planner, breaking down complex visual generation and editing tasks into simpler regional subtasks.
The researchers evaluated RPG by pairing advanced multimodal language models (such as GPT-4) with leading image generation backbones (such as Stable Diffusion XL). The framework executes a three-stage workflow: first, it decomposes and expands user prompts into descriptive subprompts; second, it uses chain-of-thought reasoning to divide the image space into non-overlapping subregions; and third, it generates regional image representations in parallel before merging them through complementary regional diffusion. The team benchmarked this approach on standard industry evaluation datasets against leading models like SDXL, DALL-E 3, and specialized layout-guided systems, while also testing closed-loop image editing capabilities.
The evaluation revealed four key findings. First, RPG achieved the highest overall score on the compositional benchmark T2I-CompBench, outperforming all baseline models in attribute binding, spatial relationships, and complex compositions. In spatial relationship accuracy, RPG achieved a score of 0.4781, more than doubling SDXL's score of 0.2032. Second, RPG significantly outperformed commercial state-of-the-art models like DALL-E 3 and SDXL in numeric precision and attribute isolation, successfully preventing attribute bleeding across distinct objects. Third, the framework effectively unified text-guided image generation and iterative editing into a closed-loop system, achieving precise targeted modifications typically within three refinement rounds. Fourth, RPG demonstrated broad architectural flexibility, working successfully across various language models and diffusion backbones, as well as spatial conditioning tools like ControlNet.
These findings indicate that high-fidelity image composition does not necessarily require retraining massive generative models from scratch. Organizations can dramatically enhance image generation accuracy and reduce manual revisions by integrating intelligent language-based planners on top of existing generation backbones. Because the framework is training-free, it minimizes infrastructure and development costs while improving reliability in high-precision visual generation tasks.
Decision-makers should consider adopting modular, planning-guided architectures like RPG when deploying generative visual tools that demand high accuracy across complex prompts. Implementation teams should optimize the balance between global base prompts and localized subprompts depending on whether an image requires shared stylistic context or strict separation between distinct subjects.
While the results demonstrate high confidence and clear empirical advantages across multiple benchmarks, performance remains dependent on the underlying reasoning quality of the multimodal language model and the baseline rendering capacity of the chosen diffusion model. Users should note that excessive reliance on base prompt weighting can introduce semantic confusion, requiring careful calibration of blending parameters during production deployment.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). This foundational work establishes attention manipulation in text-to-image diffusion, providing the mechanistic baseline for the region-specific cross-attention control employed in RPG.
- Paper: Blended Diffusion for Text-driven Editing of Natural Images, Omri Avrahami et al. (2021). It introduces spatial region blending during iterative diffusion steps, laying key conceptual groundwork for complementary regional diffusion across decomposed subregions.
- Paper: InstructPix2Pix: Learning to Follow Image Editing Instructions, Tim Brooks et al. (2023). It pioneers instruction-guided image editing with diffusion models, which RPG extends into an iterative, closed-loop generation and editing framework guided by multimodal reasoning.
- Paper: What the DAAM: Interpreting Stable Diffusion Using Cross Attention, Raphael Tang et al. (2023). This paper analyzes the spatial grounding of text tokens in diffusion cross-attention maps, explaining the semantic binding challenges that RPG directly addresses through multimodal LLM planning.
- Paper: Optimizing Prompts for Text-to-Image Generation, Yaru Hao et al. (2023). It demonstrates how language models can optimize and elaborate user prompts for text-to-image synthesis, a prerequisite principle behind the recaptioning stage in RPG.
- Paper: MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis, Dewei Zhou et al. (2024). This work advances multi-instance compositional generation by integrating instance-level coordinate embeddings directly into cross-attention layers, building beyond high-level regional LLM decomposition.
- Paper: SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models, Yuzhou Huang et al. (2024). It builds upon MLLM-directed visual reasoning by creating bidirectional interaction mechanisms to carry out complex, multi-step instruction-based image editing.
- Paper: UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics, Xi Chen et al. (2025). It unifies compositional multi-image generation and instruction editing into a single large-scale diffusion transformer architecture leveraging learned physical dynamics.
- Paper: Dysen-VDM: Empowering Dynamics-Aware Text-to-Video Diffusion with LLMs, Hao Fei et al. (2024). It extends the concept of LLM-based structured scene planning from static regional image composition to dynamic temporal scene graphs in video diffusion.
