MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
Dewei ZhouYou LiFan MaXiaoting ZhangYi Yang
Proposes a divide-and-conquer framework for Stable Diffusion that enables precise multi-instance text-to-image synthesis by shading individual instances separately through dedicated attention mechanisms before aggregating them into a unified layout.
Generating complex images that contain multiple distinct objects remains a major hurdle for modern text-to-image systems. While existing models excel at creating single subjects from simple text prompts, they frequently fail when asked to generate multiple instances within specified bounding areas. Typical failures include attribute leakage—where colors, textures, or shapes bleed across separate objects—as well as missing instances and unwanted object merging caused by weak spatial guidance in standard text encoders and attention layers.
The article introduces and evaluates the Multi-Instance Generation Controller, a framework designed to enable standard diffusion models to accurately synthesize multiple instances at defined positions while preserving distinct attributes, quantities, and harmonious global compositions. To rigorously measure performance on this task, the authors also develop a dedicated evaluation benchmark known as COCO-MIG alongside evaluations on standard vision benchmarks.
The controller uses a divide-and-conquer strategy operating directly within the cross-attention layers of a pre-trained diffusion model. First, it breaks down the complex multi-object scene into single-instance shading subtasks bounded by spatial masks. Second, it resolves missing and merged instances through an Enhancement Attention layer that pairs text descriptions with Fourier-embedded position tokens. Third, it synthesizes the isolated instances, the background, and a Layout Attention template into a coherent final image using an attention-based Shading Aggregation Controller, further supported by an inhibition loss to prevent background artifacts. Testing was conducted across 6,400 synthetic evaluations on the COCO-MIG and COCO-Position datasets, as well as 512 test samples on DrawBench.
The evaluation demonstrates substantial improvements over prior leading methods across key metrics. On the primary benchmark, the instance generation success rate rose from 32.39% in the best baseline to 58.43%, while mean intersection over union improved from 32.25 to 51.48. On standard spatial layout benchmarks, the average precision metric increased from 40.68 to 54.69, and position success reached 80.29% compared to the prior 70.52%. On the DrawBench attribute tests, human-evaluated success climbed dramatically from 48.20% to 97.50%. Crucially, the approach achieves these control improvements without degrading overall image quality or significantly slowing down generation runtimes.
These results show that precise spatial and compositional control can be added to pre-trained generative diffusion models without requiring full retraining from scratch or heavy computational overhead during generation. By decomposing multi-object interactions into modular attention subtasks, organizations can significantly reduce generation failure rates, streamline automated digital design pipelines, and eliminate costly trial-and-error prompting cycles.
For practical adoption, engineering teams developing image generation tools should integrate localized attention conditioning modules rather than relying strictly on standard text-prompt engineering. Moving forward, the research suggests extending this architectural framework to better model complex physical interactions between adjacent instances and deploying targeted validation pilots in commercial workflows to assess complex composition needs.
Confidence in the reported improvements is high based on consistent automated and human evaluation gains across multiple public datasets. However, stakeholders should note that bounding boxes currently require explicit pre-definition or upstream extraction via language and detection models, and high inhibition loss settings must be carefully tuned to avoid subtle trade-offs in image quality.
- Paper: T2I-CompBench: A Comprehensive Benchmark for Open-world Compositional Text-to-image Generation, Kaiyi Huang et al. (2023). This benchmark establishes the standard evaluative framework and metrics for compositional text-to-image generation, directly motivating MIGC's focus on resolving attribute binding and spatial layout failures.
- Paper: Prompt-to-Prompt Image Editing with Cross Attention Control, Amir Hertz et al. (2022). This foundational work establishes the methodology of manipulating internal cross-attention maps to control spatial layout and semantics in diffusion models, a core principle leveraged in MIGC's attention controller design.
- Paper: Adding Conditional Control to Text-to-Image Diffusion Models, Lvmin Zhang et al. (2023). ControlNet provides essential background on incorporating spatial conditioning into pre-trained diffusion models without requiring full retraining from scratch.
- Paper: What the DAAM: Interpreting Stable Diffusion Using Cross Attention, Raphael Tang et al. (2023). DAAM demonstrates how cross-attention maps reflect linguistic syntax and instance-level attribution in Stable Diffusion, providing key interpretability insights that MIGC builds upon for instance-specific shading.
- Paper: T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models, Chong Mou et al. (2023). T2I-Adapter introduces lightweight conditioning mechanisms for pre-trained diffusion models, serving as key prior work for modular attention-based spatial guidance.
- Paper: IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models, Hu Ye et al. (2023). IP-Adapter outlines how decoupling cross-attention pathways enables modular conditioning in diffusion models without altering base network weights.
- Paper: Multi-Concept Customization of Text-to-Image Diffusion, Nupur Kumari et al. (2022). Custom Diffusion analyzes cross-attention parameter updates to bind and compose multiple distinct concepts in diffusion generation, addressing related multi-concept challenges.
- Paper: Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding, Chitwan Saharia et al. (2022). Imagen introduces the DrawBench evaluation suite and demonstrates language-conditioned diffusion architectures, both of which serve as standard baselines and benchmarks in MIGC.
- Paper: Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models, Zengbin Wang et al. (2026). SpatialGenEval broadens the evaluation of multi-object spatial reasoning and arrangement in generative models, providing a comprehensive benchmark to evaluate advanced compositional controllers like MIGC.
- Paper: Initno: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization, Xiefan Guo et al. (2024). INITNO offers a complementary inference-time approach to multi-instance alignment by optimizing initial latent noise to prevent subject omission and blending in diffusion models.
- Paper: DiffEditor: Boosting Accuracy and Flexibility on Diffusion-Based Image Editing, Chong Mou et al. (2024). DiffEditor applies localized diffusion guidance principles to fine-grained, interactive multi-object spatial editing and manipulation.
- Paper: SmartEdit: Exploring Complex Instruction-Based Image Editing with Multimodal Large Language Models, Yuzhou Huang et al. (2024). SmartEdit integrates multimodal reasoning into diffusion pipelines to identify and manipulate specific spatial instances based on complex compositional instructions.
- Paper: Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models, Chang Liu et al. (2024). StoryGen extends multi-instance visual and character control from single-frame synthesis to sequential storytelling across multi-frame contexts.
