MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

Dewei ZhouYou LiFan MaXiaoting ZhangYi Yang

article2024CVPR135 citations

Proposes a divide-and-conquer framework for Stable Diffusion that enables precise multi-instance text-to-image synthesis by shading individual instances separately through dedicated attention mechanisms before aggregating them into a unified layout.

Listen

Generating complex images that contain multiple distinct objects remains a major hurdle for modern text-to-image systems. While existing models excel at creating single subjects from simple text prompts, they frequently fail when asked to generate multiple instances within specified bounding areas. Typical failures include attribute leakage—where colors, textures, or shapes bleed across separate objects—as well as missing instances and unwanted object merging caused by weak spatial guidance in standard text encoders and attention layers.

The article introduces and evaluates the Multi-Instance Generation Controller, a framework designed to enable standard diffusion models to accurately synthesize multiple instances at defined positions while preserving distinct attributes, quantities, and harmonious global compositions. To rigorously measure performance on this task, the authors also develop a dedicated evaluation benchmark known as COCO-MIG alongside evaluations on standard vision benchmarks.

The controller uses a divide-and-conquer strategy operating directly within the cross-attention layers of a pre-trained diffusion model. First, it breaks down the complex multi-object scene into single-instance shading subtasks bounded by spatial masks. Second, it resolves missing and merged instances through an Enhancement Attention layer that pairs text descriptions with Fourier-embedded position tokens. Third, it synthesizes the isolated instances, the background, and a Layout Attention template into a coherent final image using an attention-based Shading Aggregation Controller, further supported by an inhibition loss to prevent background artifacts. Testing was conducted across 6,400 synthetic evaluations on the COCO-MIG and COCO-Position datasets, as well as 512 test samples on DrawBench.

The evaluation demonstrates substantial improvements over prior leading methods across key metrics. On the primary benchmark, the instance generation success rate rose from 32.39% in the best baseline to 58.43%, while mean intersection over union improved from 32.25 to 51.48. On standard spatial layout benchmarks, the average precision metric increased from 40.68 to 54.69, and position success reached 80.29% compared to the prior 70.52%. On the DrawBench attribute tests, human-evaluated success climbed dramatically from 48.20% to 97.50%. Crucially, the approach achieves these control improvements without degrading overall image quality or significantly slowing down generation runtimes.

These results show that precise spatial and compositional control can be added to pre-trained generative diffusion models without requiring full retraining from scratch or heavy computational overhead during generation. By decomposing multi-object interactions into modular attention subtasks, organizations can significantly reduce generation failure rates, streamline automated digital design pipelines, and eliminate costly trial-and-error prompting cycles.

For practical adoption, engineering teams developing image generation tools should integrate localized attention conditioning modules rather than relying strictly on standard text-prompt engineering. Moving forward, the research suggests extending this architectural framework to better model complex physical interactions between adjacent instances and deploying targeted validation pilots in commercial workflows to assess complex composition needs.

Confidence in the reported improvements is high based on consistent automated and human evaluation gains across multiple public datasets. However, stakeholders should note that bounding boxes currently require explicit pre-definition or upstream extraction via language and detection models, and high inhibition loss settings must be carefully tuned to avoid subtle trade-offs in image quality.

arXiv: 2402.05408
Cover for MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis

Abstract

We present a Multi-Instance Generation (MIG) task, simultaneously generating multiple instances with diverse controls in one image. Given a set of predefined coordinates and their corresponding descriptions, the task is to ensure that generated instances are accurately at the designated locations and that all instances' attributes adhere to their corresponding description. This broadens the scope of current research on Single-instance generation, elevating it to a more versatile and practical dimension. Inspired by the idea of divide and conquer, we introduce an innovative approach named Multi-Instance Generation Controller (MIGC) to address the challenges of the MIG task. Initially, we break down the MIG task into several subtasks, each involving the shading of a single instance. To ensure precise shading for each instance, we introduce an instance enhancement attention mechanism. Lastly, we aggregate all the shaded instances to provide the necessary information for accurately generating multiple instances in stable diffusion (SD). To evaluate how well generation models perform on the MIG task, we provide a COCO-MIG benchmark along with an evaluation pipeline. Extensive experiments were conducted on the proposed COCO-MIG benchmark, as well as on various commonly used benchmarks. The evaluation results illustrate the exceptional control capabilities of our model in terms of quantity, position, attribute, and interaction. Code and demos will be released at https://migcproject.github.io/.

Table of Contents

  • 1. Introduction
  • 2. Related work
  • 2.1. Text-to-Image Generation
  • 2.2. Layout-to-Image Generation
  • 3. Method
  • 3.1. Preliminaries
  • 3.2. Overview
  • 3.3. Divide MIG into Instance Shading Subtasks
  • 3.4. Conquer Instance Shading
  • 3.5. Combine Shading Results
  • 3.6. Summary
  • 4. Experiments
  • 4.1. Benchmarks
  • 4.2. Evaluation Metrics
  • 4.3. Baselines
  • 4.4. Quantitative Results
  • 4.5. Qualitative Results
  • 4.6. Analysis of Shading Aggregation Controller
  • 4.7. Ablation Study
  • 5. Conclusion
  • References

Knowls

  1. Knowl 1 — Multi-Instance Generation (MIG) Task Formulation

    definition

    The Multi-Instance Generation (MIG) task is defined as synthesizing an image II conditioned on three user inputs:

    • A global text prompt P\mathcal{P};
    • A set of NN instance bounding box coordinates B={b1,b2,…,bN}\mathbb{B} = \{b^1, b^2, \dots, b^N\}, where each bounding box is specified by normalized coordinates bi=[x1i,y1i,x2i,y2i]∈[0,1]4b^i = [x_1^i, y_1^i, x_2^i, y_2^i] \in [0, 1]^4;
    • A set of corresponding instance descriptions D={d1,d2,…,dN}\mathbb{D} = \{d^1, d^2, \dots, d^N\}.

    The goal is to generate an image II such that for each i∈{1,…,N}i \in \{1, \dots, N\}, the visual object generated strictly within bounding box bib^i conforms to the instance description did^i (including attributes like color, count, shape, material, texture, and style), while ensuring global visual coherence across instances and the scene context P\mathcal{P}. The core difficulties of MIG in diffusion models are twofold:

    1. Textual Leakage: Semantic confusion occurring during prompt encoding where attributes of one instance erroneously attach to other instances.
    2. Spatial Leakage: Cross-attention bleeding where instance generation overflows designated spatial boundaries and interferes with neighboring regions.
  2. Knowl 2 — Single-Instance Shading Enhancement with Grounded Phrase Tokens

    model/method

    The Multi-Instance Generation Controller (MIGC) divides the multi-instance generation task into independent single-instance shading subtasks executed directly within the Cross-Attention space of a latent diffusion model's UNet. For each instance i∈{1,…,N}i \in \{1, \dots, N\} defined by bounding box bib^i and text description did^i:

    1. Initial Masked Cross-Attention Shading: An initial instance shading feature Rfi\mathbf{R}^i_f is computed using the pre-trained Cross-Attention layer: Rfi=Softmax(Q(Ki)Td)Vi⋅Mi\mathbf{R}^i_f = \text{Softmax}\left(\frac{\mathbf{Q} (\mathbf{K}^i)^T}{\sqrt{d}}\right) \mathbf{V}^i \cdot \mathbf{M}^i where Q∈R(H⋅W)×d\mathbf{Q} \in \mathbb{R}^{(H \cdot W) \times d} is projected from the image feature map, Ki,Vi\mathbf{K}^i, \mathbf{V}^i are projected from the CLIP text embedding of did^i, and Mi∈{0,1}H×W\mathbf{M}^i \in \{0, 1\}^{H \times W} is a binary spatial mask with 11 inside box bib^i and 00 elsewhere.

    2. Grounded Phrase Tokens for Instance Disambiguation: To eliminate instance merge when separate bounding boxes share identical descriptions did^i, text tokens are augmented with spatial position tokens: Gi=[CLIP(di),MLP(Fourier(bi))]\mathbf{G}^i = [\text{CLIP}(d^i), \text{MLP}(\text{Fourier}(b^i))] where Fourier(⋅)\text{Fourier}(\cdot) maps box coordinates to Fourier embeddings, MLP(⋅)\text{MLP}(\cdot) projects them to position tokens, and [⋅,⋅][\cdot, \cdot] denotes concatenation.

    3. Enhancement Attention (EA) Layer: To resolve instance missing caused by unfavorable initial noise distributions, a trainable Enhancement Attention Cross-Attention layer computes an additive residual constrained to Mi\mathbf{M}^i: Rsi=Rfi+Softmax(Qea(Keai)Td)Veai⋅Mi\mathbf{R}^i_s = \mathbf{R}^i_f + \text{Softmax}\left(\frac{\mathbf{Q}_{ea} (\mathbf{K}^i_{ea})^T}{\sqrt{d}}\right) \mathbf{V}^i_{ea} \cdot \mathbf{M}^i where Qea\mathbf{Q}_{ea} is projected from the image feature map, and Keai,Veai\mathbf{K}^i_{ea}, \mathbf{V}^i_{ea} are projected from the grounded phrase tokens Gi\mathbf{G}^i. The resulting tensor Rsi\mathbf{R}^i_s forms the solved shading feature for instance ii.

  3. Knowl 3 — Layout Attention for Harmonized Shading Template Generation

    model/method

    To bridge the representational gap between independently computed instance shading results {Rs1,…,RsN}\{\mathbf{R}^1_s, \dots, \mathbf{R}^N_s\} and the background shading result Rbg\mathbf{R}^{bg} (computed from the global prompt P\mathcal{P} with background mask Mbg=1−⋃i=1NMi\mathbf{M}^{bg} = 1 - \bigcup_{i=1}^N \mathbf{M}^i), MIGC utilizes a Layout Attention (LA) layer to extract a global shading template RLA\mathbf{R}_{LA} conditioned on image features while preventing attribute leakage between regions.

    The Layout Attention residual is formulated as: RLA=Softmax(QLAKLATd⊙A)VLA\mathbf{R}_{LA} = \text{Softmax}\left(\frac{\mathbf{Q}_{LA} \mathbf{K}_{LA}^T}{\sqrt{d}} \odot \mathbf{A}\right) \mathbf{V}_{LA} where QLA,KLA,VLA∈R(H⋅W)×d\mathbf{Q}_{LA}, \mathbf{K}_{LA}, \mathbf{V}_{LA} \in \mathbb{R}^{(H \cdot W) \times d} are projected from the UNet intermediate image feature map, ⊙\odot represents the Hadamard product, and A∈R(H⋅W)×(H⋅W)\mathbf{A} \in \mathbb{R}^{(H \cdot W) \times (H \cdot W)} is a structured attention mask defined across pixel spatial coordinates (a,b)(a, b) and (c,d)(c, d) by: A(a,b),(c,d)={1,if ∃m∈Minst such that ma,b=mc,d=1−∞,otherwise\mathbf{A}_{(a,b),(c,d)} = \begin{cases} 1, & \text{if } \exists \mathbf{m} \in \mathbb{M}_{inst} \text{ such that } \mathbf{m}_{a,b} = \mathbf{m}_{c,d} = 1 \\ -\infty, & \text{otherwise} \end{cases} where Minst={Mbg,M1,M2,…,MN}\mathbb{M}_{inst} = \{\mathbf{M}^{bg}, \mathbf{M}^1, \mathbf{M}^2, \dots, \mathbf{M}^N\}. The mask A\mathbf{A} restricts self-attention at each pixel strictly to other pixels belonging to the same instance or the shared background, eliminating cross-instance feature bleeding.

  4. Knowl 4 — Shading Aggregation Controller (SAC)

    model/method

    The Shading Aggregation Controller (SAC) dynamically fuses foreground instance shadings, background shading, and the layout attention template across diffusion sampling timesteps into a final Cross-Attention residual Rfinal∈RH×W×C\mathbf{R}_{final} \in \mathbb{R}^{H \times W \times C}.

    Given the set of N+2N+2 candidate shading feature maps and corresponding spatial guidance masks: Rs={Rs1,Rs2,…,RsN,Rbg,RLA}∈R(N+2)×C×H×W\mathbb{R}_s = \{\mathbf{R}^1_s, \mathbf{R}^2_s, \dots, \mathbf{R}^N_s, \mathbf{R}^{bg}, \mathbf{R}_{LA}\} \in \mathbb{R}^{(N+2) \times C \times H \times W} M={M1,M2,…,MN,Mbg,MLA}∈R(N+2)×1×H×W\mathbb{M} = \{\mathbf{M}^1, \mathbf{M}^2, \dots, \mathbf{M}^N, \mathbf{M}^{bg}, \mathbf{M}_{LA}\} \in \mathbb{R}^{(N+2) \times 1 \times H \times W} where MLA\mathbf{M}_{LA} is an all-ones mask 1H×W\mathbf{1}^{H \times W}, SAC concatenates Rs\mathbb{R}_s and M\mathbb{M} along the channel dimension to form a tensor in R(N+2)×(C+1)×H×W\mathbb{R}^{(N+2) \times (C+1) \times H \times W}.

    SAC sequentially executes:

    1. Instance Intra-Attention: Self-attention within individual instance feature slices.
    2. Instance Inter-Attention: Cross-instance interaction across all N+2N+2 shading candidates.
    3. Spatial Softmax Normalization: Computes aggregation weights W∈R(N+2)×1×H×W\mathbf{W} \in \mathbb{R}^{(N+2) \times 1 \times H \times W} normalized along the instance dimension (dimension 0) such that ∑j=1N+2Wj,1,h,w=1\sum_{j=1}^{N+2} \mathbf{W}_{j, 1, h, w} = 1.

    The final integrated shading residual is computed via weighted summation: Rfinal=∑j=1N+2Wj⊙Rs,j\mathbf{R}_{final} = \sum_{j=1}^{N+2} \mathbf{W}_j \odot \mathbb{R}_{s, j} During early sampling steps, SAC assigns higher aggregation weights to foreground Enhancement Attention features Rsi\mathbf{R}_s^i and background Layout Attention templates RLA\mathbf{R}_{LA}, gradually shifting background attention to the global context Rbg\mathbf{R}^{bg} in later steps.

  5. Knowl 5 — Multi-Instance Generation Controller Training Loss and Background Inhibition

    equation

    The parameters θ′\theta' of the Multi-Instance Generation Controller (MIGC) are trained while keeping the underlying pre-trained latent diffusion model parameters θ\theta frozen. The total optimization objective is: min⁡θ′L=LLDM+λLihbt\min_{\theta'} \mathcal{L} = \mathcal{L}_{\text{LDM}} + \lambda \mathcal{L}_{\text{ihbt}} with loss weight λ=0.1\lambda = 0.1, where:

    1. Latent Denoising Loss (LLDM\mathcal{L}_{\text{LDM}}): LLDM=Ez,ϵ∼N(0,I),t[∥ϵ−fθ,θ′(zt,t,P,B,D)∥22]\mathcal{L}_{\text{LDM}} = \mathbb{E}_{\mathbf{z}, \epsilon \sim \mathcal{N}(\mathbf{0}, \mathbf{I}), t}\left[ \left\| \epsilon - f_{\theta, \theta'}(\mathbf{z}_t, t, \mathcal{P}, \mathbb{B}, \mathbb{D}) \right\|_2^2 \right] where zt\mathbf{z}_t is the noisy latent at timestep tt, ϵ\epsilon is standard Gaussian noise, P\mathcal{P} is the global prompt, B\mathbb{B} is the bounding box set, and D\mathbb{D} is the set of instance descriptions.

    2. Background Inhibition Loss (Lihbt\mathcal{L}_{\text{ihbt}}): To constrain instances within their designated bounding boxes and suppress the emergence of spurious objects in the background, an attention inhibition loss is applied: Lihbt=∑i=1N∣Aci−DNR(Aci)∣⊙Mbg\mathcal{L}_{\text{ihbt}} = \sum_{i=1}^N \left| \mathbf{A}_c^i - \text{DNR}(\mathbf{A}_c^i) \right| \odot \mathbf{M}^{bg} where Aci\mathbf{A}_c^i denotes the cross-attention map for the ii-th instance in the frozen 16×1616 \times 16 Cross-Attention layer of the UNet decoder, Mbg\mathbf{M}^{bg} is the background mask (00 on all instance boxes, 11 elsewhere), and DNR(⋅)\text{DNR}(\cdot) denotes background region denoising using a spatial average operation.

  6. Knowl 6 — COCO-MIG Benchmark and Evaluation Protocol

    experimental setup

    The COCO-MIG benchmark evaluates simultaneous control over position, color attribute, and instance quantity in text-to-image synthesis:

    • Dataset Construction: 800 images are sampled from MS COCO 2014, retaining original instance bounding box layouts. Each instance is assigned a designated color attribute, and global prompts are constructed via the template "a <attr1> <obj1> and a <attr2> <obj2> and a ...".
    • Difficulty Levels: The benchmark is split into five levels L2,L3,L4,L5,L6L_2, L_3, L_4, L_5, L_6, where level LkL_k requires generating exactly kk controlled instances. With 8 random seeds per prompt, models generate 6,400 images in total.
    • Position Evaluation: Grounding-DINO detects instances in the generated image. An instance is classified as Position Correctly Generated if the maximum Intersection over Union (IoU) between the predicted box and the ground truth box is ≥0.5\ge 0.5.
    • Attribute Evaluation: For position-correct instances, Grounded-SAM segments the instance mask. An instance is classified as Fully Correctly Generated if the designated target color occupies at least S=20%S = 20\% (S=0.2S = 0.2) of the segmented pixels in HSV color space.
    • Metrics:
      • Instance Success Rate (%): Percentage of specified instances across all images that are Fully Correctly Generated.
      • mIoU: Mean of maximum IoUs across all target instances, where instances failing the color attribute criterion are penalized with an IoU of 00.
  7. Knowl 7 — Quantitative Performance on the COCO-MIG Benchmark

    data/table

    The table below compares the Multi-Instance Generation Controller (MIGC) against text-to-image and layout-to-image baselines on the COCO-MIG benchmark across difficulty levels L2L_2 through L6L_6 (where LiL_i indicates ii target instances), evaluating Instance Success Rate, mean Intersection over Union (mIoU), and average per-image inference latency:

    Method Instance Success Rate (%) ↑\uparrow mIoU ↑\uparrow Time (s) ↓\downarrow
    L2L_2 L3L_3 L4L_4 L5L_5 L6L_6 Avg L2L_2 L3L_3 L4L_4 L5L_5 L6L_6 Avg
    Stable Diffusion 6.87 5.01 3.45 3.27 2.21 3.61 18.92 17.44 15.85 15.17 14.42 15.80 9.18
    TFLCG 20.47 12.71 8.36 6.72 4.36 8.62 29.34 25.06 20.82 18.81 17.86 20.92 19.92
    BOX-Diffusion 24.61 19.22 14.20 11.92 9.31 13.96 32.64 29.88 25.39 23.81 21.19 25.14 44.17
    Multi Diffusion 24.88 22.14 19.88 18.97 18.60 20.12 29.41 28.06 25.59 24.83 24.71 25.89 25.15
    GLIGEN 42.30 35.55 32.66 28.18 30.84 32.39 37.58 32.34 29.95 26.60 27.70 32.25 22.00
    Ours (MIGC) 67.70 59.61 58.09 56.16 56.88 58.43 59.39 52.73 51.45 49.52 49.89 51.48 15.61

    MIGC improves the average Instance Success Rate to 58.43% (a +26.04% gain over GLIGEN's 32.39%) and average mIoU to 51.48 (a +19.23 gain over GLIGEN's 32.25), while running in 15.61 seconds per image.

  8. Knowl 8 — Spatial Accuracy and Image Quality on the COCO-Position Benchmark

    data/table

    The table below evaluates layout control models on the COCO-Position benchmark (800 sampled images, 6,400 generated images) measuring spatial layout precision (Success Ratio, mIoU, AP, AP50, AP75 via Grounding-DINO), image-text consistency (CLIP and Local CLIP scores), and visual quality (FID-6K):

    Method Spatial Accuracy (%) Image Text Consistency Image Quality
    Success Ratio ↑\uparrow mIoU ↑\uparrow AP ↑\uparrow AP50 ↑\uparrow AP75 ↑\uparrow CLIP ↑\uparrow Local CLIP ↑\uparrow FID-6K ↓\downarrow
    Real Image 83.75 85.49 65.97 79.11 71.22 24.22 19.74 -
    Stable Diffusion 5.95 21.60 0.80 2.71 0.42 25.69 17.34 23.56
    TFLCG 13.54 28.01 1.75 6.77 0.56 25.07 17.97 24.65
    BOX-Diffusion 17.84 33.38 3.29 12.27 1.08 23.79 18.70 25.15
    Multi Diffusion 23.86 38.82 6.72 18.65 3.63 22.10 19.13 33.20
    Layout Diffusion 50.53 57.49 23.45 48.10 20.70 18.28 19.08 25.94
    GLIGEN 70.52 71.61 40.68 68.26 42.85 24.61 19.69 26.80
    Ours (MIGC) 80.29 77.38 54.69 84.17 61.71 24.66 20.25 24.52

    MIGC outperforms GLIGEN in Success Ratio (80.29% vs. 70.52%), mIoU (77.38 vs. 71.61), AP (54.69 vs. 40.68), and AP50 (84.17 vs. 68.26) while maintaining an FID of 24.52, comparable to base Stable Diffusion (23.56).

  9. Knowl 9 — Quantitative and Human Evaluation on DrawBench

    data/table

    The table below presents automated evaluation (RR, using Grounding-DINO and Grounded-SAM) and manual human evaluation on the DrawBench benchmark across 64 prompts (25 color/attribute, 19 counting, and 20 spatial position prompts, generating 512 images total):

    Method Spatial (%) ↑\uparrow Attribute (%) ↑\uparrow Count (%) ↑\uparrow
    R Human R Human R Human
    SD1.4 - 13.30 - 57.52 - 23.70
    AAE - 23.13 - 51.50 - 30.92
    Struc-D - 13.12 - 56.50 - 30.26
    Box-D 11.88 50.00 28.50 57.50 9.21 39.47
    TFLCG 9.38 53.13 35.00 60.00 15.79 31.58
    Multi-D 10.63 55.63 18.50 65.50 17.76 36.18
    GLIGEN 61.25 78.80 51.00 48.20 44.08 55.90
    Ours (MIGC) 69.38 93.13 79.00 97.50 67.76 67.50

    MIGC attains the highest success rates in both automated and human metrics across all categories, achieving an attribute success rate of 97.50% (human) / 79.00% (RR), count success rate of 67.50% (human) / 67.76% (RR), and spatial success rate of 93.13% (human) / 69.38% (RR).

  10. Knowl 10 — Ablation Study of MIGC Architectural Components and Loss Functions

    empirical result

    Ablation studies on the COCO-Position dataset evaluate the individual contributions of Enhancement Attention (EA), Layout Attention (LA), the Shading Aggregation Controller (SAC), and background inhibition loss (Lihbt\mathcal{L}_{\text{ihbt}}):

    1. Module Ablation on COCO-Position:
    SAC EA LA R (%) ↑\uparrow mIoU ↑\uparrow AP ↑\uparrow AP50 ↑\uparrow AP75 ↑\uparrow
    7.66 22.71 0.91 3.18 0.35
    ✓ 12.10 29.55 1.89 7.64 0.49
    ✓ 34.70 44.08 11.02 28.64 6.83
    ✓ ✓ 80.16 76.63 53.03 84.05 58.67
    ✓ ✓ 78.12 75.47 52.05 83.48 57.16
    ✓ ✓ ✓ 80.29 77.38 54.69 84.17 61.71
    • Integrating Enhancement Attention (EA) produces the primary gain in spatial accuracy, raising the Success Rate from 12.10% to 80.16% and AP from 1.89 to 53.03 by mitigating instance missing.
    • Layout Attention (LA) and SAC refine boundaries and template alignment, further boosting AP to 54.69 and AP75 to 61.71.
    1. Inhibition Loss Weight λ\lambda:
    • λ=0.0\lambda = 0.0 (no inhibition loss): R=80.20%\text{R} = 80.20\%, mIoU=77.03\text{mIoU} = 77.03, AP=52.46\text{AP} = 52.46, FID=24.73\text{FID} = 24.73.
    • λ=0.1\lambda = 0.1: R=80.29%\text{R} = 80.29\%, mIoU=77.38\text{mIoU} = 77.38, AP=54.69\text{AP} = 54.69, FID=24.52\text{FID} = 24.52.
    • λ=1.0\lambda = 1.0: R=80.61%\text{R} = 80.61\%, mIoU=77.79\text{mIoU} = 77.79, AP=55.62\text{AP} = 55.62, FID=26.94\text{FID} = 26.94. Setting λ=0.1\lambda = 0.1 delivers significant AP gains without degrading image quality (FID).

Coverage note — None. All core methodological contributions, equations, benchmarks, quantitative findings, and ablation studies have been included as self-contained knowls.

References

  1. 1.Omri Avrahami, Thomas Hayes, Oran Gafni, Sonal Gupta, Yaniv Taigman, Devi Parikh, Dani Lischinski, Ohad Fried, and Xi Yin. SpaText: Spatio-textual representation for controllable image generation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2023.
  2. 2.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, Tero Karras, and Ming-Yu Liu. ediff-i: Text-to-image diffusion models with ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022.
  3. 3.Omer Bar-Tal, Lior Yariv, Yaron Lipman, and Tali Dekel. Multidiffusion: Fusing diffusion paths for controlled image generation. arXiv preprint arXiv:2302.08113, 2023.
  4. 4.Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin P. Murphy, William T. Freeman, Michael Rubinstein, Yuanzhen Li, and Dilip Krishnan. Muse: Text-to-image generation via masked generative transformers. 2023.
  5. 5.Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models, 2023.
  6. 6.Minghao Chen, Iro Laina, and Andrea Vedaldi. Training-free layout control with cross-attention guidance. arXiv preprint arXiv:2304.03373, 2023.
  7. 7.Wenhu Chen, Hexiang Hu, Chitwan Saharia, and William W. Cohen. Re-imagen: Retrieval-augmented text-to-image generator, 2022.
  8. 8.Xi Chen, Lianghua Huang, Yu Liu, Yujun Shen, Deli Zhao, and Hengshuang Zhao. Anydoor: Zero-shot object-level image customization. arXiv preprint arXiv:2307.09481, 2023.
  9. 9.Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Diffedit: Diffusion-based semantic image editing with mask guidance. In ICLR 2023 (Eleventh International Conference on Learning Representations), 2023.
  10. 10.Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, and Jie Tang. Cogview: Mastering text-to-image generation via transformers. arXiv preprint arXiv:2105.13290, 2021.
  11. 11.Zheng Ding, Xuaner Zhang, Zhihao Xia, Lars Jebe, Zhuowen Tu, and Xiuming Zhang. Diffusionrig: Learning personalized priors for facial appearance editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12736–12746, 2023.
  12. 12.Dave Epstein, Allan Jabri, Ben Poole, Alexei A. Efros, and Aleksander Holynski. Diffusion self-guidance for controllable image generation, 2023.
  13. 13.Aditya Ramesh et al. Hierarchical text-conditional image generation with clip latents, 2022.
  14. 14.Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun Reddy Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang. Training-free structured diffusion guidance for compositional text-to-image synthesis. In The Eleventh International Conference on Learning Representations, 2023.
  15. 15.Weixi Feng, Wanrong Zhu, Tsu-jui Fu, Varun Jampani, Arjun Akula, Xuehai He, Sugato Basu, Xin Eric Wang, and William Yang Wang. Layoutgpt: Compositional visual planning and generation with large language models. arXiv preprint arXiv:2305.15393, 2023.
  16. 16.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium, 2018.
  17. 17.Jonathan Ho. Classifier-free diffusion guidance. ArXiv, abs/2207.12598, 2022.
  18. 18.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  19. 19.Shuo Huang, Zongxin Yang, Liangting Li, Yi Yang, and Jia Jia. Avatarfusion: Zero-shot generation of clothing-decoupled 3d avatars using 2d diffusion. In Proceedings of the 31st ACM International Conference on Multimedia, pages 5734–5745, 2023.
  20. 20.Wenjing Huang, Shikui Tu, and Lei Xu. Pfb-diff: Progressive feature blending diffusion for text-driven image editing, 2023.
  21. 21.Animesh Karnewar, Andrea Vedaldi, David Novotny, and Niloy J Mitra. Holodiffusion: Training a 3d diffusion model using 2d images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18423–18433, 2023.
  22. 22.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models, 2022.
  23. 23.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models, 2023.
  24. 24.Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization, 2017.
  25. 25.Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment anything. arXiv:2304.02643, 2023.
  26. 26.Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee. Gligen: Open-set grounded text-to-image generation. CVPR, 2023.
  27. 27.Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. Microsoft coco: Common objects in context, 2015.
  28. 28.Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, and Carl Vondrick. Zero-1-to-3: Zero-shot one image to 3d object, 2023.
  29. 29.Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499, 2023.
  30. 30.Shilin Lu, Yanzhu Liu, and Adams Wai-Kin Kong. Tf-icon: Diffusion-based training-free cross-domain image composition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2294–2305, 2023.
  31. 31.Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong. Mace: Mass concept erasure in diffusion models. arXiv preprint arXiv:2403.06135, 2024.
  32. 32.Jiafeng Mao, Xueting Wang, and Kiyoharu Aizawa. Guided image synthesis via initial image editing in diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia. ACM, 2023.
  33. 33.Chong Mou, Xintao Wang, Jiechong Song, Ying Shan, and Jian Zhang. Dragondiffusion: Enabling drag-style manipulation on diffusion models. arXiv preprint arXiv:2307.02421, 2023.
  34. 34.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022.
  35. 35.OpenAI. Gpt-4 technical report, 2023.
  36. 36.Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. Stanza: A Python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, 2020.
  37. 37.Ruijie Quan, Wenguan Wang, Zhibo Tian, Fan Ma, and Yi Yang. Psychometry: An omnifit model for image reconstruction from human brain activity. In CVPR, 2024.
  38. 38.Alec Radford, Jong Wook Kim, Chris Hallacy, A. Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021.
  39. 39.Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee. Generative adversarial text to image synthesis, 2016.
  40. 40.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models, 2021.
  41. 41.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500–22510, 2023.
  42. 42.Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. Palette: Image-to-image diffusion models. In ACM SIGGRAPH 2022 Conference Proceedings, pages 1–10, 2022.
  43. 43.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. 2022.
  44. 44.Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. Self-attention with relative position representations. arXiv preprint arXiv:1803.02155, 2018.
  45. 45.Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung. Instantbooth: Personalized text-to-image generation without test-time finetuning, 2023.
  46. 46.Yujun Shi, Chuhui Xue, Jiachun Pan, Wenqing Zhang, Vincent YF Tan, and Song Bai. Dragdiffusion: Harnessing diffusion models for interactive point-based image editing. arXiv preprint arXiv:2306.14435, 2023.
  47. 47.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  48. 48.Weijia Wu, Yuzhong Zhao, Mike Zheng Shou, Hong Zhou, and Chunhua Shen. Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models. arXiv preprint arXiv:2303.11681, 2023.
  49. 49.Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou. Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion. arXiv preprint arXiv:2307.10816, 2023.
  50. 50.Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks, 2017.
  51. 51.Yuanyou Xu, Zongxin Yang, and Yi Yang. Seeavatar: Photorealistic text-to-3d avatar generation with constrained geometry and appearance. arXiv preprint arXiv:2312.08889, 2023.
  52. 52.Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen. Paint by example: Exemplar-based image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18381–18391, 2023.
  53. 53.Zongxin Yang, Linchao Zhu, Yu Wu, and Yi Yang. Gated channel transformation for visual recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11794–11803, 2020.
  54. 54.Zongxin Yang, Guikun Chen, Xiaodi Li, Wenguan Wang, and Yi Yang. Doraemongpt: Toward understanding dynamic scenes with large language models. arXiv preprint arXiv:2401.08392, 2024.
  55. 55.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, Ben Hutchinson, Wei Han, Zarana Parekh, Xin Li, Han Zhang, Jason Baldridge, and Yonghui Wu. Scaling autoregressive models for content-rich text-to-image generation, 2022.
  56. 56.Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. Cross-modal contrastive learning for text-to-image generation, 2022.
  57. 57.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023.
  58. 58.Zechuan Zhang, Zongxin Yang, and Yi Yang. Sifu: Side-view conditioned implicit function for real-world usable clothed human reconstruction. In CVPR, 2024.
  59. 59.Chen Zhao, Weiling Cai, Chenyu Dong, and Chengwei Hu. Wavelet-based fourier information interaction with frequency diffusion adjustment for underwater image restoration. arXiv preprint arXiv:2311.16845, 2023.
  60. 60.Chen Zhao, Chenyu Dong, and Weiling Cai. Learning a physical-aware diffusion model based on transformer for underwater image enhancement. arXiv preprint arXiv:2403.01497, 2024.
  61. 61.Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li. Layoutdiffusion: Controllable diffusion model for layout-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 22490–22499, 2023.
  62. 62.Dewei Zhou, Zongxin Yang, and Yi Yang. Pyramid diffusion models for low-light image enhancement. In IJCAI, 2023.

Citation

MLA
Zhou, D., et al. “MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis”. arXiv, 2024, http://arxiv.org/abs/2402.05408v2.
APA
Zhou, D., Li, Y., Ma, F., Zhang, X., & Yang, Y. (2024). MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis. arXiv. http://arxiv.org/abs/2402.05408v2
Chicago
Zhou, D., Y. Li, F. Ma, X. Zhang, and Y. Yang. 2024. “MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis”. arXiv. http://arxiv.org/abs/2402.05408v2.
Harvard
Zhou, D. et al. (2024) “MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.05408v2.
Vancouver
1. Zhou D, Li Y, Ma F, Zhang X, Yang Y (2024) MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis. arXiv

BibTeX

@article{zhou2024migc,
  title = {MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis},
  author = {Zhou, Dewei and Li, You and Ma, Fan and Zhang, Xiaoting and Yang, Yi},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.05408v2},
  eprint = {2402.05408}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE