Directed Diffusion: Direct Control of Object Placement through Attention Guidance

Wan-Duo Kurt MaAvisek LahiriJohn P. LewisThomas LeungW. Bastiaan Kleijn

article2024AAAI91 citations

Introduces a training-free technique for text-to-image diffusion models that directly guides cross-attention maps during early denoising steps to accurately place multiple specified objects within user-defined bounding boxes.

Listen

Modern text-to-image artificial intelligence systems can generate high-quality visual content from descriptive prompts, but they routinely struggle to compose complex scenes involving multiple objects in precise arrangements. Because simple text descriptions cannot reliably specify spatial layouts, creators face tedious and costly trial-and-error cycles. This limitation presents a major barrier for practical storytelling applications, such as illustrated books and storyboards, where characters and objects must maintain intentional positional relationships.

The article introduces and evaluates Directed Diffusion, a lightweight method designed to give creators intuitive, high-level control over where multiple objects appear in synthesized images. The technique works on existing, pre-trained image generation systems without requiring expensive model retraining or fine-tuning.

The researchers developed an approach that operates during the early stages of the image generation process, when general layouts are established. By adjusting internal cross-attention maps—the representations that connect specific words in a prompt to corresponding image regions—the system steers object generation toward user-defined bounding boxes. To evaluate the method, the authors conducted comparative experiments against standard Stable Diffusion and existing state-of-the-art layout control methods, measuring prompt fidelity using standardized text-image similarity scores across single-object and complex multi-object scenes.

The evaluation revealed several key findings. First, Directed Diffusion successfully places specified objects within user-defined bounding boxes while producing natural contextual interactions, such as realistic shadows, lighting, and occlusions with the background. Second, the method achieved quantitative prompt-alignment scores (CLIP similarity) on par with or exceeding competing approaches, reaching 0.824 in complex scene composition compared to 0.821 for baseline Stable Diffusion and 0.791 for Composable Diffusion. Third, the authors demonstrated a companion technique, Placement Finetuning, which enables users to reposition an object within an existing scene while preserving the object's identity and background consistency without retraining.

These findings demonstrate that high-level spatial control can be achieved with minimal computational overhead, requiring only a few lines of code added to standard image generation pipelines. For organizations and creative teams, this significantly reduces the time and computing costs associated with repeated prompt engineering and eliminates the heavy infrastructure requirements of training specialized models from scratch.

The authors recommend integrating Directed Diffusion into creative production workflows for static visual storytelling formats, such as comic books and illustrated literature. Practitioners should pair this positional framework with complementary identity-preservation tools to maintain consistent character appearances across multiple narrative scenes.

The method inherits certain baseline limitations from underlying diffusion models, including occasional generation failures that still require seed exploration, as well as grammatical parsing limitations in language understanding models. Confidence in the reported results is high for static, multi-object image composition, though the authors caution that substantial technical advances remain necessary before these capabilities can be extended to video production.

arXiv: 2302.13153
Cover for Directed Diffusion: Direct Control of Object Placement through Attention Guidance

Abstract

Text-guided diffusion models such as DALL-E 2, Imagen, eDiff-I, and Stable Diffusion are able to generate an effectively endless variety of images given only a short text prompt describing the desired image content. In many cases the images are of very high quality. However, these models often struggle to compose scenes containing several key objects such as characters in specified positional relationships. The missing capability to “direct” the placement of characters and objects both within and across images is crucial in storytelling, as recognized in the literature on film and animation theory. In this work, we take a particularly straightforward approach to provide the needed direction. Drawing on the observation that the cross-attention maps for prompt words reflect the spatial layout of objects denoted by those words, we introduce an optimization objective that produces “activation” at desired positions in these cross-attention maps. The resulting approach is a step toward generalizing the applicability of text-guided diffusion models beyond single images to collections of related images, as in storybooks. Directed Diffusion provides easy high-level positional control over multiple objects, while making use of an existing pre-trained model and maintaining a coherent blend between the positioned objects and the background. Moreover, it requires only a few lines to implement.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Pipeline
  • 3.2 Cross-Attention Map Guidance
  • 4 Applications
  • 4.1 Scene compositing
  • 4.2 Placement Finetuning
  • 5 Experiments and Comparisons
  • 5.1 Comparison: Scene compositing
  • 5.2 Comparison: One Object
  • 5.3 Placement finetuning
  • 5.4 Comparison: Two Objects
  • 5.5 Comparison: object interactions from prompt verbs
  • 6 Limitations and Conclusion
  • 7 Acknowledgements
  • References
  • 8 Implementation
  • 9 Analysis and Ablation
  • 10 Placement finetuning
  • 11 Scene compositing
  • 12 Qualitative Evaluations
  • 13 Quantitative Evaluation and Limitations
  • 14 Societal Impact
  • 14.1 Scene compositing: A white church under lightning in the pink sky
  • 14.2 Scene compositing: Mystical trees next to a dark magical pond
  • 14.3 Scene compositing: A castle in a forest with grainy fog
  • 14.4 Scene compositing: cherry blossoms next to the lake and a mountain
  • 14.5 One Object: bonfire next to a man
  • 14.6 One Object: A white horse in front of the erupting volcano
  • 14.7 One Object: A gravestone next to a man in red shirt
  • 14.8 Two Objects: A white dog chasing a red sphere
  • 14.9 Two Objects: A red cube above the blue sphere in the supermarket
  • 14.10 Two Objects: MultiDiffusion
  • 14.11 Intersection over Union

Knowls

  1. Knowl 1 — Directed Diffusion Framework and Two-Stage Denoising Pipeline

    model/method

    Directed Diffusion (DD) provides spatial control over the placement of multiple objects in text-to-image diffusion models without requiring model retraining, fine-tuning, or per-word weight parameter tuning. The model takes a text prompt P\mathcal{P} and spatial region information R={B,I}\mathcal{R} = \{\mathcal{B}, \mathcal{I}\}, where B\mathcal{B} specifies bounding boxes for desired object positions and I⊂{1,…,∣P∣}\mathcal{I} \subset \{1, \dots, |\mathcal{P}|\} denotes the token indices of corresponding directed words.

    The reverse diffusion process spanning timesteps t∈{T,…,0}t \in \{T, \dots, 0\} is partitioned into two distinct phases:

    1. Attention Editing Stage (t∈[T,T−N)t \in [T, T-N)): Modifies the cross-attention maps during the initial NN denoising steps (typically N=10N=10). In this stage, neural activation within the bounding boxes B\mathcal{B} is amplified while surrounding activations are suppressed via optimization, establishing the coarse spatial layout and rough shape of directed objects early in synthesis.

    2. Conventional Denoising Stage (t∈[T−N,0]t \in [T-N, 0]): Resumes standard Latent Diffusion Model denoising with classifier-free guidance without further attention manipulation. This stage refines object details, textures, and contextual physical interactions (such as shadows and illumination) to seamlessly integrate placed objects into the background.

  2. Knowl 2 — Target Cross-Attention Map Construction with Spatial Masks

    model/method

    In Directed Diffusion, target cross-attention maps D(i)\mathbf{D}^{(i)} are constructed for each directed token index ii by modifying the existing cross-attention slice A(i)\mathbf{A}^{(i)} using a weaken-mask W(B′)\mathbf{W}(\mathcal{B}') and a strengthen-mask S(B)\mathbf{S}(\mathcal{B}).

    Given a bounding box B\mathcal{B} with normalized coordinates b=(bleft,bright,btop,bbottom)∈[0,1]4b = (b_{\text{left}}, b_{\text{right}}, b_{\text{top}}, b_{\text{bottom}}) \in [0, 1]^4 on an attention map of spatial resolution w×hw \times h, the masks are defined coordinate-wise as:

    W(B′)xy={c,(x,y)∈B′1,otherwise\mathbf{W}(\mathcal{B}')_{xy} = \begin{cases} c, & (x, y) \in \mathcal{B}' \\ 1, & \text{otherwise} \end{cases}

    S(B)xy={f(x,y),(x,y)∈B0,otherwise\mathbf{S}(\mathcal{B})_{xy} = \begin{cases} f(x, y), & (x, y) \in \mathcal{B} \\ 0, & \text{otherwise} \end{cases}

    where B′\mathcal{B}' is the complement of region B\mathcal{B}, c<1c < 1 is an attenuation scalar that suppresses attention outside the bounding box, and f(x,y)f(x, y) is a 2D Gaussian activation window centered in B\mathcal{B} with standard deviations σx=bw/2\sigma_x = b_w / 2 and σy=bh/2\sigma_y = b_h / 2. Here, bw=⌈(bright−bleft)×w⌉b_w = \lceil (b_{\text{right}} - b_{\text{left}}) \times w \rceil and bh=⌈(btop−bbottom)×h⌉b_h = \lceil (b_{\text{top}} - b_{\text{bottom}}) \times h \rceil are the width and height of B\mathcal{B}.

    The target attention map is then formed by:

    D(i)=A(i)⊙W(B′)+S(B)\mathbf{D}^{(i)} = \mathbf{A}^{(i)} \odot \mathbf{W}(\mathcal{B}') + \mathbf{S}(\mathcal{B})

    where ⊙\odot represents element-wise multiplication.

  3. Knowl 3 — Cross-Attention Optimization via Trailing Attention Maps

    model/method

    In Stable Diffusion, the CLIP text encoder produces fixed-length token representations ∣W∣=77|W| = 77. For a prompt of length ∣P∣|\mathcal{P}|, the cross-attention tensor contains prompt maps A(i)\mathbf{A}^{(i)} for i≤∣P∣i \le |\mathcal{P}| and trailing attention maps A(i)\mathbf{A}^{(i)} for indices i∈T={∣P∣+1,…,∣W∣}i \in \mathcal{T} = \{|\mathcal{P}|+1, \dots, |W|\} corresponding to unused padding tokens.

    To avoid instabilities and out-of-distribution latents caused by direct attention injection, Directed Diffusion optimizes a weight vector at∈R∣W∣−∣P∣\mathbf{a}_t \in \mathbb{R}^{|W|-|\mathcal{P}|} over the trailing attention maps at timestep tt to steer the directed prompt cross-attention maps At−1(i)\mathbf{A}_{t-1}^{(i)} at step t−1t-1 toward target maps D(i)\mathbf{D}^{(i)}:

    a∗=arg⁡min⁡at∑i∈I∥At−1(i)(At(∣P∣+1:∣W∣)⊙at)−D(i)∥22\mathbf{a}^* = \arg\min_{\mathbf{a}_t} \sum_{i \in \mathcal{I}} \left\| \mathbf{A}_{t-1}^{(i)}\left(\mathbf{A}_t^{(|\mathcal{P}|+1:|W|)} \odot \mathbf{a}_t\right) - \mathbf{D}^{(i)} \right\|_2^2

    where I\mathcal{I} is the set of directed prompt word indices, ⊙\odot denotes element-wise scaling of each trailing map by its corresponding weight in at\mathbf{a}_t, and At−1(i)(⋅)\mathbf{A}_{t-1}^{(i)}(\cdot) denotes the indirect functional dependency of the attention map at time t−1t-1 on the re-weighted trailing maps applied during the denoising step at time tt. The optimized weights a∗\mathbf{a}^* reweight the trailing maps to influence U-Net conditioning at step tt.

  4. Knowl 4 — Directed Diffusion Cross-Attention Editing Algorithm

    algorithm

    The cross-attention editing algorithm modifies the forward pass through the cross-attention layers of the denoising U-Net during the attention editing stage:

    procedure DDCrossAttnEdit(DM, P, B)
        Input: Denoising model DM with latent z_t, prompt P, and bounding box specification B
        Output: Updated latent z_t
        for each layer l in layer(DM(z_t, P)) do
            if type(l) in CrossAttn then
                Compute cross-attention A = Softmax(Q_l(z_t) * K_l(P)^T / sqrt(d))
                for each token index i in T union I do
                    Compute target D^(i) = A^(i) * W(B') + S(B)
                end for
                Optimize weight vector a* = argmin_a L_a
                Update trailing attention slices A'^(|P|+1:77) = D^(|P|+1:77) * a*
                Update latent z_t = l(z_t, A' * V_l(P))
            else
                Update latent z_t = l(z_t)
            end if
        end for
        return z_t

    Here, Ql,Kl,VlQ_l, K_l, V_l denote the query, key, and value projection layers in cross-attention layer ll, I\mathcal{I} denotes the directed prompt token indices, T={∣P∣+1,…,77}\mathcal{T} = \{|\mathcal{P}|+1, \dots, 77\} denotes trailing attention map indices, W(B′)W(\mathcal{B}') and S(B)S(\mathcal{B}) are weakening and strengthening masks, and LaL_a is the trailing map optimization loss.

  5. Knowl 5 — Multi-Object Scene Compositing via Latent Interpolation

    model/method

    When directing more than two bounding boxes simultaneously, direct multi-box attention editing can become unreliable. Directed Diffusion resolves this by independently generating latents for each directed object and performing linear spatial interpolation during the initial NN denoising steps.

    For RR directed objects with bounding boxes Br\mathcal{B}_r (r∈{1,…,R}r \in \{1, \dots, R\}), individual denoising paths record the intermediate latents zt(r)\mathbf{z}_t^{(r)}. The combined scene latent zt\mathbf{z}_t conditioned on the global prompt P\mathcal{P} is updated for spatial coordinates (x,y)∈Br(x, y) \in \mathcal{B}_r during timesteps t∈[T,T−N]t \in [T, T-N] according to:

    zt(x,y):=1R∑r=1R[wrzt(x,y)+(1−wr)zt(r)(x,y)]\mathbf{z}_t(x, y) := \frac{1}{R} \sum_{r=1}^R \left[ w_r \mathbf{z}_t(x, y) + (1 - w_r) \mathbf{z}_t^{(r)}(x, y) \right]

    where wr∈Rw_r \in \mathbb{R} is a weighting scalar (typically set to 0.10.1) and N≈10N \approx 10. Subsequent denoising steps (t<T−Nt < T-N) follow standard diffusion without interpolation, allowing the model to produce unified lighting and contextual interactions across all objects.

  6. Knowl 6 — Placement Finetuning for Identity-Preserving Object Repositioning

    model/method

    Placement Finetuning (PF) reposition an object within an already generated image while preserving its synthesized identity and visual characteristics without model retraining.

    The algorithm comprises three components:

    1. Segmentation and Spatial Transformation: An object mask Mo\mathbf{M}_o is extracted by thresholding the final t=0t=0 cross-attention map and clipping to bounding box B\mathcal{B}. At step T−NT-N, a spatial translation transformation X(⋅)\mathbf{X}(\cdot) (implemented via cyclic tensor shifting) is applied to the stored latent zT−N\mathbf{z}_{T-N} and mask Mo\mathbf{M}_o, producing translated mask X(Mo)\mathbf{X}(\mathbf{M}_o) and translated latent X(zT−N)\mathbf{X}(\mathbf{z}_{T-N}).

    2. Background Inpainting: The hole left by moving the foreground object is initialized using the transformed latent X(zT)\mathbf{X}(\mathbf{z}_T) masked by Mo\mathbf{M}_o against the background complement ¬Mo\neg \mathbf{M}_o, followed by a single noise-and-denoise step to synthesize background texture.

    3. Latent Blending: For remaining timesteps t∈{T−N−1,…,0}t \in \{T-N-1, \dots, 0\}, the latent is composited at each step via:

    zt′:=zt′⊙¬X(Mo)+X(zt)⊙X(Mo)\mathbf{z}'_t := \mathbf{z}'_t \odot \neg \mathbf{X}(\mathbf{M}_o) + \mathbf{X}(\mathbf{z}_t) \odot \mathbf{X}(\mathbf{M}_o)

    followed by standard denoising. Setting N≈10N \approx 10 provides a trade-off where larger NN preserves foreground and background identity while smaller NN increases contextual interaction.

  7. Knowl 7 — Success-at-k (SS@k) Evaluation Protocol

    experimental setup

    To model realistic generative text-to-image usage where users iterate over random seeds rather than relying on a single output, the Success-at-kk (SS@kk) protocol generates kk candidate images using consecutive random seeds beginning from an initial seed. The highest-quality image among the kk candidates is subjectively selected for qualitative and quantitative evaluation. In the evaluations of Directed Diffusion, k=12k=12 (SS@12) is used across all comparative benchmarks.

  8. Knowl 8 — Quantitative Evaluation of Text-Image Alignment Across Placement Benchmarks

    data/table

    Text-image alignment was quantitatively assessed using CLIP similarity scores between prompt embeddings and generated images under the SS@12 protocol across three categories: Scene Composition (SceneComp), single-object placement (OneMask), and two-object placement (TwoMasks).

    Eval Type SD GLIGEN CD DD
    SceneComp 0.821 0.802 0.791 0.824
    OneMask 0.834 0.810 - 0.807
    TwoMasks 0.842 0.802 - 0.851

    Directed Diffusion (DD) outperforms GLIGEN and Composable Diffusion (CD) on multi-object benchmarks (achieving 0.824 on SceneComp and 0.851 on TwoMasks), while remaining competitive with standard Stable Diffusion (SD) and GLIGEN on single-mask generation (0.807 vs 0.834 for SD and 0.810 for GLIGEN).

  9. Knowl 9 — Limitations of Directed Diffusion

    limitation

    Directed Diffusion exhibits several limitations:

    1. Dependency on Base Diffusion Model Robustness: It inherits baseline Stable Diffusion failure modes, including seed sensitivity where particular seeds yield distorted object anatomy, missing limbs, or omitted prompt subjects.

    2. Sensitivity to the Editing Step Hyperparameter: The number of active editing steps NN must be manually specified. Excessively large NN suppresses natural contextual interaction (e.g., lighting and shadows) with the scene, while overly small NN fails to enforce positional placement.

    3. Scalability with Simultaneous Bounding Boxes: Direct attention guidance becomes unreliable when guiding more than two bounding boxes simultaneously without employing multi-pass latent interpolation.

    4. Text Encoder Semantic and Grammatical Constraints: Due to CLIP's limited understanding of complex grammar, prepositional bindings, and transitive verbs, specifying action relationships between directed objects (e.g., distinguishing 'a dog chasing a ball' from 'a dog and a ball') remains inconsistent despite accurate spatial positioning.

Coverage note — Visual qualitative figure demonstrations from the paper were omitted as standalone knowls because their core conceptual and empirical claims are fully captured in the methodology, experimental setup, evaluation table, and limitation knowls.

References

  1. 1.Arijon, D. 1976. Grammar of the Film Language. Focal Press.
  2. 2.Avrahami, O.; Fried, O.; and Lischinski, D. 2022. Blended Latent Diffusion. CoRR, abs/2206.02779.
  3. 3.Avrahami, O.; Hayes, T.; Gafni, O.; Gupta, S.; Taigman, Y.; Parikh, D.; Lischinski, D.; Fried, O.; and Yin, X. 2022. SpaText: Spatio-Textual Representation for Controllable Image Generation. CoRR, abs/2211.14305.
  4. 4.Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; Catanzaro, B.; Karras, T.; and Liu, M. 2022. eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers. CoRR, abs/2211.01324.
  5. 5.Bar-Tal, O.; Yariv, L.; Lipman, Y.; and Dekel, T. 2023. MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation. CoRR, abs/2302.08113.
  6. 6.Chang, H.; Zhang, H.; Barber, J.; Maschinot, A.; Lezama, J.; Jiang, L.; Yang, M.; Murphy, K.; Freeman, W. T.; Rubinstein, M.; Li, Y.; and Krishnan, D. 2023. Muse: Text-To-Image Generation via Masked Generative Transformers. CoRR, abs/2301.00704.
  7. 7.Feng, W.; He, X.; Fu, T.; Jampani, V.; Akula, A. R.; Narayana, P.; Basu, S.; Wang, X. E.; and Wang, W. Y. 2022. Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis. CoRR, abs/2212.05032.
  8. 8.Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. CoRR, abs/2208.01618.
  9. 9.Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In EMNLP.
  10. 10.Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33.
  11. 11.Ho, J.; and Salimans, T. 2021. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications.
  12. 12.Hyv"arinen, A. 2005. Estimation of Non-Normalized Statistical Models by Score Matching. J. Mach. Learning Research, 6: 695–709.
  13. 13.Kingma, D. P.; and Welling, M. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114.
  14. 14.Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; and Lee, Y. J. 2023. GLIGEN: Open-Set Grounded Text-to-Image Generation. CoRR, abs/2301.07093.
  15. 15.Liew, J. H.; Yan, H.; Zhou, D.; and Feng, J. 2022. MagicMix: Semantic Mixing with Diffusion Models. CoRR, abs/2210.16056.
  16. 16.Liu, N.; Li, S.; Du, Y.; Torralba, A.; and Tenenbaum, J. B. 2022. Compositional Visual Generation with Composable Diffusion Models. In ECCV.
  17. 17.Ma, W.-D. K.; Lewis, J. P.; Lahiri, A.; Leung, T.; and Kleijn, W. B. 2023. Directed Diffusion: Direct Control of Object Placement through Attention Guidance. arXiv:2302.13153.
  18. 18.Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.; and Ermon, S. 2022. SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations. In Int. Conf. on Learning Representations (ICLR).
  19. 19.Nichol, A. Q.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2022. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models. In ICML.
  20. 20.Park, D. H.; Luo, G.; Toste, C.; Azadi, S.; Liu, X.; Karalashvili, M.; Rohrbach, A.; and Darrell, T. 2022. Shape-Guided Diffusion with Inside-Outside Attention.
  21. 21.Patashnik, O.; Wu, Z.; Shechtman, E.; Cohen-Or, D.; and Lischinski, D. 2021. Styleclip: Text-driven manipulation of stylegan imagery. In ICCV.
  22. 22.Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Proc. ICML.
  23. 23.Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022. Hierarchical Text-Conditional Image Generation with CLIP Latents. CoRR, abs/2204.06125.
  24. 24.Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022. High-Resolution Image Synthesis With Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 10684–10695.
  25. 25.Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2022. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. CoRR, abs/2208.12242.
  26. 26.Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.; Ghasemipour, S. K. S.; Ayan, B. K.; Mahdavi, S. S.; Lopes, R. G.; Salimans, T.; Ho, J.; Fleet, D. J.; and Norouzi, M. 2022. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. CoRR, abs/2205.11487.
  27. 27.Smith, E. 2022. A Traveler’s Guide to the Latent Space.
  28. 28.Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning.
  29. 29.Song, J.; Meng, C.; and Ermon, S. 2021. Denoising Diffusion Implicit Models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021.
  30. 30.Song, Y.; and Ermon, S. 2019. Generative modeling by estimating gradients of the data distribution. In NeurIPS, volume 32.
  31. 31.StabilityAI. 2023. Stable Diffusion 2.1 Demo. https://huggingface.co/spaces/stabilityai/stable-diffusion. Accessed: 2024-01-01.
  32. 32.Thomas, F.; and Johnston, O. 1981. Disney animation : the illusion of life. Abbeville Press New York, 1st ed. edition. ISBN 0896592332 0896592324.
  33. 33.Weng, L. 2021. What are diffusion models?
  34. 34.Xie, J.; Li, Y.; Huang, Y.; Liu, H.; Zhang, W.; Zheng, Y.; and Shou, M. Z. 2023. BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion. CoRR, abs/2307.10816.
  35. 35.Yu, J.; Xu, Y.; Koh, J. Y.; Luong, T.; Baid, G.; Wang, Z.; Vasudevan, V.; Ku, A.; Yang, Y.; Ayan, B. K.; Hutchinson, B.; Han, W.; Parekh, Z.; Li, X.; Zhang, H.; Baldridge, J.; and Wu, Y. 2022. Parti: Pathways Autoregressive Text-to-Image model.
  36. 36.Zhang, L.; and Agrawala, M. 2023. Adding Conditional Control to Text-to-Image Diffusion Models. arXiv:2302.05543.

Citation

MLA
Ma, W.-D. K., et al. “Directed Diffusion: Direct Control of Object Placement Through Attention Guidance”. arXiv, 2023, http://arxiv.org/abs/2302.13153v3.
APA
Ma, W.-D. K., Lewis, J. P., Lahiri, A., Leung, T., & Kleijn, W. B. (2023). Directed Diffusion: Direct Control of Object Placement through Attention Guidance. arXiv. http://arxiv.org/abs/2302.13153v3
Chicago
Ma, W.-D. K., J. P. Lewis, A. Lahiri, T. Leung, and W. B. Kleijn. 2023. “Directed Diffusion: Direct Control of Object Placement Through Attention Guidance”. arXiv. http://arxiv.org/abs/2302.13153v3.
Harvard
Ma, W.-D.K. et al. (2023) “Directed Diffusion: Direct Control of Object Placement through Attention Guidance”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2302.13153v3.
Vancouver
1. Ma W-DK, Lewis JP, Lahiri A, Leung T, Kleijn WB (2023) Directed Diffusion: Direct Control of Object Placement through Attention Guidance. arXiv

BibTeX

@article{ma2023directed,
  title = {Directed Diffusion: Direct Control of Object Placement through Attention Guidance},
  author = {Ma, Wan-Duo Kurt and Lewis, J. P. and Lahiri, Avisek and Leung, Thomas and Kleijn, W. Bastiaan},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2302.13153v3},
  eprint = {2302.13153}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF