Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing

Bingyan LiuChengyu WangTingfeng CaoKui JiaJun Huang

article2024CVPR186 citations

Reveals why cross-attention replacement causes editing failures in Stable Diffusion through probing analysis and introduces Free-Prompt-Editing, an efficient tuning-free method that outperforms existing approaches by modifying only self-attention maps without requiring a source prompt.

Listen

Deep text-to-image synthesis models, such as Stable Diffusion, have achieved remarkable success in generating realistic visuals, but precise text-guided image editing remains challenging for practical deployment. Existing editing techniques often attempt to manipulate the underlying internal attention mechanisms of these models without a clear understanding of what semantic data these layers actually encode. Consequently, popular methods frequently suffer from editing failures, such as failing to change object colors or structures, or they require heavy computational fine-tuning and original text descriptions that are unavailable when editing authentic real-world photos.

The article evaluates the distinct functional roles of cross-attention and self-attention mechanisms in diffusion models during image editing. Its objective is to explain why current editing techniques fail and to develop a streamlined, tuning-free editing approach that modifies target images accurately while preserving the structure of the original image without requiring source text prompts.

To conduct this evaluation, the researchers applied probing analysis—a technique borrowed from natural language processing that trains shallow classification networks on intermediate model representations—to determine what semantic information is captured within cross-attention and self-attention layers. They evaluated these representations using newly constructed datasets comprising color adjectives and animal categories across generated and real-world image benchmarks. Building on these empirical insights, the authors designed a simplified editing method called Free-Prompt-Editing and evaluated its visual fidelity, prompt alignment, and computational runtime against eight competitive baseline algorithms across multiple benchmark datasets, including Stanford Cars and ImageNet subsets.

The analysis yielded several critical findings. First, probing classifiers achieved very high accuracy (up to 98% for animals and 93% for colors) when classifying cross-attention maps, proving that cross-attention carries rich object attribution and semantic token information rather than merely spatial weights. Replacing cross-attention maps directly transfers unintended source attributes, explaining frequent editing failures in existing tools. Second, self-attention maps contain structural and contour details rather than categorical identities, establishing that self-attention is the primary driver for preserving geometric shapes. Third, restricting self-attention map replacement specifically to intermediate model layers (layers 4 through 14) successfully retains the original spatial structure while allowing the model to adapt to new target text instructions. Finally, the proposed Free-Prompt-Editing framework achieved superior alignment and structure preservation scores while operating significantly faster than leading alternatives, reducing generation time per synthetic image from approximately 335.6 seconds under complex multi-step methods to just 6.3 seconds.

These findings indicate that image-editing pipelines do not need complex cross-attention manipulations or source prompt tracking to achieve high-quality results. Organizations deploying generative image editing can achieve better visual accuracy with lower compute overhead, reducing inference latency and cloud costs by orders of magnitude while enabling direct editing of real photographs. The results challenge prevailing assumptions that cross-attention control is necessary for diffusion-based image modification.

Engineering teams building computer vision and image generation products should adopt intermediate self-attention injection methods and eliminate unnecessary cross-attention replacements. This enables streamlined text-guided image editing workflows that do not depend on original image captions. Future development should focus on testing this approach on emerging generative architectures and integrating enhanced autoencoders to minimize subtle detail loss during the real-image reconstruction stage.

The conclusions are supported by extensive quantitative metrics and visual evaluations across diverse image categories and multiple model variants. However, decision-makers should note that editing success remains constrained by the underlying text-to-image base model's capacity to recognize requested concepts, and fine facial details in real-world images can occasionally degrade during the initial reconstruction step.

Cover for Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing

Abstract

Deep Text-to-Image Synthesis (TIS) models such as Stable Diffusion have recently gained significant popularity for creative text-to-image generation. However, for domain-specific scenarios, tuning-free Text-guided Image Editing (TIE) is of greater importance for application developers. This approach modifies objects or object properties in images by manipulating feature components in attention layers during the generation process. Nevertheless, little is known about the semantic meanings that these attention layers have learned and which parts of the attention maps contribute to the success of image editing. In this paper, we conduct an in-depth probing analysis and demonstrate that cross-attention maps in Stable Diffusion often contain object attribution information, which can result in editing failures. In contrast, self-attention maps play a crucial role in preserving the geometric and shape details of the source image during the transformation to the target image. Our analysis offers valuable insights into understanding cross and self-attention mechanisms in diffusion models. Furthermore, based on our findings, we propose a simplified, yet more stable and efficient, tuning-free procedure that modifies only the self-attention maps of specified attention layers during the denoising process. Experimental results show that our simplified method consistently surpasses the performance of popular approaches on multiple datasets.

Table of Contents

  • 1. Introduction
  • 2. Related Works
  • 2.1. Tuning-free Methods
  • 2.2. Fine-tuning Based Methods
  • 3. Analysis on Cross and Self-Attention
  • 3.1. Cross-Attention in Stable Diffusion
  • 3.2. Self-Attention in Stable Diffusion
  • 3.3. Probing Analysis
  • 3.4. Probing Results on Cross-Attention Maps
  • 3.5. Probing Results on Self-Attention Maps
  • 3.6. Probing Results for Other Tokens
  • 4. Our Approach
  • 5. Experiments
  • 5.1. Experimental Settings
  • 5.2. Image Editing Results
  • 5.2.1 Comparison to Prior/Concurrent Work
  • 5.2.2 Results in Other TIS Models
  • 5.3. Limitations and Discussion
  • 6. Conclusion
  • References

Knowls

  1. Knowl 1 — Free-Prompt-Editing Algorithm for Generated Image Editing

    algorithm

    Free-Prompt-Editing (FPE) for synthetic image editing modifies an image generated from a source text prompt PsrcP_{\text{src}} into an edited image described by a target text prompt PdstP_{\text{dst}} without modifying cross-attention maps. The method injects the self-attention maps MselfM_{\text{self}} extracted from the source image denoising trajectory into the target image denoising process specifically within intermediate attention layers (layers 4 to 14 of the diffusion UNet).

    Input: Source prompt PsrcP_{\text{src}}, target prompt PdstP_{\text{dst}}, random seed SS, total denoising steps TT, latent diffusion denoising model DM\text{DM}, variational autoencoder decoder Decoder\text{Decoder}
    Output: Source image IsrcI_{\text{src}}, edited target image IdstI_{\text{dst}}
    zT∼N(0,I)z_T \sim \mathcal{N}(0, I) sampled with random seed SS
    zT∗←zTz_T^* \leftarrow z_T
    for t=T,T−1,…,1t = T, T - 1, \dots, 1 do
        zt−1,Mself←DM(zt,Psrc,t)z_{t-1}, M_{\text{self}} \leftarrow \text{DM}(z_t, P_{\text{src}}, t)
        zt−1∗←DM(zt∗,Pdst,t) with Mself∗←Mself in layers 4 to 14z_{t-1}^* \leftarrow \text{DM}(z_t^*, P_{\text{dst}}, t) \text{ with } M_{\text{self}}^* \leftarrow M_{\text{self}} \text{ in layers 4 to 14}
    end for
    Isrc←Decoder(z0)I_{\text{src}} \leftarrow \text{Decoder}(z_0)
    Idst←Decoder(z0∗)I_{\text{dst}} \leftarrow \text{Decoder}(z_0^*)
    return Isrc,IdstI_{\text{src}}, I_{\text{dst}}

    In this procedure, ztz_t and zt∗z_t^* represent the latent states at timestep tt for the source and target generation paths, respectively. Replacing Mself∗M_{\text{self}}^* with MselfM_{\text{self}} in layers 4 to 14 preserves the geometric layout and shape details of the source image while enabling the cross-attention mechanism to freely synthesize the semantic concepts specified in PdstP_{\text{dst}}.

  2. Knowl 2 — Free-Prompt-Editing Algorithm for Real Image Editing via DDIM Inversion

    algorithm

    Free-Prompt-Editing (FPE) for real image editing modifies an existing real image IsrcI_{\text{src}} based on a target text prompt PdstP_{\text{dst}} without requiring a source text prompt describing the original image. The latent trajectory is obtained via Denoising Diffusion Implicit Models (DDIM) inversion, and self-attention maps from the reconstruction path are injected into the target path across layers 4 to 14.

    Input: Target prompt PdstP_{\text{dst}}, real input image IsrcI_{\text{src}}, total denoising steps TT, diffusion model DM\text{DM}, DDIM inversion operator DDIM_inv\text{DDIM\_inv}, variational autoencoder decoder Decoder\text{Decoder}
    Output: Reconstructed source image IresI_{\text{res}}, edited target image IdstI_{\text{dst}}
    {zt}t=0T←DDIM_inv(Isrc)\{z_t\}_{t=0}^T \leftarrow \text{DDIM\_inv}(I_{\text{src}})
    zT∗←zTz_T^* \leftarrow z_T
    for t=T,T−1,…,1t = T, T - 1, \dots, 1 do
        zt−1,Mself←DM(zt,t)z_{t-1}, M_{\text{self}} \leftarrow \text{DM}(z_t, t)
        zt−1∗←DM(zt∗,Pdst,t) with Mself∗←Mself in layers 4 to 14z_{t-1}^* \leftarrow \text{DM}(z_t^*, P_{\text{dst}}, t) \text{ with } M_{\text{self}}^* \leftarrow M_{\text{self}} \text{ in layers 4 to 14}
    end for
    Ires←Decoder(z0)I_{\text{res}} \leftarrow \text{Decoder}(z_0)
    Idst←Decoder(z0∗)I_{\text{dst}} \leftarrow \text{Decoder}(z_0^*)
    return Ires,IdstI_{\text{res}}, I_{\text{dst}}

    Here, {zt}t=0T\{z_t\}_{t=0}^T is the deterministic sequence of inverted latents. The reconstruction path DM(zt,t)\text{DM}(z_t, t) executes unconditional diffusion denoising to extract the spatial self-attention matrices MselfM_{\text{self}}. Injecting these matrices into layers 4 to 14 of the editing branch DM(zt∗,Pdst,t)\text{DM}(z_t^*, P_{\text{dst}}, t) retains the real image's spatial structure while altering semantics according to PdstP_{\text{dst}}.

  3. Knowl 3 — Attention Formulations in Latent Diffusion Models

    equation

    In latent text-to-image diffusion models such as Stable Diffusion, image editing operates on cross-attention and self-attention layers within the UNet denoising backbone. Let ztz_t denote the noisy latent at timestep tt, ϕcross(zt)\phi_{\text{cross}}(z_t) and ϕself(zt)\phi_{\text{self}}(z_t) denote intermediate spatial feature maps, and PembP_{\text{emb}} denote the text token embeddings produced by a language encoder from an input text prompt PP.

    Cross-attention maps McrossM_{\text{cross}} and output features ϕ^(zt)\hat{\phi}(z_t) are computed via linear projections ℓq,ℓk,ℓv\ell_q, \ell_k, \ell_v:

    Qcross=ℓq(ϕcross(zt)),Kcross=ℓk(Pemb),Vcross=ℓv(Pemb)Q_{\text{cross}} = \ell_q(\phi_{\text{cross}}(z_t)), \quad K_{\text{cross}} = \ell_k(P_{\text{emb}}), \quad V_{\text{cross}} = \ell_v(P_{\text{emb}})

    Mcross=Softmax(QcrossKcrossTdcross),ϕ^(zt)=McrossVcrossM_{\text{cross}} = \text{Softmax}\left(\frac{Q_{\text{cross}} K_{\text{cross}}^T}{\sqrt{d_{\text{cross}}}}\right), \quad \hat{\phi}(z_t) = M_{\text{cross}} V_{\text{cross}}

    where dcrossd_{\text{cross}} is the projection dimension of the keys and queries, and element MijM_{ij} of McrossM_{\text{cross}} denotes the attention weight assigned to the jj-th prompt token relative to the ii-th spatial image feature.

    Self-attention maps MselfM_{\text{self}} are computed using learned linear projections ℓˉQ\bar{\ell}_Q and ℓˉK\bar{\ell}_K acting on the image spatial features:

    Qself=ℓˉQ(ϕself(zt)),Kself=ℓˉK(ϕself(zt))Q_{\text{self}} = \bar{\ell}_Q(\phi_{\text{self}}(z_t)), \quad K_{\text{self}} = \bar{\ell}_K(\phi_{\text{self}}(z_t))

    Mself=Softmax(QselfKselfTdself)M_{\text{self}} = \text{Softmax}\left(\frac{Q_{\text{self}} K_{\text{self}}^T}{\sqrt{d_{\text{self}}}}\right)

    where dselfd_{\text{self}} is the feature dimension of QselfQ_{\text{self}} and KselfK_{\text{self}}, and MselfM_{\text{self}} quantifies pairwise spatial feature affinities.

  4. Knowl 4 — Category Attribute Encoding in Cross-Attention Maps and Editing Failure

    empirical result

    Probing analysis using a two-layer Multi-Layer Perceptron (MLP) classifier trained directly on cross-attention maps McrossM_{\text{cross}} demonstrates that cross-attention maps do not merely act as spatial weighting masks; they encode dense semantic and category-level representations of their associated text tokens.

    When evaluated on prompts involving animal categories (e.g., "a/an <animal> standing in the park") and color attributes (e.g., "a <color> car"), linear/MLP probing classifiers achieve high classification accuracies on McrossM_{\text{cross}} across UNet layers (reaching average accuracies of 0.980.98 for sheep, 0.970.97 for tiger, 0.950.95 for dog, 0.930.93 for orange, and 0.900.90 for white).

    Because McrossM_{\text{cross}} carries category-specific feature representations from KcrossK_{\text{cross}} and QcrossQ_{\text{cross}}, replacing the cross-attention maps of a target generation with those of a source image (as done in methods like Prompt-to-Prompt) inadvertently re-injects the source category and attribute features. This causes image editing failures, such as incomplete animal conversions (e.g., edited leopards retaining sheep coat features) or failure to alter vehicle colors.

  5. Knowl 5 — Structural Preservation via Intermediate Self-Attention Layers

    empirical result

    Probing classifiers trained on self-attention maps MselfM_{\text{self}} achieve low accuracy for color categories (averaging 0.000.00 to 0.190.19 across UNet layers) and moderate accuracy for animal shapes (0.360.36 to 0.590.59 average accuracy). This indicates that MselfM_{\text{self}} primarily encodes spatial layout, geometric outlines, and shape contours rather than fine-grained category semantics.

    Manipulating self-attention maps across different layers during diffusion denoising reveals distinct behavioral trade-offs:

    1. Replacing MselfM_{\text{self}} across all UNet layers (layers 1 to 16) freezes the source image layout and appearance completely, preventing target prompt modifications.
    2. Leaving MselfM_{\text{self}} entirely unmodified produces unconstrained direct generation aligned with the target prompt, discarding the source image's structure.
    3. Replacing MselfM_{\text{self}} selectively in intermediate layers (layers 4 to 14) preserves the spatial layout and shape contours of the source image while enabling the target prompt to alter object categories, styles, and attributes.
  6. Knowl 6 — Contextual Category Entanglement in Non-Edited Word Cross-Attention Maps

    empirical result

    Probing analysis of cross-attention maps for non-edited words within prompt sentences reveals that transformer-based text encoders propagate semantic attribute features across neighboring token embeddings.

    In prompt structures of the form "a <color> car", probing classifiers trained on the cross-attention map of the article "a" achieve near-zero category accuracy (average 0.000.00 to 0.280.28 across layers). In contrast, classifiers trained on the cross-attention map of the unmodified noun "car" achieve high color classification accuracy (averaging 0.800.80 for orange, 0.540.54 for white, and 0.380.38 for red across layers).

    Because the cross-attention map for an unedited noun carries the contextual attribute (color) of the adjective modifying it, replacing cross-attention maps corresponding to non-edited tokens transfers source attributes into the target generation, leading to editing failures.

  7. Knowl 7 — Layer-Wise Probing Classification Accuracies for Cross- vs Self-Attention

    data/table

    Linear MLP probing classification accuracy evaluated across distinct UNet layers (down, middle, and up blocks: Layers 3, 6, 9, 10, 12, 14, 16) for cross-attention maps (McrossM_{\text{cross}}) and self-attention maps (MselfM_{\text{self}}) across five animal classes and five color classes.

    Cross-Attention Map (McrossM_{\text{cross}}) Self-Attention Map (MselfM_{\text{self}})
    Class L3 L6 L9 L10 L12 L14 L16 Avg. L3 L6 L9 L10 L12 L14 L16 Avg.
    dog 1.00 1.00 1.00 1.00 0.89 0.76 1.00 0.95 0.53 0.60 0.78 0.60 0.53 0.47 0.38 0.55
    horse 0.96 1.00 1.00 1.00 0.64 1.00 0.91 0.93 0.50 0.70 0.82 0.65 0.68 0.53 0.28 0.59
    sheep 0.97 1.00 1.00 1.00 1.00 0.90 0.97 0.98 0.53 0.45 0.25 0.45 0.62 0.53 0.25 0.44
    leopard 0.97 1.00 1.00 1.00 0.97 0.79 0.87 0.94 0.47 0.65 0.57 0.60 0.47 0.65 0.60 0.57
    tiger 1.00 1.00 0.97 1.00 0.88 1.00 0.97 0.97 0.23 0.12 0.55 0.20 0.45 0.42 0.53 0.36
    green 0.93 0.91 0.91 0.96 0.67 0.38 0.60 0.77 0.00 0.00 0.05 0.00 0.05 0.00 0.12 0.03
    white 0.97 1.00 0.94 0.97 0.97 0.61 0.85 0.90 0.00 0.05 0.30 0.55 0.03 0.15 0.25 0.19
    orange 0.97 1.00 0.94 0.92 0.89 0.94 0.83 0.93 0.00 0.00 0.00 0.00 0.00 0.00 0.00 0.00
    yellow 0.96 0.77 1.00 0.98 1.00 0.36 0.68 0.82 0.00 0.42 0.07 0.05 0.00 0.30 0.07 0.13
    red 0.97 0.97 0.93 0.85 0.70 0.23 0.65 0.76 0.00 0.15 0.28 0.20 0.00 0.20 0.10 0.13

    The cross-attention maps maintain high classification accuracy across categories and layers, demonstrating strong category feature representations. In contrast, self-attention maps yield near-zero accuracy on colors and modest accuracy on animal outlines, supporting the strategy of manipulating self-attention rather than cross-attention for text-guided editing.

  8. Knowl 8 — Quantitative Comparison of Text-Guided Image Editing Algorithms

    data/table

    Performance comparison of tuning-free text-guided image editing methods across the ImageNet-R-TI2I and Wild-TI2I benchmarks on both synthetic (fake) and real images using Stable Diffusion 1.5. Metrics include CLIP Score (CS), CLIP Directional Similarity (CDS), and average per-image editing runtime in seconds on an NVIDIA A100 40GB GPU.

    Method ImageNet-R-TI2I fake ImageNet-R-TI2I real Wild-fake Wild-real Editing Times (s)
    CS ↑\uparrow CDS ↑\uparrow CS ↑\uparrow CDS ↑\uparrow CS ↑\uparrow CDS ↑\uparrow CS ↑\uparrow CDS ↑\uparrow fake ↓\downarrow real ↓\downarrow
    SDEdit (0.5) – – 28.37 0.1415 – – 27.48 0.1220 – 2.59
    SDEdit (0.75) – – 30.17 0.2171 – – 29.79 0.2007 – 3.35
    Shape-Guided – – 26.01 0.1090 – – 26.53 0.1330 – 16.02
    DiffEdit 26.68 0.0748 26.50 0.0909 25.59 0.0794 26.33 0.0879 9.02 4.85
    Pix2pix-zero 27.94 0.2271 28.96 0.1415 28.19 0.2864 29.55 0.1462 24.92 36.76
    P2P 28.88 0.3394 28.56 0.2146 27.85 0.2796 28.42 0.1930 6.41 55.32
    PnP 28.83 0.2318 28.76 0.2073 28.20 0.2838 28.46 0.2020 335.65 384.26
    MasaCtrl 29.66 0.3024 31.40 0.2170 29.96 0.3474 29.33 0.2101 6.18 10.90
    Ours (FPE) 29.79 0.3559 29.05 0.2271 27.88 0.3116 29.04 0.2234 6.30 10.75

    FPE achieves the highest CDS scores on both real image benchmarks (0.22710.2271 on ImageNet-R-TI2I real, 0.22340.2234 on Wild-real) and the highest CDS on ImageNet-R-TI2I fake (0.35590.3559). It executes in 6.306.30 seconds for fake editing and 10.7510.75 seconds for real editing, avoiding the computational overhead of feature-injection methods like PnP (335.65s/384.26s335.65\text{s} / 384.26\text{s}) and null-text optimization in P2P (55.32s55.32\text{s}).

  9. Knowl 9 — Performance Comparison of Free-Prompt-Editing vs. Prompt-to-Prompt on Dedicated Edit Datasets

    data/table

    Quantitative comparison between Prompt-to-Prompt (P2P) and Free-Prompt-Editing (FPE) across four constructed datasets: Car-fake-edit (756 prompt pairs), Car-real-edit (3,321 pairs from CARS196), ImageNet-fake-edit (1,182 pairs from FlexIT and ImageNet), and ImageNet-real-edit (1,092 pairs). Evaluated on Stable Diffusion 1.5 using CLIP Score (CS) and CLIP Directional Similarity (CDS).

    CLIP Score (CS) ↑\uparrow CLIP Directional Similarity (CDS) ↑\uparrow
    Dataset P2P Ours (FPE) P2P Ours (FPE)
    Car-fake-edit 25.96 26.02 0.2451 0.2659
    Car-real-edit 24.64 24.85 0.2288 0.2605
    ImageNet-fake-edit 27.42 27.80 0.2401 0.2560
    ImageNet-real-edit 26.17 26.35 0.2426 0.2468

    FPE consistently improves over P2P across both synthetic and real image domains, showing notable gains in CDS (e.g., 0.26050.2605 vs. 0.22880.2288 on Car-real-edit and 0.26590.2659 vs. 0.24510.2451 on Car-fake-edit), reflecting better adherence to target edits while maintaining source image structures.

  10. Knowl 10 — Limitations of Free-Prompt-Editing

    limitation

    Free-Prompt-Editing (FPE) is subject to two main constraints:

    1. Generative Capacity of Base Model: FPE relies entirely on the pre-trained text-to-image synthesis backbone (e.g., Stable Diffusion). If the base model cannot generate the specific concept or attribute described in the target prompt PdstP_{\text{dst}}, the editing operation fails.
    2. Autoencoder Reconstruction Fidelity: Real image editing requires reconstructing the original image latent through DDIM inversion. High-frequency spatial details, particularly fine facial features, can be degraded or lost during autoencoding and latent inversion due to the reconstruction constraints of the pre-trained vector-quantized (VQ) autoencoder.

Coverage note — Qualitative visual comparison figures (Figures 1, 3, 4, 6, 7, 8, 9) demonstrating specific image editing outputs across models were omitted as visual examples whose underlying algorithmic methods and quantitative findings are fully captured in the knowls.

References

  1. 1.Tim Brooks, Aleksander Holynski, and Alexei A Efros. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18392–18402, 2023. 1, 2, 3, 6, 7, 8
  2. 2.Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22560–22570, 2023. 1, 2, 6, 7, 8
  3. 3.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286. Association for Computational Linguistics, 2019. 2, 3
  4. 4.Guillaume Couairon, Asya Grechka, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Flexit: Towards flexible semantic image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 18270–18279, 2022. 6, 2
  5. 5.Guillaume Couairon, Jakob Verbeek, Holger Schwenk, and Matthieu Cord. Diffedit: Diffusion-based semantic image editing with mask guidance. In The Eleventh International Conference on Learning Representations, 2023. 1, 2, 6, 7, 8
  6. 6.Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. 1
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers), pages 4171–4186. Association for Computational Linguistics, 2019. 5
  8. 8.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022. 2
  9. 9.Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. Stylegan-nada: Clipguided domain adaptation of image generators. ACM Transactions on Graphics (TOG), 41(4):1–13, 2022. 6, 8
  10. 10.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. Prompt-to-prompt image editing with cross-attention control. In The Eleventh International Conference on Learning Representations, 2023. 1, 2, 5, 6, 7, 8
  11. 11.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020. 1
  12. 12.Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685, 2021. 2
  13. 13.Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani. Imagic: Text-based real image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6007–6017, 2023. 2
  14. 14.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In International Conference on Learning Representations, 2014. 1, 8
  15. 15.Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In Proceedings of the IEEE international conference on computer vision workshops, pages 554–561, 2013. 6, 1
  16. 16.Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. Linguistic knowledge and transferability of contextual representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1073–1094. Association for Computational Linguistics, 2019. 2, 3
  17. 17.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022. 1, 2, 6, 7, 8
  18. 18.Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Null-text inversion for editing real images using guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6038–6047, 2023. 1, 2, 6, 4, 5
  19. 19.Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, Ying Shan, and Xiaohu Qie. T2iadapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. arXiv preprint arXiv:2302.08453, 2023. 2
  20. 20.OpenAI. Improving image generation with better captions. https://cdn.openai.com/papers/dall-e-3.pdf, 2023. 1
  21. 21.Dong Huk Park, Grace Luo, Clayton Toste, Samaneh Azadi, Xihui Liu, Maka Karalashvili, Anna Rohrbach, and Trevor Darrell. Shape-guided diffusion with inside-outside attention. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 4198–4207, 2024. 1, 6, 7, 8
  22. 22.Gaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li, Jingwan Lu, and Jun-Yan Zhu. Zero-shot image-to-image translation. In ACM SIGGRAPH 2023 Conference Proceedings, pages 1–11, 2023. 1, 6, 7, 8
  23. 23.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PMLR, 2021. 1, 5, 6, 8
  24. 24.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020. 1
  25. 25.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 1
  26. 26.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. 1, 5
  27. 27.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22500–22510, 2023. 2
  28. 28.Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115 (3):211–252, 2015. 6, 2
  29. 29.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494, 2022. 1
  30. 30.Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. 1
  31. 31.Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in Neural Information Processing Systems, 35:25278–25294, 2022. 1
  32. 32.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015. 1
  33. 33.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021. 5
  34. 34.Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-toimage translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1921–1930, 2023. 1, 2, 6, 7, 8
  35. 35.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 5
  36. 36.Michael E Wall, Andreas Rechtsteiner, and Luis M Rocha. Singular value decomposition and principal component analysis. In A practical approach to microarray data analysis, pages 91–109. Springer, 2003. 4
  37. 37.Chengyu Wang, Zhongjie Duan, Bingyan Liu, Xinyi Zou, Cen Chen, Kui Jia, and Jun Huang. Pai-diffusion: Constructing and serving a family of open chinese diffusion models for text-to-image synthesis on the cloud. arXiv preprint arXiv:2309.05534, 2023. 1
  38. 38.Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and MingHsuan Yang. Diffusion models: A comprehensive survey of methods and applications. ACM Computing Surveys, 2022. 1
  39. 39.Fangneng Zhan, Yingchen Yu, Rongliang Wu, Jiahui Zhang, Shijian Lu, Lingjie Liu, Adam Kortylewski, Christian Theobalt, and Eric Xing. Multimodal image synthesis and editing: A survey and taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023. 2
  40. 40.Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 2

Citation

MLA
Liu, B., et al. “Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing”. arXiv, 2024, http://arxiv.org/abs/2403.03431v1.
APA
Liu, B., Wang, C., Cao, T., Jia, K., & Huang, J. (2024). Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing. arXiv. http://arxiv.org/abs/2403.03431v1
Chicago
Liu, B., C. Wang, T. Cao, K. Jia, and J. Huang. 2024. “Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing”. arXiv. http://arxiv.org/abs/2403.03431v1.
Harvard
Liu, B. et al. (2024) “Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2403.03431v1.
Vancouver
1. Liu B, Wang C, Cao T, Jia K, Huang J (2024) Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing. arXiv

BibTeX

@article{liu2024towards,
  title = {Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing},
  author = {Liu, Bingyan and Wang, Chengyu and Cao, Tingfeng and Jia, Kui and Huang, Jun},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2403.03431v1},
  eprint = {2403.03431}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE