What the DAAM: Interpreting Stable Diffusion Using Cross Attention

Raphael TangLinqing LiuAkshat PandeyZhiying JiangGefei YangKarun KumarPontus StenetorpJimmy LinFerhan Ture

article2023ACL259 citationsBest Paper Award

Introduces DAAM, a method that aggregates cross-attention maps in Stable Diffusion to interpret how individual prompt words influence generated pixels and uncover visual-linguistic failure modes such as feature entanglement.

Listen

Modern text-to-image generative systems, such as diffusion models, can synthesize high-fidelity and photorealistic images from textual prompts. Despite their widespread adoption and commercial relevance, the internal mechanics driving how specific words shape visual outputs remain poorly understood. Because these models function as black boxes, practitioners face challenges diagnosing generation errors, controlling layout precision, and understanding semantic failures.

The article demonstrates a novel interpretability method called Diffusion Attentive Attribution Maps (DAAM) to evaluate how individual input words directly influence specific regions of generated images in Stable Diffusion. It investigates the visual realization of linguistic syntax and analyzes the core semantic failure modes that degrade image quality.

To conduct this evaluation, the researchers aggregated and upscaled cross-attention scores across all network layers and denoising time steps, producing two-dimensional attribution heat maps without requiring computationally intractable gradient calculations. The approach was evaluated on two synthetic benchmarks derived from standard image caption datasets (COCO-Gen and Unreal-Gen) for object segmentation, followed by crowdsourced human evaluations covering multiple parts of speech. The researchers also parsed 8,000 syntactic head–dependent relationships across 1,000 prompts to measure visual overlap and tested semantic prompts containing pairs of cohyponyms (words sharing a broader category, such as "giraffe" and "zebra") and varying descriptive adjectives.

The analysis produced three primary findings. First, the attribution method achieves competitive unsupervised noun segmentation performance (58.8 to 64.8 mean intersection-over-union) and generates reliable visual maps across diverse grammatical categories, with over 80% to 90% of human ratings scored as fair to excellent. Second, visual interactions between words often reflect syntactic dependencies, showing that verbs spatially encompass their subject and object nouns, whereas descriptive adjectives unexpectedly attend more broadly across the image than the specific nouns they modify. Third, semantically related words suffer from severe feature entanglement: prompts featuring cohyponyms suffer a 9% absolute reduction in generation accuracy compared to unrelated noun pairs (52% versus 61%), frequently generating only one intended entity while their attribution maps heavily overlap.

These findings indicate that current diffusion models struggle with compositional disentanglement. When models process visually related concepts or modifying adjectives, features bleed across the image rather than binding cleanly to specific objects, leading to missing subjects, corrupted backgrounds, and inconsistent prompt fidelity. This behavioral analysis explains why prompt engineering often yields unpredictable results and provides a diagnostic basis for improving generative control and compliance in real-world deployment.

To improve generative performance, future engineering and model development should leverage cross-attention maps as internal feature representations to explicitly disentangle concept attributes and object boundaries. Organizations deploying generative pipelines should implement pre-generation prompt screening or post-generation attention monitoring to detect cohyponym overlap and prevent rendering errors. Further research is necessary to design training techniques or architectural constraints that actively decouple entangled semantic features.

These conclusions are bounded by automated dependency parsing quality and an evaluation benchmark focused primarily on concrete, visual entities derived from caption datasets rather than highly abstract concepts. Additionally, while the attribution maps offer high-confidence diagnostic value, attention scores serve as an empirical proxy for model behavior rather than an exact causal guarantee of model outputs.

Cover for What the DAAM: Interpreting Stable Diffusion Using Cross Attention

Abstract

Diffusion models are a milestone in text-to-image generation, but they remain poorly understood, lacking interpretability analyses. In this paper, we perform a text–image attribution analysis on Stable Diffusion, a recently open-sourced model. To produce attribution maps, we upscale and aggregate cross-attention maps in the denoising module, naming our method DAAM. We validate it by testing its segmentation ability on nouns, as well as its generalized attribution quality on all parts of speech, rated by humans. On two generated datasets, we attain a competitive 58.8–64.8 mIoU on noun segmentation and fair to good mean opinion scores (3.4–4.2) on generalized attribution. Then, we apply DAAM to study the role of syntax in the pixel space across head–dependent heat map interaction patterns for ten common dependency relations. We show that, for some relations, the head map consistently subsumes the dependent, while the opposite is true for others. Finally, we study several semantic phenomena, focusing on feature entanglement; we find that the presence of cohyponyms worsens generation quality by 9%, and descriptive adjectives attend too broadly. We are the first to interpret large diffusion models from a visuolinguistic perspective, which enables future research. Our code is at https://github.com/castorini/daam.

Table of Contents

  • 1 Introduction
  • 2 Our Approach
  • 2.1 Preliminaries
  • 2.2 Diffusion Attentive Attribution Maps
  • 3 Attribution Analyses
  • 3.1 Object Attribution
  • 3.2 Generalized Attribution
  • 4 Visuosyntactic Analysis
  • 5 Visuosemantic Analyses
  • 5.1 Cohyponym Entanglement
  • 5.2 Adjectival Entanglement
  • 6 Related Work and Future Directions
  • 7 Conclusions
  • Limitations
  • Acknowledgments
  • References
  • A Supplements for Attribution Analyses
  • A.1 Object Attribution
  • A.2 Generalized Attribution
  • B Supplements for Syntactic Analyses
  • C Supplements for Semantic Analyses

Knowls

  1. Knowl 1 — Diffusion Attentive Attribution Maps (DAAM) for Text-to-Image Diffusion Models

    model/method

    Diffusion Attentive Attribution Maps (DAAM) is an interpretability method that produces 2D pixel-level attribution heat maps for each word in an input prompt by aggregating cross-attention arrays across the denoising U-Net in text-to-image latent diffusion models (such as Stable Diffusion).

    In latent diffusion models, multi-headed cross-attention layers at downsampling block ii, head ℓ\ell, and denoising time step tt compute token–image attention score arrays Ft(i)↓∈R⌈w/ci⌉×⌈h/ci⌉×lH×lWF_t^{(i)\downarrow} \in \mathbb{R}^{\lceil w/c_i \rceil \times \lceil h/c_i \rceil \times l_H \times l_W} (and symmetrically Ft(i)↑F_t^{(i)\uparrow} for upsampling blocks), where w×hw \times h is the spatial latent dimension, ci>1c_i > 1 is the downsampling factor, lHl_H is the number of attention heads, and lWl_W is the number of text tokens:

    Ft(i)↓(h^i,t↓,X):=softmax((Wq(i)h^i,t↓)(Wk(i)X)Td)F_t^{(i)\downarrow}(\hat{h}_{i,t}^\downarrow, X) := \text{softmax}\left(\frac{(W_q^{(i)} \hat{h}_{i,t}^\downarrow)(W_k^{(i)} X)^T}{\sqrt{d}}\right)

    where X∈RlW×dX \in \mathbb{R}^{l_W \times d} contains the text word embeddings, h^i,t↓\hat{h}_{i,t}^\downarrow is the intermediate latent feature map, and Wq(i),Wk(i)W_q^{(i)}, W_k^{(i)} are learned projection matrices.

    To construct the unthresholded attribution map DkR∈Rw×hD_k^R \in \mathbb{R}^{w \times h} for the kk-th word, DAAM upscales every intermediate attention array Ftj(i)↓[⋅,⋅,ℓ,k]F_{t_j}^{(i)\downarrow}[\cdot, \cdot, \ell, k] and Ftj(i)↑[⋅,⋅,ℓ,k]F_{t_j}^{(i)\uparrow}[\cdot, \cdot, \ell, k] to the target image dimensions (w,h)(w, h) using bicubic interpolation (denoted F~\tilde{F}), and sums them across all layers ii, time steps tjt_j, and attention heads ℓ\ell:

    DkR[x,y]:=∑i,j,ℓ(F~tj,k,ℓ(i)↓[x,y]+F~tj,k,ℓ(i)↑[x,y])D_k^R[x, y] := \sum_{i, j, \ell} \left(\tilde{F}_{t_j, k, \ell}^{(i)\downarrow}[x, y] + \tilde{F}_{t_j, k, \ell}^{(i)\uparrow}[x, y]\right)

    To produce a binary segmentation mask DkIτ[x,y]∈{0,1}D_k^{I\tau}[x, y] \in \{0, 1\}, DkRD_k^R is thresholded relative to its maximum value using a threshold parameter τ∈[0,1]\tau \in [0, 1]:

    DkIτ[x,y]:=I(DkR[x,y]≥τmax⁡u,vDkR[u,v])D_k^{I\tau}[x, y] := \mathbb{I}\left(D_k^R[x, y] \ge \tau \max_{u, v} D_k^R[u, v]\right)

    where I(⋅)\mathbb{I}(\cdot) is the indicator function.

  2. Knowl 2 — Noun Semantic Segmentation Performance on COCO-Gen and Unreal-Gen Datasets

    data/table

    To evaluate the veracity of DAAM attribution maps as semantic segmentations without explicit segmentation training, Stable Diffusion 2.0-base was evaluated on two datasets of 500 prompt–image pairs each: COCO-Gen (prompts drawn from MS-COCO validation captions) and Unreal-Gen (prompts with nouns randomly swapped to test compositional generalization). All countable nouns were hand-annotated with ground-truth segmentation masks. Predictions were evaluated using mean Intersection over Union restricted to the 80 COCO classes (mIoU80\text{mIoU}^{80}) and unrestricted over open vocabulary (mIoU∞\text{mIoU}^{\infty}).

    Method COCO-Gen Unreal-Gen
    mIoU80\text{mIoU}^{80} mIoU∞\text{mIoU}^{\infty} mIoU80\text{mIoU}^{80} mIoU∞\text{mIoU}^{\infty}
    Supervised Methods
    Mask R-CNN (ResNet-101) 80.4 26.9 84.0 25.7
    QueryInst (ResNet-101-FPN) 81.2 27.1 83.6 25.5
    Mask2Former (Swin-S) 82.0 27.4 85.0 25.9
    CLIPSeg 74.2 67.0 79.0 64.5
    Unsupervised Methods
    Whole image mask 21.7 20.6 24.8 24.0
    PiCIE + H 31.7 22.3 35.9 29.2
    STEGO (DINO ViT-B) 42.0 61.3 38.2 56.6
    Our DAAM-0.3 62.7 57.0 64.7 62.6
    Our DAAM-0.4 62.8 58.8 64.8 62.2
    Our DAAM-0.5 59.6 55.9 60.0 57.1

    DAAM-0.4 achieves 62.8 mIoU8062.8\text{ mIoU}^{80} and 58.8 mIoU∞58.8\text{ mIoU}^{\infty} on COCO-Gen, outperforming unsupervised segmentation baselines by 6–27 mIoU points. Because DAAM operates open-vocabulary, it surpasses closed-set COCO-supervised baselines on mIoU∞\text{mIoU}^{\infty} by 22–28 points. Consistent performance on Unreal-Gen (64.8 mIoU8064.8\text{ mIoU}^{80} / 62.2 mIoU∞62.2\text{ mIoU}^{\infty}) demonstrates that DAAM attribution remains reliable when the model generates unseen, unnatural compositions.

  3. Knowl 3 — Metrics for Head–Dependent Attentive Overlap and Dominance

    definition

    Given a pair of binarized DAAM attribution maps (D(i1)Iτ,D(i2)Iτ)(D_{(i_1)}^{I\tau}, D_{(i_2)}^{I\tau}) corresponding to a dependent token (index i1i_1) and its syntactic head token (index i2i_2) across nn syntactic dependency pairs, the interaction and directional dominance between their visual attribution areas are quantified by three set-based overlap statistics:

    1. Mean Intersection over Union (mIoU): mIoU=1n∑i=1n∑x,y(D(i1)Iτ[x,y]∧D(i2)Iτ[x,y])∑x,y(D(i1)Iτ[x,y]∨D(i2)Iτ[x,y])\text{mIoU} = \frac{1}{n} \sum_{i=1}^n \frac{\sum_{x, y} \left(D_{(i_1)}^{I\tau}[x, y] \land D_{(i_2)}^{I\tau}[x, y]\right)}{\sum_{x, y} \left(D_{(i_1)}^{I\tau}[x, y] \lor D_{(i_2)}^{I\tau}[x, y]\right)}

    2. Mean Intersection over Dependent (mIoD): mIoD=1n∑i=1n∑x,y(D(i1)Iτ[x,y]∧D(i2)Iτ[x,y])∑x,yD(i1)Iτ[x,y]\text{mIoD} = \frac{1}{n} \sum_{i=1}^n \frac{\sum_{x, y} \left(D_{(i_1)}^{I\tau}[x, y] \land D_{(i_2)}^{I\tau}[x, y]\right)}{\sum_{x, y} D_{(i_1)}^{I\tau}[x, y]}

    3. Mean Intersection over Head (mIoH): mIoH=1n∑i=1n∑x,y(D(i1)Iτ[x,y]∧D(i2)Iτ[x,y])∑x,yD(i2)Iτ[x,y]\text{mIoH} = \frac{1}{n} \sum_{i=1}^n \frac{\sum_{x, y} \left(D_{(i_1)}^{I\tau}[x, y] \land D_{(i_2)}^{I\tau}[x, y]\right)}{\sum_{x, y} D_{(i_2)}^{I\tau}[x, y]}

    where ∧\land and ∨\lor denote the logical AND and OR operators on binary pixel masks.

    Dominance is defined as the absolute difference Δ:=∣mIoD−mIoH∣\Delta := |\text{mIoD} - \text{mIoH}|. When mIoD>mIoH\text{mIoD} > \text{mIoH}, the head's attribution map covers a greater proportion of the dependent's area (the head spatially subsumes the dependent). When mIoH>mIoD\text{mIoH} > \text{mIoD}, the dependent's attribution map spatially subsumes the head.

  4. Knowl 4 — Head–Dependent Cross-Attention Attribution Overlap Across Syntactic Relations

    data/table

    Pairwise interactions between head and dependent DAAM maps were evaluated across 8,000 dependency pairs sampled from 1,000 MS-COCO prompts parsed with Stanford CoreNLP Enhanced Universal Dependencies and synthesized with Stable Diffusion 2.0-base:

    # Relation mIoD mIoH Δ\Delta mIoU
    1 Unrelated pairs 65.1 66.1 1.0 47.5
    2 All head–dependent pairs 62.3 62.0 0.3 43.4
    3 compound 71.3 71.5 0.2 51.1
    4 punct 68.2 70.0 1.8 49.5
    5 nconj:and 58.0 56.1 1.9 38.2
    6 det 54.8 52.2 2.6 35.0
    7 case 51.7 58.1 6.4 36.9
    8 acl 67.4 79.3 12.0 55.4
    9 nsubj 76.4 63.9 12.5 52.2
    10 amod 62.4 77.6 15.2 51.1
    11 nmod:of 73.5 57.9 15.6 47.5
    12 obj 75.6 46.3 29.3 55.4
    13 Coreferent word pairs 84.8 77.4 7.4 66.6

    Dominant relations (∥mIoD−mIoH∥>10\|\text{mIoD} - \text{mIoH}\| > 10, p<0.01p < 0.01 with Holm–Bonferroni correction) reveal structural visual patterns:

    • For verbal arguments (nsubj, obj), the head verb dominates its noun subject (Δ=12.5\Delta = 12.5) and direct object (Δ=29.3\Delta = 29.3), indicating that verbs visually contextualize both the arguments and their surrounding scene.
    • For nominal modifiers, collective nominal heads (nmod:of, e.g., "pile of oranges") dominate their dependents (Δ=15.6\Delta = 15.6), whereas adjectival modifiers (amod) counterintuitively exhibit dependent dominance (Δ=15.2\Delta = 15.2), where descriptive adjectives visually subsume the nouns they modify.
    • Coreferent word pairs show the highest overall overlap (66.6 mIoU), reflecting coreference grounding during synthesis.
  5. Knowl 5 — Cohyponym Feature Entanglement and Generation Failure in Stable Diffusion

    empirical result

    In Stable Diffusion 2.0-base, prompts containing semantically related cohyponyms (nouns sharing an immediate hypernym in a WordNet taxonomy) suffer from visual feature entanglement and significantly higher generation failure rates than prompts containing non-cohyponyms.

    In an experiment generating 1,000 images using the prompt template "a(n) <noun> and a(n) <noun>" over 28 MS-COCO noun classes spanning 16 hypernyms:

    1. Generation Accuracy: The generation accuracy (the fraction of images where human raters confirmed both objects were present) was 52% for cohyponyms (e.g., "a giraffe and a zebra") versus 61% for non-cohyponyms (e.g., "a cake and a bus"), a statistically significant reduction of 9% (p<0.01p < 0.01, exact test).
    2. Attribution Overlap: The mean DAAM overlap (mIoU at τ=0.4\tau = 0.4) between the two nouns was 46.7 for cohyponyms compared to 22.9 for non-cohyponyms.
    3. Overlap Breakdown by Correctness:
      • Correctly generated non-cohyponyms: 5.43 mIoU5.43\text{ mIoU}
      • Incorrectly generated non-cohyponyms: 29.5 mIoU29.5\text{ mIoU}
      • Correctly generated cohyponyms: 43.5 mIoU43.5\text{ mIoU}
      • Incorrectly generated cohyponyms: 60.9 mIoU60.9\text{ mIoU}
    4. Accuracy by Overlap Level: Image generation accuracy strongly degrades as DAAM attribution overlap increases:
      • Low overlap (mIoU≤0.4\text{mIoU} \le 0.4): 77.5% accuracy for non-cohyponyms; 71.7% for cohyponyms.
      • Moderate overlap (0.4<mIoU<0.60.4 < \text{mIoU} < 0.6): 56.1% for non-cohyponyms; 19.3% for cohyponyms.
      • High overlap (mIoU≥0.6\text{mIoU} \ge 0.6): 36.0% for non-cohyponyms; 9.75% for cohyponyms.
  6. Knowl 6 — Adjectival Feature Entanglement and Cross-Attention Leakage Across Image Backgrounds

    empirical result

    In Stable Diffusion 2.0, descriptive adjectival modifiers attend broadly across the generated image rather than remaining localized to the noun they modify, leaking their semantic features into the global background.

    Using prompts of the structure "a <adj> <noun> <verb phrase>" while fixing the scene layout by holding cross-attention maps constant to a seed prompt:

    • In "a {rusty, metallic, wooden} shovel sitting in a clean shed", changing only the adjective causes the surrounding background shed to become visibly rusty/grey/wooden instead of remaining clean.
    • In "a {bumpy, smooth, spiky} ball rolling down a hill", altering the adjective changes the terrain texture (producing rugged ground for bumpy, flat ground for smooth, and grass blades for spiky).
    • In "a {blue, green, red} car driving down the streets", changing the color adjective introduces quantifiable excess hue of that color across the background street pixels outside the cropped car region.

    This broad cross-attention leakage accounts for why adjectival modifiers (amod) exhibit visual dominance over their head nouns (mIoH=77.6\text{mIoH} = 77.6 vs. mIoD=62.4\text{mIoD} = 62.4, Δ=15.2\Delta = 15.2).

  7. Knowl 7 — Human Evaluation of DAAM Attribution Quality Across Parts of Speech

    empirical result

    In a human evaluation on Amazon Mechanical Turk covering 2,800 unique prompt–word pairs (200 words randomly sampled across 14 parts of speech in MS-COCO) scored on a 5-point Likert scale (1=Bad to 5=Excellent by three independent master-level raters):

    • Interpretable open-class words achieve high attribution fidelity: adjectives (ADJ), verbs (VERB), nouns (NOUN), and proper nouns (PROPN) obtain median opinion scores between 3.4 and 4.2 ("good" to above "good").
    • Numerals (NUM) and adverbs (ADV) achieve scores closer to "fair" (MOS 3.0–3.4), as their visual localization is naturally more diffuse (e.g., motion blur regions or counting boundaries).
    • The proportion of fair-to-excellent ratings is >80%>80\% for numerals and adverbs, and >90%>90\% for adjectives, verbs, nouns, and proper nouns.
    • In direct comparison with CLIPSeg across generalized POS attribution, DAAM significantly outperforms CLIPSeg on verbs, proper nouns, and adverbs (p<0.02p < 0.02, unpaired tt-test), achieving an overall Mean Opinion Score of 3.4 compared to CLIPSeg's 2.5.
  8. Knowl 8 — Necessity of Full Spatiotemporal and Multi-Layer Cross-Attention Aggregation

    empirical result

    Ablation experiments on Stable Diffusion 2.0-base (30 DPM solver inference steps) evaluating noun segmentation on COCO-Gen demonstrate that aggregating across all denoising time steps and all U-Net layer resolutions is necessary for optimal attribution quality:

    1. Time Step Truncation: Restricting the temporal summation in DAAM to either the first nn time steps (j≤j∗j \le j^*) or the last nn time steps (j≥j∗j \ge j^*) for j∗∈{1,2,5,10,15,20,25,30}j^* \in \{1, 2, 5, 10, 15, 20, 25, 30\} produces strictly inferior mIoU80\text{mIoU}^{80} relative to summing over all 30 time steps (which achieves 62.8 mIoU8062.8\text{ mIoU}^{80}).
    2. Layer Resolution Truncation: Restricting cross-attention aggregation to individual spatial resolutions (16×1616\times 16, 32×3232\times 32, or 64×6464\times 64) degrades segmentation quality compared to summing over both downsampling and upsampling layers across all resolutions.
    3. Alternative Segmentation Metrics: On COCO-Gen, DAAM-0.4 achieves 90% pixel accuracy (matching CLIPSeg at 90% and outperforming Mask2Former at 85%) and a Dice score of 68 (compared to CLIPSeg's 72 and Mask2Former's 30), confirming the validity of DAAM across multiple evaluation metrics.
  9. Knowl 9 — Limitations of Visuolinguistic Attribution in Latent Diffusion Models

    limitation

    The methodology and empirical findings of DAAM are subject to four main limitations:

    1. Syntactic Tooling Dependencies: Visuosyntactic interaction analysis relies on automated dependency parsers (Stanford CoreNLP), whose potential parsing errors or paradigm constraints could impact syntactic head–dependent boundary assignments.
    2. Concrete Concept Dataset Bias: Evaluations are conducted on MS-COCO image captions, creating an inherent domain bias toward concrete physical nouns and visible properties, leaving the attribution behavior of abstract concepts (e.g., "love", "dignity") unverified.
    3. Attention Weight as Attribution Proxy: Cross-attention score aggregation serves as a tractable surrogate for gradient-based or perturbation attribution across diffusion steps, but attention weights may not strictly reflect exact input-output causal gradients.
    4. Observational vs. Remediation Scope: The study identifies and analyzes failure modes (such as cohyponym entanglement and adjectival cross-attention leakage) observationally without providing algorithmic or architectural training modifications to remediate them.

Coverage note — None was omitted; all key contributions—DAAM mathematical formulation, segmentation benchmark on COCO-Gen/Unreal-Gen, generalized POS human study, visuosyntactic overlap statistics, cohyponym and adjectival entanglement analyses, ablation studies, and limitations—are included.

References

  1. 1.David Alvarez-Melis and Tommi S. Jaakkola. 2018. On the robustness of interpretability methods. arXiv:1806.08049.
  2. 2.Kai Chen, Jiaqi Wang, Jiangmiao Pang, Yuhang Cao, Yu Xiong, Xiaoxiao Li, Shuyang Sun, Wansen Feng, Ziwei Liu, Jiarui Xu, et al. 2019. MMDetection: Open MMLab detection toolbox and benchmark. arXiv:1906.07155.
  3. 3.Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. 2022. Masked-attention mask transformer for universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  4. 4.Jang Hyun Cho, Utkarsh Mall, Kavita Bala, and Bharath Hariharan. 2021. PiCIE: Unsupervised semantic segmentation using invariance and equivariance in clustering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  5. 5.Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. What does BERT look at? an analysis of BERT's attention. In Proceedings of BlackboxNLP.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers).
  7. 7.Yuxin Fang, Shusheng Yang, Xinggang Wang, Yu Li, Chen Fang, Ying Shan, Bin Feng, and Wenyu Liu. 2021. Instances as queries. In Proceedings of the IEEE/CVF International Conference on Computer Vision.
  8. 8.Mark Hamilton, Zhoutong Zhang, Bharath Hariharan, Noah Snavely, and William T. Freeman. 2021. Unsupervised semantic segmentation by distilling feature correspondences. In International Conference on Learning Representations.
  9. 9.Charles R. Harris, K. Jarrod Millman, Stéfan J. Van Der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, et al. 2020. Array programming with NumPy. Nature.
  10. 10.David J. Hauser and Norbert Schwarz. 2016. Attentive turkers: MTurk participants perform better on online attention checks than do subject pool participants. Behavior research methods.
  11. 11.Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017. Mask R-CNN. In Proceedings of the IEEE/CVF International Conference on Computer Vision.
  12. 12.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  13. 13.Lisa Anne Hendricks and Aida Nematzadeh. 2021. Probing image-language transformers for verb understanding. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021.
  14. 14.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2022. Prompt-to-prompt image editing with cross attention control. arXiv:2208.01626.
  15. 15.John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers).
  16. 16.Nikolai Ilinykh and Simon Dobnik. 2022. Attention as grounding: Exploring textual and cross-modal attention on entities and relations in language-and-vision transformer. In Findings of the Association for Computational Linguistics: ACL 2022.
  17. 17.Zhiying Jiang, Raphael Tang, Ji Xin, and Jimmy Lin. 2020. Inserting information bottlenecks for attribution in transformers. In Findings of the Association for Computational Linguistics: EMNLP 2020.
  18. 18.Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  19. 19.Diederik P. Kingma and Max Welling. 2013. Autoencoding variational bayes. arXiv:1312.6114.
  20. 20.Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019. Revealing the dark secrets of BERT. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP).
  21. 21.Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017. Feature pyramid networks for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  22. 22.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014. Microsoft COCO: Common objects in context. In European Conference on Computer Vision.
  23. 23.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision.
  24. 24.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2022. DPM-solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. arXiv:2206.00927.
  25. 25.Timo Lüddecke and Alexander Ecker. 2022. Image segmentation using text and image prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  26. 26.Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Rose Finkel, Steven Bethard, and David McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In Proceedings of 52nd annual meeting of the Association for Computational Linguistics: System Demonstrations.
  27. 27.Mary Ann Marcinkiewicz. 1994. Building a large annotated corpus of English: The Penn treebank. Using Large Corpora.
  28. 28.Joanna Materzyńska, Antonio Torralba, and David Bau. 2022. Disentangling visual and written concepts in CLIP. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  29. 29.George A. Miller. 1995. WordNet: a lexical database for English. Communications of the ACM.
  30. 30.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning.
  31. 31.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with CLIP latents. arXiv:2204.06125.
  32. 32.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  33. 33.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention.
  34. 34.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. arXiv:2205.11487.
  35. 35.Christoph Schuhmann, Romain Beaumont, Cade W Gordon, Ross Wightman, Theo Coombes, et al. 2022. LAION-5B: An open large-scale dataset for training next generation image-text models.
  36. 36.Sebastian Schuster and Christopher D. Manning. 2016. Enhanced English universal dependencies: An improved representation for natural language understanding tasks. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16).
  37. 37.Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision.
  38. 38.Sofia Serrano and Noah A. Smith. 2019. Is attention interpretable? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.
  39. 39.Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, and Sebastian Riedel. 2016. Neural architectures for fine-grained entity type classification. arXiv preprint arXiv:1606.01341.
  40. 40.Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations.
  41. 41.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems.
  42. 42.Jesse Vig. 2019. BertViz: A tool for visualizing multihead self-attention in the BERT model. In ICLR Workshop: Debugging Machine Learning Models.
  43. 43.Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf. 2022. Diffusers: State-of-the-art diffusion models.
  44. 44.Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019. AllenNLP interpret: A framework for explaining predictions of NLP models. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations.
  45. 45.Max Woolf. 2022. Stable Diffusion 2.0 and the importance of negative prompts for good results.
  46. 46.Chenyun Wu, Zhe Lin, Scott Cohen, Trung Bui, and Subhransu Maji. 2020. PhraseCut: Language-based image segmentation in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  47. 47.Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Yingxia Shao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2022. Diffusion models: A comprehensive survey of methods and applications. arXiv:2209.00796.

Citation

MLA
Tang, R., et al. “What the DAAM: Interpreting Stable Diffusion Using Cross Attention”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 5644–59, https://doi.org/10.18653/v1/2023.acl-long.310.
APA
Tang, R., Liu, L., Pandey, A., Jiang, Z., Yang, G., Kumar, K., Stenetorp, P., Lin, J., & Türe, F. (2023). What the DAAM: Interpreting Stable Diffusion Using Cross Attention. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5644–5659. https://doi.org/10.18653/v1/2023.acl-long.310
Chicago
Tang, R., L. Liu, A. Pandey, et al. 2023. “What the DAAM: Interpreting Stable Diffusion Using Cross Attention”. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 5644–59. https://doi.org/10.18653/v1/2023.acl-long.310.
Harvard
Tang, R. et al. (2023) “What the DAAM: Interpreting Stable Diffusion Using Cross Attention”, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 5644–5659. Available at: https://doi.org/10.18653/v1/2023.acl-long.310.
Vancouver
1. Tang R, Liu L, Pandey A, Jiang Z, Yang G, Kumar K, Stenetorp P, Lin J, Türe F (2023) What the DAAM: Interpreting Stable Diffusion Using Cross Attention. In: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 5644–5659

BibTeX

@inproceedings{tang-etal-2023-daam,
    title = "What the {DAAM}: Interpreting Stable Diffusion Using Cross Attention",
    author = "Tang, Raphael  and
      Liu, Linqing  and
      Pandey, Akshat  and
      Jiang, Zhiying  and
      Yang, Gefei  and
      Kumar, Karun  and
      Stenetorp, Pontus  and
      Lin, Jimmy  and
      Ture, Ferhan",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.310/",
    doi = "10.18653/v1/2023.acl-long.310",
    pages = "5644--5659"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/