Initno: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization

Xiefan GuoJinlin LiuMiaomiao CuiJiankai LiHongyu YangDi Huang

article2024CVPR142 citations

Proposes a training-free optimization method that evaluates attention scores to steer initial noise into semantically valid latent regions, preventing common text-to-image synthesis failures such as subject neglect, mixing, and incorrect attribute binding.

Listen

Modern text-to-image diffusion models produce visually impressive imagery but frequently fail to align precisely with user prompts. Common issues include omitting requested subjects, blending distinct subjects together, and incorrectly assigning attributes such as colors. These semantic alignment errors limit the reliability and usability of image generation systems in professional and commercial applications.

The article evaluates the root cause of these alignment failures and demonstrates an optimization framework called Initial Noise Optimization (INITNO). The primary objective is to steer random starting noise into valid latent space regions prior to image synthesis, ensuring that generated images faithfully adhere to prompt instructions without requiring model retraining.

To accomplish this, the researchers analyzed attention layers within latent diffusion models, identifying cross-attention response as a measure of subject omission and self-attention conflict as a measure of subject blending. Using these metrics alongside a distribution alignment constraint, the authors formulated an optimization pipeline that adjusts the mean and standard deviation of the initial noise. The approach was tested on benchmark datasets comprising combinations of animals and objects, and evaluated using automated image-text similarity metrics and a structured user study comparing against leading alternatives.

Key findings show substantial improvements in image-text alignment. In objective evaluations, the proposed method consistently achieved higher image-text and text-text similarity scores than standard Stable Diffusion and existing step-wise guidance techniques across all benchmark categories. In a blind user preference study with image processing specialists, the proposed method received 63.33% of favorable votes, vastly outperforming the baseline Stable Diffusion (4.17%) and competing methods (which scored between 2.50% and 14.17%). Furthermore, the framework successfully prevented out-of-distribution image distortions by constraining the optimized noise to standard statistical distributions.

These results indicate that adjusting the initial starting noise is a highly effective, plug-and-play solution for improving generative accuracy. By performing full optimization on the starting noise rather than fine-tuning every subsequent step of image creation, the method avoids delicate parameter tuning and reduces the risk of generating distorted, out-of-domain artifacts. This significantly enhances the control and fidelity of existing diffusion systems without the high cost of retraining massive neural networks.

Organizations deploying text-to-image systems should consider integrating initial noise optimization to improve output reliability for complex compositional prompts and grounded generation tasks. However, stakeholders must account for computational trade-offs: image generation time increased from approximately 8.34 seconds for baseline Stable Diffusion to 18.93 seconds with the proposed approach. Future work should focus on reducing optimization latency and validating the technique across wider domains and larger foundational diffusion architectures.

No sufficiently relevant recommendations were found.

Cover for Initno: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization

Abstract

Recent strides in the development of diffusion models, exemplified by advancements such as Stable Diffusion, have underscored their remarkable prowess in generating visually compelling images. However, the imperative of achieving a seamless alignment between the generated image and the provided prompt persists as a formidable challenge. This paper traces the root of these difficulties to invalid initial noise, and proposes a solution in the form of Initial Noise Optimization (INITNO), a paradigm that refines this noise. Considering text prompts, not all random noises are effective in synthesizing semantically-faithful images. We design the cross-attention response score and the self-attention conflict score to evaluate the initial noise, bifurcating the initial latent space into valid and invalid sectors. A strategically crafted noise optimization pipeline is developed to guide the initial noise towards valid regions. Our method, validated through rigorous experimentation, shows a commendable proficiency in generating images in strict accordance with text prompts. Our code is available at https://github.com/xiefan-guo/initno.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminaries
  • 4. INITNO
  • 4.1. Initial latent space partitioning
  • 4.2. Noise optimization pipeline
  • 5. Experiments
  • 5.1. Experimental settings
  • 5.2. Qualitative comparison
  • 5.3. Quantitative comparison
  • 5.4. Ablation study
  • 5.5. Grounded Text-to-Image
  • 5.6. More results
  • 6. Conclusion
  • Acknowledgment
  • References

Knowls

  1. Knowl 1 — Initial Latent Space Partitioning in Text-to-Image Diffusion

    definition

    In text-to-image (T2I) diffusion models such as Stable Diffusion, the initial Gaussian noise space N(0,I)\mathcal{N}(\mathbf{0}, \mathbf{I}) at timestep TT can be partitioned into valid and invalid regions based on whether a sampled initial noise latent zT\mathbf{z}_T yields a semantically aligned image that avoids subject neglect and subject mixing.

    An initial noise vector zT\mathbf{z}_T is defined as belonging to the valid region if and only if it satisfies two conditions evaluated during the initial denoising step:

    1. The cross-attention response score SCrossAttn\mathcal{S}_{\text{CrossAttn}} is strictly below a predefined threshold τc\tau_c (empirically set to 0.20.2).
    2. The self-attention conflict score SSelfAttn\mathcal{S}_{\text{SelfAttn}} is strictly below a predefined threshold τs\tau_s (empirically set to 0.30.3).

    zT∈Valid  ⟺  SCrossAttn(zT)<τcandSSelfAttn(zT)<τs\mathbf{z}_T \in \text{Valid} \iff \mathcal{S}_{\text{CrossAttn}}(\mathbf{z}_T) < \tau_c \quad \text{and} \quad \mathcal{S}_{\text{SelfAttn}}(\mathbf{z}_T) < \tau_s

    Any initial noise latent failing either condition is classified as invalid, as it leads to failure modes including omitted prompt entities or blended entity representations.

  2. Knowl 2 — Cross-Attention Response Score

    equation

    The cross-attention response score SCrossAttn\mathcal{S}_{\text{CrossAttn}} quantifies the presence or neglect of prompt subject tokens during the initial denoising stage. Let Y\mathcal{Y} denote the set of target subject text tokens extracted from prompt y={y1,y2,…,yn}\mathbf{y} = \{y_1, y_2, \dots, y_n\}. The aggregated cross-attention map Ac∈R16×16×n\mathbf{A}^c \in \mathbb{R}^{16 \times 16 \times n} is computed by averaging the cross-attention maps across all attention layers and heads at a resolution of 16×1616 \times 16 pixels, omitting the special start-of-text token sot\text{sot} and applying a softmax operation across remaining tokens.

    For each target subject token yi∈Y\mathbf{y}_i \in \mathcal{Y}, Ayic\mathbf{A}^c_{\mathbf{y}_i} denotes its 16×1616 \times 16 spatial attention map. The cross-attention response score is defined as:

    SCrossAttn=1−min⁡yi∈Ymax⁡x,y(Ayic[x,y])\mathcal{S}_{\text{CrossAttn}} = 1 - \min_{\mathbf{y}_i \in \mathcal{Y}} \max_{x, y} \left( \mathbf{A}^c_{\mathbf{y}_i}[x, y] \right)

    where max⁡x,y(Ayic[x,y])\max_{x, y} \left(\mathbf{A}^c_{\mathbf{y}_i}[x, y]\right) captures the maximum activation for token yi\mathbf{y}_i across all spatial locations (x,y)(x, y). A low maximum cross-attention response implies subject neglect, driving SCrossAttn\mathcal{S}_{\text{CrossAttn}} toward 11. An initial latent is considered valid with respect to subject presence if SCrossAttn<τc\mathcal{S}_{\text{CrossAttn}} < \tau_c, where τc=0.2\tau_c = 0.2.

  3. Knowl 3 — Self-Attention Conflict Score

    equation

    The self-attention conflict score SSelfAttn\mathcal{S}_{\text{SelfAttn}} quantifies semantic entanglement and subject mixing between distinct subject tokens in the image latent space. Given the target subject tokens Y\mathcal{Y} and aggregated cross-attention maps Ac\mathbf{A}^c, the spatial coordinate (xi,yi)(x_i, y_i) corresponding to the peak cross-attention activation for subject token yi\mathbf{y}_i is queried as:

    xi,yi=arg⁡max⁡x,yAyic[x,y]x_i, y_i = \arg\max_{x, y} \mathbf{A}^c_{\mathbf{y}_i}[x, y]

    Let Axi,yis∈R16×16\mathbf{A}^s_{x_i, y_i} \in \mathbb{R}^{16 \times 16} denote the self-attention map corresponding to the query patch (xi,yi)(x_i, y_i), extracted by averaging self-attention maps across layers and heads at the 16×1616 \times 16 resolution. For every pair of distinct subject tokens (yi,yj)(\mathbf{y}_i, \mathbf{y}_j), their spatial self-attention overlap ratio f(yi,yj)f(\mathbf{y}_i, \mathbf{y}_j) is defined as:

    f(yi,yj)=∑x,ymin⁡(Axi,yis[x,y],Axj,yjs[x,y])∑x,y(Axi,yis[x,y]+Axj,yjs[x,y])f(\mathbf{y}_i, \mathbf{y}_j) = \frac{\sum_{x, y} \min\left( \mathbf{A}^s_{x_i, y_i}[x, y], \mathbf{A}^s_{x_j, y_j}[x, y] \right)}{\sum_{x, y} \left( \mathbf{A}^s_{x_i, y_i}[x, y] + \mathbf{A}^s_{x_j, y_j}[x, y] \right)}

    The self-attention conflict score over all NN pairs of target subjects (N=(∣Y∣2)N = \binom{|\mathcal{Y}|}{2}) is:

    SSelfAttn=∑yi,yj∈Y,i<jf(yi,yj)N\mathcal{S}_{\text{SelfAttn}} = \sum_{\mathbf{y}_i, \mathbf{y}_j \in \mathcal{Y}, i < j} \frac{f(\mathbf{y}_i, \mathbf{y}_j)}{N}

    A higher score indicates mutual spatial overlap between the feature activations of distinct entities (causing subject mixing). An initial latent is considered non-conflicting if SSelfAttn<τs\mathcal{S}_{\text{SelfAttn}} < \tau_s, where τs=0.3\tau_s = 0.3.

  4. Knowl 4 — Initial Noise Distribution Optimization Strategy and Joint Loss

    model/method

    Rather than modifying noise via additive point updates (ϵ′←ϵ+Δϵ\epsilon' \leftarrow \epsilon + \Delta \epsilon), Initial Noise Optimization (INITNO) updates the initial latent zT∼N(0,I)\mathbf{z}_T \sim \mathcal{N}(\mathbf{0}, \mathbf{I}) by optimizing parameter shifts on the underlying Gaussian distribution parameters μ\mu and σ\sigma:

    μ′←μ+Δμ,σ′←σ+Δσ\mu' \leftarrow \mu + \Delta \mu, \qquad \sigma' \leftarrow \sigma + \Delta \sigma zT′=μ′+σ′zT,zT′∼N(μ′,σ′2)\mathbf{z}_T' = \mu' + \sigma' \mathbf{z}_T, \qquad \mathbf{z}_T' \sim \mathcal{N}(\mu', {\sigma'}^2)

    where μ\mu and σ\sigma are learnable parameters initialized to 00 and 11, respectively.

    The parameters μ\mu and σ\sigma are updated using a joint objective function that balances attention constraints with a distribution alignment regularizer:

    Ljoint=λ1LCrossAttn+λ2LSelfAttn+λ3LKL\mathcal{L}_{\text{joint}} = \lambda_1 \mathcal{L}_{\text{CrossAttn}} + \lambda_2 \mathcal{L}_{\text{SelfAttn}} + \lambda_3 \mathcal{L}_{\text{KL}}

    where:

    • LCrossAttn=SCrossAttn\mathcal{L}_{\text{CrossAttn}} = \mathcal{S}_{\text{CrossAttn}} (the cross-attention response score),
    • LSelfAttn=SSelfAttn\mathcal{L}_{\text{SelfAttn}} = \mathcal{S}_{\text{SelfAttn}} (the self-attention conflict score),
    • LKL=KL(N(μ,σ2) ∥ N(0,I))\mathcal{L}_{\text{KL}} = \text{KL}\left( \mathcal{N}(\mu, \sigma^2) \,\parallel\, \mathcal{N}(\mathbf{0}, \mathbf{I}) \right) is the Kullback-Leibler divergence between the transformed distribution and the standard normal distribution, preventing the optimized noise from generating out-of-domain distorted latents.

    The weighting hyperparameters are set to λ1=1\lambda_1 = 1, λ2=1\lambda_2 = 1, and λ3=500\lambda_3 = 500.

  5. Knowl 5 — INITNO Initial Noise Optimization Algorithm

    algorithm

    Initial Noise Optimization (INITNO) iteratively optimizes the distribution parameters μ\mu and σ\sigma of an initial noise latent zT\mathbf{z}_T until it satisfies the validity thresholds or reaches computational resource limits.

    Input: Pre-trained T2I diffusion model SD(⋅)SD(\cdot), text prompt yy, thresholds τc=0.2\tau_c = 0.2, τs=0.3\tau_s = 0.3, max step τMaxStep=50\tau_{\text{MaxStep}} = 50, max rounds τMaxRound=5\tau_{\text{MaxRound}} = 5
    Output: Generated image xx
    Initialize noise pool P←∅\mathcal{P} \leftarrow \emptyset
    for i=1i = 1 to τMaxRound\tau_{\text{MaxRound}} do
        Initialize zT∼N(0,1)z_T \sim \mathcal{N}(0, 1), μ←0\mu \leftarrow 0, σ←1\sigma \leftarrow 1
        for j=1j = 1 to τMaxStep\tau_{\text{MaxStep}} do
            Extract attention maps: _,Ac,As←SD(μ+σzT,y)\_, \mathbf{A}^c, \mathbf{A}^s \leftarrow SD(\mu + \sigma z_T, y)
            Calculate SCrossAttn\mathcal{S}_{\text{CrossAttn}} and SSelfAttn\mathcal{S}_{\text{SelfAttn}}
            if SCrossAttn<τc\mathcal{S}_{\text{CrossAttn}} < \tau_c and SSelfAttn<τs\mathcal{S}_{\text{SelfAttn}} < \tau_s then
                z^T←μ+σzT\hat{z}_T \leftarrow \mu + \sigma z_T
                Synthesize image: x←SD(z^T,y)x \leftarrow SD(\hat{z}_T, y)
                return xx
            else
                Calculate Ljoint=λ1LCrossAttn+λ2LSelfAttn+λ3LKL\mathcal{L}_{\text{joint}} = \lambda_1 \mathcal{L}_{\text{CrossAttn}} + \lambda_2 \mathcal{L}_{\text{SelfAttn}} + \lambda_3 \mathcal{L}_{\text{KL}}
                Update μ,σ←Adam(μ,σ,Ljoint)\mu, \sigma \leftarrow \text{Adam}(\mu, \sigma, \mathcal{L}_{\text{joint}}) with learning rate 0.010.01
                P←P∪{μ+σzT}\mathcal{P} \leftarrow \mathcal{P} \cup \{\mu + \sigma z_T\}
            end if
        end for
    end for
    z^T←arg⁡min⁡z∈P(SCrossAttn(z)+SSelfAttn(z))\hat{z}_T \leftarrow \arg\min_{z \in \mathcal{P}} \left( \mathcal{S}_{\text{CrossAttn}}(z) + \mathcal{S}_{\text{SelfAttn}}(z) \right)
    Synthesize image: x←SD(z^T,y)x \leftarrow SD(\hat{z}_T, y)
    return xx

    Attention maps Ac\mathbf{A}^c and As\mathbf{A}^s are smoothed using a Gaussian filter (kernel size 33, standard deviation 0.50.5) before computing scores. Optimization uses the Adam optimizer with a learning rate of 1×10−21 \times 10^{-2}.

  6. Knowl 6 — Comparative Effectiveness of Self-Attention vs Cross-Attention Conflict Loss

    empirical result

    Subject mixing in diffusion models arises from overlapping self-attention maps between distinct prompt concepts in the spatial domain. Replacing the self-attention conflict loss LSelfAttn\mathcal{L}_{\text{SelfAttn}} with a cross-attention conflict loss (which minimizes cross-attention spatial overlap between different token embeddings) fails to adequately resolve subject mixing.

    The failure of cross-attention conflict minimization is attributed to inaccuracies in CLIP-derived text conditioning embeddings. In contrast, operating directly on UNet intermediate self-attention maps As\mathbf{A}^s explicitly decouples the spatial features of distinct subjects, effectively preventing conceptual bleeding (such as fusing cat and rabbit facial features or apple and pear textures).

  7. Knowl 7 — Ablation on Distribution Alignment Loss in Initial Noise Optimization

    empirical result

    Optimizing initial latent variables solely via cross-attention and self-attention objectives (λ3=0\lambda_3 = 0) causes the parameters (μ,σ)(\mu, \sigma) to drift away from the standard normal distribution N(0,I)\mathcal{N}(\mathbf{0}, \mathbf{I}). In t-SNE visualizations, latents optimized without the Kullback-Leibler alignment loss LKL\mathcal{L}_{\text{KL}} form an out-of-domain cluster separate from randomly sampled Gaussian noise, which leads to severely degraded, distorted, and un-photorealistic image synthesis.

    Enforcing LKL=KL(N(μ,σ2) ∥ N(0,I))\mathcal{L}_{\text{KL}} = \text{KL}\left( \mathcal{N}(\mu, \sigma^2) \,\parallel\, \mathcal{N}(\mathbf{0}, \mathbf{I}) \right) with λ3=500\lambda_3 = 500 ensures the optimized noise latent z^T\hat{\mathbf{z}}_T remains strictly within the nominal prior distribution of the latent diffusion model while satisfying semantic attention constraints.

  8. Knowl 8 — Human Preference and Computational Overhead of INITNO

    data/table

    In a user evaluation study where 12 participants with image processing expertise evaluated 20 questions each, choosing the most visually appealing and semantically faithful generated image among competing training-free T2I methods, INITNO achieved a 63.33%63.33\% overall selection preference.

    Method User Study Preference (%)
    Stable Diffusion (CVPR 2022) 4.17%
    Composable Diffusion (ECCV 2022) 2.50%
    Structure Diffusion (ICLR 2023) 3.33%
    Attend-and-Excite (SIGGRAPH 2023) 14.17%
    Divide-and-Bind (BMVC 2023) 6.67%
    A-STAR (ICCV 2023) 5.83%
    INITNO (Ours) 63.33%

    In terms of runtime overhead evaluated on a single 32GB Tesla V100 GPU for 512×512512 \times 512 resolution generation (5050 denoising steps), baseline Stable Diffusion generates an image in an average of 8.348.34 seconds, while INITNO takes 18.9318.93 seconds due to initial latent search and optimization iterations.

  9. Knowl 9 — Plug-and-Play Integration with Layout-to-Image Generation (BoxDiff)

    model/method

    INITNO is compatible with spatial conditioning frameworks such as BoxDiff. Standard BoxDiff performs spatial-constraint updates across intermediate noisy latents at each denoising step tt. By integrating INITNO, spatial layout box constraints and attention alignment are transferred entirely to the initial latent space optimization pipeline at timestep TT.

    Optimizing the initial noise zT\mathbf{z}_T to satisfy spatial bounding box cross-attention localization prior to sequential denoising achieves training-free, location-accurate image synthesis with high semantic compliance.

Coverage note — No substantial contributed material was omitted; the full methodology, mathematical definitions, optimization algorithms, loss formulation, ablation insights, benchmark comparisons, and grounded extensions are fully represented.

References

  1. 1.Aishwarya Agarwal, Srikrishna Karanam, KJ Joseph, Apoorv Saxena, Koustava Goswami, and Balaji Vasan Srinivasan. A-star: Test-time attention segregation and retention for text-to-image synthesis. In ICCV, 2023. 2, 6, 7
  2. 2.Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022. 1, 2
  3. 3.Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. In NeurIPS, 2020. 5
  4. 4.Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In ICCV, 2023. 3
  5. 5.Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Murphy, William T Freeman, Michael Rubinstein, et al. Muse: Text-to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023. 2
  6. 6.Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, and Daniel Cohen-Or. Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. In SIGGRAPH, 2023. 2, 3, 4, 5, 6, 7, 8
  7. 7.Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In ICML, 2020. 1
  8. 8.Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. In NeurIPS, 2021. 1, 2
  9. 9.Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al. Cogview: Mastering text-to-image generation via transformers. In NeurIPS, 2021. 2
  10. 10.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In CVPR, 2021. 1
  11. 11.Mengyang Feng, Jinlin Liu, Kai Yu, Yuan Yao, Zheng Hui, Xiefan Guo, Xianhui Lin, Haolan Xue, Chen Shi, Xiaowen Li, et al. Dreamoving: A human dance video generation framework based on diffusion models. arXiv preprint arXiv:2312.05107, 2023. 1
  12. 12.Weixi Feng, Xuehai He, Tsu-Jui Fu, Varun Jampani, Arjun Reddy Akula, Pradyumna Narayana, Sugato Basu, Xin Eric Wang, and William Yang Wang. Training-free structured diffusion guidance for compositional text-to-image synthesis. In ICLR, 2023. 2, 6, 7
  13. 13.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. In NeurIPS, 2014. 1
  14. 14.Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Josh Susskind, and Navdeep Jaitly. Matryoshka diffusion models. arXiv preprint arXiv:2310.15111, 2023. 1, 2
  15. 15.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-or. Prompt-to-prompt image editing with cross-attention control. In ICLR, 2023. 3
  16. 16.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In NeurIPS, 2020. 1, 2, 3
  17. 17.Minguk Kang, Jun-Yan Zhu, Richard Zhang, Jaesik Park, Eli Shechtman, Sylvain Paris, and Taesung Park. Scaling up gans for text-to-image synthesis. In CVPR, 2023. 1, 2
  18. 18.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, 2019. 1
  19. 19.Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Training generative adversarial networks with limited data. In NeurIPS, 2020.
  20. 20.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In CVPR, 2020.
  21. 21.Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Alias-free generative adversarial networks. In NeurIPS, 2021. 1
  22. 22.Diederik P Kingma and Max Welling. Auto-encoding variational bayes. In ICLR, 2014. 1, 5
  23. 23.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In ICML, 2022. 6
  24. 24.Yumeng Li, Margret Keuper, Dan Zhang, and Anna Khoreva. Divide & bind your attention for improved generative semantic nursing. In BMVC, 2023. 2, 6, 7
  25. 25.Nan Liu, Shuang Li, Yilun Du, Antonio Torralba, and Joshua B Tenenbaum. Compositional visual generation with composable diffusion models. In ECCV, 2022. 2, 6, 7
  26. 26.Gaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li, Jingwan Lu, and Jun-Yan Zhu. Zero-shot image-to-image translation. In SIGGRAPH, 2023. 3
  27. 27.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 3
  28. 28.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020. 2
  29. 29.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In ICML, 2021. 2
  30. 30.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 1, 2
  31. 31.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 2, 3, 6, 7
  32. 32.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. In NeurIPS, 2022. 1, 2
  33. 33.Eyal Segalis, Dani Valevski, Danny Lumen, Yossi Matias, and Yaniv Leviathan. A picture is worth a thousand words: Principled recaptioning improves image generation. arXiv preprint arXiv:2310.16656, 2023. 2
  34. 34.Ming Tao, Hao Tang, Fei Wu, Xiao-Yuan Jing, Bing-Kun Bao, and Changsheng Xu. Df-gan: A simple and effective baseline for text-to-image synthesis. In CVPR, 2022. 2
  35. 35.Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. In CVPR, 2023. 3
  36. 36.Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al. Conditional image generation with pixelcnn decoders. In NeurIPS, 2016. 1
  37. 37.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NeurIPS, 2017. 1
  38. 38.Jinheng Xie, Yuexiang Li, Yawen Huang, Haozhe Liu, Wentian Zhang, Yefeng Zheng, and Mike Zheng Shou. Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion. In ICCV, 2023. 7
  39. 39.Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In CVPR, 2018. 2
  40. 40.Zeyue Xue, Guanglu Song, Qiushan Guo, Boxiao Liu, Zhuofan Zong, Yu Liu, and Ping Luo. Raphael: Text-to-image generation via large mixture of diffusion paths. In NeurIPS, 2023. 1, 2
  41. 41.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. TMLR, 2022. 1, 2
  42. 42.Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In ICCV, 2017. 2
  43. 43.Han Zhang, Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. Cross-modal contrastive learning for text-to-image generation. In CVPR, 2021.
  44. 44.Minfeng Zhu, Pingbo Pan, Wei Chen, and Yi Yang. Dm-gan: Dynamic memory generative adversarial networks for text-to-image synthesis. In CVPR, 2019. 2

Citation

MLA
Guo, X., et al. “InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization”. arXiv, 2024, http://arxiv.org/abs/2404.04650v1.
APA
Guo, X., Liu, J., Cui, M., Li, J., Yang, H., & Huang, D. (2024). InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization. arXiv. http://arxiv.org/abs/2404.04650v1
Chicago
Guo, X., J. Liu, M. Cui, J. Li, H. Yang, and D. Huang. 2024. “InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization”. arXiv. http://arxiv.org/abs/2404.04650v1.
Harvard
Guo, X. et al. (2024) “InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2404.04650v1.
Vancouver
1. Guo X, Liu J, Cui M, Li J, Yang H, Huang D (2024) InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization. arXiv

BibTeX

@article{guo2024initno,
  title = {InitNO: Boosting Text-to-Image Diffusion Models via Initial Noise Optimization},
  author = {Guo, Xiefan and Liu, Jinlin and Cui, Miaomiao and Li, Jiankai and Yang, Hongyu and Huang, Di},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2404.04650v1},
  eprint = {2404.04650}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE