High-Fidelity Guided Image Synthesis with Latent Diffusion Models

Jaskirat SinghStephen GouldLiang Zheng

article2023CVPR61 citations

Proposes an optimization framework and cross-attention mechanism that enable precise, region-specific semantic control and photorealistic image synthesis from coarse user scribbles without requiring additional model training or fine-tuning.

Listen

Generating realistic, customized images from user sketches and text prompts has become an essential capability in modern artificial intelligence, yet existing diffusion-based systems face significant limitations. Prior methods either require massive paired segmentation datasets for conditional training or rely on inversion techniques that suffer from an intrinsic domain shift problem, which produces overly simplistic, cartoonish, and blurry images lacking photorealistic detail. While iterative refinement techniques can restore some realism, they dramatically slow down processing times and degrade visual faithfulness to the original sketch.

The article demonstrates a guided image synthesis framework that generates high-fidelity, realistic images matching both text prompts and rough user sketches in a single reverse diffusion pass, without requiring specialized training data or model fine-tuning. It models the generation task as a constrained optimization problem, simultaneously ensuring that the synthesized image aligns with the target domain indicated by the prompt and preserves the structure defined by the user's color sketch.

To overcome the computational bottleneck of solving this constrained problem directly, the authors introduce two gradient-based approximation techniques called GradOP and its enhanced single-pass variant, GradOP+. The method injects optimization gradients directly into the latent reverse diffusion process, steering generation toward the reference painting while relying on forward diffusion steps to keep the intermediate representations within the realistic data distribution. Furthermore, the framework introduces cross-attention control between the text tokens and painted regions, allowing users to explicitly assign semantic identities to specific painted areas without additional model retraining.

Quantitative and qualitative assessments confirm substantial improvements over prior leading methods. The proposed method achieves strong realism scores, with a Fisher Inception Distance of 134.2 compared to 223.8 for standard SDEdit baselines, while retaining high faithfulness to the input sketch. In human user studies, the proposed approach surpassed existing state-of-the-art methods by over 85.32% in overall user satisfaction, achieving a 94.09% user preference on realism against standard diffusion inversion and 85.32% against iterative loopback methods.

These findings indicate that organizations can achieve highly controlled, studio-grade image synthesis using standard, pre-trained latent diffusion models without incurring expensive computational retraining costs or dataset collection pipelines. By eliminating multi-pass sampling requirements, the approach lowers computational costs, enhances workflow responsiveness, and prevents the loss of structural control seen in iterative tools.

Teams developing interactive design, media generation, or visual prototyping systems should adopt single-pass gradient-guided diffusion workflows to maximize rendering quality and user satisfaction. When deployed, user interfaces should allow region-specific cross-attention labeling alongside sketch tools to give creators precise semantic control over visual scenes. Future work should focus on extending these gradient-guidance pipelines to complex out-of-distribution prompts, as the evaluation noted occasional failures when handling highly unusual, compositional scenarios, such as depicting a rat chasing a lion.

arXiv: 2211.17084
Cover for High-Fidelity Guided Image Synthesis with Latent Diffusion Models

Abstract

Controllable image synthesis with user scribbles has gained huge public interest with the recent advent of text-conditioned latent diffusion models. The user scribbles control the color composition while the text prompt provides control over the overall image semantics. However, we note that prior works in this direction suffer from an intrinsic domain shift problem wherein the generated outputs often lack details and resemble simplistic representations of the target domain. In this paper, we propose a novel guided image synthesis framework, which addresses this problem by modelling the output image as the solution of a constrained optimization problem. We show that while computing an exact solution to the optimization is infeasible, an approximation of the same can be achieved while just requiring a single pass of the reverse diffusion process. Additionally, we show that by simply defining a cross-attention based correspondence between the input text tokens and the user stroke-painting, the user is also able to control the semantics of different painted regions without requiring any conditional training or finetuning. Human user study results show that the proposed approach outperforms the previous state-of-the-art by over 85.32% on the overall user satisfaction scores. Project page for our paper is available at https://ljsingh.github.io/gradop.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Our Method
  • 3.1. GradOP: Obtaining an Approximate Solution
  • 3.2. GradOP+: Improving Sampling Efficiency
  • 3.3. Controlling Semantics of Painted Regions
  • 4. Experiments
  • 4.1. Stroke Guided Image Synthesis
  • 4.2. Controlling Semantics of Painted Regions
  • 5. Analysis
  • 5.1. Variation in Target Domain
  • 5.2. Variation with Number of Gradient Steps
  • 5.3. Out-of-Distribution Generalization
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — Constrained formulation for faithful, text-domain-consistent synthesis

    model/method

    The paper models guided synthesis from a colored stroke painting as a constrained optimization problem. Let y∈Dpainty\in\mathcal D_{\mathrm{paint}} be the reference painting, let τtext\tau_{\mathrm{text}} be the text prompt, and let f:Dreal→Dpaintf:\mathcal D_{\mathrm{real}}\rightarrow\mathcal D_{\mathrm{paint}} be an autonomous painting function that maps a generated image to its painted representation. Let Sτtext\mathcal S_{\tau_{\mathrm{text}}} denote the subspace of images that the text-conditioned diffusion model can generate using only τtext\tau_{\mathrm{text}}, and let L\mathcal L measure the discrepancy between two paintings. The desired image is

    x⋆=argmin⁡x  L(f(x),y)subject tox∈Sτtext.x^\star=\underset{x}{\operatorname{argmin}}\;\mathcal L(f(x),y) \quad\text{subject to}\quad x\in\mathcal S_{\tau_{\mathrm{text}}}.

    This objective requires the generated image to recover the user’s stroke composition after painting while remaining in the target domain specified by the prompt; for example, a prompt requesting a realistic photograph should not produce a cartoon-like rendering merely because the input painting is simplistic.

  2. Knowl 2 — Latent-space approximation used by GradOP

    equation

    Because directly measuring distance to the entire text-conditioned image subspace Sτtext\mathcal S_{\tau_{\mathrm{text}}} would require many text-only samples, the paper approximates that subspace with one randomly sampled text-conditioned image xτtext∈Sτtextx_{\tau_{\mathrm{text}}}\in\mathcal S_{\tau_{\mathrm{text}}}. With latent-diffusion encoder EE, decoder DD, latent variable zz, reference latent zτtext=E(xτtext)z_{\tau_{\mathrm{text}}}=E(x_{\tau_{\mathrm{text}}}), and nonnegative tradeoff parameter γ\gamma, GradOP optimizes

    z⋆=argmin⁡z  L(f(D(z)),y)+γ∥z−zτtext∥2,x⋆=D(z⋆).z^\star=\underset{z}{\operatorname{argmin}}\;\mathcal L\bigl(f(D(z)),y\bigr)+\gamma\lVert z-z_{\tau_{\mathrm{text}}}\rVert_2, \qquad x^\star=D(z^\star).

    The first term preserves faithfulness to the stroke painting, while the latent proximity term keeps the solution near a text-only diffusion sample. Since the unconstrained latent optimum is not guaranteed to remain on the text-conditioned diffusion manifold, the decoded result is subsequently passed through a diffusion-based forward-then-reverse inversion procedure to map it back toward Sτtext\mathcal S_{\tau_{\mathrm{text}}}. The paper presents this as an approximation rather than an exact solution; computing the exact constrained optimum is considered infeasible.

  3. Knowl 3 — GradOP solution-approximation procedure

    algorithm

    GradOP takes a stroke painting yy, text prompt τtext\tau_{\mathrm{text}}, differentiable painting function ff, differentiable painting distance L\mathcal L, tradeoff weight γ\gamma, gradient step size λ\lambda, optimization-step count MM, and diffusion timestep t0t_0. It returns a generated image whose painted output is close to yy while its appearance is restored toward the text-conditioned domain.

    Input: Stroke painting yy, text prompt τtext\tau_{\mathrm{text}}
    Require: Differentiable ff and L\mathcal L, γ\gamma, step size λ\lambda, timestep t0t_0, and MM
    Sample a text-only image xτtextx_{\tau_{\mathrm{text}}} from the diffusion model
    Set zτtext=E(xτtext)z_{\tau_{\mathrm{text}}}=E(x_{\tau_{\mathrm{text}}}) and initialize z=zτtextz=z_{\tau_{\mathrm{text}}}
    for i=0,1,…,Mi=0,1,\ldots,M do
        Set Ltotal=L(f(D(z)),y)+γ∥z−zτtext∥2\mathcal L_{\mathrm{total}}=\mathcal L(f(D(z)),y)+\gamma\lVert z-z_{\tau_{\mathrm{text}}}\rVert_2
        Update z=z−λ∇zLtotalz=z-\lambda\nabla_z\mathcal L_{\mathrm{total}}
    end for
    Set zt0=FORWARDDIFF⁡(z,0→t0)z_{t_0}=\operatorname{FORWARDDIFF}(z,0\mathbin{\to}t_0)
    Set z=REVERSEDIFF⁡(zt0,t0→0)z=\operatorname{REVERSEDIFF}(z_{t_0},t_0\mathbin{\to}0)
    return xout=D(z)x_{\mathrm{out}}=D(z)

    The optimization first moves a text-conditioned latent toward a painting-faithful point, unlike methods that directly diffuse the stroke painting despite its potentially large domain gap from natural images. In the experiments, gradient descent is implemented with Adam and typically uses 2020–6060 optimization steps.

  4. Knowl 4 — GradOP+ gradient injection during one reverse diffusion pass

    algorithm

    GradOP+ avoids separately sampling a complete text-only image and then running a second diffusion restoration. It begins with Gaussian noise zT∼N(0,I)z_T\sim\mathcal N(0,I) and performs one reverse diffusion trajectory. At selected timesteps t∈[tstart,tend]t\in[t_{\mathrm{start}},t_{\mathrm{end}}], it optimizes the current latent using

    zt⋆=argmin⁡z  L(f(D(z)),y)+γ∥z−zt∥2,z_t^\star=\underset{z}{\operatorname{argmin}}\;\mathcal L\bigl(f(D(z)),y\bigr)+\gamma\lVert z-z_t\rVert_2,

    then forward-diffuses the optimized latent back to timestep tt so that it approximately matches the expected latent distribution at that timestep. Here TT is the initial diffusion timestep, ztz_t is the current reverse-diffusion latent, and zt⋆z_t^\star is the locally optimized latent.

    Input: Stroke painting yy, text prompt τtext\tau_{\mathrm{text}}
    Require: Differentiable ff and L\mathcal L, γ\gamma, step size λ\lambda, MM, TT, tstartt_{\mathrm{start}}, and tendt_{\mathrm{end}}
    Sample zT∼N(0,I)z_T\sim\mathcal N(0,I)
    for t=T−1,T−2,…,0t=T-1,T-2,\ldots,0 do
        Set zt=REVERSEDIFF⁡(zt+1,t+1→t)z_t=\operatorname{REVERSEDIFF}(z_{t+1},t+1\mathbin{\to}t)
        if tstart≤t≤tendt_{\mathrm{start}}\leq t\leq t_{\mathrm{end}} then
            Initialize z=ztz=z_t
            for i=0,1,…,Mi=0,1,\ldots,M do
                Set Ltotal=L(f(D(z)),y)+γ∥z−zt∥2\mathcal L_{\mathrm{total}}=\mathcal L(f(D(z)),y)+\gamma\lVert z-z_t\rVert_2
                Update z=z−λ∇zLtotalz=z-\lambda\nabla_z\mathcal L_{\mathrm{total}}
            end for
            Set zt=FORWARDDIFF⁡(z,0→t)z_t=\operatorname{FORWARDDIFF}(z,0\mathbin{\to}t)
        end if
    end for
    return xout=D(z0)x_{\mathrm{out}}=D(z_0)

    The forward-diffusion correction prevents the painting-gradient update from moving the latent away from the distribution expected by the reverse diffusion model. The method therefore injects painting-recovery optimization into a single reverse diffusion pass rather than repeatedly regenerating an output.

  5. Knowl 5 — Cross-attention control of painted-region semantics

    model/method

    To let users assign explicit meanings to painted regions, the paper uses binary masks and cross-attention rather than semantic-segmentation training. Let {B1,…,BN}\{B_1,\ldots,B_N\} be binary masks for painted regions, let {u1,…,uN}\{u_1,\ldots,u_N\} be their desired semantic labels, and let τ\tau be the original prompt token set. The prompt is augmented with CLIP representations of the labels:

    τmodified=τ+{CLIP⁡(ui)∣i∈{1,…,N}}.\tau_{\mathrm{modified}}=\tau+\{\operatorname{CLIP}(u_i)\mid i\in\{1,\ldots,N\}\}.

    At reverse-diffusion timestep tt, let AtiA_t^i be the cross-attention map for label uiu_i, with the same spatial support as mask BiB_i. The map is replaced by

    A~ti=wi[(1−κt)Ati+κtBi∥Bi∥F∥Ati∥F],κt=tT∈[0,1],\widetilde A_t^i=w_i\left[(1-\kappa_t)A_t^i+\kappa_t\frac{B_i}{\lVert B_i\rVert_F}\lVert A_t^i\rVert_F\right], \qquad \kappa_t=\frac{t}{T}\in[0,1],

    where TT is the total diffusion horizon, ∥⋅∥F\lVert\cdot\rVert_F is the Frobenius norm, and wiw_i controls the relative strength of semantic concept uiu_i. The interpolation progressively encourages the label’s attention to overlap the user-specified region during reverse diffusion. This provides region-level semantic control without paired segmentation data, conditional retraining, or model fine-tuning.

  6. Knowl 6 — Evaluation measures and experimental configuration

    experimental setup

    The evaluation uses unpaired stroke-guided synthesis: each output image is conditioned on a stroke painting yy and text prompt τtext\tau_{\mathrm{text}}, while no paired ground-truth photograph is available. Faithfulness of an output image xx is measured by the painting discrepancy

    F(x,y)=L2(f(x),y),\mathcal F(x,y)=L_2(f(x),y),

    where f(x)f(x) is the painted output and lower values indicate closer recovery of the reference painting. Let S(y,τtext)\mathcal S(y,\tau_{\mathrm{text}}) be generated samples conditioned on both painting and text, and let S(τtext)\mathcal S(\tau_{\mathrm{text}}) be text-only samples. Target-domain realism is measured by

    R(S(y,τtext))=FID⁡(S(y,τtext),S(τtext)),\mathcal R\bigl(\mathcal S(y,\tau_{\mathrm{text}})\bigr)=\operatorname{FID}\bigl(\mathcal S(y,\tau_{\mathrm{text}}),\mathcal S(\tau_{\mathrm{text}})\bigr),

    so a lower FID means that painting-guided outputs are closer to the text-only target distribution. The comparisons use SDEdit, SDEdit with Loopback, and latent-diffusion-adapted ILVR. The proposed method uses GradOP+ for quantitative evaluation unless otherwise stated. Public text-conditioned latent diffusion models are used; Adam performs the latent optimization with Ngrad∈[20,60]N_{\mathrm{grad}}\in[20,60] steps. A mean-squared distance and Gaussian-convolution painting function are used for fast differentiable optimization, whereas quantitative comparisons use the nondifferentiable SDEdit painting function for consistency. The standard diffusion noise level is t0=0.8t_0=0.8 for SDEdit comparisons.

  7. Knowl 7 — Quantitative faithfulness, realism, and user preference results

    data/table

    The comparison measures the tradeoff between recovering the stroke composition and matching the text-only target domain. The first metric is painting faithfulness F(x,y)\mathcal F(x,y), the second is target-domain FID R\mathcal R, and the last two columns report the percentages of human comparisons in which the proposed approach is preferred over the named baseline for realism and overall satisfaction.

    Method F(x,y)↓\mathcal F(x,y)\downarrow R(⋅)↓\mathcal R(\cdot)\downarrow Realism↑\uparrow Satisfaction↑\uparrow
    SDEdit 88.93 223.8 94.09% 91.98%
    Loopback 104.6 132.9 54.28% 85.32%
    ILVR 108.2 161.7 76.54% 93.47%
    Ours 94.40 134.2 N/A N/A

    SDEdit has the lowest faithfulness error but the worst target-domain FID, indicating that direct inversion preserves strokes at the cost of realism. Loopback has the best FID by a small margin but substantially worse faithfulness than SDEdit and the proposed method. The proposed method obtains a substantially better realism–faithfulness balance than the baselines, and the reported overall-satisfaction preference over prior methods is at least 85.32%85.32\%.

  8. Knowl 8 — Qualitative realism–faithfulness tradeoff

    empirical result

    Across landscape, animal, tree, and other stroke-guided examples, GradOP and GradOP+ produce outputs that retain the reference painting’s global color composition while adding details characteristic of the text-specified target domain. Direct SDEdit outputs remain close to the painting but often look like simplified pictorial art rather than realistic photographs. Loopback improves visual realism by repeatedly reapplying guided synthesis, but each additional iteration increases generation time and progressively reduces faithfulness to the original painting. ILVR produces realistic images but can fail to preserve the reference’s overall color arrangement. The proposed optimization is therefore reported to provide the strongest qualitative compromise between stroke faithfulness and target-domain realism.

  9. Knowl 9 — Semantic-control behavior and region-level results

    empirical result

    Without explicit attention control, the diffusion model infers painted-region meanings from its learned priors, so the same blue region may become a river, waterfall, or valley, while brown strokes intended as a hut or castle may become muddy or rocky terrain. Merely appending the desired object word to the text prompt can produce an object, but its placement and interpretation may still disagree with the user’s painted region. Applying the cross-attention intervention causes the requested concept to appear in the corresponding mask: brown ground regions can be made into huts or castles, blue regions into rivers or waterfalls, and orange sky regions into suns or moons. These results demonstrate that the semantic mapping can be controlled without adding segmentation-conditioned training or fine-tuning.

  10. Knowl 10 — Target-domain and optimization-step analysis

    empirical result

    The method is qualitatively evaluated across target domains including realistic photographs, generic paintings, Van Gogh-style drawings, children’s watercolor paintings, and Disney-like scenes. It adapts the generated appearance to the requested domain while maintaining strong correspondence with the stroke reference. SDEdit tends to produce similarly simplified outputs across domains, whereas Loopback requires multiple reverse-diffusion passes and loses reference faithfulness as iterations accumulate.

    The number of latent optimization steps NgradN_{\mathrm{grad}} controls the balance within the proposed method. With Ngrad=0N_{\mathrm{grad}}=0, the output is effectively a random sample from the text-conditioned subspace and does not use the painting strongly. Increasing the steps through the tested values 2020, 4040, and 6060 moves the result toward a subset of text-conditioned solutions with greater faithfulness to the reference while retaining target-domain realism. This differs from SDEdit’s reported behavior, where stronger reference adherence is associated with reduced realism.

  11. Knowl 11 — Out-of-distribution limitation

    limitation

    The method does not reliably solve every out-of-distribution text prompt. In the reported examples, it generates realistic-looking photographs for the unusual prompt “a photo of a cat with six legs,” whereas SDEdit and Loopback tend either to produce cartoon-like outputs or ordinary cats. However, it performs poorly for “a photo of a rat chasing a lion.” Thus, the framework can preserve realism and provide semantic control for some prompts outside the training distribution, but success is prompt-dependent and is not guaranteed.

Coverage note — No substantial contributed material was omitted; background, related work, references, and acknowledgements were excluded.

References

  1. 1.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2stylegan: How to embed images into the stylegan latent space? In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4432–4441, 2019. 2
  2. 2.Rameen Abdal, Yipeng Qin, and Peter Wonka. Image2stylegan++: How to edit the embedded images? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8296–8305, 2020. 2
  3. 3.Yuval Alaluf, Or Patashnik, and Daniel Cohen-Or. Restyle: A residual-based stylegan encoder via iterative refinement. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2021. 2
  4. 4.AUTOMATIC1111. Stable-diffusion-webui. https : / / github . com / AUTOMATIC1111 / stable - diffusion-webui, 2022. 2, 5, 6, 7, 8
  5. 5.Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. Ilvr: Conditioning method for denoising diffusion probabilistic models. arXiv preprint arXiv:2108.02938, 2021. 2, 4, 5, 6, 8
  6. 6.Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al. Cogview: Mastering text-to-image generation via transformers. Advances in Neural Information Processing Systems, 34:19822–19835, 2021. 2
  7. 7.Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12873–12883, 2021. 2
  8. 8.Oran Gafni, Adam Polyak, Oron Ashual, Shelly Sheynin, Devi Parikh, and Yaniv Taigman. Make-a-scene: Scene-based text-to-image generation with human priors. arXiv preprint arXiv:2203.13131, 2022. 2
  9. 9.Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022. 2
  10. 10.Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. Prompt-to-prompt image editing with cross attention control. arXiv preprint arXiv:2208.01626, 2022. 2
  11. 11.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017. 5
  12. 12.Nisha Huang, Fan Tang, Weiming Dong, and Changsheng Xu. Draw your art dream: Diverse digital art synthesis with multimodal guided diffusion. In Proceedings of the 30th ACM International Conference on Multimedia, pages 1085–1094, 2022. 2
  13. 13.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1125–1134, 2017. 2
  14. 14.Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. Training generative adversarial networks with limited data. Advances in Neural Information Processing Systems, 33:12104–12114, 2020. 2
  15. 15.Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. Diffusionclip: Text-guided diffusion models for robust image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2426–2435, 2022. 2, 5
  16. 16.Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014. 5
  17. 17.Gihyun Kwon and Jong Chul Ye. Diffusion-based image translation using disentangled style and content representation. arXiv preprint arXiv:2209.15264, 2022. 2
  18. 18.Cheng-Han Lee, Ziwei Liu, Lingyun Wu, and Ping Luo. Maskgan: Towards diverse and interactive facial image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5549–5558, 2020. 2
  19. 19.Songhua Liu, Tianwei Lin, Dongliang He, Fu Li, Ruifeng Deng, Xin Li, Errui Ding, and Hao Wang. Paint transformer: Feed forward neural painting with stroke prediction. arXiv preprint arXiv:2108.03798, 2021. 2
  20. 20.Xihui Liu, Dong Huk Park, Samaneh Azadi, Gong Zhang, Arman Chopikyan, Yuxiao Hu, Humphrey Shi, Anna Rohrbach, and Trevor Darrell. More control for free! image synthesis with semantic diffusion guidance. arXiv preprint arXiv:2112.05744, 2021. 2
  21. 21.Xihui Liu, Guojun Yin, Jing Shao, Xiaogang Wang, et al. Learning to predict layout-to-image conditional convolutions for semantic image synthesis. Advances in Neural Information Processing Systems, 32, 2019. 2
  22. 22.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. SDEdit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022. 2, 4, 5, 6, 7, 8
  23. 23.Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741, 2021. 1, 2
  24. 24.Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. Semantic image synthesis with spatially-adaptive normalization. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2337–2346, 2019. 2
  25. 25.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning, pages 8748–8763. PMLR, 2021. 5
  26. 26.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022. 1, 2
  27. 27.Elad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan, Yaniv Azar, Stav Shapiro, and Daniel Cohen-Or. Encoding in style: a stylegan encoder for image-to-image translation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2021. 2
  28. 28.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjorn Ommer. High-resolution image synthesis with latent diffusion models, 2021. 1, 2, 5
  29. 29.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv preprint arXiv:2208.12242, 2022. 2
  30. 30.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022. 1, 2
  31. 31.Junyoung Seo, Gyuseong Lee, Seokju Cho, Jiyoung Lee, and Seungryong Kim. Midms: Matching interleaved diffusion models for exemplar-based image translation. arXiv preprint arXiv:2209.11047, 2022. 2
  32. 32.Jaskirat Singh, Cameron Smith, Jose Echevarria, and Liang Zheng. Intelli-paint: Towards developing human-like painting agents. In European conference on computer vision. Springer, 2022. 2
  33. 33.Jaskirat Singh and Liang Zheng. Combining semantic guidance and deep reinforcement learning for generating human level paintings. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021. 2
  34. 34.Jaskirat Singh, Liang Zheng, Cameron Smith, and Jose Echevarria. Paint2pix: Interactive painting based progressive image synthesis and editing. In European conference on computer vision. Springer, 2022. 2
  35. 35.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020. 2, 3, 5
  36. 36.Vadim Sushko, Edgar Schonfeld, Dan Zhang, Juergen Gall, Bernt Schiele, and Anna Khoreva. You only need adversarial supervision for semantic image synthesis. arXiv preprint arXiv:2012.04781, 2020. 2
  37. 37.Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or. Designing an encoder for stylegan image manipulation. arXiv preprint arXiv:2102.02766, 2021. 2
  38. 38.Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf. Diffusers: State-of-the-art diffusion models. https://github.com/huggingface/diffusers, 2022. 5
  39. 39.Weilun Wang, Jianmin Bao, Wengang Zhou, Dongdong Chen, Dong Chen, Lu Yuan, and Houqiang Li. Semantic image synthesis via diffusion models. arXiv preprint arXiv:2207.00050, 2022. 2
  40. 40.Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.10789, 2022. 1, 2
  41. 41.Yufan Zhou, Ruiyi Zhang, Changyou Chen, Chunyuan Li, Chris Tensmeyer, Tong Yu, Jiuxiang Gu, Jinhui Xu, and Tong Sun. Towards language-free training for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17907–17917, 2022. 2
  42. 42.Jun-Yan Zhu, Philipp Krahenbühl, Eli Shechtman, and Alexei A Efros. Generative visual manipulation on the natural image manifold. In European conference on computer vision, pages 597–613. Springer, 2016. 2
  43. 43.Peihao Zhu, Rameen Abdal, Yipeng Qin, and Peter Wonka. Sean: Image synthesis with semantic region-adaptive normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5104–5113, 2020. 2
  44. 44.Zhengxia Zou, Tianyang Shi, Shuang Qiu, Yi Yuan, and Zhenwei Shi. Stylized neural painting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15689–15698, 2021. 2

Citation

MLA
Singh, J., et al. “High-Fidelity Guided Image Synthesis with Latent Diffusion Models”. arXiv, 2022, http://arxiv.org/abs/2211.17084v1.
APA
Singh, J., Gould, S., & Zheng, L. (2022). High-Fidelity Guided Image Synthesis with Latent Diffusion Models. arXiv. http://arxiv.org/abs/2211.17084v1
Chicago
Singh, J., S. Gould, and L. Zheng. 2022. “High-Fidelity Guided Image Synthesis with Latent Diffusion Models”. arXiv. http://arxiv.org/abs/2211.17084v1.
Harvard
Singh, J., Gould, S. and Zheng, L. (2022) “High-Fidelity Guided Image Synthesis with Latent Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2211.17084v1.
Vancouver
1. Singh J, Gould S, Zheng L (2022) High-Fidelity Guided Image Synthesis with Latent Diffusion Models. arXiv

BibTeX

@article{singh2022high,
  title = {High-Fidelity Guided Image Synthesis with Latent Diffusion Models},
  author = {Singh, Jaskirat and Gould, Stephen and Zheng, Liang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2211.17084v1},
  eprint = {2211.17084}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE