Diffusion Models in Vision: A Survey

Florinel-Alin CroitoruVlad HondruRadu Tudor IonescuMubarak Shah

article2022TPAMI2,338 citations

Systematizes the theoretical foundations of diffusion models across probabilistic, score-based, and stochastic differential equation frameworks while analyzing their vision applications, generative trade-offs, and computational bottlenecks.

Listen

Deep generative models have seen rapid growth in artificial intelligence, with diffusion models emerging as a leading approach for synthesizing high-fidelity visual content. Understanding the theoretical foundations, operational trade-offs, and practical capabilities of these models is critical for leaders evaluating next-generation visual AI deployment. The article provides a comprehensive review and multi-perspective taxonomy of denoising diffusion models in computer vision, analyzing their mathematical formulations, structural connections to other generative paradigms, and applications across diverse vision tasks.

The article conducts a systematic literature review and comparative analysis of the diffusion modeling landscape. It formalizes three generic frameworks: Denoising Diffusion Probabilistic Models (inspired by non-equilibrium thermodynamics), Noise Conditioned Score Networks (trained via score matching to estimate data gradients), and continuous Stochastic Differential Equations, which generalize the first two formulations. The authors evaluate how these frameworks map to a wide spectrum of tasks—including unconditional image generation, conditional and text-to-image synthesis, super-resolution, inpainting, segmentation, medical imaging, and video generation—while analyzing the structural trade-offs against conventional generative models such as generative adversarial networks, variational auto-encoders, and normalizing flows.

The review establishes several core findings. First, diffusion models have matched and in many domains surpassed generative adversarial networks in sample quality, fine detail, and output diversity, notably avoiding common training instabilities like mode collapse. Second, the fundamental operating principle across all frameworks consists of a two-phase process: a forward stage that gradually degrades data into standard Gaussian noise, and a parameterized reverse stage where neural networks learn to iteratively remove the noise. Third, text-to-image models (such as Latent Diffusion Models and Imagen) exhibit high generalization capabilities, synthesizing complex, out-of-distribution concepts with minimal visual artifacts. Fourth, the primary operational bottleneck of diffusion models is poor inference speed, historically requiring hundreds to thousands of iterative network evaluations to produce a single image.

These findings have direct operational and commercial implications. The high visual quality and training stability make diffusion models dependable for critical applications, including medical anomaly detection, image super-resolution, and artistic generation. However, the high computational cost and latency at inference time present a significant barrier for real-time applications and low-latency production pipelines compared to single-step generative methods. Furthermore, text-conditioned models inherit specific failure modes from external components, such as spelling errors caused by image-text embedding encoders lacking granular character representations.

To overcome these limitations, organizations and researchers should pursue algorithmic and architectural optimizations. Key recommended strategies include adopting accelerated differential equation solvers, distilling knowledge into student networks with significantly fewer sampling steps (e.g., reducing steps by factors of 20 to 40 or distilling down to single-digit steps), and executing diffusion processes in lower-dimensional latent spaces rather than raw pixel spaces. Future development should also focus on extending diffusion representations to discriminative tasks, long-term video synthesis, and unified multi-task architectures. The findings are backed by an extensive body of recent empirical research, though practitioners should remain cautious regarding inference latency constraints and domain-specific conditioning limits.

arXiv: 2209.04747
Cover for Diffusion Models in Vision: A Survey

Abstract

Denoising diffusion models represent a recent emerging topic in computer vision, demonstrating remarkable results in the area of generative modeling. A diffusion model is a deep generative model that is based on two stages, a forward diffusion stage and a reverse diffusion stage. In the forward diffusion stage, the input data is gradually perturbed over several steps by adding Gaussian noise. In the reverse stage, a model is tasked at recovering the original input data by learning to gradually reverse the diffusion process, step by step. Diffusion models are widely appreciated for the quality and diversity of the generated samples, despite their known computational burdens, i.e. low speeds due to the high number of steps involved during sampling. In this survey, we provide a comprehensive review of articles on denoising diffusion models applied in vision, comprising both theoretical and practical contributions in the field. First, we identify and present three generic diffusion modeling frameworks, which are based on denoising diffusion probabilistic models, noise conditioned score networks, and stochastic differential equations. We further discuss the relations between diffusion models and other deep generative models, including variational auto-encoders, generative adversarial networks, energy-based models, autoregressive models and normalizing flows. Then, we introduce a multi-perspective categorization of diffusion models applied in computer vision. Finally, we illustrate the current limitations of diffusion models and envision some interesting directions for future research.

Table of Contents

  • I Introduction
  • II Generic Framework
  • II-A Denoising Diffusion Probabilistic Models (DDPMs)
  • II-B Noise Conditioned Score Networks (NCSNs)
  • II-C Stochastic Differential Equations (SDEs)
  • II-D Relation to Other Generative Models
  • III A Categorization of Diffusion Models
  • III-A Unconditional Image Generation
  • III-A1 Denoising Diffusion Probabilistic Models
  • III-A2 Score-Based Generative Models
  • III-A3 Stochastic Differential Equations
  • III-B Conditional Image Generation
  • III-B1 Denoising Diffusion Probabilistic Models
  • III-B2 Score-Based Generative Models
  • III-B3 Stochastic Differential Equations
  • III-C Image-to-Image Translation
  • III-D Text-to-Image Synthesis
  • III-E Image Super-Resolution
  • III-F Image Editing
  • III-G Image Inpainting
  • III-H Image Segmentation
  • III-I Multi-Task Approaches
  • III-J Medical Image Generation and Translation
  • III-K Anomaly Detection in Medical Images
  • III-L Video Generation
  • III-M Other Tasks
  • III-N Theoretical Contributions
  • IV Closing Remarks and Future Directions
  • References
  • A Variational bound.
  • B Noise estimation.

Knowls

  1. Knowl 1 — Three Generic Frameworks of Diffusion Modeling

    model/method

    Diffusion models are deep generative models structured around two complementary stages:

    1. Forward diffusion process: An input data point x0∼p(x0)x_0 \sim p(x_0) is progressively corrupted across time steps by adding noise (typically Gaussian), gradually destroying data structure until pure random noise is obtained.
    2. Reverse (backward) denoising process: A parameterized neural network (typically based on a U-Net architecture) is trained to iteratively invert the corruption trajectory, synthesizing new data samples starting from pure noise.

    Three foundational formulations define the field:

    • Denoising Diffusion Probabilistic Models (DDPMs): Discrete latent variable Markovian models grounded in non-equilibrium thermodynamics, optimizing variational bounds with simplified noise-prediction objectives.
    • Noise Conditioned Score Networks (NCSNs): Score-based models estimating the score function ∇xlog⁡p(x)\nabla_x \log p(x) across a geometric sequence of noise scales using denoising score matching and sampling via annealed Langevin dynamics.
    • Stochastic Differential Equations (SDEs): A continuous-time framework generalizing DDPMs and NCSNs that models data corruption via forward Itô SDEs and generation via reverse-time SDEs or deterministic probability flow ordinary differential equations (ODEs).
  2. Knowl 2 — Denoising Diffusion Probabilistic Model Forward and Reverse Formulation

    equation

    In Denoising Diffusion Probabilistic Models (DDPMs), the forward process is a discrete Markov chain corrupting data x0∼p(x0)x_0 \sim p(x_0) over TT steps with a fixed variance schedule β1,…,βT∈[0,1)\beta_1, \dots, \beta_T \in [0, 1):

    p(xt∣xt−1)=N(xt;1−βtxt−1,βtI),∀t∈{1,…,T}p(x_t | x_{t-1}) = \mathcal{N}\left(x_t; \sqrt{1 - \beta_t} x_{t-1}, \beta_t I\right), \quad \forall t \in \{1, \dots, T\}

    Defining αt=1−βt\alpha_t = 1 - \beta_t and cumulative product β^t=∏i=1tαi\hat{\beta}_t = \prod_{i=1}^t \alpha_i, the marginal distribution at any step tt conditioned on uncorrupted input x0x_0 is:

    p(xt∣x0)=N(xt;β^tx0,(1−β^t)I)p(x_t | x_0) = \mathcal{N}\left(x_t; \sqrt{\hat{\beta}_t} x_0, (1 - \hat{\beta}_t) I\right)

    This permits direct sampling of xtx_t via the reparameterization trick:

    xt=β^tx0+1−β^tzt,zt∼N(0,I)x_t = \sqrt{\hat{\beta}_t} x_0 + \sqrt{1 - \hat{\beta}_t} z_t, \quad z_t \sim \mathcal{N}(0, I)

    The reverse transition pθ(xt−1∣xt)=N(xt−1;μθ(xt,t),Σθ(xt,t))p_\theta(x_{t-1}|x_t) = \mathcal{N}(x_{t-1}; \mu_\theta(x_t, t), \Sigma_\theta(x_t, t)) fixes the covariance to Σθ(xt,t)=βtI\Sigma_\theta(x_t, t) = \beta_t I and parameterizes the mean as a function of the noise zθ(xt,t)z_\theta(x_t, t) predicted by a neural network:

    μθ(xt,t)=1αt(xt−1−αt1−β^tzθ(xt,t))\mu_\theta(x_t, t) = \frac{1}{\sqrt{\alpha_t}} \left( x_t - \frac{1 - \alpha_t}{\sqrt{1 - \hat{\beta}_t}} z_\theta(x_t, t) \right)

    Training minimizes the simplified mean squared error objective:

    Lsimple=Et∼[1,T],x0∼p(x0),zt∼N(0,I)[∥zt−zθ(xt,t)∥2]L_{\text{simple}} = \mathbb{E}_{t \sim [1, T], x_0 \sim p(x_0), z_t \sim \mathcal{N}(0, I)} \left[ \| z_t - z_\theta(x_t, t) \|^2 \right]

    where tt is drawn uniformly from {1,…,T}\{1, \dots, T\}.

  3. Knowl 3 — DDPM Iterative Sampling Method

    algorithm

    The DDPM sampling procedure generates an image x0x_0 starting from standard Gaussian white noise by iteratively applying learned reverse transitions.

    Input: Number of diffusion steps TT; transition standard deviations σ1,…,σT\sigma_1, \dots, \sigma_T.
    Output: Generated sample x0x_0.
    xT∼N(0,I)x_T \sim \mathcal{N}(0, I)
    for t=T,…,1t = T, \dots, 1 do
        if t>1t > 1 then
            z∼N(0,I)z \sim \mathcal{N}(0, I)
        else
            z=0z = 0
        μθ=1αt⋅(xt−1−αt1−β^t⋅zθ(xt,t))\mu_\theta = \frac{1}{\sqrt{\alpha_t}} \cdot \left( x_t - \frac{1 - \alpha_t}{\sqrt{1 - \hat{\beta}_t}} \cdot z_\theta(x_t, t) \right)
        xt−1=μθ+σt⋅zx_{t-1} = \mu_\theta + \sigma_t \cdot z
    return x0x_0
  4. Knowl 4 — Noise Conditioned Score Networks and Denoising Score Matching

    equation

    The score function of a probability density p(x)p(x) is defined as the gradient of the log density with respect to the input: ∇xlog⁡p(x)\nabla_x \log p(x). To prevent score estimation failure on low-dimensional manifolds, Noise Conditioned Score Networks (NCSNs) perturb data with a geometric progression of Gaussian noise scales σ1<σ2<⋯<σT\sigma_1 < \sigma_2 < \dots < \sigma_T, where σ1\sigma_1 satisfies pσ1(x)≈p(x0)p_{\sigma_1}(x) \approx p(x_0) and σT\sigma_T satisfies pσT(x)≈N(0,I)p_{\sigma_T}(x) \approx \mathcal{N}(0, I).

    The perturbation distribution is:

    pσt(xt∣x)=N(xt;x,σt2I)=1σt2πexp⁡(−12∥xt−xσt∥2)p_{\sigma_t}(x_t | x) = \mathcal{N}\left(x_t; x, \sigma_t^2 I\right) = \frac{1}{\sigma_t \sqrt{2\pi}} \exp\left(-\frac{1}{2} \left\| \frac{x_t - x}{\sigma_t} \right\|^2\right)

    with analytical gradient:

    ∇xtlog⁡pσt(xt∣x)=−xt−xσt2\nabla_{x_t} \log p_{\sigma_t}(x_t | x) = -\frac{x_t - x}{\sigma_t^2}

    A shared noise-conditioned score network sθ(x,σt)s_\theta(x, \sigma_t) is trained across all noise scales by minimizing the denoising score matching loss:

    Ldsm=1T∑t=1Tλ(σt)Ep(x)Ext∼pσt(xt∣x)[∥sθ(xt,σt)+xt−xσt2∥22]L_{\text{dsm}} = \frac{1}{T} \sum_{t=1}^T \lambda(\sigma_t) \mathbb{E}_{p(x)} \mathbb{E}_{x_t \sim p_{\sigma_t}(x_t | x)} \left[ \left\| s_\theta(x_t, \sigma_t) + \frac{x_t - x}{\sigma_t^2} \right\|_2^2 \right]

    where λ(σt)\lambda(\sigma_t) is a positive weighting function.

  5. Knowl 5 — Annealed Langevin Dynamics Sampling Algorithm

    algorithm

    Annealed Langevin dynamics synthesizes samples from a trained noise conditioned score network by running recursive Langevin steps at decreasing noise scales.

    Input: Sequence of Gaussian noise scales σ1,…,σT\sigma_1, \dots, \sigma_T; number of iterations per scale NN; step sizes γ1,…,γT\gamma_1, \dots, \gamma_T.
    Output: Generated sample x00x_0^0.
    xT0∼N(0,I)x_T^0 \sim \mathcal{N}(0, I)
    for t=T,…,1t = T, \dots, 1 do
        for i=1,…,Ni = 1, \dots, N do
            ω∼N(0,I)\omega \sim \mathcal{N}(0, I)
            xti=xti−1+γt2⋅sθ(xti−1,σt)+γt⋅ωx_t^i = x_t^{i-1} + \frac{\gamma_t}{2} \cdot s_\theta(x_t^{i-1}, \sigma_t) + \sqrt{\gamma_t} \cdot \omega
        xt−10=xtNx_{t-1}^0 = x_t^N
    return x00x_0^0
  6. Knowl 6 — Continuous-Time Stochastic Differential Equations Formulation

    equation

    The continuous diffusion framework models the forward perturbation of data over continuous time t∈[0,T]t \in [0, T] via an Itô SDE:

    dx=f(x,t)dt+σ(t)dω\mathrm{d}x = f(x, t) \mathrm{d}t + \sigma(t) \mathrm{d}\omega

    where f(x,t)f(x, t) is the drift coefficient designed to nullify the initial data x0x_0, σ(t)\sigma(t) is the diffusion coefficient controlling noise addition, and ω\omega is standard Brownian motion with dω∼N(0,dtI)\mathrm{d}\omega \sim \mathcal{N}(0, \mathrm{d}t I).

    The generative reverse process is described by the reverse-time SDE:

    dx=[f(x,t)−σ(t)2∇xlog⁡pt(x)]dt+σ(t)dω^\mathrm{d}x = \left[ f(x, t) - \sigma(t)^2 \nabla_x \log p_t(x) \right] \mathrm{d}t + \sigma(t) \mathrm{d}\hat{\omega}

    where ω^\hat{\omega} is standard Brownian motion running backward in time from t=Tt = T to t=0t = 0.

    A continuous neural score model sθ(x,t)≈∇xlog⁡pt(x)s_\theta(x, t) \approx \nabla_x \log p_t(x) is trained using continuous score matching:

    Ldsm∗=Et∼U([0,T])[λ(t)Ep(x0)Ept(xt∣x0)[∥sθ(xt,t)−∇xtlog⁡pt(xt∣x0)∥22]]L^*_{\text{dsm}} = \mathbb{E}_{t \sim \mathcal{U}([0, T])} \left[ \lambda(t) \mathbb{E}_{p(x_0)} \mathbb{E}_{p_t(x_t|x_0)} \left[ \left\| s_\theta(x_t, t) - \nabla_{x_t} \log p_t(x_t|x_0) \right\|_2^2 \right] \right]

    where λ(t)\lambda(t) is a positive weighting function. When f(x,t)f(x, t) is affine, the transition kernel pt(xt∣x0)p_t(x_t|x_0) is Gaussian, allowing closed-form conditional score computation.

  7. Knowl 7 — Euler-Maruyama Numerical Sampling for Reverse-Time SDEs

    algorithm

    The Euler-Maruyama discretization method integrates the reverse-time SDE backward from t=Tt = T to t=0t = 0 using discrete negative time increments Δt\Delta t.

    Input: Negative step size Δt<0\Delta t < 0; drift coefficient function f(x,t)f(x, t); diffusion coefficient function σ(t)\sigma(t); score model ∇xlog⁡pt(x)\nabla_x \log p_t(x); terminal time TT.
    Output: Sampled image xx.
    t=Tt = T
    while t>0t > 0 do
        z∼N(0,I)z \sim \mathcal{N}(0, I)
        Δω^=∣Δt∣⋅z\Delta\hat{\omega} = \sqrt{|\Delta t|} \cdot z
        Δx=(f(x,t)−σ(t)2⋅∇xlog⁡pt(x))⋅Δt+σ(t)⋅Δω^\Delta x = \left( f(x, t) - \sigma(t)^2 \cdot \nabla_x \log p_t(x) \right) \cdot \Delta t + \sigma(t) \cdot \Delta\hat{\omega}
        x=x+Δxx = x + \Delta x
        t=t+Δtt = t + \Delta t
    return xx
  8. Knowl 8 — Variational Lower Bound Objective for Diffusion Models

    equation

    For a Markovian forward process p(x1:T∣x0)=∏t=1Tp(xt∣xt−1)p(x_{1:T}|x_0) = \prod_{t=1}^T p(x_t|x_{t-1}) and reverse model pθ(x0:T)=pθ(xT)∏t=1Tpθ(xt−1∣xt)p_\theta(x_{0:T}) = p_\theta(x_T) \prod_{t=1}^T p_\theta(x_{t-1}|x_t) with standard normal prior π(xT)=N(0,I)\pi(x_T) = \mathcal{N}(0, I), the negative log-likelihood −log⁡pθ(x0)-\log p_\theta(x_0) is bounded by the variational lower bound LvlbL_{\text{vlb}}:

    Lvlb=−log⁡pθ(x0∣x1)+KL(p(xT∣x0)∥π(xT))+∑t>1KL(p(xt−1∣xt,x0)∥pθ(xt−1∣xt))L_{\text{vlb}} = -\log p_\theta(x_0|x_1) + \mathrm{KL}\left(p(x_T|x_0) \parallel \pi(x_T)\right) + \sum_{t > 1} \mathrm{KL}\left(p(x_{t-1}|x_t, x_0) \parallel p_\theta(x_{t-1}|x_t)\right)

    Applying Bayes rule yields the tractable Gaussian forward posterior conditioned on x0x_0:

    p(xt−1∣xt,x0)=p(xt∣xt−1,x0)p(xt−1∣x0)p(xt∣x0)p(x_{t-1}|x_t, x_0) = \frac{p(x_t|x_{t-1}, x_0) p(x_{t-1}|x_0)}{p(x_t|x_0)}

    Because both p(xt−1∣xt,x0)p(x_{t-1}|x_t, x_0) and pθ(xt−1∣xt)p_\theta(x_{t-1}|x_t) are Gaussian with fixed identical covariances σt2I\sigma_t^2 I, the KL divergence simplifies to the distance between their means:

    Lkl=12σt2∥μ~(xt,x0)−μθ(xt,t)∥2+CL_{\text{kl}} = \frac{1}{2\sigma_t^2} \|\tilde{\mu}(x_t, x_0) - \mu_\theta(x_t, t)\|^2 + C

    where μ~(xt,x0)=1αt(xt−βt1−β^tzt)\tilde{\mu}(x_t, x_0) = \frac{1}{\sqrt{\alpha_t}} \left( x_t - \frac{\beta_t}{\sqrt{1 - \hat{\beta}_t}} z_t \right) is the analytical mean of p(xt−1∣xt,x0)p(x_{t-1}|x_t, x_0), μθ(xt,t)\mu_\theta(x_t, t) is the parameterization predicted by the model, and CC is a constant independent of θ\theta.

  9. Knowl 9 — Comparison Between Diffusion Models and Alternative Generative Frameworks

    model/method

    Diffusion models relate to and differ from five major generative modeling paradigms:

    • Variational Auto-Encoders (VAEs): Both optimize variational bounds on data log-likelihood. However, VAEs learn trainable encoding mappings into a lower-dimensional compressed latent space, whereas standard diffusion models use a non-trainable, fixed forward process that preserves the spatial dimensions of the data and completely erases information by step TT.
    • Generative Adversarial Networks (GANs): GANs perform single-step generation via adversarial minimax training but are prone to training instability and mode collapse. Diffusion models optimize likelihood-based objectives, offering stable training and superior sample diversity at the cost of multi-step iterative inference.
    • Normalizing Flows: Both map data distributions to Gaussian distributions. Flows enforce exact invertibility and tractable Jacobian determinants through constrained neural architectures on deterministic paths, whereas diffusion models accommodate unconstrained architectures (e.g., U-Nets) via stochastic trajectories.
    • Energy-Based Models (EBMs): EBMs parameterize unnormalized data densities via scalar energy functions. Noise-conditioned score matching represents a specific case of energy-based modeling where training and Markov Chain Monte Carlo (MCMC) sampling depend entirely on the score function without partition function computation.
    • Autoregressive Models: Autoregressive architectures generate images sequentially pixel-by-pixel with unidirectional bias. Diffusion models synthesize all pixels concurrently across iterative denoising steps.
  10. Knowl 10 — Multi-Perspective Categorization Taxonomy for Vision Diffusion Models

    model/method

    Diffusion models applied in computer vision are categorized along four primary axes:

    1. Task:
      • Unconditional image generation: Generating images from Gaussian noise without conditioning signals.
      • Conditional image generation: Synthesis conditioned on class labels, text descriptions, semantic layouts, source images, or auxiliary signals.
      • Image-to-image translation: Modifying domain attributes while preserving domain-agnostic content.
      • Text-to-image generation: Cascaded diffusion pipelines, Latent Diffusion Models (LDMs), and discrete vector-quantized diffusion models conditioned on multimodal language representations (e.g., CLIP, Sentence-BERT).
      • Image restoration and editing: Super-resolution, inpainting, deblurring, weather artifact removal, and local masked editing.
      • Discriminative and inverse vision tasks: Semantic segmentation using U-Net intermediate feature maps, medical image reconstruction (CT, MRI), image registration, and unsupervised anomaly detection.
      • Video and 3D vision: Autoregressive/3D U-Net video generation, frame infilling, and point-voxel 3D shape generation.
    2. Denoising Condition: Unconditional, class label, classifier guidance (gradient-based), classifier-free guidance, cross-attention CLIP embeddings, or physical measurement operators.
    3. Underlying Architecture: Continuous DDPM/DDIM, score-based NCSN/NCSN++, SDE/ODE continuous solvers (Euler-Maruyama, Predictor-Corrector, DPM-Solver, exponential integrators), discrete-state formulations (D3PM, VQ-DDM), or latent diffusion models (LDM, DiVAE).
    4. Datasets: Standard benchmarks (CIFAR-10/100, ImageNet, CelebA/CelebA-HQ, FFHQ, LSUN categories, MS-COCO, BRATS, fastMRI).
  11. Knowl 11 — Inference Latency and Semantic Limitations of Diffusion Models in Vision

    limitation

    Current diffusion models exhibit several notable limitations in vision applications:

    • Inference Latency: The primary drawback is slow generation speed caused by the requirement of hundreds to thousands of sequential network evaluation steps per sample during reverse sampling, making them considerably slower than single-pass GAN generators.
    • Text Rendering in Text-to-Image Synthesis: Conditioning pipelines on frozen cross-modal embeddings (such as CLIP) often fail to generate legible text or correct spelling in generated scenes because the conditioning embeddings lack character-level orthographic details.
    • Step Size vs. Uncertainty Trade-off: Diffusion trajectories require small discrete steps to ensure intermediate states remain within the support of the learned Gaussian distributions; taking larger step sizes introduces approximation errors and instability during sampling.
    • Temporal Coherence in Video: Video diffusion models remain constrained to relatively short sequences, where capturing long-term temporal dependencies, object persistence, and complex physical interactions remains an open challenge.

Coverage note — Individual summaries of the 70+ specific cited vision papers reviewed in the survey are organized into the unified multi-perspective categorization and task taxonomy knowls rather than extracted as separate knowls.

References

  1. 1.J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using non-equilibrium thermodynamics,” in Proceedings of ICML, pp. 2256–2265, 2015.
  2. 2.J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Proceedings of NeurIPS, vol. 33, pp. 6840–6851, 2020.
  3. 3.Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” in Proceedings of NeurIPS, vol. 32, pp. 11918–11930, 2019.
  4. 4.Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-Based Generative Modeling through Stochastic Differential Equations,” in Proceedings of ICLR, 2021.
  5. 5.P. Dhariwal and A. Nichol, “Diffusion models beat GANs on image synthesis,” in Proceedings of NeurIPS, vol. 34, pp. 8780–8794, 2021.
  6. 6.A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in Proceedings of ICML, pp. 8162–8171, 2021.
  7. 7.J. Song, C. Meng, and S. Ermon, “Denoising Diffusion Implicit Models,” in Proceedings of ICLR, 2021.
  8. 8.D. Watson, W. Chan, J. Ho, and M. Norouzi, “Learning fast samplers for diffusion models by differentiating through sample quality,” in Proceedings of ICLR, 2021.
  9. 9.J. Shi, C. Wu, J. Liang, X. Liu, and N. Duan, “XiVAE: Photorealistic Images Synthesis with Denoising Diffusion Decoder,” arXiv preprint arXiv:2206.00386, 2022.
  10. 10.R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-Resolution Image Synthesis with Latent Diffusion Models,” in Proceedings of CVPR, pp. 10684–10695, 2022.
  11. 11.R. Rombach, A. Blattmann, and B. Ommer, “Text-Guided Synthesis of Artistic Images with Retrieval-Augmented Diffusion Models,” arXiv preprint arXiv:2207.13038, 2022.
  12. 12.C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. Denton, S. K. S. Ghasemipour, B. K. Ayan, S. S. Mahdavi, R. G. Lopes, et al., “Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding,” arXiv preprint arXiv:2205.11487, 2022.
  13. 13.Y. Song and S. Ermon, “Improved techniques for training score-based generative models,” in Proceedings of NeurIPS, vol. 33, pp. 12438–12448, 2020.
  14. 14.A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen, “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models,” in Proceedings of ICML, pp. 16784–16804, 2021.
  15. 15.Y. Song, C. Durkan, I. Murray, and S. Ermon, “Maximum likelihood training of score-based diffusion models,” in Proceedings of NeurIPS, vol. 34, pp. 1415–1428, 2021.
  16. 16.A. Sinha, J. Song, C. Meng, and S. Ermon, “XiC: Diffusion-decoding models for few-shot conditional generation,” in Proceedings of NeurIPS, vol. 34, pp. 12533–12548, 2021.
  17. 17.A. Vahdat, K. Kreis, and J. Kautz, “Score-based generative modeling in latent space,” in Proceedings of NeurIPS, vol. 34, pp. 11287–11302, 2021.
  18. 18.C. Saharia, J. Ho, W. Chan, T. Salimans, D. J. Fleet, and M. Norouzi, “Image super-resolution via iterative refinement,” arXiv preprint arXiv:2104.07636, 2021.
  19. 19.K. Pandey, A. Mukherjee, P. Rai, and A. Kumar, “VAEs meet diffusion models: Efficient and high-fidelity generation,” in Proceedings of NeurIPS Workshop on DGMs and Applications, 2021.
  20. 20.F. Bao, C. Li, J. Zhu, and B. Zhang, “Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models,” in Proceedings of ICLR, 2022.
  21. 21.T. Dockhorn, A. Vahdat, and K. Kreis, “Score-based generative modeling with critically-damped Langevin diffusion,” in Proceedings of ICLR, 2022.
  22. 22.N. Liu, S. Li, Y. Du, A. Torralba, and J. B. Tenenbaum, “Compositional Visual Generation with Composable Diffusion Models,” in Proceedings of ECCV, 2022.
  23. 23.Y. Jiang, S. Yang, H. Qiu, W. Wu, C. C. Loy, and Z. Liu, “Xixt2Human: Text-Driven Controllable Human Image Generation,” ACM Transactions on Graphics, vol. 41, no. 4, pp. 1–11, 2022.
  24. 24.G. Batzolis, J. Stanczuk, C.-B. Schönlieb, and C. Etmann, “Xi-ditional image generation with score-based diffusion models,” arXiv preprint arXiv:2111.13606, 2021.
  25. 25.M. Daniels, T. Maunu, and P. Hand, “Score-based generative neural networks for large-scale optimal transport,” in Proceedings of NeurIPS, pp. 12955–12965, 2021.
  26. 26.H. Chung, B. Sim, and J. C. Ye, “Xiome-Closer-Diffuse-Faster: Accelerating Conditional Diffusion Models for Inverse Problems through Stochastic Contraction,” in Proceedings of CVPR, pp. 12413–12422, 2022.
  27. 27.B. Kawar, M. Elad, S. Ermon, and J. Song, “Xi-noising diffusion restoration models,” in Proceedings of DGM4HSD, 2022.
  28. 28.P. Esser, R. Rombach, A. Blattmann, and B. Ommer, “ImageBART: Bidirectional Context with Multinomial Diffusion for Autoregressive Image Synthesis,” in Proceedings of NeurIPS, vol. 34, pp. 3518–3532, 2021.
  29. 29.A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool, “RePaint: Inpainting using Denoising Diffusion Probabilistic Models,” in Proceedings of CVPR, pp. 11461–11471, 2022.
  30. 30.B. Jing, G. Corso, R. Berlinghieri, and T. Jaakkola, “Subspace diffusion generative models,” arXiv preprint arXiv:2205.01490, 2022.
  31. 31.O. Avrahami, D. Lischinski, and O. Fried, “Blended diffusion for text-driven editing of natural images,” in Proceedings of CVPR, pp. 18208–18218, 2022.
  32. 32.J. Choi, S. Kim, Y. Jeong, Y. Gwon, and S. Yoon, “ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models,” in Proceedings of ICCV, pp. 14347–14356, 2021.
  33. 33.C. Meng, Y. Song, J. Song, J. Wu, J.-Y. Zhu, and S. Ermon, “Xi-dit: Guided Image Synthesis and Editing with Stochastic Differential Equations,” in Proceedings of ICLR, 2021.
  34. 34.C. Saharia, W. Chan, H. Chang, C. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi, “Xi-lette: Image-to-image diffusion models,” in Proceedings of SIGGRAPH, pp. 1–10, 2022.
  35. 35.M. Zhao, F. Bao, C. Li, and J. Zhu, “Xi-SDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations,” arXiv preprint arXiv:2207.06635, 2022.
  36. 36.T. Wang, T. Zhang, B. Zhang, H. Ouyang, D. Chen, Q. Chen, and F. Wen, “Pretraining is All You Need for Image-to-Image Translation,” arXiv preprint arXiv:2205.12952, 2022.
  37. 37.B. Li, K. Xue, B. Liu, and Y.-K. Lai, “XiBB: Image-to-image Translation with Vector Quantized Brownian Bridge,” arXiv preprint arXiv:2205.07680, 2022.
  38. 38.J. Wolleb, R. Sandköhler, F. Bieder, and P. C. Cattin, “The Swiss Army Knife for Image-to-Image Translation: Multi-Task Diffusion Models,” arXiv preprint arXiv:2204.02641, 2022.
  39. 39.D. Baranchuk, I. Rubachev, A. Voynov, V. Khrulkov, and A. Babenko, “Xi-bel-Efficient Semantic Segmentation with Diffusion Models,” in Proceedings of ICLR, 2022.
  40. 40.A. Graikos, N. Malkin, N. Jojic, and D. Samaras, “Xi-fusion models as plug-and-play priors,” arXiv preprint arXiv:2206.09012, 2022.
  41. 41.J. Wolleb, R. Sandköhler, F. Bieder, P. Valmaggia, and P. C. Cattin, “Xi-fusion Models for Implicit Image Segmentation Ensembles,” in Proceedings of MIDL, 2022.
  42. 42.T. Amit, E. Nachmani, T. Shaharbany, and L. Wolf, “Xi-gDiff: Image Segmentation with Diffusion Probabilistic Models,” arXiv preprint arXiv:2112.00390, 2021.
  43. 43.R. S. Zimmermann, L. Schott, Y. Song, B. A. Dunn, and D. A. Klindt, “Score-based generative classifiers,” in Proceedings of NeurIPS Workshop on DGMs and Applications, 2021.
  44. 44.W. H. Pinaya, M. S. Graham, R. Gray, P. F. Da Costa, P.-D. Tudosiu, P. Wright, Y. H. Mah, A. D. MacKinnon, J. T. Teo, R. Jager, et al., “Xi-st Unsupervised Brain Anomaly Detection and Segmentation with Diffusion Models,” arXiv preprint arXiv:2206.03461, 2022.
  45. 45.J. Wolleb, F. Bieder, R. Sandköhler, and P. C. Cattin, “Xi-fusion Models for Medical Anomaly Detection,” arXiv preprint arXiv:2203.04306, 2022.
  46. 46.J. Wyatt, A. Leach, S. M. Schmon, and C. G. Willcocks, “Xi-oDDPM: Anomaly Detection With Denoising Diffusion Probabilistic Models Using Simplex Noise,” in Proceedings of CVPRW, pp. 650–656, 2022.
  47. 47.Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 8, pp. 1798–1828, 2013.
  48. 48.I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. MIT Press, 2016.
  49. 49.G. E. Hinton and R. R. Salakhutdinov, “Xi-ducing the dimensionality of data with neural networks,” Science, vol. 313, no. 5786, pp. 504–507, 2006.
  50. 50.D. P. Kingma and M. Welling, “Auto-Encoding Variational Bayes,” in Proceedings of ICLR, 2014.
  51. 51.I. Higgins, L. Matthey, A. Pal, C. P. Burgess, X. Glorot, M. M. Botvinick, S. Mohamed, and A. Lerchner, “beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework,” in Proceedings of ICLR, 2017.
  52. 52.I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Proceedings of NIPS, pp. 2672–2680, 2014.
  53. 53.M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin, “Unsupervised learning of visual features by contrasting cluster assignments,” in Proceedings of NeurIPS, vol. 33, pp. 9912–9924, 2020.
  54. 54.T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “Xi simple framework for contrastive learning of visual representations,” in Proceedings of ICML, vol. 119, pp. 1597–1607, 2020.
  55. 55.F.-A. Croitoru, D.-N. Grigore, and R. T. Ionescu, “Discriminability-enforcing loss to improve representation learning,” in Proceedings of CVPRW, pp. 2598–2602, 2022.
  56. 56.A. van den Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv preprint arXiv:1807.03748, 2018.
  57. 57.S. Laine and T. Aila, “Xi-mporal ensembling for semi-supervised learning,” in Proceedings of ICLR, 2017.
  58. 58.A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,” in Proceedings of NIPS, vol. 30, pp. 1195–1204, 2017.
  59. 59.C.-W. Huang, J. H. Lim, and A. C. Courville, “Xi variational perspective on diffusion-based generative models and score matching,” in Proceedings of NeurIPS, vol. 34, pp. 22863–22876, 2021.
  60. 60.Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, and F. J. Huang, “Xi tutorial on energy-based learning,” in Predicting Structured Data, MIT Press, 2006.
  61. 61.J. Ngiam, Z. Chen, P. W. Koh, and A. Y. Ng, “Xi-arning Deep Energy Models,” in Proceedings of ICML, pp. 1105–1112, 2011.
  62. 62.A. van den Oord, N. Kalchbrenner, L. Espeholt, K. Kavukcuoglu, O. Vinyals, and A. Graves, “Xi-nditional Image Generation with PixelCNN Decoders,” in Proceedings of NeurIPS, vol. 29, pp. 4797–4805, 2016.
  63. 63.L. Dinh, D. Krueger, and Y. Bengio, “Xi-CE: Non-linear Independent Components Estimation,” in Proceedings of ICLR, 2015.
  64. 64.L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Xi-nsity estimation using Real NVP,” in Proceedings of ICLR, 2017.
  65. 65.O. Ronneberger, P. Fischer, and T. Brox, “Xi-Net: Convolutional Networks for Biomedical Image Segmentation,” in Proceedings of MICCAI, pp. 234–241, 2015.
  66. 66.W. Feller, “Xi the Theory of Stochastic Processes, with Particular Reference to Applications,” in First Berkeley Symposium on Mathematical Statistics and Probability, pp. 403–432, 1949.
  67. 67.P. Vincent, “Xi Connection Between Score Matching and Denoising Autoencoders,” Neural Computation, vol. 23, pp. 1661–1674, 2011.
  68. 68.Y. Song, S. Garg, J. Shi, and S. Ermon, “Xi-ced Score Matching: A Scalable Approach to Density and Score Estimation,” in Proceedings of UAI, p. 204, 2019.
  69. 69.B. D. Anderson, “Xi-verse-time diffusion equation models,” Stochastic Processes and their Applications, vol. 12, no. 3, pp. 313–326, 1982.
  70. 70.T. Salimans, A. Karpathy, X. Chen, and D. P. Kingma, “Xi-xelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications,” in Proceedings of ICLR, 2017.
  71. 71.Q. Zhang and Y. Chen, “Xi-ffusion normalizing flow,” in Proceedings of NeurIPS, vol. 34, pp. 16280–16291, 2021.
  72. 72.K. Swersky, M. Ranzato, D. Buchman, B. M. Marlin, and N. Freitas, “Xi Autoencoders and Score Matching for Energy Based Models,” in Proceedings of ICML, pp. 1201–1208, 2011.
  73. 73.F. Bao, K. Xu, C. Li, L. Hong, J. Zhu, and B. Zhang, “Variational (gradient) estimate of the score function in energy-based latent variable models,” in Proceedings of ICML, pp. 651–661, 2021.
  74. 74.T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen, “Xi-proved Techniques for Training GANs,” in Proceedings of NeurIPS, pp. 2234–2242, 2016.
  75. 75.Y. Shen, J. Gu, X. Tang, and B. Zhou, “Xi-terpreting the Latent Space of GANs for Semantic Face Editing,” in Proceedings of CVPR, pp. 9240–9249, 2020.
  76. 76.A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” in Proceedings of ICLR, 2016.
  77. 77.J. Ho and T. Salimans, “Classifier-Free Diffusion Guidance,” in Proceedings of NeurIPS Workshop on DGMs and Applications, 2021.
  78. 78.J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. van den Berg, “Xi-ructured denoising diffusion models in discrete state-spaces,” in Proceedings of NeurIPS, vol. 34, pp. 17981–17993, 2021.
  79. 79.Y. Benny and L. Wolf, “Xi-namic Dual-Output Diffusion Models,” in Proceedings of CVPR, pp. 11482–11491, 2022.
  80. 80.S. Bond-Taylor, P. Hessey, H. Sasaki, T. P. Breckon, and C. G. Willcocks, “Unleashing Transformers: Parallel Token Prediction with Discrete Absorbing Diffusion for Fast High-Resolution Image Generation from Vector-Quantized Codes,” in Proceedings of ECCV, 2022.
  81. 81.J. Choi, J. Lee, C. Shin, S. Kim, H. Kim, and S. Yoon, “Perception Prioritized Training of Diffusion Models,” in Proceedings of CVPR, pp. 11472–11481, 2022.
  82. 82.V. De Bortoli, J. Thornton, J. Heng, and A. Doucet, “Xi-ffusion Schrödinger bridge with applications to score-based generative modeling,” in Proceedings of NeurIPS, vol. 34, pp. 17695–17709, 2021.
  83. 83.J. Deasy, N. Simidjievski, and P. Lio, “Heavy-tailed denoising score matching,” arXiv preprint arXiv:2112.09788, 2021.
  84. 84.K. Deja, A. Kuzina, T. Trzciński, and J. M. Tomczak, “Xi Analyzing Generative and Denoising Capabilities of Diffusion-Based Deep Generative Models,” arXiv preprint arXiv:2206.00070, 2022.
  85. 85.A. Jolicoeur-Martineau, R. Piché-Taillefer, I. Mitliagkas, and R. T. des Combes, “Xi-versarial score matching and improved sampling for image generation,” in Proceedings of ICLR, 2021.
  86. 86.A. Jolicoeur-Martineau, K. Li, R. Piché-Taillefer, T. Kachman, and I. Mitliagkas, “Xi-tta go fast when generating data with score-based models,” arXiv preprint arXiv:2105.14080, 2021.
  87. 87.D. Kim, B. Na, S. J. Kwon, D. Lee, W. Kang, and I.-C. Moon, “Xi-ximum Likelihood Training of Implicit Nonlinear Diffusion Models,” arXiv preprint arXiv:2205.13699, 2022.
  88. 88.D. Kingma, T. Salimans, B. Poole, and J. Ho, “Variational diffusion models,” in Proceedings of NeurIPS, vol. 34, pp. 21696–21707, 2021.
  89. 89.Z. Kong and W. Ping, “Xi Fast Sampling of Diffusion Probabilistic Models,” in Proceedings of INNF+, 2021.
  90. 90.M. W. Lam, J. Wang, R. Huang, D. Su, and D. Yu, “Xi-lateral denoising diffusion models,” arXiv preprint arXiv:2108.11514, 2021.
  91. 91.L. Liu, Y. Ren, Z. Lin, and Z. Zhao, “Xi-eudo Numerical Methods for Diffusion Models on Manifolds,” in Proceedings of ICLR, 2022.
  92. 92.H. Ma, L. Zhang, X. Zhu, J. Zhang, and J. Feng, “Accelerating Score-Based Generative Models for High-Resolution Image Synthesis,” arXiv preprint arXiv:2206.04029, 2022.
  93. 93.E. Nachmani, R. S. Roman, and L. Wolf, “Xi-n-Gaussian denoising diffusion models,” arXiv preprint arXiv:2106.07582, 2021.
  94. 94.R. San-Roman, E. Nachmani, and L. Wolf, “Xi-ise estimation for generative diffusion models,” arXiv preprint arXiv:2104.02600, 2021.
  95. 95.V. Sehwag, C. Hazirbas, A. Gordo, F. Ozgenel, and C. Canton, “Xi-nerating High Fidelity Data from Low-Density Regions Using Diffusion Models,” in Proceedings of CVPR, pp. 11492–11501, 2022.
  96. 96.G. Wang, Y. Jiao, Q. Xu, Y. Wang, and C. Yang, “Deep generative learning via Schrödinger bridge,” in Proceedings of ICML, pp. 10794–10804, 2021.
  97. 97.Z. Wang, H. Zheng, P. He, W. Chen, and M. Zhou, “Xi-ffusion-GAN: Training GANs with Diffusion,” arXiv preprint arXiv:2206.02262, 2022.
  98. 98.D. Watson, J. Ho, M. Norouzi, and W. Chan, “Xi-arning to efficiently sample from diffusion probabilistic models,” arXiv preprint arXiv:2106.03802, 2021.
  99. 99.Z. Xiao, K. Kreis, and A. Vahdat, “Xi-ckling the generative learning trilemma with denoising diffusion GANs,” in Proceedings of ICLR, 2022.
  100. 100.H. Zheng, P. He, W. Chen, and M. Zhou, “Xi-uncated diffusion probabilistic models,” arXiv preprint arXiv:2202.09671, 2022.
  101. 101.F. Bordes, R. Balestriero, and P. Vincent, “Xi-gh fidelity visualization of what your self-supervised representation knows about,” Transactions on Machine Learning Research, 2022.
  102. 102.A. Campbell, J. Benton, V. De Bortoli, T. Rainforth, G. Deligiannidis, and A. Doucet, “Xi Continuous Time Framework for Discrete Denoising Models,” arXiv preprint arXiv:2205.14987, 2022.
  103. 103.C.-H. Chao, W.-F. Sun, B.-W. Cheng, Y.-C. Lo, C.-C. Chang, Y.-L. Liu, Y.-L. Chang, C.-P. Chen, and C.-Y. Lee, “Xi-noising Likelihood Score Matching for Conditional Score-Based Data Generation,” in Proceedings of ICLR, 2022.
  104. 104.J. Ho, C. Saharia, W. Chan, D. J. Fleet, M. Norouzi, and T. Salimans, “Xi-scaded Diffusion Models for High Fidelity Image Generation,” Journal of Machine Learning Research, vol. 23, no. 47, pp. 1–33, 2022.
  105. 105.T. Karras, M. Aittala, T. Aila, and S. Laine, “Xi-ucidating the Design Space of Diffusion-Based Generative Models,” arXiv preprint arXiv:2206.00364, 2022.
  106. 106.X. Liu, D. H. Park, S. Azadi, G. Zhang, A. Chopikyan, Y. Hu, H. Shi, A. Rohrbach, and T. Darrell, “Xi-re control for free! Image synthesis with semantic diffusion guidance,” arXiv preprint arXiv:2112.05744, 2021.
  107. 107.C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu, “Xi-M-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps,” arXiv preprint arXiv:2206.00927, 2022.
  108. 108.T. Salimans and J. Ho, “Xi-ogressive distillation for fast sampling of diffusion models,” in Proceedings of ICLR, 2022.
  109. 109.V. Singh, S. Jandial, A. Chopra, S. Ramesh, B. Krishnamurthy, and V. N. Balasubramanian, “Xi Conditioning the Input Noise for Controlled Image Generation with Diffusion Models,” arXiv preprint arXiv:2205.03859, 2022.
  110. 110.H. Sasaki, C. G. Willcocks, and T. P. Breckon, “UNIT-DDPM: UNpaired Image Translation with Denoising Diffusion Probabilistic Models,” arXiv preprint arXiv:2104.05358, 2021.
  111. 111.S. Gu, D. Chen, J. Bao, F. Wen, B. Zhang, D. Chen, L. Yuan, and B. Guo, “Xi-ctor quantized diffusion model for text-to-image synthesis,” in Proceedings of CVPR, pp. 10696–10706, 2022.
  112. 112.A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Xi-erarchical text-conditional image generation with CLIP latents,” arXiv preprint arXiv:2204.06125, 2022.
  113. 113.Q. Zhang and Y. Chen, “Xi-st Sampling of Diffusion Models with Exponential Integrator,” arXiv preprint arXiv:2204.13902, 2022.
  114. 114.O. Avrahami, O. Fried, and D. Lischinski, “Xi-ended latent diffusion,” arXiv preprint arXiv:2206.02779, 2022.
  115. 115.G. Batzolis, J. Stanczuk, C.-B. Schönlieb, and C. Etmann, “Xi-n-Uniform Diffusion Models,” arXiv preprint arXiv:2207.09786, 2022.
  116. 116.A. Blattmann, R. Rombach, K. Oktay, and B. Ommer, “Xi-trieval-Augmented Diffusion Models,” arXiv preprint arXiv:2204.11824, 2022.
  117. 117.R. Gao, Y. Song, B. Poole, Y. N. Wu, and D. P. Kingma, “Xi-arning Energy-Based Models by Diffusion Recovery Likelihood,” in Proceedings of ICLR, 2021.
  118. 118.M. Hu, Y. Wang, T.-J. Cham, J. Yang, and P. N. Suganthan, “Xi-obal Context with Discrete Diffusion in Vector Quantised Modelling for Image Generation,” in Proceedings of CVPR, pp. 11502–11511, 2022.
  119. 119.V. Khrulkov and I. Oseledets, “Understanding DDPM Latent Codes Through Optimal Transport,” arXiv preprint arXiv:2202.07477, 2022.
  120. 120.G. Kim, T. Kwon, and J. C. Ye, “Xi-ffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation,” in Proceedings of CVPR, pp. 2426–2435, 2022.
  121. 121.S. Luo and W. Hu, “Xi-ffusion probabilistic models for 3D point cloud generation,” in Proceedings of CVPR, pp. 2837–2845, 2021.
  122. 122.Z. Lyu, X. Xu, C. Yang, D. Lin, and B. Dai, “Xi-celerating Diffusion Models via Early Stop of the Diffusion Process,” arXiv preprint arXiv:2205.12524, 2022.
  123. 123.K. Preechakul, N. Chatthee, S. Wizadwongsa, and S. Suwajanakorn, “Xi-ffusion Autoencoders: Toward a Meaningful and Decodable Representation,” in Proceedings of CVPR, pp. 10619–10629, 2022.
  124. 124.Y. Shi, V. De Bortoli, G. Deligiannidis, and A. Doucet, “Xi-nditional Simulation Using Diffusion Schrödinger Bridges,” in Proceedings of UAI, 2022.
  125. 125.Z. Kadkhodaie and E. P. Simoncelli, “Xi-ochastic solutions for linear inverse problems using the prior implicit in a denoiser,” in Proceedings of NeurIPS, vol. 34, pp. 13242–13254, 2021.
  126. 126.D. Hu, Y. K. Tao, and I. Oguz, “Unsupervised denoising of retinal OCT with diffusion probabilistic model,” in Proceedings of SPIE Medical Imaging, vol. 12032, pp. 25–34, 2022.
  127. 127.H. Chung and J. C. Ye, “Score-based diffusion models for accelerated MRI,” Medical Image Analysis, vol. 80, p. 102479, 2022.
  128. 128.M. Özbey, S. U. Dar, H. A. Bedel, O. Dalmaz, Ş. Özturk, A. Güngör, and T. Çukur, “Unsupervised Medical Image Translation with Adversarial Diffusion Models,” arXiv preprint arXiv:2207.08208, 2022.
  129. 129.Y. Song, L. Shen, L. Xing, and S. Ermon, “Xi-lving inverse problems in medical imaging with score-based generative models,” in Proceedings of ICLR, 2022.
  130. 130.P. Sanchez, A. Kascenas, X. Liu, A. Q. O’Neil, and S. A. Tsaftaris, “What is healthy? generative counterfactual diffusion for lesion localization,” in Proceedings of DGM4MICCAI, 2022.
  131. 131.W. Harvey, S. Naderiparizi, V. Masrani, C. Weilbach, and F. Wood, “Xi-exible Diffusion Modeling of Long Videos,” arXiv preprint arXiv:2205.11495, 2022.
  132. 132.J. Ho, T. Salimans, A. A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” in Proceedings of DGM4HSD, 2022.
  133. 133.R. Yang, P. Srivastava, and S. Mandt, “Xi-ffusion Probabilistic Modeling for Video Generation,” arXiv preprint arXiv:2203.09481, 2022.
  134. 134.T. Höppe, A. Mehrjou, S. Bauer, D. Nielsen, and A. Dittadi, “Xi-ffusion Models for Video Prediction and Infilling,” arXiv preprint arXiv:2206.07696, 2022.
  135. 135.G. Giannone, D. Nielsen, and O. Winther, “Xi-w-Shot Diffusion Models,” arXiv preprint arXiv:2205.15463, 2022.
  136. 136.G. Jeanneret, L. Simon, and F. Jurie, “Xi-ffusion Models for Counterfactual Explanations,” arXiv preprint arXiv:2203.15636, 2022.
  137. 137.P. Sanchez and S. A. Tsaftaris, “Xi-ffusion causal models for counterfactual estimation,” in Proceedings of CLeaR, vol. 140, pp. 1–21, 2022.
  138. 138.O. Özdenizci and R. Legenstein, “Xi-storing Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models,” arXiv preprint arXiv:2207.14626, 2022.
  139. 139.B. Kim, I. Han, and J. C. Ye, “Xi-ffuseMorph: Unsupervised Deformable Image Registration Along Continuous Trajectory Using Diffusion Models,” arXiv preprint arXiv:2112.05149, 2021.
  140. 140.W. Nie, B. Guo, Y. Huang, C. Xiao, A. Vahdat, and A. Anandkumar, “Xi-ffusion models for adversarial purification,” in Proceedings of ICML, 2022.
  141. 141.W. Wang, J. Bao, W. Zhou, D. Chen, D. Chen, L. Yuan, and H. Li, “Xi-mantic image synthesis via diffusion models,” arXiv preprint arXiv:2207.00050, 2022.
  142. 142.L. Zhou, Y. Du, and J. Wu, “3D shape generation and completion through point-voxel diffusion,” in Proceedings of ICCV, pp. 5826–5835, 2021.
  143. 143.A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Xi-tention is all you need,” in Proceedings of NIPS, vol. 30, pp. 6000–6010, 2017.
  144. 144.M. Shannon, B. Poole, S. Mariooryad, T. Bagby, E. Battenberg, D. Kao, D. Stanton, and R. Skerry-Ryan, “Xi-n-saturating GAN training as divergence minimization,” arXiv preprint arXiv:2010.08029, 2020.
  145. 145.M. Arjovsky and L. Bottou, “Xi-wards principled methods for training generative adversarial networks,” in Proceedings of ICLR, 2017.
  146. 146.C. K. Sønderby, J. Caballero, L. Theis, W. Shi, and F. Huszár, “Xi-ortised map inference for image super-resolution,” in Proceedings of ICLR, 2017.
  147. 147.J. Geiping, H. Bauermeister, H. Dröge, and M. Moeller, “Xi-verting gradients – How easy is it to break privacy in federated learning?,” in Proceedings of NeurIPS, vol. 33, pp. 16937–16947, 2020.
  148. 148.H. Tachibana, M. Go, M. Inahara, Y. Katayama, and Y. Watanabe, “Itô-Taylor Sampling Scheme for Denoising Diffusion Probabilistic Models using Ideal Derivatives,” arXiv preprint arXiv:2112.13339, 2021.
  149. 149.J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Xi-paired image-to-image translation using cycle-consistent adversarial networks,” in Proceedings of ICCV, pp. 2223–2232, 2017.
  150. 150.P. Esser, R. Rombach, and B. Ommer, “Xi-ming transformers for high-resolution image synthesis,” in Proceedings of CVPR, pp. 12873–12883, 2021.
  151. 151.A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Xi-arning transferable visual models from natural language supervision,” in Proceedings of ICML, vol. 139, pp. 8748–8763, 2021.
  152. 152.A. Van Den Oord, O. Vinyals, and K. Kavukcuoglu, “Xi-ural discrete representation learning,” Proceedings of NIPS, vol. 30, pp. 6309–6318, 2017.
  153. 153.N. Reimers and I. Gurevych, “Xi-ntence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proceedings of EMNLP, pp. 3982–3992, 2019.
  154. 154.X. Wang, K. Yu, S. Wu, J. Gu, Y. Liu, C. Dong, Y. Qiao, and C. Change Loy, “Xi-RGAN: Enhanced Super-Resolution Generative Adversarial Networks,” in Proceedings of ECCVW, pp. 63–79, 2018.
  155. 155.G. Lin, A. Milan, C. Shen, and I. Reid, “Xi-fineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation,” in Proceeding of CVPR, pp. 5168–5177, 2017.
  156. 156.K. Miyasawa, “Xi empirical Bayes estimator of the mean of a normal population,” Bulletin of the International Statistical Institute, vol. 38, pp. 181–188, 1961.
  157. 157.N. Kovachki, R. Baptista, B. Hosseini, and Y. Marzouk, “Xi-nditional sampling with monotone GANs,” arXiv preprint arXiv:2006.06755, 2020.
  158. 158.Y. Marzouk, T. Moselhy, M. Parno, and A. Spantini, “Xi-mpling via Measure Transport: An Introduction,” in Handbook of Uncertainty Quantification, (Cham), pp. 1–41, Springer, 2016.
  159. 159.S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. B. Van Den Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. De Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. Rae, E. Elsen, and L. Sifre, “Improving Language Models by Retrieving from Trillions of Tokens,” in Proceedings of ICML, vol. 162, pp. 2206–2240, 2022.
  160. 160.I. Oguz, J. D. Malone, Y. Atay, and Y. K. Tao, “Xi-lf-fusion for OCT noise reduction,” in Proceedings of SPIE Medical Imaging, vol. 11313, p. 113130C, SPIE, 2020.
  161. 161.R. T. Ionescu, F. S. Khan, M.-I. Georgescu, and L. Shao, “Xi-ject-Centric Auto-Encoders and Dummy Anomalies for Abnormal Event Detection in Video,” in Proceedings of CVPR, pp. 7842–7851, 2019.
  162. 162.O. Çiçek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger, “3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation,” in Proceedings of MICCAI, pp. 424–432, 2016.
  163. 163.R. Q. Charles, H. Su, M. Kaichun, and L. J. Guibas, “Xi-intNet: Deep Learning on Point Sets for 3D Classification and Segmentation,” in Proceedings of CVPR, pp. 77–85, 2017.
  164. 164.G. Balakrishnan, A. Zhao, M. R. Sabuncu, J. Guttag, and A. V. Dalca, “Xi unsupervised learning model for deformable medical image registration,” in Proceedings of CVPR, pp. 9252–9260, 2018.
  165. 165.X. Li, T.-K. L. Wong, R. T. Chen, and D. Duvenaud, “Xi-alable gradients for stochastic differential equations,” in Proceedings of AISTATS, pp. 3870–3882, 2020.
  166. 166.T. Park, M.-Y. Liu, T.-C. Wang, and J.-Y. Zhu, “Xi-mantic image synthesis with spatially-adaptive normalization,” in Proceedings of CVPR, pp. 2337–2346, 2019.
  167. 167.B. Kawar, G. Vaksman, and M. Elad, “Xi-IPS: Solving noisy inverse problems stochastically,” in Proceedings of NeurIPS, vol. 34, pp. 21757–21769, 2021.

Citation

MLA
Croitoru, F.-A., et al. “Diffusion Models in Vision: A Survey”. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, no. 9, 2023, pp. 10850–69, https://doi.org/10.1109/TPAMI.2023.3261988.
APA
Croitoru, F.-A., Hondru, V., Ionescu, R. T., & Shah, M. (2023). Diffusion Models in Vision: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9), 10850–10869. https://doi.org/10.1109/TPAMI.2023.3261988
Chicago
Croitoru, F.-A., V. Hondru, R. T. Ionescu, and M. Shah. 2023. “Diffusion Models in Vision: A Survey”. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (9): 10850–69. https://doi.org/10.1109/TPAMI.2023.3261988.
Harvard
Croitoru, F.-A. et al. (2023) “Diffusion Models in Vision: A Survey”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(9), pp. 10850–10869. Available at: https://doi.org/10.1109/TPAMI.2023.3261988.
Vancouver
1. Croitoru F-A, Hondru V, Ionescu RT, Shah M (2023) Diffusion Models in Vision: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45:10850–10869

BibTeX

@article{Croitoru_2023, title={Diffusion Models in Vision: A Survey}, volume={45}, ISSN={1939-3539}, url={http://dx.doi.org/10.1109/TPAMI.2023.3261988}, DOI={10.1109/tpami.2023.3261988}, number={9}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, publisher={Institute of Electrical and Electronics Engineers (IEEE)}, author={Croitoru, Florinel-Alin and Hondru, Vlad and Ionescu, Radu Tudor and Shah, Mubarak}, year={2023}, month=Sept, pages={10850–10869} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF