Diffusion Models: A Comprehensive Survey of Methods and Applications

Ling YangZhilong ZhangShenda HongRunsheng XuYue ZhaoYingxia ShaoWentao ZhangMing-Hsuan YangBin Cui

article2022ACM Computing Surveys2,452 citations

Presents a structured taxonomy of diffusion model research, detailing core algorithmic advancements in efficient sampling and likelihood estimation while systematically mapping their applications across computer vision, natural language processing, and the natural sciences.

Listen

Deep generative artificial intelligence models have rapidly advanced, transforming automated content creation, scientific modeling, and pattern recognition. Historically, Generative Adversarial Networks (GANs) served as the primary technology for high-fidelity synthetic image generation, but they suffer from notorious training instability and limited sample diversity. Diffusion models have recently emerged as a dominant alternative, surpassing previous architectures across image synthesis, molecule design, and multi-modal generation. However, the rapid expansion of literature has obscured technical distinctions, core operational mechanics, and optimization trade-offs necessary for strategic technical decision-making.

The article provides an exhaustive, contextualized evaluation of diffusion models, establishing a formal taxonomy of their theoretical foundations, core methodological enhancements, structural adaptations, and wide-ranging domain applications. To do this, the authors conduct an extensive analytical survey synthesizing foundational mathematical formulations—primarily denoising diffusion probabilistic models, score-based generative models, and stochastic differential equations—alongside advanced algorithmic optimizations and integration strategies across multiple fields of applied artificial intelligence.

The findings show that all major diffusion frameworks operate on a unified mathematical principle: progressively corrupting structured data into random noise across a forward trajectory, then learning a parameterized reverse trajectory to reconstruct clean data. Second, while traditional diffusion generation suffered from severe computational latency requiring hundreds of iterative steps, recent innovations—such as higher-order differential equation solvers, knowledge distillation, and latent-space diffusion—reduce inference requirements down to approximately 10 to 20 steps without major drops in sample fidelity. Third, optimizing noise schedules and learning reverse transition variances significantly tightens the mathematical bounds on data likelihood, bridging the performance gap with exact density estimators. Fourth, adapting diffusion frameworks to discrete spaces, non-Euclidean manifolds, and geometric equivariance enables successful deployment on complex domain structures, including molecular graphs, 3D point clouds, and geospatial coordinates.

These developments carry major implications for computational cost, operational latency, and deployment feasibility. By separating the noisy diffusion process into latent spaces or leveraging accelerated solvers, organizations can drastically cut training and serving expenses while matching or exceeding GAN image fidelity. Furthermore, diffusion-based systems offer robust defense capabilities against adversarial noise attacks and enable high-fidelity synthetic data generation for specialized applications in medical imaging and pharmaceutical material discovery. However, practitioners must account for remaining limitations: high training resource demands, potential amplification of dataset biases, privacy risks regarding training data extraction, and an under-developed theoretical framework regarding optimal latent representations and hyperparameter selection.

Moving forward, technical decision-makers should consider adopting latent-space architectures and higher-order deterministic solvers to minimize inference latency in production environments. Organizations deploying these models should implement rigorous data auditing and filtering pipelines to mitigate bias and privacy vulnerabilities. Finally, further exploratory research and pilot evaluations are needed to refine finite-time diffusion formulations—such as Schrödinger bridges—and establish clear theoretical principles for automated hyperparameter tuning.

  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Introduces Denoising Diffusion Probabilistic Models (DDPM), establishing the core mathematical formulation and training objectives surveyed in the source paper.
  • Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Unifies score-based modeling and diffusion models via continuous stochastic differential equations, forming the theoretical backbone for continuous-time diffusion discussed extensively in the survey.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Presents Denoising Diffusion Implicit Models (DDIM), the foundational deterministic sampling method that underpins the survey's discussion of fast sampling techniques.
  • Paper: Improved Denoising Diffusion Probabilistic Models, Alex Nichol et al. (2021). Develops improved variance parameterizations and noise schedules that are central to the survey's coverage of likelihood estimation and efficient sampling.
  • Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). Demonstrates that diffusion models surpass GANs using classifier guidance and architectural improvements, providing key context for the survey's section on conditional generation.
  • Paper: Variational Diffusion Models, Diederik P. Kingma et al. (2021). Formulates variational diffusion models optimizing likelihood via signal-to-noise schedules, serving as a primary foundation for the survey's analysis of density estimation.
  • Paper: Structured Denoising Diffusion Models in Discrete State-Spaces, Jacob Austin et al. (2021). Introduces discrete denoising diffusion probabilistic models (D3PM), establishing the foundational framework for handling non-continuous data reviewed in the survey.
  • Paper: DiffWave: A Versatile Diffusion Model for Audio Synthesis, Zhifeng Kong et al. (2021). Demonstrates the initial expansion of diffusion probabilistic models into high-fidelity audio synthesis, directly supporting the survey's review of temporal data modeling.
  • Paper: Argmax Flows and Multinomial Diffusion: Learning Categorical Distributions, Emiel Hoogeboom et al. (2021). Presents multinomial diffusion for categorical variables, providing early theoretical grounding for the survey's analysis of discrete state-space generative models.
Cover for Diffusion Models: A Comprehensive Survey of Methods and Applications

Abstract

Diffusion models have emerged as a powerful new family of deep generative models with record-breaking performance in many applications, including image synthesis, video generation, and molecule design. In this survey, we provide an overview of the rapidly expanding body of work on diffusion models, categorizing the research into three key areas: efficient sampling, improved likelihood estimation, and handling data with special structures. We also discuss the potential for combining diffusion models with other generative models for enhanced results. We further review the wide-ranging applications of diffusion models in fields spanning from computer vision, natural language generation, temporal data modeling, to interdisciplinary applications in other scientific disciplines. This survey aims to provide a contextualized, in-depth look at the state of diffusion models, identifying the key areas of focus and pointing to potential areas for further exploration. Github: this https URL.

Table of Contents

  • 1 Introduction
  • 2 Foundations of Diffusion Models
  • 2.1 Denoising Diffusion Probabilistic Models (DDPMs)
  • 2.2 Score-Based Generative Models (SGMs)
  • 2.3 Stochastic Differential Equations (Score SDEs)
  • 3 Diffusion Models with Efficient Sampling
  • 3.1 Learning-Free Sampling
  • 3.1.1 SDE Solvers
  • 3.1.2 ODE solvers
  • 3.2 Learning-Based Sampling
  • 3.2.1 Optimized Discretization
  • 3.2.2 Truncated Diffusion
  • 3.2.3 Knowledge Distillation
  • 4 Diffusion Models with Improved Likelihood
  • 4.1 Noise Schedule Optimization
  • 4.2 Reverse Variance Learning
  • 4.3 Exact Likelihood Computation
  • 5 Diffusion Models for Data with Special Structures
  • 5.1 Discrete Data
  • 5.2 Data with Invariant Structures
  • 5.3 Data with Manifold Structures
  • 5.3.1 Known Manifolds
  • 5.3.2 Learned Manifolds
  • 6 Connections with Other Generative Models
  • 6.1 Large Language Models and Connections with Diffusion Models
  • 6.2 Variational Autoencoders and Connections with Diffusion Models
  • 6.3 Generative Adversarial Networks and Connections with Diffusion Models
  • 6.4 Normalizing Flows and Connections with Diffusion Models
  • 6.5 Autoregressive Models and Connections with Diffusion Models
  • 6.6 Energy-based Models and Connections with Diffusion Models
  • 7 Applications of Diffusion Models
  • 7.1 Unconditional and Conditional Diffusion Models
  • 7.1.1 Conditioning Mechanisms in Diffusion Models
  • 7.1.2 Diffusion with DPO/RLHF
  • 7.1.3 Condition Diffusion on Labels and Classifiers
  • 7.1.4 Condition Diffusion on Texts, Images, and Semantic Maps
  • 7.1.5 Condition Diffusion on Graphs
  • 7.2 Computer Vision
  • 7.2.1 Image Super Resolution, Inpainting, Restoration, Translation, and Editing
  • 7.2.2 Semantic Segmentation
  • 7.2.3 Video Generation
  • 7.2.4 Generating Data from Diffusion Models
  • 7.2.5 Point Cloud Completion and Generation
  • 7.2.6 Anomaly Detection
  • 7.3 Natural Language Generation
  • 7.4 Multi-Modal Generation
  • 7.4.1 Text-to-Image Generation
  • 7.4.2 Scene Graph-to-Image Generation
  • 7.4.3 Text-to-3D Generation
  • 7.4.4 Text-to-Motion Generation
  • 7.4.5 Text-to-Video Generation
  • 7.4.6 Text-to-Audio Generation
  • 7.5 Temporal Data Modeling
  • 7.5.1 Time Series Imputation
  • 7.5.2 Time Series Forecasting
  • 7.5.3 Waveform Signal Processing
  • 7.6 Robust Learning
  • 7.7 Interdisciplinary Applications
  • 7.7.1 Drug Design and Life Science
  • 7.7.2 Material Design
  • 7.7.3 Medical Image Reconstruction
  • 8 Future Directions
  • 9 Conclusion
  • References

Knowls

  1. Knowl 1 — Denoising Diffusion Probabilistic Model (DDPM) Formulation and Training

    model/method

    A Denoising Diffusion Probabilistic Model (DDPM) models data generation through two discrete Markov chains: a forward perturbation process and a reverse denoising process.

    Given a data distribution x0∼q(x0)x_0 \sim q(x_0), the forward Markov chain progressively injects Gaussian noise across TT discrete steps according to the transition kernel:

    q(xt∣xt−1)=N(xt;1−βtxt−1,βtI)q(x_t \mid x_{t-1}) = \mathcal{N}(x_t; \sqrt{1 - \beta_t} x_{t-1}, \beta_t I)

    where βt∈(0,1)\beta_t \in (0, 1) is a predetermined variance schedule. Defining αt=1−βt\alpha_t = 1 - \beta_t and αˉt=∏s=1tαs\bar{\alpha}_t = \prod_{s=1}^t \alpha_s, the marginal distribution of xtx_t given x0x_0 can be evaluated analytically at any arbitrary step t∈{1,…,T}t \in \{1, \dots, T\} as:

    q(xt∣x0)=N(xt;αˉtx0,(1−αˉt)I)q(x_t \mid x_0) = \mathcal{N}(x_t; \sqrt{\bar{\alpha}_t} x_0, (1 - \bar{\alpha}_t) I)

    which permits direct sampling via xt=αˉtx0+1−αˉtϵx_t = \sqrt{\bar{\alpha}_t} x_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon for ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I).

    The reverse generative process is defined by a prior p(xT)=N(xT;0,I)p(x_T) = \mathcal{N}(x_T; 0, I) and parameterized transition kernels:

    pθ(xt−1∣xt)=N(xt−1;μθ(xt,t),Σθ(xt,t))p_\theta(x_{t-1} \mid x_t) = \mathcal{N}(x_{t-1}; \mu_\theta(x_t, t), \Sigma_\theta(x_t, t))

    Training minimizes the Kullback-Leibler (KL) divergence between the forward joint distribution q(x0,…,xT)q(x_0, \dots, x_T) and the reverse joint distribution pθ(x0,…,xT)p_\theta(x_0, \dots, x_T), which is equivalent to maximizing the variational lower bound (VLB) on data log-likelihood E[log⁡pθ(x0)]\mathbb{E}[\log p_\theta(x_0)]. With a reweighted VLB objective, the training simplifies to optimizing a noise-prediction neural network ϵθ(xt,t)\epsilon_\theta(x_t, t):

    Lsimple(θ)=Et∼U{1,…,T},x0∼q(x0),ϵ∼N(0,I)[λ(t)∥ϵ−ϵθ(xt,t)∥2]\mathcal{L}_{\text{simple}}(\theta) = \mathbb{E}_{t \sim \mathcal{U}\{1, \dots, T\}, x_0 \sim q(x_0), \epsilon \sim \mathcal{N}(0, I)} \left[ \lambda(t) \|\epsilon - \epsilon_\theta(x_t, t)\|^2 \right]

    where U{1,…,T}\mathcal{U}\{1, \dots, T\} is a discrete uniform distribution and λ(t)>0\lambda(t) > 0 is a positive weighting function.

  2. Knowl 2 — Score-Based Generative Models (SGMs) and Annealed Langevin Dynamics

    model/method

    Score-based Generative Models (SGMs) perturb a data distribution q(x0)q(x_0) with a sequence of intensifying Gaussian noise levels 0<σ1<σ2<⋯<σT0 < \sigma_1 < \sigma_2 < \dots < \sigma_T, producing noisy distributions q(xt∣x0)=N(xt;x0,σt2I)q(x_t \mid x_0) = \mathcal{N}(x_t; x_0, \sigma_t^2 I). The model estimates the Stein score (the gradient of the log probability density ∇xtlog⁡q(xt)\nabla_{x_t} \log q(x_t)) across all noise scales simultaneously using a noise-conditional score network sθ(xt,t)s_\theta(x_t, t).

    The network is trained via denoising score matching:

    LDSM(θ)=Et,x0∼q(x0),xt∼q(xt∣x0)[λ(t)σt2∥∇xtlog⁡q(xt∣x0)−sθ(xt,t)∥2]\mathcal{L}_{\text{DSM}}(\theta) = \mathbb{E}_{t, x_0 \sim q(x_0), x_t \sim q(x_t \mid x_0)} \left[ \lambda(t) \sigma_t^2 \|\nabla_{x_t} \log q(x_t \mid x_0) - s_\theta(x_t, t)\|^2 \right]

    Because ∇xtlog⁡q(xt∣x0)=−xt−x0σt2=−ϵσt\nabla_{x_t} \log q(x_t \mid x_0) = -\frac{x_t - x_0}{\sigma_t^2} = -\frac{\epsilon}{\sigma_t}, the loss reduces to:

    LDSM(θ)=Et,x0∼q(x0),ϵ∼N(0,I)[λ(t)∥ϵ+σtsθ(xt,t)∥2]+const\mathcal{L}_{\text{DSM}}(\theta) = \mathbb{E}_{t, x_0 \sim q(x_0), \epsilon \sim \mathcal{N}(0, I)} \left[ \lambda(t) \|\epsilon + \sigma_t s_\theta(x_t, t)\|^2 \right] + \text{const}

    This establishes an equivalence between the training objectives of DDPMs and SGMs under the reparameterization ϵθ(x,t)=−σtsθ(x,t)\epsilon_\theta(x, t) = -\sigma_t s_\theta(x, t).

    Sample generation is achieved by Annealed Langevin Dynamics (ALD), chaining score updates across noise levels from t=Tt = T down to t=1t = 1:

    Input: Noise levels σ1<⋯<σT\sigma_1 < \dots < \sigma_T, steps per scale NN, step sizes st>0s_t > 0, score model sθ(x,t)s_\theta(x, t)
    Output: Sample x0x_0 approximating q(x0)q(x_0)
    Sample xT(0)∼N(0,σT2I)x_T^{(0)} \sim \mathcal{N}(0, \sigma_T^2 I)
    for t=T,T−1,…,1t = T, T-1, \dots, 1 do
        if t<Tt < T then
            xt(0)←xt+1(N)x_t^{(0)} \leftarrow x_{t+1}^{(N)}
        for i=0,1,…,N−1i = 0, 1, \dots, N-1 do
            Sample ϵ(i)∼N(0,I)\epsilon^{(i)} \sim \mathcal{N}(0, I)
            xt(i+1)←xt(i)+12stsθ(xt(i),t)+stϵ(i)x_t^{(i+1)} \leftarrow x_t^{(i)} + \frac{1}{2} s_t s_\theta(x_t^{(i)}, t) + \sqrt{s_t} \epsilon^{(i)}
    return x1(N)x_1^{(N)}
  3. Knowl 3 — Continuous-Time Score SDE and Probability Flow ODE Formulations

    theoretical result

    In continuous time (T→∞T \to \infty), data perturbation is formalized by a forward stochastic differential equation (SDE):

    dx=f(x,t)dt+g(t)dw\mathrm{d}x = f(x, t)\mathrm{d}t + g(t)\mathrm{d}w

    where f(x,t)f(x, t) is the drift function, g(t)g(t) is the diffusion coefficient, and ww is a standard Wiener process. Discrete diffusion models correspond to specific discretizations:

    • Variance Preserving SDE (VP-SDE, continuous limit of DDPM): dx=−12β(t)xdt+β(t)dw\mathrm{d}x = -\frac{1}{2}\beta(t)x\mathrm{d}t + \sqrt{\beta(t)}\mathrm{d}w, where β(t/T)=Tβt\beta(t/T) = T\beta_t.
    • Variance Exploding SDE (VE-SDE, continuous limit of SGM): dx=d[σ(t)2]dtdw\mathrm{d}x = \sqrt{\frac{\mathrm{d}[\sigma(t)^2]}{\mathrm{d}t}}\mathrm{d}w, where σ(t/T)=σt\sigma(t/T) = \sigma_t.

    For any forward SDE of this form, the time-reversal is governed by the reverse-time SDE:

    dx=[f(x,t)−g(t)2∇xlog⁡qt(x)]dt+g(t)dwˉ\mathrm{d}x = \left[ f(x, t) - g(t)^2 \nabla_x \log q_t(x) \right] \mathrm{d}t + g(t)\mathrm{d}\bar{w}

    where wˉ\bar{w} is a standard Wiener process evolving backwards in time, dt\mathrm{d}t is an infinitesimal negative time step, and qt(x)q_t(x) is the marginal density of x(t)x(t).

    Furthermore, there exists an associated deterministic ordinary differential equation, termed the probability flow ODE:

    dx=[f(x,t)−12g(t)2∇xlog⁡qt(x)]dt\mathrm{d}x = \left[ f(x, t) - \frac{1}{2}g(t)^2 \nabla_x \log q_t(x) \right] \mathrm{d}t

    The trajectories of both the reverse-time SDE and the probability flow ODE have the exact same marginal probability densities qt(x)q_t(x) as the forward SDE at all times t∈[0,T]t \in [0, T], enabling sample generation from the data distribution q0(x)q_0(x) when integrated backwards from t=Tt=T to t=0t=0.

    The score function ∇xlog⁡qt(x)\nabla_x \log q_t(x) is estimated by optimizing a score model sθ(xt,t)s_\theta(x_t, t) with the continuous-time score matching objective:

    Et∼U[0,T],x0∼q(x0),xt∼q(xt∣x0)[λ(t)∥sθ(xt,t)−∇xtlog⁡q0t(xt∣x0)∥2]\mathbb{E}_{t \sim \mathcal{U}[0, T], x_0 \sim q(x_0), x_t \sim q(x_t \mid x_0)} \left[ \lambda(t) \| s_\theta(x_t, t) - \nabla_{x_t} \log q_{0t}(x_t \mid x_0) \|^2 \right]

  4. Knowl 4 — Non-Markovian Forward Perturbation and Deterministic Sampling in DDIM

    model/method

    Denoising Diffusion Implicit Models (DDIM) generalize DDPMs to a family of non-Markovian forward perturbation processes that share the same marginal distributions q(xt∣x0)=N(xt;αˉtx0,(1−αˉt)I)q(x_t \mid x_0) = \mathcal{N}(x_t; \sqrt{\bar{\alpha}_t} x_0, (1 - \bar{\alpha}_t) I) for all t∈{1,…,T}t \in \{1, \dots, T\}.

    The joint forward distribution is factorized as q(x1:T∣x0)=∏t=1Tq(xt∣xt−1,x0)q(x_{1:T} \mid x_0) = \prod_{t=1}^T q(x_t \mid x_{t-1}, x_0), where the transition kernels are specified via a parameter σt\sigma_t:

    qσ(xt−1∣xt,x0)=N(xt−1;αˉt−1x0+1−αˉt−1−σt2xt−αˉtx01−αˉt,σt2I)q_\sigma(x_{t-1} \mid x_t, x_0) = \mathcal{N}\left(x_{t-1}; \sqrt{\bar{\alpha}_{t-1}} x_0 + \sqrt{1 - \bar{\alpha}_{t-1} - \sigma_t^2} \frac{x_t - \sqrt{\bar{\alpha}_t} x_0}{\sqrt{1 - \bar{\alpha}_t}}, \sigma_t^2 I\right)

    This formulation encapsulates both stochastic and deterministic generative processes:

    • Choosing σt=1−αˉt−11−αˉtβt\sigma_t = \sqrt{\frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t} \beta_t} recovers the standard DDPM reverse Markov chain.
    • Choosing σt=0\sigma_t = 0 for all tt produces a completely deterministic generative process.

    When σt=0\sigma_t = 0, the DDIM update step corresponds to a numerical discretization of the continuous probability flow ODE. Because it uses the exact same score network or noise predictor ϵθ(xt,t)\epsilon_\theta(x_t, t) trained via standard DDPM objectives, DDIM enables fast deterministic sampling and latent trajectory inversion over a reduced number of discretization steps without model retraining.

  5. Knowl 5 — Analytical and Learned Reverse Variance Formulations for Likelihood Maximization

    model/method

    Standard DDPM models often fix the reverse process variance to a constant schedule Σθ(xt,t)=βtI\Sigma_\theta(x_t, t) = \beta_t I or β~tI\tilde{\beta}_t I, where β~t=1−αˉt−11−αˉtβt\tilde{\beta}_t = \frac{1 - \bar{\alpha}_{t-1}}{1 - \bar{\alpha}_t} \beta_t. To improve likelihood estimation and reduce discretization steps, the reverse variance can be optimized or evaluated analytically:

    1. Learned Variance Interpolation (iDDPM): The reverse variance is parameterized as an interpolation in the logarithmic domain:

    Σθ(xt,t)=exp⁡(vlog⁡βt+(1−v)log⁡β~t)\Sigma_\theta(x_t, t) = \exp\left( v \log \beta_t + (1 - v) \log \tilde{\beta}_t \right)

    where v∈[0,1]v \in [0, 1] is a vector predicted by an auxiliary head of the neural network, trained jointly using the variational lower bound (VLB) objective.

    1. Optimal Analytic Variance (Analytic-DPM): The optimal reverse covariance Σ∗(xt,t)\Sigma^*(x_t, t) for a dd-dimensional continuous diffusion model can be calculated directly in closed form from a pre-trained score network ∇xtlog⁡qt(xt)\nabla_{x_t} \log q_t(x_t) without retraining:

    Σ∗(xt,t)=σt2I+(βtαt−βt−1−σt2)2(1−βtdEqt(xt)[∥∇xtlog⁡qt(xt)∥2])I\Sigma^*(x_t, t) = \sigma_t^2 I + \left( \sqrt{\frac{\beta_t}{\alpha_t}} - \sqrt{\beta_{t-1} - \sigma_t^2} \right)^2 \left( 1 - \frac{\beta_t}{d} \mathbb{E}_{q_t(x_t)}\left[ \|\nabla_{x_t} \log q_t(x_t)\|^2 \right] \right) I

    Plugging these analytically estimated variances into the reverse transitions yields tighter variational lower bounds and improved log-likelihoods.

  6. Knowl 6 — Exact Likelihood Computation via Probability Flow ODE and Score SDE Bounds

    theoretical result

    Exact data log-likelihood and upper bounds can be computed for continuous-time diffusion models using differential equation formulations:

    1. Exact Likelihood via Probability Flow ODE: The probability flow ODE dxtdt=f~θ(xt,t)=f(xt,t)−12g(t)2sθ(xt,t)\frac{\mathrm{d}x_t}{\mathrm{d}t} = \tilde{f}_\theta(x_t, t) = f(x_t, t) - \frac{1}{2}g(t)^2 s_\theta(x_t, t) represents a continuous normalizing flow. The exact log-likelihood of a data sample x0x_0 under the ODE distribution pθodep_\theta^{\text{ode}} is given by the instantaneous change of variables formula:

    log⁡pθode(x0)=log⁡pT(xT)+∫0T∇⋅f~θ(xt,t)dt\log p_\theta^{\text{ode}}(x_0) = \log p_T(x_T) + \int_0^T \nabla \cdot \tilde{f}_\theta(x_t, t) \mathrm{d}t

    where the divergence term ∇⋅f~θ(xt,t)\nabla \cdot \tilde{f}_\theta(x_t, t) is evaluated efficiently using the Skilling-Hutchinson trace estimator with random noise vectors ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) via ϵ⊤∇f~θ(xt,t)ϵ\epsilon^\top \nabla \tilde{f}_\theta(x_t, t) \epsilon.

    1. Reverse SDE Likelihood Bound (ScoreFlows): Training a continuous score SDE with the likelihood weighting function λ(t)=g(t)2\lambda(t) = g(t)^2 implicitly minimizes the KL divergence DKL(q0∥pθsde)D_{\text{KL}}(q_0 \parallel p_\theta^{\text{sde}}) to the data distribution q0q_0:

    DKL(q0∥pθsde)≤L(θ;g(⋅)2)+DKL(qT∥π)D_{\text{KL}}(q_0 \parallel p_\theta^{\text{sde}}) \le \mathcal{L}(\theta; g(\cdot)^2) + D_{\text{KL}}(q_T \parallel \pi)

    Furthermore, the negative log-likelihood −log⁡pθsde(x)-\log p_\theta^{\text{sde}}(x) for an individual point xx is bounded above by L′(x)\mathcal{L}'(x):

    L′(x)=∫0TE[12∥g(t)sθ(xt,t)∥2+∇⋅(g(t)2sθ(xt,t)−f(xt,t))  |  x0=x]dt−ExT[log⁡pθsde(xT)∣x0=x]\mathcal{L}'(x) = \int_0^T \mathbb{E}\left[ \frac{1}{2}\|g(t)s_\theta(x_t, t)\|^2 + \nabla \cdot (g(t)^2 s_\theta(x_t, t) - f(x_t, t)) \;\middle|\; x_0 = x \right] \mathrm{d}t - \mathbb{E}_{x_T}\left[ \log p_\theta^{\text{sde}}(x_T) \mid x_0 = x \right]

  7. Knowl 7 — Signal-to-Noise Ratio Decomposition of the Continuous-Time Variational Lower Bound

    theoretical result

    In Variational Diffusion Models (VDMs), the continuous forward noising process is defined by variance σt2=sigmoid(γη(t))\sigma_t^2 = \text{sigmoid}(\gamma_\eta(t)) and scale αˉt=1−σt2\bar{\alpha}_t = \sqrt{1 - \sigma_t^2}, parameterized by a monotonic neural network γη(t)\gamma_\eta(t). The Signal-to-Noise Ratio (SNR) is defined as R(t)=αˉt2σt2R(t) = \frac{\bar{\alpha}_t^2}{\sigma_t^2}.

    The continuous-time Variational Lower Bound (VLB) on data log-likelihood LVLB\mathcal{L}_{\text{VLB}} decomposes into:

    LVLB=−Ex0[DKL(q(xT∣x0)∥p(xT))]+Ex0,x1[log⁡p(x0∣x1)]−LD\mathcal{L}_{\text{VLB}} = -\mathbb{E}_{x_0}\left[ D_{\text{KL}}(q(x_T \mid x_0) \parallel p(x_T)) \right] + \mathbb{E}_{x_0, x_1}\left[ \log p(x_0 \mid x_1) \right] - \mathcal{L}_D

    where the diffusion loss LD\mathcal{L}_D is given by an integral over SNR values:

    LD=12Ex0,ϵ∼N(0,I)[∫Rmin⁡Rmax⁡∥x0−x~θ(xv,v)∥22 dv]\mathcal{L}_D = \frac{1}{2} \mathbb{E}_{x_0, \epsilon \sim \mathcal{N}(0, I)} \left[ \int_{R_{\min}}^{R_{\max}} \|x_0 - \tilde{x}_\theta(x_v, v)\|_2^2 \, \mathrm{d}v \right]

    where Rmax⁡=R(1)R_{\max} = R(1), Rmin⁡=R(T)R_{\min} = R(T), vv is the integration variable in SNR space, xv=αˉvx0+σvϵx_v = \sqrt{\bar{\alpha}_v} x_0 + \sigma_v \epsilon at t=R−1(v)t = R^{-1}(v), and x~θ\tilde{x}_\theta denotes the model's estimate of the clean data point x0x_0.

    Because the integral depends solely on the endpoint SNRs Rmin⁡R_{\min} and Rmax⁡R_{\max}, intermediate choices of the forward noise schedule γη(t)\gamma_\eta(t) do not alter the theoretical value of the continuous VLB, affecting only the variance of Monte Carlo estimators used in empirical training.

  8. Knowl 8 — Invariance Preservation in Diffusion Models via Group Equivariant Markov Kernels

    theoretical result

    For structured data exhibiting global symmetries (such as graphs under node permutation or 3D molecular structures under 3D Euclidean transformations T∈SE(3)\mathcal{T} \in \mathrm{SE}(3) including rotations and translations), diffusion models preserve distributional invariance under transformation group T\mathcal{T} if the following two conditions hold:

    1. Invariant Prior Distribution: The prior noise distribution p(xT)p(x_T) satisfies:

    p(xT)=p(T(xT)),∀Tp(x_T) = p(\mathcal{T}(x_T)), \quad \forall \mathcal{T}

    1. Equivariant Reverse Kernels: The parameterized reverse transition kernels pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t) satisfy:

    pθ(xt−1∣xt)=pθ(T(xt−1)∣T(xt)),∀Tp_\theta(x_{t-1} \mid x_t) = p_\theta(\mathcal{T}(x_{t-1}) \mid \mathcal{T}(x_t)), \quad \forall \mathcal{T}

    When both conditions are satisfied, the marginal distribution of generated samples p0(x)p_0(x) is mathematically guaranteed to be strictly invariant under transformation T\mathcal{T}:

    p0(x)=p0(T(x))p_0(x) = p_0(\mathcal{T}(x))

    This principle ensures exact symmetry preservation in geometric diffusion models (e.g., GeoDiff and EDP-GNN) by utilizing SE(3)\mathrm{SE}(3)-equivariant and permutation-equivariant graph neural networks for the score function.

  9. Knowl 9 — Fast ODE Sampling via Semi-Linear Exponential Integrators

    model/method

    Fast ODE solvers for diffusion models (such as DEIS and DPM-Solver) exploit the semi-linear formulation of the probability flow ODE. Under diffusion SDEs where the forward drift f(x,t)f(x, t) is linear in xx (such as f(x,t)=−12β(t)xf(x, t) = -\frac{1}{2}\beta(t)x in VP-SDE), the probability flow ODE can be written as:

    dxtdt=f(xt,t)−12g(t)2sθ(xt,t)\frac{\mathrm{d}x_t}{\mathrm{d}t} = f(x_t, t) - \frac{1}{2}g(t)^2 s_\theta(x_t, t)

    By applying an integrating factor (exponential integrator), the linear component f(xt,t)f(x_t, t) is integrated analytically without introducing discretization error:

    xt=e∫stf0(τ)dτxs−∫ste∫utf0(τ)dτ12g(u)2sθ(xu,u)dux_t = e^{\int_s^t f_0(\tau)\mathrm{d}\tau} x_s - \int_s^t e^{\int_u^t f_0(\tau)\mathrm{d}\tau} \frac{1}{2}g(u)^2 s_\theta(x_u, u) \mathrm{d}u

    where f0(t)f_0(t) is the scalar multiplier in the linear drift f(x,t)=f0(t)xf(x, t) = f_0(t)x. Numerical approximation is restricted to the integral containing the non-linear score function sθ(x,t)s_\theta(x, t).

    First-order polynomial approximation of the score function recovers the deterministic DDIM update rule. Higher-order approximations (such as second-order and third-order exponential integrators) drastically reduce discretization error, enabling high-quality image generation in 10 to 20 evaluation steps.

  10. Knowl 10 — Discrete Diffusion Models via Transition Matrices and Continuous-Time Markov Chains

    model/method

    Discrete diffusion models (such as D3PM and VQ-Diffusion) generalize diffusion models to discrete state spaces where continuous Gaussian perturbations and Euclidean Stein score functions are ill-defined.

    For a discrete variable x∈{1,…,K}x \in \{1, \dots, K\} represented as a one-hot column vector v(x)∈{0,1}Kv(x) \in \{0, 1\}^K, the discrete forward perturbation is governed by a sequence of categorical transition matrices Qt∈RK×KQ_t \in \mathbb{R}^{K \times K}:

    q(xt∣xt−1)=v(xt)⊤Qtv(xt−1)q(x_t \mid x_{t-1}) = v(x_t)^\top Q_t v(x_{t-1})

    where [Qt]ij=q(xt=i∣xt−1=j)[Q_t]_{ij} = q(x_t = i \mid x_{t-1} = j) is a row-stochastic matrix. The multi-step forward marginal q(xt∣x0)q(x_t \mid x_0) is computed analytically via matrix products:

    q(xt∣x0)=v(xt)⊤Qˉtv(x0),Qˉt=QtQt−1…Q1q(x_t \mid x_0) = v(x_t)^\top \bar{Q}_t v(x_0), \quad \bar{Q}_t = Q_t Q_{t-1} \dots Q_1

    Choices for QtQ_t include uniform categorical transitions, discretized Gaussian transitions, lazy random walks, or absorbing state kernels (where tokens transition to a special [MASK] absorbing token with probability βt\beta_t).

    In continuous time, discrete diffusion is formulated via Continuous Time Markov Chains (CTMCs) characterized by transition rate matrices, allowing continuous-time sampling with bounded total variation error relative to the data distribution.

  11. Knowl 11 — Riemannian Score-Based Generative Models and Manifold Diffusion

    model/method

    For data constrained to lie on known compact Riemannian manifolds M\mathcal{M} (such as spheres Sd\mathbb{S}^d or toruses Td\mathbb{T}^d), diffusion models are adapted through two primary formulations:

    1. Intrinsic Formulation (Riemannian Score-Based Generative Model / RSGM): The continuous-time Score SDE is extended to M\mathcal{M} by replacing Euclidean Brownian motion with Riemannian Brownian motion generated by the Laplace-Beltrami operator ΔM\Delta_\mathcal{M}. The score function ∇Mlog⁡qt(x)\nabla_\mathcal{M} \log q_t(x) is defined as a vector field on the tangent bundle TMT\mathcal{M}. Sampling is conducted intrinsically along the manifold geodesics using a Geodesic Random Walk, and the network is trained using generalized Riemannian denoising score matching.

    2. Extrinsic Formulation (Riemannian Diffusion Model / RDM): The Riemannian manifold M\mathcal{M} is treated as being embedded in an ambient Euclidean space RN\mathbb{R}^N. RDM defines a variational objective matching the continuous evidence lower bound on the embedded manifold and proves that maximizing this lower bound is equivalent to minimizing a projected Riemannian score-matching error.

  12. Knowl 12 — Theoretical Connections Between Diffusion Models and Other Generative Frameworks

    theoretical result

    Diffusion models share formal structural and mathematical relationships with five major classes of deep generative models:

    1. Variational Autoencoders (VAEs): DDPM is equivalent to an LL-step hierarchical Markovian VAE with fixed linear Gaussian encoders and shared decoder parameters across time steps. In the continuous-time limit, optimizing a Score SDE is equivalent to maximizing the Evidence Lower Bound (ELBO) of an infinitely deep hierarchical VAE.

    2. Generative Adversarial Networks (GANs): GAN training instability caused by disjoint data and generator supports is mitigated by injecting adaptive diffusion noise into discriminator inputs (Diffusion-GAN). Conversely, replacing Gaussian reverse steps with conditional GAN generators allows multi-step denoising with large step sizes (Denoising Diffusion GANs).

    3. Normalizing Flows: The deterministic probability flow ODE is a continuous normalizing flow (CNF) that enables exact likelihood computation. Stochastic normalizing flows (such as DiffFlow) incorporate stochastic diffusion noise directly into invertible forward and backward flow dynamics.

    4. Energy-Based Models (EBMs): The score function ∇xlog⁡p(x)\nabla_x \log p(x) corresponds to the negative gradient of an energy function −∇xEθ(x)-\nabla_x E_\theta(x). Diffusion Recovery Likelihood models a sequence of EBMs conditioned on noisy observations from higher noise levels, avoiding intractable partition functions through tractable recovery likelihood optimization.

    5. Autoregressive Models (ARMs): Autoregressive Diffusion Models (ARDMs) generalize order-agnostic ARMs and discrete diffusion models by replacing causal sequential factorization with parallelizable score-matching objectives.

Coverage note — Application-specific implementation details across individual downstream tasks (in computer vision, NLP, time series, multi-modal generation, and medical imaging) and general discussions of future directions were omitted, as they represent contextual summaries of external literature rather than foundational methodological formulations of the survey.

References

  1. 1.Juan Miguel Lopez Alcaraz and Nils Strodthoff. 2022. Diffusion-based time series imputation and forecasting with structured state space models. arXiv preprint arXiv:2208.09399 (2022).
  2. 2.Tomer Amit, Eliya Nachmani, Tal Shaharbany, and Lior Wolf. 2021. SegDiff: Image segmentation with diffusion probabilistic models. arXiv preprint arXiv:2112.00390 (2021).
  3. 3.Namrata Anand and Tudor Achim. 2022. Protein structure and sequence generation with equivariant denoising diffusion probabilistic models. arXiv preprint arXiv:2205.15019 (2022).
  4. 4.Brian D. O. Anderson. 1982. Reverse-time diffusion equation models. Stochastic Processes and Their Applications 12, 3 (1982), 313–326.
  5. 5.Uri M. Ascher and Linda R. Petzold. 1998. Computer Methods for Ordinary Differential Equations and Differential-Algebraic Equations. SIAM.
  6. 6.Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg. 2021. Structured denoising diffusion models in discrete state-spaces. In Advances in Neural Information Processing Systems.
  7. 7.Omri Avrahami, Dani Lischinski, and Ohad Fried. 2022. Blended diffusion for text-driven editing of natural images. In IEEE Conference on Computer Vision and Pattern Recognition. 18208–18218.
  8. 8.Wele Gedara Chaminda Bandara, Nithin Gopalakrishnan Nair, and Vishal M. Patel. 2022. DDPM-CD: Remote sensing change detection using denoising diffusion probabilistic models. arXiv preprint arXiv:2206.11892 (2022).
  9. 9.Hritik Bansal and Aditya Grover. 2023. Leaving reality to imagination: Robust classification via generated datasets. In International Conference on Learning Representations.
  10. 10.Fan Bao, Chongxuan Li, Jun Zhu, and Bo Zhang. 2021. Analytic-DPM: An analytic estimate of the optimal reverse variance in diffusion probabilistic models. In International Conference on Learning Representations.
  11. 11.Dmitry Baranchuk, Andrey Voynov, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. 2021. Label-efficient semantic segmentation with diffusion models. In International Conference on Learning Representations.
  12. 12.Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schönlieb, and Christian Etmann. 2021. Conditional image generation with score-based diffusion models. arXiv preprint arXiv:2111.13606 (2021).
  13. 13.Samy Bengio and Yoshua Bengio. 2000. Taking on the curse of dimensionality in joint distributions using neural networks. IEEE Transactions on Neural Networks (2000).
  14. 14.Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003. A neural probabilistic language model. Journal of Machine Learning Research 3 (2003), 1137–1155.
  15. 15.Helen M. Berman, John Westbrook, Zukang Feng, Gary Gilliland, Talapady N. Bhat, Helge Weissig, Ilya N. Shindyalov, and Philip E. Bourne. 2000. The protein data bank. Nucleic Acids Research 28, 1 (2000), 235–242.
  16. 16.Piotr Bielak, Tomasz Kajdanowicz, and Nitesh V. Chawla. 2021. Graph Barlow Twins: A self-supervised representation learning framework for graphs. arXiv preprint arXiv:2106.02466 (2021).
  17. 17.Mikołaj Bińkowski, Dougal J. Sutherland, Michael Arbel, and Arthur Gretton. 2018. Demystifying MMD GANs. In International Conference on Learning Representations.
  18. 18.Tsachi Blau, Roy Ganz, Bahjat Kawar, Alex Bronstein, and Michael Elad. 2022. Threat model-agnostic adversarial defense using diffusion models. arXiv preprint arXiv:2207.08089 (2022).
  19. 19.Emmanuel Asiedu Brempong, Simon Kornblith, Ting Chen, Niki Parmar, Matthias Minderer, and Mohammad Norouzi. 2022. Denoising pretraining for semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition. 4175–4186.
  20. 20.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems.
  21. 21.Blake Bullwinkel, Kristen Grabarz, Lily Ke, Scarlett Gong, Chris Tanner, and Joshua Allen. 2022. Evaluating the fairness impact of differentially private synthetic data. arXiv preprint arXiv:2205.04321 (2022).
  22. 22.Keith T. Butler, Daniel W. Davies, Hugh Cartwright, Olexandr Isayev, and Aron Walsh. 2018. Machine learning for molecular and materials science. Nature 559, 7715 (2018), 547–555.
  23. 23.Ruojin Cai, Guandao Yang, Hadar Averbuch-Elor, Zekun Hao, Serge Belongie, Noah Snavely, and Bharath Hariharan. 2020. Learning gradient fields for shape generation. In European Conference on Computer Vision. 364–381.
  24. 24.Andrew Campbell, Joe Benton, Valentin De Bortoli, Tom Rainforth, George Deligiannidis, and Arnaud Doucet. 2022. A continuous time framework for discrete denoising models. arXiv preprint arXiv:2205.14987 (2022).
  25. 25.Chentao Cao, Zhuo-Xu Cui, Shaonan Liu, Dong Liang, and Yanjie Zhu. 2022. High-frequency space diffusion models for accelerated MRI. arXiv preprint arXiv:2208.05481 (2022).
  26. 26.Wei Cao, Dong Wang, Jian Li, Hao Zhou, Lei Li, and Yitan Li. 2018. BRITS: Bidirectional recurrent imputation for time series. In Advances in Neural Information Processing Systems, Vol. 31.
  27. 27.Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, and Eric Wallace. 2023. Extracting training data from diffusion models. arXiv preprint arXiv:2301.13188 (2023).
  28. 28.Nicholas Carlini, Florian Tramer, Krishnamurthy Dvijotham, and Kolter J. Zico. 2022. (Certified!!) Adversarial robustness for free! arXiv preprint arXiv:2206.10550 (2022).
  29. 29.Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. 2022. MaskGIT: Masked generative image transformer. In IEEE Conference on Computer Vision and Pattern Recognition. 11315–11325.
  30. 30.Zhengping Che, Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu. 2018. Recurrent neural networks for multivariate time series with missing values. Scientific Reports 8, 1 (2018), 1–12.
  31. 31.Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013. One billion word benchmark for measuring progress in statistical language modeling. arXiv preprint arXiv:1312.3005 (2013).
  32. 32.Dong Chen, Xinda Qi, Yu Zheng, Yuzhen Lu, and Zhaojian Li. 2022. Deep data augmentation for weed recognition enhancement: A diffusion probabilistic model and transfer learning based approach. arXiv preprint arXiv:2210.09509 (2022).
  33. 33.Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan. 2020. WaveGrad: Estimating gradients for waveform generation. arXiv preprint arXiv:2009.00713 (2020).
  34. 34.Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud. 2018. Neural ordinary differential equations. arXiv preprint arXiv:1806.07366 (2018).
  35. 35.Tianrong Chen, Guan-Horng Liu, and Evangelos Theodorou. 2021. Likelihood training of Schrödinger bridge using forward-backward SDEs theory. In International Conference on Learning Representations.
  36. 36.Ting Chen, Ruixiang Zhang, and Geoffrey Hinton. 2022. Analog bits: Generating discrete data using diffusion models with self-conditioning. arXiv preprint arXiv:2208.04202 (2022).
  37. 37.Rewon Child. 2020. Very deep VAEs generalize autoregressive models and can outperform them on images. In International Conference on Learning Representations.
  38. 38.Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019. Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509 (2019).
  39. 39.Jaemin Cho, Abhay Zala, and Mohit Bansal. 2022. DALL-EVAL: Probing the reasoning skills and social biases of text-to-image generative models. arXiv preprint arXiv:2202.04053 (2022).
  40. 40.Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022. PaLM: Scaling language modeling with pathways. arXiv preprint arXiv:2204.02311 (2022).
  41. 41.Hyungjin Chung, Eun Sun Lee, and Jong Chul Ye. 2022. MR image denoising and super-resolution using regularized reverse diffusion. arXiv preprint arXiv:2203.12621 (2022).
  42. 42.Hyungjin Chung, Byeongsu Sim, and Jong Chul Ye. 2022. Come-closer-diffuse-faster: Accelerating conditional diffusion models for inverse problems through stochastic contraction. In IEEE Conference on Computer Vision and Pattern Recognition. 12413–12422.
  43. 43.Hyungjin Chung and Jong Chul Ye. 2022. Score-based diffusion models for accelerated MRI. Medical Image Analysis (2022), 102479.
  44. 44.Rob Cornish, Anthony Caterini, George Deligiannidis, and Arnaud Doucet. 2020. Relaxing bijectivity constraints with continuously indexed normalising flows. In International Conference on Machine Learning. 2133–2143.
  45. 45.Antonia Creswell, Tom White, Vincent Dumoulin, Kai Arulkumaran, Biswa Sengupta, and Anil A. Bharath. 2018. Generative adversarial networks: An overview. IEEE Signal Processing Magazine 35, 1 (2018), 53–65.
  46. 46.Katherine Crowson, Stella Biderman, Daniel Kornis, Dashiell Stander, Eric Hallahan, Louis Castricato, and Edward Raff. 2022. VQGAN-CLIP: Open domain image generation and editing with natural language guidance. arXiv preprint arXiv:2204.08583 (2022).
  47. 47.Koller Daphne and Friedman Nir. 2009. Probabilistic Graphical Models: Principles and Techniques. MIT Press.
  48. 48.Salman U. H. Dar, Şaban Öztürk, Yilmaz Korkmaz, Gokberk Elmas, Muzaffer Özbey, Alper Güngör, and Tolga Çukur. 2022. Adaptive diffusion priors for accelerated MRI reconstruction. arXiv preprint arXiv:2207.05876 (2022).
  49. 49.Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2019. Plug and play language models: A simple approach to controlled text generation. In International Conference on Learning Representations.
  50. 50.Valentin De Bortoli, Arnaud Doucet, Jeremy Heng, and James Thornton. 2021. Simulating diffusion bridges with score matching. arXiv preprint arXiv:2111.07243 (2021).
  51. 51.Valentin De Bortoli, Emile Mathieu, Michael Hutchinson, James Thornton, Yee Whye Teh, and Arnaud Doucet. 2022. Riemannian score-based generative modeling. arXiv preprint arXiv:2202.02763 (2022).
  52. 52.Valentin De Bortoli, James Thornton, Jeremy Heng, and Arnaud Doucet. 2021. Diffusion Schrödinger bridge with applications to score-based generative modeling. In Advances in Neural Information Processing Systems, Vol. 34. 17695–17709.
  53. 53.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition. 248–255.
  54. 54.Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, Vol. 34. 8780–8794.
  55. 55.Laurent Dinh, David Krueger, and Yoshua Bengio. 2014. NICE: Non-linear independent components estimation. arXiv preprint arXiv:1410.8516 (2014).
  56. 56.Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. 2017. Density estimation using Real NVP. In International Conference on Learning Representations.
  57. 57.Laurent Dinh, Jascha Sohl-Dickstein, Hugo Larochelle, and Razvan Pascanu. 2019. A RAD approach to deep mixture models. arXiv preprint arXiv:1903.07714 (2019).
  58. 58.Tim Dockhorn, Tianshi Cao, Arash Vahdat, and Karsten Kreis. 2022. Differentially private diffusion models. arXiv preprint arXiv:2210.09929 (2022).
  59. 59.Tim Dockhorn, Arash Vahdat, and Karsten Kreis. 2021. Score-based generative modeling with critically-damped Langevin diffusion. In International Conference on Learning Representations.
  60. 60.Tim Dockhorn, Arash Vahdat, and Karsten Kreis. 2022. GENIE: Higher-order denoising diffusion solvers. Advances in Neural Information Processing Systems (2022).
  61. 61.Carl Doersch. 2016. Tutorial on variational autoencoders. arXiv preprint arXiv:1606.05908 (2016).
  62. 62.Yifan Du, Zikang Liu, Junyi Li, and Wayne Xin Zhao. 2022. A survey of vision-language pre-trained models. arXiv preprint arXiv:2202.10936 (2022).
  63. 63.David K. Duvenaud, Dougal Maclaurin, Jorge Iparraguirre, Rafael Bombarell, Timothy Hirzel, Alán Aspuru-Guzik, and Ryan P. Adams. 2015. Convolutional networks on graphs for learning molecular fingerprints. In Advances in Neural Information Processing Systems, Vol. 28.
  64. 64.Emadeldeen Eldele, Mohamed Ragab, Zhenghua Chen, Min Wu, Chee Keong Kwoh, Xiaoli Li, and Cuntai Guan. 2021. Time-series representation learning via temporal and contextual contrasting. arXiv preprint arXiv:2106.14112 (2021).
  65. 65.Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. In IEEE Conference on Computer Vision and Pattern Recognition. 12873–12883.
  66. 66.Charles Fefferman, Sanjoy Mitter, and Hariharan Narayanan. 2016. Testing the manifold hypothesis. Journal of the American Mathematical Society 29, 4 (2016), 983–1049.
  67. 67.Vincent Fortuin, Dmitry Baranchuk, Gunnar Ratsch, and Stephan Mandt. 2020. GP-VAE: Deep probabilistic time series imputation. In International Conference on Artificial Intelligence and Statistics. 1651–1661.
  68. 68.Giulio Franzese, Simone Rossi, Lixuan Yang, Alessandro Finamore, Dario Rossi, Maurizio Filippone, and Pietro Michiardi. 2022. How much is enough? A study on diffusion times in score-based generative models. arXiv preprint arXiv:2206.05173 (2022).
  69. 69.Ruiqi Gao, Yang Song, Ben Poole, Ying Nian Wu, and Diederik P. Kingma. 2020. Learning energy-based models by diffusion recovery likelihood. arXiv preprint arXiv:2012.08125 (2020).
  70. 70.Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. 2017. Neural message passing for quantum chemistry. In International Conference on Machine Learning. 1263–1272.
  71. 71.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in Neural Information Processing Systems, Vol. 27. 139–144.
  72. 72.Marco Gori, Gabriele Monfardini, and Franco Scarselli. 2005. A new model for learning in graph domains. In International Joint Conference on Neural Networks, Vol. 2. 729–734.
  73. 73.Sven Gowal, Sylvestre-Alvise Rebuffi, Olivia Wiles, Florian Stimberg, Dan Andrei Calian, and Timothy A. Mann. 2021. Improving robustness using generated data. Advances in Neural Information Processing Systems 34 (2021), 4218–4233.
  74. 74.Will Grathwohl, Ricky T. Q. Chen, Jesse Bettencourt, and David Duvenaud. 2019. Scalable reversible generative models with free-form continuous dynamics. In International Conference on Learning Representations.
  75. 75.Alex Graves. 2013. Generating sequences with recurrent neural networks. arXiv preprint arXiv:1308.0850 (2013).
  76. 76.Ulf Grenander and Michael I. Miller. 1994. Representations of knowledge in complex systems. Journal of the Royal Statistical Society: Series B (Methodological) 56, 4 (1994), 549–581.
  77. 77.Albert Gu, Karan Goel, and Christopher Re. 2021. Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations.
  78. 78.Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. 2022. Vector quantized diffusion model for text-to-image synthesis. In IEEE Conference on Computer Vision and Pattern Recognition. 10696–10706.
  79. 79.Jie Gui, Zhenan Sun, Yonggang Wen, Dacheng Tao, and Jieping Ye. 2021. A review on generative adversarial networks: Algorithms, theory, and applications. IEEE Transactions on Knowledge and Data Engineering (2021).
  80. 80.William L. Hamilton, Rex Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. In Advances in Neural Information Processing Systems. 1025–1035.
  81. 81.William L. Hamilton, Rex Ying, and Jure Leskovec. 2017. Representation learning on graphs: Methods and applications. arXiv preprint arXiv:1709.05584 (2017).
  82. 82.Songqiao Han, Xiyang Hu, Hailiang Huang, Mingqi Jiang, and Yue Zhao. 2022. ADBench: Anomaly detection benchmark. arXiv preprint arXiv:2206.09426 (2022).
  83. 83.William Harvey, Saeid Naderiparizi, Vaden Masrani, Christian Weilbach, and Frank Wood. 2022. Flexible diffusion modeling of long videos. arXiv preprint arXiv:2205.11495 (2022).
  84. 84.Ruifei He, Shuyang Sun, Xin Yu, Chuhui Xue, Wenqing Zhang, Philip Torr, Song Bai, and Xiaojuan Qi. 2022. Is synthetic data from generative models ready for image recognition?. In International Conference on Learning Representations.
  85. 85.Roei Herzig, Amir Bar, Huijuan Xu, Gal Chechik, Trevor Darrell, and Amir Globerson. 2020. Learning canonical representations for scene graph to image generation. 210–227.
  86. 86.Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. 2022. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303 (2022).
  87. 87.Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, Vol. 33. 6840–6851.
  88. 88.Jonathan Ho, Chitwan Saharia, William Chan, David J. Fleet, Mohammad Norouzi, and Tim Salimans. 2022. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research 23 (2022), 47–1.
  89. 89.Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022).
  90. 90.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. 2022. Video diffusion models. arXiv preprint arXiv:2204.03458 (2022).
  91. 91.Emiel Hoogeboom, Victor Garcia Satorras, Clement Vignac, and Max Welling. 2022. Equivariant diffusion for molecule generation in 3D. arXiv print arXiv:2203.17003 (2022).
  92. 92.Emiel Hoogeboom, Alexey A. Gritsenko, Jasmijn Bastings, Ben Poole, Rianne van den Berg, and Tim Salimans. 2021. Autoregressive diffusion models. In International Conference on Learning Representations.
  93. 93.Emiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré, and Max Welling. 2021. Argmax flows and multinomial diffusion: Learning categorical distributions. In Advances in Neural Information Processing Systems, Vol. 34. 12454–12465.
  94. 94.Chin-Wei Huang, Milad Aghajohari, Joey Bose, Prakash Panangaden, and Aaron C. Courville. 2022. Riemannian diffusion models. Advances in Neural Information Processing Systems 35 (2022), 2750–2761.
  95. 95.Chin-Wei Huang, Jae Hyun Lim, and Aaron C. Courville. 2021. A variational perspective on diffusion-based generative models and score matching. In Advances in Neural Information Processing Systems, Vol. 34. 22863–22876.
  96. 96.Rongjie Huang, Zhou Zhao, Huadai Liu, Jinglin Liu, Chenye Cui, and Yi Ren. 2022. ProDiff: Progressive fast diffusion model for high-quality text-to-speech. arXiv preprint arXiv:2207.06389 (2022).
  97. 97.Michael F. Hutchinson. 1989. A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines. Communications in Statistics-Simulation and Computation 18, 3 (1989), 1059–1076.
  98. 98.Aapo Hyvärinen. 2005. Estimation of non-normalized statistical models by score matching. Journal of Machine Learning Research 6 (2005), 695–709.
  99. 99.Touseef Iqbal and Shaima Qureshi. 2020. The survey: Text generation models in deep learning. Journal of King Saud University-Computer and Information Sciences (2020).
  100. 100.Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. 2017. Image-to-image translation with conditional adversarial networks. In IEEE Conference on Computer Vision and Pattern Recognition. 1125–1134.
  101. 101.Long Jin, Justin Lazarow, and Zhuowen Tu. 2017. Introspective classification with convolutional nets. In Advances in Neural Information Processing Systems, Vol. 30. 823–833.
  102. 102.Wengong Jin, Regina Barzilay, and Tommi Jaakkola. 2018. Junction tree variational autoencoder for molecular graph generation. In International Conference on Machine Learning. 2323–2332.
  103. 103.Bowen Jing, Gabriele Corso, Renato Berlinghieri, and Tommi Jaakkola. 2022. Subspace diffusion generative models. arXiv preprint arXiv:2205.01490 (2022).
  104. 104.Bowen Jing, Gabriele Corso, Jeffrey Chang, Regina Barzilay, and Tommi Jaakkola. 2022. Torsional diffusion for molecular conformer generation. arXiv preprint arXiv:2206.01729 (2022).
  105. 105.Jaehyeong Jo, Seul Lee, and Sung Ju Hwang. 2022. Score-based generative modeling of graphs via the system of stochastic differential equations. arXiv preprint arXiv:2202.02514 (2022).
  106. 106.Justin Johnson, Agrim Gupta, and Li Fei-Fei. 2018. Image generation from scene graphs. In IEEE Conference on Computer Vision and Pattern Recognition. 1219–1228.
  107. 107.Alexia Jolicoeur-Martineau, Ke Li, Rémi Piché-Taillefer, Tal Kachman, and Ioannis Mitliagkas. 2021. Gotta go fast when generating data with score-based models. arXiv preprint arXiv:2105.14080 (2021).
  108. 108.Alexia Jolicoeur-Martineau, Rémi Piché-Taillefer, Rémi Tachet des Combes, and Ioannis Mitliagkas. 2020. Adversarial score matching and improved sampling for image generation. arXiv preprint arXiv:2009.05475 (2020).
  109. 109.John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, Anna Potapenko, Alex Bridgland, Clemens Meyer, Simon A. A. Kohl, Andrew J. Ballard, Andrew Cowie, Bernardino Romera-Paredes, Stanislav Nikolov, Rishub Jain, Jonas Adler, Trevor Back, Stig Petersen, David Reiman, Ellen Clancy, Michal Zielinski, Martin Steinegger, Michalina Pacholska, Tamas Berghammer, Sebastian Bodenstein, David Silver, Oriol Vinyals, Andrew W. Senior, Koray Kavukcuoglu, Pushmeet Kohli, and Demis Hassabis. 2021. Highly accurate protein structure prediction with AlphaFold. Nature 596, 7873 (2021), 583–589.
  110. 110.Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu. 2018. Efficient neural audio synthesis. In International Conference on Machine Learning. 2410–2419.
  111. 111.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. 2022. Elucidating the design space of diffusion-based generative models. arXiv preprint arXiv:2206.00364 (2022).
  112. 112.Bahjat Kawar, Roy Ganz, and Michael Elad. 2022. Enhancing diffusion-based image synthesis with robust classifier guidance. arXiv preprint arXiv:2208.08664 (2022).
  113. 113.Bahjat Kawar, Gregory Vaksman, and Michael Elad. 2021. Stochastic image denoising by sampling from the posterior distribution. In International Conference on Computer Vision. 1866–1875.
  114. 114.Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019. CTRL: A conditional transformer language model for controllable generation. arXiv preprint arXiv:1909.05858 (2019).
  115. 115.Boah Kim, Inhwa Han, and Jong Chul Ye. 2021. DiffuseMorph: Unsupervised deformable image registration along continuous trajectory using diffusion models. arXiv preprint arXiv:2112.05149 (2021).
  116. 116.Gwanghyun Kim and Se Young Chun. 2023. DATID-3D: Diversity-preserved domain adaptation using text-to-image diffusion for 3D generative model. In IEEE Conference on Computer Vision and Pattern Recognition. 14203–14213.
  117. 117.Gwanghyun Kim, Taesung Kwon, and Jong Chul Ye. 2022. DiffusionCLIP: Text-guided diffusion models for robust image manipulation. In IEEE Conference on Computer Vision and Pattern Recognition. 2426–2435.
  118. 118.Sungwon Kim, Heeseung Kim, and Sungroh Yoon. 2022. Guided-TTS 2: A diffusion model for high-quality adaptive text-to-speech with untranscribed data. arXiv preprint arXiv:2205.15370 (2022).
  119. 119.Taesup Kim and Yoshua Bengio. 2016. Deep directed generative models with energy-based probability estimation. arXiv preprint arXiv:1606.03439 (2016).
  120. 120.Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. 2021. Variational diffusion models. In Advances in Neural Information Processing Systems, Vol. 34. 21696–21707.
  121. 121.Diederik P. Kingma and Prafulla Dhariwal. 2018. Glow: Generative flow with invertible 1x1 convolutions. arXiv preprint arXiv:1807.03039 (2018).
  122. 122.Diederik P. Kingma and Max Welling. 2013. Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114 (2013).
  123. 123.Diederik P. Kingma and Max Welling. 2019. An introduction to variational autoencoders. Foundations and Trends in Machine Learning 12, 4 (2019), 307–392.
  124. 124.Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. 2020. DiffWave: A versatile diffusion model for audio synthesis. arXiv preprint arXiv:2009.09761 (2020).
  125. 125.Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq Joty, Richard Socher, and Nazneen Fatema Rajani. 2020. GeDi: Generative discriminator guided sequence generation. arXiv preprint arXiv:2009.06367 (2020).
  126. 126.Alex Krizhevsky. 2009. Learning multiple layers of features from tiny images. (2009).
  127. 127.Hugo Larochelle and Iain Murray. 2011. The neural autoregressive distribution estimator. In International Conference on Artificial Intelligence and Statistics.
  128. 128.Justin Lazarow, Long Jin, and Zhuowen Tu. 2017. Introspective neural networks for generative modeling. In International Conference on Computer Vision. 2774–2783.
  129. 129.Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fujie Huang. 2006. A tutorial on energy-based learning. Predicting Structured Data (2006).
  130. 130.Jin Sub Lee, Jisun Kim, and Philip M. Kim. 2022. ProteinSGM: Score-based generative modeling for de novo protein design. bioRxiv (2022), 2022–07.
  131. 131.Kwonjoon Lee, Weijian Xu, Fan Fan, and Zhuowen Tu. 2018. Wasserstein introspective neural networks. In IEEE Conference on Computer Vision and Pattern Recognition. 3702–3711.
  132. 132.Seul Lee, Jaehyeong Jo, and Sung Ju Hwang. 2022. Exploring chemical space with score-based out-of-distribution generation. arXiv preprint arXiv:2206.07632 (2022).
  133. 133.Alon Levkovitch, Eliya Nachmani, and Lior Wolf. 2022. Zero-shot voice conditioning for denoising diffusion TTS models. arXiv preprint arXiv:2206.02246 (2022).
  134. 134.Haoying Li, Yifan Yang, Meng Chang, Huajun Feng, Zhi hai Xu, Qi Li, and Yue ting Chen. 2022. SRDiff: Single image super-resolution with diffusion probabilistic models. Neurocomputing 479 (2022), 47–59.
  135. 135.Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning. PMLR, 12888–12900.
  136. 136.Junyi Li, Tianyi Tang, Gaole He, Jinhao Jiang, Xiaoxuan Hu, Puzhao Xie, Zhipeng Chen, Zhuohao Yu, Wayne Xin Zhao, and Ji-Rong Wen. 2021. TextBox: A unified, modularized, and extensible framework for text generation. arXiv preprint arXiv:2101.02046 (2021).
  137. 137.Junyi Li, Tianyi Tang, Wayne Xin Zhao, and Ji-Rong Wen. 2021. Pretrained language models for text generation: A survey. arXiv preprint arXiv:2105.10311 (2021).
  138. 138.Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan. 2019. Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting. In Advances in Neural Information Processing Systems, Vol. 32.
  139. 139.Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, and Tatsunori B. Hashimoto. 2022. Diffusion-LM improves controllable text generation. arXiv preprint arXiv:2205.14217 (2022).
  140. 140.Yikang Li, Tao Ma, Yeqi Bai, Nan Duan, Sining Wei, and Xiaogang Wang. 2019. PasteGAN: A semi-parametric method to generate image from scene graph. Advances in Neural Information Processing Systems 32.
  141. 141.Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. 2022. Magic3D: High-resolution text-to-3D content creation. arXiv preprint arXiv:2211.10440 (2022).
  142. 142.Luping Liu, Yi Ren, Zhijie Lin, and Zhou Zhao. 2021. Pseudo numerical methods for diffusion models on manifolds. In International Conference on Learning Representations.
  143. 143.Aaron Lou, Derek Lim, Isay Katsman, Leo Huang, Qingxuan Jiang, Ser Nam Lim, and Christopher M. De Sa. 2020. Neural manifold ordinary differential equations. Advances in Neural Information Processing Systems 33 (2020), 17548–17558.
  144. 144.Cheng Lu, Kaiwen Zheng, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2022. Maximum likelihood training for score-based diffusion ODEs by high order denoising score matching. In International Conference on Machine Learning. 14429–14460.
  145. 145.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2022. DPM-solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927 (2022).
  146. 146.Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. 2022. Repaint: Inpainting using denoising diffusion probabilistic models. In IEEE Conference on Computer Vision and Pattern Recognition. 11461–11471.
  147. 147.Eric Luhman and Troy Luhman. 2021. Knowledge distillation in iterative generative models for improved sampling speed. arXiv preprint arXiv:2101.02388 (2021).
  148. 148.Calvin Luo. 2022. Understanding diffusion models: A unified perspective. arXiv preprint arXiv:2208.11970 (2022).
  149. 149.Shitong Luo and Wei Hu. 2021. Diffusion probabilistic models for 3D point cloud generation. In IEEE Conference on Computer Vision and Pattern Recognition. 2837–2845.
  150. 150.Shitong Luo and Wei Hu. 2021. Score-based point cloud denoising. In International Conference on Computer Vision. 4583–4592.
  151. 151.Shitong Luo, Chence Shi, Minkai Xu, and Jian Tang. 2021. Predicting molecular conformation via dynamic graph score matching. In Advances in Neural Information Processing Systems, Vol. 34. 19784–19795.
  152. 152.Shitong Luo, Yufeng Su, Xingang Peng, Sheng Wang, Jian Peng, and Jianzhu Ma. 2022. Antigen-specific antibody design and optimization with diffusion-based generative models for protein structures. Advances in Neural Information Processing Systems 35 (2022), 9754–9767.
  153. 153.Yonghong Luo, Xiangrui Cai, Ying Zhang, Jun Xu, and Xiaojie Yuan. 2018. Multivariate time series imputation with generative adversarial networks. In Advances in Neural Information Processing Systems, Vol. 31.
  154. 154.Zhaoyang Lyu, Zhifeng Kong, X. U. Xudong, Liang Pan, and Dahua Lin. 2021. A conditional point diffusion-refinement paradigm for 3D point cloud completion. In International Conference on Learning Representations.
  155. 155.Zhaoyang Lyu, Xudong Xu, Ceyuan Yang, Dahua Lin, and Bo Dai. 2022. Accelerating diffusion models via early stop of the diffusion process. arXiv preprint arXiv:2205.12524 (2022).
  156. 156.Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards deep learning models resistant to adversarial attacks. In International Conference on Learning Representations.
  157. 157.Emile Mathieu and Maximilian Nickel. 2020. Riemannian continuous normalizing flows. Advances in Neural Information Processing Systems 33 (2020), 2503–2515.
  158. 158.Siyuan Mei, Fuxin Fan, and Andreas Maier. 2022. Metal inpainting in CBCT projections using score-based generative model. arXiv preprint arXiv:2209.09733 (2022).
  159. 159.Gábor Melis, Chris Dyer, and Phil Blunsom. 2018. On the state of the art of evaluation in neural language models. In International Conference on Learning Representations.
  160. 160.Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon. 2021. SDEdit: Guided image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations.
  161. 161.Chenlin Meng, Jiaming Song, Yang Song, Shengjia Zhao, and Stefano Ermon. 2020. Improved autoregressive modeling with distribution smoothing. In International Conference on Learning Representations.
  162. 162.Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018. Regularizing and optimizing LSTM language models. In International Conference on Learning Representations.
  163. 163.Nicholas Metropolis and Stanislaw Ulam. 1949. The Monte Carlo method. Journal of the American Statistical Association 44, 247 (1949), 335–341.
  164. 164.Gal Metzer, Elad Richardson, Or Patashnik, Raja Giryes, and Daniel Cohen-Or. 2023. Latent-NeRF for shape-guided generation of 3D shapes and textures. In IEEE Conference on Computer Vision and Pattern Recognition. 12663–12673.
  165. 165.Alexander Quinn Nichol and Prafulla Dhariwal. 2021. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning. 8162–8171.
  166. 166.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob Mcgrew, Ilya Sutskever, and Mark Chen. 2022. GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning. 16784–16804.
  167. 167.Weili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao, Arash Vahdat, and Anima Anandkumar. 2022. Diffusion models for adversarial purification. arXiv preprint arXiv:2205.07460 (2022).
  168. 168.Erik Nijkamp, Mitch Hill, Tian Han, Song-Chun Zhu, and Ying Nian Wu. 2019. On the anatomy of MCMC-based maximum likelihood learning of energy-based models. arXiv preprint arXiv:1903.12370 (2019).
  169. 169.Chenhao Niu, Yang Song, Jiaming Song, Shengjia Zhao, Aditya Grover, and Stefano Ermon. 2020. Permutation invariant graph generation via score-based generative modeling. In International Conference on Artificial Intelligence and Statistics. 4474–4484.
  170. 170.Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. WaveNet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016).
  171. 171.Boris N. Oreshkin, Dmitri Carpov, Nicolas Chapados, and Yoshua Bengio. 2019. N-BEATS: Neural basis expansion analysis for interpretable time series forecasting. In International Conference on Learning Representations.
  172. 172.Irfan Pratama, Adhistya Erna Permanasari, Igi Ardiyanto, and Rini Indrayani. 2016. A review of missing values handling methods on time-series data. In International Conference on Information Technology Systems and Innovation, IEEE, 1–6.
  173. 173.Muzaffer Özbey, Salman U. H. Dar, Hasan A. Bedel, Onat Dalmaz, Şaban Özturk, Alper Güngör, and Tolga Çukur. 2022. Unsupervised medical image translation with adversarial diffusion models. arXiv preprint arXiv:2207.08208 (2022).
  174. 174.George Papamakarios, Eric T. Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. 2021. Normalizing flows for probabilistic modeling and inference. Journal of Machine Learning Research 22, 57 (2021), 1–64.
  175. 175.Giorgio Parisi. 1981. Correlation functions and computer simulations. Nuclear Physics B 180, 3 (1981), 378–384.
  176. 176.Sung Woo Park, Kyungjae Lee, and Junseok Kwon. 2021. Neural Markov controlled SDE: Stochastic optimization for continuous-time data. In International Conference on Learning Representations.
  177. 177.Cheng Peng, Pengfei Guo, S. Kevin Zhou, Vishal M. Patel, and Rama Chellappa. 2022. Towards performant and reliable undersampled MR reconstruction via diffusion model sampling. In International Conference on Medical Image Computing and Computer-Assisted Intervention. 623–633.
  178. 178.Stanislav Pidhorskyi, Donald A. Adjeroh, and Gianfranco Doretto. 2020. Adversarial latent autoencoders. In IEEE Conference on Computer Vision and Pattern Recognition. 14104–14113.
  179. 179.Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. 2022. DreamFusion: Text-to-3D using 2D diffusion. arXiv preprint arXiv:2209.14988 (2022).
  180. 180.Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov. 2021. Grad-TTS: A diffusion probabilistic model for text-to-speech. In International Conference on Machine Learning. 8599–8608.
  181. 181.Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. 2023. FateZero: Fusing attentions for zero-shot text-based video editing. In International Conference on Computer Vision.
  182. 182.Lawrence R. Rabiner. 1989. A tutorial on hidden Markov models and selected applications in speech recognition. Proceedings of the IEEE 77, 2 (1989), 257–286.
  183. 183.Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning. 8748–8763.
  184. 184.Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9.
  185. 185.Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21, 1 (2020), 5485–5551.
  186. 186.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125 (2022).
  187. 187.Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International Conference on Machine Learning. 8821–8831.
  188. 188.Martin Raphan and Eero P. Simoncelli. 2007. Learning to be Bayesian without supervision. In Advances in Neural Information Processing Systems. 1145–1152.
  189. 189.Martin Raphan and Eero P. Simoncelli. 2011. Least squares estimation without priors or supervision. Neural Computation 23, 2 (2011), 374–420.
  190. 190.Kashif Rasul, Calvin Seward, Ingmar Schuster, and Roland Vollgraf. 2021. Autoregressive denoising diffusion models for multivariate probabilistic time series forecasting. In International Conference on Machine Learning. 8857–8868.
  191. 191.Kashif Rasul, Abdul-Saboor Sheikh, Ingmar Schuster, Urs M. Bergmann, and Roland Vollgraf. 2020. Multivariate probabilistic time series forecasting via conditioned normalizing flows. In International Conference on Learning Representations.
  192. 192.Lillian J. Ratliff, Samuel A. Burden, and S. Shankar Sastry. 2013. Characterization and computation of local Nash equilibria in continuous games. In Annual Allerton Conference on Communication, Control, and Computing. 917–924.
  193. 193.Danilo Rezende and Shakir Mohamed. 2015. Variational inference with normalizing flows. In International Conference on Machine Learning. 1530–1538.
  194. 194.Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra. 2014. Stochastic backpropagation and approximate inference in deep generative models. In International Conference on Machine Learning. 1278–1286.
  195. 195.Oren Rippel and Ryan Prescott Adams. 2013. High-dimensional probability estimation with deep density models. arXiv preprint arXiv:1302.5125 (2013).
  196. 196.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition. 10684–10695.
  197. 197.Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2022. DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation. arXiv preprint arXiv:2208.12242 (2022).
  198. 198.Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. 2022. Palette: Image-to-image diffusion models. In Special Interest Group on Computer Graphics and Interactive Techniques Conference Proceedings. 1–10.
  199. 199.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J. Fleet, and Mohammad Norouzi. 2022. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487 (2022).
  200. 200.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J. Fleet, and Mohammad Norouzi. 2022. Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022).
  201. 201.Tim Salimans and Jonathan Ho. 2021. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations.
  202. 202.Tim Salimans and Jonathan Ho. 2021. Should EBMs model the energy or the score?. In International Conference on Learning Representations.
  203. 203.David Salinas, Michael Bohlke-Schneider, Laurent Callot, Roberto Medico, and Jan Gasthaus. 2019. High-dimensional multivariate forecasting with low-rank Gaussian copula processes. In Advances in Neural Information Processing Systems, Vol. 32.
  204. 204.David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski. 2020. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. International Journal of Forecasting 36, 3 (2020), 1181–1191.
  205. 205.Nikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen, and Aaron van den Oord. 2021. Step-unrolled denoising autoencoders for text generation. In International Conference on Learning Representations.
  206. 206.Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. 2008. The graph neural network model. IEEE Transactions on Neural Networks 20, 1 (2008), 61–80.
  207. 207.Thomas Schlegl, Philipp Seeböck, Sebastian M. Waldstein, Ursula Schmidt-Erfurth, and Georg Langs. 2017. Unsupervised anomaly detection with generative adversarial networks to guide marker discovery. In International Conference on Information Processing in Medical Imaging. 146–157.
  208. 208.Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting. 2023. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition. 22522–22531.
  209. 209.Vikash Sehwag, Caner Hazirbas, Albert Gordo, Firat Ozgenel, and Cristian Canton. 2022. Generating high fidelity data from low-density regions using diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition. 11492–11501.
  210. 210.Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina, Sihui Dai, Chong Xiang, Mung Chiang, and Prateek Mittal. 2021. Robust learning meets generative models: Can proxy distributions improve adversarial robustness?. In International Conference on Learning Representations.
  211. 211.Zhuchen Shao, Liuxi Dai, Yifeng Wang, Haoqian Wang, and Yongbing Zhang. 2023. AugDiff: Diffusion based feature augmentation for multiple instance learning in whole slide image. arXiv preprint arXiv:2303.06371 (2023).
  212. 212.Chence Shi, Shitong Luo, Minkai Xu, and Jian Tang. 2021. Learning gradient fields for molecular conformation generation. In International Conference on Machine Learning. 9558–9568.
  213. 213.Chence Shi, Minkai Xu, Zhaocheng Zhu, Weinan Zhang, Ming Zhang, and Jian Tang. 2020. GraphAF: A flow-based autoregressive model for molecular graph generation. arXiv preprint arXiv:2001.09382 (2020).
  214. 214.Yuyang Shi, Valentin De Bortoli, George Deligiannidis, and Arnaud Doucet. 2022. Conditional simulation using diffusion Schrödinger bridges. arXiv preprint arXiv:2202.13460 (2022).
  215. 215.Ikaro Silva, George Moody, Daniel J. Scott, Leo A. Celi, and Roger G. Mark. 2012. Predicting in-hospital mortality of ICU patients: The physionet/computing in cardiology challenge 2012. In 2012 Computing in Cardiology. 245–248.
  216. 216.Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, and Oran Gafni. 2022. Make-a-Video: Text-to-video generation without text-video data. In International Conference on Learning Representations.
  217. 217.John Skilling. 1989. The eigenvalues of mega-dimensional matrices. In Maximum Entropy and Bayesian Methods. 455–466.
  218. 218.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning. 2256–2265.
  219. 219.Xizewen Han, Huangjie Zheng, and Mingyuan Zhou. 2022. Card: Classification and regression diffusion models. Advances in Neural Information Processing Systems 35, (2022), 18100–18115.
  220. 220.Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion implicit models. In International Conference on Learning Representations.
  221. 221.Ki-Ung Song. 2022. Applying regularized Schrödinger-Bridge-based stochastic process in generative modeling. arXiv preprint arXiv:2208.07131 (2022).
  222. 222.Yang Song, Conor Durkan, Iain Murray, and Stefano Ermon. 2021. Maximum likelihood training of score-based diffusion models. In Advances in Neural Information Processing Systems, Vol. 34. 1415–1428.
  223. 223.Yang Song and Stefano Ermon. 2019. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems, Vol. 32.
  224. 224.Yang Song and Stefano Ermon. 2020. Improved techniques for training score-based generative models. In Advances in Neural Information Processing Systems, Vol. 33. 12438–12448.
  225. 225.Yang Song, Sahaj Garg, Jiaxin Shi, and Stefano Ermon. 2019. Sliced score matching: A scalable approach to density and score estimation. In The Conference on Uncertainty in Artificial Intelligence. 204.
  226. 226.Yang Song and Diederik P. Kingma. 2021. How to train your energy-based models. arXiv preprint arXiv:2101.03288 (2021).
  227. 227.Yang Song, Liyue Shen, Lei Xing, and Stefano Ermon. 2021. Solving inverse problems in medical imaging with score-based generative models. In International Conference on Learning Representations.
  228. 228.Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2020. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations.
  229. 229.James C. Spall. 2012. Stochastic optimization. In Handbook of Computational Statistics. 173–201.
  230. 230.Xuan Su, Jiaming Song, Chenlin Meng, and Stefano Ermon. 2022. Dual diffusion implicit bridges for image-to-image translation. In International Conference on Learning Representations.
  231. 231.Jaesung Tae, Hyeongju Kim, and Taesu Kim. 2021. EdiTTS: Score-based editing for controllable text-to-speech. arXiv preprint arXiv:2110.02584 (2021).
  232. 232.Huachun Tan, Guangdong Feng, Jianshuai Feng, Wuhong Wang, Yu-Jin Zhang, and Feng Li. 2013. A tensor-based method for missing traffic data completion. Transportation Research Part C: Emerging Technologies 28 (2013), 15–27.
  233. 233.Yusuke Tashiro, Jiaming Song, Yang Song, and Stefano Ermon. 2021. CSDI: Conditional score-based diffusion models for probabilistic time series imputation. In Advances in Neural Information Processing Systems, Vol. 34. 24804–24816.
  234. 234.Shantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Rémi Munos, Petar Veličković, and Michal Valko. 2021. Bootstrapped representation learning on graphs. arXiv preprint arXiv:2102.06514 (2021).
  235. 235.Lucas Theis, Aäron van den Oord, and Matthias Bethge. 2015. A note on the evaluation of generative models. arXiv preprint arXiv:1511.01844 (2015).
  236. 236.Brandon Trabucco, Kyle Doherty, Max Gurinas, and Ruslan Salakhutdinov. 2023. Effective data augmentation with diffusion models. In International Conference on Learning Representations.
  237. 237.Soobin Um and Jong Chul Ye. 2023. Don’t play favorites: Minority guidance for diffusion models. arXiv preprint arXiv:2301.12334 (2023).
  238. 238.Arash Vahdat, Karsten Kreis, and Jan Kautz. 2021. Score-based generative modeling in latent space. In Advances in Neural Information Processing Systems, Vol. 34. 11287–11302.
  239. 239.Aäron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu. 2016. Pixel recurrent neural networks. In International Conference on Machine Learning. 1747–1756.
  240. 240.Pascal Vincent. 2011. A connection between score matching and denoising autoencoders. Neural Computation 23, 7 (2011), 1661–1674.
  241. 241.Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. 2008. Extracting and composing robust features with denoising autoencoders. In International Conference on Machine Learning. 1096–1103.
  242. 242.Jinyi Wang, Zhaoyang Lyu, Dahua Lin, Bo Dai, and Hongfei Fu. 2022. Guided diffusion model for adversarial purification. arXiv preprint arXiv:2205.14969 (2022).
  243. 243.Yufei Wang, Jiayi Zheng, Can Xu, Xiubo Geng, Tao Shen, Chongyang Tao, and Daxin Jiang. 2022. KnowDA: All-in-one knowledge mixture model for data augmentation in few-shot NLP. arXiv preprint arXiv:2206.10265 (2022).
  244. 244.Zekai Wang, Tianyu Pang, Chao Du, Min Lin, Weiwei Liu, and Shuicheng Yan. 2023. Better diffusion models further improve adversarial training. arXiv preprint arXiv:2302.04638 (2023).
  245. 245.Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou. 2022. Diffusion-GAN: Training GANs with diffusion. arXiv preprint arXiv:2206.02262 (2022).
  246. 246.Daniel Watson, William Chan, Jonathan Ho, and Mohammad Norouzi. 2021. Learning fast samplers for diffusion models by differentiating through sample quality. In International Conference on Learning Representations.
  247. 247.Daniel Watson, Jonathan Ho, Mohammad Norouzi, and William Chan. 2021. Learning to efficiently sample from diffusion probabilistic models. arXiv preprint arXiv:2106.03802 (2021).
  248. 248.Jay Whang, Mauricio Delbracio, Hossein Talebi, Chitwan Saharia, Alexandros G. Dimakis, and Peyman Milanfar. 2022. Deblurring via stochastic refinement. In IEEE Conference on Computer Vision and Pattern Recognition. 16293–16303.
  249. 249.Hao Wu, Jonas Köhler, and Frank Noe. 2020. Stochastic normalizing flows. In Advances in Neural Information Processing Systems, Vol. 33. 5933–5944.
  250. 250.Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Weixian Lei, Yuchao Gu, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2022. Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. arXiv preprint arXiv:2212.11565 (2022).
  251. 251.Quanlin Wu, Hang Ye, and Yuntian Gu. 2022. Guided diffusion model for adversarial purification from random noise. arXiv preprint arXiv:2206.10875 (2022).
  252. 252.Shoule Wu and Ziqiang Shi. 2022. Itôwave: Itô stochastic differential equation is all you need for wave generation. In IEEE International Conference on Acoustics, Speech and Signal Processing. 8422–8426.
  253. 253.Shiwen Wu, Fei Sun, Wentao Zhang, Xu Xie, and Bin Cui. 2020. Graph neural networks in recommender systems: A survey. Comput. Surveys (2020).
  254. 254.Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, and S. Yu Philip. 2020. A comprehensive survey on graph neural networks. IEEE Transactions on Neural Networks and Learning Systems 32, 1 (2020), 4–24.
  255. 255.Julian Wyatt, Adam Leach, Sebastian M. Schmon, and Chris G. Willcocks. 2022. AnoDDPM: Anomaly detection with denoising diffusion probabilistic models using simplex noise. In IEEE Conference on Computer Vision and Pattern Recognition. 650–656.
  256. 256.Zhisheng Xiao, Karsten Kreis, and Arash Vahdat. 2021. Tackling the generative learning trilemma with denoising diffusion GANs. arXiv preprint arXiv:2112.07804 (2021).
  257. 257.Pan Xie, Qipeng Zhang, Zexian Li, Hao Tang, Yao Du, and Xiaohui Hu. 2022. Vector quantized diffusion model with CodeUnet for text-to-sign pose sequences generation. arXiv preprint arXiv:2208.09141 (2022).
  258. 258.Tian Xie, Xiang Fu, Octavian-Eugen Ganea, Regina Barzilay, and Tommi S. Jaakkola. 2021. Crystal diffusion variational autoencoder for periodic material generation. In International Conference on Learning Representations.
  259. 259.Yutong Xie and Quanzheng Li. 2022. Measurement-conditioned denoising diffusion probabilistic model for undersampled medical image reconstruction. arXiv preprint arXiv:2203.03623 (2022).
  260. 260.Minghao Xu, Hang Wang, Bingbing Ni, Hongyu Guo, and Jian Tang. 2021. Self-supervised graph-level representation learning with local and global structure. In International Conference on Machine Learning. 11548–11558.
  261. 261.Minkai Xu, Lantao Yu, Yang Song, Chence Shi, Stefano Ermon, and Jian Tang. 2021. GeoDiff: A geometric diffusion model for molecular conformation generation. In International Conference on Learning Representations.
  262. 262.Tijin Yan, Hongwei Zhang, Tong Zhou, Yufeng Zhan, and Yuanqing Xia. 2021. ScoreGrad: Multivariate probabilistic time series forecasting with continuous energy-based generative models. arXiv preprint arXiv:2106.10121 (2021).
  263. 263.Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu. 2022. Diffsound: Discrete diffusion model for text-to-sound generation. arXiv preprint arXiv:2207.09983 (2022).
  264. 264.Jie Yang, Ruijie Xu, Zhiquan Qi, and Yong Shi. 2021. Visual anomaly detection for images: A survey. arXiv preprint arXiv:2109.13157 (2021).
  265. 265.Kevin Yang and Dan Klein. 2021. FUDGE: Controlled text generation with future discriminators. arXiv preprint arXiv:2104.05218 (2021).
  266. 266.Ling Yang and Shenda Hong. 2022. Omni-granular ego-semantic propagation for self-supervised graph representation learning. arXiv preprint arXiv:2205.15746 (2022).
  267. 267.Ling Yang and Shenda Hong. 2022. Unsupervised time-series representation learning with iterative bilinear temporal-spectral fusion. In International Conference on Machine Learning. 25038–25054.
  268. 268.Ling Yang, Zhilin Huang, Yang Song, Shenda Hong, Guohao Li, Wentao Zhang, Bin Cui, Bernard Ghanem, and Ming-Hsuan Yang. 2022. Diffusion-based scene graph to image generation with masked contrastive pre-training. arXiv preprint arXiv:2211.11138 (2022).
  269. 269.Ling Yang, Liangliang Li, Zilun Zhang, Xinyu Zhou, Erjin Zhou, and Yu Liu. 2020. DPGN: Distribution propagation graph network for few-shot learning. In IEEE Conference on Computer Vision and Pattern Recognition. 13390–13399.
  270. 270.Ruihan Yang and Stephan Mandt. 2022. Lossy image compression with conditional diffusion models. arXiv preprint arXiv:2209.06950 (2022).
  271. 271.Ruihan Yang, Prakhar Srivastava, and Stephan Mandt. 2022. Diffusion probabilistic modeling for video generation. arXiv preprint arXiv:2203.09481 (2022).
  272. 272.Xiuwen Yi, Yu Zheng, Junbo Zhang, and Tianrui Li. 2016. ST-MVL: Filling missing values in geo-sensory time series data. In International Joint Conference on Artificial Intelligence.
  273. 273.Jongmin Yoon, Sung Ju Hwang, and Juho Lee. 2021. Adversarial purification with score-based generative models. In International Conference on Machine Learning. 12062–12072.
  274. 274.Jinsung Yoon, Daniel Jarrett, and Mihaela Van der Schaar. 2019. Time-series generative adversarial networks. In Advances in Neural Information Processing Systems, Vol. 32.
  275. 275.Peiyu Yu, Sirui Xie, Xiaojian Ma, Baoxiong Jia, Bo Pang, Ruiqi Gao, Yixin Zhu, Song-Chun Zhu, and Ying Nian Wu. 2022. Latent diffusion energy-based model for interpretable text modelling. In International Conference on Machine Learning. 25702–25720.
  276. 276.Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin. 2022. Generating videos with dynamics-aware implicit generative adversarial networks. arXiv preprint arXiv:2202.10571 (2022).
  277. 277.Lvmin Zhang and Maneesh Agrawala. 2023. Adding conditional control to text-to-image diffusion models. arXiv preprint arXiv:2302.05543 (2023).
  278. 278.Qinsheng Zhang and Yongxin Chen. 2021. Diffusion normalizing flow. In Advances in Neural Information Processing Systems, Vol. 34. 16280–16291.
  279. 279.Qinsheng Zhang and Yongxin Chen. 2022. Fast sampling of diffusion models with exponential integrator. arXiv preprint arXiv:2204.13902 (2022).
  280. 280.Qinsheng Zhang, Molei Tao, and Yongxin Chen. 2022. gDDIM: Generalized denoising diffusion implicit models. arXiv preprint arXiv:2206.05564 (2022).
  281. 281.Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022. OPT: Open pre-trained transformer language models. arXiv preprint arXiv:2205.01068 (2022).
  282. 282.Wenrui Zhang, Ling Yang, Shijia Geng, and Shenda Hong. 2022. Cross reconstruction transformer for self-supervised time series representation learning. arXiv preprint arXiv:2205.09928 (2022).
  283. 283.Min Zhao, Fan Bao, Chongxuan Li, and Jun Zhu. 2022. EGSDE: Unpaired image-to-image translation via energy-guided stochastic differential equations. arXiv preprint arXiv:2207.06635 (2022).
  284. 284.Yue Zhao, Zain Nasrullah, and Zheng Li. 2019. PyOD: A Python toolbox for scalable outlier detection. Journal of Machine Learning Research 20 (2019), 1–7.
  285. 285.Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou. 2022. Truncated diffusion probabilistic models. arXiv preprint arXiv:2202.09671 (2022).
  286. 286.Jie Zhou, Ganqu Cui, Shengding Hu, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun. 2020. Graph neural networks: A review of methods and applications. AI Open 1 (2020), 57–81.
  287. 287.Linqi Zhou, Yilun Du, and Jiajun Wu. 2021. 3D shape generation and completion through point-voxel diffusion. In International Conference on Computer Vision. 5826–5835.
  288. 288.Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan. 2022. Discrete contrastive diffusion for cross-modal and conditional generation. arXiv preprint arXiv:2206.07771 (2022).
  289. 289.Yanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu, Shu Wu, and Liang Wang. 2020. Deep graph contrastive representation learning. arXiv preprint arXiv:2006.04131 (2020).
  290. 290.Roland S. Zimmermann, Lukas Schott, Yang Song, Benjamin A. Dunn, and David A. Klindt. 2021. Score-based generative classifiers. arXiv preprint arXiv:2110.00473 (2021).

Citation

MLA
Yang, L., et al. “Diffusion Models: A Comprehensive Survey of Methods and Applications”. arXiv, 2022, http://arxiv.org/abs/2209.00796v15.
APA
Yang, L., Zhang, Z., Song, Y., Hong, S., Xu, R., Zhao, Y., Zhang, W., Cui, B., & Yang, M.-H. (2022). Diffusion Models: A Comprehensive Survey of Methods and Applications. arXiv. http://arxiv.org/abs/2209.00796v15
Chicago
Yang, L., Z. Zhang, Y. Song, et al. 2022. “Diffusion Models: A Comprehensive Survey of Methods and Applications”. arXiv. http://arxiv.org/abs/2209.00796v15.
Harvard
Yang, L. et al. (2022) “Diffusion Models: A Comprehensive Survey of Methods and Applications”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2209.00796v15.
Vancouver
1. Yang L, Zhang Z, Song Y, Hong S, Xu R, Zhao Y, Zhang W, Cui B, Yang M-H (2022) Diffusion Models: A Comprehensive Survey of Methods and Applications. arXiv

BibTeX

@article{yang2022diffusion,
  title = {Diffusion Models: A Comprehensive Survey of Methods and Applications},
  author = {Yang, Ling and Zhang, Zhilong and Song, Yang and Hong, Shenda and Xu, Runsheng and Zhao, Yue and Zhang, Wentao and Cui, Bin and Yang, Ming-Hsuan},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2209.00796v15},
  eprint = {2209.00796}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF