Neural Diffusion Models

Grigory BartoshDmitry P. VetrovChristian A. Naesseth

article2024ICML68 citations

Introduces Neural Diffusion Models, a generative framework that replaces rigid linear noising processes with learnable, time-dependent non-linear data transformations to achieve tighter likelihood bounds and state-of-the-art density estimation on image benchmarks.

Listen

Generative diffusion models have demonstrated remarkable success across domains such as computer vision, audio synthesis, and molecular biology. However, conventional diffusion frameworks rely on a rigid, pre-specified noise-injection process that applies simple linear scaling to data before adding Gaussian noise. This inability to adapt the forward data-corruption path limits latent space flexibility, complicates the reverse generative process, and restricts density estimation performance.

The article aims to overcome these constraints by introducing Neural Diffusion Models, a generalized framework that incorporates learnable, time-dependent, and non-linear data transformations into the forward process. The objective is to evaluate whether parameterizing these transformations with neural networks improves likelihood estimation and generative dynamics without sacrificing the simulation-free computational efficiency of modern diffusion models.

To evaluate this framework, the authors formulated both discrete-time and continuous-time versions of Neural Diffusion Models, deriving exact time derivatives using Jacobian-vector products to maintain mathematically rigorous optimization. They evaluated the approach across 2D synthetic benchmarks and standard image datasets, including MNIST, CIFAR-10, downsampled ImageNet at 32x32 and 64x64 resolutions, and CelebA-HQ at 256x256 resolution. Models were benchmarked against standard denoising diffusion baselines under controlled conditions, utilizing identical neural network architectures, variance schedules, and computational budgets.

The experimental results reveal five primary findings. First, Neural Diffusion Models consistently achieved superior density estimation, setting new state-of-the-art diffusion likelihood marks on ImageNet 32x32 (3.55 bits per dimension), ImageNet 64x64 (3.35 bits per dimension), and CelebA-HQ. Second, the learned transformations actively simplify data geometry during training—for instance, increasing contrast in natural images or separating complex patterns—allowing the model to produce terminal data predictions that align much more closely with the underlying data distribution. Third, in few-step generation regimes (e.g., 10 discrete steps), the proposed model substantially outperformed standard diffusion in sample quality, improving the CIFAR-10 quality score from 37.83 to 31.56. Fourth, when integrated with latent-space models like LSGM, the framework reduced negative log-likelihood on CelebA-HQ while maintaining competitive visual fidelity. Fifth, ablation studies confirmed that these performance gains stem directly from the learned non-linear transformations rather than the increased parameter count.

These findings indicate that adapting the forward corruption process to the data distribution bridges the gap between true data likelihood and the variational lower bound. For practical applications that depend heavily on accurate probability modeling—such as data compression, anomaly detection, semi-supervised learning, and adversarial defense—Neural Diffusion Models offer a significant performance upgrade over standard diffusion techniques with fixed corruption dynamics.

Organizations evaluating this approach should consider adoption for tasks where density estimation and few-step generation quality are paramount. However, decision-makers must account for trade-offs: learning the transformation roughly doubles model parameters and increases training time by approximately 2.3 times compared to baseline diffusion. Additionally, the learned forward dynamics prevent direct use of standard classifier guidance, and the models exhibit greater performance degradation when post-hoc reducing sampling steps from a model trained on many steps. Future work should focus on developing efficient parameterizations, enabling classifier-free or alternative conditional sampling methods, and extending restricted reverse dynamics to high-dimensional datasets.

  • Paper: Score-Based Diffusion Models in Function Space, Jae Hyun Lim 0001 et al. (2025). Building on Neural Diffusion Models’ flexible forward corruption, this work carries score-based diffusion into function spaces where the data are continuous fields rather than finite vectors.
Cover for Neural Diffusion Models

Abstract

Diffusion models have shown remarkable performance on many generative tasks. Despite recent success, most diffusion models are restricted in that they only allow linear transformation of the data distribution. In contrast, broader family of transformations can help train generative distributions more efficiently, simplifying the reverse process and closing the gap between the true negative log-likelihood and the variational approximation. In this paper, we present Neural Diffusion Models (NDMs), a generalization of conventional diffusion models that enables defining and learning time-dependent non-linear transformations of data. We show how to optimise NDMs using a variational bound in a simulation-free setting. Moreover, we derive a time-continuous formulation of NDMs, which allows fast and reliable inference using off-the-shelf numerical ODE and SDE solvers. Finally, we demonstrate the utility of NDMs through experiments on many image generation benchmarks, including MNIST, CIFAR-10, downsampled versions of ImageNet and CelebA-HQ. NDMs outperform conventional diffusion models in terms of likelihood, achieving state-of-the-art results on ImageNet and CelebA-HQ, and produces high-quality samples.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 3. Neural diffusion models
  • 3.1. Model definition and variational objective
  • 3.2. Continuous time NDMs
  • 4. Experiments
  • 4.1. Implementation details
  • 4.2. Learned transformations
  • 4.3. Image generation
  • 5. Related work
  • 6. Limitations
  • 7. Conclusion
  • Impact Statement
  • References
  • A. Derivations and proofs
  • A.1. Forward posterior
  • A.2. Objective
  • A.3. Reverse SDE and ODE
  • A.4. Continuous time objective
  • B. Connections with other works
  • B.1. Diffusion in latent space
  • B.2. Schrödinger Bridges
  • B.3. Stochastic Interpolants
  • B.4. DiffEnc
  • C. Implementation details
  • C.1. Dequantization
  • C.2. Parameterization
  • C.3. Diffusion in latent space
  • D. Additional results
  • D.1. Additional evaluation
  • D.2. Additional samples
  • D.3. Ablation studies
  • E. Dynamic optimal transport
  • E.1. Restricted reverse process
  • E.2. Objective function
  • E.3. Results and discussion

Knowls

  1. Knowl 1 — Neural diffusion uses learnable transformed data marginals

    definition

    A Neural Diffusion Model (NDM) defines the forward latent at time tt by first applying a time-dependent transformation to the data and then adding Gaussian noise. For data x∈Rdx\in\mathbb{R}^d, latent zt∈Rdz_t\in\mathbb{R}^d, time t∈[0,T]t\in[0,T], scalar schedules αt\alpha_t and σt\sigma_t, and a transformation Fϕ:Rd×[0,T]→RdF_\phi:\mathbb{R}^d\times[0,T]\to\mathbb{R}^d parameterized by ϕ\phi, the marginal is

    qϕ(zt∣x)=N ⁣(zt;αtFϕ(x,t),σt2Id),q_\phi(z_t\mid x)=\mathcal{N}\!\left(z_t;\alpha_tF_\phi(x,t),\sigma_t^2 I_d\right),

    where IdI_d is the dd-dimensional identity matrix. For 0≤s≤t≤T0\le s\le t\le T, a compatible conditional posterior is

    qϕ(zs∣zt,x)=N ⁣(zs;μs∣tF,σ~s∣t 2Id),q_\phi(z_s\mid z_t,x)=\mathcal{N}\!\left(z_s;\mu^F_{s\mid t},\widetilde\sigma_{s\mid t}^{\,2}I_d\right),

    μs∣tF=αsFϕ(x,s)+σs2−σ~s∣t 2σt(zt−αtFϕ(x,t)),\mu^F_{s\mid t}=\alpha_sF_\phi(x,s)+\frac{\sqrt{\sigma_s^2-\widetilde\sigma_{s\mid t}^{\,2}}}{\sigma_t}\left(z_t-\alpha_tF_\phi(x,t)\right),

    with design choice 0≤σ~s∣t 2≤σs20\le\widetilde\sigma_{s\mid t}^{\,2}\le\sigma_s^2. These marginals and posteriors define an implicit, generally non-Markovian forward process: each latent can be sampled directly given xx, without simulating all preceding latents. Conventional diffusion with the identity transform is a special case.

  2. Knowl 2 — A simulation-free variational bound trains the transformed forward process

    equation

    For data distribution q(x)q(x), the NDM minimizes a variational upper bound on expected negative log-likelihood. The forward joint qϕ(z0:T∣x)q_\phi(z_{0:T}\mid x) is defined by the NDM marginals and posteriors; the reverse model is pθ(z0:T)=p(zT)∏t=1Tpθ(zt−1∣zt)p_\theta(z_{0:T})=p(z_T)\prod_{t=1}^{T}p_\theta(z_{t-1}\mid z_t) with p(zT)=N(0,Id)p(z_T)=\mathcal{N}(0,I_d). The bound is

    L=Ex∼q, z0:T∼qϕ(⋅∣x) ⁣[DKL ⁣(qϕ(zT∣x) ∥ p(zT))−log⁡pθ(x∣z0)+∑t=1TDKL ⁣(qϕ(zt−1∣zt,x) ∥ pθ(zt−1∣zt))].\mathcal{L}=\mathbb{E}_{x\sim q,\,z_{0:T}\sim q_\phi(\cdot\mid x)}\!\left[ D_{\mathrm{KL}}\!\left(q_\phi(z_T\mid x)\,\|\,p(z_T)\right)-\log p_\theta(x\mid z_0)+\sum_{t=1}^{T}D_{\mathrm{KL}}\!\left(q_\phi(z_{t-1}\mid z_t,x)\,\|\,p_\theta(z_{t-1}\mid z_t)\right)\right].

    Here DKLD_{\mathrm{KL}} is Kullback–Leibler divergence, θ\theta parameterizes the reverse model, and pθ(x∣z0)p_\theta(x\mid z_0) is the reconstruction distribution. In the standard reverse parameterization, pθ(zs∣zt)=qϕ(zs∣zt,x^θ(zt,t))p_\theta(z_s\mid z_t)=q_\phi(z_s\mid z_t,\widehat x_\theta(z_t,t)), where x^θ\widehat x_\theta predicts data. The diffusion-term divergence then has a closed form:

    DKL ⁣(qϕ(zs∣zt,x) ∥ pθ(zs∣zt))=12σ~s∣t 2∥αs[Fϕ(x,s)−Fϕ(x^θ(zt,t),s)]+αtσs2−σ~s∣t 2σt[Fϕ(x^θ(zt,t),t)−Fϕ(x,t)]∥22.D_{\mathrm{KL}}\!\left(q_\phi(z_s\mid z_t,x)\,\|\,p_\theta(z_s\mid z_t)\right)=\frac{1}{2\widetilde\sigma_{s\mid t}^{\,2}}\left\|\alpha_s\left[F_\phi(x,s)-F_\phi(\widehat x_\theta(z_t,t),s)\right]+\frac{\alpha_t\sqrt{\sigma_s^2-\widetilde\sigma_{s\mid t}^{\,2}}}{\sigma_t}\left[F_\phi(\widehat x_\theta(z_t,t),t)-F_\phi(x,t)\right]\right\|_2^2.

    Thus the reverse network is trained through errors in transformed data, not merely through an untransformed data prediction. Because FϕF_\phi is learned, the prior and reconstruction terms can also depend on ϕ\phi and must be retained unless the model is constructed to remove the reconstruction term. Training is simulation-free: the conditional latent ztz_t can be sampled directly and the time index can be sampled rather than simulating the full chain or summing all time-step divergences at every update.

  3. Knowl 3 — Continuous-time NDM reverse dynamics enable ODE and SDE inference

    equation

    In the continuous-time formulation, time tt lies in [0,1][0,1], the transformation Fϕ(x,t)F_\phi(x,t) is differentiable in time, and the reverse process runs from t=1t=1 toward t=0t=0. Let F˙ϕ(x,t)=∂Fϕ(x,t)/∂t\dot F_\phi(x,t)=\partial F_\phi(x,t)/\partial t, r(t)=∂tlog⁡αtr(t)=\partial_t\log\alpha_t, and let νt\nu_t parameterize posterior noise so that σ~s∣t 2=σs2(1−eνs−νt)\widetilde\sigma_{s\mid t}^{\,2}=\sigma_s^2(1-e^{\nu_s-\nu_t}). Define g2(t)=ν˙tσt2g^2(t)=\dot\nu_t\sigma_t^2 and the model score-like quantity

    sθ(zt,t)=αtFϕ(x^θ(zt,t),t)−ztσt2.s_\theta(z_t,t)=\frac{\alpha_tF_\phi(\widehat x_\theta(z_t,t),t)-z_t}{\sigma_t^2}.

    The reverse stochastic differential equation is

    dzt=[αtF˙ϕ(x^θ(zt,t),t)+r(t)zt−12(g2(t)−2r(t)σt2)sθ(zt,t)]dt+g(t)dwt,dz_t=\left[\alpha_t\dot F_\phi(\widehat x_\theta(z_t,t),t)+r(t)z_t-\frac{1}{2}\left(g^2(t)-2r(t)\sigma_t^2\right)s_\theta(z_t,t)\right]dt+g(t)dw_t,

    where wtw_t is Brownian motion. The time derivative of the learned transform can be evaluated with a Jacobian-vector product. Choosing constant νt\nu_t gives deterministic dynamics representable by an ordinary differential equation (ODE); the authors use the ODE as a continuous normalizing flow for inference and density estimation. The formulation also permits inference with standard numerical SDE or ODE solvers.

  4. Knowl 4 — The continuous-time diffusion loss compares forward and predicted transformed dynamics

    equation

    For differentiable FϕF_\phi, continuous-time NDM training replaces the discrete sum of diffusion divergences with an integral over time. With x∼q(x)x\sim q(x), zt∼qϕ(zt∣x)z_t\sim q_\phi(z_t\mid x), r(t)=∂tlog⁡αtr(t)=\partial_t\log\alpha_t, g2(t)=ν˙tσt2g^2(t)=\dot\nu_t\sigma_t^2, and

    s(x,zt,t)=αtFϕ(x,t)−ztσt2,s(x,z_t,t)=\frac{\alpha_tF_\phi(x,t)-z_t}{\sigma_t^2},

    the diffusion contribution is

    Ldiff=Ex∼q ⁣[∫01Ezt∼qϕ(⋅∣x) ⁣[1g2(t)∥αt(F˙ϕ(x,t)−F˙ϕ(x^θ(zt,t),t))+12(∂tσt2−2r(t)σt2+g2(t))(s(x,zt,t)−s(x^θ(zt,t),zt,t))∥22]dt].\mathcal{L}_{\mathrm{diff}}=\mathbb{E}_{x\sim q}\!\left[\int_0^1\mathbb{E}_{z_t\sim q_\phi(\cdot\mid x)}\!\left[\frac{1}{g^2(t)}\left\|\alpha_t\left(\dot F_\phi(x,t)-\dot F_\phi(\widehat x_\theta(z_t,t),t)\right)+\frac{1}{2}\left(\partial_t\sigma_t^2-2r(t)\sigma_t^2+g^2(t)\right)\left(s(x,z_t,t)-s(\widehat x_\theta(z_t,t),z_t,t)\right)\right\|_2^2\right]dt\right].

    Here x^θ(zt,t)\widehat x_\theta(z_t,t) is the reverse model's data prediction, and the second score-like term uses the same definition as ss with xx replaced by that prediction. In practice, the continuous-time model samples time using importance sampling proportional to 1/g2(t)1/g^2(t). Exact time derivatives of the transformation are obtained with Jacobian-vector products rather than approximating those derivatives.

  5. Knowl 5 — The experimental transform is anchored to the identity at time zero

    model/method

    For image experiments, the learned transformation is constrained as

    Fϕ(x,t)=(1−t)x+tF‾ϕ(x,t),F_\phi(x,t)=(1-t)x+t\overline F_\phi(x,t),

    where F‾ϕ\overline F_\phi is a neural network. This makes Fϕ(x,0)=xF_\phi(x,0)=x, so the initial latent distribution satisfies qϕ(z0∣x)≈δ(z0−x)q_\phi(z_0\mid x)\approx\delta(z_0-x) and the reconstruction term can be omitted. The authors use a U-Net for image transformations, the DDPM variance-preserving schedules for αt\alpha_t and σt\sigma_t, and choose

    σ~s∣t 2=(σt2−αt2αs2σs2)σs2σt2,\widetilde\sigma_{s\mid t}^{\,2}=\left(\sigma_t^2-\frac{\alpha_t^2}{\alpha_s^2}\sigma_s^2\right)\frac{\sigma_s^2}{\sigma_t^2},

    which makes the NDM and DDPM forward processes consistent when the transform is the identity. In experiments the reverse network predicts injected-noise parameter ϵ^θ\widehat\epsilon_\theta and converts it to a data prediction by

    x^θ(zt,t)=zt−σtϵ^θ(zt,t)αt.\widehat x_\theta(z_t,t)=\frac{z_t-\sigma_t\widehat\epsilon_\theta(z_t,t)}{\alpha_t}.

    Because this conversion does not account for FϕF_\phi, ϵ^θ\widehat\epsilon_\theta need not equal the actual injected noise.

  6. Knowl 6 — Several diffusion families are special choices of the NDM transformation

    model/method

    The NDM marginal qϕ(zt∣x)=N(zt;αtFϕ(x,t),σt2Id)q_\phi(z_t\mid x)=\mathcal{N}(z_t;\alpha_tF_\phi(x,t),\sigma_t^2I_d) recovers multiple established constructions through choices of FϕF_\phi and the noise schedules. DDPM, DDIM, and variational diffusion models use the identity transform Fϕ(x,t)=xF_\phi(x,t)=x with their respective schedules. Flow Matching with an optimal-transport path also uses the identity transform, with αt=t\alpha_t=t and σt=1−(1−σmin⁡)t\sigma_t=1-(1-\sigma_{\min})t. Soft Diffusion corresponds to a time-dependent linear transform Fϕ(x,t)=CtxF_\phi(x,t)=C_tx, with αt=1\alpha_t=1 and σt2=st2\sigma_t^2=s_t^2. Latent score-based generative models use a time-independent encoder, Fϕ(x,t)=E(x)F_\phi(x,t)=E(x), and specify a decoder distribution for reconstructing xx from z0z_0. These examples illustrate that the NDM framework includes fixed linear corruption, selected time-dependent linear transformations, and nonlinear data-space mappings.

  7. Knowl 7 — Learned transformations reshape data and improve intermediate predictions

    empirical result

    On a two-dimensional checkerboard, the learned transformation turns the interleaved pattern into a non-interleaved one. On MNIST, it emphasizes digit-defining strokes, thickens lines, and can form bubbles near corners; on CIFAR-10, it increases image contrast. These observed transformations simplify the distributions presented to the reverse model. In a separate MNIST comparison, both DDPM and NDM receive terminal latents sampled from a standard normal distribution, but NDM's predicted data points at that time resemble real MNIST digits more closely than DDPM's predictions. The paper attributes this behavior to the NDM objective's dependence on the transformed prediction Fϕ(x^θ,t)F_\phi(\widehat x_\theta,t), which favors data predictions that remain plausible after transformation.

  8. Knowl 8 — Continuous-time likelihood benchmarks improve over diffusion baselines on ImageNet

    data/table

    The table compares reported test negative log-likelihood in bits per dimension (BPD) for image benchmarks. NDM likelihoods were evaluated using the continuous-time formulation and the corresponding ODE integrated with RK45. Lower values are better. NDM is competitive on CIFAR-10 and has the lowest listed diffusion-model value on both ImageNet resolutions; the CIFAR-10 VDM result is lower than NDM's.

    Model CIFAR-10 ImageNet 32 ImageNet 64
    DDPM 3.69 – –
    Improved DDPM 2.94 – 3.54
    VDM 2.65 3.72 3.40
    Score SDE 2.99 – –
    Score Flow 2.83 3.76 –
    NDM 2.70 3.55 3.35
  9. Knowl 9 — Matched DDPM comparisons show better NLL but mixed sampling quality

    data/table

    These controlled comparisons use the same objective, architecture, and hyperparameters for DDPM and NDM; DDPM is implemented as the identity-transform NDM. The table reports test NLL and NELBO in BPD, plus FID for DDPM-style sampling and FID* for DDIM-style sampling; lower is better. For training and sampling with 1,000 or 10 steps, NDM improves NLL and NELBO on CIFAR-10 and ImageNet 32, while sample-quality metrics are mixed. When models trained at 1,000 steps are sampled with 10 steps, NDM retains better likelihood but has worse DDPM-style FID on both datasets, illustrating reduced robustness to this step-count change.

    Steps Model CIFAR-10 ImageNet 32
    NLL NELBO FID FID* NLL NELBO FID FID*
    1000 DDPM 3.11 3.18 11.44 13.35 3.89 3.95 16.18 19.08
    1000 NDM 3.02 3.03 11.82 13.79 3.79 3.82 17.02 19.76
    10 DDPM 5.02 5.13 37.83 19.89 6.28 6.42 53.51 26.47
    10 NDM 4.63 4.74 31.56 22.20 5.81 5.94 45.38 29.95
    100010 DDPM 8.78 8.98 43.85 17.73 10.99 11.23 58.35 25.53
    100010 NDM 8.58 8.81 48.41 16.96 10.78 11.06 62.12 23.77
  10. Knowl 10 — Latent NDM improves the reported CelebA-HQ likelihood and FID

    empirical result

    The authors replace the linear diffusion process in the LSGM baseline with an NDM having a learnable transformation in the VAE latent space, keeping the baseline setup and transformation/diffusion network architecture consistent. On CelebA-HQ-256, LSGM reports NLL ≤0.70\le 0.70 and FID 7.227.22, while latent NDM reports NLL ≤0.65\le 0.65 and FID 7.187.18. Thus the reported latent NDM improves both the likelihood bound and sample-quality score in this experiment.

  11. Knowl 11 — A constrained NDM can learn one-dimensional optimal-transport trajectories

    empirical result

    In a proof-of-concept experiment on a one-dimensional Gaussian-mixture distribution, the reverse process is restricted to linear trajectories

    zt=(1−t)x^θ(ϵ)+tϵ,z_t=(1-t)\widehat x_\theta(\epsilon)+t\epsilon,

    where ϵ\epsilon is standard Gaussian noise and x^θ\widehat x_\theta is constrained to be monotonically increasing. Under the paper's smooth one-dimensional setting, these trajectories represent dynamic optimal transport between the Gaussian and the distribution induced by x^θ\widehat x_\theta. Jointly learning the NDM forward transformation and this restricted reverse process produces straight transport trajectories. A DDPM with its fixed forward process cannot match that restriction and instead produces curved trajectories. The demonstration is limited to one dimension; inverting the trajectory map required five Newton iterations.

  12. Knowl 12 — Learnable transforms increase training cost and constrain standard guidance

    limitation

    NDMs with learnable transformations have approximately twice as many parameters as conventional diffusion models and took about 2.3 times longer than DDPM to train in the image experiments. The authors report that simply doubling DDPM parameters did not reproduce NDM's likelihood improvements. Training must also use the full NDM variational objective: a simplified DDPM-style noise-prediction loss that ignores the transformation can cause FϕF_\phi to collapse to zero. Finally, because the generative process depends on the learned forward-process parameters, the paper states that standard classifier guidance is not supported for learnable-transform NDMs; alternative conditional-generation methods are left for future work.

Coverage note — Additional ablations, intermediate-step sample panels, and detailed optimizer and hardware settings were omitted because the principal quantitative comparisons and methodological conditions are captured in the knowls above.

References

  1. 1.Albergo, M. S. and Vanden-Eijnden, E. Building normalizing flows with stochastic interpolants. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=li7qeBbCR1t.
  2. 2.Chen, R. T., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. Advances in neural information processing systems, 31, 2018.
  3. 3.Chen, T., Liu, G.-H., and Theodorou, E. A. Likelihood training of schr\” odinger bridge using forward-backward sdes theory. arXiv preprint arXiv:2110.11291, 2021.
  4. 4.Creswell, A., White, T., Dumoulin, V., Arulkumaran, K., Sengupta, B., and Bharath, A. A. Generative adversarial networks: An overview. IEEE Signal Processing Magazine, 35(1):53–65, 2018.
  5. 5.Dai, Z., Yang, Z., Yang, F., Cohen, W. W., and Salakhutdinov, R. R. Good semi-supervised learning that requires a bad gan. Advances in neural information processing systems, 30, 2017.
  6. 6.Daras, G., Delbracio, M., Talebi, H., Dimakis, A. G., and Milanfar, P. Soft diffusion: Score matching for general corruptions. arXiv preprint arXiv:2209.05442, 2022.
  7. 7.De Bortoli, V., Thornton, J., Heng, J., and Doucet, A. Diffusion schrodinger bridge with applications to score-based generative modeling. Advances in Neural Information Processing Systems, 34:17695–17709, 2021.
  8. 8.Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee, 2009.
  9. 9.Deng, L. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE signal processing magazine, 29(6):141–142, 2012.
  10. 10.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34:8780–8794, 2021.
  11. 11.Dormand, J. R. and Prince, P. J. A family of embedded runge-kutta formulae. Journal of computational and applied mathematics, 6(1):19–26, 1980.
  12. 12.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial nets. Advances in neural information processing systems, 27, 2014.
  13. 13.Grathwohl, W., Chen, R. T. Q., Bettencourt, J., and Duvenaud, D. Scalable reversible generative models with free-form continuous dynamics. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rJxgknCcK7.
  14. 14.Gu, J., Zhai, S., Zhang, Y., Bautista, M. A., and Susskind, J. M. f-DM: A multi-stage diffusion model via progressive signal transformation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=iBdwKIsg4m.
  15. 15.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  16. 16.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020.
  17. 17.Hoogeboom, E. and Salimans, T. Blurring diffusion models. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=OjDkC57x5sz.
  18. 18.Karras, T., Aila, T., Laine, S., and Lehtinen, J. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196, 2017.
  19. 19.Kim, D., Na, B., Kwon, S. J., Lee, D., Kang, W., and Moon, I.-c. Maximum likelihood training of implicit nonlinear diffusion model. Advances in Neural Information Processing Systems, 35:32270–32284, 2022.
  20. 20.Kingma, D., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. Advances in neural information processing systems, 34:21696–21707, 2021.
  21. 21.Kingma, D. P. and Welling, M. Auto-encoding variational Bayes. International Conference on Learning Representations, 2014.
  22. 22.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  23. 23.Lee, S., Kim, B., and Ye, J. C. Minimizing trajectory curvature of ode-based generative models. In International Conference on Machine Learning, pp. 18957–18973. PMLR, 2023.
  24. 24.Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=PqvMRDCJT9t.
  25. 25.Liu, L., Ren, Y., Lin, Z., and Zhao, Z. Pseudo numerical methods for diffusion models on manifolds. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=PlKWVd2yBkY.
  26. 26.Liu, Q. Rectified flow: A marginal preserving approach to optimal transport. arXiv preprint arXiv:2209.14577, 2022.
  27. 27.Liu, X., Gong, C., and qiang liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=XVjTT1nw5z.
  28. 28.MacKay, D. J. Information theory, inference and learning algorithms. Cambridge university press, 2003.
  29. 29.Neklyudov, K., Severo, D., and Makhzani, A. Action matching: A variational method for learning stochastic dynamics from samples. arXiv preprint arXiv:2210.06662, 2022.
  30. 30.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International conference on machine learning, pp. 8162–8171. PMLR, 2021.
  31. 31.Nielsen, B. M. G., Christensen, A., Dittadi, A., and Winther, O. Diffenc: Variational diffusion with a learned encoder. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=8nxy1bQWTG.
  32. 32.Øksendal, B. and Øksendal, B. Stochastic differential equations. Springer, 2003.
  33. 33.Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B. Normalizing flows for probabilistic modeling and inference. The Journal of Machine Learning Research, 22(1):2617–2680, 2021.
  34. 34.Peluchetti, S. Non-denoising forward-time diffusions.
  35. 35.Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., and Kudinov, M. Grad-tts: A diffusion probabilistic model for text-to-speech. In International Conference on Machine Learning, pp. 8599–8608. PMLR, 2021.
  36. 36.Rezende, D. J., Mohamed, S., and Wierstra, D. Stochastic backpropagation and approximate inference in deep generative models. In International conference on machine learning, pp. 1278–1286. PMLR, 2014.
  37. 37.Rissanen, S., Heinonen, M., and Solin, A. Generative modelling with inverse heat dissipation. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=4PJUBT9f2Ol.
  38. 38.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 10684–10695, 2022.
  39. 39.Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. arXiv preprint arXiv:2104.07636, 2021.
  40. 40.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=TIdIXIpzhoI.
  41. 41.Singhal, R., Goldstein, M., and Ranganath, R. Where to diffuse, how to diffuse, and how to get back: Automated learning for multivariate diffusions. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=osei3IzUia.
  42. 42.Smale, S. and Hirsch, M. W. Differential equations, dynamical systems, and linear algebra, volume 60. Elsevier, 1974.
  43. 43.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265. PMLR, 2015.
  44. 44.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021a. URL https://openreview.net/forum?id=St1giarCHLP.
  45. 45.Song, Y., Kim, T., Nowozin, S., Ermon, S., and Kushman, N. Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. arXiv preprint arXiv:1710.10766, 2017.
  46. 46.Song, Y., Durkan, C., Murray, I., and Ermon, S. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34:1415–1428, 2021b.
  47. 47.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021c. URL https://openreview.net/forum?id=PxTIG12RRHS.
  48. 48.Tachibana, H., Go, M., Inahara, M., Katayama, Y., and Watanabe, Y. It\ˆ{o}-taylor sampling scheme for denoising diffusion probabilistic models using ideal derivatives. arXiv preprint arXiv:2112.13339, 2021.
  49. 49.Tomczak, J. M. Deep generative modeling. Springer, 2022.
  50. 50.Trippe, B. L., Yim, J., Tischer, D., Baker, D., Broderick, T., Barzilay, R., and Jaakkola, T. S. Diffusion probabilistic modeling of protein backbones in 3D for the motif-scaffolding problem. In The Eleventh International Conference on Learning Representations, 2023.
  51. 51.Vahdat, A., Kreis, K., and Kautz, J. Score-based generative modeling in latent space. Advances in Neural Information Processing Systems, 34:11287–11302, 2021.
  52. 52.Van Den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K. Pixel recurrent neural networks. In International conference on machine learning, pp. 1747–1756. PMLR, 2016.
  53. 53.Wang, G., Jiao, Y., Xu, Q., Wang, Y., and Yang, C. Deep generative learning via schrodinger bridge. In International Conference on Machine Learning, pp. 10794–10804. PMLR, 2021.
  54. 54.Watson, J. L., Juergens, D., Bennett, N. R., Trippe, B. L., Yim, J., Eisenach, H. E., Ahern, W., Borst, A. J., Ragotte, R. J., Milles, L. F., et al. Broadly applicable and accurate protein design by integrating structure prediction networks and diffusion generative models. bioRxiv, pp. 2022–12, 2022.
  55. 55.Wu, L., Trippe, B. L., Naesseth, C. A., Blei, D. M., and Cunningham, J. P. Practical and asymptotically exact conditional sampling in diffusion models. arXiv preprint arXiv:2306.17775, 2023.
  56. 56.Xiao, Z., Kreis, K., and Vahdat, A. Tackling the generative learning trilemma with denoising diffusion GANs. arXiv preprint arXiv:2112.07804, 2021.
  57. 57.Yang, L., Zhang, Z., Song, Y., Hong, S., Xu, R., Zhao, Y., Shao, Y., Zhang, W., Cui, B., and Yang, M.-H. Diffusion models: A comprehensive survey of methods and applications. arXiv preprint arXiv:2209.00796, 2022.

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/