Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs

Kaiwen ZhengCheng LuJianfei ChenJun Zhu

article2023ICML64 citations

Proposes a suite of training and evaluation techniques—including velocity parameterization, high-order flow matching finetuning, and training-free truncated-normal dequantization—that enables diffusion ODEs to achieve state-of-the-art exact likelihood estimation without variational dequantization or data augmentation.

Listen

Generative modeling and density estimation play critical roles in key data applications such as lossless data compression, anomaly detection, and out-of-distribution identification. Within deep generative modeling, diffusion models formulated as probability flow ordinary differential equations (diffusion ODEs) enable deterministic inference and exact likelihood evaluation. However, existing diffusion ODEs have consistently trailed stochastic counterparts, such as Variational Diffusion Models, in likelihood estimation performance. This lag primarily stems from slow training convergence, numerical instabilities, and significant mismatches between continuous training models and discrete real-world evaluation data.

The article develops a comprehensive framework called Improved Diffusion ODE (i-DODE) to advance maximum likelihood estimation in diffusion ODEs. Its primary objective is to demonstrate that targeted innovations across training and evaluation enable diffusion ODEs to achieve state-of-the-art likelihood performance without requiring additional data augmentation or costly learned dequantization networks.

The authors evaluated their framework through rigorous empirical testing on standard image density estimation benchmarks, specifically CIFAR-10 and ImageNet-32, across two common noise schedules: Variance Preserving and Straight Path. The methodological approach divides training into pretraining and fine-tuning stages. Pretraining incorporates normalized velocity parameterization—which predicts the drift along the diffusion path rather than the standard noise—and applies analytical importance sampling for variance reduction. Fine-tuning uses a novel second-order flow matching objective to smooth ODE trajectories. For evaluation, the framework introduces a training-free truncated-normal dequantization method coupled with an importance-weighted estimator to convert discrete data to continuous representations smoothly.

The experiments produced four major findings. First, the framework achieved state-of-the-art likelihood results on standard benchmarks, attaining 2.56 bits per dimension on CIFAR-10 (improving upon previous ODE bests of 2.90 to 2.99) and 3.43 on ImageNet-32. Second, the training strategy accelerated pretraining convergence by approximately two to three times compared to existing leading models like Variational Diffusion Models. Third, the truncated-normal dequantization closed the discrepancy between training and evaluation distributions, outperforming standard uniform dequantization by about 0.14 bits per dimension. Fourth, the second-order fine-tuning objective successfully regularized model divergence, cutting the required number of function evaluations during sampling from 248 down to 126, which substantially improves inference efficiency.

These results establish that diffusion ODEs can act as highly competitive, exact likelihood estimators while eliminating the computational complexity and training overhead of variational dequantization. By accelerating convergence and halving the required sampling steps, the framework lowers computational costs and shortens deployment timelines. Practitioners should note that optimizing specifically for likelihood slightly degrades sample visual diversity (FID scores), highlighting a standard trade-off between exact density fitting and peak generative image quality.

Organizations deploying diffusion models for density estimation, data compression, or anomaly detection should adopt velocity parameterization, analytical importance sampling, and truncated-normal dequantization. Where computational budgets permit, teams should implement the second-order fine-tuning stage to halve function evaluations during deterministic ODE sampling. For pure image synthesis applications, teams should consider integrating predictor-corrector samplers or specialized noise schedules to balance visual quality.

Confidence in the reported density estimation improvements is high across the evaluated image benchmarks. However, the study was constrained by computational limits to small image resolutions (32x32) and fixed hyperparameter configurations. Readers should exercise caution when extrapolating these likelihood gains to larger image resolutions, varied data domains like text or audio, or scenarios prioritizing raw visual realism over density accuracy before conducting domain-specific pilot studies.

  • Paper: Neural Diffusion Models, Grigory Bartosh et al. (2024). It advances diffusion likelihood estimation by learning flexible forward transformations, extending the source’s effort to improve density estimation within diffusion models.
Cover for Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs

Abstract

Diffusion models have exhibited excellent performance in various domains. The probability flow ordinary differential equation (ODE) of diffusion models (i.e., diffusion ODEs) is a particular case of continuous normalizing flows (CNFs), which enables deterministic inference and exact likelihood evaluation. However, the likelihood estimation results by diffusion ODEs are still far from those of the state-of-the-art likelihood-based generative models. In this work, we propose several improved techniques for maximum likelihood estimation for diffusion ODEs, including both training and evaluation perspectives. For training, we propose velocity parameterization and explore variance reduction techniques for faster convergence. We also derive an error-bounded high-order flow matching objective for finetuning, which improves the ODE likelihood and smooths its trajectory. For evaluation, we propose a novel training-free truncated-normal dequantization to fill the training-evaluation gap commonly existing in diffusion ODEs. Building upon these techniques, we achieve state-of-the-art likelihood estimation results on image datasets (2.56 on CIFAR-10, 3.43/3.69 on ImageNet-32) without variational dequantization or data augmentation.

Table of Contents

  • 1. Introduction
  • 2. Diffusion Models
  • 2.1. Diffusion ODEs and Maximum Likelihood Training
  • 2.2. Log-SNR Timed Diffusion Models
  • 2.3. Dequantization for Density Estimation
  • 3. Diffusion ODEs with Truncated-Normal Dequantization
  • 3.1. Challenges for Diffusion ODEs with Dequantization
  • 3.2. Training-Free Dequantization by Truncated Normal
  • 4. Practical Techniques for Improving the Likelihood of Diffusion ODEs
  • 4.1. Velocity Parameterization
  • 4.2. Error-bounded Second-Order Flow Matching
  • 4.3. Timing by Log-SNR and Normalizing Velocity
  • 4.4. Variance Reduction with Importance Sampling
  • 5. Related Work
  • 6. Experiments
  • 6.1. Likelihood and Samples
  • 6.2. Ablations
  • 7. Conclusion
  • Acknowledgements
  • References
  • A. Different perspective of diffusion ODEs for bridging the gap between discrete and continuous data
  • A.1. Dequantization perspective
  • A.2. Variational perspective
  • A.3. Practical connections and results
  • B. Equivalence of different predictors and matching objectives
  • C. Specifications under VP and SP schedule
  • D. Illustration of velocity prediction and imbalance problem
  • E. Relationship between velocity parameterization and other works
  • E.1. Interpretation by preconditioning
  • E.2. Connection to flow matching in Lipman et al. (2022)
  • E.3. Connection to v prediction
  • F. Error-bounded trace of second-order flow matching
  • G. Difference between our second-order flow matching and the previous time score matching in Choi et al. (2022)
  • H. Details of our adaptive IS
  • I. Experiment details
  • J. Additional samples

Knowls

  1. Knowl 1 — Truncated-normal dequantization matches diffusion ODE training inputs

    model/method

    For 8-bit image data scaled to [−1,1][-1,1], each discrete value x0x_0 represents a continuous cell x0+ux_0+u, with each coordinate of uu in [−1/256,1/256)[-1/256,1/256). A diffusion ODE is trained from inputs xε=αεx0+σεϵx_\varepsilon=\alpha_\varepsilon x_0+\sigma_\varepsilon\epsilon, where ϵ∼N(0,Id)\epsilon\sim\mathcal N(0,I_d) and ε>0\varepsilon>0 is the starting time. Uniform dequantization instead evaluates the model on uniformly perturbed data, producing a mismatch with those approximately Gaussian training inputs.

    The proposed training-free dequantization draws the perturbation from a truncated normal:

    q(u∣x0)=TN⁡ ⁣(0,σε2αε2Id,−1256,1256).q(u\mid x_0)=\operatorname{TN}\!\left(0,\frac{\sigma_\varepsilon^2}{\alpha_\varepsilon^2}I_d,-\frac{1}{256},\frac{1}{256}\right).

    Here dd is the data dimension, IdI_d is the dd-dimensional identity matrix, and the truncation bounds apply coordinatewise. Equivalently, draw ϵ^∼TN⁡(0,Id,−τ,τ)\hat\epsilon\sim\operatorname{TN}(0,I_d,-\tau,\tau) with τ=αε/(256σε)\tau=\alpha_\varepsilon/(256\sigma_\varepsilon) and evaluate at x^ε=αεx0+σεϵ^\hat x_\varepsilon=\alpha_\varepsilon x_0+\sigma_\varepsilon\hat\epsilon. Choosing the starting time so that the standardized truncation range is about three standard deviations gives γε=log⁡(σε2/αε2)≈−13.3\gamma_\varepsilon=\log(\sigma_\varepsilon^2/\alpha_\varepsilon^2)\approx-13.3 and τ≈3\tau\approx3; the authors report that this makes training and evaluation inputs virtually identically distributed without training an auxiliary dequantization model.

  2. Knowl 2 — Analytic likelihood bounds for truncated-normal dequantization

    theoretical result

    Let P0(x0)P_0(x_0) be the discrete probability assigned to an 8-bit datum x0x_0 by integrating a continuous diffusion-ODE density pεp_\varepsilon over its quantization cell. Under the truncated-normal dequantization with x^ε=αεx0+σεϵ^\hat x_\varepsilon=\alpha_\varepsilon x_0+\sigma_\varepsilon\hat\epsilon, ϵ^∼TN⁡(0,Id,−τ,τ)\hat\epsilon\sim\operatorname{TN}(0,I_d,-\tau,\tau), and τ=αε/(256σε)\tau=\alpha_\varepsilon/(256\sigma_\varepsilon), the paper gives the following variational lower bound:

    log⁡P0(x0)≥Eϵ^ ⁣[log⁡pε(x^ε)]+d2(1+log⁡(2πσε2))+dlog⁡Z−dτ2πZexp⁡ ⁣(−τ22),\log P_0(x_0)\geq \mathbb E_{\hat\epsilon}\!\left[\log p_\varepsilon(\hat x_\varepsilon)\right]+\frac d2\left(1+\log(2\pi\sigma_\varepsilon^2)\right)+d\log Z-\frac{d\tau}{\sqrt{2\pi}Z}\exp\!\left(-\frac{\tau^2}{2}\right),

    where dd is the data dimension, Z=erf⁡(τ/2)Z=\operatorname{erf}(\tau/\sqrt2), and erf⁡\operatorname{erf} is the error function. The density pεp_\varepsilon is the continuous model density at starting time ε\varepsilon; its likelihood is evaluated by the diffusion ODE.

    The paper also provides an importance-weighted lower bound from KK independent draws ϵ^(i)\hat\epsilon^{(i)} from the same truncated standard normal:

    log⁡P0(x0)≥E ⁣[log⁡(1K∑i=1Kpε(x^ε(i))q(ϵ^(i)))]+dlog⁡σε,\log P_0(x_0)\geq \mathbb E\!\left[\log\left(\frac1K\sum_{i=1}^K\frac{p_\varepsilon(\hat x_\varepsilon^{(i)})}{q(\hat\epsilon^{(i)})}\right)\right]+d\log\sigma_\varepsilon,

    where x^ε(i)=αεx0+σεϵ^(i)\hat x_\varepsilon^{(i)}=\alpha_\varepsilon x_0+\sigma_\varepsilon\hat\epsilon^{(i)} and qq is the density of TN⁡(0,Id,−τ,τ)\operatorname{TN}(0,I_d,-\tau,\tau):

    q(ϵ^)=1(2πZ2)d/2exp⁡ ⁣(−∥ϵ^∥222).q(\hat\epsilon)=\frac{1}{(2\pi Z^2)^{d/2}}\exp\!\left(-\frac{\|\hat\epsilon\|_2^2}{2}\right).

    Increasing KK tightens this Jensen-based bound; the experiments use K=20K=20 for the principal likelihood results.

  3. Knowl 3 — Velocity parameterization turns score matching into flow matching

    model/method

    For a diffusion with scalar noise schedules αt\alpha_t and σt\sigma_t, sample data x0x_0 and noise ϵ∼N(0,Id)\epsilon\sim\mathcal N(0,I_d), and set xt=αtx0+σtϵx_t=\alpha_t x_0+\sigma_t\epsilon. The path velocity is v=α˙tx0+σ˙tϵv=\dot\alpha_t x_0+\dot\sigma_t\epsilon, where dots denote derivatives with respect to time tt. Instead of directly predicting the score, the method parameterizes the diffusion-ODE drift as

    vθ(xt,t)=f(t)xt−12g(t)2sθ(xt,t),v_\theta(x_t,t)=f(t)x_t-\frac12g(t)^2s_\theta(x_t,t),

    where sθs_\theta is a score predictor and f,gf,g are the forward SDE drift and diffusion schedules. The resulting likelihood-weighted first-order flow-matching objective is

    JFM(θ)=∫0T2g(t)2 Ex0,ϵ ⁣[∥vθ(xt,t)−v∥22]dt.J_{\rm FM}(\theta)=\int_0^T\frac{2}{g(t)^2}\,\mathbb E_{x_0,\epsilon}\!\left[\|v_\theta(x_t,t)-v\|_2^2\right]dt.

    With an unrestricted predictor, the target is the probability-flow drift v∗(xt,t)=f(t)xt−12g(t)2∇xlog⁡qt(xt)v^*(x_t,t)=f(t)x_t-\tfrac12g(t)^2\nabla_x\log q_t(x_t), where qtq_t is the forward-process marginal. Thus the method trains a drift predictor by matching the velocity of sampled diffusion paths, while retaining the likelihood weighting associated with maximum-likelihood training.

  4. Knowl 4 — Log-SNR timing and normalized velocity define the practical ODE

    model/method

    The practical parameterization uses negative log-SNR γ=log⁡(σγ2/αγ2)\gamma=\log(\sigma_\gamma^2/\alpha_\gamma^2) as time and predicts a normalized path velocity. For xγ=αγx0+σγϵx_\gamma=\alpha_\gamma x_0+\sigma_\gamma\epsilon, define

    v~=α˙γx0+σ˙γϵα˙γ2+σ˙γ2,\tilde v=\frac{\dot\alpha_\gamma x_0+\dot\sigma_\gamma\epsilon}{\sqrt{\dot\alpha_\gamma^2+\dot\sigma_\gamma^2}},

    where dots now denote derivatives with respect to γ\gamma. The network v~θ(xγ,γ)\tilde v_\theta(x_\gamma,\gamma) is trained with

    JFM(θ)=∫γ0γT2(α˙γ2+σ˙γ2)σγ2 Ex0,ϵ ⁣[∥v~θ(xγ,γ)−v~∥22]dγ.J_{\rm FM}(\theta)=\int_{\gamma_0}^{\gamma_T}\frac{2(\dot\alpha_\gamma^2+\dot\sigma_\gamma^2)}{\sigma_\gamma^2}\,\mathbb E_{x_0,\epsilon}\!\left[\|\tilde v_\theta(x_\gamma,\gamma)-\tilde v\|_2^2\right]d\gamma.

    The corresponding learned probability-flow ODE is

    dxγdγ=α˙γ2+σ˙γ2 v~θ(xγ,γ).\frac{dx_\gamma}{d\gamma}=\sqrt{\dot\alpha_\gamma^2+\dot\sigma_\gamma^2}\,\tilde v_\theta(x_\gamma,\gamma).

    Normalization controls changes in target scale across time, while log-SNR timing separates the choice of noise schedule from the time coordinate and affects the variance of Monte Carlo training estimates.

  5. Knowl 5 — Error-bounded second-order flow matching regularizes the likelihood trace

    theoretical result

    For a diffusion path xt=αtx0+σtϵx_t=\alpha_t x_0+\sigma_t\epsilon, let v=α˙tx0+σ˙tϵv=\dot\alpha_t x_0+\dot\sigma_t\epsilon be its velocity and let v∗(xt,t)v^*(x_t,t) be the optimal first-order velocity predictor. The diffusion-ODE log-density change depends on tr⁡(∇xvθ)\operatorname{tr}(\nabla_x v_\theta), which ordinary first-order matching does not directly constrain. Given a first-order estimator v^1\hat v_1, the paper trains a trace predictor v2trace(xt,t;θ)v_2^{\rm trace}(x_t,t;\theta) toward the target

    σ˙tσtd+ℓ1(ϵ,x0,t),ℓ1(ϵ,x0,t)=2g(t)2∥v^1(xt,t)−v∥22,\frac{\dot\sigma_t}{\sigma_t}d+\ell_1(\epsilon,x_0,t),\qquad \ell_1(\epsilon,x_0,t)=\frac{2}{g(t)^2}\|\hat v_1(x_t,t)-v\|_2^2,

    by minimizing the expected squared difference from that target over x0x_0 and ϵ\epsilon. Here dd is the data dimension and g(t)g(t) is the forward SDE diffusion schedule. In practice, the trace predictor is tied to the learned field as v2trace=tr⁡(∇xvθ)v_2^{\rm trace}=\operatorname{tr}(\nabla_xv_\theta), with stop-gradient treatment of the first-order estimator in the practical objective.

    If θ∗\theta^* minimizes the trace-prediction objective and δ1(xt,t)=∥v^1(xt,t)−v∗(xt,t)∥2\delta_1(x_t,t)=\|\hat v_1(x_t,t)-v^*(x_t,t)\|_2, the paper's error bound is

    ∣v2trace(xt,t;θ)−tr⁡(∇xv∗(xt,t))∣≤∣v2trace(xt,t;θ)−v2trace(xt,t;θ∗)∣+2g(t)2δ1(xt,t)2.\left|v_2^{\rm trace}(x_t,t;\theta)-\operatorname{tr}(\nabla_xv^*(x_t,t))\right|\leq \left|v_2^{\rm trace}(x_t,t;\theta)-v_2^{\rm trace}(x_t,t;\theta^*)\right|+\frac{2}{g(t)^2}\delta_1(x_t,t)^2.

    The result bounds the trace-estimation error by the trace model's own fitting error plus a term quadratic in the first-order velocity error. The authors use a mixture of first- and second-order objectives for finetuning; they report improved likelihood and smoother ODE trajectories.

  6. Knowl 6 — Designed importance sampling reduces variance without a learned proposal

    model/method

    The normalized-velocity objective is an integral over log-SNR γ\gamma. Write its per-time loss as

    Lθ(x0,ϵ,γ)=2(α˙γ2+σ˙γ2)σγ2∥v~θ(xγ,γ)−v~∥22.L_\theta(x_0,\epsilon,\gamma)=\frac{2(\dot\alpha_\gamma^2+\dot\sigma_\gamma^2)}{\sigma_\gamma^2}\|\tilde v_\theta(x_\gamma,\gamma)-\tilde v\|_2^2.

    For any sampling density p(γ)p(\gamma) on [γ0,γT][\gamma_0,\gamma_T], the objective can be estimated unbiasedly as Eγ∼p[Lθ/p(γ)]\mathbb E_{\gamma\sim p}[L_\theta/p(\gamma)]. The paper's training-free designed proposal is

    p(γ)∝α˙γ2+σ˙γ2σγ2.p(\gamma)\propto\frac{\dot\alpha_\gamma^2+\dot\sigma_\gamma^2}{\sigma_\gamma^2}.

    This choice cancels the time-varying coefficient on the velocity-matching error in the importance-weighted estimator. Under both tested schedules—variance-preserving (VP), αγ2+σγ2=1\alpha_\gamma^2+\sigma_\gamma^2=1, and straight-path (SP), αγ+σγ=1\alpha_\gamma+\sigma_\gamma=1—the proposal is proportional to αγ2\alpha_\gamma^2. The authors sample it by inverse-transform sampling: the VP mapping has a closed form, whereas the SP mapping is obtained by bisection. Their experiments find that designed sampling accelerates training relative to uniform-time sampling and gives convergence similar to a learned proposal, without the learned proposal's extra network and optimization overhead.

  7. Knowl 7 — Training and evaluation configuration for i-DODE

    experimental setup

    The experiments train diffusion ODEs on CIFAR-10 and ImageNet-32 using VP and SP noise schedules. Models predict normalized velocity as a function of log-SNR and use a U-Net based on the VDM architecture: depth 32, 128 channels for CIFAR-10 and 256 for ImageNet-32, with dropout 0.1. The authors use Adam with learning rate 2×10−42\times10^{-4}, β1=0.9\beta_1=0.9, β2=0.99\beta_2=0.99, decoupled weight decay 0.01, and an evaluation exponential moving average rate of 0.9999.

    Training first minimizes the first-order flow-matching objective, then finetunes with the mixture JFM+λJFM,trJ_{\rm FM}+\lambda J_{\rm FM,tr} using λ=0.1\lambda=0.1. The standard time endpoints are γε=−13.3\gamma_\varepsilon=-13.3 and γT=5.0\gamma_T=5.0. CIFAR-10 uses batch size 128, 6 million pretraining iterations, and 200,000 finetuning iterations; the newer ImageNet-32 version uses batch size 128, 2 million pretraining iterations, and 250,000 finetuning iterations. The older ImageNet-32 version uses batch size 512 for 2 million pretraining iterations and batch size 128 for 500,000 finetuning iterations, with gradients accumulated over four batches.

    Likelihood is evaluated using truncated-normal dequantization; the principal reported results use K=20K=20 importance samples. For sampling and FID evaluation, the authors use an adaptive-step RK45 ODE solver with relative and absolute tolerances both 10−510^{-5} and generate 50,000 samples. The experiments do not use data augmentation or trained variational dequantization.

  8. Knowl 8 — i-DODE improves benchmark likelihood over prior diffusion ODEs

    data/table

    The comparison reports negative log-likelihood in bits per dimension (NLL, lower is better), sample quality as FID (lower is better), and sampling cost as number of function evaluations (NFE, lower is better). i-DODE likelihood uses truncated-normal dequantization and the importance-weighted estimator with K=20K=20; the reported models use no data augmentation or trained variational dequantization. A starred ImageNet-32 value refers to the older dataset version, and VDM's FID is from 1,000-step SDE sampling rather than ODE sampling.

    The table shows that i-DODE achieves the best listed ODE NLL on CIFAR-10 and ImageNet-32. Its likelihood improves over the prior ODE results while its FID is not the best listed, reflecting a likelihood/sample-quality tradeoff.

    Model CIFAR-10 NLL CIFAR-10 FID CIFAR-10 NFE ImageNet-32 NLL ImageNet-32 FID ImageNet-32 NFE
    VDM 2.65 7.60†^{\dagger} 1000†^{\dagger} 3.72∗^{*} / /
    ScoreFlow 2.90 5.40 / 3.82∗^{*} 10.18∗^{*} /
    Flow Matching 2.99 6.35 142 3.53 5.31 122
    Stochastic Interpolant 2.99 10.27 / 3.48 8.49 /
    i-DODE (SP) 2.56 11.20 162 3.44/3.69∗^{*} 10.31 138
    i-DODE (VP) 2.57 10.74 126 3.43/3.70∗^{*} 9.09 152

    The two ImageNet-32 NLL values are reported as new-version/old-version results, with the older-version value marked by an asterisk in the paper. The comparison establishes the likelihood gain; it does not establish superior sample quality, since i-DODE's FID is worse than several comparison models.

  9. Knowl 9 — Ablations isolate the gains from dequantization and finetuning

    empirical result

    On CIFAR-10 under the VP schedule, the authors compare VDM, their pretrained flow-matching model, and the same model after second-order finetuning. NLL is in bits per dimension; “U” denotes uniform dequantization and “TN” denotes truncated-normal dequantization. Finetuning improves TN NLL from 2.61 to 2.60 and reduces sampling NFE from 248 to 126, while FID changes from 10.66 to 10.74. The TN evaluation is substantially better than uniform evaluation for all rows, consistent with the proposed training-evaluation match.

    Model NLL (U) NLL (TN) FID NFE
    VDM 2.78 2.64 8.65 213
    Pretrain (ours) 2.75 2.61 10.66 248
    Pretrain + finetune (ours) 2.74 2.60 10.74 126

    The CIFAR-10 pretraining curve also shows faster convergence for velocity prediction with designed importance sampling than for the VDM noise-prediction baseline; the authors describe the speedup as approximately 2–3 times. In the ablation from scratch, both velocity prediction and importance sampling accelerate loss descent.

  10. Knowl 10 — Reported limitations and likelihood–sample-quality tradeoff

    limitation

    The paper reports that the likelihood-oriented architecture and training choices yield worse FID than state-of-the-art sample-generation systems, despite strong likelihood results. Its CIFAR-10 i-DODE FIDs are 11.20 (SP) and 10.74 (VP), and its ImageNet-32 FIDs are 10.31 (SP) and 9.09 (VP). The authors suggest that emphasizing training at low log-SNR or using higher-quality sampling procedures could improve sample quality, but do not test those remedies in this work. They also state that limited resources prevented broader tuning of network architectures and hyperparameters.

Coverage note — Appendix proofs, the alternate variational-autoencoder interpretation of dequantization, predictor-equivalence details beyond the velocity formulation, and sample galleries are omitted because they support or illustrate the extracted methods and results rather than adding a separate central contribution.

References

  1. 1.Albergo, M. S. and Vanden-Eijnden, E. Building normalizing flows with stochastic interpolants. arXiv preprint arXiv:2209.15571, 2022.
  2. 2.Anderson, B. D. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982.
  3. 3.Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al. Jax: composable transformations of python+ numpy programs. Version 0.2, 5:14–24, 2018.
  4. 4.Burda, Y., Grosse, R., and Salakhutdinov, R. Importance weighted autoencoders. arXiv preprint arXiv:1509.00519, 2015.
  5. 5.Chen, N., Zhang, Y., Zen, H., Weiss, R. J., Norouzi, M., and Chan, W. Wavegrad: Estimating gradients for waveform generation. In International Conference on Learning Representations, 2021.
  6. 6.Chen, R. T., Rubanova, Y., Bettencourt, J., and Duvenaud, D. Neural ordinary differential equations. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 6572–6583, 2018a.
  7. 7.Chen, X., Mishra, N., Rohaninejad, M., and Abbeel, P. Pixelsnail: An improved autoregressive generative model. In International Conference on Machine Learning, pp. 864–872. PMLR, 2018b.
  8. 8.Chen, Z., Yeo, C. K., Lee, B. S., and Lau, C. T. Autoencoder-based network anomaly detection. In 2018 Wireless telecommunications symposium (WTS), pp. 1–5. IEEE, 2018c.
  9. 9.Choi, K., Meng, C., Song, Y., and Ermon, S. Density ratio estimation via infinitesimal classification. In International Conference on Artificial Intelligence and Statistics, pp. 2552–2573. PMLR, 2022.
  10. 10.Chung, H., Kim, J., Mccann, M. T., Klasky, M. L., and Ye, J. C. Diffusion posterior sampling for general noisy inverse problems. arXiv preprint arXiv:2209.14687, 2022.
  11. 11.Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255. IEEE, 2009.
  12. 12.Dhariwal, P. and Nichol, A. Q. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, volume 34, pp. 8780–8794, 2021.
  13. 13.Dias, M. L., Mattos, C. L. C., da Silva, T. L., de Macedo, J. A. F., and Silva, W. C. Anomaly detection in trajectory data with normalizing flows. In 2020 International Joint Conference on Neural Networks (IJCNN), pp. 1–8. IEEE, 2020.
  14. 14.Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using real nvp. In International Conference on Learning Representations, 2017.
  15. 15.Dormand, J. R. and Prince, P. J. A family of embedded Runge-Kutta formulae. Journal of computational and applied mathematics, 6(1):19–26, 1980.
  16. 16.Finlay, C., Jacobsen, J.-H., Nurbekyan, L., and Oberman, A. How to train your neural ode: the world of jacobian and kinetic regularization. In International conference on machine learning, pp. 3154–3164. PMLR, 2020.
  17. 17.Grathwohl, W., Chen, R. T., Bettencourt, J., Sutskever, I., and Duvenaud, D. Ffjord: Free-form continuous dynamics for scalable reversible generative models. In International Conference on Learning Representations, 2019.
  18. 18.Helminger, L., Djelouah, A., Gross, M., and Schroers, C. Lossy image compression with normalizing flows. arXiv preprint arXiv:2008.10486, 2020.
  19. 19.Ho, J., Chen, X., Srinivas, A., Duan, Y., and Abbeel, P. Flow++: Improving flow-based generative models with variational dequantization and architecture design. In International Conference on Machine Learning, pp. 2722–2730. PMLR, 2019.
  20. 20.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33, pp. 6840–6851, 2020.
  21. 21.Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al. Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303, 2022.
  22. 22.Ho, Y.-H., Chan, C.-C., Peng, W.-H., Hang, H.-M., and Domanski, M. Anfic: Image compression using augmented normalizing flows. IEEE Open Journal of Circuits and Systems, 2:613–626, 2021.
  23. 23.Huang, C.-W., Lim, J. H., and Courville, A. A variational perspective on diffusion-based generative models and score matching. In Advances in Neural Information Processing Systems, 2021.
  24. 24.Hutchinson, M. F. A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines. Communications in Statistics-Simulation and Computation, 19(2):433–450, 1990.
  25. 25.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. In Advances in Neural Information Processing Systems, 2022.
  26. 26.Kawar, B., Elad, M., Ermon, S., and Song, J. Denoising diffusion restoration models. In Advances in Neural Information Processing Systems, 2022.
  27. 27.Kim, D., Shin, S., Song, K., Kang, W., and Moon, I.-C. Soft truncation: A universal training technique of score-based diffusion model for high precision score estimation. In International Conference on Machine Learning, pp. 11201–11228. PMLR, 2022.
  28. 28.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  29. 29.Kingma, D. P. and Dhariwal, P. Glow: generative flow with invertible 1× 1 convolutions. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, pp. 10236–10245, 2018.
  30. 30.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. In International Conference on Learning Representations, 2014.
  31. 31.Kingma, D. P., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. In Advances in Neural Information Processing Systems, 2021.
  32. 32.Krizhevsky, A., Hinton, G., et al. Learning multiple layers of features from tiny images. 2009.
  33. 33.Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2022.
  34. 34.Liu, J., Li, C., Ren, Y., Chen, F., and Zhao, Z. Diffsinger: Singing voice synthesis via shallow diffusion mechanism. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp. 11020–11028, 2022a.
  35. 35.Liu, X., Gong, C., and Liu, Q. Flow straight and fast: Learning to generate and transfer data with rectified flow. arXiv preprint arXiv:2209.03003, 2022b.
  36. 36.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019.
  37. 37.Lu, C., Zheng, K., Bao, F., Chen, J., Li, C., and Zhu, J. Maximum likelihood training for score-based diffusion odes by high order denoising score matching. In International Conference on Machine Learning, pp. 14429–14460. PMLR, 2022a.
  38. 38.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. In Advances in Neural Information Processing Systems, 2022b.
  39. 39.Meng, C., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S. SDEdit: Image synthesis and editing with stochastic differential equations. In International Conference on Learning Representations, 2022.
  40. 40.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pp. 8162–8171. PMLR, 2021.
  41. 41.Nichol, A. Q., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., Mcgrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, pp. 16784–16804. PMLR, 2022.
  42. 42.Oord, A. v. d., Kalchbrenner, N., Vinyals, O., Espeholt, L., Graves, A., and Kavukcuoglu, K. Conditional image generation with pixelcnn decoders. In Proceedings of the 30th International Conference on Neural Information Processing Systems, pp. 4797–4805, 2016.
  43. 43.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with CLIP latents. arXiv preprint arXiv:2204.06125, 2022.
  44. 44.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  45. 45.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, 2022.
  46. 46.Salimans, T., Karpathy, A., Chen, X., and Kingma, D. P. Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modifications. In International Conference on Learning Representations, 2017.
  47. 47.Serra, J., ` Alvarez, D., G ´ omez, V., Slizovskaia, O., N ´ u´nez, ˜ J. F., and Luque, J. Input complexity and out-of-distribution detection with likelihood-based generative models. In International Conference on Learning Representations, 2020.
  48. 48.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265. PMLR, 2015.
  49. 49.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021a.
  50. 50.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems, volume 32, pp. 11895–11907, 2019.
  51. 51.Song, Y., Garg, S., Shi, J., and Ermon, S. Sliced score matching: A scalable approach to density and score estimation. In Uncertainty in Artificial Intelligence, pp. 574–584. PMLR, 2020.
  52. 52.Song, Y., Durkan, C., Murray, I., and Ermon, S. Maximum likelihood training of score-based diffusion models. In Advances in Neural Information Processing Systems, volume 34, pp. 1415–1428, 2021b.
  53. 53.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021c.
  54. 54.Uria, B., Murray, I., and Larochelle, H. RNADE: The real-valued neural autoregressive density-estimator. Advances in Neural Information Processing Systems, 26, 2013.
  55. 55.Vahdat, A. and Kautz, J. Nvae: a deep hierarchical variational autoencoder. In Proceedings of the 34th International Conference on Neural Information Processing Systems, pp. 19667–19679, 2020.
  56. 56.Vincent, P. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661–1674, 2011.
  57. 57.Xiao, Z., Yan, Q., and Amit, Y. Likelihood regret: an out-of-distribution detection score for variational autoencoder. In Proceedings of the 34th International Conference on Neural Information Processing Systems, pp. 20685–20696, 2020.
  58. 58.Xu, Y., Liu, Z., Tegmark, M., and Jaakkola, T. S. Poisson flow generative models. In Advances in Neural Information Processing Systems, 2022.
  59. 59.Xu, Y., Liu, Z., Tian, Y., Tong, S., Tegmark, M., and Jaakkola, T. Pfgm++: Unlocking the potential of physics-inspired generative models. arXiv preprint arXiv:2302.04265, 2023.
  60. 60.Yang, R. and Mandt, S. Lossy image compression with conditional diffusion models. arXiv preprint arXiv:2209.06950, 2022.
  61. 61.Zhao, M., Bao, F., Li, C., and Zhu, J. Egsde: Unpaired image-to-image translation via energy-guided stochastic differential equations. In Advances in Neural Information Processing Systems, 2022.

Citation

MLA
Zheng, K., et al. “Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs”. International Conference on Machine Learning, vol. 202, 2023, pp. 42363–89, https://proceedings.mlr.press/v202/zheng23c.html.
APA
Zheng, K., Lu, C., Chen, J., & Zhu, J. (2023). Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs. International Conference on Machine Learning, 202, 42363–42389. https://proceedings.mlr.press/v202/zheng23c.html
Chicago
Zheng, K., C. Lu, J. Chen, and J. Zhu. 2023. “Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs”. International Conference on Machine Learning 202: 42363–89. https://proceedings.mlr.press/v202/zheng23c.html.
Harvard
Zheng, K. et al. (2023) “Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs”, International Conference on Machine Learning. PMLR, pp. 42363–42389. Available at: https://proceedings.mlr.press/v202/zheng23c.html.
Vancouver
1. Zheng K, Lu C, Chen J, Zhu J (2023) Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs. In: International Conference on Machine Learning. PMLR, pp 42363–42389

BibTeX

@InProceedings{pmlr-v202-zheng23c,
  title = 	 {Improved Techniques for Maximum Likelihood Estimation for Diffusion {ODE}s},
  author =       {Zheng, Kaiwen and Lu, Cheng and Chen, Jianfei and Zhu, Jun},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {42363--42389},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/zheng23c/zheng23c.pdf},
  url = 	 {https://proceedings.mlr.press/v202/zheng23c.html},
  abstract = 	 {Diffusion models have exhibited excellent performance in various domains. The probability flow ordinary differential equation (ODE) of diffusion models (i.e., diffusion ODEs) is a particular case of continuous normalizing flows (CNFs), which enables deterministic inference and exact likelihood evaluation. However, the likelihood estimation results by diffusion ODEs are still far from those of the state-of-the-art likelihood-based generative models. In this work, we propose several improved techniques for maximum likelihood estimation for diffusion ODEs, including both training and evaluation perspectives. For training, we propose velocity parameterization and explore variance reduction techniques for faster convergence. We also derive an error-bounded high-order flow matching objective for finetuning, which improves the ODE likelihood and smooths its trajectory. For evaluation, we propose a novel training-free truncated-normal dequantization to fill the training-evaluation gap commonly existing in diffusion ODEs. Building upon these techniques, we achieve state-of-the-art likelihood estimation results on image datasets (2.56 on CIFAR-10, 3.43/3.69 on ImageNet-32) without variational dequantization or data augmentation.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/