Improved Techniques for Maximum Likelihood Estimation for Diffusion ODEs
Kaiwen ZhengCheng LuJianfei ChenJun Zhu
Proposes a suite of training and evaluation techniques—including velocity parameterization, high-order flow matching finetuning, and training-free truncated-normal dequantization—that enables diffusion ODEs to achieve state-of-the-art exact likelihood estimation without variational dequantization or data augmentation.
Generative modeling and density estimation play critical roles in key data applications such as lossless data compression, anomaly detection, and out-of-distribution identification. Within deep generative modeling, diffusion models formulated as probability flow ordinary differential equations (diffusion ODEs) enable deterministic inference and exact likelihood evaluation. However, existing diffusion ODEs have consistently trailed stochastic counterparts, such as Variational Diffusion Models, in likelihood estimation performance. This lag primarily stems from slow training convergence, numerical instabilities, and significant mismatches between continuous training models and discrete real-world evaluation data.
The article develops a comprehensive framework called Improved Diffusion ODE (i-DODE) to advance maximum likelihood estimation in diffusion ODEs. Its primary objective is to demonstrate that targeted innovations across training and evaluation enable diffusion ODEs to achieve state-of-the-art likelihood performance without requiring additional data augmentation or costly learned dequantization networks.
The authors evaluated their framework through rigorous empirical testing on standard image density estimation benchmarks, specifically CIFAR-10 and ImageNet-32, across two common noise schedules: Variance Preserving and Straight Path. The methodological approach divides training into pretraining and fine-tuning stages. Pretraining incorporates normalized velocity parameterization—which predicts the drift along the diffusion path rather than the standard noise—and applies analytical importance sampling for variance reduction. Fine-tuning uses a novel second-order flow matching objective to smooth ODE trajectories. For evaluation, the framework introduces a training-free truncated-normal dequantization method coupled with an importance-weighted estimator to convert discrete data to continuous representations smoothly.
The experiments produced four major findings. First, the framework achieved state-of-the-art likelihood results on standard benchmarks, attaining 2.56 bits per dimension on CIFAR-10 (improving upon previous ODE bests of 2.90 to 2.99) and 3.43 on ImageNet-32. Second, the training strategy accelerated pretraining convergence by approximately two to three times compared to existing leading models like Variational Diffusion Models. Third, the truncated-normal dequantization closed the discrepancy between training and evaluation distributions, outperforming standard uniform dequantization by about 0.14 bits per dimension. Fourth, the second-order fine-tuning objective successfully regularized model divergence, cutting the required number of function evaluations during sampling from 248 down to 126, which substantially improves inference efficiency.
These results establish that diffusion ODEs can act as highly competitive, exact likelihood estimators while eliminating the computational complexity and training overhead of variational dequantization. By accelerating convergence and halving the required sampling steps, the framework lowers computational costs and shortens deployment timelines. Practitioners should note that optimizing specifically for likelihood slightly degrades sample visual diversity (FID scores), highlighting a standard trade-off between exact density fitting and peak generative image quality.
Organizations deploying diffusion models for density estimation, data compression, or anomaly detection should adopt velocity parameterization, analytical importance sampling, and truncated-normal dequantization. Where computational budgets permit, teams should implement the second-order fine-tuning stage to halve function evaluations during deterministic ODE sampling. For pure image synthesis applications, teams should consider integrating predictor-corrector samplers or specialized noise schedules to balance visual quality.
Confidence in the reported density estimation improvements is high across the evaluated image benchmarks. However, the study was constrained by computational limits to small image resolutions (32x32) and fixed hyperparameter configurations. Readers should exercise caution when extrapolating these likelihood gains to larger image resolutions, varied data domains like text or audio, or scenarios prioritizing raw visual realism over density accuracy before conducting domain-specific pilot studies.
- Paper: Maximum Likelihood Training for Score-based Diffusion ODEs by High Order Denoising Score Matching, Cheng Lu et al. (2022). It develops an earlier method for directly improving diffusion-ODE likelihood, whose limitations motivate the source’s training and evaluation innovations.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Its continuous-time SDE framework derives the probability-flow ODE and likelihood computation that the source seeks to optimize.
- Paper: Variational Diffusion Models, Diederik P. Kingma et al. (2021). It establishes the variational diffusion likelihood benchmark against which the source compares its improved ODE estimates.
- Paper: Neural Diffusion Models, Grigory Bartosh et al. (2024). It advances diffusion likelihood estimation by learning flexible forward transformations, extending the source’s effort to improve density estimation within diffusion models.
