Maximum Likelihood Training for Score-based Diffusion ODEs by High Order Denoising Score Matching
Cheng LuKaiwen ZhengFan BaoJianfei ChenChongxuan LiJun Zhu
Proves that first-order score matching fails to maximize the likelihood of score-based diffusion ODEs and proposes a high-order denoising score matching method that bounds the ODE likelihood error using higher-order score terms to achieve superior likelihood evaluation without sacrificing sample quality.
Score-based generative models have emerged as leading tools for creating high-fidelity images, audio, and complex synthetic data. These models operate through two primary mathematical representations: stochastic differential equations, which excel at generating realistic samples, and ordinary differential equations (diffusion ODEs), which enable exact evaluation of data likelihood for density estimation and model comparison. However, standard training procedures minimize only first-order score matching errors (matching the gradient of the log data distribution). While this maximizes the likelihood of the stochastic model, it does not guarantee high likelihood for the corresponding diffusion ODE, often resulting in severe density misestimations even on simple synthetic datasets.
The main objective of the article is to establish the theoretical relationship between score matching objectives and diffusion ODE likelihood, and to develop an error-bounded training method that directly maximizes ODE likelihood without sacrificing sample quality or computational scalability.
To address this objective, the authors conduct mathematical proofs to identify the gap between the Kullback-Leibler divergence (a standard measure of difference between probability distributions) and first-order score matching. They prove that ODE likelihood can be bounded by simultaneously controlling first-, second-, and third-order score matching errors (the gradient, Hessian, and gradient of the Hessian trace). Rather than using slow, step-by-step differential equation solvers during training, the authors introduce a scalable high-order denoising score matching algorithm that uses Monte Carlo random sampling and efficient trace estimators. They evaluate their method using synthetic 1-D and 2-D data distributions alongside benchmark image datasets (CIFAR-10 and ImageNet 32x32) using Variance Exploding diffusion models.
The key findings are structured as follows:
- Theoretical Gap: Minimizing standard first-order score matching leaves an uncontrolled error term in the diffusion ODE's likelihood, explaining why previous models produced poor density estimates.
- Error-Bounded Framework: The proposed high-order algorithm mathematically guarantees that higher-order score matching errors remain strictly bounded by lower-order errors and empirical training error.
- Improved Likelihood: On CIFAR-10 image benchmarks, the method improves the negative log-likelihood of deep Variance Exploding models from 3.45 to 3.27 bits per dimension, and on ImageNet 32x32 from 4.21 to 4.03 bits per dimension.
- Preserved Generation Quality: The model retains state-of-the-art sample quality (achieving competitive image quality scores of 2.61 vs. 2.19 on CIFAR-10) because the optimal model state remains aligned with the true data score function.
- Practical Computational Overhead: The training method requires no expensive differential equation solvers; memory consumption and runtime remain practical, with third-order training running in less than double the time of standard first-order training.
These findings resolve a longstanding discrepancy between high sample quality and poor likelihood evaluation in diffusion models. By establishing bounded high-order score matching, the article provides a path to train generative models that are both top-tier data generators and reliable density estimators. This dual capability reduces technical risk in applications requiring rigorous probabilistic guarantees, such as anomaly detection, data compression, and uncertainty quantification.
For practical implementation, teams deploying diffusion models should consider fine-tuning existing pretrained model checkpoints with second- or third-order objectives rather than training from scratch, which converges in approximately 100,000 iterations (roughly half a day of compute). Practitioners should also adopt standard trace estimators and variance-reduction weighting to keep compute costs manageable. Future development should focus on extending these high-order principles to other diffusion architectures, including latent-space and critically-damped Langevin models.
The study's primary limitation is that near the starting time step (t close to zero), numerical instability and unbounded score errors can diminish likelihood gains on certain architecture types (such as Variance Preserving models). However, the theoretical framework is mathematically sound, and empirical results provide strong confidence in the effectiveness of high-order score matching for scalable, high-performing density modeling.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Introduces the continuous-time stochastic differential equation and probability flow ODE formulations of score-based diffusion models that this work directly analyzes and optimizes.
- Paper: Estimation of Non-Normalized Statistical Models by Score Matching, Aapo Hyvärinen (2005). Establishes foundational score matching for non-normalized densities, providing the mathematical underpinning extended here to higher-order score derivatives.
- Paper: Generative Modeling by Estimating Gradients of the Data Distribution, Yang Song et al. (2019). Pioneers noise-conditional score networks trained via denoising score matching, which this paper enhances through higher-order objectives.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Connects discrete denoising diffusion probabilistic models to continuous score matching, establishing the base training framework examined in this paper.
- Paper: Variational Diffusion Models, Diederik P. Kingma et al. (2021). Investigates likelihood estimation bounds in continuous-time diffusion models, highlighting the likelihood gap that higher-order score matching addresses.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Develops deterministic implicit sampling trajectories that correspond to continuous diffusion ODEs evaluated for exact likelihood computation.
- Paper: Improved Techniques for Training Score-Based Generative Models, Yang Song et al. (2020). Provides key technical and architectural foundations for scaling variance-exploding score-based generative models.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). Systematizes the design space of diffusion ODEs and deterministic higher-order solvers, complementing exact likelihood-driven ODE training.
- Paper: Flow Matching for Generative Modeling, Yaron Lipman et al. (2023). Extends continuous ODE-based generative paths by directly matching continuous velocity fields rather than higher-order score gradients.
- Paper: Consistency Models, Yang Song et al. (2023). Enforces consistency along probability flow ODE trajectories to enable rapid single- and few-step sampling from models trained via continuous ODE dynamics.
- Paper: Stochastic Interpolants: A Unifying Framework for Flows and Diffusions, Michael S. Albergo et al. (2025). Unifies deterministic ODE flows and stochastic diffusions through finite-time interpolants, generalizing probability flow modeling.
- Paper: Score-Based Diffusion Models in Function Space, Jae Hyun Lim 0001 et al. (2025). Extends score-based diffusion differential equation formulations from finite Euclidean spaces to infinite-dimensional function spaces.
- Paper: A Mathematical Introduction to Diffusion Models, Jianfeng Lu (2026). Provides a comprehensive mathematical treatment of continuous reverse dynamics, discretization errors, and score-matching sample bounds.
