Neural Diffusion Models
Grigory BartoshDmitry P. VetrovChristian A. Naesseth
Introduces Neural Diffusion Models, a generative framework that replaces rigid linear noising processes with learnable, time-dependent non-linear data transformations to achieve tighter likelihood bounds and state-of-the-art density estimation on image benchmarks.
Generative diffusion models have demonstrated remarkable success across domains such as computer vision, audio synthesis, and molecular biology. However, conventional diffusion frameworks rely on a rigid, pre-specified noise-injection process that applies simple linear scaling to data before adding Gaussian noise. This inability to adapt the forward data-corruption path limits latent space flexibility, complicates the reverse generative process, and restricts density estimation performance.
The article aims to overcome these constraints by introducing Neural Diffusion Models, a generalized framework that incorporates learnable, time-dependent, and non-linear data transformations into the forward process. The objective is to evaluate whether parameterizing these transformations with neural networks improves likelihood estimation and generative dynamics without sacrificing the simulation-free computational efficiency of modern diffusion models.
To evaluate this framework, the authors formulated both discrete-time and continuous-time versions of Neural Diffusion Models, deriving exact time derivatives using Jacobian-vector products to maintain mathematically rigorous optimization. They evaluated the approach across 2D synthetic benchmarks and standard image datasets, including MNIST, CIFAR-10, downsampled ImageNet at 32x32 and 64x64 resolutions, and CelebA-HQ at 256x256 resolution. Models were benchmarked against standard denoising diffusion baselines under controlled conditions, utilizing identical neural network architectures, variance schedules, and computational budgets.
The experimental results reveal five primary findings. First, Neural Diffusion Models consistently achieved superior density estimation, setting new state-of-the-art diffusion likelihood marks on ImageNet 32x32 (3.55 bits per dimension), ImageNet 64x64 (3.35 bits per dimension), and CelebA-HQ. Second, the learned transformations actively simplify data geometry during training—for instance, increasing contrast in natural images or separating complex patterns—allowing the model to produce terminal data predictions that align much more closely with the underlying data distribution. Third, in few-step generation regimes (e.g., 10 discrete steps), the proposed model substantially outperformed standard diffusion in sample quality, improving the CIFAR-10 quality score from 37.83 to 31.56. Fourth, when integrated with latent-space models like LSGM, the framework reduced negative log-likelihood on CelebA-HQ while maintaining competitive visual fidelity. Fifth, ablation studies confirmed that these performance gains stem directly from the learned non-linear transformations rather than the increased parameter count.
These findings indicate that adapting the forward corruption process to the data distribution bridges the gap between true data likelihood and the variational lower bound. For practical applications that depend heavily on accurate probability modeling—such as data compression, anomaly detection, semi-supervised learning, and adversarial defense—Neural Diffusion Models offer a significant performance upgrade over standard diffusion techniques with fixed corruption dynamics.
Organizations evaluating this approach should consider adoption for tasks where density estimation and few-step generation quality are paramount. However, decision-makers must account for trade-offs: learning the transformation roughly doubles model parameters and increases training time by approximately 2.3 times compared to baseline diffusion. Additionally, the learned forward dynamics prevent direct use of standard classifier guidance, and the models exhibit greater performance degradation when post-hoc reducing sampling steps from a model trained on many steps. Future work should focus on developing efficient parameterizations, enabling classifier-free or alternative conditional sampling methods, and extending restricted reverse dynamics to high-dimensional datasets.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Read the foundational DDPM formulation first to understand the fixed Gaussian corruption and learned denoising process that Neural Diffusion Models generalize.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). Its continuous-time SDE framework supplies the score-based diffusion and reverse-process foundations used to situate Neural Diffusion Models’ continuous formulation.
- Paper: Variational Diffusion Models, Diederik P. Kingma et al. (2021). Its likelihood-focused treatment of diffusion and learnable noise schedules prepares readers for Neural Diffusion Models’ density-estimation objective and schedule comparisons.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). Its modular account of diffusion design choices clarifies the baseline schedules and model components against which the source isolates learned nonlinear transformations.
- Paper: Diffusion Models: A Comprehensive Survey of Methods and Applications, Ling Yang et al. (2022). This survey establishes the principal diffusion families and their shared forward-corruption and reverse-generation framework that the source extends.
- Paper: Score-Based Diffusion Models in Function Space, Jae Hyun Lim 0001 et al. (2025). Building on Neural Diffusion Models’ flexible forward corruption, this work carries score-based diffusion into function spaces where the data are continuous fields rather than finite vectors.
