Fast Sampling of Diffusion Models via Operator Learning
Hongkai ZhengWeili NieArash VahdatKamyar AzizzadenesheliAnima Anandkumar
Proposes a neural operator framework that uses Fourier-parameterized temporal convolutions to map noise to complete reverse diffusion trajectories, enabling state-of-the-art image generation in a single forward pass through parallel decoding.
Diffusion models have established themselves as a leading approach for generative artificial intelligence across domains such as image generation, audio synthesis, and molecular modeling. However, their practical adoption in time-sensitive and interactive systems is severely limited by slow generation speeds. Standard diffusion models rely on iterative numerical solvers that require tens to hundreds of sequential neural network evaluations to generate a single output, creating substantial computational overhead.
The article demonstrates and evaluates Diffusion Model Sampling with Neural Operator (DSNO), a novel framework designed to achieve high-quality generation in a single model evaluation. DSNO formulates diffusion generation as an operator learning problem, mapping initial random noise directly to the continuous trajectory of the underlying differential equation via temporal parallel decoding.
To accomplish this, the authors integrated lightweight temporal convolution layers parameterized in Fourier space into standard diffusion neural network architectures. These layers increase overall model parameters by only about 10% and process temporal correlations across the generation trajectory simultaneously. The system was trained on solution trajectories generated from pre-trained diffusion models and evaluated on standard benchmarks, including unconditional CIFAR-10 image generation and class-conditional ImageNet-64 synthesis using perceptual quality metrics (FID) and mode coverage (Recall).
The evaluation produced four key findings. First, DSNO achieved new state-of-the-art single-step visual quality, recording an FID score of 3.78 on CIFAR-10 and 7.83 on ImageNet-64. Second, in a single forward pass, DSNO outperformed established one-step distillation baselines, lowering the ImageNet-64 FID from 15.99 down to 7.83. Third, the method delivered substantial inference speedups, operating approximately 2.6 times faster than four-step progressive distillation on CIFAR-10 and 1.7 times faster on ImageNet-64 while maintaining image quality comparable to multi-step methods. Fourth, DSNO retained a strong recall metric (0.61 on ImageNet-64), indicating that the diversity and mode coverage of the original diffusion models are fully preserved.
These results demonstrate that parallel temporal decoding via neural operators can overcome the sequential evaluation bottleneck of traditional diffusion solvers. By enabling single-pass generation with minimal added parameter cost, the approach significantly lowers runtime inference latency and computational expenses, making diffusion models practical for interactive consumer tools, real-time creative software, and embedded decision-making workflows.
For future work, development teams and researchers should focus on applying this operator framework to guided diffusion models and integrating Fourier temporal convolutions into emerging transformer-based diffusion backbones. Additionally, teams adopting this framework should utilize modern high-order differential equation solvers to generate offline training trajectories more cheaply, reducing upfront distillation data collection costs.
Confidence in these findings is supported by consistent benchmark improvements and extensive ablation studies covering loss weighting, discretization schemes, and temporal resolution. However, decision-makers should recognize current operational boundaries: the framework requires pre-computed trajectories from existing diffusion models for training, and while single-step fidelity improves markedly over prior single-step baselines, it still trails the absolute peak quality achievable by unconstrained, multi-step numerical solvers.
- Paper: Progressive Distillation for Fast Sampling of Diffusion Models, Tim Salimans et al. (2022). It introduces progressive distillation for fast few-step sampling in diffusion models, serving as the foundational distillation baseline that DSNO builds on and compares against.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). It formulates diffusion sampling as continuous trajectories via differential equations, providing the underlying continuous ODE formulation that DSNO maps directly with neural operators.
- Paper: DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps, Cheng Lu et al. (2022). It establishes dedicated fast ODE solvers for diffusion trajectory sampling, which DSNO suggests using offline to construct its training solution paths.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). It presents non-Markovian deterministic sampling (DDIM) that enables deterministic ODE-like sampling trajectories foundational to operator learning distillation.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). It provides the foundational formulation and standard U-Net architecture for denoising diffusion probabilistic models whose slow sequential sampling DSNO aims to replace.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). It establishes a standardized design space and high-order ODE discretization for diffusion models that underpin trajectory generation.
- Paper: Deep Equilibrium Approaches to Diffusion Models, Ashwini Pokle et al. (2022). It explores modeling entire diffusion generation trajectories simultaneously across time, motivating parallelized trajectory estimation paradigms.
- Paper: Score-Based Diffusion Models in Function Space, Jae Hyun Lim 0001 et al. (2025). It generalizes neural operator architectures directly into the core formulation of score-based diffusion models on infinite-dimensional function spaces.
- Paper: Fast ODE-based Sampling for Diffusion Models in Around 5 Steps, Zhenyu Zhou et al. (2024). It investigates trajectory geometry in diffusion sampling to achieve ultra-fast generation in few steps using auxiliary trajectory prediction.
- Paper: One-Step Diffusion Distillation through Score Implicit Matching, Weijian Luo et al. (2024). It develops a data-free single-step distillation method via score implicit matching, offering an alternative approach to one-step sampling without pre-computing offline trajectories.
- Paper: Inductive Moment Matching, Linqi Zhou et al. (2025). It explores single-stage fast generative modeling via inductive moment matching, extending the goal of high-fidelity, single-step sample generation.
- Paper: Variational Flow Maps: Make Some Noise for One-Step Conditional Generation, Abbas Mammadov et al. (2026). It adapts flow mapping concepts to enable rapid, one-step conditional generation and inverse problem solving directly in noise space.
