Fast Sampling of Diffusion Models via Operator Learning

Hongkai ZhengWeili NieArash VahdatKamyar AzizzadenesheliAnima Anandkumar

article2023ICML193 citations

Proposes a neural operator framework that uses Fourier-parameterized temporal convolutions to map noise to complete reverse diffusion trajectories, enabling state-of-the-art image generation in a single forward pass through parallel decoding.

Listen

Diffusion models have established themselves as a leading approach for generative artificial intelligence across domains such as image generation, audio synthesis, and molecular modeling. However, their practical adoption in time-sensitive and interactive systems is severely limited by slow generation speeds. Standard diffusion models rely on iterative numerical solvers that require tens to hundreds of sequential neural network evaluations to generate a single output, creating substantial computational overhead.

The article demonstrates and evaluates Diffusion Model Sampling with Neural Operator (DSNO), a novel framework designed to achieve high-quality generation in a single model evaluation. DSNO formulates diffusion generation as an operator learning problem, mapping initial random noise directly to the continuous trajectory of the underlying differential equation via temporal parallel decoding.

To accomplish this, the authors integrated lightweight temporal convolution layers parameterized in Fourier space into standard diffusion neural network architectures. These layers increase overall model parameters by only about 10% and process temporal correlations across the generation trajectory simultaneously. The system was trained on solution trajectories generated from pre-trained diffusion models and evaluated on standard benchmarks, including unconditional CIFAR-10 image generation and class-conditional ImageNet-64 synthesis using perceptual quality metrics (FID) and mode coverage (Recall).

The evaluation produced four key findings. First, DSNO achieved new state-of-the-art single-step visual quality, recording an FID score of 3.78 on CIFAR-10 and 7.83 on ImageNet-64. Second, in a single forward pass, DSNO outperformed established one-step distillation baselines, lowering the ImageNet-64 FID from 15.99 down to 7.83. Third, the method delivered substantial inference speedups, operating approximately 2.6 times faster than four-step progressive distillation on CIFAR-10 and 1.7 times faster on ImageNet-64 while maintaining image quality comparable to multi-step methods. Fourth, DSNO retained a strong recall metric (0.61 on ImageNet-64), indicating that the diversity and mode coverage of the original diffusion models are fully preserved.

These results demonstrate that parallel temporal decoding via neural operators can overcome the sequential evaluation bottleneck of traditional diffusion solvers. By enabling single-pass generation with minimal added parameter cost, the approach significantly lowers runtime inference latency and computational expenses, making diffusion models practical for interactive consumer tools, real-time creative software, and embedded decision-making workflows.

For future work, development teams and researchers should focus on applying this operator framework to guided diffusion models and integrating Fourier temporal convolutions into emerging transformer-based diffusion backbones. Additionally, teams adopting this framework should utilize modern high-order differential equation solvers to generate offline training trajectories more cheaply, reducing upfront distillation data collection costs.

Confidence in these findings is supported by consistent benchmark improvements and extensive ablation studies covering loss weighting, discretization schemes, and temporal resolution. However, decision-makers should recognize current operational boundaries: the framework requires pre-computed trajectories from existing diffusion models for training, and while single-step fidelity improves markedly over prior single-step baselines, it still trails the absolute peak quality achievable by unconstrained, multi-step numerical solvers.

arXiv: 2211.13449
Cover for Fast Sampling of Diffusion Models via Operator Learning

Abstract

Diffusion models have found widespread adoption in various areas. However, their sampling process is slow because it requires hundreds to thousands of network evaluations to emulate a continuous process defined by differential equations. In this work, we use neural operators, an efficient method to solve the probability flow differential equations, to accelerate the sampling process of diffusion models. Compared to other fast sampling methods that have a sequential nature, we are the first to propose a parallel decoding method that generates images with only one model forward pass. We propose diffusion model sampling with neural operator (DSNO) that maps the initial condition, i.e., Gaussian distribution, to the continuous-time solution trajectory of the reverse diffusion process. To model the temporal correlations along the trajectory, we introduce temporal convolution layers that are parameterized in the Fourier space into the given diffusion model backbone. We show our method achieves state-of-the-art FID of 3.78 for CIFAR-10 and 7.83 for ImageNet-64 in the one-model-evaluation setting.

Table of Contents

  • 1. Introduction
  • Our contributions.
  • 2. Background
  • 3. Learning the trajectory with neural operator
  • 4. Experiments
  • 4.1. Experimental setup
  • 4.2. Unconditional generation: CIFAR-10
  • 4.3. Conditional generation: ImageNet-64
  • 4.4. Ablation study
  • 5. Related work
  • 6. Conclusion and discussion
  • Acknowledgements
  • References
  • A. Appendix
  • A.1. Energy spectrum
  • A.2. Background: neural operators
  • A.3. Extended set of generated samples
  • A.4. Generalization to different resolution
  • A.5. Further discussion

Knowls

  1. Knowl 1 — Diffusion Model Sampling with Neural Operator (DSNO)

    model/method

    Diffusion Model Sampling with Neural Operator (DSNO) is a generative modeling method that accelerates the sampling process of diffusion models to a single neural network forward evaluation. Instead of solving the reverse probability flow ordinary differential equation (ODE) step by step via iterative numerical integration, DSNO frames sampling as learning an operator mapping between function spaces.

    Let A\mathcal{A} denote the space of initial noise vectors x(T)∼N(0,I)x(T) \sim \mathcal{N}(0, I) at terminal time TT, and let U=U(D;Rd)\mathcal{U} = \mathcal{U}(D; \mathbb{R}^d) denote the space of continuous-time trajectory functions over temporal domain D=[0,s]D = [0, s] with 0<s≤T0 < s \le T, where x(0)∈Rdx(0) \in \mathbb{R}^d corresponds to clean data. DSNO uses a neural operator Gθ:A→UG_\theta: \mathcal{A} \to \mathcal{U} parameterized by weights θ\theta to approximate the true ODE solution operator G†:A→UG^\dagger: \mathcal{A} \to \mathcal{U} directly in one forward evaluation, simultaneously decoding trajectory states across multiple queried time points.

  2. Knowl 2 — Temporal Fourier Convolution Block for Solution Trajectories

    model/method

    To capture temporal correlations along the reverse probability flow trajectory without altering spatial convolutions, DSNO incorporates Fourier temporal convolution layers into the levels of a standard diffusion U-Net architecture.

    Given an intermediate feature trajectory function u:D→Rdu: D \to \mathbb{R}^d, the temporal convolution layer T\mathcal{T} is defined as:

    (Tu)(t)=u(t)+σ((Ku)(t))(\mathcal{T}u)(t) = u(t) + \sigma((Ku)(t))

    where σ\sigma is a pointwise activation function (such as LeakyReLU), and KK is an integral kernel operator parameterized in Fourier space:

    (Ku)(t)=F−1(R⋅(Fu))(t)=∫D(F−1R)(τ)u(t−τ)dτ(Ku)(t) = \mathcal{F}^{-1}(R \cdot (\mathcal{F}u))(t) = \int_D (\mathcal{F}^{-1}R)(\tau) u(t - \tau) d\tau

    Here, F\mathcal{F} and F−1\mathcal{F}^{-1} denote the Fourier transform and inverse Fourier transform across the discretized temporal domain {t1,…,tM}\{t_1, \dots, t_M\}, and R∈CJ×d×dR \in \mathbb{C}^{J \times d \times d} is a learnable complex-valued parameter matrix operating on the truncated lowest JJ frequency modes. For discrete inputs u∈RM×du \in \mathbb{R}^{M \times d}, the matrix product at mode j∈{1,…,J}j \in \{1, \dots, J\} and output channel k∈{1,…,d}k \in \{1, \dots, d\} is:

    (R⋅(Fu))j,k=∑l=1dRj,k,l(Fu)j,l(R \cdot (\mathcal{F}u))_{j,k} = \sum_{l=1}^d R_{j,k,l} (\mathcal{F}u)_{j,l}

    The temporal convolution blocks operate along the temporal and channel dimensions (M×CM \times C) while treating spatial dimensions (H×WH \times W) as batch dimensions, increasing model parameter count by approximately 10% over the original U-Net backbone.

  3. Knowl 3 — Parallel Temporal Decoding for Diffusion Trajectories

    model/method

    DSNO achieves single-step sampling by performing parallel temporal decoding across all discretized time locations {t1,…,tM}\{t_1, \dots, t_M\} in the trajectory. Because the solutions x(ti)x(t_i) of the probability flow ODE at different times tit_i are conditionally independent given the initial noise condition x(T)∼N(0,I)x(T) \sim \mathcal{N}(0, I), the trajectory can be evaluated simultaneously.

    During a forward pass with input noise x(T)x(T), the feature map of the initial convolution layer is repeated MM times across the temporal dimension and paired with the respective time embeddings {t1,…,tM}\{t_1, \dots, t_M\}. In each temporal Fourier convolution layer, the spectral transformation F\mathcal{F}, parameter multiplication R⋅(Fu)R \cdot (\mathcal{F}u), and inverse transformation F−1\mathcal{F}^{-1} compute features across all MM time points concurrently, generating the entire reverse trajectory {x^(ti)}i=1M\{\hat{x}(t_i)\}_{i=1}^M including the final data sample x^(0)\hat{x}(0) in one model call.

  4. Knowl 4 — Universal Approximation of the Diffusion Probability Flow ODE Operator

    theoretical result

    Let the reverse diffusion process be governed by the probability flow ODE:

    dx=f(x,t)dt−12g(t)2∇xlog⁡pt(x)dtdx = f(x, t)dt - \frac{1}{2} g(t)^2 \nabla_x \log p_t(x) dt

    with affine drift f(x,t)=h(t)xf(x, t) = h(t)x. For any t<st < s, the exact solution trajectory satisfies the integral equation:

    x(t)=ϕ(t,s)x(s)−∫stϕ(t,τ)g(τ)22∇xlog⁡pτ(x)dτx(t) = \phi(t, s)x(s) - \int_s^t \phi(t, \tau) \frac{g(\tau)^2}{2} \nabla_x \log p_\tau(x) d\tau

    where ϕ(t,s)=exp⁡(∫sth(τ)dτ)\phi(t, s) = \exp\left(\int_s^t h(\tau) d\tau\right). The solution operator G†:A→UG^\dagger: \mathcal{A} \to \mathcal{U} mapping an initial condition x(T)∼N(0,I)x(T) \sim \mathcal{N}(0, I) to the continuous trajectory {x(t)}t∈[0,s]\{x(t)\}_{t \in [0, s]} is a well-defined, unique weighted integral operator.

    By the universal approximation theorem for Fourier neural operators on Banach spaces, the class of neural operators GθG_\theta can approximate the true solution operator G†G^\dagger arbitrarily well in the operator norm.

  5. Knowl 5 — Training Objective and Empirical Risk Formulation for DSNO

    equation

    DSNO parameters θ\theta are trained using empirical risk minimization over a pre-sampled dataset of NN ground-truth probability flow trajectories {G†(xT(j))}j=1N\{G^\dagger(x_T^{(j)})\}_{j=1}^N generated by an existing numerical ODE solver (such as DDIM or progressive distillation):

    min⁡θ1N∑j=1N1M∑i=1Mλ(ti)∥Gθ(xT(j))(ti)−G†(xT(j))(ti)∥\min_\theta \frac{1}{N} \sum_{j=1}^N \frac{1}{M} \sum_{i=1}^M \lambda(t_i) \|G_\theta(x_T^{(j)})(t_i) - G^\dagger(x_T^{(j)})(t_i)\|

    where {t1,…,tM}\{t_1, \dots, t_M\} are MM discretization points on the temporal domain DD, ∥⋅∥\|\cdot\| is a norm or perceptual distance metric (such as the ℓ1\ell^1 norm or VGG-based LPIPS), and λ(t)\lambda(t) is a time-dependent loss weighting function defined by the square root of the signal-to-noise ratio (SNR):

    λ(t)=αtσt=SNR(t)\lambda(t) = \frac{\alpha_t}{\sigma_t} = \sqrt{\text{SNR}(t)}

    where αt\alpha_t and σt\sigma_t correspond to the noise schedule parameters of the diffusion model at time tt.

  6. Knowl 6 — Energy Spectrum Concentration of Diffusion ODE Trajectories

    empirical result

    The discrete-time Fourier power spectrum SjS_j of diffusion probability flow ODE trajectories x(t)x(t) with period T=1T=1 and time step Δ=1/N\Delta = 1/N sampled at 1000 Hz is given by:

    Sj=2Δ2TXjXj∗,Xj=∑i=1Nx(ti)exp⁡(−2πjitiT)S_j = \frac{2\Delta^2}{T} X_j X_j^*, \quad X_j = \sum_{i=1}^N x(t_i) \exp\left(-\frac{2\pi j i t_i}{T}\right)

    Across pixel locations and channels of trajectories generated by standard models (such as DDPM++ on CIFAR-10), the power spectrum is compact and heavily concentrated in the low-frequency regime below 5 Hz. High-frequency modes contribute negligibly to trajectory dynamics, which allows Fourier temporal convolution blocks with mode truncation to approximate the solution operator accurately using small temporal discretization resolutions (such as M=4M = 4).

  7. Knowl 7 — CIFAR-10 Fast Sampling Quality Comparison

    data/table

    The following table compares the sample generation quality (measured by Fréchet Inception Distance, FID-50K) and the number of function evaluations (NFE) on unconditional CIFAR-10 across fast sampling methods.

    Method NFE FID Model size
    DSNO (Ours) 1 3.78 65.8M
    Knowledge distillation (Luhman Luhman, 2021) 1 9.36 35.7M
    Progressive distillation (Salimans Ho, 2021) 1 9.12 60.0M
    Progressive distillation (Salimans Ho, 2021) 2 4.51 60.0M
    Progressive distillation (Salimans Ho, 2021) 4 3.00 60.0M
    LSGM (Vahdat et al., 2021) 147 2.10 475.0M
    GGDM + PRED + TIME (Watson et al., 2021) 5 13.77 35.7M
    GGDM + PRED + TIME (Watson et al., 2021) 10 8.23 35.7M
    DDIM (Song et al., 2020a) 10 13.36 35.7M
    DDIM (Song et al., 2020a) 20 6.84 35.7M
    DDIM (Song et al., 2020a) 50 4.67 35.7M
    SN-DDIM (Bao et al., 2022) 10 12.19 52.6M
    FastDPM (Kong Ping, 2021) 10 9.90 35.7M
    DPM-solver (Lu et al., 2022) 10 4.70 35.7M
    DEIS (Zhang Chen, 2022) 10 4.17 -
    TDPM (Zheng et al., 2022) 5 3.34 35.7M
    DDGAN (Xiao et al., 2021) 4 3.75 -

    In the single-model-evaluation setting (NFE = 1), DSNO achieves an FID of 3.78, substantially outperforming single-step knowledge distillation (FID 9.36) and single-step progressive distillation (FID 9.12), as well as two-step progressive distillation (FID 4.51).

  8. Knowl 8 — Class-Conditional ImageNet-64 Fast Sampling Performance

    data/table

    The following table compares image sample quality (FID-50K) and mode coverage (Recall) on class-conditional ImageNet-64×6464\times 64 across sampling techniques.

    Method Model evaluations FID score Recall Model size
    DSNO (Ours) 1 7.83 0.61 329.2M
    Progressive distillation (Salimans Ho, 2021) 1 15.99 0.60 295.9M
    Progressive distillation (Salimans Ho, 2021) 2 7.11 0.63 295.9M
    Progressive distillation (Salimans Ho, 2021) 4 3.84 0.63 295.9M
    EDM (Karras et al., 2022) 79 2.44 0.67 295.9M
    DDIM (Song et al., 2020a) 32 5.00 - 295.9M
    BigGAN-deep (Brock et al., 2018) 1 4.06 0.48 -
    ADM (Dhariwal Nichol, 2021) 250 2.07 0.63 295.9M

    With 1 model evaluation, DSNO achieves an FID score of 7.83 and Recall of 0.61 on ImageNet-64, outperforming 1-step progressive distillation (FID 15.99, Recall 0.60) and achieving sample quality close to 2-step progressive distillation (FID 7.11) while preserving mode coverage comparable to the full ADM baseline (Recall 0.63).

  9. Knowl 9 — Computational Runtime and Parameter Overhead of DSNO

    data/table

    The computational latency and parameter scaling of DSNO relative to the baseline diffusion backbone are evaluated on an NVIDIA V100 GPU (averaged over 20 runs following 20 warm-up runs):

    Backbone Runtime Model size
    CIFAR-10 baseline 0.033s 60.00M
    DSNO-CIFAR-10 (ours) 0.050s 65.77M
    ImageNet-64 baseline 0.066s 295.90M
    DSNO-ImageNet-64 (ours) 0.080s 329.23M

    Adding Fourier temporal convolution layers increases total model parameters by approximately 9.6% on CIFAR-10 and 11.3% on ImageNet-64. In wall-clock inference time, 1-step DSNO is 2.6 times faster than 4-step progressive distillation (0.132s vs 0.050s) and 1.3 times faster than 2-step progressive distillation (0.066s vs 0.050s) on CIFAR-10, and 1.7 times faster than 2-step progressive distillation on ImageNet-64.

  10. Knowl 10 — Ablation Studies on DSNO Architectural Components and Hyperparameters

    empirical result

    Ablation experiments conducted on CIFAR-10 reveal the individual contributions of DSNO design components:

    1. Temporal Convolution Layers: Incorporating temporal convolution blocks into the U-Net improves CIFAR-10 FID from 8.09 to 4.23 at 300k training steps, and from 7.85 to 4.12 at 400k training steps (batch size 256, temporal resolution M=4M=4, quadratic time steps).

    2. Loss Weighting Scheme: Applying the square-root SNR weighting function λ(t)=SNR(t)=αtσt\lambda(t) = \sqrt{\text{SNR}(t)} = \frac{\alpha_t}{\sigma_t} improves FID from 4.56 (uniform weighting) to 4.21.

    3. Time Discretization Scheme: Quadratic time step discretization yields an FID of 4.21 compared to 4.33 for uniform time step discretization.

    4. Temporal Resolution MM: Increasing the number of trajectory supervision steps during training improves sample quality: M=2M=2 gives FID 5.01, M=4M=4 gives FID 4.21, and M=8M=8 gives FID 3.98.

    5. Loss Function: Replacing the ℓ1\ell^1 loss (FID 4.12) with the VGG-based LPIPS perceptual loss further improves the CIFAR-10 FID to 3.78.

Coverage note — None was omitted; all key theoretical claims, architectural components, training procedures, spectral properties, benchmark evaluation data, runtime analyses, and ablation experiments are fully represented.

References

  1. 1.Ajay, A., Du, Y., Gupta, A., Tenenbaum, J., Jaakkola, T., and Agrawal, P. Is conditional generative modeling all you need for decision-making? arXiv preprint arXiv:2211.15657, 2022.
  2. 2.Bao, F., Li, C., Zhu, J., and Zhang, B. Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models. In International Conference on Learning Representations, 2021.
  3. 3.Bao, F., Li, C., Sun, J., Zhu, J., and Zhang, B. Estimating the optimal covariance with imperfect mean in diffusion probabilistic models. arXiv preprint arXiv:2206.07309, 2022.
  4. 4.Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J. All are worth words: A vit backbone for diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 22669–22679, 2023.
  5. 5.Brock, A., Donahue, J., and Simonyan, K. Large scale gan training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096, 2018.
  6. 6.Chang, H., Zhang, H., Barber, J., Maschinot, A., Lezama, J., Jiang, L., Yang, M.-H., Murphy, K., Freeman, W. T., Rubinstein, M., et al. Muse: Text-to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023.
  7. 7.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  8. 8.Dockhorn, T., Vahdat, A., and Kreis, K. GENIE: Higher-Order Denoising Diffusion Solvers. In Advances in Neural Information Processing Systems, 2022.
  9. 9.Ghazvininejad, M., Levy, O., Liu, Y., and Zettlemoyer, L. Mask-predict: Parallel decoding of conditional masked language models. arXiv preprint arXiv:1904.09324, 2019.
  10. 10.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020.
  11. 11.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  12. 12.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017.
  13. 13.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  14. 14.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. arXiv preprint arXiv:2206.00364, 2022.
  15. 15.Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014.
  16. 16.Kong, Z. and Ping, W. On fast sampling of diffusion probabilistic models. arXiv preprint arXiv:2106.00132, 2021.
  17. 17.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. In ICLR, 2021.
  18. 18.Kovachki, N., Lanthaler, S., and Mishra, S. On universal approximation and error bounds for fourier neural operators. Journal of Machine Learning Research, 22:Art–No, 2021a.
  19. 19.Kovachki, N., Li, Z., Liu, B., Azizzadenesheli, K., Bhattacharya, K., Stuart, A., and Anandkumar, A. Neural operator: Learning maps between function spaces. arXiv preprint arXiv:2108.08481, 2021b.
  20. 20.Kynk¨a¨anniemi, T., Karras, T., Laine, S., Lehtinen, J., and Aila, T. Improved precision and recall metric for assessing generative models. Advances in Neural Information Processing Systems, 32, 2019.
  21. 21.Lam, M. W., Wang, J., Huang, R., Su, D., and Yu, D. Bilateral denoising diffusion models. arXiv preprint arXiv:2108.11514, 2021.
  22. 22.Li, Z., Kovachki, N., Azizzadenesheli, K., Liu, B., Bhattacharya, K., Stuart, A., and Anandkumar, A. Fourier neural operator for parametric partial differential equations. arXiv preprint arXiv:2010.08895, 2020a.
  23. 23.Li, Z., Kovachki, N., Azizzadenesheli, K., Liu, B., Bhattacharya, K., Stuart, A., and Anandkumar, A. Neural operator: Graph kernel network for partial differential equations. arXiv preprint arXiv:2003.03485, 2020b.
  24. 24.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022.
  25. 25.Luhman, E. and Luhman, T. Knowledge distillation in iterative generative models for improved sampling speed. arXiv preprint arXiv:2101.02388, 2021.
  26. 26.Meng, C., Gao, R., Kingma, D. P., Ermon, S., Ho, J., and Salimans, T. On distillation of guided diffusion models. arXiv preprint arXiv:2210.03142, 2022.
  27. 27.Nie, W., Guo, B., Huang, Y., Xiao, C., Vahdat, A., and Anandkumar, A. Diffusion models for adversarial purification. In International Conference on Machine Learning (ICML), 2022.
  28. 28.Peebles, W. and Xie, S. Scalable diffusion models with transformers. arXiv preprint arXiv:2212.09748, 2022.
  29. 29.Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 2022.
  30. 30.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. URL https://github.com/CompVis/latent-diffusionhttps://arxiv.org/abs/2112.10752.
  31. 31.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, 2021.
  32. 32.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265. PMLR, 2015.
  33. 33.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020a.
  34. 34.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In ICLR, 2021.
  35. 35.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020b.
  36. 36.Song, Y., Dhariwal, P., Chen, M., and Sutskever, I. Consistency models. arXiv preprint arXiv:2303.01469, 2023.
  37. 37.Vahdat, A., Kreis, K., and Kautz, J. Score-based generative modeling in latent space. Advances in Neural Information Processing Systems, 34:11287–11302, 2021.
  38. 38.Watson, D., Chan, W., Ho, J., and Norouzi, M. Learning fast samplers for diffusion models by differentiating through sample quality. In International Conference on Learning Representations, 2021.
  39. 39.Wen, G., Li, Z., Long, Q., Azizzadenesheli, K., Anandkumar, A., and Benson, S. M. Accelerating carbon capture and storage modeling using fourier neural operators. arXiv preprint arXiv:2210.17051, 2022.
  40. 40.Xiao, Z., Kreis, K., and Vahdat, A. Tackling the generative learning trilemma with denoising diffusion gans. In International Conference on Learning Representations, 2021.
  41. 41.Xu, M., Yu, L., Song, Y., Shi, C., Ermon, S., and Tang, J. Geodiff: A geometric diffusion model for molecular conformation generation. arXiv preprint arXiv:2203.02923, 2022.
  42. 42.Yang, Y., Gao, A. F., Castellanos, J. C., Ross, Z. E., Azizzadenesheli, K., and Clayton, R. W. Seismic wave propagation and inversion with neural operators. The Seismic Record, 1(3):126–134, 2021.
  43. 43.Zhang, Q. and Chen, Y. Fast sampling of diffusion models with exponential integrator. arXiv preprint arXiv:2204.13902, 2022.
  44. 44.Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 586–595, 2018.
  45. 45.Zheng, H., He, P., Chen, W., and Zhou, M. Truncated diffusion probabilistic models and diffusion-based adversarial auto-encoders. arXiv preprint arXiv:2202.09671, 2022.

Citation

MLA
Zheng, H., et al. “Fast Sampling of Diffusion Models via Operator Learning”. International Conference on Machine Learning, vol. 202, 2023, pp. 42390–402, https://proceedings.mlr.press/v202/zheng23d.html.
APA
Zheng, H., Nie, W., Vahdat, A., Azizzadenesheli, K., & Anandkumar, A. (2023). Fast Sampling of Diffusion Models via Operator Learning. International Conference on Machine Learning, 202, 42390–42402. https://proceedings.mlr.press/v202/zheng23d.html
Chicago
Zheng, H., W. Nie, A. Vahdat, K. Azizzadenesheli, and A. Anandkumar. 2023. “Fast Sampling of Diffusion Models via Operator Learning”. International Conference on Machine Learning 202: 42390–402. https://proceedings.mlr.press/v202/zheng23d.html.
Harvard
Zheng, H. et al. (2023) “Fast Sampling of Diffusion Models via Operator Learning”, International Conference on Machine Learning. PMLR, pp. 42390–42402. Available at: https://proceedings.mlr.press/v202/zheng23d.html.
Vancouver
1. Zheng H, Nie W, Vahdat A, Azizzadenesheli K, Anandkumar A (2023) Fast Sampling of Diffusion Models via Operator Learning. In: International Conference on Machine Learning. PMLR, pp 42390–42402

BibTeX

@InProceedings{pmlr-v202-zheng23d,
  title = 	 {Fast Sampling of Diffusion Models via Operator Learning},
  author =       {Zheng, Hongkai and Nie, Weili and Vahdat, Arash and Azizzadenesheli, Kamyar and Anandkumar, Anima},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {42390--42402},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/zheng23d/zheng23d.pdf},
  url = 	 {https://proceedings.mlr.press/v202/zheng23d.html},
  abstract = 	 {Diffusion models have found widespread adoption in various areas. However, their sampling process is slow because it requires hundreds to thousands of network evaluations to emulate a continuous process defined by differential equations. In this work, we use neural operators, an efficient method to solve the probability flow differential equations, to accelerate the sampling process of diffusion models. Compared to other fast sampling methods that have a sequential nature, we are the first to propose a parallel decoding method that generates images with only one model forward pass. We propose diffusion model sampling with neural operator (DSNO) that maps the initial condition, i.e., Gaussian distribution, to the continuous-time solution trajectory of the reverse diffusion process. To model the temporal correlations along the trajectory, we introduce temporal convolution layers that are parameterized in the Fourier space into the given diffusion model backbone. We show our method achieves state-of-the-art FID of 3.78 for CIFAR-10 and 7.83 for ImageNet-64 in the one-model-evaluation setting.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/