Accelerating Diffusion Sampling with Optimized Time Steps

Shuchen XueZhaoqiang LiuFei ChenShifeng ZhangTianyang HuEnze XieZhenguo Li

article2024CVPR67 citations

Proposes a training-free optimization framework that computes non-uniform sampling schedules in under 15 seconds to minimize numerical solver approximation errors, substantially improving diffusion generation quality at very low step counts.

Listen

Diffusion models have become the leading approach for high-quality image generation, yet their practical deployment is constrained by high computational costs. Generating an image typically requires evaluating complex neural networks over many sequential steps. While modern numerical solvers have reduced the required step count, standard sampling approaches still divide generation time into uniform steps, leaving significant room for performance improvements when operating under extremely small step budgets.

The article establishes a general, training-free optimization framework to find non-uniform time steps tailored to numerical solvers for diffusion models. The primary objective is to formulate and efficiently solve an optimization problem that minimizes the mathematical distance between the exact solution of the generative trajectory and the approximation produced by the numerical solver.

To achieve this, the authors derived theoretical error bounds for generative trajectories and structured an optimization problem using standard score approximation assumptions. This framework accommodates any explicit solver, including algorithms with variable step orders. The authors solved the resulting formulation using the constrained trust region method and evaluated the optimized schedules across standard image benchmarks (such as CIFAR-10, ImageNet, FFHQ, and AFHQv2) using popular diffusion architectures (Score-SDE, ADM, EDM, and DiT) across both pixel- and latent-space models.

The findings show substantial generation quality improvements, measured by lower Fréchet Inception Distance (FID) scores, especially in few-step regimes. When paired with the high-order UniPC solver at five network evaluations, the proposed method reduced FID on CIFAR-10 from 23.22 (using uniform schedules) to 12.11, improved ImageNet 64x64 FID to 10.47, and lowered ImageNet 256x256 FID from 23.48 to 8.66. Across all tested datasets and architectures, optimized schedules consistently outperformed conventional uniform schemes. Furthermore, calculating the optimal schedule requires under 15 seconds on standard central processing units, contrasting sharply with previous reinforcement learning or search-based scheduling techniques that require hours of expensive graphics processing unit compute.

These results demonstrate that optimizing time step intervals delivers dramatic inference speedups without requiring model retraining, architecture modifications, or costly schedule searches. Organizations deploying diffusion models can directly lower computational latency and server infrastructure costs while maintaining or improving output fidelity. The framework integrates seamlessly as a plug-and-play enhancement for existing pre-trained pipelines.

Teams maintaining generative diffusion systems should adopt optimized time step schedules alongside advanced solvers such as UniPC, prioritizing applications where latency is critical. Future engineering and research efforts should explore refining the surrogate error objective for higher precision and extending the optimization framework to implicit solver components.

arXiv: 2402.17376
Cover for Accelerating Diffusion Sampling with Optimized Time Steps

Abstract

Diffusion probabilistic models (DPMs) have shown remarkable performance in high-resolution image synthesis, but their sampling efficiency is still to be desired due to the typically large number of sampling steps. Recent advancements in high-order numerical ODE solvers for DPMs have enabled the generation of high-quality images with much fewer sampling steps. While this is a significant development, most sampling methods still employ uniform time steps, which is not optimal when using a small number of steps. To address this issue, we propose a general framework for designing an optimization problem that seeks more appropriate time steps for a specific numerical ODE solver for DPMs. This optimization problem aims to minimize the distance between the ground-truth solution to the ODE and an approximate solution corresponding to the numerical solver. It can be efficiently solved using the constrained trust region method, taking less than 15 seconds. Our extensive experiments on both unconditional and conditional sampling using pixel- and latent-space DPMs demonstrate that, when combined with the state-of-the-art sampling method UniPC, our optimized time steps significantly improve image generation performance in terms of FID scores for datasets such as CIFAR-10 and ImageNet, compared to using uniform time steps.^1

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Preliminary
  • 3.1. Diffusion Models
  • 3.2. Discretization Schemes
  • 4. Problem Formulation
  • 5. Analysis and Method
  • 6. Experiments
  • 6.1. Pixel Diffusion Model Generation
  • 6.2. Latent Diffusion Model Generation
  • 6.3. Running Time Analysis
  • 7. Conclusion
  • References

Knowls

  1. Knowl 1 — Optimization-based construction of diffusion sampling schedules

    model/method

    The paper introduces a training-free framework that optimizes the intermediate sampling times for a fixed explicit numerical solver of the reverse diffusion ODE. Let T=t0>t1>⋯>tN=ϵ>0T=t_0>t_1>\cdots>t_N=\epsilon>0 be the sampling times, let λt=log⁡(αt/σt)\lambda_t=\log(\alpha_t/\sigma_t) be the half log-SNR, and write λn=λtn\lambda_n=\lambda_{t_n}. Because λt\lambda_t is strictly decreasing in tt, the sequence satisfies λ0<λ1<⋯<λN\lambda_0<\lambda_1<\cdots<\lambda_N.

    For an explicit solver whose local approximation at step nn has order knk_n, define f(λ)=xθ(xtλ,tλ)f(\lambda)=x_\theta(x_{t_\lambda},t_\lambda), where tλt_\lambda is the inverse of λt\lambda_t. If the solver uses a local polynomial

    Pn;kn−1(λ)=∑j=0kn−1ℓn;kn,j(λ)f(λn−kn+j),\mathcal P_{n;k_n-1}(\lambda)=\sum_{j=0}^{k_n-1}\ell_{n;k_n,j}(\lambda)f(\lambda_{n-k_n+j}),

    its contribution weights are

    wn;kn,j=∫λn−1λneλℓn;kn,j(λ) dλ.w_{n;k_n,j}=\int_{\lambda_{n-1}}^{\lambda_n}e^\lambda\ell_{n;k_n,j}(\lambda)\,d\lambda.

    The resulting approximate endpoint is

    x~ϵ=σϵσTxT+σϵ∑n=1N∑j=0kn−1wn;kn,jf(λn−kn+j),\widetilde x_\epsilon=\frac{\sigma_\epsilon}{\sigma_T}x_T+\sigma_\epsilon\sum_{n=1}^{N}\sum_{j=0}^{k_n-1}w_{n;k_n,j}f(\lambda_{n-k_n+j}),

    where xTx_T is the initial noisy sample and σt\sigma_t is the diffusion noise scale. The intermediate values λ1,…,λN−1\lambda_1,\ldots,\lambda_{N-1} are selected by minimizing the score-error-weighted coefficient magnitude

    min⁡λ1,…,λN−1∑i=0N−1ε~ti∣∑n−kn+j=iwn;kn,j∣subject toλn+1>λn,\min_{\lambda_1,\ldots,\lambda_{N-1}}\quad \sum_{i=0}^{N-1}\widetilde\varepsilon_{t_i}\left|\sum_{n-k_n+j=i}w_{n;k_n,j}\right| \quad\text{subject to}\quad \lambda_{n+1}>\lambda_n,

    with fixed endpoints λ0=λT\lambda_0=\lambda_T and λN=λϵ\lambda_N=\lambda_\epsilon. In the experiments, ε~t=σtp/αt\widetilde\varepsilon_t=\sigma_t^p/\alpha_t, using p=1p=1 for pixel-space models and p=2p=2 for latent-space models. The optimization is solver-specific but does not require retraining the diffusion model.

  2. Knowl 2 — Solver-independent weighted approximation with variable local orders

    model/method

    The optimized-schedule framework applies to any explicit diffusion ODE solver that represents the score-dependent integrand locally by a polynomial in the half log-SNR λ\lambda. The local order knk_n may vary with the step index, subject to 1≤kn≤n1\leq k_n\leq n; therefore, the framework covers startup steps, fixed high-order solvers, and customized order schedules.

    For Lagrange interpolation, the basis functions are

    ℓn;kn,j(λ)=∏i=0i≠jkn−1λ−λn−kn+iλn−kn+j−λn−kn+i.\ell_{n;k_n,j}(\lambda)=\prod_{\substack{i=0\\i\ne j}}^{k_n-1}\frac{\lambda-\lambda_{n-k_n+i}}{\lambda_{n-k_n+j}-\lambda_{n-k_n+i}}.

    Taylor-based solvers also fit the same representation after approximating derivatives of f(λ)f(\lambda) from previously evaluated neural-network values. The paper observes that the weights satisfy the consistency identity

    ∑j=0kn−1wn;kn,j=eλn−eλn−1,\sum_{j=0}^{k_n-1}w_{n;k_n,j}=e^{\lambda_n}-e^{\lambda_{n-1}},

    which holds when the approximated integrand is constant. This identity makes the schedule objective applicable to different explicit solvers without changing its overall form. The authors instantiate the construction for DPM-Solver++ and UniPC; for UniPC, the optimization uses the predictor-style explicit approximation and does not explicitly model the corrector.

  3. Knowl 3 — Probabilistic sampling-error bound motivating the schedule objective

    theoretical result

    Assume that the score network sθ(x,t)s_\theta(x,t) satisfies, at every grid time t∈{t0,…,tN}t\in\{t_0,\ldots,t_N\},

    Ex∼qt[∥∇xlog⁡qt(x)−sθ(x,t)∥2]≤η2εt2,\mathbb E_{x\sim q_t}\left[\left\|\nabla_x\log q_t(x)-s_\theta(x,t)\right\|^2\right]\leq \eta^2\varepsilon_t^2,

    where qtq_t is the diffusion marginal, η>0\eta>0 is constant, and εt\varepsilon_t describes the time-dependent score-estimation error scale. Let the data-prediction network be related to the score by xθ(xt,t)=(σt2sθ(xt,t)+xt)/αtx_\theta(x_t,t)=(\sigma_t^2s_\theta(x_t,t)+x_t)/\alpha_t. For any failure probability P0∈(0,1)P_0\in(0,1), define

    η~=N+1P0 η,ε~t=εtσt2αt.\widetilde\eta=\sqrt{\frac{N+1}{P_0}}\,\eta, \qquad \widetilde\varepsilon_t=\frac{\varepsilon_t\sigma_t^2}{\alpha_t}.

    If the numerical weights satisfy ∑jwn;kn,j=eλn−eλn−1\sum_jw_{n;k_n,j}=e^{\lambda_n}-e^{\lambda_{n-1}}, then with probability at least 1−P01-P_0,

    ∥x~ϵ−x0∥≤∥σϵσTxT+σϵ(eλϵ−eλT)x0−x0∥+σϵη~∑i=0N−1ε~ti∣∑n−kn+j=iwn;kn,j∣,\begin{aligned} \|\widetilde x_\epsilon-x_0\| &\leq\left\|\frac{\sigma_\epsilon}{\sigma_T}x_T+\sigma_\epsilon\left(e^{\lambda_\epsilon}-e^{\lambda_T}\right)x_0-x_0\right\|\\ &\quad+\sigma_\epsilon\widetilde\eta\sum_{i=0}^{N-1}\widetilde\varepsilon_{t_i} \left|\sum_{n-k_n+j=i}w_{n;k_n,j}\right|, \end{aligned}

    where x0x_0 is the clean data sample and x~ϵ\widetilde x_\epsilon is the numerical endpoint. With fixed TT and ϵ\epsilon, the first term is independent of the intermediate schedule and is expected to be small when αT≈0\alpha_T\approx0 and σϵ≈0\sigma_\epsilon\approx0. Consequently, minimizing the second term yields the paper's timestep optimization problem as a surrogate for reducing numerical sampling error.

  4. Knowl 4 — Constrained trust-region algorithm for finding optimized time steps

    algorithm

    The schedule is computed once for a chosen solver, number of neural function evaluations, diffusion noise schedule, and score-error weighting exponent. The output is a feasible sequence of intermediate half-log-SNR values, which can be converted back to times using the inverse function tλt_\lambda.

    Input: Number of steps NN, start time TT, end time ϵ\epsilon, an explicit solver with local polynomials Pn;kn−1\mathcal P_{n;k_n-1}, exponent pp, and initial feasible values λ1,…,λN−1\lambda_1,\ldots,\lambda_{N-1}
    Output: Optimized intermediate values λ^1,…,λ^N−1\widehat\lambda_1,\ldots,\widehat\lambda_{N-1}
    Set λ0=λT\lambda_0=\lambda_T and λN=λϵ\lambda_N=\lambda_\epsilon
    For i=0,…,N−1i=0,\ldots,N-1, evaluate ε~ti=σtip/αti\widetilde\varepsilon_{t_i}=\sigma_{t_i}^p/\alpha_{t_i}
    For every solver step nn and local index jj, calculate wn;kn,j=∫λn−1λneλℓn;kn,j(λ) dλw_{n;k_n,j}=\int_{\lambda_{n-1}}^{\lambda_n}e^\lambda\ell_{n;k_n,j}(\lambda)\,d\lambda
    Construct the objective ∑i=0N−1ε~ti∣∑n−kn+j=iwn;kn,j∣\sum_{i=0}^{N-1}\widetilde\varepsilon_{t_i}\left|\sum_{n-k_n+j=i}w_{n;k_n,j}\right|
    Minimize the objective with a constrained trust-region optimizer subject to λn+1>λn\lambda_{n+1}>\lambda_n for all nn
    Return the optimized intermediate values and, if needed, convert them to times with tn=tλnt_n=t_{\lambda_n}

    The optimization uses no image samples or additional neural-network training. Its result can be precomputed and reused for repeated sampling with the same diffusion model, solver, and number of evaluations.

  5. Knowl 5 — Reference discretization schedules used for comparison

    model/method

    The experiments compare optimized schedules with three conventional discretizations of the interval [ϵ,T][\epsilon,T]. For NN steps, the uniform-time schedule is

    tn=T+nN(ϵ−T),n=0,…,N.t_n=T+\frac{n}{N}(\epsilon-T),\qquad n=0,\ldots,N.

    The uniform-λ\lambda schedule divides the half-log-SNR interval uniformly and maps back to time:

    tn=tλ(λT+nN(λϵ−λT)).t_n=t_\lambda\left(\lambda_T+\frac{n}{N}(\lambda_\epsilon-\lambda_T)\right).

    The EDM schedule defines κt=σt/αt\kappa_t=\sigma_t/\alpha_t, assumes that κt\kappa_t is strictly increasing, and uniformly discretizes κt1/ρ\kappa_t^{1/\rho}:

    tn=tκ([κT1/ρ+nN(κϵ1/ρ−κT1/ρ)]ρ),t_n=t_\kappa\left(\left[\kappa_T^{1/\rho}+\frac{n}{N}\left(\kappa_\epsilon^{1/\rho}-\kappa_T^{1/\rho}\right)\right]^\rho\right),

    where tκt_\kappa is the inverse of κt\kappa_t and the experiments use ρ=7\rho=7. These schedules are fixed heuristics, whereas the proposed schedule adapts its intermediate points to the coefficient structure of the selected numerical solver.

  6. Knowl 6 — Experimental evaluation protocol across pixel and latent diffusion models

    experimental setup

    The optimized schedules are evaluated with the third-order DPM-Solver++ and UniPC samplers using 5, 6, 7, 8, 9, 10, 12, or 15 neural function evaluations (NFEs). Image quality is measured by Fréchet Inception Distance (FID), generally using 50,000 generated images per evaluation.

    The evaluation covers unconditional pixel-space generation on CIFAR-10 at 32×3232\times32, FFHQ at 64×6464\times64, and AFHQv2 at 64×6464\times64; conditional pixel-space generation on ImageNet at 64×6464\times64; and conditional latent-space generation on ImageNet at 256×256256\times256 and 512×512512\times512. The models include Score-SDE, ADM, EDM, and DiT-XL-2. The DiT experiments use the KL-8 latent encoder-decoder and classifier-free guidance scale s=1.5s=1.5. The method is therefore tested across different resolutions, diffusion parameterizations, unconditional and conditional sampling, and pixel and latent representations.

  7. Knowl 7 — CIFAR-10 sampling-quality comparison over NFEs

    data/table

    The following FID values compare uniform-λ\lambda, uniform-tt, EDM, and optimized schedules for DPM-Solver++ and UniPC on unconditional pixel-space CIFAR-10 at 32×3232\times32. Lower FID is better. The optimized schedule is consistently the strongest or near-strongest choice, especially at very small NFE budgets; UniPC with the optimized schedule reaches FID 12.1112.11 at 5 NFEs, compared with 23.2223.22 for UniPC with uniform-λ\lambda.

    Could not parse LaTeX table

    The optimized schedule improves both solvers at low NFE counts. UniPC is better than DPM-Solver++ under every listed schedule at 5 NFEs, and the advantage of optimization diminishes as the NFE budget becomes large.

  8. Knowl 8 — Cross-dataset image-generation improvements at five NFEs

    empirical result

    At only 5 NFEs, the optimized schedule substantially improves FID across pixel-space and latent-space diffusion models. For unconditional pixel-space CIFAR-10, UniPC obtains FID 12.1112.11 with optimized steps versus 23.2223.22 with uniform-λ\lambda. For conditional pixel-space ImageNet 64×6464\times64, the optimized UniPC schedule obtains FID 10.4710.47. For unconditional EDM models, it obtains FID 13.6613.66 on FFHQ 64×6464\times64 and 12.1112.11 on AFHQv2 64×6464\times64.

    For conditional latent-space DiT-XL-2 models with classifier-free guidance scale s=1.5s=1.5, optimized UniPC obtains FID 8.668.66 on ImageNet 256×256256\times256, improving over 23.4823.48 with uniform-tt, and FID 11.4011.40 on ImageNet 512×512512\times512, improving over 20.2820.28 with uniform-tt. On the 256×256256\times256 task, the corresponding 5-NFE FIDs for optimized, uniform-tt, EDM, and uniform-λ\lambda schedules are respectively 8.668.66, 23.4823.48, 45.8945.89, and 41.8941.89 under the same random seed. The improvements persist across the reported NFE range of 5–15, although the differences narrow as more evaluations are used.

  9. Knowl 9 — Negligible precomputation cost of schedule optimization

    data/table

    The schedule optimizer was timed on an Intel(R) Xeon(R) Gold 6278C CPU at 2.60 GHz. The reported times are the longest observed optimization times for each NFE budget, not the time required to generate images. The result can be computed before sampling and reused.

    Could not parse LaTeX table

    The optimization completes within 15 seconds for every tested budget up to 15 NFEs. This makes the schedule-search overhead negligible relative to repeated diffusion sampling and much smaller than the GPU-hour-scale costs reported for learning-based schedule-search approaches.

  10. Knowl 10 — Surrogate-objective and corrector-related limitations

    limitation

    The optimized schedule minimizes an analytically derived upper-bound surrogate based on score-estimation scales and numerical-solver weights, rather than directly minimizing FID or the exact sample-distribution error. The paper therefore notes that a more accurate optimization objective could further improve the method.

    The schedule objective is formulated for explicit local polynomial approximations and does not account for UniPC's implicit predictor-corrector component. Nevertheless, the optimized steps work effectively with UniPC in the experiments. The method is also solver-, diffusion-schedule-, and NFE-specific: changing these conditions generally requires recomputing the optimized intermediate times.

Coverage note — No substantial contributed material was omitted; the solver-specific Lagrange/Taylor variants, variable-order handling, optimization procedure, theoretical bound, runtime, limitations, and main empirical comparisons are represented.

References

  1. 1.Brian D.O. Anderson. Reverse-time diffusion equation models. Stochastic Processes and their Applications, 12(3):313–326, 1982.
  2. 2.Fan Bao, Chongxuan Li, Jun Zhu, and Bo Zhang. Analytic-DPM: An analytic estimate of the optimal reverse variance in diffusion probabilistic models. In International Conference on Learning Representations, 2022.
  3. 3.Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
  4. 4.Hongrui Chen, Holden Lee, and Jianfeng Lu. Improved analysis of score-based generative modeling: User-friendly bounds under minimal smoothness assumptions. In International Conference on Machine Learning, pages 4735–4763. PMLR, 2023.
  5. 5.Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023.
  6. 6.Minshuo Chen, Kaixuan Huang, Tuo Zhao, and Mengdi Wang. Score approximation, estimation and distribution recovery of diffusion models on low-dimensional data. arXiv preprint arXiv:2302.07194, 2023.
  7. 7.Sitan Chen, Sinho Chewi, Holden Lee, Yuanzhi Li, Jianfeng Lu, and Adil Salim. The probability flow ode is provably fast. arXiv preprint arXiv:2305.11798, 2023.
  8. 8.Sitan Chen, Sinho Chewi, Jerry Li, Yuanzhi Li, Adil Salim, and Anru Zhang. Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions. In International Conference on Learning Representations, 2023.
  9. 9.Valentin De Bortoli. Convergence of denoising diffusion models under the manifold hypothesis. Transactions on Machine Learning Research, 2022.
  10. 10.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255. IEEE, 2009.
  11. 11.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems, pages 8780–8794, 2021.
  12. 12.Yansong Gao, Zhihong Pan, Xin Zhou, Le Kang, and Pratik Chaudhari. Fast diffusion probabilistic model sampling through the lens of backward error analysis. arXiv preprint arXiv:2304.11446, 2023.
  13. 13.Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio. Generative adversarial nets. In Advances in Neural Information Processing Systems, pages 2672–2680, 2014.
  14. 14.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems, pages 6626–6637, 2017.
  15. 15.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, pages 6840–6851, 2020.
  16. 16.Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans. Cascaded diffusion models for high fidelity image generation. Journal of Machine Learning Research, 23(47):1–33, 2022.
  17. 17.Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. In Advances in Neural Information Processing Systems, 2022.
  18. 18.Alexia Jolicoeur-Martineau, Ke Li, Remi Piché-Taillefer, Tal Kachman, and Ioannis Mitliagkas. Gotta go fast when generating data with score-based models. arXiv preprint arXiv:2105.14080, 2021.
  19. 19.Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. In Proc. NeurIPS, 2022.
  20. 20.Dongjun Kim, Seungjae Shin, Kyungwoo Song, Wanmo Kang, and Il-Chul Moon. Soft truncation: A universal training technique of score-based diffusion model for high precision score estimation. In ICML, pages 11201–11228. PMLR, 2022.
  21. 21.Diederik P. Kingma and Max Welling. Auto-encoding variational bayes. In International Conference on Learning Representations, 2014.
  22. 22.Diederik P Kingma, Tim Salimans, Ben Poole, and Jonathan Ho. Variational diffusion models. In Advances in Neural Information Processing Systems, 2021.
  23. 23.Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton. The CIFAR-10 Dataset. online: http://www. cs. toronto. edu/kriz/cifar. html, 55, 2014.
  24. 24.Max WY Lam, Jun Wang, Dan Su, and Dong Yu. Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis. In International Conference on Learning Representations, 2022.
  25. 25.Holden Lee, Jianfeng Lu, and Yixin Tan. Convergence for score-based generative modeling with polynomial complexity. Advances in Neural Information Processing Systems, 35:22870–22882, 2022.
  26. 26.Holden Lee, Jianfeng Lu, and Yixin Tan. Convergence of score-based generative modeling for general data distributions. In International Conference on Algorithmic Learning Theory, pages 946–985. PMLR, 2023.
  27. 27.Lijiang Li, Huixia Li, Xiawu Zheng, Jie Wu, Xuefeng Xiao, Rui Wang, Min Zheng, Xin Pan, Fei Chao, and Rongrong Ji. Autodiffusion: Training-free optimization of time steps and architectures for automated diffusion model acceleration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7105–7114, 2023.
  28. 28.Shigui Li, Wei Chen, and Delu Zeng. Scire-solver: Accelerating diffusion models sampling by score-integrand solver with recursive difference. 2023.
  29. 29.Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13, pages 740–755. Springer, 2014.
  30. 30.Enshu Liu, Xuefei Ning, Zinan Lin, Huazhong Yang, and Yu Wang. Oms-dpm: Optimizing the model schedule for diffusion probabilistic models. arXiv preprint arXiv:2306.08860, 2023.
  31. 31.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan LI, and Jun Zhu. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. In Advances in Neural Information Processing Systems, pages 5775–5787, 2022.
  32. 32.Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models, 2023.
  33. 33.Eric Luhman and Troy Luhman. Knowledge distillation in iterative generative models for improved sampling speed. arXiv preprint arXiv:2101.02388, 2021.
  34. 34.Weijian Luo, Tianyang Hu, Shifeng Zhang, Jiacheng Sun, Zhenguo Li, and Zhihua Zhang. Diff-instruct: A universal approach for transferring knowledge from pre-trained diffusion models. Advances in Neural Information Processing Systems, 36, 2024.
  35. 35.Chenlin Meng, Ruiqi Gao, Diederik P Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans. On distillation of guided diffusion models. In NeurIPS 2022 Workshop on Score-Based Methods, 2022.
  36. 36.Alexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. GLIDE: towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning (ICML), 2022.
  37. 37.Francesco Pedrotti, Jan Maas, and Marco Mondelli. Improved convergence of score-based diffusion models via prediction-correction. arXiv preprint arXiv:2305.14164, 2023.
  38. 38.William Peebles and Saining Xie. Scalable diffusion models with transformers. arXiv preprint arXiv:2212.09748, 2022.
  39. 39.Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents, 2022.
  40. 40.Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, 2022.
  41. 41.Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, Jonathan Ho, David J Fleet, and Mohammad Norouzi. Photorealistic text-to-image diffusion models with deep language understanding. In Advances in Neural Information Processing Systems, 2022.
  42. 42.Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In International Conference on Learning Representations, 2022.
  43. 43.Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pages 2256–2265. PMLR, 2015.
  44. 44.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021.
  45. 45.Kaitao Song, Yichong Leng, Xu Tan, Yicheng Zou, Tao Qin, and Dongsheng Li. Transcormer: Transformer for sentence scoring with sliding language modeling. In Advances in Neural Information Processing Systems, 2022.
  46. 46.Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
  47. 47.Yang Song, Prafulla Dhariwal, Mark Chen, and Ilya Sutskever. Consistency models. arXiv preprint arXiv:2303.01469, 2023.
  48. 48.Yunke Wang, Xiyu Wang, Anh-Dung Dinh, Bo Du, and Charles Xu. Learning to schedule in diffusion probabilistic models. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 2478–2488, 2023.
  49. 49.Zhendong Wang, Huangjie Zheng, Pengcheng He, Weizhu Chen, and Mingyuan Zhou. Diffusion-GAN: Training GANs with diffusion. In The Eleventh International Conference on Learning Representations, 2023.
  50. 50.Daniel Watson, William Chan, Jonathan Ho, and Mohammad Norouzi. Learning fast samplers for diffusion models by differentiating through sample quality. In International Conference on Learning Representations, 2022.
  51. 51.Mengfei Xia, Yujun Shen, Changsong Lei, Yu Zhou, Ran Yi, Deli Zhao, Wenping Wang, and Yong-jin Liu. Towards more accurate diffusion model acceleration with a timestep aligner. arXiv preprint arXiv:2310.09469, 2023.
  52. 52.Zhisheng Xiao, Karsten Kreis, and Arash Vahdat. Tackling the generative learning trilemma with denoising diffusion GANs. In International Conference on Learning Representations, 2022.
  53. 53.Shuchen Xue, Mingyang Yi, Weijian Luo, Shifeng Zhang, Jiacheng Sun, Zhenguo Li, and Zhi-Ming Ma. Sa-solver: Stochastic adams solver for fast sampling of diffusion models, 2023.
  54. 54.Qinsheng Zhang and Yongxin Chen. Fast sampling of diffusion models with exponential integrator. In The Eleventh International Conference on Learning Representations, 2023.
  55. 55.Wenliang Zhao, Lujia Bai, Yongming Rao, Jie Zhou, and Jiwen Lu. Unipc: A unified predictor-corrector framework for fast sampling of diffusion models. arXiv preprint arXiv:2302.04867, 2023.

Citation

MLA
Xue, S., et al. “Accelerating Diffusion Sampling with Optimized Time Steps”. arXiv, 2024, http://arxiv.org/abs/2402.17376v3.
APA
Xue, S., Liu, Z., Chen, F., Zhang, S., Hu, T., Xie, E., & Li, Z. (2024). Accelerating Diffusion Sampling with Optimized Time Steps. arXiv. http://arxiv.org/abs/2402.17376v3
Chicago
Xue, S., Z. Liu, F. Chen, et al. 2024. “Accelerating Diffusion Sampling with Optimized Time Steps”. arXiv. http://arxiv.org/abs/2402.17376v3.
Harvard
Xue, S. et al. (2024) “Accelerating Diffusion Sampling with Optimized Time Steps”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.17376v3.
Vancouver
1. Xue S, Liu Z, Chen F, Zhang S, Hu T, Xie E, Li Z (2024) Accelerating Diffusion Sampling with Optimized Time Steps. arXiv

BibTeX

@article{xue2024accelerating,
  title = {Accelerating Diffusion Sampling with Optimized Time Steps},
  author = {Xue, Shuchen and Liu, Zhaoqiang and Chen, Fei and Zhang, Shifeng and Hu, Tianyang and Xie, Enze and Li, Zhenguo},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.17376v3},
  eprint = {2402.17376}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE