OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models

Enshu LiuXuefei NingZinan LinHuazhong YangYu Wang

article2023ICML61 citations

Proposes a predictor-based search framework that dynamically pairs differently sized pre-trained neural networks with specific denoising steps, doubling the sampling speed of models like Stable Diffusion without sacrificing generation quality.

Listen

Diffusion probabilistic models represent the state of the art in generative artificial intelligence across domains such as image, speech, and video synthesis. However, their real-world adoption in production and real-time settings remains severely limited by slow generation speeds. Standard sampling requires evaluating deep neural networks over dozens or hundreds of sequential denoising steps, creating massive computational latency and high inference costs. While prior acceleration techniques have focused almost exclusively on mathematical solvers and noise schedules, they universally rely on executing a single neural network architecture across the entire sampling trajectory.

The article introduces and evaluates an overlooked optimization dimension—the "model schedule"—to optimize the trade-off between generation quality and computation speed. The primary objective is to demonstrate that dynamically assigning different pre-trained neural networks to different denoising steps can simultaneously accelerate inference and improve output quality under arbitrary computation budgets, without requiring any model retraining.

To achieve this, the article develops a framework called OMS-DPM (Optimizing the Model Schedule for Diffusion Probabilistic Models). The approach begins by curating a collection of pre-trained networks of varying sizes or training checkpoints. Because the search space of step-by-step model assignments is astronomically large (up to 10^84 combinations), the authors construct an automated predictor model trained on a modest set of schedule-quality data pairs using a ranking loss. This predictor evaluates candidate schedules in less than one second, enabling an evolutionary search algorithm to rapidly discover optimal step-by-step model allocations and solver configurations under strict generation latency budgets. The framework was evaluated across standard benchmark datasets (CIFAR-10, CelebA, ImageNet-64, and LSUN-Church) and on public Stable Diffusion checkpoints for text-to-image synthesis.

The investigation produced several key findings: First, smaller, faster models frequently outperform larger models at specific denoising stages, disproving the assumption that a single globally superior model is optimal at every step. Second, OMS-DPM consistently outperformed state-of-the-art single-model baselines across all datasets, samplers, and compute budgets; for instance, on CIFAR-10, it improved generation quality while running 2.8 times faster than standard baselines. Third, on text-to-image synthesis using Stable Diffusion, OMS-DPM accelerated generation by more than 2x (matching 24-step quality in only 12 steps) simply by combining intermediate checkpoints saved during standard model training. Finally, structural analysis showed distinct scheduling patterns: under tight compute budgets, running smaller models with more steps yields better quality than running large models with few steps, and dataset characteristics dictate whether larger models should be placed near the initial noise stage or the final image stage.

These findings have immediate practical implications for engineering and deployment costs. Organizations can dramatically reduce inference hardware costs, lower energy consumption, and decrease user latency for generative AI applications without retraining expensive base models. Furthermore, the ability to combine historical training checkpoints means practitioners can boost performance using existing, off-the-shelf assets with zero additional training expense.

For future implementation, teams deploying diffusion models should adopt dynamic model scheduling and test OMS-DPM on existing production pipelines, especially where latency constraints are tight. Before applying the framework to new domains, organizations should account for its main limitation: training the search predictor requires generating an initial set of schedule evaluations for each target dataset and task. Overall confidence in the reported improvements is high, supported by consistent empirical gains across multiple datasets, samplers, and architectures, though users should benchmark domain-specific schedules before full rollout.

No sufficiently relevant recommendations were found.

Cover for OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models

Abstract

Diffusion probabilistic models (DPMs) are a new class of generative models that have achieved state-of-the-art generation quality in various domains. Despite the promise, one major drawback of DPMs is the slow generation speed due to the large number of neural network evaluations required in the generation process. In this paper, we reveal an overlooked dimension—model schedule—for optimizing the trade-off between generation quality and speed. More specifically, we observe that small models, though having worse generation quality when used alone, could outperform large models in certain generation steps. Therefore, unlike the traditional way of using a single model, using different models in different generation steps in a carefully designed model schedule could potentially improve generation quality and speed simultaneously. We design OMS-DPM, a predictor-based search algorithm, to optimize the model schedule given an arbitrary generation time budget and a set of pre-trained models. We demonstrate that OMS-DPM can find model schedules that improve generation quality and speed than prior state-of-the-art methods across CIFAR-10, CelebA, ImageNet, and LSUN datasets. When applied to the public checkpoints of the Stable Diffusion model, we are able to accelerate the sampling by 2× while maintaining the generation quality.

Table of Contents

  • 1. Introduction
  • 2. Background and Related Work
  • 2.1. Diffusion Probabilistic Models
  • 2.2. Training-Free Samplers
  • 2.3. AutoML
  • 2.3.1. PREDICTOR-BASED NEURAL ARCHITECTURE SEARCH
  • 3. Model Schedule: A New Dimension in DPM Design
  • 4. OMS-DPM: Optimizing the Model Schedule
  • 4.1. Problem Definition: Model Schedule Optimization
  • 4.2. Predictor-based Model Schedule Optimization
  • 4.2.1. OVERALL WORKFLOW
  • 4.2.2. PREDICTOR DESIGN
  • 5. Experiments
  • 5.1. The Effectiveness of OMS-DPM
  • 5.2. Results with Stable Diffusion
  • 5.3. Ablation Study
  • 5.4. Empirical Observations
  • 6. Limitations and Future Work
  • Acknowledgements
  • References
  • A. Additional Results
  • A.1. Results of ablation study
  • A.2. Demonstration of searched model schedules
  • B. Model Zoo information
  • B.1. Model Zoo Construction
  • B.2. Inference Latency
  • C. Experiment Details
  • C.1. Evaluation of DPMs
  • C.2. How to Conduct DPM Sampling for A Model Schedule
  • C.3. Schedule-FID Data Generation
  • C.4. Adjustment of the Sequence Predictor for DPM-Solver
  • C.5. Hyperparameter of Predictor
  • C.6. Training
  • C.7. Validating the Reliability of Predictor
  • C.8. Evolutionary Search
  • C.9. Implementation Details and Full Results of Baseline (1)
  • D. Generated Images

Knowls

  1. Knowl 1 — Denoising-step performance can reverse the ranking of models

    empirical result

    On CIFAR-10, the authors found that models with different capacities did not have a consistent ranking in denoising ability across diffusion timesteps. A model that performed worst overall when used alone could outperform most other models during late denoising. In a 90-step DPM-Solver experiment, randomly assigning models from the zoo to different steps produced a schedule with FID approximately 5.4—better sample quality and lower latency than using the best standalone model throughout. This demonstrates why a model schedule can improve the quality–speed trade-off even when smaller models perform worse as complete samplers.

  2. Knowl 2 — Model-schedule optimization under a sampling-time budget

    model/method

    Let a zoo contain NN pretrained noise-prediction models, with model ii having inference latency lil_i. A schedule assigns a model to each used denoising step; for an MM-step sampler its estimated latency is the sum of the assigned models’ latencies. The optimization objective is to minimize the sample-quality score FF (for example, FID) subject to a time budget CC and a maximum of LL steps:

    min⁡M≤L, s1,…,sM, t1<⋯<tMF([(as1,t1),…,(asM,tM)]),∑m=1Mlsm<C.\min_{M\leq L,\,s_1,\ldots,s_M,\,t_1<\cdots<t_M} F\big([(a_{s_1},t_1),\ldots,(a_{s_M},t_M)]\big), \qquad \sum_{m=1}^{M} l_{s_m}<C.

    Here asma_{s_m} is the model assigned to step mm, tmt_m is its diffusion timestep, and the sampler applies the assignments in reverse order, from noise toward data. OMS-DPM fixes LL candidate timesteps and represents a schedule by choices sl′∈{0,1,…,N}s'_l\in\{0,1,\ldots,N\}; choice 00 is a null model that skips that timestep. This encodes both model selection and the effective number and placement of steps, yielding (N+1)L(N+1)^L possible schedules. For DDIM, the candidate times are linearly discretized and non-null entries are used. For DPM-Solver, candidate entries are grouped in threes so a group can represent an inactive step or a first-, second-, or third-order solver; active solver steps are then placed using uniform log-SNR discretization.

  3. Knowl 3 — OMS-DPM trains a schedule-quality predictor before budgeted search

    model/method

    OMS-DPM takes a pretrained model zoo and a sampler, constructs training examples by sampling diverse model schedules, evaluates their sample quality, and trains a predictor to estimate schedule quality. The schedule samples are drawn using multiple manually specified multinomial model-choice distributions so that the training scores cover a range of quality; the predictor is trained once for the given task and sampler, independently of the eventual time budget. The method then searches schedules using predictor scores and the given platform’s model latencies, selecting a low-predicted-FID schedule that satisfies the budget. The authors report that predictor evaluation takes less than one GPU second per schedule. In their unconditional experiments each schedule’s training label used 5,000 generated images; for Stable Diffusion, labels used 1,500 sampled MS-COCO captions.

  4. Knowl 4 — Sequence predictor combines model and timestep information

    model/method

    For a schedule of length LL, OMS-DPM embeds each model choice and each timestep separately. Model choices use a trainable embedding table, including an embedding for the null choice. A multilayer perceptron (MLP) maps a sinusoidal embedding of each scalar timestep to a timestep embedding. The two embeddings at each position are concatenated and passed as a sequence to a one-layer LSTM; the LSTM outputs are averaged over positions and passed through an MLP to produce the predicted quality score P(q)P(q) for schedule qq. For DPM-Solver, each group of three model embeddings is first combined by an MLP into a solver embedding, which is then paired with the timestep embedding before the LSTM.

    The predictor is trained with pairwise ranking rather than direct regression. Given NsN_s training schedules qiq_i with true quality scores F(qi)F(q_i), it uses

    L=∑i=1Ns∑j: F(qj)>F(qi)max⁡(0, m−[P(qj)−P(qi)]),\mathcal{L}=\sum_{i=1}^{N_s}\sum_{j:\,F(q_j)>F(q_i)}\max\bigl(0,\,m-[P(q_j)-P(q_i)]\bigr),

    where lower FF is better and mm is a positive comparison margin. Thus, a worse-FID schedule qjq_j is trained to receive a predicted score at least mm higher than the better schedule qiq_i; lower predicted scores indicate schedules expected to have better quality. In the reported experiments, model-embedding dimensions were 32 (64 for ImageNet-64), timestep dimension 64, LSTM hidden size 128, and the final MLP had four layers of width 200.

  5. Knowl 5 — Predictor-guided evolutionary schedule search

    algorithm

    Given a trained predictor P(q)P(q), a schedule-cost function equal to the sum of assigned-model latencies, and a strict time budget CC, OMS-DPM uses an evolutionary population of schedules. It starts from a one-model schedule with cost in [0.9C,C][0.9C,C]. At each epoch it samples up to 10 candidate parents from the population, chooses the parent with the lowest predicted score, and repeatedly generates random mutations of that parent. Mutations costing at least CC are rejected. Accepted schedules form a next-generation set; they are added to the population, and schedules with the worst predicted scores are removed if the population exceeds its cap. After the search, the lowest-predicted-score schedule is returned.

    Input: Predictor P, schedule-cost function Cost, time budget C
    Parameters: epochs = 600, candidate-parent limit = 10,
                mutation-attempt limit = 200, next-generation cap = 40,
                population cap = 40
    Initialize population with a one-model schedule q0
      whose cost is in [0.9C, C]
    for each epoch:
        Sample up to 10 schedules from the population
        Choose the sampled schedule q with the lowest P(q)
        Set next generation NG to empty
        Repeat at most 200 times, stopping if NG has 40 schedules:
            Randomly mutate q to obtain qnew
            If Cost(qnew) < C, add qnew to NG
        Add NG to the population
        If the population exceeds 40 schedules:
            Remove schedules with the highest predicted scores
    Return the schedule in the population with the lowest P(q)

    The authors used 600 epochs, although they report that search commonly reached a local optimum after about 200–300 epochs. The paper specifies random mutation but does not detail the mutation operator, so that operator is not recoverable from the reported algorithm alone.

  6. Knowl 6 — OMS-DPM improves FID across datasets, budgets, and samplers

    empirical result

    The unconditional evaluations used CIFAR-10, CelebA, ImageNet-64, and LSUN-Church with both DPM-Solver and DDIM. The paper reports that OMS-DPM beat the best single-model baseline at every tested budget for these datasets and samplers; FID is the quality metric, with lower values better. Examples from DPM-Solver are CIFAR-10 at a 700 ms budget, where OMS-DPM achieved 6.08±0.006.08\pm0.00 versus 8.738.73 for the best single-model baseline; ImageNet-64 at 800 ms, 23.94±0.0023.94\pm0.00 versus 29.5929.59; and LSUN-Church at 4,000 ms, 13.94±0.0013.94\pm0.00 versus 32.2332.23. With DDIM, OMS-DPM achieved FID 12.34±0.3112.34\pm0.31 versus 16.1116.11 on CIFAR-10 at 750 ms, and 23.03±0.6723.03\pm0.67 versus 25.2725.27 on LSUN-Church at 4,000 ms. Search results are reported as mean and standard deviation over three random seeds. The comparisons support the claim that mixing models can help at both tight budgets and larger budgets, and that the method is compatible with both tested samplers.

  7. Knowl 7 — Stable Diffusion reaches better FID with half as many steps

    empirical result

    On text-to-image generation with four public Stable Diffusion checkpoints and DPM-Solver, OMS-DPM achieved FID 11.3411.34 with 12 NFEs, compared with FID 11.8111.81 for the best single-checkpoint schedule with 24 NFEs. Because all four checkpoints have the same architecture and latency, halving the NFEs approximately halves sampling time. The evaluation used 30,000 MS-COCO 256×256 validation captions and 30,000 corresponding generated images, with guidance scale 1.5; real and generated images were compared using the same captions. At other tested budgets, OMS-DPM’s FIDs were 12.90, 10.72, 10.68, and 10.57 at 9, 15, 18, and 24 NFEs, respectively.

  8. Knowl 8 — Ablations show effects of model-zoo and predictor-data size

    empirical result

    On CIFAR-10 with DPM-Solver, the authors varied the model-zoo size among 2, 4, and 6 models and the predictor training set among 915, 1,831, and 3,662 schedule–FID examples. A six-model zoo often gave the strongest searched FID, but a smaller zoo could be competitive: at a 700 ms budget, searched FIDs were 6.05±0.006.05\pm0.00, 6.25±0.366.25\pm0.36, and 6.08±0.006.08\pm0.00 for zoo sizes 2, 4, and 6. At a 4,000 ms budget, the corresponding values were 3.36±0.023.36\pm0.02, 3.42±0.003.42\pm0.00, and 3.14±0.023.14\pm0.02. Using fewer predictor-training examples generally weakened search results, though they remained promising in most tested settings; for example, at 1,400 ms, FID improved from 3.97±0.023.97\pm0.02 with 915 examples to 3.75±0.143.75\pm0.14 with 1,831 and 3.48±0.063.48\pm0.06 with 3,662. The ablation indicates that both the available model choices and the amount of schedule-quality data affect performance.

  9. Knowl 9 — Useful schedule patterns depend on dataset and budget

    empirical result

    The schedules found by OMS-DPM exhibit dataset-specific patterns rather than a single universal rule. Under tight budgets, the authors observed schedules relying on the lower-latency models; they suggest that adding steps with smaller models can be preferable to using fewer steps with larger models when solver and discretization errors rise quickly at low NFE. For ImageNet-64 with either tested sampler and CIFAR-10 with DDIM, larger models were more often assigned near the final generated data; LSUN-Church showed the opposite tendency, favoring larger models nearer the noise. In DDIM schedules, step sizes were often larger near the final image for CIFAR-10, CelebA, and ImageNet-64, while LSUN-Church schedules often had smaller steps at both ends than in the middle. DPM-Solver schedules tended to use first- or second-order solvers at tight budgets and reserve third-order solvers for larger budgets. These are observed tendencies, not guaranteed rules for other tasks.

  10. Knowl 10 — Schedule predictors must be prepared for each new task

    limitation

    The reported method does not directly transfer its schedule-quality predictor to a new dataset or downstream task: the authors state that new schedule-evaluation data must be prepared and a new predictor trained, which incurs substantial overhead. They also find that model-zoo size and quality influence results, while leaving efficient construction of a diverse model zoo as an open problem. The paper suggests, but does not demonstrate, approaches such as pruning pretrained models or changing training-loss weights to create such a zoo.

Coverage note — Detailed U-Net configurations, training recipes, full baseline tables, and generated-image examples are omitted because they support the evaluations but are not separate load-bearing contributions.

References

  1. 1.Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Kreis, K., Aittala, M., Aila, T., Laine, S., Catanzaro, B., et al. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022.
  2. 2.Bao, F., Li, C., Zhu, J., and Zhang, B. Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models. arXiv preprint arXiv:2201.06503, 2022.
  3. 3.Chen, N., Zhang, Y., Zen, H., Weiss, R. J., Norouzi, M., and Chan, W. Wavegrad: Estimating gradients for waveform generation. arXiv preprint arXiv:2009.00713, 2020.
  4. 4.Choi, J., Lee, J., Shin, C., Kim, S., Kim, H., and Yoon, S. Perception prioritized training of diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11472–11481, 2022.
  5. 5.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021.
  6. 6.Elsken, T., Metzen, J. H., and Hutter, F. Neural architecture search: A survey. The Journal of Machine Learning Research, 20(1):1997–2017, 2019.
  7. 7.Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020.
  8. 8.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33:6840–6851, 2020.
  9. 9.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022.
  10. 10.Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al. Population based training of neural networks. arXiv preprint arXiv:1711.09846, 2017.
  11. 11.Jing, B., Corso, G., Berlinghieri, R., and Jaakkola, T. Subspace diffusion generative models. arXiv preprint arXiv:2205.01490, 2022.
  12. 12.Kim, D., Shin, S., Song, K., Kang, W., and Moon, I.-C. Soft truncation: A universal training technique of score-based diffusion model for high precision score estimation. arXiv preprint arXiv:2106.05527, 2021.
  13. 13.Kingma, D. P. and Gao, R. Understanding the diffusion objective as a weighted integral of elbos. arXiv preprint arXiv:2303.00848, 2023.
  14. 14.Kingma, D. P. and Welling, M. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013.
  15. 15.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. arXiv preprint arXiv:2009.09761, 2020.
  16. 16.Li, H., Yang, Y., Chang, M., Chen, S., Feng, H., Xu, Z., Li, Q., and Chen, Y. Srdiff: Single image super-resolution with diffusion probabilistic models. Neurocomputing, 479:47–59, 2022.
  17. 17.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740–755. Springer, 2014.
  18. 18.Liu, L., Ren, Y., Lin, Z., and Zhao, Z. Pseudo numerical methods for diffusion models on manifolds. arXiv preprint arXiv:2202.09778, 2022.
  19. 19.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022.
  20. 20.Luo, R., Tian, F., Qin, T., Chen, E., and Liu, T.-Y. Neural architecture optimization. Advances in neural information processing systems, 31, 2018.
  21. 21.Luo, S. and Hu, W. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2837–2845, 2021.
  22. 22.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In International Conference on Machine Learning, pp. 8162–8171. PMLR, 2021.
  23. 23.Ning, X., Zheng, Y., Zhao, T., Wang, Y., and Yang, H. A generic graph-based neural architecture encoding scheme for predictor-based nas. In European Conference on Computer Vision, pp. 189–204. Springer, 2020.
  24. 24.Real, E., Aggarwal, A., Huang, Y., and Le, Q. V. Regularized evolution for image classifier architecture search. In Proceedings of the aaai conference on artificial intelligence, volume 33, pp. 4780–4789, 2019.
  25. 25.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  26. 26.Ronneberger, O., Fischer, P., and Brox, T. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pp. 234–241. Springer, 2015.
  27. 27.Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022.
  28. 28.Sen, P. K. Estimates of the regression coefficient based on kendall’s tau. Journal of the American statistical association, 63(324):1379–1389, 1968.
  29. 29.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning, pp. 2256–2265. PMLR, 2015.
  30. 30.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502, 2020a.
  31. 31.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. Advances in Neural Information Processing Systems, 32, 2019.
  32. 32.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456, 2020b.
  33. 33.Song, Y., Durkan, C., Murray, I., and Ermon, S. Maximum likelihood training of score-based diffusion models. Advances in Neural Information Processing Systems, 34:1415–1428, 2021.
  34. 34.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems, 30, 2017.
  35. 35.Watson, D., Ho, J., Norouzi, M., and Chan, W. Learning to efficiently sample from diffusion probabilistic models. arXiv preprint arXiv:2106.03802, 2021.
  36. 36.Watson, D., Chan, W., Ho, J., and Norouzi, M. Learning fast samplers for diffusion models by differentiating through sample quality. In International Conference on Learning Representations, 2022.
  37. 37.Yang, C., Akimoto, Y., Kim, D. W., and Udell, M. Oboe: Collaborative filtering for automl model selection. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining, pp. 1173–1183, 2019.
  38. 38.Yang, X., Zhou, D., Feng, J., and Wang, X. Diffusion probabilistic model made slim. arXiv preprint arXiv:2211.17106, 2022.
  39. 39.Zhang, Q. and Chen, Y. Fast sampling of diffusion models with exponential integrator. arXiv preprint arXiv:2204.13902, 2022.
  40. 40.Zoph, B. and Le, Q. V. Neural architecture search with reinforcement learning. arXiv preprint arXiv:1611.01578, 2016.

Citation

MLA
Liu, E., et al. “OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models”. International Conference on Machine Learning, vol. 202, 2023, pp. 21915–36, https://proceedings.mlr.press/v202/liu23ab.html.
APA
Liu, E., Ning, X., Lin, Z., Yang, H., & Wang, Y. (2023). OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models. International Conference on Machine Learning, 202, 21915–21936. https://proceedings.mlr.press/v202/liu23ab.html
Chicago
Liu, E., X. Ning, Z. Lin, H. Yang, and Y. Wang. 2023. “OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models”. International Conference on Machine Learning 202: 21915–36. https://proceedings.mlr.press/v202/liu23ab.html.
Harvard
Liu, E. et al. (2023) “OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models”, International Conference on Machine Learning. PMLR, pp. 21915–21936. Available at: https://proceedings.mlr.press/v202/liu23ab.html.
Vancouver
1. Liu E, Ning X, Lin Z, Yang H, Wang Y (2023) OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models. In: International Conference on Machine Learning. PMLR, pp 21915–21936

BibTeX

@InProceedings{pmlr-v202-liu23ab,
  title = 	 {{OMS}-{DPM}: Optimizing the Model Schedule for Diffusion Probabilistic Models},
  author =       {Liu, Enshu and Ning, Xuefei and Lin, Zinan and Yang, Huazhong and Wang, Yu},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {21915--21936},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/liu23ab/liu23ab.pdf},
  url = 	 {https://proceedings.mlr.press/v202/liu23ab.html},
  abstract = 	 {Diffusion probabilistic models (DPMs) are a new class of generative models that have achieved state-of-the-art generation quality in various domains. Despite the promise, one major drawback of DPMs is the slow generation speed due to the large number of neural network evaluations required in the generation process. In this paper, we reveal an overlooked dimension—model schedule—for optimizing the trade-off between generation quality and speed. More specifically, we observe that small models, though having worse generation quality when used alone, could outperform large models in certain generation steps. Therefore, unlike the traditional way of using a single model, using different models in different generation steps in a carefully designed model schedule could potentially improve generation quality and speed simultaneously. We design OMS-DPM, a predictor-based search algorithm, to determine the optimal model schedule given an arbitrary generation time budget and a set of pre-trained models. We demonstrate that OMS-DPM can find model schedules that improve generation quality and speed than prior state-of-the-art methods across CIFAR-10, CelebA, ImageNet, and LSUN datasets. When applied to the public checkpoints of the Stable Diffusion model, we are able to accelerate the sampling by 2x while maintaining the generation quality.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/