OMS-DPM: Optimizing the Model Schedule for Diffusion Probabilistic Models
Enshu LiuXuefei NingZinan LinHuazhong YangYu Wang
Proposes a predictor-based search framework that dynamically pairs differently sized pre-trained neural networks with specific denoising steps, doubling the sampling speed of models like Stable Diffusion without sacrificing generation quality.
Diffusion probabilistic models represent the state of the art in generative artificial intelligence across domains such as image, speech, and video synthesis. However, their real-world adoption in production and real-time settings remains severely limited by slow generation speeds. Standard sampling requires evaluating deep neural networks over dozens or hundreds of sequential denoising steps, creating massive computational latency and high inference costs. While prior acceleration techniques have focused almost exclusively on mathematical solvers and noise schedules, they universally rely on executing a single neural network architecture across the entire sampling trajectory.
The article introduces and evaluates an overlooked optimization dimension—the "model schedule"—to optimize the trade-off between generation quality and computation speed. The primary objective is to demonstrate that dynamically assigning different pre-trained neural networks to different denoising steps can simultaneously accelerate inference and improve output quality under arbitrary computation budgets, without requiring any model retraining.
To achieve this, the article develops a framework called OMS-DPM (Optimizing the Model Schedule for Diffusion Probabilistic Models). The approach begins by curating a collection of pre-trained networks of varying sizes or training checkpoints. Because the search space of step-by-step model assignments is astronomically large (up to 10^84 combinations), the authors construct an automated predictor model trained on a modest set of schedule-quality data pairs using a ranking loss. This predictor evaluates candidate schedules in less than one second, enabling an evolutionary search algorithm to rapidly discover optimal step-by-step model allocations and solver configurations under strict generation latency budgets. The framework was evaluated across standard benchmark datasets (CIFAR-10, CelebA, ImageNet-64, and LSUN-Church) and on public Stable Diffusion checkpoints for text-to-image synthesis.
The investigation produced several key findings: First, smaller, faster models frequently outperform larger models at specific denoising stages, disproving the assumption that a single globally superior model is optimal at every step. Second, OMS-DPM consistently outperformed state-of-the-art single-model baselines across all datasets, samplers, and compute budgets; for instance, on CIFAR-10, it improved generation quality while running 2.8 times faster than standard baselines. Third, on text-to-image synthesis using Stable Diffusion, OMS-DPM accelerated generation by more than 2x (matching 24-step quality in only 12 steps) simply by combining intermediate checkpoints saved during standard model training. Finally, structural analysis showed distinct scheduling patterns: under tight compute budgets, running smaller models with more steps yields better quality than running large models with few steps, and dataset characteristics dictate whether larger models should be placed near the initial noise stage or the final image stage.
These findings have immediate practical implications for engineering and deployment costs. Organizations can dramatically reduce inference hardware costs, lower energy consumption, and decrease user latency for generative AI applications without retraining expensive base models. Furthermore, the ability to combine historical training checkpoints means practitioners can boost performance using existing, off-the-shelf assets with zero additional training expense.
For future implementation, teams deploying diffusion models should adopt dynamic model scheduling and test OMS-DPM on existing production pipelines, especially where latency constraints are tight. Before applying the framework to new domains, organizations should account for its main limitation: training the search predictor requires generating an initial set of schedule evaluations for each target dataset and task. Overall confidence in the reported improvements is high, supported by consistent empirical gains across multiple datasets, samplers, and architectures, though users should benchmark domain-specific schedules before full rollout.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Read the foundational DDPM paper first to understand the iterative denoising process and pretrained diffusion networks that OMS-DPM schedules.
- Paper: DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps, Cheng Lu et al. (2022). DPM-Solver establishes how solver choice changes the cost and quality of diffusion sampling, a key part of the budgets OMS-DPM optimizes.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). DDIM explains a widely used accelerated sampling trajectory that helps situate OMS-DPM’s scheduling of models across denoising steps.
- Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). Karras et al. disentangle diffusion-model design choices, providing useful grounding in the noise schedules and sampling procedures OMS-DPM configures.
- Paper: Score-Based Generative Modeling through Stochastic Differential Equations, Yang Song et al. (2021). The SDE framework supplies the continuous-time diffusion and probability-flow foundations behind the sampling trajectories OMS-DPM schedules.
No sufficiently relevant recommendations were found.
