Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation
Jiaming SongQinsheng ZhangHongxu YinMorteza MardaniMing-Yu LiuJan KautzYongxin ChenArash Vahdat
Proposes a Monte Carlo approximation technique for plug-and-play loss guidance in pretrained diffusion models, correcting guidance scale estimation errors to achieve precise controllable generation in image synthesis and obstacle-avoiding motion planning without extra training.
Generative artificial intelligence models, particularly diffusion models, excel at creating high-quality data such as images, audio, and motion sequences. However, steering these models to satisfy specific constraints or user-defined goals typically requires retraining them on massive paired datasets or relying on narrow, hand-crafted mathematical techniques. Existing plug-and-play methods that attempt to guide pretrained models using arbitrary differentiable loss functions rely heavily on point estimates, which introduce substantial mathematical approximation errors and often fail to produce realistic, constraint-compliant outputs.
The article evaluates the Loss-Guided Diffusion framework and demonstrates a new Monte Carlo sampling technique, termed LGD-MC, designed to achieve plug-and-play controllable generation without retraining. Its primary objective is to accurately approximate the guidance term across all diffusion steps using general loss functions while preserving computational efficiency.
The authors conducted a series of comparative experiments evaluating LGD-MC against standard baselines and existing guidance techniques, such as Diffusion Posterior Sampling and Classifier Guidance. The evaluation spanned multiple domains, including synthetic probability distributions, 64x64 to 256x256 image super-resolution on ImageNet validation data, conditional text- and label-driven image synthesis, and 3D human motion synthesis under path-following and obstacle-avoidance constraints.
The findings show that standard point-estimate approaches severely miscalculate the guidance scale across noise levels, whereas LGD-MC drastically reduces estimation bias. In synthetic benchmarks, LGD-MC achieved a tenfold reduction in approximation error compared to point-estimate methods. In image super-resolution over 100 steps, LGD-MC improved image quality metrics from an FID of 74.11 down to 5.21 and increased ResNet-50 classification accuracy from 25.84% to over 72%. In controllable motion synthesis, LGD-MC enabled pretrained models to navigate around complex obstacles and follow precise trajectories—cutting collision and objective penalty metrics by roughly half compared to prior methods—while adding only negligible computational overhead of around 8% longer wall-clock time per iteration.
These results indicate that organizations can enforce fine-grained operational constraints on general-purpose pretrained models without costly model retraining or extensive hyperparameter retuning. By enabling plug-and-play control with minimal computing overhead, this approach lowers deployment costs and accelerates timelines for specialized generative applications, from robotics path planning to creative asset generation.
Teams deploying diffusion models should consider adopting multi-sample Monte Carlo guidance when adapting off-the-shelf foundation models to constrained generation tasks. Practitioners should weigh sample count trade-offs, as drawing 10 to 100 samples yields notable accuracy gains with marginal resource increases, provided the guiding loss function is cheap to compute.
A key limitation is that LGD-MC assumes a simplified Gaussian distribution during intermediate sampling steps, which remains an approximation of the true data state. Additionally, performance degrades if target conditions fall outside the underlying model's initial training distribution. Confidence in the reported gains is high for the tested domains, though further testing is recommended before applying the method to complex multimodal distributions outside standard image and motion benchmarks.
- Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). Introduces classifier guidance via explicit gradients to steer diffusion sampling, establishing the foundational guided-sampling paradigm that Loss-Guided Diffusion generalizes.
- Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). Pioneers point-estimate guidance via Tweedie's formula for inverse problems, providing the exact approximation baseline whose errors and bias Loss-Guided Diffusion explicitly analyzes and overcomes using Monte Carlo sampling.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the standard denoising diffusion probabilistic model formulation and reverse-time sampling equations utilized across the source paper.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Formulates non-Markovian deterministic sampling (DDIM), which provides the core accelerated sampling infrastructure leveraged in plug-and-play guided diffusion.
- Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). Presents classifier-free guidance, a fundamental baseline and alternative formulation for conditioning diffusion models without auxiliary classifiers.
- Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Applies diffusion models to trajectory planning and behavioral synthesis with guided perturbation functions, directly motivating the motion synthesis and obstacle avoidance benchmarks in the source paper.
- Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Introduces the human motion diffusion framework that the source paper adopts as a pretrained backbone for path-following and obstacle-avoidance experiments.
- Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). Presents an unsupervised diffusion restoration method for linear inverse problems, offering essential background on zero-shot conditioning with pretrained diffusion models.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). Develops SDE-based stroke and patch guidance for image editing without retraining, highlighting the advantages and limits of prior heuristic guidance.
- Paper: DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps, Cheng Lu et al. (2022). Derives high-order fast ODE solvers for diffusion sampling that can be integrated with training-free plug-and-play guidance mechanisms.
- Paper: Lumiere: A Space-Time Diffusion Model for Video Generation, Omer Bar-Tal et al. (2024). Extends space-time diffusion architectures to video generation and video editing, offering a richer temporal generative domain for plug-and-play loss guidance.
- Paper: A Mathematical Introduction to Diffusion Models, Jianfeng Lu (2026). Formalizes a comprehensive mathematical framework and discretization error bounds for sampling and inference-time guidance in diffusion models.
