Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation

Jiaming SongQinsheng ZhangHongxu YinMorteza MardaniMing-Yu LiuJan KautzYongxin ChenArash Vahdat

article2023ICML220 citations

Proposes a Monte Carlo approximation technique for plug-and-play loss guidance in pretrained diffusion models, correcting guidance scale estimation errors to achieve precise controllable generation in image synthesis and obstacle-avoiding motion planning without extra training.

Listen

Generative artificial intelligence models, particularly diffusion models, excel at creating high-quality data such as images, audio, and motion sequences. However, steering these models to satisfy specific constraints or user-defined goals typically requires retraining them on massive paired datasets or relying on narrow, hand-crafted mathematical techniques. Existing plug-and-play methods that attempt to guide pretrained models using arbitrary differentiable loss functions rely heavily on point estimates, which introduce substantial mathematical approximation errors and often fail to produce realistic, constraint-compliant outputs.

The article evaluates the Loss-Guided Diffusion framework and demonstrates a new Monte Carlo sampling technique, termed LGD-MC, designed to achieve plug-and-play controllable generation without retraining. Its primary objective is to accurately approximate the guidance term across all diffusion steps using general loss functions while preserving computational efficiency.

The authors conducted a series of comparative experiments evaluating LGD-MC against standard baselines and existing guidance techniques, such as Diffusion Posterior Sampling and Classifier Guidance. The evaluation spanned multiple domains, including synthetic probability distributions, 64x64 to 256x256 image super-resolution on ImageNet validation data, conditional text- and label-driven image synthesis, and 3D human motion synthesis under path-following and obstacle-avoidance constraints.

The findings show that standard point-estimate approaches severely miscalculate the guidance scale across noise levels, whereas LGD-MC drastically reduces estimation bias. In synthetic benchmarks, LGD-MC achieved a tenfold reduction in approximation error compared to point-estimate methods. In image super-resolution over 100 steps, LGD-MC improved image quality metrics from an FID of 74.11 down to 5.21 and increased ResNet-50 classification accuracy from 25.84% to over 72%. In controllable motion synthesis, LGD-MC enabled pretrained models to navigate around complex obstacles and follow precise trajectories—cutting collision and objective penalty metrics by roughly half compared to prior methods—while adding only negligible computational overhead of around 8% longer wall-clock time per iteration.

These results indicate that organizations can enforce fine-grained operational constraints on general-purpose pretrained models without costly model retraining or extensive hyperparameter retuning. By enabling plug-and-play control with minimal computing overhead, this approach lowers deployment costs and accelerates timelines for specialized generative applications, from robotics path planning to creative asset generation.

Teams deploying diffusion models should consider adopting multi-sample Monte Carlo guidance when adapting off-the-shelf foundation models to constrained generation tasks. Practitioners should weigh sample count trade-offs, as drawing 10 to 100 samples yields notable accuracy gains with marginal resource increases, provided the guiding loss function is cheap to compute.

A key limitation is that LGD-MC assumes a simplified Gaussian distribution during intermediate sampling steps, which remains an approximation of the true data state. Additionally, performance degrades if target conditions fall outside the underlying model's initial training distribution. Confidence in the reported gains is high for the tested domains, though further testing is recommended before applying the method to complex multimodal distributions outside standard image and motion benchmarks.

Song et al (2023).pdf
  • Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). Introduces classifier guidance via explicit gradients to steer diffusion sampling, establishing the foundational guided-sampling paradigm that Loss-Guided Diffusion generalizes.
  • Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). Pioneers point-estimate guidance via Tweedie's formula for inverse problems, providing the exact approximation baseline whose errors and bias Loss-Guided Diffusion explicitly analyzes and overcomes using Monte Carlo sampling.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the standard denoising diffusion probabilistic model formulation and reverse-time sampling equations utilized across the source paper.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Formulates non-Markovian deterministic sampling (DDIM), which provides the core accelerated sampling infrastructure leveraged in plug-and-play guided diffusion.
  • Paper: Classifier-Free Diffusion Guidance, Jonathan Ho et al. (2022). Presents classifier-free guidance, a fundamental baseline and alternative formulation for conditioning diffusion models without auxiliary classifiers.
  • Paper: Planning with Diffusion for Flexible Behavior Synthesis, Michael Janner et al. (2022). Applies diffusion models to trajectory planning and behavioral synthesis with guided perturbation functions, directly motivating the motion synthesis and obstacle avoidance benchmarks in the source paper.
  • Paper: Human Motion Diffusion Model, Guy Tevet et al. (2022). Introduces the human motion diffusion framework that the source paper adopts as a pretrained backbone for path-following and obstacle-avoidance experiments.
  • Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). Presents an unsupervised diffusion restoration method for linear inverse problems, offering essential background on zero-shot conditioning with pretrained diffusion models.
  • Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). Develops SDE-based stroke and patch guidance for image editing without retraining, highlighting the advantages and limits of prior heuristic guidance.
  • Paper: DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps, Cheng Lu et al. (2022). Derives high-order fast ODE solvers for diffusion sampling that can be integrated with training-free plug-and-play guidance mechanisms.
Cover for Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation

Abstract

We consider guiding denoising diffusion models with general differentiable loss functions in a plug-and-play fashion, enabling controllable generation without additional training. This paradigm, termed Loss-Guided Diffusion (LGD), can easily be integrated into all diffusion models and leverage various efficient samplers. Despite the benefits, the resulting guidance term is, unfortunately, an intractable integral and needs to be approximated. Existing methods compute the guidance term based on a point estimate. However, we show that such approaches have significant errors over the scale of the approximations. To address this issue, we propose a Monte Carlo method that uses multiple samples from a suitable distribution to reduce bias. Our method is effective in various synthetic and real-world settings, including image super-resolution, text or label-conditional image generation, and controllable motion synthesis. Notably, we show how our method can be applied to control a pretrained motion diffusion model to follow certain paths and avoid obstacles that are proven challenging to prior methods.

Table of Contents

  • 1. Introduction
  • 2. Background
  • 2.1. Diffusion Models
  • 2.2. Guidance Methods for Conditional Generation
  • 3. Loss-Guided Diffusion
  • 4. Practical Algorithms
  • 4.1. Existing Approximations with DPS
  • 4.2. DPS Miscalculates the Scale of Loss Guidance
  • 4.3. Improving the Estimation with Monte Carlo
  • 5. Related Works
  • 6. Experiments
  • 6.1. Mixture of Gaussian
  • 6.2. Image Super-resolution
  • 6.3. Controllable Motion Synthesis for Human Motion Diffusion Models
  • 6.4. Sampling Conditional Images from Unconditional Diffusion Models
  • 7. Conclusions
  • References
  • A. Additional method details
  • A.1. Preliminaries on Diffusion Models
  • A.2. Comparing Different Paradigms for Conditional Generation
  • A.3. Proof of Proposition 4.1
  • A.4. Ablation on the Number of Monte Carlo Samples
  • A.5. Computational Efficiency
  • B. Additional Experimental Details
  • B.1. Mixture of Gaussian
  • B.2. Image Super-resolution
  • B.3. Controllable Motion Synthesis
  • B.3.1. VALIDITY OF THE EMBEDDING METRIC
  • B.3.2. OBSTACLE AVOIDANCE
  • B.4. Conditional Image Generation

Knowls

  1. Knowl 1 — Loss-Guided Diffusion Formulation

    definition

    Loss-Guided Diffusion (LGD) is a framework for plug-and-play conditional generation using a pre-trained, unconditional or generic diffusion model DθD_\theta and a differentiable loss function ℓy:X→R\ell_y: \mathcal{X} \to \mathbb{R} conditioned on arbitrary target constraints yy. The goal is to sample clean data points x0∈Xx_0 \in \mathcal{X} from the posterior distribution:

    p0(ℓ)(x0∣y)=1Zp0(x0)exp⁡(−ℓy(x0))p_0^{(\ell)}(x_0|y) = \frac{1}{Z} p_0(x_0) \exp(-\ell_y(x_0))

    where p0(x0)p_0(x_0) is the unguided data distribution and Z=∫Xp0(x0)exp⁡(−ℓy(x0))dx0Z = \int_{\mathcal{X}} p_0(x_0) \exp(-\ell_y(x_0)) dx_0 is the normalizing constant. For a forward noising process xt=x0+σtϵx_t = x_0 + \sigma_t \epsilon with ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) and monotonically increasing noise schedule σt\sigma_t for t∈(0,T]t \in (0, T], the corresponding conditional diffusion score requires the loss guidance term:

    ∇xtlog⁡pt(ℓ)(y∣xt)=∇xtlog⁡∫Xp(x0∣xt)exp⁡(−ℓy(x0))dx0\nabla_{x_t} \log p_t^{(\ell)}(y|x_t) = \nabla_{x_t} \log \int_{\mathcal{X}} p(x_0|x_t) \exp(-\ell_y(x_0)) dx_0

    Evaluating this integral is intractable in the general case because the transition posterior p(x0∣xt)p(x_0|x_t) requires iterative diffusion evaluation to compute or sample from exactly.

  2. Knowl 2 — Monte Carlo Loss-Guided Diffusion Estimator (LGD-MC)

    model/method

    To approximate the intractable loss guidance term ∇xtlog⁡pt(ℓ)(y∣xt)\nabla_{x_t} \log p_t^{(\ell)}(y|x_t) while querying the diffusion network DθD_\theta only once per timestep, LGD-MC replaces the intractable conditional distribution p(x0∣xt)p(x_0|x_t) with an isotropic Gaussian variational surrogate:

    q(x0∣xt)=N(x^t,rt2I)q(x_0|x_t) = \mathcal{N}(\hat{x}_t, r_t^2 I)

    where x^t=Dθ(xt,t)=xt+σt2∇xtlog⁡pt(xt)\hat{x}_t = D_\theta(x_t, t) = x_t + \sigma_t^2 \nabla_{x_t} \log p_t(x_t) is the minimum mean squared error (MMSE) Tweedie estimate of x0x_0, and the variance parameter is set to rt=σt1+σt2r_t = \frac{\sigma_t}{\sqrt{1 + \sigma_t^2}}.

    The loss guidance score is estimated via Monte Carlo integration over nn independent and identically distributed samples x(i)∼q(x0∣xt)x^{(i)} \sim q(x_0|x_t):

    MCn(xt,y)=∇xtlog⁡(1n∑i=1nexp⁡(−ℓy(x(i))))\text{MC}_n(x_t, y) = \nabla_{x_t} \log \left( \frac{1}{n} \sum_{i=1}^n \exp(-\ell_y(x^{(i)})) \right)

    Because each sample is parameterized as x(i)=x^t+rtϵix^{(i)} = \hat{x}_t + r_t \epsilon_i with ϵi∼N(0,I)\epsilon_i \sim \mathcal{N}(0, I), backpropagation through MCn(xt,y)\text{MC}_n(x_t, y) computes gradients with respect to x^t\hat{x}_t and passes through the diffusion denoiser DθD_\theta only once, keeping the asymptotic cost per diffusion step dominated by a single neural network evaluation.

  3. Knowl 3 — Guidance Scale Miscalculation in Diffusion Posterior Sampling

    model/method

    Diffusion Posterior Sampling (DPS) approximates the loss guidance integral ∇xtlog⁡pt(ℓ)(y∣xt)\nabla_{x_t} \log p_t^{(\ell)}(y|x_t) by replacing the distribution p(x0∣xt)p(x_0|x_t) with a Dirac delta point estimate centered at the MMSE estimate x^t=Dθ(xt,t)\hat{x}_t = D_\theta(x_t, t):

    DPS(xt,y)=−∇xtℓy(x^t)\text{DPS}(x_t, y) = -\nabla_{x_t} \ell_y(\hat{x}_t)

    In multimodal distributions (such as mixtures of separated components), this point estimate produces systematic scale errors across noise levels σt\sigma_t:

    1. At high noise levels (large σt\sigma_t), the MMSE estimate x^t=E[x0∣xt]\hat{x}_t = \mathbb{E}[x_0|x_t] shrinks toward the center of the prior distribution where decision boundaries lie. In these regions, loss gradients ∇x0ℓy(x0)\nabla_{x_0} \ell_y(x_0) are large, causing DPS to significantly overestimate the scale of the true guidance term.
    2. At low noise levels (small σt\sigma_t), x^t\hat{x}_t lands close to component modes where loss gradients flatten out, causing DPS to underestimate the required guidance magnitude.

    In contrast, drawing samples across a non-degenerate distribution q(x0∣xt)q(x_0|x_t) averages gradients across high-probability regions, mitigating point-estimate scale distortion.

  4. Knowl 4 — Total Variation Bias Bound for Loss Guidance Approximation

    theoretical result

    Let the likelihood factor p0(ℓ)(y∣x0)=exp⁡(−ℓy(x0))p_0^{(\ell)}(y|x_0) = \exp(-\ell_y(x_0)) be bounded for all x0∈Xx_0 \in \mathcal{X} and conditions yy, such that M=max⁡x0p0(ℓ)(y∣x0)<∞M = \max_{x_0} p_0^{(\ell)}(y|x_0) < \infty. For any t∈(0,T]t \in (0, T], let p(x0∣xt)p(x_0|x_t) and q(x0∣xt)q(x_0|x_t) be probability density functions supported on X\mathcal{X}.

    For all xt∈Xx_t \in \mathcal{X}, the difference between the expectations of the likelihood factor under the true posterior and the approximate distribution is bounded by the total variation distance TV(p,q)=12∫X∣p(x0∣xt)−q(x0∣xt)∣dx0\text{TV}(p, q) = \frac{1}{2} \int_\mathcal{X} |p(x_0|x_t) - q(x_0|x_t)| dx_0:

    ∣Ep(x0∣xt)[p0(ℓ)(y∣x0)]−Eq(x0∣xt)[p0(ℓ)(y∣x0)]∣≤2M⋅TV(p(x0∣xt),q(x0∣xt))\left| \mathbb{E}_{p(x_0|x_t)}[p_0^{(\ell)}(y|x_0)] - \mathbb{E}_{q(x_0|x_t)}[p_0^{(\ell)}(y|x_0)] \right| \le 2M \cdot \text{TV}(p(x_0|x_t), q(x_0|x_t))

    Choosing a continuous distribution q(x0∣xt)=N(x^t,rt2I)q(x_0|x_t) = \mathcal{N}(\hat{x}_t, r_t^2 I) yields a smaller total variation distance to p(x0∣xt)p(x_0|x_t) than choosing a Dirac delta point mass δ(x0−x^t)\delta(x_0 - \hat{x}_t), thereby reducing the asymptotic expectation bias.

  5. Knowl 5 — Loss Formulations for Controllable Human Motion Synthesis

    model/method

    For text-guided human motion diffusion models that generate clean motion sequences x0x_0, controllable trajectory generation is formulated through differentiable root motion losses on root(x0)={x0(i)}\text{root}(x_0) = \{x_0^{(i)}\}, the sequence of 3D global root joint coordinates across time frames ii:

    1. Path following loss: To guide the avatar along a prescribed 3D coordinate trajectory yy, the loss is: ℓy(1)(x0)=∥y−root(x0)∥22\ell_y^{(1)}(x_0) = \|y - \text{root}(x_0)\|_2^2

    2. Obstacle avoidance loss: To direct the avatar to reach destination yty_t while evading an obstacle located at yobsy_{\text{obs}}, the loss is: ℓy(2)(x0)=∥yt−target(root(x0))∥22+∑isigmoid(−(∥root(x0(i))−yobs∥22−1.0)×50)×100\ell_y^{(2)}(x_0) = \|y_t - \text{target}(\text{root}(x_0))\|_2^2 + \sum_i \text{sigmoid}\left(-(\left\|\text{root}(x_0^{(i)}) - y_{\text{obs}}\right\|_2^2 - 1.0) \times 50\right) \times 100 where target(⋅)\text{target}(\cdot) computes the avatar's final frame location, and the sigmoid obstacle penalty transitions sharply to 100 whenever the avatar's root position comes within Euclidean distance 1.0 of yobsy_{\text{obs}}.

  6. Knowl 6 — Benchmark on Bicubic Image Super-Resolution

    data/table

    The table compares plug-and-play conditional generation methods for 64×64→256×25664 \times 64 \to 256 \times 256 bicubic image super-resolution on 50,000 ImageNet validation images using 100 sampling steps with an unconditional diffusion model. The loss function is ℓy(x0)=∥y−Hx0∥222st2\ell_y(x_0) = \frac{\|y - Hx_0\|_2^2}{2 s_t^2} with st=0.25s_t = 0.25, where HH is the bicubic downsampling operator.

    Method FID ↓\downarrow ResNet50 Accuracy ↑\uparrow # Steps
    DDRM 21.3 63.2% 100
    Π\PiGDM 3.6 72.1% 100
    DPS 74.11 25.84% 100
    LGD-MC (n=1n = 1) 5.42 71.98% 100
    LGD-MC (n=10n = 10) 5.32 72.07% 100
    LGD-MC (n=100n = 100) 5.21 72.03% 100

    While Π\PiGDM leverages specialized linear closed-form integrals, LGD-MC operates on generic differentiable losses and significantly outperforms DPS (which fails to produce accurate samples at 100 steps, achieving 74.11 FID and 25.84% accuracy) and DDRM. Increasing the Monte Carlo sample count nn from 1 to 100 monotonically improves FID from 5.42 to 5.21.

  7. Knowl 7 — Benchmark on Controllable Human Motion Synthesis

    data/table

    Quantitative results for controllable motion synthesis using a pre-trained text-to-motion diffusion model with 100 DDIM sampling steps. Performance is measured by Objective loss (the trajectory/obstacle loss divided by number of frames, ↓\downarrow) and Embedding loss (Euclidean distance between motion embeddings and the text prompt embeddings, ↓\downarrow).

    Path Following Results (averaged across 3 target directions and 10 seeds per prompt):

    Prompt (i) "backwards" (ii) "balance beam" (iii) "jogging" (iv) öbject above head"
    Metric Obj. ↓\downarrow Emb. ↓\downarrow Obj. ↓\downarrow Emb. ↓\downarrow Obj. ↓\downarrow Emb. ↓\downarrow Obj. ↓\downarrow Emb. ↓\downarrow
    Baseline 11.583 6.405 1.024 3.756 5.831 5.017 7.080 3.397
    LGD-DPS 2.168 5.891 0.388 3.784 3.521 5.090 2.954 3.545
    LGD-MC (n=10n = 10) 2.150 5.887 0.395 3.778 3.505 5.093 2.857 3.515
    LGD-MC (n=100n = 100) 2.126 5.886 0.402 3.797 3.458 5.100 2.800 3.531

    Obstacle Avoidance Results:

    Prompt (v) "walking" (vi) "backwards"
    Metric Obj. ↓\downarrow Emb. ↓\downarrow Obj. ↓\downarrow Emb. ↓\downarrow
    Baseline 43.736 2.759 42.933 6.519
    LGD-DPS 7.040 4.405 5.307 6.240
    LGD-MC (n=10n = 10) 3.902 3.187 3.921 6.339
    LGD-MC (n=100n = 100) 3.666 3.100 4.160 6.633

    For path following, LGD-MC achieves lower objective loss than LGD-DPS across most prompts while matching text prompt alignment. On obstacle avoidance, where the step-like collision penalty yields vanishing gradients for point estimates, LGD-MC reduces objective loss substantially compared to LGD-DPS (e.g., from 7.040 to 3.666 for prompt v).

  8. Knowl 8 — Label-Conditioned Generation on Unconditional ImageNet Models

    data/table

    The table compares conditional generation paradigms applied to an unconditional 256×256256 \times 256 ImageNet diffusion model using pre-trained classifiers at t=10−3t = 10^{-3} as guidance. Evaluation is over 10,000 generated samples using Fréchet Inception Distance (FID, ↓\downarrow) for image quality and classification loss (Loss, ↓\downarrow) for condition fidelity.

    Method FID ↓\downarrow Loss ↓\downarrow
    Baseline 37.28 281.84
    Classifier Guidance (CG) 44.64 25.47
    Diffusion as Plug-and-Play Prior (D-PnP) 317.21 221.75
    LGD-DPS 48.24 31.95
    LGD-MC (n=1n = 1) 39.87 33.73
    LGD-MC (n=5n = 5) 38.65 32.80

    Optimization-based plug-and-play prior methods (D-PnP) fail to synthesize coherent images on complex multimodal datasets like ImageNet (FID 317.21). Standard Classifier Guidance and LGD-DPS achieve lower classification loss but degrade image fidelity (FID 44.64 and 48.24). LGD-MC (n=5n=5) balances control and visual fidelity, attaining a 32.80 classification loss while maintaining an FID of 38.65, close to the unguided baseline.

  9. Knowl 9 — Computational Overhead and Scaling of LGD-MC

    data/table

    Memory and wall-clock execution time per sampling iteration for 256×256256 \times 256 image super-resolution evaluated on a single NVIDIA RTX 3090 GPU across varying numbers of Monte Carlo samples nn.

    Resources Peak memory Wall-clock time / iter
    LGD-DPS 16.51 GB 0.380 s
    LGD-MC (n=1n = 1) 16.51 GB 0.387 s
    LGD-MC (n=10n = 10) 16.53 GB 0.390 s
    LGD-MC (n=100n = 100) 16.67 GB 0.410 s

    Because the heavy neural network DθD_\theta is evaluated only once per diffusion step and loss evaluations over the nn samples are vectorized and differentiated jointly through x^t\hat{x}_t, scaling nn from 1 to 100 increases per-iteration runtime by only 5.9%5.9\% (from 0.387 s to 0.410 s) and peak VRAM by less than 1%1\% (from 16.51 GB to 16.67 GB).

  10. Knowl 10 — Limitations of Loss-Guided Diffusion and Gaussian Variational Approximation

    limitation

    LGD-MC has two main limitations:

    1. Inaccuracy of the Gaussian assumption: The posterior approximation q(x0∣xt)=N(x^t,rt2I)q(x_0|x_t) = \mathcal{N}(\hat{x}_t, r_t^2 I) assumes an isotropic Gaussian distribution around the MMSE prediction. In reality, data distributions and transition posteriors p(x0∣xt)p(x_0|x_t) are highly non-Gaussian and multimodal, limiting the precision of the Monte Carlo score estimates.
    2. Prior coverage constraints: Like other plug-and-play conditional generation methods, LGD cannot generate valid samples when the target conditions yy specify regions outside the support of the pre-trained generative prior p0(x0)p_0(x_0).

Coverage note — None was omitted; all key theoretical formulations, approximations, algorithms, empirical benchmarks, and stated limitations are represented.

References

  1. 1.Anonymous. Pseudoinverse-guided diffusion models for inverse problems. 2022.
  2. 2.Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Kreis, K., Aittala, M., Aila, T., Laine, S., Catanzaro, B., et al. ediffi: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324, 2022.
  3. 3.Bautista, M. A., Guo, P., Abnar, S., Talbott, W., Toshev, A., Chen, Z., Dinh, L., Zhai, S., Goh, H., Ulbricht, D., et al. Gaudi: A neural architect for immersive 3d scene generation. arXiv preprint arXiv:2207.13751, 2022.
  4. 4.Biloš, M., Rasul, K., Schneider, A., Nevmyvaka, Y., and Günnemann, S. Modeling temporal data as continuous functions with process diffusion. arXiv preprint arXiv:2211.02590, 2022.
  5. 5.Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Kohd, P. W., Krass, M., Krishna, R., Kuditipudi, R., Kumar, A., Ladhak, F., Lee, M., Lee, T., Leskovec, J., Levent, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C. D., Mirchandani, S., Mitchell, E., Munyikwa, Z., Nair, S., Narayan, A., Narayanan, D., Newman, B., Nie, A., Niebles, J. C., Nilforoshan, H., Nyarko, J., Ogut, G., Orr, L., Papadimitriou, I., Park, J. S., Piech, C., Portelance, E., Potts, C., Raghunathan, A., Reich, R., Ren, H., Rong, F., Roohani, Y., Ruiz, C., Ryan, J., Re, C., Sadigh, D., Sagawa, S., Santhanam, K., Shih, A., Srinivasan, K., Tamkin, A., Taori, R., Thomas, A. W., Tramèr, F., Wang, R. E., Wang, W., Wu, B., Wu, J., Wu, Y., Xie, S. M., Yasunaga, M., You, J., Zaharia, M., Zhang, M., Zhang, T., Zhang, X., Zhang, Y., Zheng, L., Zhou, K., and Liang, P. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, August 2021.
  6. 6.Choi, J., Kim, S., Jeong, Y., Gwon, Y., and Yoon, S. ILVR: Conditioning method for denoising diffusion probabilistic models. arXiv preprint arXiv:2108.02938, August 2021.
  7. 7.Chung, H. and Ye, J. C. Score-based diffusion models for accelerated mri. Medical Image Analysis, pp. 102479, 2022.
  8. 8.Chung, H., Kim, J., Mccann, M. T., Klasky, M. L., and Ye, J. C. Diffusion posterior sampling for general noisy inverse problems. arXiv preprint arXiv:2209.14687, 2022.
  9. 9.Crowson, K., Biderman, S., Kornis, D., Stander, D., Hallahan, E., Castricato, L., and Raff, E. Vqgan-clip: Open domain image generation and editing with natural language guidance. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVII, pp. 88–105. Springer, 2022.
  10. 10.Dhariwal, P. and Nichol, A. Diffusion models beat GANs on image synthesis. arXiv preprint arXiv:2105.05233, May 2021.
  11. 11.Dockhorn, T., Vahdat, A., and Kreis, K. GENIE: Higher-Order Denoising Diffusion Solvers. In Advances in Neural Information Processing Systems, 2022.
  12. 12.Efron, B. Tweedie’s formula and selection bias. Journal of the American Statistical Association, 106(496):1602–1614, 2011.
  13. 13.Graikos, A., Malkin, N., Jojic, N., and Samaras, D. Diffusion models as plug-and-play priors. arXiv preprint arXiv:2206.09012, 2022.
  14. 14.Grathwohl, W., Wang, K.-C., Jacobsen, J.-H., Duvenaud, D., Norouzi, M., and Swersky, K. Your classifier is secretly an energy based model and you should treat it like one. arXiv preprint arXiv:1912.03263, 2019.
  15. 15.Grenander, U. and Miller, M. I. Representations of knowledge in complex systems. Journal of the Royal Statistical Society: Series B (Methodological), 56(4): 549–581, 1994.
  16. 16.Guo, C., Zou, S., Zuo, X., Wang, S., Ji, W., Li, X., and Cheng, L. Generating diverse and natural 3d human motions from text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5152–5161, 2022.
  17. 17.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. GANs trained by a two Time-Scale update rule converge to a local nash equilibrium. arXiv preprint arXiv:1706.08500, June 2017.
  18. 18.Ho, J. and Salimans, T. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022.
  19. 19.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. arXiv preprint arXiv:2006.11239, June 2020.
  20. 20.Ho, J., Salimans, T., Gritsenko, A., Chan, W., Norouzi, M., and Fleet, D. J. Video diffusion models. arXiv preprint arXiv:2204.03458, 2022.
  21. 21.Jalal, A., Arvinte, M., Daras, G., Price, E., Dimakis, A. G., and Tamir, J. I. Robust compressed sensing MRI with deep generative priors. arXiv preprint arXiv:2108.01368, August 2021.
  22. 22.Karchev, K., Montel, N. A., Coogan, A., and Weniger, C. Strong-lensing source reconstruction with denoising diffusion restoration models. arXiv preprint arXiv:2211.04365, 2022.
  23. 23.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. arXiv preprint arXiv:2206.00364, 2022.
  24. 24.Kawar, B., Vaksman, G., and Elad, M. SNIPS: Solving noisy inverse problems stochastically. arXiv preprint arXiv:2105.14951, May 2021.
  25. 25.Kawar, B., Elad, M., Ermon, S., and Song, J. Denoising diffusion restoration models. arXiv preprint arXiv:2201.11793, 2022a.
  26. 26.Kawar, B., Song, J., Ermon, S., and Elad, M. Jpeg artifact correction using denoising diffusion restoration models. arXiv preprint arXiv:2209.11888, 2022b.
  27. 27.Kong, Z., Ping, W., Huang, J., Zhao, K., and Catanzaro, B. Diffwave: A versatile diffusion model for audio synthesis. arXiv preprint arXiv:2009.09761, 2020.
  28. 28.Li, X. L., Thickstun, J., Gulrajani, I., Liang, P., and Hashimoto, T. B. Diffusion-lm improves controllable text generation. arXiv preprint arXiv:2205.14217, 2022.
  29. 29.Lin, C.-H., Gao, J., Tang, L., Takikawa, T., Zeng, X., Huang, X., Kreis, K., Fidler, S., Liu, M.-Y., and Lin, T.-Y. Magic3d: High-resolution text-to-3d content creation. arXiv preprint arXiv:2211.10440, 2022.
  30. 30.Liu, L., Ren, Y., Lin, Z., and Zhao, Z. Pseudo numerical methods for diffusion models on manifolds. arXiv preprint arXiv:2202.09778, 2022.
  31. 31.Lovelace, J., Kishore, V., Wan, C., Shekhtman, E., and Weinberger, K. Latent diffusion for language generation. arXiv preprint arXiv:2212.09462, 2022.
  32. 32.Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. arXiv preprint arXiv:2206.00927, 2022.
  33. 33.Meng, C., Gao, R., Kingma, D. P., Ermon, S., Ho, J., and Salimans, T. On distillation of guided diffusion models. arXiv preprint arXiv:2210.03142, 2022.
  34. 34.Poole, B., Jain, A., Barron, J. T., and Mildenhall, B. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988, 2022.
  35. 35.Qiao, Z., Nie, W., Vahdat, A., Miller III, T. F., and Anandkumar, A. Dynamic-backbone protein-ligand structure prediction with multiscale generative diffusion models. arXiv preprint arXiv:2209.15171, 2022.
  36. 36.Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I. Learning transferable visual models from natural language supervision. arXiv preprint arXiv:2103.00020, February 2021.
  37. 37.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10684–10695, 2022.
  38. 38.Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. Photorealistic text-to-image diffusion models with deep language understanding. arXiv preprint arXiv:2205.11487, 2022.
  39. 39.Saito, K., Murata, N., Uesaka, T., Lai, C.-H., Takida, Y., Fukui, T., and Mitsufuji, Y. Unsupervised vocal dereverberation with diffusion-based generative models. arXiv preprint arXiv:2211.04124, 2022.
  40. 40.Salimans, T. and Ho, J. Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512, February 2022.
  41. 41.Santurkar, S., Ilyas, A., Tsipras, D., Engstrom, L., Tran, B., and Madry, A. Image synthesis with a single (robust) classifier. Advances in Neural Information Processing Systems, 32, 2019.
  42. 42.Sawata, R., Murata, N., Takida, Y., Uesaka, T., Shibuya, T., Takahashi, S., and Mitsufuji, Y. A versatile diffusion-based generative refiner for speech enhancement. arXiv preprint arXiv:2210.17287, 2022.
  43. 43.Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402, 2022.
  44. 44.Shue, J. R., Chan, E. R., Po, R., Ankner, Z., Wu, J., and Wetzstein, G. 3d neural field generation using triplane diffusion. arXiv preprint arXiv:2211.16677, 2022.
  45. 45.Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. arXiv preprint arXiv:1503.03585, March 2015.
  46. 46.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In International Conference on Learning Representations, 2021a.
  47. 47.Song, Y., Shen, L., Xing, L., and Ermon, S. Solving inverse problems in medical imaging with score-based generative models. arXiv preprint arXiv:2111.08005, 2021b.
  48. 48.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021c.
  49. 49.Tashiro, Y., Song, J., Song, Y., and Ermon, S. Csdi: Conditional score-based diffusion models for probabilistic time series imputation. Advances in Neural Information Processing Systems, 34:24804–24816, 2021.
  50. 50.Tevet, G., Raab, S., Gordon, B., Shafir, Y., Cohen-Or, D., and Bermano, A. H. Human motion diffusion model. arXiv preprint arXiv:2209.14916, 2022.
  51. 51.Venkatakrishnan, S. V., Bouman, C. A., and Wohlberg, B. Plug-and-play priors for model based reconstruction. In 2013 IEEE Global Conference on Signal and Information Processing, pp. 945–948. IEEE, 2013.
  52. 52.Vincent, P. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661–1674, July 2011.
  53. 53.Wu, K. E., Yang, K. K., Berg, R. v. d., Zou, J. Y., Lu, A. X., and Amini, A. P. Protein structure generation via folding diffusion. arXiv preprint arXiv:2209.15611, 2022.
  54. 54.Xiao, Z., Kreis, K., and Vahdat, A. Tackling the generative learning trilemma with denoising diffusion GANs. In International Conference on Learning Representations (ICLR), 2022.
  55. 55.Xu, D., Jiang, Y., Wang, P., Fan, Z., Wang, Y., and Wang, Z. Neurallift-360: Lifting an in-the-wild 2d photo to a 3d object with 360° views. arXiv e-prints, pp. arXiv–2211, 2022a.
  56. 56.Xu, M., Yu, L., Song, Y., Shi, C., Ermon, S., and Tang, J. Geodiff: A geometric diffusion model for molecular conformation generation. arXiv preprint arXiv:2203.02923, 2022b.
  57. 57.Zhang, Q. and Chen, Y. Fast sampling of diffusion models with exponential integrator. arXiv preprint arXiv:2204.13902, April 2022.
  58. 58.Zhang, Q., Tao, M., and Chen, Y. gddim: Generalized denoising diffusion implicit models. arXiv preprint arXiv:2206.05564, June 2022.
  59. 59.Zheng, H., Nie, W., Vahdat, A., Azizzadenesheli, K., and Anandkumar, A. Fast sampling of diffusion models via operator learning. arXiv preprint arXiv:2211.13449, 2022.

Citation

MLA
Song, J., et al. “Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation”. International Conference on Machine Learning, vol. 202, 2023, pp. 32483–98, https://proceedings.mlr.press/v202/song23k.html.
APA
Song, J., Zhang, Q., Yin, H., Mardani, M., Liu, M.-Y., Kautz, J., Chen, Y., & Vahdat, A. (2023). Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation. International Conference on Machine Learning, 202, 32483–32498. https://proceedings.mlr.press/v202/song23k.html
Chicago
Song, J., Q. Zhang, H. Yin, et al. 2023. “Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation”. International Conference on Machine Learning 202: 32483–98. https://proceedings.mlr.press/v202/song23k.html.
Harvard
Song, J. et al. (2023) “Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation”, International Conference on Machine Learning. PMLR, pp. 32483–32498. Available at: https://proceedings.mlr.press/v202/song23k.html.
Vancouver
1. Song J, Zhang Q, Yin H, Mardani M, Liu M-Y, Kautz J, Chen Y, Vahdat A (2023) Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation. In: International Conference on Machine Learning. PMLR, pp 32483–32498

BibTeX

@InProceedings{pmlr-v202-song23k,
  title = 	 {Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation},
  author =       {Song, Jiaming and Zhang, Qinsheng and Yin, Hongxu and Mardani, Morteza and Liu, Ming-Yu and Kautz, Jan and Chen, Yongxin and Vahdat, Arash},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {32483--32498},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/song23k/song23k.pdf},
  url = 	 {https://proceedings.mlr.press/v202/song23k.html},
  abstract = 	 {We consider guiding denoising diffusion models with general differentiable loss functions in a plug-and-play fashion, enabling controllable generation without additional training. This paradigm, termed Loss-Guided Diffusion (LGD), can easily be integrated into all diffusion models and leverage various efficient samplers. Despite the benefits, the resulting guidance term is, unfortunately, an intractable integral and needs to be approximated. Existing methods compute the guidance term based on a point estimate. However, we show that such approaches have significant errors over the scale of the approximations. To address this issue, we propose a Monte Carlo method that uses multiple samples from a suitable distribution to reduce bias. Our method is effective in various synthetic and real-world settings, including image super-resolution, text or label-conditional image generation, and controllable motion synthesis. Notably, we show how our method can be applied to control a pretrained motion diffusion model to follow certain paths and avoid obstacles that are proven challenging to prior methods.}
}
Metadata:DOI registry

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/