A Variational Perspective on Solving Inverse Problems with Diffusion Models

Morteza MardaniJiaming SongJan KautzArash Vahdat

article2024ICLR290 citations

Proposes a variational framework that transforms diffusion-based posterior sampling into stochastic optimization, enabling the use of standard solvers to perform image restoration without task-specific retraining.

Listen

Inverse problems—reconstructing complete, high-quality images or signals from incomplete, blurred, or noisy observations—are foundational to fields ranging from medical imaging to computer vision. While modern generative diffusion models offer powerful data priors that can solve these tasks without requiring custom retraining, existing posterior sampling methods suffer from severe limitations. In particular, popular approaches rely on crude mathematical approximations to infer complex, multi-modal posterior scores, or require costly and numerically unstable backpropagation through the neural network. This creates significant computational bottlenecks and degrades image reconstruction fidelity.

The article demonstrates a mathematically grounded, variational framework for solving general inverse problems using pre-trained diffusion models. Specifically, the researchers formulate posterior inference as a maximum-likelihood optimization problem that directly matches distributions, bypassing the intractable posterior score approximations of previous methods.

To achieve this, the authors introduce RED-diff, a method that regularizes standard reconstruction loss through a denoising diffusion process. By casting sampling as stochastic optimization, RED-diff leverages the denoisers across all diffusion stages—from coarse semantic modeling to fine detail generation—guided by a signal-to-noise ratio weighting mechanism. The authors validated RED-diff across linear and non-linear image restoration benchmarks, including inpainting, 4x super-resolution, high dynamic range recovery, phase retrieval, and non-linear deblurring using a 1,000-image ImageNet validation subset, as well as magnetic resonance imaging (MRI) datasets.

Key findings show substantial improvements in both reconstruction fidelity and computational efficiency:

  • In standard inpainting, RED-diff improved reconstruction fidelity to 23.29 dB PSNR compared to 20.30–21.27 dB achieved by baseline methods, while drastically lowering perceptual error (KID dropped to 0.86 compared to 2.5–15.28 for alternatives).
  • For challenging non-linear tasks such as deblurring, RED-diff achieved 45.00 dB PSNR and an SSIM of 0.987, whereas leading baseline DPS collapsed at 6.40 dB and 0.19 SSIM.
  • In medical MRI reconstruction, RED-diff reduced per-iteration runtime by roughly 67% (0.114 seconds versus 0.344 seconds) while consistently improving reconstruction fidelity across acceleration factors.
  • RED-diff eliminated the need to backpropagate through the score network or compute matrix decompositions, running at 0.05 seconds per step—roughly two to four times faster than competing methods—and supporting double the maximum batch size on equivalent graphics processing units.

These results establish that sampling can be efficiently performed using standard off-the-shelf optimizers, such as Adam or stochastic gradient descent with momentum. This transition lowers hardware infrastructure demands, stabilizes convergence in non-linear workflows, and provides practical tuning parameters (such as learning rate and step counts) to balance reconstruction accuracy against visual sharpness.

Organizations implementing generative diffusion pipelines for image restoration or medical reconstruction should consider transitioning toward variational optimization frameworks like RED-diff to reduce GPU compute costs and stabilize inversion pipelines. For production deployment, engineering teams should implement descending timestep schedules and inverse signal-to-noise weighting, which yielded optimal performance in ablations.

The findings are bounded by the mode-seeking nature of the variational distribution, which targets maximum a posteriori estimates and thus limits the diversity of generated outputs compared to pure stochastic samplers. While confidence in image fidelity and efficiency gains is high across 2D visual and MRI domains, further empirical validation is necessary before applying the framework to higher-dimensional tasks such as 3D generation.

arXiv: 2305.04391
  • Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). This paper establishes the foundational sampling-based posterior approximation via Tweedie's formula for solving general inverse problems with diffusion models, providing the direct baseline and context that the source improves upon with its variational optimization approach.
  • Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). It introduces a pioneering unsupervised framework for solving linear inverse problems with pre-trained diffusion models, establishing core concepts of noise scheduling and data fidelity in diffusion-based restoration.
  • Paper: Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model, Yinhuai Wang et al. (2023). This work formulates range-null space decomposition for zero-shot diffusion inverse restoration, representing a key prior restoration paradigm that the source contrasts against.
  • Paper: Variational Diffusion Models, Diederik P. Kingma et al. (2021). It establishes the theoretical connection between continuous-time diffusion processes, variational lower bounds, and signal-to-noise-ratio (SNR) weighting that directly informs the source's variational formulation.
  • Paper: RePaint: Inpainting using Denoising Diffusion Probabilistic Models, Andreas Lugmayr et al. (2022). It introduces unconditional diffusion conditioning for image inpainting through iterative measurement replacement, serving as a primary task benchmark and baseline for diffusion inverse solvers.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). This seminal work provides the foundational mathematical formulation of denoising diffusion probabilistic models and multi-scale score matching upon which all diffusion-based inverse solving relies.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). It develops non-Markovian deterministic sampling for diffusion models, which underpins the accelerated sampling and optimization mechanics used in posterior inference.
  • Paper: Elucidating the Design Space of Diffusion-Based Generative Models, Tero Karras et al. (2022). It clarifies the design space, preconditioning, and noise scaling of diffusion models, providing crucial technical principles for denoiser weighting and score matching.
  • Paper: Learning Deep CNN Denoiser Prior for Image Restoration, Kai Zhang et al. (2017). It introduces the regularization by denoising (RED) concept using deep CNN denoisers in variable-splitting optimization frameworks, serving as the conceptual precursor to the source's RED-Diff regularizer.
Cover for A Variational Perspective on Solving Inverse Problems with Diffusion Models

Abstract

Diffusion models have emerged as a key pillar of foundation models in visual domains. One of their critical applications is to universally solve different downstream inverse tasks via a single diffusion prior without re-training for each task. Most inverse tasks can be formulated as inferring a posterior distribution over data (e.g., a full image) given a measurement (e.g., a masked image). This is however challenging in diffusion models since the nonlinear and iterative nature of the diffusion process renders the posterior intractable. To cope with this challenge, we propose a variational approach that by design seeks to approximate the true posterior distribution. We show that our approach naturally leads to regularization by denoising diffusion process (RED-Diff) where denoisers at different timesteps concurrently impose different structural constraints over the image. To gauge the contribution of denoisers from different timesteps, we propose a weighting mechanism based on signal-to-noise-ratio (SNR). Our approach provides a new variational perspective for solving inverse problems with diffusion models, allowing us to formulate sampling as stochastic optimization, where one can simply apply off-the-shelf solvers with lightweight iterates. Our experiments for image restoration tasks such as inpainting and superresolution demonstrate the strengths of our method compared with state-of-the-art sampling-based diffusion models.

Table of Contents

  • 1 Introduction
  • 2 Related works
  • 3 Background
  • 3.1 Denoising diffusion models
  • 3.2 Score approximation for inverse problems
  • 4 Variational diffusion sampling
  • 4.1 Sampling as stochastic optimization
  • 4.2 Regularization by denoising
  • 4.3 Weighting mechanism
  • 5 Experiments
  • 5.1 Image inpainting
  • 5.2 Nonlinear inverse problems
  • 5.3 Ablations
  • 5.3.1 Denoiser weighting mechanism
  • 5.3.2 Timestep sampling
  • 6 Conclusions and limitations
  • References
  • A Proofs and Derivations
  • A.1 Proof of Proposition 1
  • A.2 Proof of Proposition 2
  • A.3 Adding dispersion to variational approximation
  • B Additional experiments
  • B.1 Pretrained diffusion model
  • B.2 Image superresolution
  • B.3 More examples for comparisons
  • B.4 Diversity
  • B.5 Diffusion evolution
  • C Additional scenarios
  • C.1 Compressed sensing MRI
  • C.2 Noisy inpainting
  • D Ablations
  • D.1 Denoiser weighting mechanism
  • D.2 Optimizing RED-diff for more epochs
  • D.3 Optimization strategy
  • D.4 Timestep sampling strategy
  • D.5 Connection and differences with RED

Citation

MLA
Mardani, M., et al. “A Variational Perspective on Solving Inverse Problems with Diffusion Models”. arXiv, 2023, http://arxiv.org/abs/2305.04391v2.
APA
Mardani, M., Song, J., Kautz, J., & Vahdat, A. (2023). A Variational Perspective on Solving Inverse Problems with Diffusion Models. arXiv. http://arxiv.org/abs/2305.04391v2
Chicago
Mardani, M., J. Song, J. Kautz, and A. Vahdat. 2023. “A Variational Perspective on Solving Inverse Problems with Diffusion Models”. arXiv. http://arxiv.org/abs/2305.04391v2.
Harvard
Mardani, M. et al. (2023) “A Variational Perspective on Solving Inverse Problems with Diffusion Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2305.04391v2.
Vancouver
1. Mardani M, Song J, Kautz J, Vahdat A (2023) A Variational Perspective on Solving Inverse Problems with Diffusion Models. arXiv

BibTeX

@article{mardani2023variational,
  title = {A Variational Perspective on Solving Inverse Problems with Diffusion Models},
  author = {Mardani, Morteza and Song, Jiaming and Kautz, Jan and Vahdat, Arash},
  year = {2023},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2305.04391v2},
  eprint = {2305.04391}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors