Denoising Diffusion Restoration Models

Bahjat KawarMichael EladS. ErmonJiaming Song

article2022NeurIPS1,395 citations

Proposes Denoising Diffusion Restoration Models (DDRM), an unsupervised method that leverages pre-trained unconditional diffusion models to efficiently solve diverse linear inverse problems—such as super-resolution, deblurring, and inpainting—in as few as 20 steps without requiring task-specific training.

Listen

Real-world image restoration tasks, such as enhancing resolution, removing blur, filling in missing regions, and adding color, are fundamental across fields like medical imaging and consumer photography. Existing machine learning solutions present a difficult trade-off: supervised methods deliver fast results but require specialized retraining for every new type of degradation, while flexible unsupervised methods that use pre-trained image models typically rely on slow, computationally demanding iterative procedures requiring hundreds or thousands of steps.

The article demonstrates an unsupervised restoration framework called Denoising Diffusion Restoration Models (DDRM). The objective of the work is to show that a single pre-trained generative diffusion model can solve diverse linear inverse image restoration tasks efficiently and accurately without requiring task-specific retraining.

The authors designed a mathematical approach that transforms degraded image measurements into their underlying spectral components, aligning the measurement noise directly with the diffusion model's denoising schedule. To validate performance, the authors conducted extensive experiments across standard benchmarks, evaluating image fidelity, structural similarity, and perceptual visual quality. The evaluation tested tasks such as super-resolution, deblurring, inpainting, and colorization under both clean and heavily noisy conditions, comparing the method against existing unsupervised algorithms on diverse datasets such as ImageNet.

The evaluation revealed several key findings. First, DDRM achieved state-of-the-art reconstruction accuracy and perceptual quality using as few as 20 computational steps, operating at least 50 times faster than prior sampling-based methods that demand 1,000 or more steps. Second, the method maintained strong restoration fidelity when handling substantial measurement noise, avoiding the severe visual artifacts that degrade competing optimization techniques. Third, the framework successfully restored general natural images lying entirely outside the domain of its original training dataset, demonstrating robust cross-domain flexibility.

These findings indicate that organizations can deploy a single pre-trained foundation model to handle multiple distinct imaging problems, significantly reducing compute costs, operational latency, and the need for expensive model retraining pipelines. Furthermore, the framework's resilience against measurement noise improves reliability in mission-critical applications where input data is imperfect.

Decision-makers should consider adopting diffusion-based spectral restoration pipelines where flexible, multi-task image recovery is required at low inference budgets. For future work, development teams should explore expanding the technique to non-linear degradation settings and scenarios where the precise mathematical distortion model is unknown.

Confidence in these findings is high for linear degradation problems with known mathematical operators. However, users should exercise caution when applying the method to non-linear physical distortions or uncharacterized imaging systems, as these scenarios lie outside the current mathematical formulation.

  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Establishes the fundamental mathematical formulations and U-Net training techniques for Denoising Diffusion Probabilistic Models (DDPM) that DDRM uses as foundational generative priors.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Introduces non-Markovian sampling trajectories (DDIM) that enable accelerated diffusion inference schedules upon which DDRM builds its efficient spectral restoration steps.
  • Paper: Diffusion Models Beat GANs on Image Synthesis, Prafulla Dhariwal et al. (2021). Demonstrates scaled diffusion architectures and guidance principles that provide the high-capacity generative image priors leveraged in zero-shot restoration frameworks.
  • Paper: Improved Denoising Diffusion Probabilistic Models, Alex Nichol et al. (2021). Refines noise schedules and variance formulations in diffusion models, offering foundational insights for accelerating reverse sampling schedules.
  • Paper: RePaint: Inpainting using Denoising Diffusion Probabilistic Models, Andreas Lugmayr et al. (2022). Pioneers the use of unconditional pre-trained diffusion models for inverse problems via modified reverse sampling trajectories, serving as direct motivation for DDRM's generalized spectral framework.
  • Paper: Learning Deep CNN Denoiser Prior for Image Restoration, Kai Zhang et al. (2017). Presents the plug-and-play optimization paradigm for using deep denoisers as explicit priors to solve linear inverse problems.
  • Paper: Deep Image Prior, Dmitry Ulyanov et al. (2017). Introduces the concept of utilizing unsupervised deep network representations as natural image priors for inverse tasks like super-resolution and inpainting without task-specific retraining.
Cover for Denoising Diffusion Restoration Models

Abstract

Many interesting tasks in image restoration can be cast as linear inverse problems. A recent family of approaches for solving these problems uses stochastic algorithms that sample from the posterior distribution of natural images given the measurements. However, efficient solutions often require problem-specific supervised training to model the posterior, whereas unsupervised methods that are not problem-specific typically rely on inefficient iterative methods. This work addresses these issues by introducing Denoising Diffusion Restoration Models (DDRM), an efficient, unsupervised posterior sampling method. Motivated by variational inference, DDRM takes advantage of a pre-trained denoising diffusion generative model for solving any linear inverse problem. We demonstrate DDRM’s versatility on several image datasets for super-resolution, deblurring, inpainting, and colorization under various amounts of measurement noise. DDRM outperforms the current leading unsupervised methods on the diverse ImageNet dataset in reconstruction quality, perceptual quality, and runtime, being 5× faster than the nearest competitor. DDRM also generalizes well for natural images out of the distribution of the observed ImageNet training set.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 3 Denoising Diffusion Restoration Models
  • 3.1 Variational Objective for DDRM
  • 3.2 A Diffusion Process for Image Restoration
  • 3.3 “Learning” Image Restoration Models
  • 3.4 Accelerated Algorithms for DDRM
  • 3.5 Memory Efficient SVD
  • 4 Related Work
  • 5 Experiments
  • 5.1 Experimental Setup
  • 5.2 Quantitative Experiments
  • 5.3 Qualitative Experiments
  • 6 Conclusions
  • References
  • A Details of the DDRM ELBO objective
  • B Equivalence between “Variance Preserving” and “Variance Exploding” Diffusion Models
  • C Proofs
  • D Memory Efficient SVD
  • D.1 Denoising
  • D.2 Inpainting
  • D.3 Super Resolution
  • D.4 Colorization
  • D.5 Deblurring
  • E Ablation Studies on Hyperparameters
  • F Experimental Setup of DGP, RED, and SNIPS
  • G Runtime of Algorithms
  • H ILVR as a special case of DDRM
  • I Additional Results

Knowls

  1. Knowl 1 — Spectral Space Representation for Linear Inverse Problems

    definition

    A linear inverse problem seeks to recover a clean signal x∈Rn\mathbf{x} \in \mathbb{R}^n from degraded measurements y∈Rm\mathbf{y} \in \mathbb{R}^m governed by the observation model:

    y=Hx+z\mathbf{y} = \mathbf{H}\mathbf{x} + \mathbf{z}

    where H∈Rm×n\mathbf{H} \in \mathbb{R}^{m \times n} (with m≤nm \le n) is a known linear degradation operator, and z∼N(0,σy2Im)\mathbf{z} \sim \mathcal{N}(0, \sigma_y^2 \mathbf{I}_m) is additive Gaussian noise with known standard deviation σy\sigma_y.

    Using the Singular Value Decomposition (SVD) of the degradation matrix:

    H=UΣV⊤\mathbf{H} = \mathbf{U}\mathbf{\Sigma}\mathbf{V}^\top

    where U∈Rm×m\mathbf{U} \in \mathbb{R}^{m \times m} and V∈Rn×n\mathbf{V} \in \mathbb{R}^{n \times n} are orthogonal matrices, and Σ∈Rm×n\mathbf{\Sigma} \in \mathbb{R}^{m \times n} is a rectangular diagonal matrix with ordered singular values s1≥s2≥⋯≥sm≥0s_1 \ge s_2 \ge \dots \ge s_m \ge 0 (with si=0s_i = 0 defined for i∈[m+1,n]i \in [m+1, n]), the spectral-space representations of a diffusion state xt∈Rn\mathbf{x}_t \in \mathbb{R}^n and the measurement y∈Rm\mathbf{y} \in \mathbb{R}^m are given by:

    xˉt=V⊤xt,yˉ=Σ†U⊤y\bar{\mathbf{x}}_t = \mathbf{V}^\top \mathbf{x}_t, \quad \bar{\mathbf{y}} = \mathbf{\Sigma}^\dagger \mathbf{U}^\top \mathbf{y}

    where Σ†∈Rn×m\mathbf{\Sigma}^\dagger \in \mathbb{R}^{n \times m} is the Moore–Penrose pseudoinverse of Σ\mathbf{\Sigma}. The ii-th coordinate of these spectral vectors is denoted by xˉt(i)\bar{x}_t^{(i)} and yˉ(i)\bar{y}^{(i)} respectively, and the signal xt\mathbf{x}_t can be recovered exactly via xt=Vxˉt\mathbf{x}_t = \mathbf{V}\bar{\mathbf{x}}_t.

  2. Knowl 2 — DDRM Forward Variational Inference Distribution

    model/method

    Given a sequence of strictly increasing noise levels 0=σ0<σ1<σ2<⋯<σT0 = \sigma_0 < \sigma_1 < \sigma_2 < \dots < \sigma_T satisfying σT≥σy/si\sigma_T \ge \sigma_y / s_i for all si>0s_i > 0, Denoising Diffusion Restoration Models (DDRM) define a factorized variational distribution q(x1:T∣x0,y)=q(T)(xT∣x0,y)∏t=0T−1q(t)(xt∣xt+1,x0,y)q(\mathbf{x}_{1:T} | \mathbf{x}_0, \mathbf{y}) = q^{(T)}(\mathbf{x}_T | \mathbf{x}_0, \mathbf{y}) \prod_{t=0}^{T-1} q^{(t)}(\mathbf{x}_t | \mathbf{x}_{t+1}, \mathbf{x}_0, \mathbf{y}).

    Operating coordinate-wise on the spectral space xˉt=V⊤xt\bar{\mathbf{x}}_t = \mathbf{V}^\top \mathbf{x}_t, the forward distributions are defined by:

    q(T)(xˉT(i)∣x0,y)={N(yˉ(i), σT2−σy2si2)if si>0N(xˉ0(i), σT2)if si=0q^{(T)}(\bar{x}_T^{(i)} | \mathbf{x}_0, \mathbf{y}) = \begin{cases} \mathcal{N}\left(\bar{y}^{(i)},\, \sigma_T^2 - \frac{\sigma_y^2}{s_i^2}\right) & \text{if } s_i > 0 \\ \mathcal{N}\left(\bar{x}_0^{(i)},\, \sigma_T^2\right) & \text{if } s_i = 0 \end{cases}

    and for each timestep t∈{0,…,T−1}t \in \{0, \dots, T-1\}:

    q(t)(xˉt(i)∣xt+1,x0,y)={N(xˉ0(i)+1−η2σtxˉt+1(i)−xˉ0(i)σt+1, η2σt2)if si=0N(xˉ0(i)+1−η2σtyˉ(i)−xˉ0(i)σy/si, η2σt2)if σt<σysiN((1−ηb)xˉ0(i)+ηbyˉ(i), σt2−σy2si2ηb2)if σt≥σysiq^{(t)}(\bar{x}_t^{(i)} | \mathbf{x}_{t+1}, \mathbf{x}_0, \mathbf{y}) = \begin{cases} \mathcal{N}\left(\bar{x}_0^{(i)} + \sqrt{1 - \eta^2}\sigma_t \frac{\bar{x}_{t+1}^{(i)} - \bar{x}_0^{(i)}}{\sigma_{t+1}},\, \eta^2\sigma_t^2\right) & \text{if } s_i = 0 \\[2ex] \mathcal{N}\left(\bar{x}_0^{(i)} + \sqrt{1 - \eta^2}\sigma_t \frac{\bar{y}^{(i)} - \bar{x}_0^{(i)}}{\sigma_y / s_i},\, \eta^2\sigma_t^2\right) & \text{if } \sigma_t < \frac{\sigma_y}{s_i} \\[2ex] \mathcal{N}\left((1 - \eta_b)\bar{x}_0^{(i)} + \eta_b \bar{y}^{(i)},\, \sigma_t^2 - \frac{\sigma_y^2}{s_i^2}\eta_b^2\right) & \text{if } \sigma_t \ge \frac{\sigma_y}{s_i} \end{cases}

    where η∈(0,1]\eta \in (0, 1] is a hyperparameter governing the stochasticity of the transitions, and ηb\eta_b is a scaling hyperparameter.

  3. Knowl 3 — Marginal Gaussianity of DDRM Forward Latents

    theoretical result

    Let the forward transition distributions q(T)(xˉT∣x0,y)q^{(T)}(\bar{\mathbf{x}}_T | \mathbf{x}_0, \mathbf{y}) and q(t)(xˉt∣xt+1,x0,y)q^{(t)}(\bar{\mathbf{x}}_t | \mathbf{x}_{t+1}, \mathbf{x}_0, \mathbf{y}) be defined according to the DDRM forward variational process. For any time step t∈{0,…,T}t \in \{0, \dots, T\}, marginalizing over all intermediate future latents xt′\mathbf{x}_{t'} (for all t′>tt' > t) and over the measurement distribution q(y∣x0)=N(Hx0,σy2Im)q(\mathbf{y} | \mathbf{x}_0) = \mathcal{N}(\mathbf{H}\mathbf{x}_0, \sigma_y^2 \mathbf{I}_m) yields isotropic Gaussian marginals:

    q(xt∣x0)=N(x0,σt2In)q(\mathbf{x}_t | \mathbf{x}_0) = \mathcal{N}(\mathbf{x}_0, \sigma_t^2 \mathbf{I}_n)

    This property guarantees that the marginal distribution of latent representations at every step tt matches that of standard unconditional variance-exploding diffusion models.

  4. Knowl 4 — DDRM Reverse Transition Sampling Distribution

    model/method

    DDRM models the posterior distribution via a conditional Markov chain pθ(x0:T∣y)=pθ(T)(xT∣y)∏t=0T−1pθ(t)(xt∣xt+1,y)p_\theta(\mathbf{x}_{0:T} | \mathbf{y}) = p_\theta^{(T)}(\mathbf{x}_T | \mathbf{y}) \prod_{t=0}^{T-1} p_\theta^{(t)}(\mathbf{x}_t | \mathbf{x}_{t+1}, \mathbf{y}).

    At each step tt, an unconditional or class-conditional denoising neural network fθ(xt+1,t+1):Rn×R→Rnf_\theta(\mathbf{x}_{t+1}, t+1): \mathbb{R}^n \times \mathbb{R} \to \mathbb{R}^n produces an estimate xθ,t\mathbf{x}_{\theta,t} of the clean image x0\mathbf{x}_0. In the spectral space, xˉθ,t=V⊤xθ,t\bar{\mathbf{x}}_{\theta,t} = \mathbf{V}^\top \mathbf{x}_{\theta,t}. The reverse conditionals in spectral space are parameterized as:

    pθ(T)(xˉT(i)∣y)={N(yˉ(i), σT2−σy2si2)if si>0N(0, σT2)if si=0p_\theta^{(T)}(\bar{x}_T^{(i)} | \mathbf{y}) = \begin{cases} \mathcal{N}\left(\bar{y}^{(i)},\, \sigma_T^2 - \frac{\sigma_y^2}{s_i^2}\right) & \text{if } s_i > 0 \\ \mathcal{N}\left(0,\, \sigma_T^2\right) & \text{if } s_i = 0 \end{cases}

    and for t=T−1,…,0t = T-1, \dots, 0:

    pθ(t)(xˉt(i)∣xt+1,y)={N(xˉθ,t(i)+1−η2σtxˉt+1(i)−xˉθ,t(i)σt+1, η2σt2)if si=0N(xˉθ,t(i)+1−η2σtyˉ(i)−xˉθ,t(i)σy/si, η2σt2)if σt<σysiN((1−ηb)xˉθ,t(i)+ηbyˉ(i), σt2−σy2si2ηb2)if σt≥σysip_\theta^{(t)}(\bar{x}_t^{(i)} | \mathbf{x}_{t+1}, \mathbf{y}) = \begin{cases} \mathcal{N}\left(\bar{x}_{\theta,t}^{(i)} + \sqrt{1 - \eta^2}\sigma_t \frac{\bar{x}_{t+1}^{(i)} - \bar{x}_{\theta,t}^{(i)}}{\sigma_{t+1}},\, \eta^2\sigma_t^2\right) & \text{if } s_i = 0 \\[2ex] \mathcal{N}\left(\bar{x}_{\theta,t}^{(i)} + \sqrt{1 - \eta^2}\sigma_t \frac{\bar{y}^{(i)} - \bar{x}_{\theta,t}^{(i)}}{\sigma_y / s_i},\, \eta^2\sigma_t^2\right) & \text{if } \sigma_t < \frac{\sigma_y}{s_i} \\[2ex] \mathcal{N}\left((1 - \eta_b)\bar{x}_{\theta,t}^{(i)} + \eta_b \bar{y}^{(i)},\, \sigma_t^2 - \frac{\sigma_y^2}{s_i^2}\eta_b^2\right) & \text{if } \sigma_t \ge \frac{\sigma_y}{s_i} \end{cases}

  5. Knowl 5 — Equivalence of DDRM ELBO to Unconditional Denoising Objective

    theoretical result

    Assume that the denoising network components fθ(t)f_\theta^{(t)} and fθ(t′)f_\theta^{(t')} do not share weights whenever t≠t′t \ne t'. When setting the transition hyperparameters to η=1\eta = 1 and ηb=2σt2σt2+σy2/si2\eta_b = \frac{2\sigma_t^2}{\sigma_t^2 + \sigma_y^2 / s_i^2}, the evidence lower bound (ELBO) objective for training DDRM on a linear inverse problem reduces to the standard unconditional denoising autoencoder objective:

    ∑t=1TγtE(x0,xt)∼q(x0)q(xt∣x0)[∥x0−fθ(t)(xt)∥22]\sum_{t=1}^T \gamma_t \mathbb{E}_{(\mathbf{x}_0, \mathbf{x}_t) \sim q(\mathbf{x}_0)q(\mathbf{x}_t | \mathbf{x}_0)} \left[ \|\mathbf{x}_0 - f_\theta^{(t)}(\mathbf{x}_t)\|_2^2 \right]

    where γt>0\gamma_t > 0 are positive coefficients depending on q(x1:T∣x0)q(\mathbf{x}_{1:T} | \mathbf{x}_0).

    Because the ELBO decomposes into a weighted sum of squared errors in spectral space, pre-trained unconditional Denoising Diffusion Probabilistic Models (DDPM) or Denoising Diffusion Implicit Models (DDIM) serve as solutions to the DDRM inference problem across any linear inverse problem without retraining.

  6. Knowl 6 — DDRM Sampling Procedure

    algorithm

    DDRM solves linear inverse problems by generating a sample x0\mathbf{x}_0 from degraded measurements y\mathbf{y} across TT discrete reverse diffusion steps using a pre-trained unconditional or class-conditional denoiser fθf_\theta.

    Input: Measurements y∈Rm\mathbf{y} \in \mathbb{R}^m, degradation SVD components U,Σ,V\mathbf{U}, \mathbf{\Sigma}, \mathbf{V}, noise standard deviation σy\sigma_y, noise schedule σ1<σ2<⋯<σT\sigma_1 < \sigma_2 < \dots < \sigma_T, hyperparameters η,ηb\eta, \eta_b, denoiser fθf_\theta
    Output: Reconstructed signal x0∈Rn\mathbf{x}_0 \in \mathbb{R}^n
    Compute yˉ=Σ†U⊤y\bar{\mathbf{y}} = \mathbf{\Sigma}^\dagger \mathbf{U}^\top \mathbf{y}
    Initialize xˉT∈Rn\bar{\mathbf{x}}_T \in \mathbb{R}^n
    for each spectral index i=1,…,ni = 1, \dots, n do
        Sample ϵ∼N(0,1)\epsilon \sim \mathcal{N}(0, 1)
        if si>0s_i > 0 then
            xˉT(i)=yˉ(i)+σT2−σy2/si2⋅ϵ\bar{x}_T^{(i)} = \bar{y}^{(i)} + \sqrt{\sigma_T^2 - \sigma_y^2 / s_i^2} \cdot \epsilon
        else
            xˉT(i)=σT⋅ϵ\bar{x}_T^{(i)} = \sigma_T \cdot \epsilon
        end if
    end for
    xT=VxˉT\mathbf{x}_T = \mathbf{V}\bar{\mathbf{x}}_T
    for t=T−1t = T-1 down to 00 do
        xθ,t=fθ(xt+1,t+1)\mathbf{x}_{\theta,t} = f_\theta(\mathbf{x}_{t+1}, t+1)
        xˉθ,t=V⊤xθ,t\bar{\mathbf{x}}_{\theta,t} = \mathbf{V}^\top \mathbf{x}_{\theta,t}
        xˉt+1=V⊤xt+1\bar{\mathbf{x}}_{t+1} = \mathbf{V}^\top \mathbf{x}_{t+1}
        for each spectral index i=1,…,ni = 1, \dots, n do
            Sample ϵ∼N(0,1)\epsilon \sim \mathcal{N}(0, 1)
            if si==0s_i == 0 then
                μ=xˉθ,t(i)+1−η2σtxˉt+1(i)−xˉθ,t(i)σt+1\mu = \bar{x}_{\theta,t}^{(i)} + \sqrt{1 - \eta^2}\sigma_t \frac{\bar{x}_{t+1}^{(i)} - \bar{x}_{\theta,t}^{(i)}}{\sigma_{t+1}}
                xˉt(i)=μ+ησt⋅ϵ\bar{x}_t^{(i)} = \mu + \eta \sigma_t \cdot \epsilon
            else if σt<σy/si\sigma_t < \sigma_y / s_i then
                μ=xˉθ,t(i)+1−η2σtyˉ(i)−xˉθ,t(i)σy/si\mu = \bar{x}_{\theta,t}^{(i)} + \sqrt{1 - \eta^2}\sigma_t \frac{\bar{y}^{(i)} - \bar{x}_{\theta,t}^{(i)}}{\sigma_y / s_i}
                xˉt(i)=μ+ησt⋅ϵ\bar{x}_t^{(i)} = \mu + \eta \sigma_t \cdot \epsilon
            else
                μ=(1−ηb)xˉθ,t(i)+ηbyˉ(i)\mu = (1 - \eta_b)\bar{x}_{\theta,t}^{(i)} + \eta_b \bar{y}^{(i)}
                xˉt(i)=μ+σt2−(σy2/si2)ηb2⋅ϵ\bar{x}_t^{(i)} = \mu + \sqrt{\sigma_t^2 - (\sigma_y^2 / s_i^2)\eta_b^2} \cdot \epsilon
            end if
        end for
        xt=Vxˉt\mathbf{x}_t = \mathbf{V}\bar{\mathbf{x}}_t
    end for
    return x0\mathbf{x}_0
  7. Knowl 7 — Memory-Efficient Linear Transformations for Inverse Problems

    model/method

    Explicit storage of the right singular matrix V∈Rn×n\mathbf{V} \in \mathbb{R}^{n \times n} requires Θ(n2)\Theta(n^2) space, which presents a memory bottleneck for large images (n=d×d×Cn = d \times d \times C). For standard linear inverse tasks, structural properties of H\mathbf{H} reduce storage and computation to Θ(n)\Theta(n) space:

    1. Inpainting: H\mathbf{H} is a diagonal masking matrix, yielding V=I\mathbf{V} = \mathbf{I} with no storage overhead.
    2. Deblurring: Circular convolution operators are diagonalized by the 2D Discrete Fourier Transform (DFT). Applying V⊤\mathbf{V}^\top and V\mathbf{V} corresponds to 2D Fast Fourier Transforms (FFT), executed in O(nlog⁡n)O(n \log n) time with Θ(n)\Theta(n) auxiliary space.
    3. Super-resolution: Uniform block-averaging downsampling decomposes into 1D orthogonal block filters via 2D Kronecker products, allowing transformation of spatial dimensions independently.
    4. Colorization: Channel averaging applies an identical 1×31 \times 3 operator across all spatial locations, allowing per-pixel orthogonal transforms with O(1)O(1) extra memory.
  8. Knowl 8 — ImageNet Restoration Performance under Noiseless Conditions

    data/table

    Evaluation of unsupervised restoration methods on 256×256256 \times 256 ImageNet validation images (1 image per each of the 1000 classes) for noiseless 4×4\times super-resolution (block averaging) and uniform deblurring (9×99 \times 9 kernel). DDRM achieves state-of-the-art PSNR and SSIM within 20 neural function evaluations (NFEs), providing up to a 50×50\times reduction in NFEs compared to sampling baselines.

    Method 4×4\times super-resolution Deblurring
    PSNR↑\uparrow SSIM↑\uparrow KID↓\downarrow NFEs↓\downarrow PSNR↑\uparrow SSIM↑\uparrow KID↓\downarrow NFEs↓\downarrow
    Baseline 25.65 0.71 44.90 0 19.26 0.48 38.00 0
    DGP 23.06 0.56 21.22 1500 22.70 0.52 27.60 1500
    RED 26.08 0.73 53.55 100 26.16 0.76 21.21 500
    SNIPS 17.58 0.22 35.17 1000 34.32 0.87 0.49 1000
    DDRM 26.55 0.72 7.22 20 35.64 0.95 0.71 20
    DDRM-CC 26.55 0.74 6.56 20 35.65 0.96 0.70 20

    KID values are multiplied by 10310^3. Baseline refers to bicubic upscaling for super-resolution and the blurry image for deblurring. DDRM-CC incorporates ground-truth class labels.

  9. Knowl 9 — ImageNet Restoration Performance under Measurement Noise

    data/table

    Evaluation of unsupervised restoration methods on 256×256256 \times 256 ImageNet validation images (1 image per each of the 1000 classes) under additive Gaussian measurement noise with σy=0.05\sigma_y = 0.05. Optimization- and projection-based baselines (DGP, RED, SNIPS) degrade substantially in the presence of noise, whereas DDRM maintains high fidelity and perceptual quality at 20 NFEs.

    Method 4×4\times super-resolution Deblurring
    PSNR↑\uparrow SSIM↑\uparrow KID↓\downarrow NFEs↓\downarrow PSNR↑\uparrow SSIM↑\uparrow KID↓\downarrow NFEs↓\downarrow
    Baseline 22.55 0.46 67.86 0 18.35 0.20 75.50 0
    DGP 20.69 0.43 42.17 1500 21.20 0.45 34.02 1500
    RED 22.90 0.49 43.45 100 14.69 0.08 121.82 500
    SNIPS 16.30 0.14 67.77 1000 16.37 0.14 77.96 1000
    DDRM 25.21 0.66 12.43 20 25.45 0.66 15.24 20
    DDRM-CC 25.22 0.67 10.82 20 25.46 0.67 13.49 20

    KID values are multiplied by 10310^3. Measurements have additive Gaussian noise σy=0.05\sigma_y = 0.05 added before reconstruction.

  10. Knowl 10 — Methodological Limitations of DDRM

    limitation

    DDRM has several structural limitations:

    1. Linear Degradation Restriction: The formulation strictly assumes a linear degradation operator H\mathbf{H} with a computationally tractable Singular Value Decomposition (SVD); it is not directly applicable to non-linear inverse problems.
    2. Known Forward Model and Noise Level: The exact degradation matrix H\mathbf{H} and measurement noise variance σy2\sigma_y^2 must be known a priori at inference time, precluding blind restoration without explicit extensions.
    3. Heuristic Hyperparameter Tuning: Hyperparameters including transition stochasticity η\eta, scaling factor ηb\eta_b, and the step-skipping schedule are selected via empirical heuristics rather than derived adaptively or learned end-to-end.

Coverage note — None was omitted; all primary theoretical formulations, equivalence theorems, algorithmic definitions, empirical results on ImageNet, memory optimizations, and stated limitations were included.

References

  1. 1.Richard G Baraniuk. Compressive sensing [lecture notes]. IEEE signal processing magazine, 24(4):118–121, 2007.
  2. 2.Johnathan M Bardsley. Mcmc-based image reconstruction with uncertainty quantification. SIAM Journal on Scientific Computing, 34(3):A1316–A1332, 2012.
  3. 3.Johnathan M Bardsley, Antti Solonen, Heikki Haario, and Marko Laine. Randomize-then-optimize: A method for sampling from posterior distributions in nonlinear inverse problems. SIAM Journal on Scientific Computing, 36(4):A1895–A1910, 2014.
  4. 4.Christopher M Bishop. Pattern recognition. Machine learning, 128(9), 2006.
  5. 5.Mikołaj Bińkowski, Danica J. Sutherland, Michael Arbel, and Arthur Gretton. Demystifying MMD GANs. In International Conference on Learning Representations, 2018.
  6. 6.Yochai Blau and Tomer Michaeli. The perception-distortion tradeoff. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6228–6237, 2018.
  7. 7.Ashish Bora, Ajil Jalal, Eric Price, and Alexandros G. Dimakis. Compressed sensing using generative models. In Proceedings of the 34th International Conference on Machine Learning, volume 70, pages 537–546, 2017.
  8. 8.Daniela Calvetti and Erkki Somersalo. Hypermodels in the bayesian imaging framework. Inverse Problems, 24(3):034013, 2008.
  9. 9.Jooyoung Choi, Sungwon Kim, Yonghyun Jeong, Youngjune Gwon, and Sungroh Yoon. Ilvr: Conditioning method for denoising diffusion probabilistic models. arXiv preprint arXiv:2108.02938, 2021.
  10. 10.Hyungjin Chung, Byeongsu Sim, and Jong Chul Ye. Come-closer-diffuse-faster: Accelerating conditional diffusion models for inverse problems through stochastic contraction. arXiv preprint arXiv:2112.05146, 2021.
  11. 11.Giannis Daras, Joseph Dean, Ajil Jalal, and Alex Dimakis. Intermediate layer optimization for inverse problems using deep generative models. In Proceedings of the 38th International Conference on Machine Learning, volume 139, pages 2421–2432, 2021.
  12. 12.Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. ImageNet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition, pages 248–255, 2009.
  13. 13.Prafulla Dhariwal and Alexander Quinn Nichol. Diffusion models beat GANs on image synthesis. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  14. 14.Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. Image super-resolution using deep convolutional networks. IEEE transactions on pattern analysis and machine intelligence, 38(2):295–307, 2015.
  15. 15.Jinjin Gu, Yujun Shen, and Bolei Zhou. Image processing using multi-code gan prior. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3012–3021, 2020.
  16. 16.Bichuan Guo, Yuxing Han, and Jiangtao Wen. Agem: Solving linear inverse problems via deep priors and sampling. Advances in Neural Information Processing Systems, 32, 2019.
  17. 17.Muhammad Haris, Gregory Shakhnarovich, and Norimichi Ukita. Deep back-projection networks for super-resolution. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1664–1673, 2018.
  18. 18.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, volume 30, 2017.
  19. 19.Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems, volume 33, pages 6840–6851, 2020.
  20. 20.Ajil Jalal, Marius Arvinte, Giannis Daras, Eric Price, Alex Dimakis, and Jonathan Tamir. Robust compressed sensing mri with deep generative priors. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  21. 21.Ajil Jalal, Sushrut Karmalkar, Alex Dimakis, and Eric Price. Instance-optimal compressed sensing via posterior sampling. In Proceedings of the 38th International Conference on Machine Learning, volume 139, pages 4709–4720, 2021.
  22. 22.Zahra Kadkhodaie and Eero Simoncelli. Stochastic solutions for linear inverse problems using the prior implicit in a denoiser. Advances in Neural Information Processing Systems, 34, 2021.
  23. 23.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of GANs for improved quality, stability, and variation. In International Conference on Learning Representations, 2018.
  24. 24.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4401–4410, 2019.
  25. 25.Bahjat Kawar, Gregory Vaksman, and Michael Elad. SNIPS: Solving noisy inverse problems stochastically. In Thirty-Fifth Conference on Neural Information Processing Systems, 2021.
  26. 26.Bahjat Kawar, Gregory Vaksman, and Michael Elad. Stochastic image denoising by sampling from the posterior distribution. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, pages 1866–1875, October 2021.
  27. 27.Diederik P Kingma and Max Welling. Auto-Encoding variational bayes. arXiv preprint arXiv:1312.6114v10, December 2013.
  28. 28.Orest Kupyn, Tetiana Martyniuk, Junru Wu, and Zhangyang Wang. Deblurgan-v2: Deblurring (orders-of-magnitude) faster and better. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8878–8887, 2019.
  29. 29.Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Learning representations for automatic colorization. In European conference on computer vision, pages 577–593. Springer, 2016.
  30. 30.Remi Laumont, Valentin De Bortoli, Andres Almansa, Julie Delon, Alain Durmus, and Marcelo Pereyra. Bayesian imaging using plug & play priors: When Langevin meets Tweedie. arXiv preprint arXiv:2103.04715, 2021.
  31. 31.Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photo-realistic single image super-resolution using a generative adversarial network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4681–4690, 2017.
  32. 32.Gary Mataev, Peyman Milanfar, and Michael Elad. DeepRED: deep image prior powered by RED. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2019.
  33. 33.Sachit Menon, Alex Damian, McCourt Hu, Nikhil Ravi, and Cynthia Rudin. Pulse: Self-supervised photo upsampling via latent space exploration of generative models. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2020.
  34. 34.Chris Metzler, Ali Mousavi, and Richard Baraniuk. Learned d-amp: Principled neural network based compressive image recovery. In Advances in Neural Information Processing Systems, volume 30, 2017.
  35. 35.Alex Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. arXiv preprint arXiv:2102.09672, 2021.
  36. 36.Gregory Ongie, Ajil Jalal, Christopher A Metzler, Richard G Baraniuk, Alexandros G Dimakis, and Rebecca Willett. Deep learning techniques for inverse problems in imaging. IEEE Journal on Selected Areas in Information Theory, 1(1):39–56, 2020.
  37. 37.Gregory Ongie, Ajil Jalal, Christopher A Metzler, Richard G Baraniuk, Alexandros G Dimakis, and Rebecca Willett. Deep learning techniques for inverse problems in imaging. IEEE Journal on Selected Areas in Information Theory, 1(1):39–56, 2020.
  38. 38.Xingang Pan, Xiaohang Zhan, Bo Dai, Dahua Lin, Chen Change Loy, and Ping Luo. Exploiting deep generative prior for versatile image restoration and manipulation. In European Conference on Computer Vision (ECCV), 2020.
  39. 39.JH Rick Chang, Chun-Liang Li, Barnabas Poczos, BVK Vijaya Kumar, and Aswin C Sankaranarayanan. One network to solve them all–solving linear inverse problems using deep projection models. In Proceedings of the IEEE International Conference on Computer Vision, pages 5888–5897, 2017.
  40. 40.Yaniv Romano, Michael Elad, and Peyman Milanfar. The little engine that could: Regularization by denoising (RED). SIAM Journal on Imaging Sciences, 10(4):1804–1844, 2017.
  41. 41.Chitwan Saharia, William Chan, Huiwen Chang, Chris A Lee, Jonathan Ho, Tim Salimans, David J Fleet, and Mohammad Norouzi. Palette: Image-to-image diffusion models. arXiv preprint arXiv:2111.05826, 2021.
  42. 42.Chitwan Saharia, Jonathan Ho, William Chan, Tim Salimans, David J Fleet, and Mohammad Norouzi. Image super-resolution via iterative refinement. arXiv preprint arXiv:2104.07636, 2021.
  43. 43.Shibani Santurkar, Dimitris Tsipras, Brandon Tran, Andrew Ilyas, Logan Engstrom, and Aleksander Madry. Image synthesis with a single (robust) classifier. arXiv preprint arXiv:1906.09453, 2019.
  44. 44.Jascha Sohl-Dickstein, Eric A Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. arXiv preprint arXiv:1503.03585, March 2015.
  45. 45.Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In International Conference on Learning Representations, April 2021.
  46. 46.Yang Song, Liyue Shen, Lei Xing, and Stefano Ermon. Solving inverse problems in medical imaging with score-based generative models. arXiv preprint arXiv:2111.08005, 2021.
  47. 47.Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equations. In International Conference on Learning Representations, 2021.
  48. 48.Maitreya Suin, Kuldeep Purohit, and AN Rajagopalan. Spatially-attentive patch-hierarchical network for adaptive motion deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3606–3615, 2020.
  49. 49.Yu Sun, Brendt Wohlberg, and Ulugbek S Kamilov. An online plug-and-play algorithm for regularized image reconstruction. IEEE Transactions on Computational Imaging, 5(3):395–408, 2019.
  50. 50.Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. Deep image prior. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 9446–9454, 2018.
  51. 51.Singanallur V Venkatakrishnan, Charles A Bouman, and Brendt Wohlberg. Plug-and-play priors for model based reconstruction. In 2013 IEEE Global Conference on Signal and Information Processing, pages 945–948. IEEE, 2013.
  52. 52.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004.
  53. 53.Allan G Weber. The USC-SIPI image database version 5. USC-SIPI Report, 315(1), 1997.
  54. 54.Jay Whang, Mauricio Delbracio, Hossein Talebi, Chitwan Saharia, Alexandros G Dimakis, and Peyman Milanfar. Deblurring via stochastic refinement. arXiv preprint arXiv:2112.02475, 2021.
  55. 55.Raymond A Yeh, Chen Chen, Teck Yian Lim, Alexander G Schwing, Mark Hasegawa-Johnson, and Minh N Do. Semantic image inpainting with deep generative models. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5485–5493, 2017.
  56. 56.Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. LSUN: construction of a large-scale image dataset using deep learning with humans in the loop. arXiv preprint arXiv:1506.03365, 2015.
  57. 57.Kai Zhang, Yawei Li, Wangmeng Zuo, Lei Zhang, Luc Van Gool, and Radu Timofte. Plug-and-play image restoration with deep denoiser prior. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021.
  58. 58.Richard Zhang, Phillip Isola, and Alexei A Efros. Colorful image colorization. In European Conference on Computer Vision, pages 649–666. Springer, 2016.

Citation

MLA
Kawar, B., et al. “Denoising Diffusion Restoration Models”. arXiv, 2022, http://arxiv.org/abs/2201.11793v3.
APA
Kawar, B., Elad, M., Ermon, S., & Song, J. (2022). Denoising Diffusion Restoration Models. arXiv. http://arxiv.org/abs/2201.11793v3
Chicago
Kawar, B., M. Elad, S. Ermon, and J. Song. 2022. “Denoising Diffusion Restoration Models”. arXiv. http://arxiv.org/abs/2201.11793v3.
Harvard
Kawar, B. et al. (2022) “Denoising Diffusion Restoration Models”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2201.11793v3.
Vancouver
1. Kawar B, Elad M, Ermon S, Song J (2022) Denoising Diffusion Restoration Models. arXiv

BibTeX

@article{kawar2022denoising,
  title = {Denoising Diffusion Restoration Models},
  author = {Kawar, Bahjat and Elad, Michael and Ermon, Stefano and Song, Jiaming},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2201.11793v3},
  eprint = {2201.11793}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: Authors