Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Xinyu PengZiyang ZhengWenrui DaiNuoqian XiaoChenglin LiJunni ZouHongkai Xiong

article2024ICML62 citations

Unifies zero-shot diffusion solvers for inverse problems under a posterior covariance framework and derives maximum-likelihood-optimized covariance estimators that boost image reconstruction quality without manual hyperparameter tuning.

Listen

Diffusion models have emerged as powerful tools for solving noisy linear inverse problems—such as image inpainting, deblurring, and super-resolution—in a "zero-shot" setting that requires no task-specific model retraining. However, existing zero-shot solvers rely on disparate formulations and hand-crafted heuristics that frequently require tedious hyperparameter tuning and yield unstable or suboptimal reconstructions.

The main objective of the article is to establish a unified mathematical framework for existing diffusion-based inverse solvers and demonstrate that replacing heuristic approximations with optimal posterior covariance derived via maximum likelihood estimation significantly enhances reconstruction accuracy without manual tuning.

The authors evaluated their approach across standard benchmarks, including the FFHQ face dataset and the ImageNet dataset, testing multiple image restoration tasks under Gaussian noise. They demonstrated how existing methods uniformly approximate the intractable denoising posterior using simple isotropic Gaussian distributions. To improve upon this, the authors introduced plug-and-play strategies that derive optimal posterior covariance either by analytically converting existing model variances or by using offline Monte Carlo error estimates. Furthermore, they introduced a scalable method to learn covariance predictions directly in a wavelet transform domain, reducing computational complexity from quadratic to linear.

The key findings are as follows:

  1. Principled covariance estimation significantly outperforms existing methods across restoration tasks; for example, on ImageNet motion deblurring, the proposed conversion approach achieved a Fréchet Inception Distance (FID) of 109.61 compared to 282.21 for Diffusion Posterior Sampling (DPS) and 302.40 for Tweedie Moment Projected Diffusion (TMPD).
  2. The proposed transform-domain approach (DWT-Var) outperforms optimally tuned baselines without requiring any hyperparameter search, eliminating the sensitivity seen in competing methods.
  3. Existing heuristic methods (such as guidance step adjustments) also see substantial quality gains simply by replacing their final low-noise sampling steps with the proposed optimal variance estimates.
  4. Analytical conversion of pre-trained reverse variances proves highly accurate at lower noise levels, whereas high noise levels suffer from numerical limits where boundaries collapse.

These findings indicate that generative image restoration can achieve superior fidelity and robustness at minimal additional computational cost. By eliminating manual hyperparameter sweeps, organizations deploying diffusion pipelines for high-resolution imaging can streamline deployment workflows, reduce experimental overhead, and achieve higher consistency in downstream image reconstruction tasks.

Practitioners should integrate the conversion or Monte Carlo covariance estimators into existing pre-trained diffusion pipelines when deploying zero-shot solvers, applying spatial variance corrections primarily during low-noise sampling steps. When training new diffusion architectures, teams should consider predicting variances directly in an orthonormal transform domain (such as wavelets) to better capture pixel correlations.

Confidence in these findings is supported by rigorous mathematical derivations and extensive multi-metric empirical benchmarks. However, readers should note that the current implementation restricts covariance matrices to diagonal forms to maintain computational feasibility, which leaves some non-diagonal pixel correlations unmodeled. Future work should investigate non-linear transformations and advanced approximations to capture full covariance structures.

Cover for Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance

Abstract

Recent diffusion models provide a promising zero-shot solution to noisy linear inverse problems without retraining for specific inverse problems. In this paper, we reveal that recent methods can be uniformly interpreted as employing a Gaussian approximation with hand-crafted isotropic covariance for the intractable denoising posterior to approximate the conditional posterior mean. Inspired by this finding, we propose to improve recent methods by using more principled covariance determined by maximum likelihood estimation. To achieve posterior covariance optimization without retraining, we provide general plug-and-play solutions based on two approaches specifically designed for leveraging pre-trained models with and without reverse covariance. We further propose a scalable method for learning posterior covariance prediction based on representation with orthonormal basis. Experimental results demonstrate that the proposed methods significantly enhance reconstruction performance without requiring hyperparameter tuning.

Table of Contents

  • 1 Introduction
  • 2 Background
  • 2.1 Bayesian Framework for Solving Inverse Problems
  • 2.2 Diffusion Models and Conditioning
  • 3 Unified Interpretation of Diffusion-based Solvers to Inverse Problems
  • 3.1 Type I Guidance: Approximating the Likelihood Score Function
  • 3.2 Type II Guidance: Approximating the Conditional Posterior Mean Using Proximal Solution
  • 3.3 Solving Inverse Problems with Optimal Posterior Covariance
  • 4 Posterior Covariance Optimization
  • 4.1 Converting Optimal Reverse Variances
  • 4.2 Monte Carlo Estimation of Posterior Variances
  • 4.3 Modeling Pixel-Correlations With Latent Variances
  • 5 Related Work
  • 6 Experiments
  • 6.1 Sanity Check for Converting Reverse Variance
  • 6.2 Quantitative Results
  • 7 Conclusions
  • References
  • A Proofs
  • A.1 Derivation of the Marginal Preserving Property of Diffusion ODEs
  • A.2 Proof of Proposition
  • A.3 Proof of Proposition
  • A.4 Proof of Theorem
  • A.5 Proof of Proposition
  • B Practical Numerical Algorithms
  • B.1 Using Closed-form Solutions for Isotropic Posterior Covariance
  • B.2 Using Conjugate Gradient Method for General Posterior Covariance
  • C Additional Experimental Details and Results
  • C.1 Additional Quantitative Results
  • C.2 Training Objective for Latent Variance
  • C.3 DDNM as Noiseless DiffPIR
  • C.4 Converting Optimal Solutions Between Different Perturbation Kernels
  • C.5 Qualitative Results

Knowls

  1. Knowl 1 — A shared Gaussian-posterior interpretation of zero-shot inverse solvers

    model/method

    For a noisy linear inverse problem, measurements satisfy y=Ax0+ny=Ax_0+n, where x0∈Rdx_0\in\mathbb{R}^d is the unknown image, AA is the measurement operator, and n∼N(0,σ2I)n\sim\mathcal{N}(0,\sigma^2I) is measurement noise. Let Dt(xt)D_t(x_t) be an unconditional diffusion denoiser and pt(x0∣xt)p_t(x_0\mid x_t) the intractable denoising posterior. The paper interprets several zero-shot solvers as approximating that posterior by qt(x0∣xt)=N(Dt(xt),rt2I)q_t(x_0\mid x_t)=\mathcal{N}(D_t(x_t),r_t^2I) and then using the measurement likelihood to approximate the conditional posterior mean.

    The methods differ primarily in their chosen isotropic variance: DPS is the limiting point-mass case rt→0r_t\to0; ΠGDM uses rt2=σt2/(1+σt2)r_t^2=\sigma_t^2/(1+\sigma_t^2) under the prior assumption x0∼N(0,I)x_0\sim\mathcal{N}(0,I); DiffPIR corresponds to rt=σt/λr_t=\sigma_t/\sqrt{\lambda}, where λ\lambda is its tuning parameter; and DDNM is recovered as the zero-measurement-noise limit for any fixed rt>0r_t>0. Thus, methods motivated by likelihood-score guidance and methods motivated by proximal or null-space updates can be understood within the same posterior-approximation framework.

    For Gaussian perturbation kernels pt(xt∣x0)=N(stx0,st2σt2I)p_t(x_t\mid x_0)=\mathcal{N}(s_tx_0,s_t^2\sigma_t^2I), the conditional posterior mean also satisfies

    E[x0∣xt,y]=E[x0∣xt]+stσt2∇xtlog⁡pt(y∣xt).\mathbb{E}[x_0\mid x_t,y]=\mathbb{E}[x_0\mid x_t]+s_t\sigma_t^2\nabla_{x_t}\log p_t(y\mid x_t).

    Here sts_t and σt\sigma_t are the perturbation scale and noise standard deviation, respectively. This identity links conditional-mean estimation to likelihood-score guidance.

  2. Knowl 2 — Maximum-likelihood learning of posterior covariance

    model/method

    The paper replaces hand-chosen isotropic uncertainty with a Gaussian approximation qt(x0∣xt)=N(Dt(xt),Σt(xt))q_t(x_0\mid x_t)=\mathcal{N}(D_t(x_t),\Sigma_t(x_t)) to the denoising posterior pt(x0∣xt)p_t(x_0\mid x_t). It selects the approximation by minimizing a time-weighted expected forward Kullback–Leibler divergence, equivalently by maximizing the conditional Gaussian log-likelihood. With a full-rank covariance, the optimal mean is the MMSE denoiser, Dt(xt)=E[x0∣xt]D_t(x_t)=\mathbb{E}[x_0\mid x_t]. With a diagonal covariance Σt(xt)=diag⁡(rt2(xt))\Sigma_t(x_t)=\operatorname{diag}(r_t^2(x_t)), each covariance entry is the conditional squared error of the corresponding denoiser coordinate:

    rt,i∗2(xt)=E ⁣[(x0,i−E[x0,i∣xt])2∣xt].r_{t,i}^{*2}(x_t)=\mathbb{E}\!\left[(x_{0,i}-\mathbb{E}[x_{0,i}\mid x_t])^2\mid x_t\right].

    In this expression, ii indexes image coordinates, rt,i∗2r_{t,i}^{*2} is a variance, and the covariance is diagonal in the coordinate system in which it is parameterized. The optimized covariance is then used with a pretrained denoiser to condition sampling, without training a separate inverse-problem model.

  3. Knowl 3 — Using the learned covariance for Type I and Type II guidance

    model/method

    Let y=Ax0+ny=Ax_0+n with n∼N(0,σ2I)n\sim\mathcal{N}(0,\sigma^2I), and let a pretrained diffusion denoiser and learned posterior covariance be Dt(xt)D_t(x_t) and Σt(xt)\Sigma_t(x_t). The Gaussian approximation to the measurement likelihood is

    pt(y∣xt)≈N ⁣(ADt(xt),  σ2I+AΣt(xt)AT).p_t(y\mid x_t)\approx\mathcal{N}\!\left(AD_t(x_t),\;\sigma^2I+A\Sigma_t(x_t)A^T\right).

    Type I guidance uses the gradient of this approximate log-likelihood with respect to xtx_t to modify the diffusion trajectory. Type II guidance substitutes the approximate conditional posterior mean for the unconditional denoiser output; equivalently, it solves

    x^0(t)=arg⁡min⁡x0{∥y−Ax0∥22+σ2(x0−Dt(xt))TΣt(xt)−1(x0−Dt(xt))}.\hat{x}_0^{(t)}=\arg\min_{x_0}\left\{\|y-Ax_0\|_2^2+\sigma^2(x_0-D_t(x_t))^T\Sigma_t(x_t)^{-1}(x_0-D_t(x_t))\right\}.

    For positive-definite Σt(xt)\Sigma_t(x_t), this solution can be evaluated as

    x^0(t)=Dt(xt)+Σt(xt)AT(σ2I+AΣt(xt)AT)−1(y−ADt(xt)).\hat{x}_0^{(t)}=D_t(x_t)+\Sigma_t(x_t)A^T\bigl(\sigma^2I+A\Sigma_t(x_t)A^T\bigr)^{-1}\bigl(y-AD_t(x_t)\bigr).

    The shared linear solve can be computed with conjugate gradients when a closed form is unavailable; the paper uses tolerance 10−410^{-4} for its general-covariance experiments. For isotropic covariance Σt=rt2I\Sigma_t=r_t^2I, the paper also gives closed-form implementations for inpainting, circular-convolution deblurring, and super-resolution. The chosen conditional estimate replaces the unconditional denoiser estimate in the diffusion sampling update.

  4. Knowl 4 — Recovering posterior variance from a DDPM reverse-variance prediction

    equation

    For a DDPM with forward schedule βt\beta_t, define αt=1−βt\alpha_t=1-\beta_t, αˉt=∏j=1tαj\bar\alpha_t=\prod_{j=1}^t\alpha_j, βˉt=1−αˉt\bar\beta_t=1-\bar\alpha_t, and β~t=(βˉt−1/βˉt)βt\tilde\beta_t=(\bar\beta_{t-1}/\bar\beta_t)\beta_t. If the reverse transition covariance is diagonal, the fixed-point relationship between its optimal variance vt∗2(xt)v_t^{*2}(x_t) and the optimal diagonal denoising-posterior variance rt∗2(xt)r_t^{*2}(x_t) is

    vt∗2(xt)=β~t+(αˉt−1 βt1−αˉt)2rt∗2(xt).v_t^{*2}(x_t)=\tilde\beta_t+\left(\frac{\sqrt{\bar\alpha_{t-1}}\,\beta_t}{1-\bar\alpha_t}\right)^2r_t^{*2}(x_t).

    Consequently, a pretrained DDPM reverse-variance prediction v^t2(xt)\hat v_t^2(x_t) gives the plug-and-play estimate

    r^t2(xt)=(v^t2(xt)−β~t)(αˉt−1 βt1−αˉt)−2.\hat r_t^2(x_t)=\bigl(\hat v_t^2(x_t)-\tilde\beta_t\bigr)\left(\frac{\sqrt{\bar\alpha_{t-1}}\,\beta_t}{1-\bar\alpha_t}\right)^{-2}.

    This conversion assumes the fixed-point relationship is a good approximation for the pretrained model. In the paper’s empirical check, the converted variances predict denoising squared errors reliably mainly at low noise levels. At high noise, the optimal reverse variance approaches its schedule bounds, making the subtraction in the conversion nearly zero and the estimate numerically unstable.

  5. Knowl 5 — Monte Carlo estimation when reverse variances are unavailable

    algorithm

    When a pretrained denoiser does not provide reverse-variance predictions, the paper estimates an isotropic, time-dependent posterior variance from denoising errors. For a data dimension dd, the optimum is the expected per-coordinate squared reconstruction error:

    rt∗2=1d Ept(x0,xt)[∥x0−Dt(xt)∥22].r_t^{*2}=\frac{1}{d}\,\mathbb{E}_{p_t(x_0,x_t)}\left[\|x_0-D_t(x_t)\|_2^2\right].

    The authors implement this estimate offline: use 5% of the training dataset to form the empirical mean squared denoising error at 1,000 discrete noise-time steps; at inference, use the stored value at the discrete time nearest the current tt. This produces a plug-and-play isotropic variance and requires no additional covariance prediction during inverse-problem sampling.

  6. Knowl 6 — Scalable covariance prediction in an orthonormal transform basis

    model/method

    To represent pixel correlations without predicting a dense covariance, the paper assumes images can be expressed as x0=Ψθ0x_0=\Psi\theta_0, where Ψ\Psi is an orthonormal transform and θ0\theta_0 contains transform coefficients. It parameterizes the covariance as

    Σt(xt)=Ψdiag⁡(rt2(xt))ΨT.\Sigma_t(x_t)=\Psi\operatorname{diag}(r_t^2(x_t))\Psi^T.

    The approach treats the coefficients of θ0\theta_0 as mutually independent, so the conditional covariance is diagonal in the transform domain. It predicts dd variances rather than the d2d^2 entries of a dense covariance. Training maximizes the Gaussian posterior likelihood in this domain: the quadratic error is between ΨTx0\Psi^Tx_0 and ΨTDt(xt)\Psi^TD_t(x_t), weighted coefficientwise by reciprocal predicted variances, with the log-variance term included. The experiments use the discrete wavelet transform (DWT) for Ψ\Psi. The authors report that DWT variances can be used over a wider range of noise levels than spatial diagonal variances, although they restrict DWT-variance use to σt<1\sigma_t<1 for computational efficiency when guidance requires conjugate-gradient solves.

  7. Knowl 7 — Evaluation protocol for image inverse problems

    experimental setup

    The experiments use unconditional diffusion models on FFHQ 256×256 and ImageNet 256×256, and evaluate random inpainting, Gaussian-kernel deblurring, motion-blur deblurring, and 4× super-resolution. Inpainting masks 50% of pixels; super-resolution uses bicubic downsampling; the blur-kernel setups follow prior work cited by the paper. Every measurement has additive Gaussian noise with standard deviation σ=0.05\sigma=0.05. Results are measured by SSIM (higher is better), LPIPS (lower is better), and FID (lower is better).

    Type I methods use a 50-step deterministic Heun sampler. Type II methods use a 50-step stochastic Heun sampler with Schurn=80S_{\mathrm{churn}}=80, Smin⁡=0.05S_{\min}=0.05, Smax⁡=50S_{\max}=50, and Snoise=1.003S_{\mathrm{noise}}=1.003. For spatial plug-and-play variances, the authors use those variances only in the final 12 of 50 steps, where σt<0.2\sigma_t<0.2, and use ΠGDM covariance at earlier, noisier steps. DWT variances are used for σt<1\sigma_t<1; the restriction avoids the additional conjugate-gradient cost at larger noise levels.

  8. Knowl 8 — Type I guidance benchmark results

    data/table

    The quantitative comparison on page 7 evaluates Type I guidance on the four inverse tasks using the protocol described here: FFHQ and ImageNet at 256×256, measurement noise σ=0.05\sigma=0.05, and 50-step deterministic Heun sampling. Each cell lists SSIM (higher is better), LPIPS (lower is better), and FID (lower is better), in that order. Convert, Analytic, and DWT-Var are the paper’s covariance methods; TMPD, DPS, and ΠGDM are comparison methods. The results show that the proposed covariance estimates are competitive across tasks, while DPS has strong results on some ImageNet super-resolution metrics; no one proposed method wins every entry.

    Dataset Method Inpaint Gaussian deblur Motion deblur 4×\times SR
    SSIM LPIPS FID SSIM LPIPS FID SSIM LPIPS FID SSIM LPIPS FID
    FFHQ Convert 0.9279 0.0794 25.90 0.7905 0.1836 52.42 0.7584 0.2156 62.88 0.7878 0.1962 58.37
    FFHQ Analytic 0.9272 0.0845 28.83 0.7926 0.1850 53.09 0.7579 0.2183 64.77 0.7878 0.1968 59.83
    FFHQ DWT-Var 0.9209 0.0863 28.54 0.7968 0.1837 57.52 0.7677 0.2103 65.34 0.8025 0.1856 57.26
    FFHQ TMPD 0.8224 0.1924 70.92 0.7289 0.2523 76.52 0.7014 0.2718 83.49 0.7085 0.2701 79.58
    FFHQ DPS 0.8891 0.1323 49.46 0.6284 0.3652 136.12 0.4904 0.4924 212.48 0.7719 0.2054 61.36
    FFHQ Π\PiGDM 0.8784 0.1422 49.89 0.7890 0.1910 59.93 0.7543 0.2209 66.14 0.7850 0.2005 61.46
    ImageNet Convert 0.8559 0.1329 29.14 0.6007 0.3327 95.23 0.5634 0.3656 109.61 0.5869 0.3477 96.76
    ImageNet Analytic 0.8481 0.1446 35.51 0.6009 0.3334 93.21 0.5611 0.3668 113.39 0.5958 0.3495 95.33
    ImageNet TMPD 0.7011 0.2892 293.80 0.5430 0.4114 291.29 0.4773 0.4567 302.40 0.5186 0.4298 296.73
    ImageNet DPS 0.8623 0.1490 36.58 0.4603 0.4630 173.77 0.3582 0.5554 282.21 0.5860 0.3231 92.89
    ImageNet Π\PiGDM 0.7658 0.2328 64.96 0.5946 0.3429 102.89 0.5534 0.3781 113.89 0.5925 0.3552 100.36
  9. Knowl 9 — Type II and adaptive-guidance results

    empirical result

    In Type II comparisons on FFHQ, the paper reports that its plug-and-play covariance methods perform comparably to DiffPIR when DiffPIR is optimally tuned over its parameter λ\lambda, while DWT-Var outperforms optimally tuned DiffPIR without requiring such tuning. Spatial plug-and-play variances are applied only during low-noise sampling steps; DWT-Var is usable over a wider range, subject to the stated computational cutoff.

    The page-8 quantitative results also test ΠGDM with its adaptive likelihood-score weight, using the same four tasks and reporting SSIM, LPIPS, and FID. Each entry below gives those metrics in that order; SSIM is better when larger, and LPIPS and FID are better when smaller. Convert and Analytic improve over adaptive ΠGDM in every listed metric and task for both datasets.

    Dataset Method Inpaint Gaussian deblur Motion deblur 4×\times SR
    SSIM LPIPS FID SSIM LPIPS FID SSIM LPIPS FID SSIM LPIPS FID
    FFHQ Convert 0.9241 0.0822 27.50 0.7783 0.1969 59.31 0.7329 0.2324 66.18 0.7632 0.2183 67.17
    FFHQ Analytic 0.9232 0.0852 28.63 0.7790 0.1971 59.80 0.7331 0.2336 69.82 0.7622 0.2195 68.84
    FFHQ Π\PiGDM 0.7078 0.2605 77.46 0.7221 0.2421 71.19 0.6977 0.2607 75.15 0.7205 0.2442 72.41
    ImageNet Convert 0.8492 0.1394 33.46 0.5770 0.3568 100.50 0.5341 0.3944 131.48 0.5613 0.3846 116.27
    ImageNet Analytic 0.8417 0.1484 37.98 0.5722 0.3583 103.67 0.5341 0.3947 126.39 0.5537 0.3867 116.87
    ImageNet Π\PiGDM 0.5102 0.4293 141.80 0.5071 0.4095 127.41 0.4887 0.4267 132.38 0.5150 0.4098 127.40
  10. Knowl 10 — Limits of diagonal covariance and reverse-variance conversion

    limitation

    The proposed covariance parameterizations are diagonal either in image coordinates or in a chosen transform basis. The paper notes that this constraint prevents the covariance from matching the unrestricted optimum, even if the denoiser and covariance predictor are perfectly trained. A further practical limitation applies to conversion from DDPM reverse variances: the empirical estimates are reliable mainly at low noise and become numerically unstable at high noise because the reverse variance is close to its schedule bounds. The paper identifies richer covariance structures, nonlinear transforms, and more efficient approximations as directions for future work.

Coverage note — Omitted proof-only derivations, task-specific Fourier closed forms, and the appendix’s perturbation-kernel conversion procedure because they support implementation but are less central than the shared posterior interpretation, covariance methods, and evaluated results.

References

  1. 1.Balle, J., Chou, P. A., Minnen, D., Singh, S., Johnston, N., Agustsson, E., Hwang, S. J., and Toderici, G. Nonlinear transform coding. IEEE Journal of Selected Topics in Signal Processing, 15(2):339–353, 2020.
  2. 2.Bao, F., Li, C., Sun, J., Zhu, J., and Zhang, B. Estimating the optimal covariance with imperfect mean in diffusion probabilistic models. In Proceedings of the 39th International Conference on Machine Learning, pp. 1555–1584, 2022a.
  3. 3.Bao, F., Li, C., Zhu, J., and Zhang, B. Analytic-DPM: an analytic estimate of the optimal reverse variance in diffusion probabilistic models. In The 10th International Conference on Learning Representations, 2022b. URL https://openreview.net/forum?id=0xiJLKH-ufZ.
  4. 4.Bishop, C. M. Pattern Recognition and Machine Learning. Springer, 2006.
  5. 5.Boys, B., Girolami, M., Pidstrigach, J., Reich, S., Mosca, A., and Akyildiz, O. D. Tweedie moment projected diffusions for inverse problems. arXiv preprint arXiv:2310.06721, 2023.
  6. 6.Chan, M. A., Young, S. I., and Metzler, C. A. SUD2: Supervision by denoising diffusion models for image reconstruction. In NeurIPS 2023 Deep Inverse Workshop, 2023.
  7. 7.Choi, J., Kim, S., Jeong, Y., Gwon, Y., and Yoon, S. ILVR: Conditioning method for denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 14347–14356, 2021.
  8. 8.Chung, H., Kim, J., Mccann, M. T., Klasky, M. L., and Ye, J. C. Diffusion posterior sampling for general noisy inverse problems. In The 11th International Conference on Learning Representations, 2023a. URL https://openreview.net/forum?id=OnD9zGAGT0k.
  9. 9.Chung, H., Lee, S., and Ye, J. C. Fast diffusion sampler for inverse problems by geometric decomposition. arXiv preprint arXiv:2303.05754, 2023b.
  10. 10.Dhariwal, P. and Nichol, A. Diffusion models beat GANs on image synthesis. In Advances in Neural Information Processing Systems 34, pp. 8780–8794, 2021.
  11. 11.Dorta, G., Vicente, S., Agapito, L., Campbell, N. D., and Simpson, I. Structured uncertainty prediction networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5477–5485, 2018.
  12. 12.Feng, B. T., Smith, J., Rubinstein, M., Chang, H., Bouman, K. L., and Freeman, W. T. Score-based diffusion models as principled priors for inverse imaging. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10520–10531, 2023.
  13. 13.Goyal, V. K. Theoretical foundations of transform coding. IEEE Signal Processing Magazine, 18(5):9–21, 2001.
  14. 14.Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. GANs trained by a two time-scale update rule converge to a local Nash equilibrium. In Advances in Neural Information Processing Systems 30, pp. 6626–6637, 2017.
  15. 15.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems 33, pp. 6840–6851, 2020.
  16. 16.Karras, T., Aittala, M., Aila, T., and Laine, S. Elucidating the design space of diffusion-based generative models. In Advances in Neural Information Processing Systems 35, pp. 26565–26577, 2022.
  17. 17.Kingma, D., Salimans, T., Poole, B., and Ho, J. Variational diffusion models. In Advances in Neural Information Processing Systems 34, pp. 21696–21707, 2021.
  18. 18.Li, H., Li, S., Dai, W., Li, C., Zou, J., and Xiong, H. Frequency-aware transformer for learned image compression. In The 12th International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=HKGQDDTuvZ.
  19. 19.Lu, C., Zheng, K., Bao, F., Chen, J., Li, C., and Zhu, J. Maximum likelihood training for score-based diffusion odes by high order denoising score matching. In Proceedings of the 39th International Conference on Machine Learning, pp. 14429–14460, 2022.
  20. 20.Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Timofte, R., and Van Gool, L. Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11461–11471, 2022.
  21. 21.Luo, Z., Gustafsson, F. K., Zhao, Z., Sjolund, J., and Schon, T. B. Refusion: Enabling large-size realistic image restoration with latent-space diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1680–1691, 2023.
  22. 22.Mardani, M., Song, J., Kautz, J., and Vahdat, A. A variational perspective on solving inverse problems with diffusion models. In The 12th International Conference on Learning Representations, 2024.
  23. 23.Meng, C., Song, Y., Li, W., and Ermon, S. Estimating high order gradients of the data distribution by denoising. In Advances in Neural Information Processing Systems 34, pp. 25359–25369, 2021.
  24. 24.Nehme, E., Yair, O., and Michaeli, T. Uncertainty quantification via neural posterior principal components. In Advances in Neural Information Processing Systems 36, pp. 37128–37141, 2023.
  25. 25.Nichol, A. Q. and Dhariwal, P. Improved denoising diffusion probabilistic models. In Proceedings of the 38th International Conference on Machine Learning, pp. 8162–8171, 2021.
  26. 26.Pokle, A., Muckley, M. J., Chen, R. T., and Karrer, B. Training-free linear image inversion via flows. arXiv preprint arXiv:2310.04432, 2023.
  27. 27.Ravula, S., Levac, B., Jalal, A., Tamir, J. I., and Dimakis, A. G. Optimizing sampling patterns for compressed sensing MRI with diffusion generative models. arXiv preprint arXiv:2306.03284, 2023.
  28. 28.Rezende, D. J. and Viola, F. Taming VAEs. arXiv preprint arXiv:1810.00597, 2018.
  29. 29.Rout, L., Chen, Y., Kumar, A., Caramanis, C., Shakkottai, S., and Chu, W.-S. Beyond first-order Tweedie: Solving inverse problems using latent diffusion. arXiv preprint arXiv:2312.00852, 2023.
  30. 30.Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(4):4713–4726, 2022.
  31. 31.Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceeding of the 35th International Conference on Machine Learning, pp. 2256–2265, 2015.
  32. 32.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. In The 9th International Conference on Learning Representations, 2021a. URL https://openreview.net/forum?id=St1giarCHLP.
  33. 33.Song, J., Vahdat, A., Mardani, M., and Kautz, J. Pseudoinverse-guided diffusion models for inverse problems. In The 11th International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=9_gsMA8MRKQ.
  34. 34.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. In Advances in Neural Information Processing Systems 32, pp. 11918–11930, 2019.
  35. 35.Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. Score-based generative modeling through stochastic differential equations. In The 9th International Conference on Learning Representations, 2021b. URL https://openreview.net/forum?id=PxTIG12RRHS.
  36. 36.Song, Y., Shen, L., Xing, L., and Ermon, S. Solving inverse problems in medical imaging with score-based generative models. In The 10th International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=vaRCHVj0uGI.
  37. 37.Wang, Y., Yu, J., and Zhang, J. Zero-shot image restoration using denoising diffusion null-space model. In The 11th International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=mRieQgMtNTQ.
  38. 38.Whang, J., Delbracio, M., Talebi, H., Saharia, C., Dimakis, A. G., and Milanfar, P. Deblurring via stochastic refinement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16293–16303, 2022.
  39. 39.Xiao, Z., Kreis, K., and Vahdat, A. Tackling the generative learning trilemma with denoising diffusion GANs. In The 10th International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=JprM0p-q0Co.
  40. 40.Zhang, K., Gool, L. V., and Timofte, R. Deep unfolding network for image super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3217–3226, 2020.
  41. 41.Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 586–595, 2018.
  42. 42.Zhu, Y., Yang, Y., and Cohen, T. Transformer-based transform coding. In International Conference on Learning Representations, 2022. URL https://openreview.net/forum?id=IDwN6xjHnK8.
  43. 43.Zhu, Y., Zhang, K., Liang, J., Cao, J., Wen, B., Timofte, R., and Van Gool, L. Denoising diffusion models for plug-and-play image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1219–1229, 2023.

Citation

MLA
Peng, X., et al. “Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance”. arXiv, 2024, http://arxiv.org/abs/2402.02149v2.
APA
Peng, X., Zheng, Z., Dai, W., Xiao, N., Li, C., Zou, J., & Xiong, H. (2024). Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance. arXiv. http://arxiv.org/abs/2402.02149v2
Chicago
Peng, X., Z. Zheng, W. Dai, et al. 2024. “Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance”. arXiv. http://arxiv.org/abs/2402.02149v2.
Harvard
Peng, X. et al. (2024) “Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2402.02149v2.
Vancouver
1. Peng X, Zheng Z, Dai W, Xiao N, Li C, Zou J, Xiong H (2024) Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance. arXiv

BibTeX

@article{peng2024improving,
  title = {Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance},
  author = {Peng, Xinyu and Zheng, Ziyang and Dai, Wenrui and Xiao, Nuoqian and Li, Chenglin and Zou, Junni and Xiong, Hongkai},
  year = {2024},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2402.02149v2},
  eprint = {2402.02149}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/