Improving Diffusion Models for Inverse Problems Using Optimal Posterior Covariance
Xinyu PengZiyang ZhengWenrui DaiNuoqian XiaoChenglin LiJunni ZouHongkai Xiong
Unifies zero-shot diffusion solvers for inverse problems under a posterior covariance framework and derives maximum-likelihood-optimized covariance estimators that boost image reconstruction quality without manual hyperparameter tuning.
Diffusion models have emerged as powerful tools for solving noisy linear inverse problems—such as image inpainting, deblurring, and super-resolution—in a "zero-shot" setting that requires no task-specific model retraining. However, existing zero-shot solvers rely on disparate formulations and hand-crafted heuristics that frequently require tedious hyperparameter tuning and yield unstable or suboptimal reconstructions.
The main objective of the article is to establish a unified mathematical framework for existing diffusion-based inverse solvers and demonstrate that replacing heuristic approximations with optimal posterior covariance derived via maximum likelihood estimation significantly enhances reconstruction accuracy without manual tuning.
The authors evaluated their approach across standard benchmarks, including the FFHQ face dataset and the ImageNet dataset, testing multiple image restoration tasks under Gaussian noise. They demonstrated how existing methods uniformly approximate the intractable denoising posterior using simple isotropic Gaussian distributions. To improve upon this, the authors introduced plug-and-play strategies that derive optimal posterior covariance either by analytically converting existing model variances or by using offline Monte Carlo error estimates. Furthermore, they introduced a scalable method to learn covariance predictions directly in a wavelet transform domain, reducing computational complexity from quadratic to linear.
The key findings are as follows:
- Principled covariance estimation significantly outperforms existing methods across restoration tasks; for example, on ImageNet motion deblurring, the proposed conversion approach achieved a Fréchet Inception Distance (FID) of 109.61 compared to 282.21 for Diffusion Posterior Sampling (DPS) and 302.40 for Tweedie Moment Projected Diffusion (TMPD).
- The proposed transform-domain approach (DWT-Var) outperforms optimally tuned baselines without requiring any hyperparameter search, eliminating the sensitivity seen in competing methods.
- Existing heuristic methods (such as guidance step adjustments) also see substantial quality gains simply by replacing their final low-noise sampling steps with the proposed optimal variance estimates.
- Analytical conversion of pre-trained reverse variances proves highly accurate at lower noise levels, whereas high noise levels suffer from numerical limits where boundaries collapse.
These findings indicate that generative image restoration can achieve superior fidelity and robustness at minimal additional computational cost. By eliminating manual hyperparameter sweeps, organizations deploying diffusion pipelines for high-resolution imaging can streamline deployment workflows, reduce experimental overhead, and achieve higher consistency in downstream image reconstruction tasks.
Practitioners should integrate the conversion or Monte Carlo covariance estimators into existing pre-trained diffusion pipelines when deploying zero-shot solvers, applying spatial variance corrections primarily during low-noise sampling steps. When training new diffusion architectures, teams should consider predicting variances directly in an orthonormal transform domain (such as wavelets) to better capture pixel correlations.
Confidence in these findings is supported by rigorous mathematical derivations and extensive multi-metric empirical benchmarks. However, readers should note that the current implementation restricts covariance matrices to diagonal forms to maintain computational feasibility, which leaves some non-diagonal pixel correlations unmodeled. Future work should investigate non-linear transformations and advanced approximations to capture full covariance structures.
- Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). Read the original Diffusion Posterior Sampling method first: the source analyzes its Gaussian posterior approximation and uses it as a central baseline.
- Paper: Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model, Yinhuai Wang et al. (2023). Its null-space projection approach is one of the zero-shot inverse solvers whose heuristic uncertainty treatment the source seeks to unify and improve.
- Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). DDRM provides an influential diffusion-based restoration baseline that helps explain the source’s comparisons with existing inverse-problem solvers.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). The foundational DDPM formulation supplies the reverse-diffusion and denoising framework that the source adapts to measurement-constrained reconstruction.
- Paper: Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior, Tongda Xu et al. (2025). This later reinterprets DPS as posterior maximization and develops aligned algorithmic improvements, extending the source’s effort to clarify and improve diffusion-based inverse solvers.
