Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior
Tongda XuXiyan CaiXinjie ZhangXingtong GeDailan HeLiming SunJingjing LiuYa-Qin ZhangJian LiYan Wang
Demonstrates that Diffusion Posterior Sampling functions as maximum a posteriori optimization rather than true conditional score matching, introducing explicit posterior maximization and an ultra-lightweight estimator trained on just 100 images to improve diffusion-based inverse problem solving.
Diffusion models have become a leading framework for solving image restoration tasks, such as super-resolution and deblurring, without requiring models to be retrained for specific tasks. A widely used baseline, Diffusion Posterior Sampling (DPS), has traditionally been understood as a method that approximates the underlying conditional score distribution during reverse diffusion. However, recent theoretical findings have questioned whether this explanation accurately reflects how the algorithm operates in practice, prompting a re-evaluation of why DPS achieves high visual quality despite known approximation errors.
The article aims to evaluate the actual statistical behavior of DPS and demonstrate that it functions as a maximum a posteriori (MAP) optimization process rather than a conditional score estimator. Based on this insight, the article develops and benchmarks practical algorithmic enhancements that align DPS directly with optimization principles.
To conduct this evaluation, the authors performed empirical tests on high-resolution 512×512 ImageNet validation images using Stable Diffusion 2.0. They measured score approximation errors, score means, and sample variance across tasks such as 8x super-resolution, Gaussian deblurring, and non-linear deblurring. The study compared DPS against baseline models like StableSR and ControlNet, evaluating standard perceptual and reconstruction metrics, including Peak Signal-to-Noise Ratio (PSNR), Fréchet Inception Distance (FID), and Learned Perceptual Image Patch Similarity (LPIPS).
The investigation yielded several critical findings. First, DPS produces conditional score estimates that diverge significantly from true conditional scores, showing errors that are higher than unconditional scores. Paradoxically, tuning parameters to increase score error yielded better image quality. Second, the score function mean in high-performing DPS setups deviated drastically from zero (reaching approximately 5.86 compared to near 0.40 for valid score estimators), violating a fundamental requirement of score functions. Third, DPS output exhibited substantially lower diversity—with average per-pixel standard deviation falling to 0.0453 compared to 0.3939 for conditional models—confirming deterministic MAP-like behavior. Leveraging these findings, the authors introduced DMAP, an algorithm combining multi-step gradient ascent with spherical projection, alongside a lightweight conditional score estimator trained in only 8 GPU hours on 100 images. In 8x super-resolution benchmarks, DMAP reduced the FID score from DPS's 58.48 to 44.37 at matched computational runtime and to 39.56 in a full-step setup, while pairing DPS with the lightweight estimator reduced FID to 44.59.
These results demonstrate that treating diffusion-based image restoration as an explicit optimization problem yields superior visual reconstruction with lower computational overhead. By resolving long-standing theoretical inconsistencies, the findings show that practitioner enhancements like multi-step updates and adaptive optimizers succeed because they improve optimization trajectories. Consequently, engineering efforts can achieve state-of-the-art restoration quality using minimal data and compute rather than relying on expensive, fully trained conditional models.
Organizations deploying diffusion models for restoration tasks should transition from standard DPS to optimization-aligned frameworks like DMAP and explore lightweight estimators for few-shot adaptation. Future engineering should prioritize refining lightweight architectures and investigating few-shot initialization techniques to further reduce training overhead.
The findings are constrained by reliance on ControlNet architectures that are not fully optimized for sample efficiency and by testing conducted on a 1,000-image ImageNet benchmark. Nonetheless, the empirical and theoretical alignment provides strong confidence in the conclusion that diffusion-based restoration behaves primarily as an optimization framework.
- Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). Read the original DPS paper first to understand the method and its conditional-score interpretation that this paper directly reassesses.
- Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). DDRM establishes an influential diffusion-based approach to inverse imaging, providing a useful foundation for the source’s analysis of posterior-guided restoration.
- Paper: Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model, Yinhuai Wang et al. (2023). DDNM’s zero-shot inverse-problem formulation helps situate the posterior-guidance strategies that the source scrutinizes and improves.
No sufficiently relevant recommendations were found.
