Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

Tongda XuXiyan CaiXinjie ZhangXingtong GeDailan HeLiming SunJingjing LiuYa-Qin ZhangJian LiYan Wang

article2025ICLR25 citations

Demonstrates that Diffusion Posterior Sampling functions as maximum a posteriori optimization rather than true conditional score matching, introducing explicit posterior maximization and an ultra-lightweight estimator trained on just 100 images to improve diffusion-based inverse problem solving.

Listen

Diffusion models have become a leading framework for solving image restoration tasks, such as super-resolution and deblurring, without requiring models to be retrained for specific tasks. A widely used baseline, Diffusion Posterior Sampling (DPS), has traditionally been understood as a method that approximates the underlying conditional score distribution during reverse diffusion. However, recent theoretical findings have questioned whether this explanation accurately reflects how the algorithm operates in practice, prompting a re-evaluation of why DPS achieves high visual quality despite known approximation errors.

The article aims to evaluate the actual statistical behavior of DPS and demonstrate that it functions as a maximum a posteriori (MAP) optimization process rather than a conditional score estimator. Based on this insight, the article develops and benchmarks practical algorithmic enhancements that align DPS directly with optimization principles.

To conduct this evaluation, the authors performed empirical tests on high-resolution 512×512 ImageNet validation images using Stable Diffusion 2.0. They measured score approximation errors, score means, and sample variance across tasks such as 8x super-resolution, Gaussian deblurring, and non-linear deblurring. The study compared DPS against baseline models like StableSR and ControlNet, evaluating standard perceptual and reconstruction metrics, including Peak Signal-to-Noise Ratio (PSNR), Fréchet Inception Distance (FID), and Learned Perceptual Image Patch Similarity (LPIPS).

The investigation yielded several critical findings. First, DPS produces conditional score estimates that diverge significantly from true conditional scores, showing errors that are higher than unconditional scores. Paradoxically, tuning parameters to increase score error yielded better image quality. Second, the score function mean in high-performing DPS setups deviated drastically from zero (reaching approximately 5.86 compared to near 0.40 for valid score estimators), violating a fundamental requirement of score functions. Third, DPS output exhibited substantially lower diversity—with average per-pixel standard deviation falling to 0.0453 compared to 0.3939 for conditional models—confirming deterministic MAP-like behavior. Leveraging these findings, the authors introduced DMAP, an algorithm combining multi-step gradient ascent with spherical projection, alongside a lightweight conditional score estimator trained in only 8 GPU hours on 100 images. In 8x super-resolution benchmarks, DMAP reduced the FID score from DPS's 58.48 to 44.37 at matched computational runtime and to 39.56 in a full-step setup, while pairing DPS with the lightweight estimator reduced FID to 44.59.

These results demonstrate that treating diffusion-based image restoration as an explicit optimization problem yields superior visual reconstruction with lower computational overhead. By resolving long-standing theoretical inconsistencies, the findings show that practitioner enhancements like multi-step updates and adaptive optimizers succeed because they improve optimization trajectories. Consequently, engineering efforts can achieve state-of-the-art restoration quality using minimal data and compute rather than relying on expensive, fully trained conditional models.

Organizations deploying diffusion models for restoration tasks should transition from standard DPS to optimization-aligned frameworks like DMAP and explore lightweight estimators for few-shot adaptation. Future engineering should prioritize refining lightweight architectures and investigating few-shot initialization techniques to further reduce training overhead.

The findings are constrained by reliance on ControlNet architectures that are not fully optimized for sample efficiency and by testing conducted on a 1,000-image ImageNet benchmark. Nonetheless, the empirical and theoretical alignment provides strong confidence in the conclusion that diffusion-based restoration behaves primarily as an optimization framework.

No sufficiently relevant recommendations were found.

Cover for Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

Abstract

Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating the conditional score. While in this paper, we demonstrate that the conditional score approximation employed by DPS is not as effective as previously assumed, but rather aligns more closely with the principle of maximizing a posterior (MAP). This assertion is substantiated through an examination of DPS on 512x512 ImageNet images, revealing that: 1) DPS's conditional score estimation significantly diverges from the score of a well-trained conditional diffusion model and is even inferior to the unconditional score; 2) The mean of DPS's conditional score estimation deviates significantly from zero, rendering it an invalid score estimation; 3) DPS generates high-quality samples with significantly lower diversity. In light of the above findings, we posit that DPS more closely resembles MAP than a conditional score estimator, and accordingly propose the following enhancements to DPS: 1) we explicitly maximize the posterior through multi-step gradient ascent and projection; 2) we utilize a light-weighted conditional score estimator trained with only 100 images and 8 GPU hours. Extensive experimental results indicate that these proposed improvements significantly enhance DPS's performance. The source code for these improvements is provided in this https URL.

Citation

MLA
Xu, T., et al. “Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior”. arXiv, 2025, http://arxiv.org/abs/2501.18913v2.
APA
Xu, T., Cai, X., Zhang, X., Ge, X., He, D., Sun, M., Liu, J., Zhang, Y.-Q., Li, J., & Wang, Y. (2025). Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior. arXiv. http://arxiv.org/abs/2501.18913v2
Chicago
Xu, T., X. Cai, X. Zhang, et al. 2025. “Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior”. arXiv. http://arxiv.org/abs/2501.18913v2.
Harvard
Xu, T. et al. (2025) “Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2501.18913v2.
Vancouver
1. Xu T, Cai X, Zhang X, Ge X, He D, Sun M, Liu J, Zhang Y-Q, Li J, Wang Y (2025) Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior. arXiv

BibTeX

@article{xu2025rethinking,
  title = {Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior},
  author = {Xu, Tongda and Cai, Xiyan and Zhang, Xinjie and Ge, Xingtong and He, Dailan and Sun, Ming and Liu, Jingjing and Zhang, Ya-Qin and Li, Jian and Wang, Yan},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2501.18913v2},
  eprint = {2501.18913}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/