GibbsDDRM: A Partially Collapsed Gibbs Sampler for Solving Blind Inverse Problems with Denoising Diffusion Restoration
Naoki MurataKoichi SaitoChieh-Hsin LaiYuhta TakidaToshimitsu UesakaYuki MitsufujiStefano Ermon
Proposes a partially collapsed Gibbs sampling framework that enables pre-trained diffusion models to solve blind inverse problems like image deblurring and vocal dereverberation without requiring fine-tuning or specialized priors for the unknown measurement operator.
Restoring clean signals from corrupted and noisy measurements is a fundamental challenge across engineering domains, including image enhancement and audio processing. In many real-world scenarios, the physical corruption process—such as camera motion blur or room reverberation—is unknown, creating an ill-posed blind linear inverse problem. While recent machine learning breakthroughs utilize pre-trained generative diffusion models as data priors to reconstruct signals, existing methods generally require the corruption process to be known in advance or demand specialized, supervised neural networks trained specifically to predict the corruption operator.
The article demonstrates an effective, problem-agnostic framework called GibbsDDRM that reconstructs clean data from corrupted measurements when the linear measurement operator is entirely unknown. The objective is to evaluate whether a single pre-trained data diffusion model, combined with simple generic parameter priors rather than specialized neural networks, can reliably solve blind restoration tasks without task-specific fine-tuning.
To achieve this, the authors extend non-blind diffusion restoration models into a blind setting by formulating a joint probability distribution over the clean data, diffusion latent variables, observed measurements, and operator parameters. The system performs approximate posterior sampling using a partially collapsed Gibbs sampler, which alternately refines the data and the operator parameters within each restoration cycle. The methodology was evaluated on two challenging benchmarks: blind image deblurring across 1,000 face images (FFHQ) and 500 animal face images (AFHQ), and vocal dereverberation across 1,000 reverberant audio clips generated from the NHSS dataset.
The empirical findings demonstrate that GibbsDDRM achieves superior perceptual reconstruction quality compared to alternative methods. In blind image deblurring, GibbsDDRM outperformed all baseline methods in Learned Perceptual Image Patch Similarity (LPIPS), achieving 0.115 on FFHQ and 0.197 on AFHQ, outperforming supervised and optimization-based models in perceptual alignment to original images. In vocal dereverberation, the method surpassed all evaluated benchmarks, achieving the lowest Fréchet Audio Distance (4.21 compared to 5.69 for the best supervised baseline) and the highest speech-to-reverberation modulation energy ratio (8.40 compared to 7.23 for the baseline). Furthermore, the analysis established that sampling operator parameters via Langevin dynamics provides greater stability and fewer extreme failure cases than standard gradient-based maximum a posteriori optimization.
These results imply that organizations can deploy high-performing restoration pipelines without incurring the significant expense and time required to collect paired training datasets or train specialized operator-estimation models. By decoupling the generic data generative model from the unknown corruption process, the framework provides a versatile foundation applicable to multiple linear inverse domains while maintaining high perceptual fidelity even under significant measurement noise.
For engineering and operational deployment, stakeholders can consider GibbsDDRM for blind restoration workflows when off-the-shelf diffusion models exist for the target signal class. However, deployment decisions must account for computational trade-offs. The method requires substantial processing time—approximately 56 seconds per 256x256 image and 36 seconds per single second of audio—and remains mathematically constrained to linear operators where singular value decomposition is computationally tractable via techniques such as the Fast Fourier Transform. Future work should focus on accelerating the sampling iterations and exploring scalable approximations for non-convolutional or non-linear measurement operators.
- Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). GibbsDDRM explicitly extends DDRM to unknown measurement operators, so reading the original method first clarifies the restoration framework it adapts.
No sufficiently relevant recommendations were found.
