Built independently by an author, for readers. Read the story and support ChapterPal

keyword

denoising diffusion implicit models

Denoising diffusion implicit models are a class of generative deep learning models that generate data, such as high-resolution images, through an accelerated and deterministic sampling process. Unlike standard denoising diffusion probabilistic models, which rely on a slow, stochastic Markov chain requiring many iterative denoising steps, these implicit models generalize the diffusion formulation to a non-Markovian process. This design allows the reverse generative trajectory to operate deterministically while sharing the exact training objective and neural network weights as standard diffusion models. Consequently, denoising diffusion implicit models can synthesize high-quality samples in significantly fewer steps and facilitate direct latent-space encoding, inversion, and smooth semantic interpolation between data points.

2 items

Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Guanhua Zhang, Jiabao Ji, Yang Zhang, Mo Yu, Tommi S. Jaakkola, Shiyu Chang

OrganizationsIBMMassachusetts Institute of TechnologyMIT-IBM Watson AI LabUniversity of California, Santa Barbara

Why you should read this

Proposes CoPaint, a Bayesian framework for diffusion-based image inpainting that jointly modifies revealed and unrevealed regions to eliminate incoherence while driving approximation errors to zero to strictly match reference constraints.

Image inpainting refers to the task of generating a complete, natural image based on a partially revealed reference image. Recently, many research interests have been focused on addressing this problem using fixed diffusion models. These approaches typically directly replace the revealed region of the intermediate or final generated images with that of the reference image or its variants. However, since the unrevealed regions are not directly modified to match the context, it results in incoherence between revealed and unrevealed regions. To address the incoherence problem, a small number of methods introduce a rigorous Bayesian framework, but they tend to introduce mismatches between the generated and the reference images due to the approximation errors in computing the posterior distributions. In this paper, we propose CoPaint, which can coherently inpaint the whole image without introducing mismatches. CoPaint also uses the Bayesian framework to jointly modify both revealed and unrevealed regions, but approximates the posterior distribution in a way that allows the errors to gradually drop to zero throughout the denoising steps, thus strongly penalizing any mismatches with the reference image. Our experiments verify that CoPaint can outperform the existing diffusion-based methods under both objective and subjective metrics. The codes are available at https://github.com/UCSB-NLP-Chang/CoPaint/.

Added

2026-09-26

Denoising Diffusion Implicit Models

Denoising Diffusion Implicit Models

Jiaming Song, Chenlin Meng, Stefano Ermon

OrganizationsStanford University

Why you should read this

Develops a deterministic sampling process (non-Markovian) that allows diffusion models to generate high-quality images in 10-50 steps instead of 1000.

Denoising diffusion probabilistic models (DDPMs) have achieved high quality image generation without adversarial training, yet they require simulating a Markov chain for many steps to produce a sample. To accelerate sampling, we present denoising diffusion implicit models (DDIMs), a more efficient class of iterative implicit probabilistic models with the same training procedure as DDPMs. In DDPMs, the generative process is defined as the reverse of a Markovian diffusion process. We construct a class of non-Markovian diffusion processes that lead to the same training objective, but whose reverse process can be much faster to sample from. We empirically demonstrate that DDIMs can produce high quality samples 10×10 \times to 50×50 \times faster in terms of wall-clock time compared to DDPMs, allow us to trade off computation for sample quality, and can perform semantically meaningful image interpolation directly in the latent space.

Added

2026-02-25