Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models
Guanhua ZhangJiabao JiYang ZhangMo YuTommi S. JaakkolaShiyu Chang
Proposes CoPaint, a Bayesian framework for diffusion-based image inpainting that jointly modifies revealed and unrevealed regions to eliminate incoherence while driving approximation errors to zero to strictly match reference constraints.
Image inpainting—the process of reconstructing missing or damaged portions of an image so the result looks complete and realistic—is critical for media editing, restoration, and content creation. While recent advances use pre-trained generative diffusion models to fill missing regions without requiring task-specific retraining, existing methods frequently suffer from severe visual incoherence. Standard techniques directly overwrite or blend the known image regions into intermediate steps, leaving unmasked areas contextually disconnected and resulting in visible seams, mismatched textures, or conflicting styles. Conversely, formal statistical methods that attempt full-image alignment struggle with high computational costs and mathematical approximations that distort the original, unmasked regions.
The article demonstrates and evaluates COPAINT, an inpainting algorithm designed to achieve visual coherence across the entire image while strictly preserving the authentic reference regions. The core objective is to deliver high-quality, seamless image completions using standard, fixed diffusion models without introducing approximation mismatches or requiring expensive model retraining.
To accomplish this, the authors implemented a framework that simultaneously adjusts both the revealed and unrevealed image areas at each step of the iterative image generation process. The algorithm calculates the required adjustments using a fast, one-step estimation of the final image, progressively reducing approximation errors to zero as the generation reaches completion. Credibility was established through extensive testing on two standard computer vision benchmarks—CelebA-HQ (celebrity portraits) and ImageNet (diverse object categories)—across seven distinct degradation masks. The authors evaluated image quality using both automated visual similarity metrics and blinded human evaluations involving 1,400 image pair comparisons.
The findings show that COPAINT consistently outperforms existing diffusion-based inpainting baselines in both objective fidelity and subjective visual appeal. On the ImageNet benchmark, COPAINT achieved a 19% relative reduction in automated error metrics compared to the top-performing baseline (REPAINT) while using 31% less computational budget. When paired with an optional error-reduction technique called "time travel" (COPAINT-TT), the method won the majority of human preference votes for visual coherence and naturalness across eleven of fourteen distinct testing scenarios. A streamlined variant, COPAINT-FAST, ran approximately four times faster than standard COPAINT while maintaining competitive or superior visual quality against baseline alternatives. Furthermore, secondary experiments confirmed that the approach successfully transfers to high-resolution (512×512) inpainting and image super-resolution tasks.
These results indicate that organizations can achieve state-of-the-art visual restorations and image completions using existing, off-the-shelf generative models rather than investing in costly, specialized model training. Operating over the entire image rather than simply pasting known pixels also mitigates undesirable machine learning biases—such as defaulting to stereotyped demographic features—by forcing the model to adhere strictly to the visual evidence present in the original input.
Technical leaders deploying generative restoration systems should consider adopting COPAINT's progressive adjustment approach, selecting COPAINT-FAST for low-latency operational environments or COPAINT-TT where maximum visual fidelity is required. Prior to full-scale deployment, teams should conduct domain-specific pilot testing and implement safeguards against the potential misuse of automated realistic image generation for deceptive media.
Decision-makers should note that the approach relies on a step-by-step optimization process that can occasionally produce minor visual flaws when reconstructing complex fine details, such as small typography, or when original images are excessively masked. Nevertheless, the extensive empirical benchmarks and consistent human evaluations support a high level of confidence in the algorithm's performance advantages over current industry baselines.
- Paper: RePaint: Inpainting using Denoising Diffusion Probabilistic Models, Andreas Lugmayr et al. (2022). RePaint introduced unconditional diffusion-based inpainting by iteratively replacing revealed regions during sampling, providing the foundational baseline and motivation for CoPaint's coherent Bayesian formulation.
- Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Denoising Diffusion Implicit Models (DDIM) formalize the non-Markovian deterministic sampling framework directly leveraged and modified by CoPaint to perform coherent inpainting.
- Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). Diffusion Posterior Sampling establishes the Bayesian framework and posterior approximation techniques using Tweedie's formula for solving linear inverse problems with diffusion models.
- Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). Denoising Diffusion Restoration Models (DDRM) present an unsupervised diffusion-based inverse problem framework that addresses inpainting and data-consistency trade-offs.
- Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Denoising Diffusion Probabilistic Models (DDPM) establish the core probabilistic diffusion and reverse denoising mathematics underlying all subsequent diffusion-based inpainting models.
- Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). SDEdit introduces stochastic differential equation-guided denoising for image editing and patch compositing, establishing practical generative editing priors for diffusion models.
- Paper: Blended Diffusion for Text-driven Editing of Natural Images, Omri Avrahami et al. (2021). Blended Diffusion introduces multi-step intermediate blending for region-based diffusion editing, illustrating the boundary incoherence challenges addressed by CoPaint.
- Paper: Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model, Yinhuai Wang et al. (2023). DDNM advances zero-shot image restoration and inpainting by utilizing range-null space decomposition to enforce strict data fidelity without posterior approximation mismatches.
- Paper: InstructPix2Pix: Learning to Follow Image Editing Instructions, Tim Brooks et al. (2023). InstructPix2Pix extends diffusion-based image modification beyond masked inpainting to direct, free-form instruction-guided image editing without requiring precise reference region replacements.
- Paper: Hierarchical Fine-Grained Image Forgery Detection and Localization, Xiao Guo et al. (2023). This work explores downstream detection and localization of subtle image manipulations produced by modern diffusion-based inpainting and generation techniques.
- Paper: Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models, Gowthami Somepalli et al. (2023). This paper investigates content memorization and data replication in diffusion models, highlighting the risks of implicit reference duplication in generative synthesis.
