Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Guanhua ZhangJiabao JiYang ZhangMo YuTommi S. JaakkolaShiyu Chang

article2023ICML96 citations

Proposes CoPaint, a Bayesian framework for diffusion-based image inpainting that jointly modifies revealed and unrevealed regions to eliminate incoherence while driving approximation errors to zero to strictly match reference constraints.

Listen

Image inpainting—the process of reconstructing missing or damaged portions of an image so the result looks complete and realistic—is critical for media editing, restoration, and content creation. While recent advances use pre-trained generative diffusion models to fill missing regions without requiring task-specific retraining, existing methods frequently suffer from severe visual incoherence. Standard techniques directly overwrite or blend the known image regions into intermediate steps, leaving unmasked areas contextually disconnected and resulting in visible seams, mismatched textures, or conflicting styles. Conversely, formal statistical methods that attempt full-image alignment struggle with high computational costs and mathematical approximations that distort the original, unmasked regions.

The article demonstrates and evaluates COPAINT, an inpainting algorithm designed to achieve visual coherence across the entire image while strictly preserving the authentic reference regions. The core objective is to deliver high-quality, seamless image completions using standard, fixed diffusion models without introducing approximation mismatches or requiring expensive model retraining.

To accomplish this, the authors implemented a framework that simultaneously adjusts both the revealed and unrevealed image areas at each step of the iterative image generation process. The algorithm calculates the required adjustments using a fast, one-step estimation of the final image, progressively reducing approximation errors to zero as the generation reaches completion. Credibility was established through extensive testing on two standard computer vision benchmarks—CelebA-HQ (celebrity portraits) and ImageNet (diverse object categories)—across seven distinct degradation masks. The authors evaluated image quality using both automated visual similarity metrics and blinded human evaluations involving 1,400 image pair comparisons.

The findings show that COPAINT consistently outperforms existing diffusion-based inpainting baselines in both objective fidelity and subjective visual appeal. On the ImageNet benchmark, COPAINT achieved a 19% relative reduction in automated error metrics compared to the top-performing baseline (REPAINT) while using 31% less computational budget. When paired with an optional error-reduction technique called "time travel" (COPAINT-TT), the method won the majority of human preference votes for visual coherence and naturalness across eleven of fourteen distinct testing scenarios. A streamlined variant, COPAINT-FAST, ran approximately four times faster than standard COPAINT while maintaining competitive or superior visual quality against baseline alternatives. Furthermore, secondary experiments confirmed that the approach successfully transfers to high-resolution (512×512) inpainting and image super-resolution tasks.

These results indicate that organizations can achieve state-of-the-art visual restorations and image completions using existing, off-the-shelf generative models rather than investing in costly, specialized model training. Operating over the entire image rather than simply pasting known pixels also mitigates undesirable machine learning biases—such as defaulting to stereotyped demographic features—by forcing the model to adhere strictly to the visual evidence present in the original input.

Technical leaders deploying generative restoration systems should consider adopting COPAINT's progressive adjustment approach, selecting COPAINT-FAST for low-latency operational environments or COPAINT-TT where maximum visual fidelity is required. Prior to full-scale deployment, teams should conduct domain-specific pilot testing and implement safeguards against the potential misuse of automated realistic image generation for deceptive media.

Decision-makers should note that the approach relies on a step-by-step optimization process that can occasionally produce minor visual flaws when reconstructing complex fine details, such as small typography, or when original images are excessively masked. Nevertheless, the extensive empirical benchmarks and consistent human evaluations support a high level of confidence in the algorithm's performance advantages over current industry baselines.

  • Paper: RePaint: Inpainting using Denoising Diffusion Probabilistic Models, Andreas Lugmayr et al. (2022). RePaint introduced unconditional diffusion-based inpainting by iteratively replacing revealed regions during sampling, providing the foundational baseline and motivation for CoPaint's coherent Bayesian formulation.
  • Paper: Denoising Diffusion Implicit Models, Jiaming Song et al. (2021). Denoising Diffusion Implicit Models (DDIM) formalize the non-Markovian deterministic sampling framework directly leveraged and modified by CoPaint to perform coherent inpainting.
  • Paper: Diffusion Posterior Sampling for General Noisy Inverse Problems, Hyungjin Chung et al. (2022). Diffusion Posterior Sampling establishes the Bayesian framework and posterior approximation techniques using Tweedie's formula for solving linear inverse problems with diffusion models.
  • Paper: Denoising Diffusion Restoration Models, Bahjat Kawar et al. (2022). Denoising Diffusion Restoration Models (DDRM) present an unsupervised diffusion-based inverse problem framework that addresses inpainting and data-consistency trade-offs.
  • Paper: Denoising Diffusion Probabilistic Models, Jonathan Ho et al. (2020). Denoising Diffusion Probabilistic Models (DDPM) establish the core probabilistic diffusion and reverse denoising mathematics underlying all subsequent diffusion-based inpainting models.
  • Paper: SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations, Chenlin Meng et al. (2022). SDEdit introduces stochastic differential equation-guided denoising for image editing and patch compositing, establishing practical generative editing priors for diffusion models.
  • Paper: Blended Diffusion for Text-driven Editing of Natural Images, Omri Avrahami et al. (2021). Blended Diffusion introduces multi-step intermediate blending for region-based diffusion editing, illustrating the boundary incoherence challenges addressed by CoPaint.
Cover for Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Abstract

Image inpainting refers to the task of generating a complete, natural image based on a partially revealed reference image. Recently, many research interests have been focused on addressing this problem using fixed diffusion models. These approaches typically directly replace the revealed region of the intermediate or final generated images with that of the reference image or its variants. However, since the unrevealed regions are not directly modified to match the context, it results in incoherence between revealed and unrevealed regions. To address the incoherence problem, a small number of methods introduce a rigorous Bayesian framework, but they tend to introduce mismatches between the generated and the reference images due to the approximation errors in computing the posterior distributions. In this paper, we propose CoPaint, which can coherently inpaint the whole image without introducing mismatches. CoPaint also uses the Bayesian framework to jointly modify both revealed and unrevealed regions, but approximates the posterior distribution in a way that allows the errors to gradually drop to zero throughout the denoising steps, thus strongly penalizing any mismatches with the reference image. Our experiments verify that CoPaint can outperform the existing diffusion-based methods under both objective and subjective metrics. The codes are available at https://github.com/UCSB-NLP-Chang/CoPaint/.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Background and Notations
  • 4. The COPAINT Algorithm
  • 4.1. Problem Formulation
  • 4.2. A Prototype Approach
  • 4.3. One-Step Approximation
  • 4.4. Denoising Successive Correction
  • 4.5. Additional Algorithmic Designs
  • 5. Experiments
  • 5.1. Experiment Setup
  • 5.2. Experiment Results
  • 5.3. Coherence Study
  • 5.4. Ablation Study
  • 6. Conclusion
  • References
  • A. Experiment Setup
  • A.1. Human Evaluation
  • A.2. Implementation Details of Baselines
  • A.3. Adaptive Learning Rate for Our Method
  • B. Qualitative Results
  • C. Additional High-resolution Inpainting Experiments
  • D. Additional Super-resolution Experiments
  • E. Failure Case Study
  • F. Potential Societal Impacts

Knowls

  1. Knowl 1 — COPAINT Posterior Approximation Framework for Diffusion-Based Inpainting

    model/method

    COPAINT is an unsupervised Bayesian image inpainting method that uses a pre-trained, fixed Denoising Diffusion Implicit Model (DDIM) without model fine-tuning. Given a reference image with revealed portion s0s_0 and a masking selection operator r(⋅)r(\cdot), the inpainting constraint is defined as event C:r(X~0)=s0C: r(\tilde{X}_0) = s_0.

    Rather than directly replacing pixels in intermediate latents, COPAINT samples a latent sequence X~0:T\tilde{X}_{0:T} from the approximate joint posterior distribution: pθ′(X~0:T∣C)=pθ′(X~T∣C)∏t=1Tpθ′(X~t−1∣X~t,C)p'_\theta(\tilde{X}_{0:T}|C) = p'_\theta(\tilde{X}_T|C) \prod_{t=1}^T p'_\theta(\tilde{X}_{t-1}|\tilde{X}_t, C)

    To make the posterior tractable at each step tt, the distribution of the final revealed image r(X~0)r(\tilde{X}_0) given X~t\tilde{X}_t is approximated around the one-step deterministic generation fθ(t)(X~t)=X~t−1−αˉtϵθ(t)(X~t)αˉtf_\theta^{(t)}(\tilde{X}_t) = \frac{\tilde{X}_t - \sqrt{1 - \bar{\alpha}_t}\epsilon_\theta^{(t)}(\tilde{X}_t)}{\sqrt{\bar{\alpha}_t}} with Gaussian error: pθ′(r(X~0)∣X~t)=N(r(X~0);r(fθ(t)(X~t)),ξt′2I)p'_\theta(r(\tilde{X}_0)|\tilde{X}_t) = \mathcal{N}(r(\tilde{X}_0); r(f_\theta^{(t)}(\tilde{X}_t)), \xi_t^{\prime 2} I) where αˉt=∏i=1tαi\bar{\alpha}_t = \prod_{i=1}^t \alpha_i is the noise schedule product and ξt′2\xi_t^{\prime 2} captures the one-step approximation error.

    The transition log-posterior for sampling X~t−1\tilde{X}_{t-1} given X~t\tilde{X}_t is expressed as: log⁡pθ′(X~t−1∣X~t,C)=−12σt2∥X~t−1−μ~t∥22−12ξt−1′2∥s0−r(fθ(t−1)(X~t−1))∥22+C′\log p'_\theta(\tilde{X}_{t-1}|\tilde{X}_t, C) = -\frac{1}{2\sigma_t^2} \|\tilde{X}_{t-1} - \tilde{\mu}_t\|_2^2 - \frac{1}{2\xi_{t-1}^{\prime 2}} \|s_0 - r(f_\theta^{(t-1)}(\tilde{X}_{t-1}))\|_2^2 + C' where C′C' is a normalizing constant, σt\sigma_t is the DDIM stochasticity parameter, and μ~t\tilde{\mu}_t is the reverse DDIM mean: μ~t=αˉt−1fθ(t)(X~t)+1−αˉt−1−σt2X~t−αˉtfθ(t)(X~t)1−αˉt\tilde{\mu}_t = \sqrt{\bar{\alpha}_{t-1}} f_\theta^{(t)}(\tilde{X}_t) + \sqrt{1 - \bar{\alpha}_{t-1} - \sigma_t^2} \frac{\tilde{X}_t - \sqrt{\bar{\alpha}_t} f_\theta^{(t)}(\tilde{X}_t)}{\sqrt{1 - \bar{\alpha}_t}}

    Optimizing over the entire latent X~t−1\tilde{X}_{t-1} modifies both revealed and unrevealed regions simultaneously, eliminating boundary incoherence. As t→1t \to 1, the one-step approximation error converges to zero (fθ(1)(X~1)=X~0f_\theta^{(1)}(\tilde{X}_1) = \tilde{X}_0 when σ1=0\sigma_1 = 0), strictly enforcing the inpainting constraint at the final step.

  2. Knowl 2 — COPAINT-TT Inpainting Algorithm

    algorithm

    COPAINT with Time Travel (COPAINT-TT) generates an inpainted image by alternating between gradient-based posterior optimization of intermediate latents, DDIM reverse denoising steps, and periodic time-travel forward steps to improve consistency.

    Input: Revealed reference content s0s_0, pre-trained denoising network family {fθ(t)}t=1T\{f_\theta^{(t)}\}_{t=1}^T, time travel interval τ\tau, time travel frequency KK, gradient descent step count GG, learning rates {ηt}t=1T\{\eta_t\}_{t=1}^T
    Output: Inpainted image X~0\tilde{X}_0
    Initialize X~T∼N(0,I)\tilde{X}_T \sim \mathcal{N}(0, I)
    t←Tt \leftarrow T
    k←Kk \leftarrow K
    while t≠0t \neq 0 do
        for g=1g = 1 to GG do
            Compute loss gradient ∇X~tLt\nabla_{\tilde{X}_t} \mathcal{L}_t
            X~t←X~t−ηt∇X~tLt\tilde{X}_t \leftarrow \tilde{X}_t - \eta_t \nabla_{\tilde{X}_t} \mathcal{L}_t
        end for
        Compute X~t−1\tilde{X}_{t-1} from X~t\tilde{X}_t using DDIM reverse transition step
        t←t−1t \leftarrow t - 1
        if t(modτ)=0t \pmod \tau = 0 and t≤T−τt \le T - \tau then
            if k>0k > 0 then
                Sample X~t+τ∼q(X~t+τ∣X~t)\tilde{X}_{t+\tau} \sim q(\tilde{X}_{t+\tau}|\tilde{X}_t)
                t←t+τ−1t \leftarrow t + \tau - 1
                k←k−1k \leftarrow k - 1
            else
                k←Kk \leftarrow K
            end if
        end if
    end while
    return X~0\tilde{X}_0

    At each timestep tt, X~t\tilde{X}_t is updated via GG steps of gradient descent using the adaptive learning rate ηt=0.02αˉt\eta_t = 0.02\sqrt{\bar{\alpha}_t}. Time travel rewinds the latent state by sampling X~t+τ∼N(αˉt+τ/αˉtX~t,(1−αˉt+τ/αˉt)I)\tilde{X}_{t+\tau} \sim \mathcal{N}(\sqrt{\bar{\alpha}_{t+\tau}/\bar{\alpha}_t}\tilde{X}_t, (1 - \bar{\alpha}_{t+\tau}/\bar{\alpha}_t)I) every τ\tau timesteps for KK rounds.

  3. Knowl 3 — Adaptive Learning Rate for Latent Posterior Optimization

    model/method

    In COPAINT, the latent X~t\tilde{X}_t at step tt is optimized to minimize the negative log-posterior loss: Lt=−log⁡pθ′(X~t∣X~t+1,C)=12σt2∥X~t−μ~t∥22+12ξt′2∥s0−r(fθ(t)(X~t))∥22\mathcal{L}_t = -\log p'_\theta(\tilde{X}_t|\tilde{X}_{t+1}, C) = \frac{1}{2\sigma_t^2} \|\tilde{X}_t - \tilde{\mu}_t\|_2^2 + \frac{1}{2\xi_t^{\prime 2}} \|s_0 - r(f_\theta^{(t)}(\tilde{X}_t))\|_2^2

    The analytical gradient with respect to X~t\tilde{X}_t is: ∇X~tLt=1σt2(X~t−μ~t)+1ξt′2(∂fθ(t)(X~t)∂X~t)T(∂r(fθ(t)(X~t))∂fθ(t)(X~t))T[r(fθ(t)(X~t))−s0]\nabla_{\tilde{X}_t} \mathcal{L}_t = \frac{1}{\sigma_t^2}(\tilde{X}_t - \tilde{\mu}_t) + \frac{1}{\xi_t^{\prime 2}} \left(\frac{\partial f_\theta^{(t)}(\tilde{X}_t)}{\partial \tilde{X}_t}\right)^T \left(\frac{\partial r(f_\theta^{(t)}(\tilde{X}_t))}{\partial f_\theta^{(t)}(\tilde{X}_t)}\right)^T [r(f_\theta^{(t)}(\tilde{X}_t)) - s_0] where ∂r(⋅)∂fθ(t)(X~t)\frac{\partial r(\cdot)}{\partial f_\theta^{(t)}(\tilde{X}_t)} is a diagonal matrix with entries in {0,1}\{0, 1\}.

    Differentiating the DDIM one-step estimation fθ(t)(X~t)=X~t−1−αˉtϵθ(t)(X~t)αˉtf_\theta^{(t)}(\tilde{X}_t) = \frac{\tilde{X}_t - \sqrt{1-\bar{\alpha}_t}\epsilon_\theta^{(t)}(\tilde{X}_t)}{\sqrt{\bar{\alpha}_t}} yields: ∂fθ(t)(X~t)∂X~t=I−1−αˉt∇X~tϵθ(t)(X~t)αˉt\frac{\partial f_\theta^{(t)}(\tilde{X}_t)}{\partial \tilde{X}_t} = \frac{I - \sqrt{1 - \bar{\alpha}_t} \nabla_{\tilde{X}_t}\epsilon_\theta^{(t)}(\tilde{X}_t)}{\sqrt{\bar{\alpha}_t}}

    Since αˉt\bar{\alpha}_t decreases monotonically toward 0 as tt increases, the factor 1/αˉt1/\sqrt{\bar{\alpha}_t} becomes very large for early denoising steps, causing gradient explosion and numerical instability (NaN values). To cancel out this magnitude imbalance, an adaptive learning rate is defined as: ηt=0.02αˉt\eta_t = 0.02 \sqrt{\bar{\alpha}_t} scaling the step size inversely with gradient inflation.

  4. Knowl 4 — Quantitative Inpainting Evaluation on CelebA-HQ and ImageNet

    data/table

    Evaluations across 100 test images each on CelebA-HQ (256×256256 \times 256) and ImageNet-1K (256×256256 \times 256) compare COPAINT variants against replacement and posterior sampling baselines across seven degradation masks (Expand, Half, Altern, Super-Resolution [S.R.], Narrow, Wide, Text). Reported metrics are AlexNet LPIPS (lower is better) and Subjective Human Vote Difference (Vote %, reported as Overall / Coherence, where positive indicates evaluator preference for COPAINT-TT).

    Dataset / Method Expand Half Altern S.R. Narrow Wide Text Average
    CelebA-HQ
    BLENDED 0.557 0.228 0.047 0.269 0.078 0.102 0.011 0.185
    DDRM 0.704 0.273 0.151 0.596 0.140 0.125 0.028 0.288
    RESAMPLING 0.536 0.231 0.050 0.261 0.077 0.102 0.013 0.181
    REPAINT 0.496 0.199 0.014 0.041 0.039 0.072 0.006 0.124
    DPS 0.449 0.261 0.166 0.182 0.160 0.181 0.152 0.222
    DDNM 0.598 0.257 0.015 0.046 0.071 0.111 0.014 0.158
    COPAINT-FAST 0.483 0.203 0.057 0.084 0.068 0.096 0.036 0.147
    COPAINT 0.472 0.188 0.016 0.033 0.040 0.071 0.007 0.118
    COPAINT-TT 0.464 0.180 0.014 0.028 0.037 0.069 0.006 0.114
    ImageNet
    BLENDED 0.717 0.366 0.277 0.686 0.161 0.194 0.028 0.347
    DDRM 0.730 0.385 0.439 0.822 0.211 0.231 0.060 0.411
    RESAMPLING 0.704 0.353 0.259 0.624 0.151 0.183 0.028 0.329
    REPAINT 0.706 0.323 0.103 0.209 0.072 0.156 0.014 0.226
    DPS 0.673 0.512 0.474 0.511 0.447 0.468 0.438 0.503
    DDNM 0.805 0.408 0.051 0.107 0.101 0.185 0.012 0.238
    COPAINT-FAST 0.678 0.335 0.075 0.128 0.103 0.167 0.043 0.218
    COPAINT 0.640 0.307 0.041 0.069 0.078 0.138 0.017 0.184
    COPAINT-TT 0.636 0.294 0.039 0.069 0.074 0.133 0.015 0.180

    COPAINT-TT achieves the lowest average LPIPS on both datasets (0.114 on CelebA-HQ and 0.180 on ImageNet), outperforming REPAINT by 8% and 20% relative margins, respectively. Human evaluation vote difference scores between COPAINT-TT and baselines on average across masks are: vs BLENDED (51%/57% on CelebA-HQ, 64%/65% on ImageNet), vs DDRM (79%/81%, 75%/71%), vs RESAMPLING (42%/56%, 61%/67%), vs REPAINT (0%/6%, 34%/29%), vs DPS (41%/45%, 87%/86%), vs DDNM (27%/39%, 33%/44%).

  5. Knowl 5 — High-Resolution (512x512) Image Inpainting Performance

    data/table

    Inpainting performance evaluated on ImageNet at 512×512512 \times 512 resolution using a pre-trained Guided Diffusion backbone across seven degradation masks:

    Method Expand Half Altern S.R. Narrow Wide Text Average
    BLENDED 0.739 0.377 0.210 0.495 0.157 0.179 0.038 0.313
    DDRM 0.859 0.391 0.339 0.712 0.204 0.197 0.073 0.396
    RESAMPLING 0.799 0.366 0.205 0.482 0.157 0.173 0.039 0.317
    REPAINT 0.835 0.351 0.066 0.158 0.083 0.146 0.019 0.237
    DPS 0.750 0.575 0.513 0.543 0.496 0.519 0.480 0.554
    DDNM 0.850 0.406 0.033 0.079 0.173 0.193 0.044 0.254
    COPAINT-FAST 0.678 0.335 0.075 0.128 0.103 0.167 0.043 0.218
    COPAINT 0.732 0.310 0.033 0.067 0.100 0.146 0.026 0.202
    COPAINT-TT 0.726 0.292 0.022 0.043 0.093 0.136 0.025 0.191

    COPAINT-TT achieves the lowest overall average LPIPS (0.191), representing a 19.4% relative error reduction over the strongest baseline REPAINT (0.237). COPAINT-FAST (0.218) outperforms all baselines including REPAINT while requiring less runtime.

  6. Knowl 6 — Zero-Shot Super-Resolution Performance across Degradation Scales

    data/table

    COPAINT was evaluated on zero-shot image super-resolution on ImageNet and CelebA-HQ. Low-resolution inputs were generated via average pooling at 2×2\times, 4×4\times, and 8×8\times downsampling factors, and reconstructed back to 256×256256 \times 256 using pre-trained diffusion models. Results are measured in LPIPS (lower is better):

    Dataset Scale DPS DDRM DDNM COPAINT COPAINT-TT
    ImageNet 2×\times 0.156 0.054 0.031 0.037 0.025
    4×\times 0.190 0.228 0.141 0.113 0.082
    8×\times 0.235 0.360 0.250 0.293 0.170
    CelebA-HQ 2×\times 0.417 0.121 0.113 0.063 0.042
    4×\times 0.483 0.345 0.328 0.252 0.204
    8×\times 0.531 0.480 0.528 0.511 0.423

    COPAINT-TT achieves the best performance across all scales and datasets. In extreme 8×8\times downsampling, where competing baselines like DDNM and DDRM generate blurry results lacking texture, COPAINT-TT reduces LPIPS from 0.250 to 0.170 on ImageNet and from 0.480 to 0.423 on CelebA-HQ.

  7. Knowl 7 — Hyperparameter Ablation of Gradient Steps, Time Travel, and Multi-Step Estimation

    data/table

    Ablation experiments conducted on the CelebA-HQ test set with Half mask evaluating gradient descent steps GG, time travel frequency KK, time travel interval τ\tau, and multi-step approximation depth HH:

    Parameter Setting LPIPS ↓\downarrow Time (s) ↓\downarrow
    Gradient steps (GG) 1 0.187 326
    2 0.180 562
    5 0.192 1365
    Travel freq. (KK) 1 0.180 562
    2 0.179 721
    5 0.181 1428
    Travel interval (τ\tau) 2 0.186 567
    5 0.178 569
    10 0.180 562
    20 0.181 564
    Approximation steps (HH) 1 0.180 562
    2 0.176 1491
    5 0.177 3346

    Key findings:

    1. Setting G=2G=2 yields optimal performance; G=5G=5 degrades LPIPS due to overfitting the gradient updates to the revealed region.
    2. A small time-travel frequency K=1K=1 is sufficient (unlike REPAINT which requires K=9K=9).
    3. Time travel interval τ≥5\tau \ge 5 yields stable performance.
    4. Multi-step estimation (H>1H > 1) yields minor LPIPS gains but increases computational cost up to sixfold.
  8. Knowl 8 — Variance Schedule Parameterization and Multi-Step Generation

    model/method

    In COPAINT, the optimal variance ξt′2\xi_t^{\prime 2} of the approximate conditional likelihood pθ′(r(X~0)∣X~t)=N(r(X~0);r(fθ(t)(X~t)),ξt′2I)p'_\theta(r(\tilde{X}_0)|\tilde{X}_t) = \mathcal{N}(r(\tilde{X}_0); r(f_\theta^{(t)}(\tilde{X}_t)), \xi_t^{\prime 2}I) analytically equals the expected mean squared error: ξt′2=1NEpθ[∥r(fθ(t)(X~t))−r(X~0)∥22]\xi_t^{\prime 2} = \frac{1}{N} \mathbb{E}_{p_\theta} [\|r(f_\theta^{(t)}(\tilde{X}_t)) - r(\tilde{X}_0)\|_2^2] where NN is the dimension of the revealed region s0s_0.

    To avoid computing this expectation explicitly at runtime, COPAINT parameterizes the variance as a monotonically decreasing function: ξt′2=(11.012)T−t\xi_t^{\prime 2} = \left(\frac{1}{1.012}\right)^{T-t} This formulation increases the penalty weight 12ξt′2\frac{1}{2\xi_t^{\prime 2}} as denoising progresses toward t=0t = 0, enforcing the inpainting constraint with increasing precision as the one-step approximation error vanishes.

    To further improve accuracy during early denoising steps (large tt) where the one-step approximation error is largest, one-step generation fθ(t)(X~t)f_\theta^{(t)}(\tilde{X}_t) can optionally be replaced with an HH-step approximation, rolling out HH deterministic DDIM denoising steps over a subsampled sequence of timesteps to estimate X0X_0.

  9. Knowl 9 — Computational Trade-offs and Fast Inpainting Variant

    empirical result

    COPAINT achieves superior inpainting quality with improved time efficiency relative to iterative sampling baselines:

    1. Compared to REPAINT, standard COPAINT (G=2G=2, 250 sampling steps) reduces average per-image processing time by approximately 60% on both CelebA-HQ and ImageNet while attaining lower average LPIPS (0.118 vs 0.124 on CelebA-HQ; 0.184 vs 0.226 on ImageNet).
    2. COPAINT-FAST, configured with G=1G=1 gradient descent step per timestep and 100 total reverse sampling steps, runs four times faster than standard COPAINT. It achieves an average LPIPS of 0.147 on CelebA-HQ (outperforming BLENDED, DDRM, RESAMPLING, DPS, and DDNM) and 0.218 on ImageNet (outperforming all baseline methods including REPAINT).
  10. Knowl 10 — Methodological and Generative Limitations of COPAINT

    limitation

    COPAINT exhibits the following limitations:

    1. Suboptimal Greedy Trajectory Optimization: Sampling X~0:T\tilde{X}_{0:T} sequentially via step-wise maximum a posteriori (MAP) gradient descent is locally greedy and does not guarantee convergence to the global posterior mode.
    2. Early Denoising Approximation Errors: In the earliest denoising steps (near t=Tt = T), the one-step estimation fθ(t)(X~t)f_\theta^{(t)}(\tilde{X}_t) has large approximation error, which can misguide intermediate latent trajectories before corrections take effect.
    3. Fine Detail and Text Synthesis Deficits: The framework struggles to synthesize fine high-frequency structures, such as legible text or fine hat insignia, reflecting resolution and detail limits inherited from the pre-trained diffusion backbones.
    4. Under-Constrained Large Mask Failures: When degradation masks conceal the vast majority of image content, the remaining context provides insufficient semantic conditioning, leading to unnatural completions or reliance on pre-trained dataset biases (e.g., demographic associations).

Coverage note — None was omitted; all core contributions, algorithmic procedures, derivations, quantitative results across datasets and tasks (CelebA-HQ, ImageNet, 512x512 inpainting, super-resolution), ablations, and stated limitations are fully represented.

References

  1. 1.Avrahami, O., Lischinski, D., and Fried, O. Blended diffusion for text-driven editing of natural images. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 18187–18197, 2021.
  2. 2.Bansal, A., Borgnia, E., Chu, H.-M., Li, J., Kazemi, H., Huang, F., Goldblum, M., Geiping, J., and Goldstein, T. Cold diffusion: Inverting arbitrary image transforms without noise. ArXiv, abs/2208.09392, 2022.
  3. 3.Batzolis, G., Stanczuk, J., Schonlieb, C.-B., and Etmann, C. Non-uniform diffusion models. ArXiv, abs/2207.09786, 2022.
  4. 4.Benton, J., Shi, Y., Bortoli, V. D., Deligiannidis, G., and Doucet, A. From denoising diffusions to denoising markov models. ArXiv, abs/2211.03595, 2022.
  5. 5.Bond-Taylor, S., Hessey, P., Sasaki, H., Breckon, T., and Willcocks, C. G. Unleashing transformers: Parallel token prediction with discrete absorbing diffusion for fast high-resolution image generation from vector-quantized codes. In European Conference on Computer Vision, 2021.
  6. 6.Chung, H., Sim, B., and Ye, J.-C. Come-closer-diffuse-faster: Accelerating conditional diffusion models for inverse problems through stochastic contraction. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12403–12412, 2021.
  7. 7.Chung, H., Kim, J., Mccann, M. T., Klasky, M. L., and Ye, J. C. Diffusion posterior sampling for general noisy inverse problems. arXiv preprint arXiv:2209.14687, 2022a.
  8. 8.Chung, H., Sim, B., Ryu, D., and Ye, J. C. Improving diffusion models for inverse problems using manifold constraints. ArXiv, abs/2206.00941, 2022b.
  9. 9.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. ArXiv, abs/2105.05233, 2021a.
  10. 10.Dhariwal, P. and Nichol, A. Diffusion models beat gans on image synthesis. Advances in Neural Information Processing Systems, 34:8780–8794, 2021b.
  11. 11.Guo, Z., Chen, Z., Yu, T., Chen, J., and Liu, S. Progressive image inpainting with full-resolution residual network. Proceedings of the 27th ACM International Conference on Multimedia, 2019.
  12. 12.Ho, J., Jain, A., and Abbeel, P. Denoising diffusion probabilistic models. ArXiv, abs/2006.11239, 2020.
  13. 13.Hong, X., Xiong, P., Ji, R., and Fan, H. Deep fusion network for image completion. Proceedings of the 27th ACM International Conference on Multimedia, 2019.
  14. 14.Horita, D., Yang, J., Chen, D., Koyama, Y., and Aizawa, K. A structure-guided diffusion model for large-hole diverse image completion. ArXiv, abs/2211.10437, 2022.
  15. 15.Horwitz, E. and Hoshen, Y. Conffusion: Confidence intervals for diffusion models. ArXiv, abs/2211.09795, 2022.
  16. 16.Iizuka, S., Simo-Serra, E., and Ishikawa, H. Globally and locally consistent image completion. ACM Transactions on Graphics (TOG), 36:1 – 14, 2017.
  17. 17.Kawar, B., Elad, M., Ermon, S., and Song, J. Denoising diffusion restoration models. arXiv preprint arXiv:2201.11793, 2022.
  18. 18.Krizhevsky, A. One weird trick for parallelizing convolutional neural networks. ArXiv, abs/1404.5997, 2014.
  19. 19.Ku, W.-F., Siu, W. C., Cheng, X., and Chan, H. A. Intelligent painter: Picture composition with resampling diffusion model. ArXiv, abs/2210.17106, 2022.
  20. 20.Li, W., Yu, X., Zhou, K., Song, Y., Lin, Z., and Jia, J. Sdm: Spatial diffusion model for large hole image inpainting. ArXiv, abs/2212.02963, 2022.
  21. 21.Liu, E. Z., Haghgoo, B., Chen, A. S., Raghunathan, A., Koh, P. W., Sagawa, S., Liang, P., and Finn, C. Just train twice: Improving group robustness without training group information. In International Conference on Machine Learning, pp. 6781–6792. PMLR, 2021.
  22. 22.Liu, G., Reda, F. A., Shih, K. J., Wang, T.-C., Tao, A., and Catanzaro, B. Image inpainting for irregular holes using partial convolutions. In European Conference on Computer Vision, 2018.
  23. 23.Liu, H., Jiang, B., Song, Y., Huang, W., and Yang, C. Correction to: Rethinking image inpainting via a mutual encoder-decoder with feature equalizations. Computer Vision – ECCV 2020, 12347:C1 – C1, 2020.
  24. 24.Liu, H., Wang, Y., Wang, M., and Rui, Y. Delving globally into texture and structure for image inpainting. Proceedings of the 30th ACM International Conference on Multimedia, 2022.
  25. 25.Liu, Z., Luo, P., Wang, X., and Tang, X. Deep learning face attributes in the wild. 2015 IEEE International Conference on Computer Vision (ICCV), pp. 3730–3738, 2014.
  26. 26.Lugmayr, A., Danelljan, M., Romero, A., Yu, F., Timofte, R., and Gool, L. V. Repaint: Inpainting using denoising diffusion probabilistic models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11451–11461, 2022.
  27. 27.Nazeri, K., Ng, E., Joseph, T., Qureshi, F. Z., and Ebrahimi, M. Edgeconnect: Structure guided image inpainting using edge prediction. 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pp. 3265–3274, 2019.
  28. 28.Neal, R. M. Mcmc using hamiltonian dynamics. arXiv: Computation, pp. 139–188, 2011.
  29. 29.Nichol, A., Dhariwal, P., Ramesh, A., Shyam, P., Mishkin, P., McGrew, B., Sutskever, I., and Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. In International Conference on Machine Learning, 2021.
  30. 30.Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A. A. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2536–2544, 2016.
  31. 31.Peng, J., Liu, D., Xu, S., and Li, H. Generating diverse structure for image inpainting with hierarchical vq-vae. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10770–10779, 2021.
  32. 32.Pokle, A., Geng, Z., and Kolter, Z. Deep equilibrium approaches to diffusion models. ArXiv, abs/2210.12867, 2022.
  33. 33.Reddy, V. R., Priya, B. L., Vinuthna, P., Reddy, K. P., and Reddy, D. S. Exploration of image inpainting approaches and challenges: A survey. International Journal of Computer Engineering in Research Trends, 2022.
  34. 34.Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with latent diffusion models. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10674–10685, 2021.
  35. 35.Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115(3): 211–252, 2015.
  36. 36.Saharia, C., Chan, W., Chang, H., Lee, C. A., Ho, J., Salimans, T., Fleet, D. J., and Norouzi, M. Palette: Image-to-image diffusion models. ACM SIGGRAPH 2022 Conference Proceedings, 2021a.
  37. 37.Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D. J., and Norouzi, M. Image super-resolution via iterative refinement. IEEE transactions on pattern analysis and machine intelligence, PP, 2021b.
  38. 38.Shah, R., Gautam, A., and Singh, S. K. Overview of image inpainting techniques: A survey. 2022 IEEE Region 10 Symposium (TENSYMP), pp. 1–6, 2022.
  39. 39.Sohl-Dickstein, J. N., Weiss, E. A., Maheswaranathan, N., and Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. ArXiv, abs/1503.03585, 2015.
  40. 40.Song, J., Meng, C., and Ermon, S. Denoising diffusion implicit models. ArXiv, abs/2010.02502, 2020.
  41. 41.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. ArXiv, abs/1907.05600, 2019a.
  42. 42.Song, Y. and Ermon, S. Generative modeling by estimating gradients of the data distribution. ArXiv, abs/1907.05600, 2019b.
  43. 43.Song, Y., Yang, C., Shen, Y., Wang, P., Huang, Q., and Kuo, C.-C. J. Spg-net: Segmentation prediction and guidance network for image inpainting. ArXiv, abs/1805.03356, 2018.
  44. 44.Suvorov, R., Logacheva, E., Mashikhin, A., Remizova, A., Ashukha, A., Silvestrov, A., Kong, N., Goka, H., Park, K., and Lempitsky, V. Resolution-robust large mask inpainting with fourier convolutions. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 2149–2159, 2022.
  45. 45.Trippe, B. L., Yim, J., Tischer, D. K., Broderick, T., Baker, D., Barzilay, R., and Jaakkola, T. Diffusion probabilistic modeling of protein backbones in 3d for the motif-scaffolding problem. ArXiv, abs/2206.04119, 2022.
  46. 46.Vo, H. V., Duong, N. Q. K., and Pérez, P. Structural inpainting. Proceedings of the 26th ACM international conference on Multimedia, 2018.
  47. 47.Wan, Z., Zhang, J., Chen, D., and Liao, J. High-fidelity pluralistic image completion with transformers. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 4672–4681, 2021.
  48. 48.Wang, Y., Yu, J., and Zhang, J. Zero-shot image restoration using denoising diffusion null-space model. arXiv preprint arXiv:2212.00490, 2022.
  49. 49.Weng, Y., Ding, S., and Zhou, T. A survey on improved gan based image inpainting. 2022 2nd International Conference on Consumer Electronics and Computer Engineering (ICCECE), pp. 319–322, 2022.
  50. 50.Whang, J., Delbracio, M., Talebi, H., Saharia, C., Dimakis, A. G., and Milanfar, P. Deblurring via stochastic refinement. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 16272–16282, 2021.
  51. 51.Xiang, H., Zou, Q., Nawaz, M. A., Huang, X., Zhang, F., and Yu, H. Deep learning for image inpainting: A survey. Pattern Recognit., 134:109046, 2022.
  52. 52.Xiao, Q., Li, G., and Chen, Q. Deep inception generative network for cognitive image inpainting. ArXiv, abs/1812.01458, 2018.
  53. 53.Yang, L., Zhang, Z., Hong, S., Xu, R., Zhao, Y., Shao, Y., Zhang, W., Yang, M.-H., and Cui, B. Diffusion models: A comprehensive survey of methods and applications. ArXiv, abs/2209.00796, 2022.
  54. 54.Yu, Y., Zhan, F., Wu, R., Pan, J., Cui, K., Lu, S., Ma, F., Xie, X., and Miao, C. Diverse image inpainting with bidirectional and autoregressive transformers. Proceedings of the 29th ACM International Conference on Multimedia, 2021.
  55. 55.Zhao, L., Mo, Q., Lin, S., Wang, Z., Zuo, Z., Chen, H., Xing, W., and Lu, D. Uctgan: Diverse image inpainting based on unsupervised cross-space translation. 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 5740–5749, 2020.
  56. 56.Zhao, S., Cui, J., Sheng, Y., Dong, Y., Liang, X., Chang, E. I.-C., and Xu, Y. Large scale image completion via co-modulated generative adversarial networks. ArXiv, abs/2103.10428, 2021.
  57. 57.Zheng, C., Cham, T.-J., and Cai, J. Pluralistic image completion. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1438–1447, 2019.

Citation

MLA
Zhang, G., et al. “Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models”. International Conference on Machine Learning, vol. 202, 2023, pp. 41164–93, https://proceedings.mlr.press/v202/zhang23q.html.
APA
Zhang, G., Ji, J., Zhang, Y., Yu, M., Jaakkola, T., & Chang, S. (2023). Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models. International Conference on Machine Learning, 202, 41164–41193. https://proceedings.mlr.press/v202/zhang23q.html
Chicago
Zhang, G., J. Ji, Y. Zhang, M. Yu, T. Jaakkola, and S. Chang. 2023. “Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models”. International Conference on Machine Learning 202: 41164–93. https://proceedings.mlr.press/v202/zhang23q.html.
Harvard
Zhang, G. et al. (2023) “Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models”, International Conference on Machine Learning. PMLR, pp. 41164–41193. Available at: https://proceedings.mlr.press/v202/zhang23q.html.
Vancouver
1. Zhang G, Ji J, Zhang Y, Yu M, Jaakkola T, Chang S (2023) Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models. In: International Conference on Machine Learning. PMLR, pp 41164–41193

BibTeX

@InProceedings{pmlr-v202-zhang23q,
  title = 	 {Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models},
  author =       {Zhang, Guanhua and Ji, Jiabao and Zhang, Yang and Yu, Mo and Jaakkola, Tommi and Chang, Shiyu},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {41164--41193},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/zhang23q/zhang23q.pdf},
  url = 	 {https://proceedings.mlr.press/v202/zhang23q.html},
  abstract = 	 {Image inpainting refers to the task of generating a complete, natural image based on a partially revealed reference image. Recently, many research interests have been focused on addressing this problem using fixed diffusion models. These approaches typically directly replace the revealed region of the intermediate or final generated images with that of the reference image or its variants. However, since the unrevealed regions are not directly modified to match the context, it results in incoherence between revealed and unrevealed regions. To address the incoherence problem, a small number of methods introduce a rigorous Bayesian framework, but they tend to introduce mismatches between the generated and the reference images due to the approximation errors in computing the posterior distributions. In this paper, we propose CoPaint, which can coherently inpaint the whole image without introducing mismatches. CoPaint also uses the Bayesian framework to jointly modify both revealed and unrevealed regions but approximates the posterior distribution in a way that allows the errors to gradually drop to zero throughout the denoising steps, thus strongly penalizing any mismatches with the reference image. Our experiments verify that CoPaint can outperform the existing diffusion-based methods under both objective and subjective metrics.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/