Built independently by an author, for readers. Read the story and support ChapterPal

keyword

image completion

Image completion, also commonly referred to as image inpainting, is a computer vision and digital image processing task that involves synthesizing visually plausible and semantically consistent content to fill in missing, corrupted, or masked regions of an image. The primary objective is to reconstruct the missing areas so that the synthesized content seamlessly integrates with the surrounding visual context in structure, color, and texture without introducing visible artifacts. Techniques for image completion range from classical patch-matching and texture-synthesis algorithms, which borrow and propagate pixel patterns from known regions of the image, to deep generative models such as convolutional neural networks, transformers, generative adversarial networks, and diffusion models, which leverage learned priors to invent complex objects and coherent scenes. This process is widely utilized in digital restoration, computational photography, watermark and object removal, and interactive image editing.

9 items

Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models

Guanhua Zhang, Jiabao Ji, Yang Zhang, Mo Yu, Tommi S. Jaakkola, Shiyu Chang

OrganizationsIBMMassachusetts Institute of TechnologyMIT-IBM Watson AI LabUniversity of California, Santa Barbara

Why you should read this

Proposes CoPaint, a Bayesian framework for diffusion-based image inpainting that jointly modifies revealed and unrevealed regions to eliminate incoherence while driving approximation errors to zero to strictly match reference constraints.

Image inpainting refers to the task of generating a complete, natural image based on a partially revealed reference image. Recently, many research interests have been focused on addressing this problem using fixed diffusion models. These approaches typically directly replace the revealed region of the intermediate or final generated images with that of the reference image or its variants. However, since the unrevealed regions are not directly modified to match the context, it results in incoherence between revealed and unrevealed regions. To address the incoherence problem, a small number of methods introduce a rigorous Bayesian framework, but they tend to introduce mismatches between the generated and the reference images due to the approximation errors in computing the posterior distributions. In this paper, we propose CoPaint, which can coherently inpaint the whole image without introducing mismatches. CoPaint also uses the Bayesian framework to jointly modify both revealed and unrevealed regions, but approximates the posterior distribution in a way that allows the errors to gradually drop to zero throughout the denoising steps, thus strongly penalizing any mismatches with the reference image. Our experiments verify that CoPaint can outperform the existing diffusion-based methods under both objective and subjective metrics. The codes are available at https://github.com/UCSB-NLP-Chang/CoPaint/.

Added

2026-09-26

Image Transformer

Image Transformer

Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, Dustin Tran

OrganizationsGoogleUniversity of California Berkeley

Why you should read this

Adapts the Transformer architecture to autoregressive image generation by restricting self-attention to local neighborhoods, outperforming convolutional networks in both density estimation on ImageNet and large-scale super-resolution.

Image generation has been successfully cast as an autoregressive sequence generation or transformation problem. Recent work has shown that self-attention is an effective way of modeling textual sequences. In this work, we generalize a recently proposed model architecture based on self-attention, the Transformer, to a sequence modeling formulation of image generation with a tractable likelihood. By restricting the self-attention mechanism to attend to local neighborhoods we significantly increase the size of images the model can process in practice, despite maintaining significantly larger receptive fields per layer than typical convolutional neural networks. While conceptually simple, our generative models significantly outperform the current state of the art in image generation on ImageNet, improving the best published negative log-likelihood on ImageNet from 3.83 to 3.77. We also present results on image super-resolution with a large magnification ratio, applying an encoder-decoder configuration of our architecture. In a human evaluation study, we find that images generated by our super-resolution model fool human observers three times more often than the previous state of the art.

Added

2026-09-18

Free-Form Image Inpainting With Gated Convolution

Free-Form Image Inpainting With Gated Convolution

Jiahui Yu, Zhe L. Lin, Jimei Yang, Xiaohui Shen, Xin Lu, Thomas S. Huang

OrganizationsAdobeByteDanceUniversity of Illinois Urbana-Champaign

Why you should read this

Proposes gated convolutions and a spectral-normalized patch discriminator to handle arbitrary mask shapes, enabling dynamic feature selection for realistic free-form image inpainting and user-guided editing.

We present a generative image inpainting system to complete images with free-form mask and guidance. The system is based on gated convolutions learned from millions of images without additional labelling efforts. The proposed gated convolution solves the issue of vanilla convolution that treats all input pixels as valid ones, generalizes partial convolution by providing a learnable dynamic feature selection mechanism for each channel at each spatial location across all layers. Moreover, as free-form masks may appear anywhere in images with any shape, global and local GANs designed for a single rectangular mask are not applicable. Thus, we also present a patch-based GAN loss, named SN-PatchGAN, by applying spectral-normalized discriminator on dense image patches. SN-PatchGAN is simple in formulation, fast and stable in training. Results on automatic image inpainting and user-guided extension demonstrate that our system generates higher-quality and more flexible results than previous methods. Our system helps user quickly remove distracting objects, modify image layouts, clear watermarks and edit faces. Code, demo and models are available at: this https URL

Added

2026-09-16

RePaint: Inpainting using Denoising Diffusion Probabilistic Models

RePaint: Inpainting using Denoising Diffusion Probabilistic Models

Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, Luc Van Gool

OrganizationsETH Zurich

Why you should read this

Proposes RePaint, an image inpainting method that conditions pretrained unconditional diffusion models during reverse sampling without retraining, achieving state-of-the-art results on arbitrary and extreme masks.

Free-form inpainting is the task of adding new content to an image in the regions specified by an arbitrary binary mask. Most existing approaches train for a certain distribution of masks, which limits their generalization capabilities to unseen mask types. Furthermore, training with pixel-wise and perceptual losses often leads to simple textural extensions towards the missing areas instead of semantically meaningful generation. In this work, we propose RePaint: A Denoising Diffusion Probabilistic Model (DDPM) based inpainting approach that is applicable to even extreme masks. We employ a pretrained unconditional DDPM as the generative prior. To condition the generation process, we only alter the reverse diffusion iterations by sampling the unmasked regions using the given image information. Since this technique does not modify or condition the original DDPM network itself, the model produces high-quality and diverse output images for any inpainting form. We validate our method for both faces and general-purpose image inpainting using standard and extreme masks. RePaint outperforms state-of-the-art Autoregressive, and GAN approaches for at least five out of six mask distributions. Github Repository: this http URL

Added

2026-09-16

Palette: Image-to-Image Diffusion Models

Palette: Image-to-Image Diffusion Models

Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David J. Fleet, Mohammad Norouzi

OrganizationsGoogle

Why you should read this

Develops a unified conditional diffusion framework that outperforms task-specific GAN baselines across colorization, inpainting, uncropping, and restoration without specialized architectures, auxiliary losses, or hyperparameter tuning.

This paper develops a unified framework for image-to-image translation based on conditional diffusion models and evaluates this framework on four challenging image-to-image translation tasks, namely colorization, inpainting, uncropping, and JPEG restoration. Our simple implementation of image-to-image diffusion models outperforms strong GAN and regression baselines on all tasks, without task-specific hyper-parameter tuning, architecture customization, or any auxiliary loss or sophisticated new techniques needed. We uncover the impact of an L2 vs. L1 loss in the denoising diffusion objective on sample diversity, and demonstrate the importance of self-attention in the neural architecture through empirical studies. Importantly, we advocate a unified evaluation protocol based on ImageNet, with human evaluation and sample quality scores (FID, Inception Score, Classification Accuracy of a pre-trained ResNet-50, and Perceptual Distance against original images). We expect this standardized evaluation protocol to play a role in advancing image-to-image translation research. Finally, we show that a generalist, multi-task diffusion model performs as well or better than task-specific specialist counterparts. Check out this https URL for an overview of the results.

Added

2026-09-15

Generative Image Inpainting with Contextual Attention

Generative Image Inpainting with Contextual Attention

Jiahui Yu, Zhe L. Lin, Jimei Yang, Xiaohui Shen, Xin Lu, Thomas S. Huang

OrganizationsAdobeUniversity of Illinois Urbana-Champaign

Why you should read this

Introduces a contextual attention mechanism for deep generative image inpainting, enabling feed-forward networks to explicitly borrow distant image features for reconstructing missing regions with sharp textures and realistic structures.

Recent deep learning based approaches have shown promising results for the challenging task of inpainting large missing regions in an image. These methods can generate visually plausible image structures and textures, but often create distorted structures or blurry textures inconsistent with surrounding areas. This is mainly due to ineffectiveness of convolutional neural networks in explicitly borrowing or copying information from distant spatial locations. On the other hand, traditional texture and patch synthesis approaches are particularly suitable when it needs to borrow textures from the surrounding regions. Motivated by these observations, we propose a new deep generative model-based approach which can not only synthesize novel image structures but also explicitly utilize surrounding image features as references during network training to make better predictions. The model is a feed-forward, fully convolutional neural network which can process images with multiple holes at arbitrary locations and with variable sizes during the test time. Experiments on multiple datasets including faces (CelebA, CelebA-HQ), textures (DTD) and natural images (ImageNet, Places2) demonstrate that our proposed approach generates higher-quality inpainting results than existing ones. Code, demo and models are available at: this https URL.

Added

2026-09-14

PatchMatch: a randomized correspondence algorithm for structural image editing

PatchMatch: a randomized correspondence algorithm for structural image editing

Connelly Barnes, Eli Shechtman, Adam Finkelstein, Dan B Goldman

OrganizationsAdobePrinceton UniversityUniversity of Washington

Why you should read this

Introduces a fast randomized correspondence algorithm that rapidly finds approximate nearest-neighbor patch matches across images, accelerating structural editing tasks such as inpainting, reshuffling, and retargeting to interactive speeds.

This paper presents interactive image editing tools using a new randomized algorithm for quickly finding approximate nearest-neighbor matches between image patches. Previous research in graphics and vision has leveraged such nearest-neighbor searches to provide a variety of high-level digital image editing tools. However, the cost of computing a field of such matches for an entire image has eluded previous efforts to provide interactive performance. Our algorithm offers substantial performance improvements over the previous state of the art (20-100x), enabling its use in interactive editing tools. The key insights driving the algorithm are that some good patch matches can be found via random sampling, and that natural coherence in the imagery allows us to propagate such matches quickly to surrounding areas. We offer theoretical analysis of the convergence properties of the algorithm, as well as empirical and practical evidence for its high quality and performance. This one simple algorithm forms the basis for a variety of tools – image retargeting, completion and reshuffling – that can be used together in the context of a high-level image editing application. Finally, we propose additional intuitive constraints on the synthesis process that offer the user a level of control unavailable in previous methods.

Added

2026-09-11