Globally and locally consistent image completion
SATOSHI IIZUKAEDGAR SIMO-SERRAHIROSHI ISHIKAWA
Proposes a fully convolutional image completion framework trained with dual global and local context discriminators, enabling realistic synthesis of arbitrary-shaped missing regions while preserving both fine local details and overall semantic coherence across diverse scenes.
Image completion—the process of seamlessly filling in missing, damaged, or unwanted regions within photographs—is essential for digital media editing, object removal, and computer vision applications. Traditional techniques often struggle to generate plausible content for large missing regions because they merely copy existing patches from elsewhere in the image, failing when required features do not already exist in the source. Meanwhile, early deep-learning approaches frequently produced blurry outputs, failed to maintain consistency with surrounding details, and were confined to fixed, low-resolution masks.
The article introduces a deep learning framework capable of completing missing regions of arbitrary shapes and resolutions within diverse images. The objective was to demonstrate that a fully convolutional neural network, trained using dual local and global context discriminators, can synthesize entirely novel visual structures while preserving overall scene harmony and fine detail.
The researchers developed an architecture composed of a completion network and two auxiliary discriminator networks used during training. The completion network utilizes dilated convolutional layers to capture broad contextual information across a 303x303-pixel support area without sacrificing resolution. To guide realistic synthesis, a global discriminator evaluates the overall scene coherence across the entire image, while a local discriminator focuses specifically on a 128x128-pixel window around the filled region. The system was trained on approximately 8.1 million images from the Places2 dataset over a two-month period, complemented by fine-tuning on specialized datasets for human faces (CelebA) and architectural facades (CMP Facade), followed by standard color blending post-processing.
The experimental findings demonstrate significant performance advantages over existing methods. First, the dual-discriminator setup successfully generates entirely new image fragments—such as missing eyes, noses, or architectural elements—which patch-based tools cannot do. Second, in a blind user study evaluating the naturalness of completed facial images, participants perceived the generated faces as real 77% of the time, approaching the 96.5% baseline rating for authentic photographs. Third, the fully convolutional model processes images across arbitrary resolutions with high operational efficiency, completing large 1024x1024-pixel images in 0.56 seconds on a standard graphics processing unit, representing a roughly 15-fold speedup over central processing unit execution.
These results show that combining global scene comprehension with local texture validation resolves the trade-off between image sharpness and contextual logic. For organizations developing automated photo editing, visual restoration, or computer-generated graphics workflows, this approach reduces the labor needed for complex manual touch-ups and delivers consistent, production-ready outputs at interactive speeds without requiring per-image optimization.
To adopt and build on this technology, organizations should deploy GPU-accelerated pipelines for image repair tools and fine-tune specialized models on domain-specific datasets when dealing with recurring structured subjects like portraits or architectural assets. Further research and development should explore expanded network receptive fields to improve performance on large-scale hole extrapolation at image borders, where contextual cues are limited to a single side.
The primary operational limitations stem from fixed receptive field boundaries and structured semantics. While the network reliably repairs diverse textures and landscapes, it struggles with extremely large masks that exceed its spatial support and can fail on complex, heavily structured subjects, such as attempting to recreate partially occluded animals or human bodies against detailed backgrounds.
- Paper: Generative Adversarial Networks, Ian J. Goodfellow et al. (2014). Reading the foundational Generative Adversarial Networks paper provides the essential adversarial loss framework built upon by the source article's global and local context discriminators.
- Paper: Context Encoders: Feature Learning by Inpainting, Deepak Pathak et al. (2016). Understanding context encoders for unsupervised feature learning by inpainting establishes the primary baseline and task formulation that the source paper extends.
- Paper: PatchMatch: a randomized correspondence algorithm for structural image editing, Connelly Barnes et al. (2009). Familiarity with the PatchMatch randomized correspondence algorithm clarifies the traditional patch-based methods that the source paper contrasts against and improves upon.
- Paper: Generative Image Inpainting with Contextual Attention, Jiahui Yu et al. (2018). This paper builds directly upon the source's image completion framework by introducing a contextual attention layer that further improves the synthesis of complex structures.
- Paper: Image Inpainting for Irregular Holes Using Partial Convolutions, Guilin Liu et al. (2018). This work extends the source's image completion concepts to handle irregular holes using partial convolutions in a single feedforward pass.
- Paper: Image-to-Image Translation with Conditional Adversarial Networks, Phillip Isola et al. (2017). This paper applies conditional adversarial networks broadly to image-to-image translation tasks, continuing the adversarial loss principles explored in the source.
- Paper: Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network, Christian Ledig et al. (2017). This research continues the use of generative adversarial networks with perceptual losses to achieve photo-realistic results in related image restoration tasks.
