MAT: Mask-Aware Transformer for Large Hole Image Inpainting
Wenbo LiZhe LinKun ZhouLu QiYi WangJiaya Jia
Develops a mask-aware transformer that restricts attention computation to valid image tokens, enabling efficient high-resolution inpainting of large missing regions with high fidelity and diversity.
Image inpainting—the process of filling in missing, damaged, or masked areas of an image—is essential for digital photo restoration, object removal, and automated content editing. However, completing large missing regions in complex, high-resolution images remains a significant challenge. Traditional convolutional neural networks struggle to capture distant context across large image gaps, while standard transformer-based attention models suffer from prohibitive computational costs and unstable training when high proportions of pixel tokens are masked out or invalid.
The article introduces and evaluates the Mask-Aware Transformer, an architecture designed to perform high-resolution image completion across large missing areas while supporting pluralistic generation, which refers to the ability to produce multiple visually realistic and diverse outputs for a single input. The authors investigate whether combining convolutional layers with a customized transformer mechanism can establish long-range semantic relationships efficiently and stably without requiring costly low-resolution bottlenecks or heavy pre-training.
The proposed method adopts a multi-stage hybrid design evaluated on standard benchmark datasets, including Places365-Standard and CelebA-HQ at resolutions up to 512x512 and 1024x1024. A convolutional head extracts visual tokens, which are processed by an adjusted transformer body containing multi-head contextual attention guided by a dynamic mask. This mask updates iteratively so attention calculations only compute relationships between valid, informative tokens. The system replaces standard layer normalization and residual connections with feature concatenation to prevent gradient instability during adversarial training. A style manipulation module then modulates convolutional layers to inject diversity, followed by a convolutional refinement network that sharpens local textures.
The evaluation yields several key findings demonstrating superior performance and efficiency. First, the proposed model sets a new state-of-the-art across benchmark datasets; under large mask conditions on Places (512x512), it achieves a Fréchet Inception Distance of 1.96 compared to 2.92 for the leading alternative, CoModGAN. Second, the system achieves these gains with high parameter and data efficiency, operating at 62 million parameters—roughly 43% fewer than CoModGAN’s 109 million—while requiring only 1.8 million training images to match or outperform models trained on 4.5 to 8 million images. Third, ablation analyses confirm that restricting attention to valid tokens and replacing conventional transformer residual blocks with concatenation improves perceptual quality and stabilizes adversarial training. Fourth, the style manipulation module successfully delivers distinct, plausible image variations without compromising structural consistency.
These findings indicate that hybrid convolutional-transformer architectures can eliminate the traditional trade-off between computational overhead and global contextual reasoning in image synthesis. For organizations deploying computer vision tools in media production, digital restoration, or content moderation, this approach reduces the compute infrastructure and training data volume needed to achieve high-fidelity generative editing. By supporting varied, realistic completions, the framework also enhances creative workflows where multiple plausible variations are preferred over a single deterministic output.
Organizations evaluating this technology should pilot the model for automated editing and content synthesis pipelines, particularly where large image areas must be reconstructed. Subsequent development should focus on extending structural awareness to complex articulated shapes, such as dynamic animals or non-rigid objects, where the current model occasionally falters due to a lack of explicit semantic annotations. Additionally, operational workflows must account for fixed windowing requirements, such as padding inputs to standard dimensions, to maximize visual fidelity in production environments.
- Paper: Generative Image Inpainting with Contextual Attention, Jiahui Yu et al. (2018). This foundational work introduces the contextual attention mechanism for inpainting, which MAT adapts and refines into a mask-aware transformer formulation.
- Paper: Image Inpainting for Irregular Holes Using Partial Convolutions, Guilin Liu et al. (2018). It establishes the concept of dynamic mask updating and conditioning operations strictly on valid pixels, directly motivating MAT's dynamic mask-aware attention.
- Paper: Free-Form Image Inpainting With Gated Convolution, Jiahui Yu et al. (2018). It presents learnable gating and free-form hole filling in deep inpainting architectures, foundational techniques upon which MAT's transformer framework improves.
- Paper: Globally and locally consistent image completion, SATOSHI IIZUKA et al. (2017). It introduces global and local discriminator architectures and benchmarks for large-hole image completion that MAT uses for adversarial evaluation.
- Paper: Taming Transformers for High-Resolution Image Synthesis, Patrick Esser et al. (2020). It demonstrates combining convolutional token representations with transformer backbones for high-resolution image synthesis, providing the architectural paradigm adopted by MAT.
- Paper: A Style-Based Generator Architecture for Generative Adversarial Networks, Tero Karras et al. (2019). It outlines style modulation techniques in generative convolutional networks, serving as the basis for MAT's style manipulation module to enable pluralistic generation.
- Paper: SwinIR: Image Restoration Using Swin Transformer, Jingyun Liang et al. (2021). It provides a key benchmark for adapting vision transformer architectures to dense image restoration tasks using hybrid convolutional-attention blocks.
- Paper: Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models, Guanhua Zhang et al. (2023). This work advances beyond transformer-GAN hybrid inpainting architectures like MAT by utilizing denoising diffusion implicit models to achieve coherent and seamless hole filling.
- Paper: RePaint: Inpainting using Denoising Diffusion Probabilistic Models, Andreas Lugmayr et al. (2022). It extends free-form image completion across arbitrary mask distributions by employing conditioning strategies within iterative diffusion models rather than feed-forward transformers.
- Paper: Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting, Su Wang et al. (2023). It extends masked image editing to high-resolution multimodal settings by combining text prompts with object-mask guidance and establishing rigorous evaluation benchmarks.
- Paper: pix2gestalt: Amodal Segmentation by Synthesizing Wholes, Ege Ozguroglu et al. (2024). It applies generative completion principles from inpainting to the challenging domain of zero-shot amodal perception and whole-object synthesis.
- Paper: RandAR: Decoder-only Autoregressive Visual Generation in Random Orders, Ziqi Pang et al. (2025). It generalizes masked token generation in vision transformers by exploring random-order autoregressive modeling for zero-shot visual editing.
