MAT: Mask-Aware Transformer for Large Hole Image Inpainting

Wenbo LiZhe LinKun ZhouLu QiYi WangJiaya Jia

article2022CVPR461 citationsBest Paper Finalist

Develops a mask-aware transformer that restricts attention computation to valid image tokens, enabling efficient high-resolution inpainting of large missing regions with high fidelity and diversity.

Listen

Image inpainting—the process of filling in missing, damaged, or masked areas of an image—is essential for digital photo restoration, object removal, and automated content editing. However, completing large missing regions in complex, high-resolution images remains a significant challenge. Traditional convolutional neural networks struggle to capture distant context across large image gaps, while standard transformer-based attention models suffer from prohibitive computational costs and unstable training when high proportions of pixel tokens are masked out or invalid.

The article introduces and evaluates the Mask-Aware Transformer, an architecture designed to perform high-resolution image completion across large missing areas while supporting pluralistic generation, which refers to the ability to produce multiple visually realistic and diverse outputs for a single input. The authors investigate whether combining convolutional layers with a customized transformer mechanism can establish long-range semantic relationships efficiently and stably without requiring costly low-resolution bottlenecks or heavy pre-training.

The proposed method adopts a multi-stage hybrid design evaluated on standard benchmark datasets, including Places365-Standard and CelebA-HQ at resolutions up to 512x512 and 1024x1024. A convolutional head extracts visual tokens, which are processed by an adjusted transformer body containing multi-head contextual attention guided by a dynamic mask. This mask updates iteratively so attention calculations only compute relationships between valid, informative tokens. The system replaces standard layer normalization and residual connections with feature concatenation to prevent gradient instability during adversarial training. A style manipulation module then modulates convolutional layers to inject diversity, followed by a convolutional refinement network that sharpens local textures.

The evaluation yields several key findings demonstrating superior performance and efficiency. First, the proposed model sets a new state-of-the-art across benchmark datasets; under large mask conditions on Places (512x512), it achieves a Fréchet Inception Distance of 1.96 compared to 2.92 for the leading alternative, CoModGAN. Second, the system achieves these gains with high parameter and data efficiency, operating at 62 million parameters—roughly 43% fewer than CoModGAN’s 109 million—while requiring only 1.8 million training images to match or outperform models trained on 4.5 to 8 million images. Third, ablation analyses confirm that restricting attention to valid tokens and replacing conventional transformer residual blocks with concatenation improves perceptual quality and stabilizes adversarial training. Fourth, the style manipulation module successfully delivers distinct, plausible image variations without compromising structural consistency.

These findings indicate that hybrid convolutional-transformer architectures can eliminate the traditional trade-off between computational overhead and global contextual reasoning in image synthesis. For organizations deploying computer vision tools in media production, digital restoration, or content moderation, this approach reduces the compute infrastructure and training data volume needed to achieve high-fidelity generative editing. By supporting varied, realistic completions, the framework also enhances creative workflows where multiple plausible variations are preferred over a single deterministic output.

Organizations evaluating this technology should pilot the model for automated editing and content synthesis pipelines, particularly where large image areas must be reconstructed. Subsequent development should focus on extending structural awareness to complex articulated shapes, such as dynamic animals or non-rigid objects, where the current model occasionally falters due to a lack of explicit semantic annotations. Additionally, operational workflows must account for fixed windowing requirements, such as padding inputs to standard dimensions, to maximize visual fidelity in production environments.

Cover for MAT: Mask-Aware Transformer for Large Hole Image Inpainting

Abstract

Recent studies have shown the importance of modeling long-range interactions in the inpainting problem. To achieve this goal, existing approaches exploit either standalone attention techniques or transformers, but usually under a low resolution in consideration of computational cost. In this paper, we present a novel transformer-based model for large hole inpainting, which unifies the merits of transformers and convolutions to efficiently process high-resolution images. We carefully design each component of our framework to guarantee the high fidelity and diversity of recovered images. Specifically, we customize an inpainting-oriented transformer block, where the attention module aggregates non-local information only from partial valid tokens, indicated by a dynamic mask. Extensive experiments demonstrate the state-of-the-art performance of the new model on multiple benchmark datasets. Code is released at this https URL.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Method
  • 3.1 Overall Architecture
  • 3.2 Convolutional Head
  • 3.3 Transformer Body
  • 3.3.1 Adjusted Transformer Block
  • 3.3.2 Multi-Head Contextual Attention
  • 3.4 Style Manipulation Module
  • 3.5 Loss Functions
  • 4 Experiments
  • 4.1 Datasets and Metrics
  • 4.2 Implementation Details
  • 4.3 Ablation Study
  • 4.4 Comparison with State of the Arts
  • 4.5 Pluralistic Generation
  • 4.6 Limitations and Failure Cases
  • 5 Conclusion
  • References
  • A Network Architecture
  • B Free-Form Mask Sampling and Statistics
  • C Tokenization
  • D Model Configuration
  • E CelebA-HQ 256 ×\times 256 Results
  • F LPIPS Results
  • G Generalization to A Higher Resolution
  • H Diversity-Fidelity Tradeoff
  • I Additional Qualitative Results
  • J Licenses of Face Images

Knowls

  1. Knowl 1 — Mask-Aware Transformer Architecture for Large Hole Inpainting

    model/method

    The Mask-Aware Transformer (MAT) is an end-to-end framework designed for high-resolution (512×512512 \times 512 or higher) pluralistic image completion with large missing regions. The architecture is organized in two stages:

    1. First-Stage Transformer-Convolution Pipeline:

      • Convolutional Head: Takes the masked image IM=I⊙M∈RH×W×3\mathbf{I}_M = \mathbf{I} \odot \mathbf{M} \in \mathbb{R}^{H \times W \times 3} and the binary mask M∈{0,1}H×W×1\mathbf{M} \in \{0, 1\}^{H \times W \times 1} (where 1 indicates valid pixels and 0 indicates holes). An initial convolution maps the 4-channel input to 180 channels, followed by three stride-2 convolutions downsampling the spatial dimension to H8×W8\frac{H}{8} \times \frac{W}{8} to produce 180-dimensional visual tokens.
      • Transformer Body: Five stages of adjusted transformer blocks with block numbers {2,3,4,3,2}\{2, 3, 4, 3, 2\} operating at resolutions {H8×W8,H16×W16,H32×W32,H16×W16,H8×W8}\{\frac{H}{8} \times \frac{W}{8}, \frac{H}{16} \times \frac{W}{16}, \frac{H}{32} \times \frac{W}{32}, \frac{H}{16} \times \frac{W}{16}, \frac{H}{8} \times \frac{W}{8}\}. Convolutions perform downsampling and upsampling between stages. Positional embeddings are excluded since convolutions supply adequate spatial inductive bias. Long-range relationships are modeled using Multi-Head Contextual Attention (MCA).
      • Convolutional Reconstruction Tail: Upsamples the output tokens from H8×W8\frac{H}{8} \times \frac{W}{8} to the original H×WH \times W image space, producing a coarse completed image.
    2. Second-Stage Refinement (Conv-U-Net):

      • Takes the coarse prediction and the input mask, downsamples features to H32×W32\frac{H}{32} \times \frac{W}{32}, and upsamples them back to H×WH \times W via a symmetric U-Net with skip connections to refine high-frequency details. Convolution channel capacities start at 64, doubling per downsampling step up to 512.
      • Incorporates a Style Manipulation Module across decoding layers in both stages to inject diversity.
  2. Knowl 2 — Multi-Head Contextual Attention with Dynamic Mask Updating

    model/method

    Multi-Head Contextual Attention (MCA) adapts self-attention for image completion by restricting token aggregation to valid tokens and dynamically tracking token validity across successive layers.

    Attention Computation

    Given query Q\mathbf{Q}, key K\mathbf{K}, and value V\mathbf{V} matrices with per-head feature dimension dkd_k, the contextual attention is formulated as: Att(Q,K,V)=Softmax(QKT+M′dk)V\text{Att}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{Softmax}\left(\frac{\mathbf{Q}\mathbf{K}^T + \mathbf{M}'}{\sqrt{d_k}}\right)\mathbf{V} where M′\mathbf{M}' is an attention mask matrix whose elements are defined by: Mij′={0,if token j is valid−τ,if token j is invalid\mathbf{M}'_{ij} = \begin{cases} 0, & \text{if token } j \text{ is valid} \\ -\tau, & \text{if token } j \text{ is invalid} \end{cases} with τ=100\tau = 100. This penalizes attention weights corresponding to invalid tokens to virtually zero, preventing invalid hole tokens from contaminating representations.

    Shifted Window and Mask Updating Strategy

    Tokens are partitioned into local w×ww \times w windows (window sizes {8,16,16,16,8}\{8, 16, 16, 16, 8\} across the five transformer stages). After each attention layer, window positions shift by (⌊w2⌋,⌊w2⌋)(\lfloor \frac{w}{2} \rfloor, \lfloor \frac{w}{2} \rfloor) pixels to enable cross-window connections.

    The dynamic mask M′\mathbf{M}' is initialized from the downsampled input mask and updated layer-by-layer according to the propagation rule: all tokens within a window become marked as valid after attention if that window contained at least one valid token prior to attention. If all tokens in a window are invalid, they remain invalid. Over successive window shifts and attention layers, the mask progressively becomes fully valid.

  3. Knowl 3 — Adjusted Transformer Block with Fusion Learning

    model/method

    Standard transformer blocks (Layer Normalization →\to Multi-Head Self-Attention →\to Layer Normalization →\to Feed-Forward Network with residual connections) suffer from gradient explosion and instability during adversarial training on large missing masks, because Layer Normalization scales uninformative zero-filled tokens and additive residual connections encourage learning high-frequency details prior to establishing low-frequency structures.

    The adjusted transformer block resolves this by removing Layer Normalization and replacing residual additions with concatenation-based fusion learning: Xk,ℓ′=FC([MCA(Xk,ℓ−1),  Xk,ℓ−1])\mathbf{X}'_{k, \ell} = \text{FC}\left(\left[\text{MCA}(\mathbf{X}_{k, \ell-1}), \; \mathbf{X}_{k, \ell-1}\right]\right) Xk,ℓ=MLP(Xk,ℓ′)\mathbf{X}_{k, \ell} = \text{MLP}(\mathbf{X}'_{k, \ell}) where Xk,ℓ−1\mathbf{X}_{k, \ell-1} and Xk,ℓ\mathbf{X}_{k, \ell} denote the input and output token representations of the ℓ\ell-th block in the kk-th stage, [⋅,⋅][\cdot, \cdot] represents feature concatenation along the channel dimension, FC\text{FC} is a fully connected linear layer, and MLP\text{MLP} is a multi-layer perceptron.

    At the end of each transformer stage, a 3×33 \times 3 convolution with a global residual connection is applied.

  4. Knowl 4 — Style Manipulation Module for Pluralistic Inpainting

    model/method

    To enable diverse completions for large missing regions, the Style Manipulation Module generates and modulates convolution weights using stochastic noise inputs combined with image features.

    1. Style Representation Extraction:

      • A noise vector n\mathbf{n} is transformed into an unconditional style vector su\mathbf{s}_u via mapping network E\mathcal{E}: su=E(n)\mathbf{s}_u = \mathcal{E}(\mathbf{n})
      • Image feature X\mathbf{X} is fused with spatially resized su\mathbf{s}_u using a random binary mask B\mathbf{B} where elements are drawn with Bernoulli probability pp: X′=B⊙X+(1−B)⊙Resize(su)\mathbf{X}' = \mathbf{B} \odot \mathbf{X} + (\mathbf{1} - \mathbf{B}) \odot \text{Resize}(\mathbf{s}_u)
      • An image-conditional style sc\mathbf{s}_c is extracted via mapping network F\mathcal{F}: sc=F(X′)\mathbf{s}_c = \mathcal{F}(\mathbf{X}')
      • The styles are merged via fusion network A\mathcal{A}: s=A(su,sc)\mathbf{s} = \mathcal{A}(\mathbf{s}_u, \mathbf{s}_c)
    2. Weight Modulation and Demodulation: The convolution weights W∈RCout×Cin×Kh×Kw\mathbf{W} \in \mathbb{R}^{C_{\text{out}} \times C_{\text{in}} \times K_h \times K_w} in the reconstruction tail and refinement decoder are modulated and normalized by the combined style vector s\mathbf{s}: Wijk′=Wijk⋅si\mathbf{W}'_{ijk} = \mathbf{W}_{ijk} \cdot \mathbf{s}_i Wijk′′=Wijk′∑i,kWijk′2+ϵ\mathbf{W}''_{ijk} = \frac{\mathbf{W}'_{ijk}}{\sqrt{\sum_{i,k} {\mathbf{W}'_{ijk}}^2 + \epsilon}} where ii, jj, and kk index the input channels, output channels, and spatial kernel footprint, and ϵ>0\epsilon > 0 is a small constant preventing division by zero.

  5. Knowl 5 — Convolutional Tokenization Head vs. Linear Projection

    model/method

    Rather than using the patch-level linear projection typical of standard Vision Transformers, the token extraction head uses a stack of convolutional layers (one channel-expansion layer to 180 channels followed by three stride-2 downsampling convolutions to reach H8×W8\frac{H}{8} \times \frac{W}{8}).

    This provides two advantages for inpainting:

    1. Progressive Hole Filling: Convolutions propagate valid contextual information into missing regions early in the visual processing hierarchy, providing higher-fidelity tokens for transformer self-attention.
    2. Multi-Scale Skip Pathways: Downsampled intermediate convolutional representations can be passed directly to the decoding modules via shortcut connections.

    In empirical evaluations at 256×256256 \times 256, stacked convolutional tokenization achieved FID of 5.97, P-IDS of 13.17%, and U-IDS of 29.23%, whereas linear projection degraded to FID 10.54, P-IDS 5.77%, and U-IDS 20.86%.

  6. Knowl 6 — Loss Formulation for MAT Inpainting

    equation

    The overall training objective for the MAT generator avoids pixel-wise ℓ1\ell_1 or ℓ2\ell_2 reconstruction losses (which cause blurriness in large-hole inpainting) and relies on adversarial, gradient regularization, and high-level perceptual objectives: L=LG+γR1+λLP\mathcal{L} = \mathcal{L}_G + \gamma R_1 + \lambda \mathcal{L}_P where γ=10\gamma = 10 and λ=0.1\lambda = 0.1.

    Non-Saturating Adversarial Loss

    Applied to both the first-stage coarse generation and the second-stage refined generation x^\hat{x} against real images xx: LG=−Ex^[log⁡(D(x^))]\mathcal{L}_G = -\mathbb{E}_{\hat{x}}\left[\log\left(D(\hat{x})\right)\right] LD=−Ex[log⁡(D(x))]−Ex^[log⁡(1−D(x^))]\mathcal{L}_D = -\mathbb{E}_{x}\left[\log\left(D(x)\right)\right] - \mathbb{E}_{\hat{x}}\left[\log\left(1 - D(\hat{x})\right)\right]

    Discriminator Regularization

    R1R_1 zero-centered gradient penalty regularizes real data gradients: R1=Ex[∥∇D(x)∥2]R_1 = \mathbb{E}_{x}\left[\|\nabla D(x)\|^2\right]

    High-Level Perceptual Loss

    Constrains activations using layer outputs ϕi\phi_i of a pre-trained VGG-19 network restricted to deep features (conv4_4\text{conv4\_4} and conv5_4\text{conv5\_4}) to permit pluralistic valid completions: LP=∑iηi∥ϕi(x^)−ϕi(x)∥1\mathcal{L}_P = \sum_{i} \eta_i \|\phi_i(\hat{x}) - \phi_i(x)\|_1 with scaling coefficients ηconv4_4=14\eta_{\text{conv4\_4}} = \frac{1}{4} and ηconv5_4=12\eta_{\text{conv5\_4}} = \frac{1}{2}.

  7. Knowl 7 — Quantitative Evaluation on Places and CelebA-HQ Benchmarks

    data/table

    Quantitative comparison of MAT with state-of-the-art inpainting methods on Places365-Standard and CelebA-HQ at 512×512512 \times 512 resolution under Small Mask and Large Mask regimes. Performance is evaluated using Fr'echet Inception Distance (FID, lower is better), Paired Inception Discriminative Score (P-IDS in %, higher is better), and Unpaired Inception Discriminative Score (U-IDS in %, higher is better).

    Method #Param
    (×106\times 10^6)
    Places (512×512512 \times 512) CelebA-HQ (512×512512 \times 512)
    Small Mask Large Mask Small Mask Large Mask
    FID↓\downarrow P-IDS↑\uparrow U-IDS↑\uparrow FID↓\downarrow P-IDS↑\uparrow U-IDS↑\uparrow FID↓\downarrow P-IDS↑\uparrow U-IDS↑\uparrow FID↓\downarrow P-IDS↑\uparrow U-IDS↑\uparrow
    MAT (Ours)†^\dagger 62 0.78 31.72 43.71 1.96 23.42 38.34 2.86 21.15 32.56 4.86 13.83 25.33
    MAT (Ours) 62 1.07 27.42 41.93 2.90 19.03 35.36 - - - - - -
    CoModGAN†^\dagger 109 1.10 26.95 41.88 2.92 19.64 35.78 3.26 19.65 31.41 5.65 11.23 22.54
    LaMa†^\dagger 51/27 0.99 22.79 40.58 2.97 13.09 32.29 4.05 9.72 21.57 8.15 2.07 7.58
    ICT 150 - - - - - - 6.28 2.24 9.99 12.84 0.13 0.58
    MADF 85 2.24 14.85 35.03 7.53 6.00 23.78 3.39 12.06 24.61 6.83 3.41 11.26
    AOT GAN 15 3.19 8.07 30.94 10.64 3.07 19.92 4.65 7.92 20.45 10.82 1.94 6.97
    HFill 3 7.94 3.98 23.60 28.92 1.24 11.24 - - - - - -
    DeepFill v2 4 3.02 9.17 32.56 9.27 4.01 21.32 10.11 3.11 9.52 24.42 0.17 0.42
    EdgeConnect 22 4.03 5.88 27.56 12.66 1.93 15.87 10.58 4.14 12.45 39.99 0.10 0.22

    †^\daggerIndicates training on larger subsets (8M images for MAT and CoModGAN on Places, 4.5M for LaMa), while MAT without dagger is trained on Places365-Standard (1.8M images).

  8. Knowl 8 — Ablation Study on MAT Core Components

    data/table

    Ablation study analyzing each architectural component of MAT on Places365-Standard (256×256256 \times 256 images, trained on 100K images for 5M samples and evaluated on 10K validation images).

    Type Model Variant FID↓\downarrow P-IDS (%)↑\uparrow U-IDS (%)↑\uparrow
    A Full Model (MAT) 5.97 13.17 29.23
    B - Tran. (replaces transformers with convolutions) 6.21 11.30 27.39
    C - Adjusted Tran. Block (replaces with standard LN + residual) 6.36 12.30 28.05
    D - MCA (replaces MCA with conventional self-attention) 6.08 13.13 29.19
    E - Style Mani. Module (removes noise style manipulation) 6.10 11.88 27.94
    F - High-Res. Gen. (limits stage-1 output to 64×6464 \times 64) 6.32 12.57 28.21

    The results show that removing long-range transformer modeling (Model B) degrades realism (P-IDS drops from 13.17% to 11.30%), using vanilla transformer blocks (Model C) degrades FID by 0.39 due to optimization instability, omitting MCA (Model D) introduces color and blur artifacts, and low-resolution intermediate representation (Model F) hurts final image generation quality.

  9. Knowl 9 — Ablation on Transformer Body Hyperparameter Configurations

    data/table

    Investigation of the effect of transformer feature dimension (width), block numbers per stage (depth), and attention window sizes in the transformer body on inpainting performance (Places dataset, 256×256256 \times 256 resolution):

    Model Feature Dim. Block Num. Window Size FID↓\downarrow
    Ours (Default) 180 {2, 3, 4, 3, 2} {8, 16, 16, 16, 8} 5.97
    V1 90 {2, 3, 4, 3, 2} {8, 16, 16, 16, 8} 6.28
    V2 180 {1, 1, 2, 1, 1} {8, 16, 16, 16, 8} 6.18
    V3 180 {2, 3, 4, 3, 2} {8, 8, 8, 8, 8} 6.09

    Performance improves consistently with increased feature channel capacity, deeper transformer stages, and larger contextual attention windows.

  10. Knowl 10 — Free-Form Mask Generation Protocol

    experimental setup

    Synthetic free-form masks are generated during training and evaluation by combining randomly placed rectangles and brush strokes across multiple scales:

    1. Training Mask Distribution (Large Masks):

      • Number of up to full-size rectangles: uniformly sampled from [0,3][0, 3].
      • Number of up to half-size rectangles: uniformly sampled from [0,5][0, 5].
      • Number of free-form brush strokes: uniformly sampled from [0,9][0, 9].
      • Brush stroke width: uniformly sampled from [12,48][12, 48] pixels.
      • Number of vertices per stroke: uniformly sampled from [4,18][4, 18].
    2. Testing Mask Distributions:

      • Large Masks: Identical configuration to the training distribution, masking up to 90% of the image area.
      • Small Masks: Number of up to full-size rectangles in [0,2][0, 2], half-size rectangles in [0,3][0, 3], and strokes in [0,4][0, 4], with all other stroke parameters unchanged.
  11. Knowl 11 — Limitations and Failure Modes of MAT

    limitation

    The Mask-Aware Transformer framework has two main limitations:

    1. Semantic Structure Ambiguity on Non-Rigid Objects: Because the network is trained without explicit semantic segmentation or object-level annotations, it struggles to hallucinate complex objects with high geometric variation and articulated poses (such as running animals or specific vehicle configurations) when contextual clues are ruled out by large masks.
    2. Resolution Divisibility Constraint: The multi-scale windowed attention structure requires the input spatial dimensions to be multiples of 512 (or padded/resized accordingly) so that window partitions align properly across the five transformer stages.

Coverage note — None. All major contributed architectural components, attention mechanisms, loss formulations, experimental setups, comparative benchmark tables, ablation studies, and limitations from the paper and supplementary material are fully captured.

References

  1. 1.Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450, 2016.
  2. 2.Coloma Ballester, Marcelo Bertalmio, Vicent Caselles, Guillermo Sapiro, and Joan Verdera. Filling-in by joint interpolation of vector fields and gray levels. TIP, 10(8):1200–1211, 2001.
  3. 3.Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. Patchmatch: A randomized correspondence algorithm for structural image editing. ToG, 28(3):24, 2009.
  4. 4.Dana Berman, Shai Avidan, et al. Non-local image dehazing. In CVPR, pages 1674–1682, 2016.
  5. 5.Marcelo Bertalmio. Strong-continuation, contrast-invariant inpainting with a third-order optimal pde. TIP, 15(7):1934–1938, 2006.
  6. 6.Marcelo Bertalmio, Guillermo Sapiro, Vincent Caselles, and Coloma Ballester. Image inpainting. In Proceedings of the 27th annual conference on Computer graphics and interactive techniques, pages 417–424, 2000.
  7. 7.Antoni Buades, Bartomeu Coll, and J-M Morel. A non-local algorithm for image denoising. In CVPR, volume 2, pages 60–65. IEEE, 2005.
  8. 8.Ting Chen, Mario Lucic, Neil Houlsby, and Sylvain Gelly. On self modulation for generative adversarial networks. In ICLR, 2018.
  9. 9.Donghyeon Cho, Jinsun Park, Tae-Hyun Oh, Yu-Wing Tai, and In So Kweon. Weakly-and self-supervised learning for content-aware deep image retargeting. In ICCV, pages 4558–4567, 2017.
  10. 10.Antonio Criminisi, Patrick Perez, and Kentaro Toyama. Object removal by exemplar-based inpainting. In CVPR, volume 2, pages II–II. IEEE, 2003.
  11. 11.Antonio Criminisi, Patrick Perez, and Kentaro Toyama. Region filling and object removal by exemplar-based image inpainting. TIP, 13(9):1200–1212, 2004.
  12. 12.Ding Ding, Sundaresh Ram, and Jeffrey J Rodrıguez. Image inpainting using nonlocal texture matching and nonlinear filtering. TIP, 28(4):1705–1719, 2018.
  13. 13.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR, 2020.
  14. 14.Mohamed-Jalal Fadili, J-L Starck, and Fionn Murtagh. Inpainting and zooming using sparse representations. The Computer Journal, 52(1):64–79, 2009.
  15. 15.Mrinmoy Ghorai, Soumitra Samanta, Sekhar Mandal, and Bhabatosh Chanda. Multiple pyramids based image inpainting using local patch statistics and steering kernel feature. TIP, 28(11):5495–5509, 2019.
  16. 16.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. NIPS, 27, 2014.
  17. 17.Harald Grossauer. A combined pde and texture synthesis approach to inpainting. In ECCV, pages 214–224. Springer, 2004.
  18. 18.Zongyu Guo, Zhibo Chen, Tao Yu, Jiale Chen, and Sen Liu. Progressive image inpainting with full-resolution residual network. In ACMMM, pages 2496–2504, 2019.
  19. 19.James Hays and Alexei A Efros. Scene completion using millions of photographs. ToG, 26(3):4–es, 2007.
  20. 20.Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016.
  21. 21.Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. NIPS, 30, 2017.
  22. 22.Satoshi Iizuka, Edgar Simo-Serra, and Hiroshi Ishikawa. Globally and locally consistent image completion. ToG, 36(4):1–14, 2017.
  23. 23.Youngjoo Jo and Jongyoul Park. Sc-fegan: Face editing generative adversarial network with user’s sketch and color. In ICCV, pages 1745–1753, 2019.
  24. 24.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In ECCV, pages 694–711. Springer, 2016.
  25. 25.Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. Progressive growing of gans for improved quality, stability, and variation. In ICLR, 2018.
  26. 26.Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. In CVPR, pages 4401–4410, 2019.
  27. 27.Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. In CVPR, pages 8110–8119, 2020.
  28. 28.Olivier Le Meur, Josselin Gautier, and Christine Guillemot. Examplar-based inpainting based on local geometry. In ICIP, pages 3401–3404. IEEE, 2011.
  29. 29.Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, et al. Photo-realistic single image super-resolution using a generative adversarial network. In CVPR, pages 4681–4690, 2017.
  30. 30.Joo Ho Lee, Inchang Choi, and Min H Kim. Laplacian patch-based image synthesis. In CVPR, pages 2727–2735, 2016.
  31. 31.Anat Levin, Assaf Zomet, and Yair Weiss. Learning how to inpaint from global image statistics. In ICCV, volume 1, pages 305–312, 2003.
  32. 32.Jingyuan Li, Ning Wang, Lefei Zhang, Bo Du, and Dacheng Tao. Recurrent feature reasoning for image inpainting. In CVPR, pages 7760–7768, 2020.
  33. 33.Guilin Liu, Fitsum A Reda, Kevin J Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro. Image inpainting for irregular holes using partial convolutions. In ECCV, pages 85–100, 2018.
  34. 34.Hongyu Liu, Bin Jiang, Yibing Song, Wei Huang, and Chao Yang. Rethinking image inpainting via a mutual encoder-decoder with feature equalizations. In ECCV, pages 725–741. Springer, 2020.
  35. 35.Hongyu Liu, Bin Jiang, Yi Xiao, and Chao Yang. Coherent semantic attention for image inpainting. In ICCV, pages 4170–4179, 2019.
  36. 36.Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted windows. ICCV, 2021.
  37. 37.Yuqing Ma, Xianglong Liu, Shihao Bai, Lei Wang, Aishan Liu, Dacheng Tao, and Edwin Hancock. Region-wise generative adversarial imageinpainting for large missing areas. arXiv preprint arXiv:1909.12507, 2019.
  38. 38.Julien Mairal, Francis Bach, Jean Ponce, Guillermo Sapiro, and Andrew Zisserman. Non-local sparse models for image restoration. In ICCV, pages 2272–2279. IEEE, 2009.
  39. 39.Lars Mescheder, Andreas Geiger, and Sebastian Nowozin. Which training methods for gans do actually converge? In ICML, pages 3481–3490. PMLR, 2018.
  40. 40.Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Z Qureshi, and Mehran Ebrahimi. Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212, 2019.
  41. 41.Seoung Wug Oh, Sungho Lee, Joon-Young Lee, and Seon Joo Kim. Onion-peel networks for deep video completion. In ICCV, pages 4403–4412, 2019.
  42. 42.Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros. Context encoders: Feature learning by inpainting. In CVPR, pages 2536–2544, 2016.
  43. 43.Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. Do vision transformers see like convolutional neural networks? arXiv preprint arXiv:2108.08810, 2021.
  44. 44.Yurui Ren, Xiaoming Yu, Ruonan Zhang, Thomas H Li, Shan Liu, and Ge Li. Structureflow: Image inpainting via structure-aware appearance flow. In ICCV, pages 181–190, 2019.
  45. 45.Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015.
  46. 46.Andrew Slavin Ross and Finale Doshi-Velez. Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients. In AAAI, 2018.
  47. 47.Mehdi SM Sajjadi, Bernhard Scholkopf, and Michael Hirsch. Enhancenet: Single image super-resolution through automated texture synthesis. In ICCV, pages 4491–4500, 2017.
  48. 48.Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. In ICLR, 2015.
  49. 49.Yuhang Song, Chao Yang, Yeji Shen, Peng Wang, Qin Huang, and C-C Jay Kuo. Spg-net: Segmentation prediction and guidance network for image inpainting. arXiv preprint arXiv:1805.03356, 2018.
  50. 50.Jian Sun, Lu Yuan, Jiaya Jia, and Heung-Yeung Shum. Image completion with structure propagation. In ACM SIGGRAPH 2005 Papers, pages 861–868. 2005.
  51. 51.Roman Suvorov, Elizaveta Logacheva, Anton Mashikhin, Anastasia Remizova, Arsenii Ashukha, Aleksei Silvestrov, Naejin Kong, Harshith Goka, Kiwoong Park, and Victor Lempitsky. Resolution-robust large mask inpainting with fourier convolutions. arXiv preprint arXiv:2109.07161, 2021.
  52. 52.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. In NIPS, pages 5998–6008, 2017.
  53. 53.Ziyu Wan, Bo Zhang, Dongdong Chen, Pan Zhang, Dong Chen, Jing Liao, and Fang Wen. Bringing old photos back to life. In CVPR, pages 2747–2757, 2020.
  54. 54.Ziyu Wan, Bo Zhang, Dongdong Chen, Pan Zhang, Dong Chen, Jing Liao, and Fang Wen. Old photo restoration via deep latent space translation. arXiv preprint arXiv:2009.07047, 2020.
  55. 55.Ziyu Wan, Jingbo Zhang, Dongdong Chen, and Jing Liao. High-fidelity pluralistic image completion with transformers. arXiv preprint arXiv:2103.14031, 2021.
  56. 56.Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. Non-local neural networks. In CVPR, pages 7794–7803, 2018.
  57. 57.Yi Wang, Xin Tao, Xiaojuan Qi, Xiaoyong Shen, and Jiaya Jia. Image inpainting via generative multi-column convolutional neural networks. NIPS, 2018.
  58. 58.Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. TIP, 13(4):600–612, 2004.
  59. 59.Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang. Cvt: Introducing convolutions to vision transformers. In ICCV, pages 22–31, 2021.
  60. 60.Tete Xiao, Mannat Singh, Eric Mintun, Trevor Darrell, Piotr Dollar, and Ross Girshick. Early convolutions help transformers see better. arXiv preprint arXiv:2106.14881, 2021.
  61. 61.Chaohao Xie, Shaohui Liu, Chao Li, Ming-Ming Cheng, Wangmeng Zuo, Xiao Liu, Shilei Wen, and Errui Ding. Image inpainting with learnable bidirectional attention maps. In ICCV, pages 8858–8867, 2019.
  62. 62.Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. Segformer: Simple and efficient design for semantic segmentation with transformers. NeurIPS, 34, 2021.
  63. 63.Wei Xiong, Jiahui Yu, Zhe Lin, Jimei Yang, Xin Lu, Connelly Barnes, and Jiebo Luo. Foreground-aware image inpainting. In CVPR, pages 5840–5848, 2019.
  64. 64.Zhaoyi Yan, Xiaoming Li, Mu Li, Wangmeng Zuo, and Shiguang Shan. Shift-net: Image inpainting via deep feature rearrangement. In ECCV, pages 1–17, 2018.
  65. 65.Zili Yi, Qiang Tang, Shekoofeh Azizi, Daesik Jang, and Zhan Xu. Contextual residual aggregation for ultra high-resolution image inpainting. In CVPR, pages 7508–7517, 2020.
  66. 66.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Generative image inpainting with contextual attention. In CVPR, pages 5505–5514, 2018.
  67. 67.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Free-form image inpainting with gated convolution. In ICCV, pages 4471–4480, 2019.
  68. 68.Yingchen Yu, Fangneng Zhan, Rongliang Wu, Jianxiong Pan, Kaiwen Cui, Shijian Lu, Feiying Ma, Xuansong Xie, and Chunyan Miao. Diverse image inpainting with bidirectional and autoregressive transformers. arXiv preprint arXiv:2104.12335, 2021.
  69. 69.Yanhong Zeng, Jianlong Fu, Hongyang Chao, and Baining Guo. Learning pyramid-context encoder network for high-quality image inpainting. In CVPR, pages 1486–1494, 2019.
  70. 70.Yanhong Zeng, Jianlong Fu, Hongyang Chao, and Baining Guo. Aggregated contextual transformations for high-resolution image inpainting. arXiv preprint arXiv:2104.01431, 2021.
  71. 71.Yu Zeng, Zhe Lin, Jimei Yang, Jianming Zhang, Eli Shechtman, and Huchuan Lu. High-resolution image inpainting with iterative confidence feedback and guided upsampling. In ECCV, pages 1–17. Springer, 2020.
  72. 72.Haoran Zhang, Zhenzhen Hu, Changzhi Luo, Wangmeng Zuo, and Meng Wang. Semantic image inpainting with progressive generative networks. In ACMMM, pages 1939–1947, 2018.
  73. 73.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, pages 586–595, 2018.
  74. 74.Lei Zhao, Qihang Mo, Sihuan Lin, Zhizhong Wang, Zhiwen Zuo, Haibo Chen, Wei Xing, and Dongming Lu. Uctgan: Diverse image inpainting based on unsupervised cross-space translation. In CVPR, pages 5741–5750, 2020.
  75. 75.Shengyu Zhao, Jonathan Cui, Yilun Sheng, Yue Dong, Xiao Liang, I Eric, Chao Chang, and Yan Xu. Large scale image completion via co-modulated generative adversarial networks. In ICLR, 2020.
  76. 76.Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai. Pluralistic image completion. In CVPR, pages 1438–1447, 2019.
  77. 77.Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai. Tfill: Image completion via a transformer-based architecture. arXiv preprint arXiv:2104.00845, 2021.
  78. 78.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. PAMI, 40(6):1452–1464, 2017.
  79. 79.Manyu Zhu, Dongliang He, Xin Li, Chao Li, Fu Li, Xiao Liu, Errui Ding, and Zhaoxiang Zhang. Image inpainting by end-to-end cascaded refinement with mask awareness. TIP, 30:4855–4866, 2021.

Citation

MLA
Li, W., et al. “MAT: Mask-Aware Transformer for Large Hole Image Inpainting”. arXiv, 2022, http://arxiv.org/abs/2203.15270v3.
APA
Li, W., Lin, Z., Zhou, K., Qi, L., Wang, Y., & Jia, J. (2022). MAT: Mask-Aware Transformer for Large Hole Image Inpainting. arXiv. http://arxiv.org/abs/2203.15270v3
Chicago
Li, W., Z. Lin, K. Zhou, L. Qi, Y. Wang, and J. Jia. 2022. “MAT: Mask-Aware Transformer for Large Hole Image Inpainting”. arXiv. http://arxiv.org/abs/2203.15270v3.
Harvard
Li, W. et al. (2022) “MAT: Mask-Aware Transformer for Large Hole Image Inpainting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.15270v3.
Vancouver
1. Li W, Lin Z, Zhou K, Qi L, Wang Y, Jia J (2022) MAT: Mask-Aware Transformer for Large Hole Image Inpainting. arXiv

BibTeX

@article{li2022mat,
  title = {MAT: Mask-Aware Transformer for Large Hole Image Inpainting},
  author = {Li, Wenbo and Lin, Zhe and Zhou, Kun and Qi, Lu and Wang, Yi and Jia, Jiaya},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.15270v3},
  eprint = {2203.15270}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE