MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting

Xiaoguang LiQing GuoDi LinPing LiWei FengSong Wang

article2022CVPR129 citations

Proposes Multi-level Interactive Siamese Filtering, a dual-branch framework that couples dynamic kernel prediction with feature- and image-level filtering to reduce visual artifacts and improve generalization across diverse inpainting benchmarks.

Listen

Digital image inpainting—the process of restoring missing or damaged parts of an image—is a critical computer vision capability for media restoration, editing, and enhancement. Existing deep generative methods often fail to generalize across diverse scenes and mask geometries, producing noticeable visual distortions, blurred textures, or structures that deviate significantly from ground truth. Conversely, traditional predictive filtering maintains local smoothness but fails to reconstruct large missing areas. The article demonstrates that formulating inpainting as an interactive, multi-level predictive filtering task bridges this gap, achieving both structural coherence and fine-grained visual fidelity.

The researchers developed Multi-level Interactive Siamese Filtering (MISF), an architecture composed of two interlinked branches: a Kernel Prediction Branch and a Semantic and Image Filtering Branch. The network was evaluated against several state-of-the-art baselines across standard public benchmarks, including natural scenes (Places2), facial portraits (CelebA), and cultural heritage artifacts (Dunhuang Challenge), under varying degrees of image corruption.

The evaluation revealed three principal findings. First, MISF consistently outperformed existing methods across all evaluated image quality and perceptual fidelity metrics. On the Places2 benchmark, MISF achieved relative peak signal-to-noise ratio improvements of 7.01% to 7.90% over strong baselines across corruption levels ranging from 0% to 60%. Second, MISF delivered a 47.12% relative reduction in perceptual distortion compared to hybrid filtering-generative baselines for small corruptions, demonstrating superior texture realism. Third, ablation studies showed that semantic filtering on deep features successfully restores high-level layout, while image-level filtering preserves sharp edge details, confirming that dual-level dynamic convolution is essential for generalizability.

These results establish that incorporating explicit neighborhood smoothness priors through dynamic, adaptive filtering overcomes the visual inconsistency inherent in pure generative reconstruction. For organizations relying on image restoration pipelines, adopting multi-level filtering reduces structural artifacts and minimizes manual post-processing across diverse image types.

Decision-makers and engineering teams should consider piloting this dynamic filtering approach within production inpainting workflows, particularly where high fidelity to original visual structures is mandatory. Future development should evaluate and extend this architecture beyond standard benchmarks into specialized domain challenges, such as cloud removal in satellite and remote sensing imagery. While the experimental findings offer high confidence across benchmark domains, caution is warranted when deploying the model to novel operational environments that fall outside the distribution of standard public datasets.

Cover for MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting

Abstract

Although achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applications. Image-level predictive filtering is a widely used restoration technique by predicting suitable kernels adaptively according to different input scenes. Inspired by this inherent advantage, we explore the possibility of addressing image inpainting as a filtering task. To this end, we first study the advantages and challenges of the image-level predictive filtering for inpainting: the method can preserve local structures and avoid artifacts but fails to fill large missing areas. Then, we propose the semantic filtering by conducting filtering on deep feature level, which fills the missing semantic information but fails to recover the details. To address the issues while adopting the respective advantages, we propose a novel filtering technique, i.e., Multi-level Interactive Siamese Filtering (MISF) containing two branches: kernel prediction branch (KPB) and semantic & image filtering branch (SIFB). These two branches are interactively linked: SIFB provides multi-level features for KPB while KPB predicts dynamic kernels for SIFB. As a result, the final method takes the advantage of effective semantic & image-level filling for high-fidelity inpainting. Moreover, we discuss the relationship between MISF and the naive encoder-decoder-based inpainting, inferring that MISF provides novel dynamic convolutional operations to enhance the high generalization capability across scenes. We validate our method on three challenging datasets, i.e., Dunhuang, Places2, and CelebA. Our method outperforms state-of-the-art baselines on four metrics, i.e., L1, PSNR, SSIM, and LPIPS.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3.2. Challenges and Motivations
  • 3. Discussion and Motivation
  • 3.1. Predictive Filtering for Image Inpainting
  • 4. Methodology
  • 4.1. Semantic Filtering for Image Inpainting
  • 4.2. Multi-level Interactive Siamese Filtering
  • 4.3. Relationship to Encoder-Decoder Network
  • 4.4. Implementation Details
  • 5. Experiments
  • 5.1. Setups
  • 5.2. Comparison Results
  • 5.3. Ablation Study
  • 5.4. Discussion
  • 6. Conclusions
  • References

Knowls

  1. Knowl 1 — Multi-level Interactive Siamese Filtering architecture

    model/method

    Multi-level Interactive Siamese Filtering (MISF) is an image-inpainting architecture with two coupled encoder-decoder branches: a kernel prediction branch (KPB) and a semantic-and-image filtering branch (SIFB). Given a corrupted image II and binary missing-region mask MM, SIFB extracts multi-level features FjF_j, while KPB extracts corresponding features EjE_j and predicts spatially varying kernels from the raw input, the KPB features, and the SIFB features. The two branches therefore influence each other rather than predicting independent kernels.

    For a feature or image map XX, a kernel field KK performs local filtering by

    FK(X)[p]=∑q∈NpKp[q−p]X[q],\mathcal{F}_{K}(X)[p]=\sum_{q\in\mathcal{N}_{p}}K_{p}[q-p]X[q],

    where pp and qq are pixel or feature coordinates, Np\mathcal{N}_{p} is the fixed neighborhood of N2N^{2} locations around pp, and KpK_{p} contains the weights used at location pp. If FlF_l is the SIFB feature at the selected semantic level, KlK_l is its predicted feature-level kernel, K0K_0 is the predicted image-level kernel, and DD denotes the remaining decoder operations, MISF produces

    I^=FK0(D(FKl(Fl))),Fl=φl(⋯φ2(φ1(I,M))).\hat{I}=\mathcal{F}_{K_0}\left(D\left(\mathcal{F}_{K_l}(F_l)\right)\right),\qquad F_l=\varphi_l(\cdots\varphi_2(\varphi_1(I,M))).

    The semantic filtering at FlF_l supplies information needed to fill large holes, whereas the final image-level filtering restores local pixels and fine details. The KPB kernels are dynamically conditioned on the input scene and multi-level SIFB features, giving MISF an explicit neighboring-pixel smoothness prior together with scene-level semantic information.

  2. Knowl 2 — Predictive image-level filtering and its failure mode

    model/method

    The paper first formulates image inpainting as pixel-wise predictive filtering. For a corrupted grayscale image I∈RH×WI\in\mathbb{R}^{H\times W}, completed image I^∈RH×W\hat I\in\mathbb{R}^{H\times W}, and kernel tensor K∈RH×W×N2K\in\mathbb{R}^{H\times W\times N^{2}}, the completion is I^=I⊛K\hat I=I\mathbin{\circledast}K, where each output pixel is a weighted combination of neighboring input pixels:

    I^[p]=∑q∈NpKp[q−p]I[q].\hat I[p]=\sum_{q\in\mathcal{N}_{p}}K_p[q-p]I[q].

    Here pp and qq are image coordinates, Np\mathcal{N}_p contains N2N^{2} neighboring coordinates, and Kp∈RN×NK_p\in\mathbb{R}^{N\times N} is the reshaped kernel for output location pp. A predictive encoder-decoder ϕ\phi estimates all kernels from the corrupted image and its binary mask according to K=ϕ(I,M)K=\phi(I,M).

    This image-level filter preserves local structures and avoids many generative artifacts when missing regions are thin or small, because pixels near the hole boundary provide useful neighboring evidence. When holes become large or thick, the missing region breaks local continuity, so pixels far from the boundary cannot be recovered reliably and the method cannot infer scene-level content. Repeatedly applying the filter does not solve this problem: errors accumulate inward from the boundary and the center of a large hole becomes vague.

  3. Knowl 3 — Semantic filtering on deep features

    model/method

    To recover content inside large missing regions, the paper extends predictive filtering from image pixels to deep feature maps. An encoder φ\varphi receives a corrupted image II and mask MM and produces the feature at layer ll as Fl=φl(⋯φ2(φ1(I,M)))F_l=\varphi_l(\cdots\varphi_2(\varphi_1(I,M))). A kernel predictor ψl\psi_l produces a feature-level kernel field Kl=ψl(I,M)K_l=\psi_l(I,M). Each feature location is then filtered by

    F^l[p]=∑q∈NpKl,p[q−p]Fl[q],\widehat{F}_l[p]=\sum_{q\in\mathcal{N}_p}K_{l,p}[q-p]F_l[q],

    where pp and qq index locations in the layer-ll feature map, Np\mathcal{N}_p is its local neighborhood, and Kl,pK_{l,p} is the kernel predicted for location pp. The filtered feature map F^l\widehat{F}_l replaces FlF_l in the subsequent encoder-decoder computation that generates the completed image.

    Deep features retain semantic information even when a large image region is absent, so semantic filtering can reconstruct major structures that image-level filtering cannot. Its output, however, is spatially coarse and tends to lose details: the main structure is recovered but local regions can remain blurred or contain artifacts. In the standalone semantic-filtering implementation, the kernel predictor is a three-layer convolutional network and filtering is applied at the third encoder layer.

  4. Knowl 4 — MISF as input-conditioned dynamic convolution

    theoretical result

    From the encoder-decoder viewpoint, semantic filtering adds an element-wise dynamic convolutional operation to an otherwise ordinary encoder-decoder. MISF makes this operation conditional on multiple feature levels: the KPB receives its own feature EjE_j and the SIFB feature FjF_j at layer jj, concatenates them, and predicts the feature-level kernel KlK_l and image-level kernel K0K_0. Thus, the convolutional weights used to reconstruct each location vary with both the raw corrupted input and the scene-level representation.

    This interpretation explains the complementary behavior of MISF. Feature-level dynamic filtering adapts the reconstruction to semantic content in large holes, while image-level dynamic filtering adapts local pixel reconstruction to fine structures. The paper's controlled comparisons support this interpretation: a naive encoder-decoder produces artifacts and structure mismatches, adding one dynamic filtering operation improves the result, and MISF with both semantic and image filtering gives the strongest semantic and detail recovery.

  5. Knowl 5 — Training objective and implementation configuration

    experimental setup

    MISF and its comparison variants are trained using a weighted combination of reconstruction, adversarial, perceptual, and style losses. For predicted completion I^\hat I and ground-truth image I∗I^{*}, the objective is

    L(I^,I∗)=λ1L1+λ2Lgan+λ3Lperc+λ4Lstyle,\mathcal{L}(\hat I,I^{*})=\lambda_1\mathcal{L}_1+\lambda_2\mathcal{L}_{\mathrm{gan}}+\lambda_3\mathcal{L}_{\mathrm{perc}}+\lambda_4\mathcal{L}_{\mathrm{style}},

    where L1\mathcal{L}_1 measures pixel reconstruction error, Lgan\mathcal{L}_{\mathrm{gan}} is the adversarial loss, Lperc\mathcal{L}_{\mathrm{perc}} is the perceptual loss, and Lstyle\mathcal{L}_{\mathrm{style}} is the style loss. The weights used are λ1=1\lambda_1=1, λ2=0.1\lambda_2=0.1, λ3=0.1\lambda_3=0.1, and λ4=250\lambda_4=250.

    The implementation uses a 15-layer encoder-decoder and applies semantic filtering only at the third encoder feature F3F_3 to limit memory and computation. Adam optimization uses learning rate 0.00010.0001, batch size 1616, and approximately 350,000350{,}000 iterations. Experiments use two NVIDIA Tesla V100 GPUs.

  6. Knowl 6 — Datasets, masks, and evaluation protocol

    experimental setup

    The method is evaluated on Places2, CelebA, and the Dunhuang Challenge dataset. Places2 contains more than eight million images spanning more than 365 scenes, CelebA contains more than 180,000 face images, and Dunhuang provides practical grotto-painting restoration data. Standard test sets are used for CelebA and Dunhuang; for Places2, 30,000 random images are used for testing.

    Places2 and CelebA images are corrupted with irregular masks grouped by missing-area ratio: 0%0\%--20%20\%, 20%20\%--40%40\%, and 40%40\%--60%60\% of the image area. The official mask protocol is used for Dunhuang. Quality is measured with PSNR and SSIM, for which larger values are better, and L1L_1 error and LPIPS, for which smaller values are better. PSNR, SSIM, and L1L_1 evaluate image fidelity, while LPIPS evaluates perceptual similarity to the ground truth.

  7. Knowl 7 — Quantitative superiority across datasets and mask sizes

    data/table

    MISF outperforms the compared inpainting methods across the available datasets and mask ratios. The following entries reproduce representative values for MISF, RFRNet, and JPGNet; the paper also compares StructFlow, EdgeConnect, PConv, and CTSDG. Places2 and CelebA use the three mask-ratio groups, while Dunhuang uses its default protocol.

    Metric Method P2 –20 P2 0–40 P2 0–60 CelebA –20 CelebA 0–40 CelebA 0–60 Dunhuang
    PSNR RFRNet 29.281 22.589 18.581 33.573 25.635 20.539 36.485
    PSNR JPGNet 30.673 23.937 19.884 34.401 26.543 21.297 37.646
    PSNR MISF 31.335 24.239 20.044 34.494 26.635 21.553 38.383
    L1L_1 RFRNet 1.009 3.218 6.719 0.521 1.811 4.346 0.401
    L1L_1 JPGNet 0.830 2.581 5.294 0.477 1.651 4.042 0.353
    L1L_1 MISF 0.726 2.340 4.965 0.474 1.616 3.826 0.341
    SSIM RFRNet 0.9283 0.7868 0.6137 0.9626 0.8746 0.7400 0.9648
    SSIM JPGNet 0.9452 0.8348 0.6915 0.9674 0.8908 0.7697 0.9724
    SSIM MISF 0.9506 0.8435 0.6931 0.9680 0.8911 0.7698 0.9735
    LPIPS RFRNet 0.0825 0.2161 0.3571 0.0400 0.1215 0.2335 0.0463
    LPIPS JPGNet 0.0817 0.2145 0.3535 0.0440 0.1316 0.2502 0.0469
    LPIPS MISF 0.0432 0.1298 0.2499 0.0315 0.0949 0.1911 0.0330

    The complete comparison reports that MISF has the best PSNR, SSIM, and L1L_1 values across all listed datasets and mask ratios, with consistent LPIPS improvements as well. Relative to RFRNet on Places2, MISF raises PSNR by 7.01%, 7.3%, and 7.9% for the three increasing mask-ratio groups. Relative to JPGNet, MISF reduces LPIPS by 47.12% for Places2 with a 0%0\%--20%20\% mask. The gains remain present for large missing regions, supporting the claimed cross-scene and cross-mask generalization.

  8. Knowl 8 — Complementarity of image and semantic filtering

    empirical result

    The ablation study compares image-only filtering, semantic-only filtering, and MISF. Image-only filtering performs poorly on Places2 and CelebA, where missing regions are often large, but performs relatively well on Dunhuang, whose default holes are more compatible with local reconstruction; its Dunhuang PSNR is 37.02137.021 with L1=0.386L_1=0.386. Semantic-only filtering improves large-hole completion, reaching CelebA PSNR values of 34.25334.253, 26.51826.518, and 21.48621.486 for the 0%0\%--20%20\%, 20%20\%--40%40\%, and 40%40\%--60%60\% mask groups. MISF improves these values to 34.49434.494, 26.63526.635, and 21.55321.553, respectively, and reaches Dunhuang PSNR 38.38338.383 with L1=0.341L_1=0.341.

    The combined design is also superior on the other metrics. On Places2, MISF obtains SSIM 0.95060.9506, 0.84350.8435, and 0.69310.6931 and LPIPS 0.04320.0432, 0.12980.1298, and 0.24990.2499 across the three mask groups, outperforming both single-level variants. Qualitative comparisons show that image-only filtering preserves small local structures but fails inside large holes, semantic-only filtering recovers plausible global structures but smooths details, and MISF recovers both natural structures and fine-grained local content.

  9. Knowl 9 — Feature-depth and dynamic-operation analyses

    empirical result

    A feature-similarity analysis compares corrupted-image features with ground-truth features before and after MISF filtering. Across 1,000 randomly sampled examples for each mask-ratio grouping and for all mask ratios combined, filtering moves the corrupted features closer to the ground-truth features under cross-correlation. The improvement margin becomes larger as the missing-area ratio increases, indicating that semantic filtering is particularly useful for severe corruption.

    On CelebA, applying semantic filtering at progressively deeper encoder features generally improves completion. The PSNR values for filtering at F1F_1, F2F_2, and F3F_3 are 33.98133.981, 34.12834.128, and 34.25334.253 for the smallest mask group; 26.36426.364, 26.42926.429, and 26.51826.518 for the middle group; and 21.36021.360, 21.35321.353, and 21.48621.486 for the largest group. The corresponding LPIPS values for the three groups are (0.0359,0.0355,0.0343)(0.0359,0.0355,0.0343), (0.1038,0.1021,0.1008)(0.1038,0.1021,0.1008), and (0.2043,0.2028,0.2009)(0.2043,0.2028,0.2009), respectively. The paper cautions that it is unknown whether this improvement continues indefinitely with deeper features.

    From the encoder-decoder ablation, the variants are ordered by the number of dynamic filtering operations: a naive encoder-decoder has none, En-decoder-Filter and semantic filtering have one, and MISF has two. Networks with more dynamic convolutional operations achieve better inpainting under all four metrics across datasets and missing-area sizes.

  10. Knowl 10 — Stated limitation and scope of validation

    limitation

    The paper trains and evaluates MISF only on widely used public datasets—Places2, CelebA, and Dunhuang—which cover natural scenes, faces, and grotto paintings but may represent only part of real-world image-inpainting conditions. Consequently, the breadth of the claimed generalization beyond these scene types is not established. The authors identify extension to additional domains, such as cloud removal in remote-sensing images, as future work.

Coverage note — No substantial contributed material was omitted; detailed per-image visual examples and full baseline rows were compressed because their conclusions are represented by the quantitative, ablation, and qualitative-result knowls.

References

  1. 1.Steve Bako, Thijs Vogels, Brian McWilliams, Mark Meyer, Jan Novák, Alex Harvill, Pradeep Sen, Tony DeRose, and Fabrice Rousselle. Kernel-predicting convolutional networks for denoising monte carlo renderings. ACM Transactions on Graphics, 36(4):97:1–97:14, 2017.
  2. 2.Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. Patchmatch: A randomized correspondence algorithm for structural image editing. ACM Transactions on Graphics, 28(3):24, 2009.
  3. 3.Marcelo Bertalmio, Luminita Vese, Guillermo Sapiro, and Stanley Osher. Simultaneous structure and texture image inpainting. IEEE Transactions on Image Processing, 12(8):882–889, 2003.
  4. 4.T. Brooks and J. T. Barron. Learning to synthesize motion blur. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 6833–6841, 2019.
  5. 5.Tony F. Chan and Jianhong Shen. Nontexture inpainting by curvature-driven diffusions. Journal of Visual Communication and Image Representation, 12(4):436–449, 2001.
  6. 6.Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. Dynamic convolution: Attention over convolution kernels. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11030–11039, 2020.
  7. 7.Ding Ding, Sundaresh Ram, and Jeffrey J Rodríguez. Image inpainting using nonlocal texture matching and nonlinear filtering. IEEE Transactions on Image Processing, 28(4):1705–1719, 2018.
  8. 8.Yixiang Fang, Kaiqiang Yu, Reynold Cheng, Laks VS Lakshmanan, and Xuemin Lin. Efficient algorithms for densest subgraph discovery. arXiv preprint arXiv:1906.00341, 2019.
  9. 9.Lan Fu, Changqing Zhou, Qing Guo, Felix Juefei-Xu, Hongkai Yu, Wei Feng, Yang Liu, and Song Wang. Auto-exposure fusion for single-image shadow removal. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021.
  10. 10.Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets. Advances in Neural Information Processing Systems, 27, 2014.
  11. 11.Qing Guo, Felix Juefei-Xu, Changqing Zhou, Yang Liu, and Song Wang. Sparta: Spatially attentive and adversarially robust activation. arXiv preprint arXiv:2105.08269, 2021.
  12. 12.Qing Guo, Xiaoguang Li, Felix Juefei-Xu, Hongkai Yu, Yang Liu, et al. Jpgnet: Joint predictive filtering and generative network for image inpainting. arXiv e-prints, pages arXiv–2107, 2021.
  13. 13.Qing Guo, Jingyang Sun, Felix Juefei-Xu, Lei Ma, Xiaofei Xie, Wei Feng, Jianjun Zhao, and Yang Liu. Efficientderain: Learning pixel-wise dilation filtering for high-efficiency single-image deraining. In AAAI Conference on Artificial Intelligence, 2021.
  14. 14.Xiefan Guo, Hongyu Yang, and Di Huang. Image inpainting via conditional texture and structure dual generation. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 14134–14143, 2021.
  15. 15.Satoshi Iizuka, Edgar Simo-Serra, and Hiroshi Ishikawa. Globally and locally consistent image completion. ACM Transactions on Graphics, 36(4):1–14, 2017.
  16. 16.Justin Johnson, Alexandre Alahi, and Li Fei-Fei. Perceptual losses for real-time style transfer and super-resolution. In European Conference on Computer Vision (ECCV), pages 694–711. Springer, 2016.
  17. 17.Jingyuan Li, Ning Wang, Lefei Zhang, Bo Du, and Dacheng Tao. Recurrent Feature Reasoning for Image Inpainting. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7757–7765, 2020.
  18. 18.Kangshun Li, Yunshan Wei, Zhen Yang, and Wenhua Wei. Image inpainting algorithm based on tv model and evolutionary algorithm. Soft Computing, 20(3):885–893, 2016.
  19. 19.Guilin Liu, Fitsum A Reda, Kevin J Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro. Image inpainting for irregular holes using partial convolutions. In European Conference on Computer Vision (ECCV), pages 85–100, 2018.
  20. 20.Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. Deep learning face attributes in the wild. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 3730–3738, 2015.
  21. 21.Ben Mildenhall, Jonathan T Barron, Jiawen Chen, Dillon Sharlet, Ren Ng, and Robert Carroll. Burst denoising with kernel prediction networks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2502–2510, 2018.
  22. 22.Mehdi Mirza and Simon Osindero. Conditional generative adversarial nets. arXiv preprint arXiv:1411.1784, 2014.
  23. 23.Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Qureshi, and Mehran Ebrahimi. Edgeconnect: Structure guided image inpainting using edge prediction. In IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Oct 2019.
  24. 24.Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A. Efros. Context Encoders: Feature Learning by Inpainting. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2536–2544, 2016.
  25. 25.Yurui Ren, Xiaoming Yu, Ruonan Zhang, Thomas H. Li, Shan Liu, and Ge Li. Structureflow: Image inpainting via structure-aware appearance flow. In IEEE/CVF International Conference on Computer Vision (ICCV), 2019.
  26. 26.Mehdi SM Sajjadi, Bernhard Scholkopf, and Michael Hirsch. Enhancenet: Single image super-resolution through automated texture synthesis. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 4491–4500, 2017.
  27. 27.Linsen Song, Jie Cao, Lingxiao Song, Yibo Hu, and Ran He. Geometry-aware face completion and editing. In AAAI Conference on Artificial Intelligence, volume 33, pages 2506–2513, 2019.
  28. 28.Zhaoyi Yan, Xiaoming Li, Mu Li, Wangmeng Zuo, and Shiguang Shan. Shift-net: Image inpainting via deep feature rearrangement. In European Conference on Computer Vision (ECCV), pages 1–17, 2018.
  29. 29.Jingyu Yang, Jianhua Guo, Huanjing Yue, Zhiheng Liu, Haofeng Hu, and Kun Li. Cdnet: Cnn-based cloud detection for remote sensing imagery. IEEE Transactions on Geoscience and Remote Sensing, 57(8):6195–6211, 2019.
  30. 30.Raymond A Yeh, Chen Chen, Teck Yian Lim, Alexander G Schwing, Mark Hasegawa-Johnson, and Minh N Do. Semantic image inpainting with deep generative models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5485–5493, 2017.
  31. 31.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Generative image inpainting with contextual attention. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5505–5514, 2018.
  32. 32.Jiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen, Xin Lu, and Thomas S Huang. Free-form image inpainting with gated convolution. In IEEE/CVF International Conference on Computer Vision (ICCV), pages 4471–4480, 2019.
  33. 33.Tianxiu Yu, Shijie Zhang, Cong Lin, Shaodi You, Jian Wu, Jiawan Zhang, Xiaohong Ding, and Huili An. Dunhuang grottoes painting dataset and benchmark. arXiv preprint arXiv:1907.04589, 2019.
  34. 34.Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 586–595, 2018.
  35. 35.Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(6):1452–1464, 2017.

Citation

MLA
Li, X., et al. “MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting”. arXiv, 2022, http://arxiv.org/abs/2203.06304v1.
APA
Li, X., Guo, Q., Lin, D., Li, P., Feng, W., & Wang, S. (2022). MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting. arXiv. http://arxiv.org/abs/2203.06304v1
Chicago
Li, X., Q. Guo, D. Lin, P. Li, W. Feng, and S. Wang. 2022. “MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting”. arXiv. http://arxiv.org/abs/2203.06304v1.
Harvard
Li, X. et al. (2022) “MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2203.06304v1.
Vancouver
1. Li X, Guo Q, Lin D, Li P, Feng W, Wang S (2022) MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting. arXiv

BibTeX

@article{li2022misf,
  title = {MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting},
  author = {Li, Xiaoguang and Guo, Qing and Lin, Di and Li, Ping and Feng, Wei and Wang, Song},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2203.06304v1},
  eprint = {2203.06304}
}
Metadata:arXiv

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: IEEE