MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting
Xiaoguang LiQing GuoDi LinPing LiWei FengSong Wang
Proposes Multi-level Interactive Siamese Filtering, a dual-branch framework that couples dynamic kernel prediction with feature- and image-level filtering to reduce visual artifacts and improve generalization across diverse inpainting benchmarks.
Digital image inpainting—the process of restoring missing or damaged parts of an image—is a critical computer vision capability for media restoration, editing, and enhancement. Existing deep generative methods often fail to generalize across diverse scenes and mask geometries, producing noticeable visual distortions, blurred textures, or structures that deviate significantly from ground truth. Conversely, traditional predictive filtering maintains local smoothness but fails to reconstruct large missing areas. The article demonstrates that formulating inpainting as an interactive, multi-level predictive filtering task bridges this gap, achieving both structural coherence and fine-grained visual fidelity.
The researchers developed Multi-level Interactive Siamese Filtering (MISF), an architecture composed of two interlinked branches: a Kernel Prediction Branch and a Semantic and Image Filtering Branch. The network was evaluated against several state-of-the-art baselines across standard public benchmarks, including natural scenes (Places2), facial portraits (CelebA), and cultural heritage artifacts (Dunhuang Challenge), under varying degrees of image corruption.
The evaluation revealed three principal findings. First, MISF consistently outperformed existing methods across all evaluated image quality and perceptual fidelity metrics. On the Places2 benchmark, MISF achieved relative peak signal-to-noise ratio improvements of 7.01% to 7.90% over strong baselines across corruption levels ranging from 0% to 60%. Second, MISF delivered a 47.12% relative reduction in perceptual distortion compared to hybrid filtering-generative baselines for small corruptions, demonstrating superior texture realism. Third, ablation studies showed that semantic filtering on deep features successfully restores high-level layout, while image-level filtering preserves sharp edge details, confirming that dual-level dynamic convolution is essential for generalizability.
These results establish that incorporating explicit neighborhood smoothness priors through dynamic, adaptive filtering overcomes the visual inconsistency inherent in pure generative reconstruction. For organizations relying on image restoration pipelines, adopting multi-level filtering reduces structural artifacts and minimizes manual post-processing across diverse image types.
Decision-makers and engineering teams should consider piloting this dynamic filtering approach within production inpainting workflows, particularly where high fidelity to original visual structures is mandatory. Future development should evaluate and extend this architecture beyond standard benchmarks into specialized domain challenges, such as cloud removal in satellite and remote sensing imagery. While the experimental findings offer high confidence across benchmark domains, caution is warranted when deploying the model to novel operational environments that fall outside the distribution of standard public datasets.
- Paper: Free-Form Image Inpainting With Gated Convolution, Jiahui Yu et al. (2018). Introduces gated convolutions for free-form inpainting, providing foundational dynamic convolutional operations that MISF seeks to advance through interactive siamese predictive filtering.
- Paper: Image Inpainting for Irregular Holes Using Partial Convolutions, Guilin Liu et al. (2018). Establishes partial convolutions with automated mask updating for irregular hole inpainting, which serves as a primary baseline and motivation for predictive filtering approaches.
- Paper: Generative Image Inpainting with Contextual Attention, Jiahui Yu et al. (2018). Pioneers contextual attention to treat feature patches as dynamic convolutional filters for inpainting, directly preceding MISF's dual-branch dynamic kernel prediction.
- Paper: Globally and locally consistent image completion, SATOSHI IIZUKA et al. (2017). Presents the foundational global and local context discriminator framework for multi-scale generative image inpainting.
- Paper: Context Encoders: Feature Learning by Inpainting, Deepak Pathak et al. (2016). Introduces deep encoder-decoder architectures trained with adversarial and reconstruction losses for image inpainting that MISF analyzes and reformulates.
- Paper: Multi-Stage Progressive Image Restoration, Syed Waqas Zamir et al. (2021). Develops multi-stage feature aggregation and cross-stage feature fusion mechanisms for balancing deep semantics and high-resolution spatial details in image restoration.
- Paper: Towards Coherent Image Inpainting Using Denoising Diffusion Implicit Models, Guanhua Zhang et al. (2023). Extends generative image inpainting by addressing spatial and contextual incoherence through progressive iterative diffusion without task-specific architectural filtering.
- Paper: Zero-Shot Image Restoration Using Denoising Diffusion Null-Space Model, Yinhuai Wang et al. (2023). Generalizes image restoration and inpainting to a zero-shot range-null space decomposition framework using pre-trained diffusion models.
- Paper: Imagen Editor and EditBench: Advancing and Evaluating Text-Guided Image Inpainting, Su Wang et al. (2023). Advances deep inpainting pipelines by conditioning the generative completion on precise text instructions and object-level masking benchmarks.
- Paper: Hierarchical Fine-Grained Image Forgery Detection and Localization, Xiao Guo et al. (2023). Applies forensic multi-branch hierarchical analysis to detect and localize fine-grained manipulations produced by modern inpainting and generative methods.
