Architecture-Agnostic Masked Image Modeling - From ViT back to CNN
Siyuan LiDi WuFang WuZelin ZangStan Z. Li
Presents a unified masked image modeling framework that operates across both vision transformers and convolutional neural networks by masking intermediate feature representations rather than input pixels, enabling CNNs to achieve superior self-supervised pre-training performance without specialized transformer components.
Training high-performance computer vision systems usually requires massive amounts of manually annotated data, which is costly and labor-intensive to produce. To overcome this limitation, self-supervised learning methods have gained popularity, especially masked image modeling. In this approach, portions of an image are hidden, and the artificial intelligence model learns by reconstructing the missing visual information. While masked image modeling has demonstrated remarkable success when paired with modern Vision Transformers, prior research widely believed it was fundamentally incompatible with traditional Convolutional Neural Networks.
The article aims to uncover the underlying operational principles of masked image modeling and demonstrate a unified framework that functions effectively across both Vision Transformers and Convolutional Neural Networks without requiring architecture-specific modifications.
To achieve this, the authors evaluated model robustness under varying levels of visual occlusion and analyzed multi-order interactions, which measure how effectively models combine information from multiple image patches. Using these empirical insights, they developed the Architecture-Agnostic Masked Image Modeling framework. Instead of masking tokens directly at the initial input layer, the proposed method fills occluded image regions with the average color values and introduces learned mask tokens into the intermediate layers of the network. Additionally, the approach incorporates an adaptive frequency loss using the Fourier transform to train the model on medium-frequency structural patterns such as contours and edges. The evaluation encompassed standard benchmarks, including ImageNet classification, COCO object detection and segmentation, and ADE20K semantic parsing, across multiple model scales.
The investigation produced several key findings. First, the core benefit of masked image modeling is teaching models to capture middle-order patch interactions—such as shapes and boundaries—rather than merely achieving visual reconstruction. Second, the proposed framework successfully bridged the architecture gap, allowing standard Convolutional Neural Networks like ResNet-50 to achieve an 80.4% top-1 accuracy on ImageNet, outperforming established contrastive learning alternatives without complex tokenizers. Third, Vision Transformers pre-trained with the framework achieved strong classification performance, reaching up to 86.3% top-1 accuracy on ViT-Large. Finally, in downstream transfer evaluations on COCO and ADE20K, the method matched or exceeded prevailing contrastive and masked learning baselines across both convolutional and transformer architectures.
These findings indicate that organizations can leverage modern self-supervised masked pre-training techniques without abandoning established convolutional network deployments. This flexibility enables teams to reuse existing hardware-optimized network pipelines while achieving competitive accuracy, thereby lowering transition costs and reducing dependence on expensive labeled datasets.
Technical leaders looking to modernize their visual recognition infrastructure should consider adopting intermediate masking and frequency-based loss functions when training unlabeled visual datasets. If deployment pipelines strictly require standard convolutional backbones, the proposed framework provides an effective, drop-in self-supervised pre-training strategy. For applications prioritizing maximum ultimate accuracy and scale, Vision Transformers remain the preferred choice.
The article notes specific limitations: Convolutional Neural Networks experience diminishing returns and plateau after roughly 300 pre-training epochs, gaining less overall performance (around 1%) compared to Vision Transformers (which gain over 2% and continue improving with extended training). While confidence in the experimental results is high across standard vision benchmarks, stakeholders should be aware that architectural inductive biases in convolution inherently limit performance gains relative to transformer-based alternatives under extended training regimes.
- Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Read MAE first to understand the masked-patch reconstruction paradigm and ViT-based baseline that this work adapts beyond transformers.
- Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). SimMIM establishes simple pixel-regression masking as a strong alternative to tokenizers, providing a direct foundation for the source’s reconstruction and masking choices.
- Paper: Revealing the Dark Secrets of Masked Image Modeling, Zhenda Xie et al. (2023). Its analysis of masked-image-modeling representations and internal mechanisms prepares readers for the source’s investigation of what MIM teaches and how it transfers across architectures.
No sufficiently relevant recommendations were found.
