Architecture-Agnostic Masked Image Modeling - From ViT back to CNN

Siyuan LiDi WuFang WuZelin ZangStan Z. Li

article2023ICML63 citations

Presents a unified masked image modeling framework that operates across both vision transformers and convolutional neural networks by masking intermediate feature representations rather than input pixels, enabling CNNs to achieve superior self-supervised pre-training performance without specialized transformer components.

Listen

Training high-performance computer vision systems usually requires massive amounts of manually annotated data, which is costly and labor-intensive to produce. To overcome this limitation, self-supervised learning methods have gained popularity, especially masked image modeling. In this approach, portions of an image are hidden, and the artificial intelligence model learns by reconstructing the missing visual information. While masked image modeling has demonstrated remarkable success when paired with modern Vision Transformers, prior research widely believed it was fundamentally incompatible with traditional Convolutional Neural Networks.

The article aims to uncover the underlying operational principles of masked image modeling and demonstrate a unified framework that functions effectively across both Vision Transformers and Convolutional Neural Networks without requiring architecture-specific modifications.

To achieve this, the authors evaluated model robustness under varying levels of visual occlusion and analyzed multi-order interactions, which measure how effectively models combine information from multiple image patches. Using these empirical insights, they developed the Architecture-Agnostic Masked Image Modeling framework. Instead of masking tokens directly at the initial input layer, the proposed method fills occluded image regions with the average color values and introduces learned mask tokens into the intermediate layers of the network. Additionally, the approach incorporates an adaptive frequency loss using the Fourier transform to train the model on medium-frequency structural patterns such as contours and edges. The evaluation encompassed standard benchmarks, including ImageNet classification, COCO object detection and segmentation, and ADE20K semantic parsing, across multiple model scales.

The investigation produced several key findings. First, the core benefit of masked image modeling is teaching models to capture middle-order patch interactions—such as shapes and boundaries—rather than merely achieving visual reconstruction. Second, the proposed framework successfully bridged the architecture gap, allowing standard Convolutional Neural Networks like ResNet-50 to achieve an 80.4% top-1 accuracy on ImageNet, outperforming established contrastive learning alternatives without complex tokenizers. Third, Vision Transformers pre-trained with the framework achieved strong classification performance, reaching up to 86.3% top-1 accuracy on ViT-Large. Finally, in downstream transfer evaluations on COCO and ADE20K, the method matched or exceeded prevailing contrastive and masked learning baselines across both convolutional and transformer architectures.

These findings indicate that organizations can leverage modern self-supervised masked pre-training techniques without abandoning established convolutional network deployments. This flexibility enables teams to reuse existing hardware-optimized network pipelines while achieving competitive accuracy, thereby lowering transition costs and reducing dependence on expensive labeled datasets.

Technical leaders looking to modernize their visual recognition infrastructure should consider adopting intermediate masking and frequency-based loss functions when training unlabeled visual datasets. If deployment pipelines strictly require standard convolutional backbones, the proposed framework provides an effective, drop-in self-supervised pre-training strategy. For applications prioritizing maximum ultimate accuracy and scale, Vision Transformers remain the preferred choice.

The article notes specific limitations: Convolutional Neural Networks experience diminishing returns and plateau after roughly 300 pre-training epochs, gaining less overall performance (around 1%) compared to Vision Transformers (which gain over 2% and continue improving with extended training). While confidence in the experimental results is high across standard vision benchmarks, stakeholders should be aware that architectural inductive biases in convolution inherently limit performance gains relative to transformer-based alternatives under extended training regimes.

  • Paper: Masked Autoencoders Are Scalable Vision Learners, Kaiming He et al. (2022). Read MAE first to understand the masked-patch reconstruction paradigm and ViT-based baseline that this work adapts beyond transformers.
  • Paper: SimMIM: a Simple Framework for Masked Image Modeling, Zhenda Xie et al. (2021). SimMIM establishes simple pixel-regression masking as a strong alternative to tokenizers, providing a direct foundation for the source’s reconstruction and masking choices.
  • Paper: Revealing the Dark Secrets of Masked Image Modeling, Zhenda Xie et al. (2023). Its analysis of masked-image-modeling representations and internal mechanisms prepares readers for the source’s investigation of what MIM teaches and how it transfers across architectures.

No sufficiently relevant recommendations were found.

Cover for Architecture-Agnostic Masked Image Modeling - From ViT back to CNN

Abstract

Masked image modeling (MIM), an emerging self-supervised pre-training paradigm, has achieved impressive success in computer vision tasks. However, most existing approaches are designed for Vision Transformers (ViTs), while the exploration of applying MIM to Convolutional Neural Networks (CNNs) is still in its infancy. In this paper, we propose a simple yet effective approach, termed Architecture-Agnostic Masked Image Modeling (A2MIM), which unifies the MIM pre-training for both ViTs and CNNs. The key insight of A2MIM is to perform MIM in the feature space rather than the input space, which naturally avoids the architectural discrepancy between ViTs and CNNs. Specifically, A2MIM first extracts the feature maps from the intermediate layer of the backbone, and then applies the mask modeling on the feature maps, which is architecture-agnostic. We conduct extensive experiments on various vision tasks, including image classification, object detection, and semantic segmentation. The results demonstrate that A2MIM consistently improves the performance of both ViTs and CNNs, and achieves state-of-the-art results. Our code will be available at https://github.com/alibaba/A2MIM.

Table of Contents

  • 1. Introduction
  • 2. Related Work
  • 3. Midst of Masked Image Modeling
  • 3.1. Is MIM Better Image Augmentation?
  • 3.2. Middle-order Interactions for Generalized Feature Extraction
  • 4. Approach
  • 4.1. Architecture Agnostic Framework
  • 4.2. Middle-order Interactions from Fourier Perspective
  • 5. Experiments
  • 5.1. Pre-training Setup
  • 5.2. Image Classification on ImageNet-1K
  • 5.3. Transfer Learning Experiments
  • 5.4. Ablation Study
  • 5.5. Verification of A 2 MIMDesign Rules
  • 6. Conclusion and Limitation
  • Acknowledgement
  • References
  • A. Details of Comparison Experiments
  • A.1. ImageNet-1K Experiments
  • A.2. Object Detection and Segmentation on COCO
  • A.3. Semantic Segmentation on ADE-20K
  • B. Empirical Experiments
  • B.1. Occlusion Robustness
  • B.2. Multi-order Interaction
  • B.3. MIM from Frequency Perspective
  • C. More Experiment Results
  • C.1. Ablation of Layers for Mask Token
  • C.2. Ablation of the Proposed Modules
  • D. Visualization Experimental Details
  • E. Extended Related Work

Knowls

  1. Knowl 1 — A2MIM is a unified masked-pretraining framework for CNNs and Transformers

    model/method

    Architecture-Agnostic Masked Image Modeling (A2MIM) applies masked image reconstruction to either convolutional or Transformer encoders without requiring architecture-specific sparse convolutions or other designs native to Vision Transformers. Its central design is to use mean-RGB-filled masked inputs, add a learnable mask token at an intermediate encoder representation, and train with a spatial reconstruction loss augmented by a Fourier-domain loss. The framework is intended to improve middle-order interactions among image patches, which the authors associate with more general visual features.

  2. Knowl 2 — MIM pretraining is associated with stronger middle-order patch interactions

    empirical result

    The paper’s interaction analysis indicates that masked image modeling (MIM) shifts learned patch interactions toward middle orders, rather than merely improving pixel reconstruction. On ImageNet-100 with 224 × 224 images represented as 16 × 16 patches, MIM-pretrained ViT-S models showed stronger interactions at orders from about 0.05n0.05n to 0.5n0.5n, and MIM-pretrained ResNet-50 models from about 0.1n0.1n to 0.55n0.55n, compared with randomly initialized models; here nn is the number of patches. The authors interpret these interactions as supporting features such as shapes and edges, in contrast to the local texture features associated with low-order interactions. In occlusion tests on the same dataset, MIM-pretrained models were more robust to random patch removal than models trained from random initialization with patch-removing augmentations. These experiments support the authors’ account of what MIM teaches, but do not establish it as the sole mechanism behind MIM’s performance.

  3. Knowl 3 — Multi-order patch interaction and its normalized strength

    definition

    For an image with nn patch locations N={1,…,n}N=\{1,\ldots,n\}, let f(A)f(A) be a scalar model-output score when patches in A⊆NA\subseteq N are retained and all other locations are replaced by a baseline. For distinct patches i,ji,j and a context S⊆N∖{i,j}S\subseteq N\setminus\{i,j\} containing mm patches, their interaction in that context is

    Δf(i,j,S)=f(S∪{i,j})−f(S∪{i})−f(S∪{j})+f(S).\Delta f(i,j,S)=f(S\cup\{i,j\})-f(S\cup\{i\})-f(S\cup\{j\})+f(S).

    The order-mm interaction I(m)(i,j)I^{(m)}(i,j) is the average of this quantity over contexts of size mm:

    I(m)(i,j)=ES⊆N∖{i,j}, ∣S∣=m[Δf(i,j,S)].I^{(m)}(i,j)=\mathbb{E}_{S\subseteq N\setminus\{i,j\},\,|S|=m}[\Delta f(i,j,S)].

    The relative strength of order mm is

    J(m)=Ex∈ΩEi,j∣I(m)(i,j∣x)∣Em′Ex∈ΩEi,j∣I(m′)(i,j∣x)∣,J^{(m)}=\frac{\mathbb{E}_{x\in\Omega}\mathbb{E}_{i,j}|I^{(m)}(i,j\mid x)|}{\mathbb{E}_{m'}\mathbb{E}_{x\in\Omega}\mathbb{E}_{i,j}|I^{(m')}(i,j\mid x)|},

    where xx ranges over samples in dataset Ω\Omega, the patch-pair average is over distinct pairs, and m,m′m,m' range from 00 to n−2n-2. Thus J(m)J^{(m)} describes the normalized distribution of interaction strength across context sizes.

  4. Knowl 4 — A2MIM adds mask tokens at intermediate feature representations

    model/method

    Rather than inserting mask tokens only at the encoder input or decoder, A2MIM adds a learnable token to an intermediate feature representation, where features contain both spatial and semantic information. If zlz^l is the feature map at CNN stage or Transformer layer ll, MM is the input occlusion mask, and D(M)D(M) is that mask downsampled to the spatial resolution of zlz^l, the operation is

    zmaskl=zl+T⊙D(M),z^l_{\mathrm{mask}}=z^l+T\odot D(M),

    where TT is the learned mask token, broadcast to the feature-map shape, and ⊙\odot is elementwise multiplication. The token is added at masked locations, preserving the encoder’s intermediate features rather than replacing input patches with a token. The authors’ layer ablation found best performance when the token was introduced at a medium-depth representation, around stage 3 of ResNet-50 or layer 5 of ViT-S.

  5. Knowl 5 — Mean-RGB input masking preserves low-level image statistics

    model/method

    A2MIM fills masked input patches with the input image’s mean RGB value, rather than with zeros or a learnable mask token. For image xx and a binary spatial mask MM that is 11 at occluded locations, the masked input is xmask=x⊙(1−M)+μRGB⊙Mx_{\mathrm{mask}}=x\odot(1-M)+\mu_{\mathrm{RGB}}\odot M, where μRGB\mu_{\mathrm{RGB}} is the image’s per-channel mean, broadcast spatially. The authors identify the RGB mean with the image’s DC component and propose that this replacement reduces distortion of local statistics while encouraging the encoder to model informative middle frequencies instead of reconstructing masked regions as blurry low-frequency color blocks.

  6. Knowl 6 — A dynamically weighted Fourier loss complements spatial reconstruction

    equation

    For RGB image xx of height HH and width WW, A2MIM uses a discrete Fourier transform (DFT) on each channel in addition to its masked spatial reconstruction loss Lspa\mathcal{L}_{\mathrm{spa}}. For channel c∈{1,2,3}c\in\{1,2,3\} and frequency indices u,vu,v, the DFT is

    Fc(u,v)=∑h=1H∑w=1Wxc(h,w)exp⁡ ⁣[−2πj(uhH+vwW)],F_c(u,v)=\sum_{h=1}^{H}\sum_{w=1}^{W}x_c(h,w)\exp\!\left[-2\pi j\left(\frac{uh}{H}+\frac{vw}{W}\right)\right],

    where jj is the imaginary unit. Let x^\hat{x} be the predicted image, let MM be the binary mask at image resolution, and let sg⁡(x^)\operatorname{sg}(\hat{x}) denote a stop-gradient copy of the prediction. Define the frequency residual as the DFT of x^c⊙M+sg⁡(x^c)⊙(1−M)\hat{x}_c\odot M+\operatorname{sg}(\hat{x}_c)\odot(1-M) minus the DFT of xcx_c. The adaptive frequency weight ω(u,v)\omega(u,v) is the magnitude of this prediction–target discrepancy at that frequency. The Fourier loss sums the weighted discrepancy over RGB channels and frequencies, and the total loss is L=Lspa+λLfreq\mathcal{L}=\mathcal{L}_{\mathrm{spa}}+\lambda\mathcal{L}_{\mathrm{freq}}, with λ=0.1\lambda=0.1 by default. The authors report that the adaptive weighting emphasizes informative frequency components and that computing the transform with FFT algorithms has O(Nlog⁡N)O(N\log N) complexity for NN image elements.

  7. Knowl 7 — A2MIM improves ImageNet-1K fine-tuning accuracy for ViT models

    empirical result

    On ImageNet-1K, the authors pretrained ViT-S, ViT-B, and ViT-L models with A2MIM for 800 epochs, then evaluated them by 200-epoch end-to-end fine-tuning. A2MIM with an RGB target achieved top-1 accuracies of 82.1%, 84.2%, and 86.1%, respectively, compared with 81.7%, 83.8%, and 85.6% for SimMIM under its reported 800-epoch pretraining. The A2MIM+ variant, which uses HOG targets and an MLP decoder with depth-wise convolutions, achieved 82.3%, 84.4%, and 86.3%. These results show improvements over the RGB-target SimMIM baseline and competitive accuracy with methods using more elaborate supervision or architectures.

  8. Knowl 8 — A2MIM improves CNN classification results across model scales

    empirical result

    In ImageNet-1K comparisons using the reported 300-epoch fine-tuning protocol, A2MIM outperformed the listed MoCoV3 and SimMIM results across ResNet-50/101/152/200 and ConvNeXt-T/S/B, although SparK was higher in every listed architecture. Top-1 accuracies (%) for A2MIM, MoCoV3, SimMIM, and SparK were: ResNet-50, 80.4, 79.8, 79.9, 80.6; ResNet-101, 81.9, 81.3, 81.3, 82.2; ResNet-152, 82.5, 81.8, 81.9, 82.7; ResNet-200, 83.0, 82.1, 82.2, 83.1; ConvNeXt-T, 82.5, 82.1, 82.1, 82.7; ConvNeXt-S, 83.7, 83.1, 83.2, 84.1; and ConvNeXt-B, 84.1, 83.5, 83.6, 84.8. The paper’s ResNet-50 experiments also report 80.4% for A2MIM versus 79.9% for SimMIM after the longer-pretraining comparison.

  9. Knowl 9 — A2MIM transfers well to detection and segmentation

    empirical result

    On transfer benchmarks, A2MIM pretrained ViT-B for 800 epochs and ResNet-50 for 300 epochs on ImageNet-1K, then evaluated downstream detection and segmentation. For ViT-B, A2MIM obtained COCO box AP 49.4, mask AP 43.5, and ADE20K mIoU 49.0; corresponding MAE results were 48.5, 42.8, and 48.1, and SimMIM results were 48.9, 43.0, and 48.4. For ResNet-50, A2MIM reached COCO box AP 39.8, mask AP 34.9, and ADE20K mIoU 38.3; MoCoV2 obtained 39.2, 34.3, and 37.5, while CIM reported 38.0 mIoU and no COCO values in the comparison. The CNN evaluation used Mask R-CNN with a ResNet-50-C4 backbone on COCO and 160K-iteration fine-tuning for ADE20K; the ViT evaluation used Mask R-CNN with ViT-B on COCO and 160K-iteration fine-tuning on ADE20K.

  10. Knowl 10 — Component ablations support intermediate masking and adaptive frequency weighting

    empirical result

    Fine-tuning ablations show stepwise improvements from A2MIM’s components. In the reported column order—ResNet-50 on ImageNet-100, ResNet-50 on ImageNet-1K, ViT-S on ImageNet-100, and ViT-B on ImageNet-1K—the SimMIM baseline scored 87.75, 78.2, 85.10, and 83.1. Using the spatial reconstruction setup with mean-RGB masking raised these to 88.19, 78.4, 85.27, and 83.2. Adding the Fourier loss without adaptive frequency weights gave 88.47, 78.4, 86.05, and 83.3; with adaptive weights the scores were 88.73, 78.6, 86.41, and 83.4; adding the intermediate mask token as well raised them to 88.86, 78.8, 86.62, and 83.5. Additional ImageNet-1K ablations found that HOG and DINO-feature targets, and depth-wise-convolution additions to MLP or Transformer decoders, could improve on the linear-decoder RGB-target baseline, supporting the framework’s ability to accommodate alternative targets and decoder designs.

  11. Knowl 11 — CNNs gain less from A2MIM and longer pretraining than ViTs

    limitation

    The authors report that CNNs benefit less from A2MIM pretraining than Vision Transformers: ResNet and ConvNeXt gain around 1 percentage point in ImageNet accuracy, whereas ViTs gain more than 2 points. They hypothesize that CNN inductive biases limit the middle-order interactions induced by MIM. The models also differ in how they benefit from longer pretraining: ViTs continue to improve, while CNNs show no significant gain after 300 pretraining epochs; ResNet-50 and ResNet-152 show limited or negative accuracy gains at 800 or more epochs.

Coverage note — No substantial contribution is omitted; appendix visualization examples and detailed per-configuration training recipes are left out because they illustrate or operationalize the main findings rather than add a separate major result.

References

  1. 1.Akbari, H., Yuan, L., Qian, R., Chuang, W.-H., Chang, S.-F., Cui, Y., and Gong, B. Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  2. 2.Ancona, M., Oztireli, C., and Gross, M. Explaining deep neural networks with a polynomial time algorithm for shapley value approximation. In International Conference on Machine Learning (ICML), pp. 272–281. PMLR, 2019.
  3. 3.Assran, M., Caron, M., Misra, I., Bojanowski, P., Bordes, F., Vincent, P., Joulin, A., Rabbat, M., and Ballas, N. Masked siamese networks for label-efficient learning. arXiv preprint arXiv:2204.07141, 2022.
  4. 4.Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M. Data2vec: A general framework for self-supervised learning in speech, vision and language. arXiv preprint arXiv:2202.03555, 2022.
  5. 5.Bao, H., Dong, L., and Wei, F. Beit: Bert pre-training of image transformers. In International Conference on Learning Representations (ICLR), 2022.
  6. 6.Bardes, A., Ponce, J., and LeCun, Y. Vicreg: Variance-invariance-covariance regularization for self-supervised learning. In International Conference on Learning Representations (ICLR), 2022.
  7. 7.Caron, M., Misra, I., Mairal, J., Goyal, P., Bojanowski, P., and Joulin, A. Unsupervised learning of visual features by contrasting cluster assignments. Advances in Neural Information Processing Systems (NeurIPS), 33:9912–9924, 2020.
  8. 8.Caron, M., Touvron, H., Misra, I., Jegou, H., Mairal, J., Bojanowski, P., and Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the International Conference on Computer Vision (ICCV), 2021.
  9. 9.Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. Generative pretraining from pixels. In International Conference on Machine Learning (ICML), pp. 1691–1703. PMLR, 2020a.
  10. 10.Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. A simple framework for contrastive learning of visual representations. arXiv preprint arXiv:2002.05709, 2020b.
  11. 11.Chen, X. and He, K. Exploring simple siamese representation learning. arXiv preprint arXiv:2011.10566, 2020.
  12. 12.Chen, X., Fan, H., Girshick, R., and He, K. Improved baselines with momentum contrastive learning. arXiv preprint arXiv:2003.04297, 2020c.
  13. 13.Chen, X., Xie, S., and He, K. An empirical study of training self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 9640–9649, 2021.
  14. 14.Chen, X., Ding, M., Wang, X., Xin, Y., Mo, S., Wang, Y., Han, S., Luo, P., Zeng, G., and Wang, J. Context autoencoder for self-supervised representation learning. arXiv preprint arXiv:2202.03026, 2022.
  15. 15.Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), pp. 702–703, 2020.
  16. 16.Deng, H., Ren, Q., Chen, X., Zhang, H., Ren, J., and Zhang, Q. Discovering and explaining the representation bottleneck of dnns. In International Conference on Learning Representations (ICLR), 2022.
  17. 17.Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018.
  18. 18.Doersch, C., Gupta, A., and Efros, A. A. Unsupervised visual representation learning by context prediction. In Proceedings of the International Conference on Computer Vision (ICCV), 2015.
  19. 19.Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR), 2021.
  20. 20.El-Nouby, A., Izacard, G., Touvron, H., Laptev, I., Jegou, H., and Grave, E. Are large-scale datasets necessary for self-supervised pre-training? arXiv preprint arXiv:2112.10740, 2021.
  21. 21.Ermolov, A., Siarohin, A., Sangineto, E., and Sebe, N. Whitening for self-supervised representation learning. In International Conference on Machine Learning (ICML), 2021.
  22. 22.Esser, P., Rombach, R., and Ommer, B. Taming transformers for high-resolution image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12873–12883, June 2021.
  23. 23.Fang, Y., Dong, L., Bao, H., Wang, X., and Wei, F. Corrupted image modeling for self-supervised visual pre-training. arXiv preprint arXiv:2202.03382, 2022.
  24. 24.Ge, C., Liang, Y., Song, Y., Jiao, J., Wang, J., and Luo, P. Revitalizing cnn attentions via transformers in self-supervised visual representation learning. In Advances in Neural Information Processing Systems (NeurIPS), 2021.
  25. 25.Gidaris, S., Singh, P., and Komodakis, N. Unsupervised representation learning by predicting image rotations. In International Conference on Learning Representations (ICLR), 2018.
  26. 26.Goyal, P., Dollar, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677, 2020.
  27. 27.Grill, J.-B., Strub, F., Altche, F., Tallec, C., Richemond, P. H., Buchatskaya, E., Doersch, C., Pires, B. A., Guo, Z. D., Azar, M. G., et al. Bootstrap your own latent: A new approach to self-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), 2020.
  28. 28.He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), pp. 770–778, 2016.
  29. 29.He, K., Gkioxari, G., Dollar, P., and Girshick, R. Mask r-cnn. In Proceedings of the International Conference on Computer Vision (ICCV), 2017.
  30. 30.He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9729–9738, 2020.
  31. 31.He, K., Chen, X., Xie, S., Li, Y., Dollar, P., and Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), 2022.
  32. 32.Hinton, G. E. and Zemel, R. Autoencoders, minimum description length and helmholtz free energy. Advances in neural information processing systems (NeurIPS), 6, 1993.
  33. 33.Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q. Deep networks with stochastic depth. In Proceedings of the European Conference on Computer Vision (ECCV), 2016.
  34. 34.Jiang, L., Dai, B., Wu, W., and Loy, C. C. Focal frequency loss for image reconstruction and synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 13919–13929, 2021.
  35. 35.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems (NeurIPS), volume 25. Curran Associates, Inc., 2012a.
  36. 36.Krizhevsky, A., Sutskever, I., and Hinton, G. E. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems (NeurIPS), pp. 1097–1105, 2012b.
  37. 37.Li, S., Liu, Z., Wang, Z., Wu, D., Liu, Z., and Li, S. Z. Boosting discriminative visual representation learning with scenario-agnostic mixup. ArXiv, abs/2111.15454, 2021.
  38. 38.Li, S., Wang, Z., Liu, Z., Wu, D., and Li, S. Z. Openmixup: Open mixup toolbox and benchmark for visual representation learning. https://github.com/Westlake-AI/openmixup, 2022.
  39. 39.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision (ECCV), 2014.
  40. 40.Lin, T.-Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), July 2017.
  41. 41.Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
  42. 42.Liu, Z., Li, S., Wu, D., Chen, Z., Wu, L., Guo, J., and Li, S. Z. Automix: Unveiling the power of mixup for stronger classifiers. In Proceedings of the European Conference on Computer Vision (ECCV), 2022a.
  43. 43.Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S. A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022b.
  44. 44.Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), 2019.
  45. 45.Naseer, M. M., Ranasinghe, K., Khan, S. H., Hayat, M., Shahbaz Khan, F., and Yang, M.-H. Intriguing properties of vision transformers. In Advances in Neural Information Processing Systems (NeurIPS), volume 34, 2021.
  46. 46.Park, N. and Kim, S. Blurs behave like ensembles: Spatial smoothings to improve accuracy, uncertainty, and robustness. arXiv preprint arXiv:2105.12639, 2021.
  47. 47.Park, N. and Kim, S. How do vision transformers work? In International Conference on Learning Representations (ICLR), 2022.
  48. 48.Pathak, D., Krahenbuhl, P., Donahue, J., Darrell, T., and Efros, A. A. Context encoders: Feature learning by inpainting. In Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR), pp. 2536–2544, 2016.
  49. 49.Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. Improving language understanding by generative pre-training, 2018.
  50. 50.Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. Zero-shot text-to-image generation. ArXiv, abs/2102.12092, 2021.
  51. 51.Sameni, S., Jenni, S., and Favaro, P. Dilemma: Self-supervised shape and texture learning with transformers. arXiv preprint arXiv:2204.04788, 2022.
  52. 52.Selvaraju, R. R., Desai, K., Johnson, J., and Naik, N. Casting your model: Learning to localize improves self-supervised representations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11058–11067, 2021.
  53. 53.Shelhamer, E., Long, J., and Darrell, T. Fully convolutional networks for semantic segmentation. IEEE transactions on pattern analysis and machine intelligence, 39(4):640–651, 2017.
  54. 54.Song, Z., Xiao, G., Hu, G., and Zhao, C. Deep perturbation learning: Improve the network performance via image perturbations. In Proceedings of the 40th international conference on Machine learning (ICML), 2023.
  55. 55.Tian, K., Jiang, Y., Diao, Q., Lin, C., Wang, L., and Yuan, Z. Designing bert for convolutional networks: Sparse and hierarchical masked modeling. In International Conference on Learning Representations (ICLR), 2023.
  56. 56.Tian, Y., Krishnan, D., and Isola, P. Contrastive multiview coding. In Proceedings of the European Conference on Computer Vision (ECCV), 2020.
  57. 57.Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jegou, H. Training data-efficient image transformers & distillation through attention. In International Conference on Machine Learning (ICML), pp. 10347–10357. PMLR, 2021.
  58. 58.van den Oord, A., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. ArXiv, abs/1807.03748, 2018.
  59. 59.Van den Oord, A., Li, Y., and Vinyals, O. Representation learning with contrastive predictive coding. arXiv e-prints, pp. arXiv–1807, 2018.
  60. 60.Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. Attention is all you need. Advances in neural information processing systems (NeurIPS), 30, 2017.
  61. 61.Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A. Extracting and composing robust features with denoising autoencoders. In Proceedings of the 25th international conference on Machine learning (ICML), pp. 1096–1103, 2008.
  62. 62.Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., Manzagol, P.-A., and Bottou, L. Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of machine learning research, 11(12), 2010.
  63. 63.Wang, W., Xie, E., Li, X., Fan, D.-P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. Pvtv2: Improved baselines with pyramid vision transformer. arXiv preprint arXiv:2106.13797, 2022.
  64. 64.Wei, C., Fan, H., Xie, S., Wu, C.-Y., Yuille, A., and Feichtenhofer, C. Masked feature prediction for self-supervised visual pre-training. arXiv preprint arXiv:2112.09133, 2021.
  65. 65.Wightman, R., Touvron, H., and Jegou, H. Resnet strikes back: An improved training procedure in timm, 2021.
  66. 66.Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I.-S., and Xie, S. Convnext v2: Co-designing and scaling convnets with masked autoencoders. ArXiv, abs/2301.00808, 2023.
  67. 67.Wu, D., Li, S., Zang, Z., and Li, S. Z. Exploring localization for self-supervised fine-grained contrastive learning. In Proceedings of the British Machine Vision Conference (BMVC), 2022.
  68. 68.Wu, F., Li, S., Jin, X., Jiang, Y., Radev, D., Niu, Z., and Li, S. Z. Explaining graph neural networks via non-parametric subgraph matching. In Proceedings of the 40th international conference on Machine learning (ICML), 2023.
  69. 69.Wu, Z., Xiong, Y., Stella, X. Y., and Lin, D. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
  70. 70.Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J. Unified perceptual parsing for scene understanding. In European Conference on Computer Vision (ECCV). Springer, 2018.
  71. 71.Xiao, T., Reed, C. J., Wang, X., Keutzer, K., and Darrell, T. Region similarity representation learning. arXiv preprint arXiv:2103.12902, 2021.
  72. 72.Xie, E., Ding, J., Wang, W., Zhan, X., Xu, H., Li, Z., and Luo, P. Detco: Unsupervised contrastive learning for object detection. In Proceedings of the International Conference on Computer Vision (ICCV), 2021a.
  73. 73.Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. Simmim: A simple framework for masked image modeling. arXiv preprint arXiv:2111.09886, 2021b.
  74. 74.Xu, K., Hu, W., Leskovec, J., and Jegelka, S. How powerful are graph neural networks? In International Conference on Learning Representations (ICLR), 2019.
  75. 75.You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J. Large batch optimization for deep learning: Training BERT in 76 minutes. In International Conference on Learning Representations (ICLR), 2020.
  76. 76.Yun, S., Han, D., Oh, S. J., Chun, S., Choe, J., and Yoo, Y. Cutmix: Regularization strategy to train strong classifiers with localizable features. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 6023–6032, 2019.
  77. 77.Zbontar, J., Jing, L., Misra, I., LeCun, Y., and Deny, S. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning (ICML), pp. 12310–12320. PMLR, 2021.
  78. 78.Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D. mixup: Beyond empirical risk minimization. In International Conference on Learning Representations (ICLR), 2018.
  79. 79.Zhang, H., Li, S., Ma, Y., Li, M., Xie, Y., and Zhang, Q. Interpreting and boosting dropout from a game-theoretic view. arXiv preprint arXiv:2009.11729, 2020.
  80. 80.Zhang, R., Isola, P., and Efros, A. A. Colorful image colorization. In Proceedings of the European Conference on Computer Vision (ECCV), 2016.
  81. 81.Zhong, Z., Zheng, L., Kang, G., Li, S., and Yang, Y. Random erasing data augmentation. In Proceedings of the AAAI conference on artificial intelligence (AAAI), pp. 13001–13008, 2020.
  82. 82.Zhou, B., Zhao, H., Puig, X., Xiao, T., Fidler, S., Barriuso, A., and Torralba, A. Semantic understanding of scenes through the ade20k dataset. International Journal of Computer Vision (IJCV), 2019.
  83. 83.Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. ibot: Image bert pre-training with online tokenizer. arXiv preprint arXiv:2111.07832, 2021.

Citation

MLA
Li, S., et al. “Architecture-Agnostic Masked Image Modeling -- From ViT Back to CNN”. International Conference on Machine Learning, vol. 202, 2023, pp. 20149–67, https://proceedings.mlr.press/v202/li23af.html.
APA
Li, S., Wu, D., Wu, F., Zang, Z., & Li, S. Z. (2023). Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN. International Conference on Machine Learning, 202, 20149–20167. https://proceedings.mlr.press/v202/li23af.html
Chicago
Li, S., D. Wu, F. Wu, Z. Zang, and S. Z. Li. 2023. “Architecture-Agnostic Masked Image Modeling -- From ViT Back to CNN”. International Conference on Machine Learning 202: 20149–67. https://proceedings.mlr.press/v202/li23af.html.
Harvard
Li, S. et al. (2023) “Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN”, International Conference on Machine Learning. PMLR, pp. 20149–20167. Available at: https://proceedings.mlr.press/v202/li23af.html.
Vancouver
1. Li S, Wu D, Wu F, Zang Z, Li SZ (2023) Architecture-Agnostic Masked Image Modeling -- From ViT back to CNN. In: International Conference on Machine Learning. PMLR, pp 20149–20167

BibTeX

@InProceedings{pmlr-v202-li23af,
  title = 	 {Architecture-Agnostic Masked Image Modeling -- From {V}i{T} back to {CNN}},
  author =       {Li, Siyuan and Wu, Di and Wu, Fang and Zang, Zelin and Li, Stan Z.},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {20149--20167},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/li23af/li23af.pdf},
  url = 	 {https://proceedings.mlr.press/v202/li23af.html},
  abstract = 	 {Masked image modeling, an emerging self-supervised pre-training method, has shown impressive success across numerous downstream vision tasks with Vision transformers. Its underlying idea is simple: a portion of the input image is masked out and then reconstructed via a pre-text task. However, the working principle behind MIM is not well explained, and previous studies insist that MIM primarily works for the Transformer family but is incompatible with CNNs. In this work, we observe that MIM essentially teaches the model to learn better middle-order interactions among patches for more generalized feature extraction. We then propose an Architecture-Agnostic Masked Image Modeling framework (A$^2$MIM), which is compatible with both Transformers and CNNs in a unified way. Extensive experiments on popular benchmarks show that A$^2$MIM learns better representations without explicit design and endows the backbone model with the stronger capability to transfer to various downstream tasks.}
}
Metadata:DOI registry

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/